Presents an end-to-end event-camera pipeline for real-time pedestrian crossing anticipation that handles motion blur, high dynamic range, and class imbalance more efficiently than frame-based networks.

Topological visualization of Pedestrian Crossing Intent Classification From Event-Based Vision Using Convolutional Spiking Neural Networks With Temporal Augmentation
Brave API

Henok Teklu et al. (September 2026) present a Convolutional Spiking Neural Network (Conv-SNN) pipeline that converts JAAD dataset footage into synthetic Dynamic Vision Sensor (DVS) event streams using the v2e simulator. To address class imbalance and environmental challenges, the model is trained on DVS-PedX synthetic data (normal and adverse weather) with clip-consistent augmentation and JAAD oversampling.

The architecture achieves 95.83% accuracy and an F1-score of 0.9695 on the JAAD DVS test set, outperforming prior frame-based approaches while operating on sparse temporal representations. It also demonstrates robustness in simulated adverse weather, reaching 94.78% accuracy on CARLA DVS sequences with only 1.07M parameters.

Generated 19d ago
Open-Weights Reasoning

This paper addresses pedestrian crossing intent classification using event-based vision, targeting the challenge of anticipating whether a pedestrian is about to cross a road in real time. Rather than relying on conventional frame-based cameras, which can suffer from motion blur, limited dynamic range, and higher latency, the work leverages event cameras that output asynchronous, pixel-level change events. The paper proposes an end-to-end pipeline built around convolutional spiking neural networks (CSNNs), which process the temporally rich event stream directly and are better suited to the sparse, high-frequency nature of event data.

A central contribution is the use of temporal augmentation within a CSNN framework to improve the model’s ability to capture motion dynamics and overcome class imbalance, a common issue in crossing-intent datasets where non-crossing or ambiguous cases dominate. By combining event-native representations with spiking neural computation, the approach aims to preserve fine-grained temporal information while remaining computationally efficient. The key insight is that pedestrian anticipation can be improved by exploiting the temporal structure of event data rather than reconstructing conventional frames, allowing the system to respond more quickly to rapid motion and high-contrast visual conditions.

The work matters because pedestrian crossing anticipation is a safety-critical perception task for autonomous vehicles and robotics, especially in challenging lighting or fast-moving scenes. Event cameras and spiking neural networks are also attractive for low-latency, low-power edge deployment, potentially enabling more responsive and energy-efficient perception systems. Overall, the paper contributes a promising direction for real-time, event-driven pedestrian understanding that may outperform frame-based pipelines in both temporal responsiveness and robustness.

Generated 19d ago
Sources