Tengyu Long

dblp:333/7083 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Image recognition and object detection · 70% Video understanding and tracking · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
event-based object detection
0.912025
Rethinking Scale-Aware Temporal Encoding for Event-based Object Detection · NeurIPS 2025
Computer vision › Image recognition and object detection
object detection
0.912025
Rethinking Scale-Aware Temporal Encoding for Event-based Object Detection · NeurIPS 2025
Computer vision › Video understanding and tracking
temporal modeling
0.912025
Rethinking Scale-Aware Temporal Encoding for Event-based Object Detection · NeurIPS 2025
Computer vision › Image recognition and object detection › object detection
feature pyramid network
0.312025
Rethinking Scale-Aware Temporal Encoding for Event-based Object Detection · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

feature pyramid network · 0.9decoupled deformable-enhanced recurrent layer · 0.9CNN-RNN hybrid · 0.9
YearPublicationVenuePosition
2025 Rethinking Scale-Aware Temporal Encoding for Event-based Object Detection
abstract
Event cameras provide asynchronous, low-latency, and high-dynamic-range visual signals, making them ideal for real-time perception tasks such as object detection. However, effectively modeling the temporal dynamics of event streams remains a core challenge. Most existing methods follow frame-based detection paradigms, applying temporal modules only at high-level features, which limits early-stage temporal modeling. Transformer-based approaches introduce global attention to capture long-range dependencies, but often add unnecessary complexity and overlook fine-grained temporal cues. In this paper, we propose a CNN-RNN hybrid framework that rethinks temporal modeling for event-based object detection. Our approach is based on two key insights: (1) introducing recurrent modules at lower spatial scales to preserve detailed temporal information where events are most dense, and (2) utilizing Decoupled Deformable-enhanced Recurrent Layers specifically designed according to the inherent motion characteristics of event cameras to extract multiple spatiotemporal features, and performing independent downsampling at multiple spatiotemporal scales to enable flexible, scale-aware representation learning. These multi-scale features are then fused via a feature pyramid network to produce robust detection outputs. Experiments on Gen1, 1 Mpx and eTram dataset demonstrate that our approach achieves superior accuracy over recent transformer-based models, highlighting the importance of precise temporal feature extraction in early stages. This work offers a new perspective on designing architectures for event-driven vision beyond attention-centric paradigms. Code: https://github.com/BIT-Vision/SATE.
Lin Zhu 0012, Tengyu Long, Xiao Wang 0014, Lizhi Wang 0001, Hua Huang 0001
NeurIPS2
2022 A fusion approach based on evidential reasoning rule considering the reliability of digital quantities
Jie Wang 0071, Zhi-Jie Zhou 0001, Shuaiwen Tang, Wei He 0008, Tengyu Long
Inf. Sci.6