Yansong Peng

dblp:349/0991 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Efficient Event-Based Semantic Segmentation via Exploiting Frame-Event Fusion: A Hybrid Neural Network Approach
abstract
Event cameras have recently been introduced into image semantic segmentation, owing to their high temporal resolution and other advantageous properties. However, existing event-based semantic segmentation methods often fail to fully exploit the complementary information provided by frames and events, resulting in complex training strategies and increased computational costs. To address these challenges, we propose an efficient hybrid framework for image semantic segmentation, comprising a Spiking Neural Network branch for events and an Artificial Neural Network branch for frames. Specifically, we introduce three specialized modules to facilitate the interaction between these two branches: the Adaptive Temporal Weighting (ATW) Injector, the Event-Driven Sparse (EDS) Injector, and the Channel Selection Fusion (CSF) module. The ATW Injector dynamically integrates temporal features from event data into frame features, enhancing segmentation accuracy by leveraging critical dynamic temporal information. The EDS Injector effectively combines sparse event data with rich frame features, ensuring precise temporal and spatial information alignment. The CSF module selectively merges these features to optimize segmentation performance. Experimental results demonstrate that our framework not only achieves state-of-the-art accuracy across the DDD17-Seg, DSEC-Semantic, and M3ED-Semantic datasets but also significantly reduces energy consumption, achieving a 65% reduction on the DSEC-Semantic dataset.
Hebei Li, Yansong Peng, Jiahui Yuan, Peixi Wu, Jin Wang 0023, Yueyi Zhang 0001, Xiaoyan Sun 0001
AAAI2
2025 Spiking Point Transformer for Point Cloud Classification
abstract
Spiking Neural Networks (SNNs) offer an attractive and energy-efficient alternative to conventional Artificial Neural Networks (ANNs) due to their sparse binary activation. When SNN meets Transformer, it shows great potential in 2D image processing. However, their application for 3D point cloud remains underexplored. To this end, we present Spiking Point Transformer (SPT), the first transformer-based SNN framework for point cloud classification. Specifically, we first design Queue-Driven Sampling Direct Encoding for point cloud to reduce computational costs while retaining the most effective support points at each time step. We introduce the Hybrid Dynamics Integrate-and-Fire Neuron (HD-IF), designed to simulate selective neuron activation and reduce over-reliance on specific artificial neurons. SPT attains state-of-the-art results on three benchmark datasets that span both real-world and synthetic datasets in the SNN domain. Meanwhile, the theoretical energy consumption of SPT is at least 6.4x less than its ANN counterpart.
Peixi Wu, Bosong Chai, Hebei Li, Menghua Zheng, Yansong Peng, Xuan Nie, Yueyi Zhang 0001, Xiaoyan Sun 0001
AAAI5
2025 Enhancing Visual Question Answering Via Clustered In-Context Sequence Configuration
abstract
Recent advances in Multimodal In-Context Learning (M-ICL) for Multimodal Large Language Models (MLLMs) have attracted considerable attention. These developments primarily focus on configuring an in-context sequence for a given test case based on instance-level semantic similarity. However, high similarity among demonstrations in the sequence introduces inductive biases, which may mislead MLLMs and ultimately degrade their overall performance. To address this, we propose a novel cluster-based in-context configuration method that adaptively groups candidate data and selects demonstrations from each cluster. This method enhances the diversity within the sequence while preserving semantic consistency, enabling MLLMs to focus on the main intent of the demonstrations. The experimental results on four Visual Question Answering (VQA) benchmarks, including OK-VQA, VQAv2, VizWiz, and TextVQA, demonstrate the effectiveness of our proposed method.
Yijun Pan, Hebei Li, Feipeng Ma, Yansong Peng, Siying Wu, Xiaoyan Sun 0001
ICIP5
2025 D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution Refinement
abstract
We introduce D-FINE, a powerful real-time object detector that achieves outstanding localization precision by redefining the bounding box regression task in DETR models. D-FINE comprises two key components: Fine-grained Distribution Refinement (FDR) and Global Optimal Localization Self-Distillation (GO-LSD). FDR transforms the regression process from predicting fixed coordinates to iteratively refining probability distributions, providing a fine-grained intermediate representation that significantly enhances localization accuracy. GO-LSD is a bidirectional optimization strategy that transfers localization knowledge from refined distributions to shallower layers through self-distillation, while also simplifying the residual prediction tasks for deeper layers. Additionally, D-FINE incorporates lightweight optimizations in computationally intensive modules and operations, achieving a better balance between speed and accuracy. Specifically, D-FINE-L / X achieves 54.0% / 55.8% AP on the COCO dataset at 124 / 78 FPS on an NVIDIA T4 GPU. When pretrained on Objects365, D-FINE-L / X attains 57.1% / 59.3% AP, surpassing all existing real-time detectors. Furthermore, our method significantly enhances the performance of a wide range of DETR models by up to 5.3% AP with negligible extra parameters and training costs. Our code and models: https://github.com/Peterande/D-FINE.
Yansong Peng, Hebei Li, Peixi Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0001
ICLR1
2025 Create Anything Anywhere: Layout-Controllable Personalized Diffusion Model for Multiple Subjects
abstract
Diffusion models have significantly advanced text-to-image generation, laying the foundation for the development of personalized generative frameworks. However, existing methods lack precise layout controllability and overlook the potential of dynamic features of reference subjects in improving fidelity. In this work, we propose Layout-Controllable Personalized Diffusion (LCP-Diffusion) model, a novel framework that integrates subject identity preservation with flexible layout guidance in a tuning-free approach. Our model employs a Dynamic-Static Complementary Visual Refining module to comprehensively capture the intricate details of reference subjects, and introduces a Dual Layout Control mechanism to enforce robust spatial control across both training and inference stages. Extensive experiments validate that LCP-Diffusion excels in both identity preservation and layout controllability. To the best of our knowledge, this is a pioneering work enabling users to "create anything anywhere".
Hebei Li, Yansong Peng, Siying Wu, Yueyi Zhang 0001, Xiaoyan Sun 0001
ICME3
2025 Dome-DETR: DETR with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object Detection
abstract
Tiny object detection plays a vital role in drone surveillance, remote sensing, and autonomous systems, enabling the identification of small targets across vast landscapes. However, existing methods suffer from inefficient feature leverage and high computational costs due to redundant feature processing and rigid query allocation. To address these challenges, we propose Dome-DETR, a novel framework with Density-Oriented Feature-Query Manipulation for Efficient Tiny Object Detection. To reduce feature redundancies, we introduce a lightweight Density-Focal Extractor (DeFE) to produce clustered compact foreground masks. Leveraging these masks, we incorporate Masked Window Attention Sparsification (MWAS) to focus computational resources on the most informative regions via sparse attention. Besides, we propose Progressive Adaptive Query Initialization (PAQI), which adaptively modulates query density across spatial areas for better query allocation. Extensive experiments demonstrate that Dome-DETR achieves state-of-the-art performance (+3.3 AP on AI-TOD-V2 and +2.5 AP on VisDrone) while maintaining low computational complexity and a compact model size. Code is available at https://github.com/RicePasteM/Dome-DETR.
Zhangchi Hu, Peixi Wu, Jie Chen 0001, Huyue Zhu, Yansong Peng, Hebei Li, Xiaoyan Sun 0001
ACM Multimedia6
2024 Event-Assisted Low-Light Video Object Segmentation
abstract
In the realm of video object segmentation (VOS), the challenge of operating under low-light conditions persists, resulting in notably degraded image quality and compromised accuracy when comparing query and memory frames for similarity computation. Event cameras, characterized by their high dynamic range and ability to capture motion information of objects, offer promise in enhancing object visibility and aiding VOS methods under such low-light conditions. This paper introduces a pioneering framework tai-lored for low-light VOS, leveraging event camera data to elevate segmentation accuracy. Our approach hinges on two pivotal components: the Adaptive Cross-Modal Fusion (ACMF) module, aimed at extracting pertinent features while fusing image and event modalities to mitigate noise interference, and the Event-Guided Memory Matching (EGMM) module, designed to rectify the issue of in-accurate matching prevalent in low-light settings. Additionally, we present the creation of a synthetic LLE-DAVIS dataset and the curation of a real-world LLE-vas dataset, encompassing frames and events. Experimental evaluations corroborate the efficacy of our method across both datasets, affirming its effectiveness in low-light scenarios. The datasets are available at https://github.com/HebeiFast/EventLowLightVOS.
Hebei Li, Jin Wang 0023, Jiahui Yuan, Wenming Weng, Yansong Peng, Yueyi Zhang 0001, Zhiwei Xiong, Xiaoyan Sun 0001
CVPR6
2024 Scene Adaptive Sparse Transformer for Event-based Object Detection
abstract
While recent Transformer-based approaches have shown impressive performances on event-based object detection tasks, their high computational costs still diminish the low power consumption advantage of event cameras. Image-based works attempt to reduce these costs by introducing sparse Transformers. However, they display inade-quate sparsity and adaptability when applied to event-based object detection, since these approaches cannot balance the fine granularity of token-level sparsification and the efficiency of window-based Transformers, leading to re-duced performance and efficiency. Furthermore, they lack scene-specific sparsity optimization, resulting in information loss and a lower recall rate. To overcome these limi-tations, we propose the Scene Adaptive Sparse Transformer (SAST). SAST enables window-token co-sparsification, sig-nificantly enhancing fault tolerance and reducing compu-tational overhead. Leveraging the innovative scoring and selection modules, along with the Masked Sparse Window Self-Attention, SAST showcases remarkable scene-aware adaptability: It focuses only on important objects and dy-namically optimizes sparsity level according to scene complexity, maintaining a remarkable balance between performance and computational cost. The evaluation results show that SAST outperforms all other dense and sparse networks in both performance and efficiency on two large-scale event-based object detection datasets (1 Mpx and Genl). Code: https://github.com/Peterande/SAST.
Yansong Peng, Hebei Li, Yueyi Zhang 0001, Xiaoyan Sun 0001, Feng Wu 0005
CVPR1
2024 Event-Based Head Pose Estimation: Benchmark and Method
Jiahui Yuan, Hebei Li, Yansong Peng, Jin Wang 0023, Yuheng Jiang, Yueyi Zhang 0001, Xiaoyan Sun 0001
ECCV (15)3
2024 ESTME: Event-driven Spatio-temporal Motion Enhancement for Micro-Expression Recognition
abstract
The inherently rapid and subtle changes in micro-expressions pose significant challenges for micro-expression recognition (MER). Previous methods, typically relying on frame aggregation or optical flow, struggle to accurately capture subtle changes because of low frame rate. In this paper, we propose an Event-driven Spatio-temporal Motion Enhancement Network, which incorporates event signals captured by an event camera, to assist MER. Specifically, we introduce an Event-Enhanced Motion Extractor module to exploit event signals’ high temporal resolution property, enhancing subtle motion details. We also propose an Event-Guided Attention module to focus on subtle changes in specific areas, capturing more precise spatial features of micro-expressions. Experimental results on synthetic and real-world datasets demonstrate the superiority of our method on MER, showcasing its strong ability to capture subtle motion changes.
Peilin Xiao, Yueyi Zhang 0001, Dachun Kai, Yansong Peng, Zheyu Zhang 0002, Xiaoyan Sun 0001
ICME4
2023 Better and Faster: Adaptive Event Conversion for Event-Based Object Detection
abstract
Event cameras are a kind of bio-inspired imaging sensor, which asynchronously collect sparse event streams with many advantages. In this paper, we focus on building better and faster event-based object detectors. To this end, we first propose a computationally efficient event representation Hyper Histogram, which adequately preserves both the polarity and temporal information of events. Then we devise an Adaptive Event Conversion module, which converts events into Hyper Histograms according to event density via an adaptive queue. Moreover, we introduce a novel event-based augmentation method Shadow Mosaic, which significantly improves the event sample diversity and enhances the generalization ability of detection models. We equip our proposed modules on three representative object detection models: YOLOv5, Deformable-DETR, and RetinaNet. Experimental results on three event-based detection datasets (1Mpx, Gen1, and MVSEC-NIGHTL21) demonstrate that our proposed approach outperforms other state-of-the-art methods by a large margin, while achieving a much faster running speed (< 14 ms and < 4 ms for 50 ms event data on the 1Mpx and Gen1 datasets).
Yansong Peng, Yueyi Zhang 0001, Peilin Xiao, Xiaoyan Sun 0001, Feng Wu 0001
AAAI1
2023 GET: Group Event Transformer for Event-Based Vision
abstract
Event cameras are a type of novel neuromorphic sensor that has been gaining increasing attention. Existing event-based backbones mainly rely on image-based designs to extract spatial information within the image transformed from events, overlooking important event properties like time and polarity. To address this issue, we propose a novel Group-based vision Transformer backbone for Event-based vision, called Group Event Transformer (GET), which decouples temporal-polarity information from spatial information throughout the feature extraction process. Specifically, we first propose a new event representation for GET, named Group Token, which groups asynchronous events based on their timestamps and polarities. Then, GET applies the Event Dual Self-Attention block, and Group Token Aggregation module to facilitate effective feature communication and integration in both the spatial and temporal-polarity domains. After that, GET can be integrated with different downstream tasks by connecting it with various heads. We evaluate our method on four event-based classification datasets (Cifar10-DVS, N-MNIST, N-CARS, and DVS128Gesture) and two event-based object detection datasets (1Mpx and Gen1), and the results demonstrate that GET outperforms other state-of-the-art methods. The code is available at https://github.com/Peterande/GET-Group-Event-Transformer.
Yansong Peng, Yueyi Zhang 0001, Zhiwei Xiong, Xiaoyan Sun 0001, Feng Wu 0001
ICCV1