Peng He 0004

dblp:84/6016-4 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2025
0009-0004-5455-3812ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 EventMamba: Enhancing Spatio-Temporal Locality with State Space Models for Event-Based Video Reconstruction
abstract
Leveraging its robust linear global modeling capability, Mamba has notably excelled in computer vision. Despite its success, existing Mamba-based vision models have overlooked the nuances of event-driven tasks, especially in video reconstruction. Event-based video reconstruction (EBVR) demands spatial translation invariance and close attention to local event relationships in the spatio-temporal domain. Unfortunately, conventional Mamba algorithms apply static window partitions and standard reshape scanning methods, leading to significant losses in local connectivity. To overcome these limitations, we introduce EventMamba—a specialized model designed for EBVR task. EventMamba innovates by incorporating random window offset (RWO) in the spatial domain, moving away from the restrictive fixed partitioning. Additionally, it features a new consistent traversal serialization approach in the spatio-temporal domain, which maintains the proximity of adjacent events both spatially and temporally. These enhancements enable EventMamba to retain Mamba’s robust modeling capabilities while significantly preserving the spatio-temporal locality of event data. Comprehensive testing on multiple datasets shows that EventMamba markedly enhances video reconstruction, drastically improving computation speed while delivering superior visual quality compared to Transformer-based methods.
Chengjie Ge, Xueyang Fu, Peng He 0004, Chengzhi Cao, Zhengjun Zha
AAAI3
2025 Fine-grained hierarchical dynamics for image harmonization
Peng He 0004, Jun Yu 0001, Liuxue Ju, Fang Gao 0001
Neural Networks1
2025 Domain-Separated Bottleneck Attention Fusion Framework for Multimodal Emotion Recognition
abstract
As a focal point of research in various fields, human body language understanding has long been a subject of intense interest. Within this realm, the exploration of emotion recognition through the analysis of facial expressions, voice patterns, and physiological signals holds significant practical value. Compared with unimodal approaches, multimodal emotion recognition models leverage complementary information from vision, acoustic, and language modalities to robust perceive the human sentiment attitudes. However, the heterogeneity among modality signals leads to significant domain shifts, posing challenges for achieving balanced fusion. In this article, we propose a Domain-Separated Bottleneck Attention (DBA) Fusion Framework for human multimodal emotion recognition with lower computational complexity. Specifically, we partition each modality into two distinct domains: the invariant/private domain. The invariant domain contains crucial shared information, while the private domain aims to capture modality-specific representations. For the decomposed features, we introduce two sets of bottleneck cross-attention modules to effectively utilize the complementarity between domains to reduce redundant information. In each module, we interweave two Fusion Adapter blocks into the Self-Attention Transformer backbone. Each Fusion Adapter block integrates a small group of latent tokens as bridges for inter-modal and inter-domain interactions, mitigating the adverse effects of modality distribution differences and lowering computational costs. Extensive experimental results demonstrate that our method outperforms State-of-the-Art (SOTA) approaches across three widely used benchmark datasets.
Peng He 0004, Jun Yu 0001, Chengjie Ge, Lei Wang 0203, Zhen Kan
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Neuromorphic Event Signal-Driven Network for Video De-raining
abstract
Convolutional neural networks-based video de-raining methods commonly rely on dense intensity frames captured by CMOS sensors. However, the limited temporal resolution of these sensors hinders the capture of dynamic rainfall information, limiting further improvement in de-raining performance. This study aims to overcome this issue by incorporating the neuromorphic event signal into the video de-raining to enhance the dynamic information perception. Specifically, we first utilize the dynamic information from the event signal as prior knowledge, and integrate it into existing de-raining objectives to better constrain the solution space. We then design an optimization algorithm to solve the objective, and construct a de-raining network with CNNs as the backbone architecture using a modular strategy to mimic the optimization process. To further explore the temporal correlation of the event signal, we incorporate a spiking self-attention module into our network. By leveraging the low latency and high temporal resolution of the event signal, along with the spatial and temporal representation capabilities of convolutional and spiking neural networks, our model captures more accurate dynamic information and significantly improves de-raining performance. For example, our network achieves a 1.24dB improvement on the SynHeavy25 dataset compared to the previous state-of-the-art method, while utilizing only 39% of the parameters.
Chengjie Ge, Xueyang Fu, Peng He 0004, Chengzhi Cao, Zhengjun Zha
AAAI3
2024 Micro-Expression Spotting Based on Optical Flow Feature with Boundary Calibration
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Zhongpeng Cai, Jianqing Sun, Jiaen Liang
ACM Multimedia4
2024 Temporal-Informative Adapters in VideoMAE V2 and Multi-Scale Feature Fusion for Micro-Expression Spotting-then-Recognize
Jun Yu 0001, Gongpeng Zhao, Peng He 0004, Jianqing Sun, Jiaen Liang
ACM Multimedia4
2023 FSR-Net: Deep Fourier Network for Shadow Removal
abstract
The presence of shadows degrades the performance of various multimedia tasks. Image shadow removal aims at restoring the background of shadow regions, which is generally an open challenge. Unlike most existing deep learning-based methods that focus on restoring such degradations in the spatial domain, we introduce a novel shadow removal method that also exploits frequency domain information. Specifically, we firstly revisit the frequency characteristics of shadow images via Fourier transform, where amplitude components contain most lightness information and phase components are related to structure information. To this end, we propose a two-stage deep Fourier shadow removal network (FSR-Net) to enhance the brightness of shadow regions, and correspondingly improve the shadow removal performance of whole images. For each stage, it consists of an amplitude recovery network and a phase recovery network to progressively reconstruct the lightness and structure components. To facilitate the learning of these two representations, we introduce the frequency and spatial interaction blocks to process the local spatial features and the global frequency information separately. Extensive experiments demonstrate that FSR-Net achieves superior results than other approaches with fewer parameters. For example, our method obtains a 1.05dB improvement on ISTD[34] dataset over the previous state-of-the-art method [43] with 0.30M parameters.
Jun Yu 0001, Peng He 0004, Ziqi Peng
ACM Multimedia2
2022 Facial Expression Spotting Based on Optical Flow Features
abstract
The purpose of micro expression (ME) and macro expression (MaE) spotting task is to locate the onset and offset frames of MaE and ME clips. Compared with MaEs, MEs are shorter in duration and lower in intensity, which makes MEs harder to be spotted. In this paper, we propose an efficient pipeline based on optical flow features to spot MEs and MaEs. We crop and align the faces and select the eyebrows area, nose area, and mouth area as our regions of interest to exclude the interference of extraneous factors on the face expression representation. Then, we extract the optical flow in these regions and enhance the features of the expressions in the optical flow with low-pass filter and EMD method. Finally, the sliding window method is used to locate the peaks of optical flow features and get the intervals containing MEs or MaEs. We evaluate the performance of our method on the MEGC2022-TestSet including 10 long videos from SAMM and CAS(ME)3 and achieve the first place in the MEGC2022 Challenge. The results prove the effectiveness of our method.
Jun Yu 0001, Zhongpeng Cai, Guochen Xie, Peng He 0004
ACM Multimedia5
2022 Micro Expression Generation with Thin-plate Spline Motion Model and Face Parsing
abstract
Micro-expression generation aims at transfering the expression from the driving videos to the source images, which can be viewed as a motion transfer task. Recently, several works have been proposed to tackle this problem and achieve great performance. However, due to the intrinsic complexity of the face motion and different attributes of face regions, the task still remains challenging. In this paper, we propose an end-to-end unsupervised motion transfer network to tackle this challenge. As the motion of the face is non-rigid, we adopt an effective and flexible thin-plate spline motion estimation method to estimate the optical flow of the face motion. What's more, we find that several faces with eyeglasses show weird deformation in motion transfering. Thus, we introduce face parsing method to pay specific attention to the eyeglasses regions to ensure the reasonability of the deformation. We conduct several experiments on the provided datasets of the ACM MM 2022 micro-expression grand challenge (MEGC2022) and compare our method with several other typical methods. In comparison, our method shows the best performance. We (Team: USTC-IAT-United) also compare our method with other competitors' in MEGC2022, and the expert evaluation results show that our method performs best, which verifies the effectiveness of our method. Our code is available at https://github.com/HowToNameMe/micro-expression
Jun Yu 0001, Guochen Xie, Zhongpeng Cai, Peng He 0004, Fang Gao 0001, Qiang Ling 0001
ACM Multimedia4