Xinhua Zeng

dblp:168/2117 · DBLP profile ↗
← Back
23ranked-venue papers
0as first author
23since 2021 · last 2026
0000-0002-3903-0392ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 STNMamba: Mamba-Based Spatial-Temporal Normality Learning for Video Anomaly Detection
abstract
Video anomaly detection (VAD) has been extensively researched due to its potential for intelligent video systems. However, most existing methods based on CNNs and transformers still suffer from substantial computational burdens and have room for improvement in learning spatial-temporal normality. Recently, Mamba has shown great potential for modeling long-range dependencies with linear complexity, providing an effective solution to the above dilemma. To this end, we propose a lightweight and effective Mamba-based network named STNMamba, which incorporates carefully designed Mamba modules to enhance the learning of spatial-temporal normality. Firstly, we develop a dual-encoder architecture, where the spatial encoder equipped with Multi-Scale Vision Space State Blocks (MS-VSSB) extracts multi-scale appearance features, and the temporal encoder employs Channel-Aware Vision Space State Blocks (CA-VSSB) to capture significant motion patterns. Secondly, a Spatial-Temporal Interaction Module (STIM) is introduced to integrate spatial and temporal information across multiple levels, enabling effective modeling of intrinsic spatial-temporal consistency. Within this module, the Spatial-Temporal Fusion Block (STFB) is proposed to fuse the spatial and temporal features into a unified feature space, and the memory bank is utilized to store spatial-temporal prototypes of normal patterns, restricting the model's ability to represent anomalies. Extensive experiments on three benchmark datasets demonstrate that our STNMamba achieves competitive performance with fewer parameters and lower computational costs than existing methods.
Zhangxun Li, Mengyang Zhao 0002, Yang Liu 0246, Jiamu Sheng, Xinhua Zeng, Tian Wang 0002, Kewei Wu, Yu-Gang Jiang 0001
IEEE Trans. Multim.6
2025 Guiding Inter-domain Class Balancing With Salient Features For Domain Adaptive Object Detection
abstract
Although multi-scale alignment has improved domain adaptive object detection by addressing data distribution differences and annotation challenges, little attention has been given to class distribution differences between domains. Additionally, the utilization of feature information across different alignment levels is limited. To alleviate these issues, this paper proposes a novel domain adaptive method that leverages salient features for inter-domain class balancing. Our method consists of three core modules. Specifically, 1) the pixel feature salience-guided module enhances target focus and guides alignment at other scales, improving overall alignment capability; 2) the spatial domain feature purification module filters the noise, extracts salient features, and provides high-quality samples for alignment; 3) the instance relationship adaptive adjustment module adjusts instance weights for different classes to alleviate class distribution differences between different domains. Extensive experiments on multiple datasets demonstrate the effectiveness of our method.
Haiming Peng, Dingkang Yang, Weilong Lin, Xinhua Zeng
ICASSP5
2025 Towards Consistent and Accurate Multi-agent Trajectory Prediction via TrajFusionNet
Haiming Peng, Xinhua Zeng
ICIC (14)3
2025 Enhancing H&E-to-IHC Virtual Staining via Multi-Channel Correlation Learning
Hao Yang 0055, Xinhua Zeng, Juncen Guo, Guibing He, Zihui Li, Xiaochuan Zhang, Chengsheng Liao, Run Fang
ICIC (1)3
2025 Pipeline-VO: A High-Speed and Lightweight Stereo Visual Odometry with an Optimized Pipeline Architecture
Yazhou Yin, ShenChong Li, Xinhua Zeng
ICIC (14)4
2025 RDT-SLAM: Real-Time SLAM with Fast Area Division and Tracking in Dynamic Environments
Yazhou Yin, Xiao Jun, Zheng Jie, Shenchong Li, Xinhua Zeng
ICIC (14)5
2025 Rethinking prediction-based video anomaly detection from local-global normality perspective
Mengyang Zhao 0002, Xinhua Zeng, Yang Liu 0246, Jing Liu 0050, Chengxin Pang
Expert Syst. Appl.2
2024 XRNeRF: View-guided Neural Radiance Fields for Occlusion Removal
abstract
The Neural Radiance Fields (NeRF) technique has gained significant popularity as a means of generating new views by representing 3D scenes using an implicit space. This work builds a controlled viewpoint model of a scene by using a multi-view stereo geometry approach, which advances the development of multi-camera group collaboration techniques. Despite NeRF being a spatial query function based on spatial coordinates and observation directions, it cannot efficiently acquire points in the region behind the occlusion when there is occlusion interference in the dataset. For obtaining images that are free from occlusion, we propose the utilization of a neural point cloud to impose constraints on the scene. Our method involves filtering out occlusion by fitting the light distribution and restoring the obscured region by leveraging the isotropic characteristics of the target points obtained from the multiview point feature. The experimental results on existing datasets validate the effectiveness of our method.
Aoxue Li, Xinhua Zeng, Yunlong Du, Chengxin Pang
CSCWD2
2024 A Lightweight Multi-Scale Efficient Model for Breast Cancer Detection and Classification in Mammograms
Xinhua Zeng, Kai Cheng 0001, Run Fang, Chengsheng Liao, Jerome Plain
ICONIP (4)2
2024 DyHGDAT: Dynamic Hypergraph Dual Attention Network for multi-agent trajectory prediction
abstract
Modeling the interactions among agents based on their historical trajectories is key to precise multi-agent trajectory prediction. Hypergraph Convolutional Networks (HGCN) have become a proper choice for capturing high-order interactions among agents in this field. However, most existing works only consider static hypergraphs, and ignore that in a hypergraph, the power of influence varies between vertices (or hyperedges). Therefore, we propose DyHGDAT, a dynamic hypergraph dual attention network to capture the high-order interactions among agents, which not only models the evolution of hypergraph over time but also highlights the vertices and hyperedges with larger impacts. We apply DyHGDAT to a CVAE-based prediction system for predicting plausible trajectories. To validate the effectiveness of prediction, we evaluate our proposed method on two well-established trajectory prediction datasets: the ETH/UCY datasets and the Stanford Drone Dataset (SDD). The experimental results show that with DyHGDAT, the CVAE-based prediction system out-performs state-of-the-art methods by 12.5%/5.3% in ADE/FDE on ETH/UCY, and the improvement on SDD is 6.4%/7.4%.
Weilong Lin, Xinhua Zeng, Chengxin Pang, Jing Teng, Jing Liu 0050
ICRA2
2024 Denoising Diffusion-Augmented Hybrid Video Anomaly Detection via Reconstructing Noised Frames
Kai Cheng 0001, Yaning Pan, Yang Liu 0246, Xinhua Zeng
IJCAI4
2024 MAFNet: Multi-domain Features Attention-Based Fusion Network for Cross-Subject Motor Imagery Classification
abstract
Electroencephalography (EEG) motor imagery(MI) is a typical brain-machine interface signal that has been widely used in medical rehabilitation and devices control. Many solutions have achieved satisfied results in subject-dependent MI classification, however, most of the methods provide less generalization ability over cross-subject tasks. In this article, we propose a multi-domain features attention-based fusion network (MAFNet) for MI cross-subject classification. The end-to-end model is based on deep learning, leveraging the strategy of data deconstruction, followed by attention fusion to deal with highly non-stationary EEG signals. The raw signal components are firstly separated from frequency domain. Subsequently, feature extraction and encoding are performed from the spatial and temporal domains. After which the most discriminative multi-domain features are extracted by the inter-branch attention mechanism. In summary, our approach effectively decouples complex, abstract EEG signal, highlights key elements, and integrates them to discern generic patterns of MI across different subjects. The proposed model demonstrates superior performance over the state-of-the-art algorithms in two open-access datasets with an accuracy of 66.2% and 73.2% for MI cross-subject tasks, respectively.
Xinhua Zeng
IJCNN2
2024 Normality learning reinforcement for anomaly detection in surveillance videos
Kai Cheng 0001, Xinhua Zeng, Yang Liu 0246, Yaning Pan, Xinzhe Li 0004
Knowl. Based Syst.2
2023 Spatial-Temporal Graph Convolutional Network Boosted Flow-Frame Prediction For Video Anomaly Detection
abstract
Video Anomaly Detection (VAD) is a critical technology for intelligent surveillance systems and remains a challenging task in the signal processing community. An intuitive idea for VAD is to use a two-stream network to learn appearance and motion normality, respectively. However, existing approaches usually design a network architecture for the appearance stream with effort, then apply a similar architecture to the motion stream, ignoring the unique appearance and motion characteristics. In this paper, we propose STGCN-FFP, an unsupervised Spatial-Temporal Graph Convolutional Networks (STGCN) boosted Flow-Frame Prediction model. Specifically, we first design an STGCN-based memory module to extract and memorize normal patterns for optical flow, which is more suitable for learning motion normality. Then, we use a memory-augmented auto-encoder to model normal appearance patterns. Finally, the latent representation of two streams is fused to predict future frames, boosting the model to learn spatial-temporal normality. To our knowledge, STGCN-FFP is the first work applying STGCN to uniquely model the motion normality. Our method performs comparably to the state-of-the-art methods on three benchmarks.
Kai Cheng 0001, Xinhua Zeng, Yang Liu 0246, Mengyang Zhao 0002, Chengxin Pang, Xing Hu 0006
ICASSP2
2023 Memory-Augmented Spatial-Temporal Consistency Network for Video Anomaly Detection
Zhangxun Li, Mengyang Zhao 0002, Xinhua Zeng, Tian Wang 0002, Chengxin Pang
PRCV (6)3
2023 Learning Graph Enhanced Spatial-Temporal Coherence for Video Anomaly Detection
abstract
Video Anomaly Detection (VAD) is a critical yet challenging task in the signal processing community. Since part abnormal events cannot be detected by analyzing spatial or temporal information alone, learning spatial-temporal coherence has been proven the key to effective VAD. To this end, we propose a Graph Enhanced Spatial-Temporal Attention (GESTA) to address unsupervised VAD by learning the spatial-temporal coherence of normal events. Firstly, we propose a Dynamic Graph Recurrent Neural Network (DGRNN) to extract the motion patterns. Then, we propose a Spatial-Temporal Attention Module (STAM) to better model spatial-temporal coherence by integrating the prototypical spatial and temporal information. Finally, the fused spatial-temporal features are fed into the decoder to predict future frames. In testing phase, the anomaly with irregular information will result in poor prediction results. Experiments on three benchmarks demonstrate that our GESTA performs comparably to the state-of-the-art methods, and extensive analysis proves the effectiveness of DGRNN and STAM.
Kai Cheng 0001, Yang Liu 0246, Xinhua Zeng
IEEE Signal Process. Lett.3
2022 Look, Listen and Pay More Attention: Fusing Multi-Modal Information for Video Violence Detection
abstract
Violence detection is an essential and challenging problem in the computer vision community. Most existing works focus on single modal data analysis, which is not effective when multi-modality is available. Therefore, we propose a two-stage multi-modal information fusion method for violence detection: 1) the first stage adopts multiple instance learning strategies to refine video-level hard labels into clip-level soft labels, and 2) the next stage uses multi-modal information fused attention module to achieve fusion, and supervised learning is carried out using the soft labels generated at the first stage. Extensive empirical evidence on the XD-Violence dataset shows that our method outperforms the state-of-the-art methods.
Donglai Wei 0002, Chen-Geng Liu, Yang Liu 0246, Jing Liu 0050, Xiao-Guang Zhu, Xinhua Zeng
ICASSP6
2022 SGFusion: Camera-LiDAR Semantic and Geometric Fusion for 3D Object Detection
Xuhua Chen, Xinhua Zeng, Chengxin Pang
ICONIP (3)2
2022 Exploiting Spatial-temporal Correlations for Video Anomaly Detection
abstract
Video anomaly detection (VAD) remains a challenging task in the pattern recognition community due to the ambiguity and diversity of abnormal events. Existing deep learning-based VAD methods usually leverage proxy tasks to learn the normal patterns and discriminate the instances that deviate from such patterns as abnormal. However, most of them do not take full advantage of spatial-temporal correlations among video frames, which is critical for understanding normal patterns. In this paper, we address unsupervised VAD by learning the evolution regularity of appearance and motion in the long and short-term and exploit the spatial-temporal correlations among consecutive frames in normal videos more adequately. Specifically, we proposed to utilize the spatiotemporal long short-term memory (ST-LSTM) to extract and memorize spatial appearances and temporal variations in a unified memory cell. In addition, inspired by the generative adversarial network, we introduce a discriminator to perform adversarial learning with the ST-LSTM to enhance the learning capability. Experimental results on standard benchmarks demonstrate the effectiveness of spatial-temporal correlations for unsupervised VAD. Our method achieves competitive performance compared to the state-of-the-art methods with AUCs of 96.7%, 87.8%, and 73.1% on the UCSD Ped2, CUHK Avenue, and ShanghaiTech, respectively.
Mengyang Zhao 0002, Yang Liu 0246, Jing Liu 0050, Xinhua Zeng
ICPR4
2022 Attention-Based Auto-Encoder Framework for Abnormal Driving Detection
abstract
With the popularity of smartphones, abnormal driving detection via smartphone sensors has been proposed in recent years. However, existing methods are insufficient in exploring feature extraction, so the practical value is limited due to the low accuracy. To address this problem, we propose an attention-based auto-encoder framework for abnormal driving detection that combines the advantages of bi-directional long short-term memory and self-attention. Specifically, these two modules are embedded in the auto-encoder for modeling latent vector and exploring the internal correlations of spatial-temporal features, respectively, so as to improve the capability of reconstructing driving time series using small and representative features. We conduct experiments on the real-world datasets, and the results show that the proposed framework achieves significant performance with recall and F1-score of 96.2% and 95.0%, superior to the other baselines.
Jing Liu 0050, Yang Liu 0246, Donglai Wei 0002, Xinhua Zeng
ISCAS5
2022 Multi-level Attention Fusion for Multimodal Driving Maneuver Recognition
abstract
Sensor-based driving maneuver recognition (DMR) is a fundamental and challenging task in ubiquitous computing, which uses multimodal signals from embedded sensors such as accelerometers and gyroscopes to recognize driving maneuvers. However, the spatial-temporal features from neural networks are often treated equally, which may limit the performance of the model in predicting maneuvers. In this paper, we propose a novel hybrid neural network model based on multi-level attention fusion for multimodal DMR. The proposed model utilizes convolutional neural networks and gated recurrent unit to extract temporal-spatial features from multimodal sensing signals and propose the multi-level attention fusion to explore the significant patterns over local and global periods. In addition, We design three different levels of fusion (early, late, and full fusion) to explore the effects of different attention fusions on the model. Extensive experiments on the real-world dataset show that the proposed model achieves superior performance to the baseline methods, and multi-level attention fusion brings 6.17% gain to the F1-score.
Jing Liu 0050, Yang Liu 0246, Chengwen Tian, Mengyang Zhao 0002, Xinhua Zeng
ISCAS5
2022 TOP-ALCM: A novel video analysis method for violence detection in crowded scenes
Xing Hu 0006, Zhe Fan, Linhua Jiang, Guoqiang Li 0001, Wenming Chen 0001, Xinhua Zeng, Genke Yang, Dawei Zhang 0009
Inf. Sci.7
2022 MSAF: Multimodal Supervise-Attention Enhanced Fusion for Video Anomaly Detection
abstract
The complementarity of multimodal signal is essential for video anomaly detection. However, existing methods either lack exploration to multimodal data or ignore the implicit alignment of multimodal features. In our work, we address this problem using a novel fusion method and propose a Multimodal Supervise-Attention enhanced Fusion (MSAF) framework under weak supervision. Our framework can be divided into two parts: 1) the multimodal labels refinement part refines video-level ground truth into pseudo clip-level labels for subsequent training, 2) the multimodal supervise-attention fusion network enhances features via implicitly aligning different information, then fusing them effectively to predict anomaly scores with the help of refined labels. We validate our framework on four challenging datasets: ShanghaiTech, UCF-Crime, LAD, and XD-Violence. Extensive experiments on the benchmarks demonstrate the effectiveness of our framework, which achieves comparable results on several benchmarks and outperforms current state-of-the-art methods on the XD-Violence audiovisual multimodal dataset.
Donglai Wei 0002, Yang Liu 0246, Xiaoguang Zhu, Jing Liu 0050, Xinhua Zeng
IEEE Signal Process. Lett.5