Menghao Zhang 0004

dblp:188/5710-4 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0001-9713-9938ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 VALU: A Benchmark for Video Anomaly Temporal Localization and Understanding at Multiple Semantic Levels
abstract
Yixiao He, Menghao Zhang, Haifeng Sun, Jing Wang, Kangheng Lin, Jinghan Wang, Chenye Xu, Pengfei Ren, Qi Qi, Jingyu Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yixiao He, Menghao Zhang 0004, Haifeng Sun 0001, Jing Wang 0039, Kangheng Lin, Chenye Xu, Pengfei Ren 0001, Qi Qi 0001, Jingyu Wang 0001
ACL (1)2
2026 Enhancing MLLMs for Online Understanding in Video Services via Preference Optimization
abstract
Online video understanding is pivotal for emerging video streaming services. However, existing Multimodal Large Language Models (MLLMs) encounter significant challenges in this domain, specifically in maintaining holistic visual perception under strict token budgets and attaining grounded semantic reasoning amidst dynamic context changes. To mitigate these issues, we propose OV-DPO, a systematic multimodal preference optimization approach. To ensure holistic visual perception, we introduce a visual preference objective that compels the model to strictly ground its reasoning in clear visual evidence. This is supported by the QGDFR strategy, which dynamically reallocates resolution budgets to preserve critical visual details for chosen samples while intentionally degrading them for rejected ones to construct valid visual contrasts. Concurrently, to promote grounded semantic reasoning, we employ a textual preference objective. We design the SROVA framework to construct high-quality preference pairs, utilizing self-refinement to generate factual chosen responses and visual degradation to induce language-prior-driven hallucinations as rejected samples. By jointly optimizing these objectives alongside anchored and supervised fine-tuning terms, OV-DPO ensures training stability and robust alignment. Experimental results demonstrate that our approach significantly outperforms baselines such as SFT and DPO. Notably, utilizing only 29K high-quality samples, our 3B-scale model achieves results competitive with significantly larger-scale models on both online and offline benchmarks, offering an efficient solution for online video understanding.
Qi Qi 0001, Yixiao He, Menghao Zhang 0004, Haifeng Sun 0001, Pengfei Ren 0001, Huazheng Wang, Jianxin Liao, Jingyu Wang 0001
IEEE Trans. Serv. Comput.3
2025 Prior-Aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose Estimation
Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Menghao Zhang 0004, Lei Zhang 0094, Jing Wang 0039, Jianxin Liao
ICCV6
2025 Masked Self-Supervised Learning and Semantic Noise Separation for Video Anomaly Detection
abstract
Recent progress in video anomaly detection assumes that anomalies cannot be effectively reconstructed because they remain unseen during training. However, we observe that most existing methods excessively rely on appearance features, resulting in the accurate reconstruction of anomalies with subtle short-term appearance variations, which we refer to as appearance confusion. Meanwhile, many approaches fail to exploit sufficient semantic distinction, resulting in motion confusion for anomalies with motion patterns similar to normal ones. In this paper, we propose a masked self-supervised learning-based framework, which effectively addresses the two confusions by exploring context-aware motion patterns and discriminative semantic normality representations. First, we introduce reconstructing multi-pattern masked spatiotemporal information to motivate the model to capture motion patterns that focus on long-term context. Then, we design a semantic noise separation network to address motion confusion, facilitating the construction of semantic normality boundaries through semantic-aware separation. Extensive experiments on the Avenue and ShanghaiTech datasets validate the effectiveness of our proposed method.
Menghao Zhang 0004, Lei Zhang 0094, Qi Qi 0001, Haifeng Sun 0001, Pengfei Ren 0001, Bo He 0003, Jing Wang 0039, Jingyu Wang 0001
ICME2
2025 A Dual-Branch 3D Spatial-Aware Latent Diffusion for Realistic Depth Image Synthesis
abstract
Synthetic images serve as a promising alternative to real images in 3D hand pose estimation, providing accurate annotations at a lower cost. However, the domain gap between real and synthetic images constrains the generalization ability of hand pose estimation trained on synthetic data. Previous methods rely on Generative Adversarial Networks (GANs) for domain translation; however, they fail to achieve realistic depth synthesis due to instability and limited image quality. Diffusion models provide high-quality synthesis due to their stability and controllability. However, existing methods often ignore the 3D structure awareness in hand image generation. In this paper, we propose a Dual-Branch 3D Spatial-Aware Latent Diffusion (DSW-LD) for realistic depth image generation. The Global Structure Module (GSM) and the Local Geometry Module (LGM) complement each other, with GSM capturing global spatial structure through coarse-grained 3D joint features and LGM focusing on local geometric details using fine-grained 3D mesh representations. To maintain the global structure consistency, we adopt a layer-aware injection mechanism that enables the model to adaptively learn the optimal representation from fused 2D latent representations and 3D joint features. To explicitly align 3D and 2D features of local regions and enhance the flexibility of feature matching, we design a dynamic depth-aware interpolation to project 3D mesh features into 2D image space. Both quantitative and qualitative experimental results demonstrate the superiority of our method over the state-of-the-arts for realistic depth synthesis. Compared to training only on real depth images, our method enables the hand pose estimator to achieve significantly better performance with our synthetic data and less real data (10%).
Shuang Hao 0017, Pengfei Ren 0001, Lei Zhang 0094, Haifeng Sun 0001, Pan Ting, Menghao Zhang 0004, Cong Liu 0046, Qi Qi 0001, Jianxin Liao, Jingyu Wang 0001
ACM Multimedia6
2025 Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose Estimation
abstract
Multi-view 3D human pose estimation (HPE) leverages complementary information across views to improve accuracy and robustness. Traditional methods rely on camera calibration to establish geometric correspondences, which is sensitive to calibration accuracy and lacks flexibility in dynamic settings. Calibration-free approaches address these limitations by learning adaptive view interactions, typically leveraging expressive and flexible continuous representations. However, as the multiview interaction relationship is learned entirely from data without constraint, they are vulnerable to noisy input, which can propagate, amplify and accumulate errors across all views, severely corrupting the final estimated pose. To mitigate this, we propose a novel framework that integrates a noise-resilient discrete prior into the continuous representation-based model. Specifically, we introduce the \textit{UniCodebook}, a unified, compact, robust, and discrete representation complementary to continuous features, allowing the model to benefit from robustness to noise while preserving regression capability. Furthermore, we further propose an attribute-preserving and complementarity-enhancing Discrete-Continuous Spatial Attention (DCSA) mechanism to facilitate interaction between discrete priors and continuous pose features. Extensive experiments on three representative datasets demonstrate that our approach outperforms both calibration-required and calibration-free methods, achieving state-of-the-art performance.
Geng Chen 0006, Pengfei Ren 0001, Xufeng Jian, Haifeng Sun 0001, Menghao Zhang 0004, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001
NeurIPS5
2025 Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence Patterns
abstract
Large Vision-Language Models (LVLMs) pretrained on large-scale multimodal data have shown promising capabilities in Video Anomaly Detection (VAD). However, their ability to reason about abnormal events based on scene semantics remains underexplored. In this paper, we investigate LVLMs’ behavior in VAD from a visual-textual co-occurrence perspective, focusing on whether their decisions are driven by statistical shortcuts between visual instances and textual phrases. By analyzing visual-textual co-occurrence in pretraining data and conducting experiments under different data settings, we reveal a hallucination phenomenon: LVLMs tend to rely on co-occurrence patterns between visual instances and textual phrases associated with either normality or abnormality, leading to incorrect predictions when these high-frequency objects appear in semantically mismatched contexts. To address this issue, we propose VAD-DPO, a direct preference optimization method supervised with counter-example pairs. By constructing visually similar but semantically contrasting video clips, VAD-DPO encourages the model to align its predictions with the semantics of scene rather than relying on co-occurrence patterns. Extensive experiments on six benchmark datasets demonstrate the effectiveness of VAD-DPO in enhancing both anomaly detection and reasoning performance, particularly in scene-dependent scenarios.
Menghao Zhang 0004, Huazheng Wang, Pengfei Ren 0001, Kangheng Lin, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao, Jingyu Wang 0001
NeurIPS1
2024 Multi-Scale Video Anomaly Detection by Multi-Grained Spatio-Temporal Representation Learning
abstract
Recent progress in video anomaly detection suggests that the features of appearance and motion play crucial roles in distinguishing abnormal patterns from normal ones. However, we note that the effect of spatial scales of anomalies is ignored. The fact that many abnormal events occur in limited localized regions and severe background noise in-terferes with the learning of anomalous changes. Mean-while, most existing methods are limited by coarse-grained modeling approaches, which are inadequate for learning highly discriminative features to discriminate subtle differences between small-scale anomalies and normal patterns. To this end, this paper address multi-scale video anomaly detection by multi-grained spatiotemporal representation learning. We utilize video continuity to design three proxy tasks to perform feature learning at both coarse-grained and fine-grained levels, i.e., continuity judgment, discontinuity localization, and missing frame estimation. In particular, we formulate missing frame estimation as a contrastive learning task in feature space instead of a reconstruction task in RGB space to learn highly discriminative features. Experiments show that our proposed method outperforms state-of-the-art methods on four datasets, especially in scenes with small-scale anomalies.
Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Pengfei Ren 0001, Ruilong Ma, Jianxin Liao
CVPR1
2024 Safeguarding Sustainable Cities: Unsupervised Video Anomaly Detection through Diffusion-based Latent Pattern Learning
Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Pengfei Ren 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao
IJCAI1
2024 Enhanced Anomaly Detection in Dashcam Videos: Dual GAN Approach with Swin-Unet for Optical Flow and Region of Interest Analysis
abstract
Video anomaly detection plays a crucial role in the field of autonomous driving to ensure driving safety. Most existing video anomaly detection methods exhibit mediocre performance when analyzing video frames captured by dynamic cameras. To enhance their performance on dynamic cameras, we propose a video anomaly detection method based on Swin-Unet to separately predict optical flow and images cropped with ROI in the form of dual GAN. By integrating the predicted results of optical flow and ROI-cropped images, the model’s ability to learn dynamic information is significantly improved. The experimental results indicate that our proposed method outperforms state-of-the-art methods in terms of AUC on the Car Crash dataset and RetroTrucks dataset.
Haodong Ru, Menghao Zhang 0004, Pengfei Ren 0001, Haifeng Sun 0001, Qi Qi 0001, Lejian Zhang, Jingyu Wang 0001
IJCNN2
2024 Video Anomaly Detection via Progressive Learning of Multiple Proxy Tasks
abstract
Learning multiple proxy tasks is a popular training strategy in semi-supervised video anomaly detection. However, the traditional method of learning multiple proxy tasks simultaneously is prone to suboptimal solutions, and simply executing multiple proxy tasks sequentially cannot ensure continuous performance improvement. In this paper, we thoroughly investigate the impact of task composition and training order on performance enhancement. We find that ensuring continuous performance improvement in multi-task learning requires different but continuous optimization objectives in different training phases. To this end, a training strategy based on progressive learning is proposed to enhance the multi-task learning in VAD. The learning objectives of the model in previous phases contribute to the training in subsequent phases. Specifically, we decompose video anomaly detection into three phases: perception, comprehension, and inference, continuously refining the learning objectives to enhance model performance. In the three phases, we perform the visual task, the semantic task and the open-set task in turn to train the model. The model learns different levels of features and focuses on different types of anomalies in different phases. Extensive experiments demonstrate the effectiveness of our method, highlighting that the benefits derived from the progressive learning transcend specific proxy tasks.
Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Pengfei Ren 0001, Haifeng Sun 0001, Zirui Zhuang, Huazheng Wang, Lei Zhang 0094, Jianxin Liao
ACM Multimedia1
2024 Cognition Guided Video Anomaly Detection Framework for Surveillance Services
abstract
The aim of surveillance services is to detect anomalous events that occur in given surveillance videos. Most existing video anomaly detection methods rely on minimizing reconstruction or prediction errors due to the lack of abnormal data, which results in poor generalization and overfitting. In fact, cognitions for anomalies in surveillance videos mainly relies on crucial relationships, including ones between objects and ones between objects and scenes. Focusing on this property of anomaly detection, aCognitionGuidedVideoAnomalyDetection framework based on prior knowledge is proposed, calledCG-VAD. CG-VAD introduces both explicit and implicit prior knowledge into the frame prediction network to let the model exploit crucial relationships. Explicit knowledge containing crucial relationships related to anomaly is introduced into the anomaly detection model through a proposed embedding network based on multi-layer Graph Convolutional Networks. Implicit knowledge in the form of learnable parameters enhances the ability of the model to learn crucial relationships through prompt tuning. By integrating prior knowledge to focus the model on the relationships associated with the anomaly, we find that CG-VAD is not only quick to adapt to new real-world scenarios, but it is also able to recognize the type of anomaly. We have conducted extensive experiments on four benchmark datasets and the results indicate that the proposed method outperforms previous methods. Specifically, CG-VAD achieves an AUROC score of 87.2$\%$on the ShanghaiTech dataset. Code is available athttps://github.com/zmh0124/CG-VAD.
Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001, Jianxin Liao
IEEE Trans. Serv. Comput.1
2023 Robust Video Anomaly Detection Framework via Prior Knowledge and Multi-Path Frame Prediction
abstract
Video anomaly detection aims to automatically detect abnormal objects or behaviors. Most existing methods tackle the problem by minimizing the reconstruction errors stemming from the lack of anomalous data, which leads to poor interpretability and robustness. Focus on the context-dependent nature of anomaly detection, a robust unsupervised Video Anomaly Detection framework based on Knowledge and Frame Prediction is proposed, called VAD-KFP. Prior knowledge which contains the context of anomaly is introduced into the multi-path frame prediction network through multi-layer Graph Convolutional Networks. By integrating the prior knowledge to accurately define anomalies, VAD-KFP is robust to different scenarios and is able to recognize the type of anomaly. An extensive range of experiments have been conducted on three benchmarks, the results of which indicate that our method outperforms strong baselines. Specifically, VAD-KFP obtains an AUROC score of 91.6% for the Avenue dataset.
Menghao Zhang 0004, Jingyu Wang 0001, Jing Wang 0039, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001
ICASSP1