EDBT 2026 Demo / reviewers in the wild / expert
Mengyang Zhao 0002
dblp:60/10173-2
· DBLP profile ↗
16ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0001-8322-0479ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and UnderstandingabstractLVLMs have been shown to perform excellently in image-level tasks such as VQA and caption. However, in many instance-level tasks, such as visual grounding and object detection, LVLMs still show performance gaps compared to previous expert models. Meanwhile, although pedestrian tracking is a classical task, there have been a number of new topics in combining object tracking and natural language, such as Referring MOT, Cross-view Referring MOT, and Semantic MOT. These tasks emphasize that models should understand the tracked object at an advanced semantic level, which is exactly where LVLMs excel. In this paper, we propose a new unified Pedestrian Tracking framework, namely OmniPT, which can track, track based on reference and generate semantic understanding of tracked objects interactively. We address two issues: how to model the tracking task into a task that foundation models can perform, and how to make the model output formatted answers. To this end, we implement a training phase consisting of RL-Mid Training-SFT-RL. Based on the pre-trained weights of the LVLM, we first perform a simple RL phase to enable the model to output fixed and supervisable bounding box format. Subsequently, we conduct a mid-training phase using a large number of pedestrian-related datasets. Finally, we perform supervised fine-tuning on several pedestrian tracking datasets, and then carry out another RL phase to improve the model's tracking performance and enhance its ability to follow instructions. We conduct experiments on tracking benchmarks and the experimental results demonstrate that the proposed method can perform better than the previous methods. Teng Fu 0001, Mengyang Zhao 0002, Ke Niu 0004, Kaixin Peng, Bin Li 0015 |
AAAI | 2 |
| 2026 | From Intent to Execution: Multimodal Chain-of-Thought Reinforcement Learning for Precise CAD Code GenerationabstractComputer-Aided Design (CAD) plays a vital role in engineering and manufacturing, yet current CAD workflows require extensive domain expertise and manual modeling effort. Recent advances in large language models (LLMs) have made it possible to generate code from natural language, opening new opportunities for automating parametric 3D modeling. However, directly translating human design intent into executable CAD code remains highly challenging, due to the need for logical reasoning, syntactic correctness, and numerical precision. In this work, we propose CAD-RL, a multimodal Chain-of-Thought (CoT) guided reinforcement learning post training framework for CAD modeling code generation. Our method combines CoT-based Cold Start with goal-driven reinforcement learning post training using three task-specific rewards: executability reward, geometric accuracy reward, and external evaluation reward. To ensure stable policy learning under sparse and high-variance reward conditions, we introduce three targeted optimization strategies: Trust Region Stretch for improved exploration, Precision Token Loss for enhanced dimensions parameter accuracy, and Overlong Filtering to reduce noisy supervision. To support training and benchmarking, we release ExeCAD, a noval dataset comprising 16,540 real-world CAD examples with paired natural language and structured design language descriptions, executable CADQuery scripts, and rendered 3D models. Experiments demonstrate that CAD-RL achieves significant improvements in reasoning quality, output precision, and code executability over existing VLMs. Ke Niu 0004, Haiyang Yu 0004, Mengyang Zhao 0002, Teng Fu 0001, Bin Li 0015, Xiangyang Xue 0001 |
AAAI | 4 |
| 2026 | FDGReID: Federated Domain Generalization for Person Re-identification
Ke Niu 0004, Haiyang Yu 0004, Teng Fu 0001, Mengyang Zhao 0002, Bin Li 0015, Xuelin Qian, Xiangyang Xue 0001 |
Mach. Learn. | 4 |
| 2026 | STNMamba: Mamba-Based Spatial-Temporal Normality Learning for Video Anomaly DetectionabstractVideo anomaly detection (VAD) has been extensively researched due to its potential for intelligent video systems. However, most existing methods based on CNNs and transformers still suffer from substantial computational burdens and have room for improvement in learning spatial-temporal normality. Recently, Mamba has shown great potential for modeling long-range dependencies with linear complexity, providing an effective solution to the above dilemma. To this end, we propose a lightweight and effective Mamba-based network named STNMamba, which incorporates carefully designed Mamba modules to enhance the learning of spatial-temporal normality. Firstly, we develop a dual-encoder architecture, where the spatial encoder equipped with Multi-Scale Vision Space State Blocks (MS-VSSB) extracts multi-scale appearance features, and the temporal encoder employs Channel-Aware Vision Space State Blocks (CA-VSSB) to capture significant motion patterns. Secondly, a Spatial-Temporal Interaction Module (STIM) is introduced to integrate spatial and temporal information across multiple levels, enabling effective modeling of intrinsic spatial-temporal consistency. Within this module, the Spatial-Temporal Fusion Block (STFB) is proposed to fuse the spatial and temporal features into a unified feature space, and the memory bank is utilized to store spatial-temporal prototypes of normal patterns, restricting the model's ability to represent anomalies. Extensive experiments on three benchmark datasets demonstrate that our STNMamba achieves competitive performance with fewer parameters and lower computational costs than existing methods. Zhangxun Li, Mengyang Zhao 0002, Yang Liu 0246, Jiamu Sheng, Xinhua Zeng, Tian Wang 0002, Kewei Wu, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 2 |
| 2026 | A New Semi-Supervised Video Anomaly Detection Baseline in Lack of Anomalous SamplesabstractVideo anomaly detection (VAD) has been widely studied for its important applications in multimedia community. Recently, many Weakly Supervised VAD (WS-VAD) methods have been proposed, which tend to treat VAD as a classification task through multiple instance learning and result in the need to collect sufficient anomaly classes and samples to be used for training a classifier. However, anomaly events tend to be open-set and rare in real-world applications, so we often have difficulty collecting all anomaly classes and enough sample anomalies, which is a difficult situation for WS-VAD to cope with. To this end, we consider to treat VAD as an out-of-distribution detection task rather than a classification task and propose a simple but effective semi-supervised baseline method. First, we leverage the powerful zero-shot capability of large visual language models to generate summary text descriptions for videos and extract visual features as intermediates for subsequent use. Next, we use a text encoder to extract language features and combine them with visual features to obtain robust multimodal features. Finally, we introduce an out-of-distribution detection method that learns the center of normality in multimodal space from normal and unlabeled samples, while deviating abnormal samples from the center to cope with the scarcity of abnormal samples. To implement our baseline method, we also provide a new semi-supervised dataset by reorganizing an existing benchmark, which is the first available dataset in the VAD community that provides trimmed videos consisting of complete abnormal events. Experiments demonstrate that our method performs more robustly when fewer anomaly classes and anomaly samples collected. Mengyang Zhao 0002, Haiyang Yu 0004, Teng Fu 0001, Yang Liu 0246, Wei Zhou 0021, Bin Li 0015, Xiangyang Xue 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | ChatReID: Open-Ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language ModelsabstractPerson re-identification (Re-ID) is a crucial task in computer vision, aiming to recognize individuals across non-overlapping camera views. While recent advanced vision-language models (VLMs) excel in logical reasoning and multi-task generalization, their applications in Re-ID tasks remain limited. They either struggle to perform accurate matching based on identity-relevant features or assist image-dominated branches as auxiliary semantics. In this paper, we propose a novel framework ChatReID, that shifts the focus towards a text-side-dominated retrieval paradigm, enabling flexible and interactive re-identification. To integrate the reasoning abilities of language models into Re-ID pipelines, We first present a large-scale instruction dataset, which contains more than 8 million prompts to promote the model fine-tuning. Next. we introduce a hierarchical progressive tuning strategy, which endows Re-ID ability through three stages of tuning, i.e., from person attribute understanding to fine-grained image retrieval and to multi-modal task reasoning. Extensive experiments across ten popular benchmarks demonstrate that ChatReID outperforms existing methods, achieving state-of-the-art performance in all Re-ID tasks. More experiments demonstrate that ChatReID not only has the ability to recognize fine-grained details but also to integrate them into a coherent reasoning process. Ke Niu 0004, Haiyang Yu 0004, Mengyang Zhao 0002, Teng Fu 0001, Siyang Yi, Bin Li 0015, Xuelin Qian, Xiangyang Xue 0001 |
ICCV | 3 |
| 2025 | CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-TuningabstractComputer-Aided Design (CAD) is pivotal in industrial manufacturing, with orthographic projection reasoning foundational to its entire workflow—encompassing design, manufacturing, and simulation. However, prevailing deep-learning approaches employ standard 3D reconstruction pipelines as an alternative, which often introduce imprecise dimensions and limit the parametric editability required for CAD workflows. Recently, some researchers adopt vision–language models (VLMs), particularly supervised fine-tuning (SFT), to tackle CAD-related challenges. SFT shows promise but often devolves into pattern memorization, resulting in poor out-of-distribution (OOD) performance on complex reasoning tasks. To tackle these limitations, we introduce CReFT-CAD, a two-stage fine-tuning paradigm: first, a curriculum-driven reinforcement learning stage with difficulty-aware rewards to steadily build reasoning abilities; second, supervised post-tuning to refine instruction following and semantic extraction. Complementing this, we release TriView2CAD, the first large-scale, open-source benchmark for orthographic projection reasoning, comprising 200,000 synthetic and 3,000 real-world orthographic projections with precise dimensional annotations and six interoperable data modalities. Benchmarking leading VLMs on orthographic projection reasoning, we show that CReFT-CAD significantly improves reasoning accuracy and OOD generalizability in real-world scenarios, providing valuable insights to advance CAD reasoning research. The code and adopted datasets are available at \url{https://github.com/KeNiu042/CReFT-CAD}. Ke Niu 0004, Haiyang Yu 0004, Teng Fu 0001, Mengyang Zhao 0002, Bin Li 0015, Xiangyang Xue 0001 |
NeurIPS | 6 |
| 2025 | Rethinking prediction-based video anomaly detection from local-global normality perspective
Mengyang Zhao 0002, Xinhua Zeng, Yang Liu 0246, Jing Liu 0050, Chengxin Pang |
Expert Syst. Appl. | 1 |
| 2023 | Spatial-Temporal Graph Convolutional Network Boosted Flow-Frame Prediction For Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) is a critical technology for intelligent surveillance systems and remains a challenging task in the signal processing community. An intuitive idea for VAD is to use a two-stream network to learn appearance and motion normality, respectively. However, existing approaches usually design a network architecture for the appearance stream with effort, then apply a similar architecture to the motion stream, ignoring the unique appearance and motion characteristics. In this paper, we propose STGCN-FFP, an unsupervised Spatial-Temporal Graph Convolutional Networks (STGCN) boosted Flow-Frame Prediction model. Specifically, we first design an STGCN-based memory module to extract and memorize normal patterns for optical flow, which is more suitable for learning motion normality. Then, we use a memory-augmented auto-encoder to model normal appearance patterns. Finally, the latent representation of two streams is fused to predict future frames, boosting the model to learn spatial-temporal normality. To our knowledge, STGCN-FFP is the first work applying STGCN to uniquely model the motion normality. Our method performs comparably to the state-of-the-art methods on three benchmarks. Kai Cheng 0001, Xinhua Zeng, Yang Liu 0246, Mengyang Zhao 0002, Chengxin Pang, Xing Hu 0006 |
ICASSP | 4 |
| 2023 | Learning Causality-inspired Representation Consistency for Video Anomaly DetectionabstractVideo anomaly detection is an essential yet challenging task in the multimedia community, with promising applications in smart cities and secure communities. Existing methods attempt to learn abstract representations of regular events with statistical dependence to model the endogenous normality, which discriminates anomalies by measuring the deviations to the learned distribution. However, conventional representation learning is only a crude description of video normality and lacks an exploration of its underlying causality. The learned statistical dependence is unreliable for diverse regular events in the real world and may cause high false alarms due to over generalization. Inspired by causal representation learning, we think that there exists a causal variable capable of adequately representing the general patterns of regular events in which anomalies will present significant variations. Therefore, we design a causality-inspired representation consistency (CRC) framework to implicitly learn the unobservable causal variables of normality directly from available normal videos and detect abnormal events with the learned representation consistency. Extensive experiments show that the causality-inspired normality is robust to regular events with label-independent shifts, and the proposed CRC framework can quickly and accurately detect various complicated anomalies from real-world surveillance videos. Yang Liu 0246, Zhaoyang Xia, Mengyang Zhao 0002, Donglai Wei 0002, Siao Liu, Bobo Ju, Gaoyun Fang, Jing Liu 0050 |
ACM Multimedia | 3 |
| 2023 | Memory-Augmented Spatial-Temporal Consistency Network for Video Anomaly Detection
Zhangxun Li, Mengyang Zhao 0002, Xinhua Zeng, Tian Wang 0002, Chengxin Pang |
PRCV (6) | 2 |
| 2023 | Stochastic video normality network for abnormal event detection in surveillance videos
Yang Liu 0246, Dingkang Yang, Gaoyun Fang, Donglai Wei 0002, Mengyang Zhao 0002, Kai Cheng 0001, Jing Liu 0050 |
Knowl. Based Syst. | 6 |
| 2023 | Two-Stage Alignments Framework for Unsupervised Domain Adaptation on Time Series DataabstractUnsupervised Domain Adaptation (UDA) aims to free models from labeled information of target domain by minimizing the discrepancy of distributions between different domains. Most existing methods are designed to learn domain-invariant features either by domain discrimination or by matching lower-order moments. However, these methods are not robust due to the limited representation of statistical characteristics for non-Gaussian distributions and thus fail in domain matching. In addition, they often focus on matching distributions while not considering class decision boundaries between domains. To address these issues, we propose a novel Two-Stage Alignments Framework (TSAF) for UAD, which not only performs arbitrary-order moment matching to approximately characterize complex non-Gaussian distributions, but also utilizes domain-specific decision boundaries to align the probabilistic outputs of classifiers. Moreover, the reconstruction-based task is introduced to enhance the representation of the inherent characteristics for specific distribution. Extensive experiments on three real-world time series datasets demonstrate that: 1) our model evidently outperforms many state-of-the-art domain adaptation methods in cross-domain classification tasks; 2) TSAF can learn domain-invariant features efficiently. Xiaowei Xiang, Yang Liu 0246, Gaoyun Fang, Jing Liu 0050, Mengyang Zhao 0002 |
IEEE Signal Process. Lett. | 5 |
| 2022 | Learning Appearance-Motion Normality for Video Anomaly DetectionabstractVideo anomaly detection is a challenging task in the Computer vision community. Most single task-based methods do not consider the independence of unique spatial and temporal patterns, while two-stream structures lack the exploration of the correlations. In this paper, we propose spatial-temporal memories augmented two-stream auto-encoder framework, which learns the appearance normality and motion normal-ity independently and explores the correlations via adversar-ial learning. Specifically, we first design two proxy tasks to train the two-stream structure to extract appearance and motion features in isolation. Then, the prototypical features are recorded in the corresponding spatial and temporal memory pools. Finally, the encoding-decoding network performs ad-versariallearning with the discriminator to explore the corre-lations between spatial and temporal patterns. Experimental results show that our framework outperforms the state-of-the-art methods, achieving AUCs of 98.1% and 89.8% on UCSD Ped2 and CUHK Avenue datasets. Yang Liu 0246, Jing Liu 0050, Mengyang Zhao 0002, Dingkang Yang, Xiaoguang Zhu |
ICME | 3 |
| 2022 | Exploiting Spatial-temporal Correlations for Video Anomaly DetectionabstractVideo anomaly detection (VAD) remains a challenging task in the pattern recognition community due to the ambiguity and diversity of abnormal events. Existing deep learning-based VAD methods usually leverage proxy tasks to learn the normal patterns and discriminate the instances that deviate from such patterns as abnormal. However, most of them do not take full advantage of spatial-temporal correlations among video frames, which is critical for understanding normal patterns. In this paper, we address unsupervised VAD by learning the evolution regularity of appearance and motion in the long and short-term and exploit the spatial-temporal correlations among consecutive frames in normal videos more adequately. Specifically, we proposed to utilize the spatiotemporal long short-term memory (ST-LSTM) to extract and memorize spatial appearances and temporal variations in a unified memory cell. In addition, inspired by the generative adversarial network, we introduce a discriminator to perform adversarial learning with the ST-LSTM to enhance the learning capability. Experimental results on standard benchmarks demonstrate the effectiveness of spatial-temporal correlations for unsupervised VAD. Our method achieves competitive performance compared to the state-of-the-art methods with AUCs of 96.7%, 87.8%, and 73.1% on the UCSD Ped2, CUHK Avenue, and ShanghaiTech, respectively. Mengyang Zhao 0002, Yang Liu 0246, Jing Liu 0050, Xinhua Zeng |
ICPR | 1 |
| 2022 | Multi-level Attention Fusion for Multimodal Driving Maneuver RecognitionabstractSensor-based driving maneuver recognition (DMR) is a fundamental and challenging task in ubiquitous computing, which uses multimodal signals from embedded sensors such as accelerometers and gyroscopes to recognize driving maneuvers. However, the spatial-temporal features from neural networks are often treated equally, which may limit the performance of the model in predicting maneuvers. In this paper, we propose a novel hybrid neural network model based on multi-level attention fusion for multimodal DMR. The proposed model utilizes convolutional neural networks and gated recurrent unit to extract temporal-spatial features from multimodal sensing signals and propose the multi-level attention fusion to explore the significant patterns over local and global periods. In addition, We design three different levels of fusion (early, late, and full fusion) to explore the effects of different attention fusions on the model. Extensive experiments on the real-world dataset show that the proposed model achieves superior performance to the baseline methods, and multi-level attention fusion brings 6.17% gain to the F1-score. Jing Liu 0050, Yang Liu 0246, Chengwen Tian, Mengyang Zhao 0002, Xinhua Zeng |
ISCAS | 4 |