VLDB 2026 Research / reviewers in the wild / expert
Mengxi Jia
dblp:268/7984
· DBLP profile ↗
20ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-0979-9803ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NER-AD: Noise-Robust Reconstruction Enhanced by Representation-Learning for Metric Anomaly Detection in Online Service Systems
Xiaosong Huang, Mengxi Jia, Zhonghai Wu, Ying Li 0012, Yu-an Tan 0001, Liehuang Zhu, Wanlei Zhou 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV SystemsabstractUnmanned Aerial Vehicles (UAVs) are widely adopted across various fields, yet they raise significant privacy and safety concerns, demanding robust monitoring solutions. Existing anti-UAV methods primarily focus on position tracking but fail to capture UAV behavior and intent. To address this, we introduce a novel task—UAV Tracking and Intent Understanding (UTIU)—which aims to track UAVs while inferring and describing their motion states and intent for a more comprehensive monitoring approach. To tackle the task, we propose JTD-UAV, the first joint tracking, and intent description framework based on large language models. Our dual-branch architecture integrates UAV tracking with Visual Question Answering (VQA), allowing simultaneous localization and behavior description. To benchmark this task, we introduce the TDUAV dataset, the largest dataset for joint UAV tracking and intent understanding, featuring 1,328 challenging video sequences, over 163K annotated thermal frames, and 3K VQA pairs. Our benchmark demonstrates the effectiveness of JTD-UAV. Jian Zhao 0006, Zhaoxin Fan, Xin Zhang 0093, Yudian Zhang, Lei Jin 0003, Gang Wang 0031, Mengxi Jia, Xuelong Li 0001 |
CVPR | 10 |
| 2025 | ScalaLog: Scalable Log-Based Failure Diagnosis Using LLMabstractAs Industrial Internet of Things (IIoT) software systems become increasingly complex, precise failure diagnosis has become both essential and challenging. Current log-based failure diagnosis methods lack scalability for different failure types. In IIoT software systems, the number of failure types is constantly growing, and retraining the model each time a new failure type is introduced is highly resource-intensive. Additionally, traditional log-based failure diagnosis models often require log parsing as a preliminary step, which can also be resource-consuming. To address these challenges, we propose a scalable log-based failure diagnosis method named ScalaLog. ScalaLog builds on RAG by utilizing LLM-based summarization to extract key log information, applying sample augmentation to increase the number of samples, and using CoT prompts to guide the LLM in failure diagnosis. Experiments on various public and real-world datasets demonstrate that ScalaLog significantly enhances failure diagnosis accuracy without the need for training or log parsing. Lingzhe Zhang, Mengxi Jia, Yifan Wu 0002, Ying Li 0012 |
ICASSP | 3 |
| 2025 | AAAD: Asynchronous Inter-Variable Relationship-Aware Anomaly Detection for Multivariate Time SeriesabstractAnomaly detection in multivariate time series (MTS) plays a crucial role in various domains, particularly in multimedia. While significant progress has been made in modeling normal data patterns and detecting anomalies based on deviations, existing methods face challenges in capturing complex inter-variable relationships, particularly in the presence of asynchronous dependencies. To address the challenges of detecting anomalies in MTS with complex inter-variable relationships, we propose an asynchronous inter-variable relationship-aware anomaly detection method. This approach simultaneously extracts both synchronous and asynchronous feature pairs between variables and leverages attention mechanisms to automatically learn unified inter-variable relationships across these dependencies. Additionally, we utilize a memory network to store and dynamically update the normal patterns of inter-variable relationships, computing anomaly scores based on deviations in these relationships. Extensive experiments demonstrate that our method outperforms existing baseline approaches, achieving an average F1 score of 97.05% across five benchmark datasets. Ablation studies further validate the effectiveness of each component of our method. Xiaosong Huang, Mengxi Jia, Lingzhe Zhang, Zhonghai Wu, Ying Li 0012 |
ICME | 3 |
| 2025 | EraseAnything: Enabling Concept Erasure in Rectified Flow TransformersabstractRemoving unwanted concepts from large-scale text-to-image (T2I) diffusion models while maintaining their overall generative quality remains an open challenge. This difficulty is especially pronounced in emerging paradigms, such as Stable Diffusion (SD) v3 and Flux, which incorporate flow matching and transformer-based architectures. These advancements limit the transferability of existing concept-erasure techniques that were originally designed for the previous T2I paradigm (e.g., SD v1.4). In this work, we introduce EraseAnything, the first method specifically developed to address concept erasure within the latest flow-based T2I framework. We formulate concept erasure as a bi-level optimization problem, employing LoRA-based parameter tuning and an attention map regularizer to selectively suppress undesirable activations. Furthermore, we propose a self-contrastive learning strategy to ensure that removing unwanted concepts does not inadvertently harm performance on unrelated ones. Experimental results demonstrate that EraseAnything successfully fills the research gap left by earlier methods in this new T2I paradigm, achieving state-of-the-art performance across a wide range of concept erasure tasks. Daiheng Gao, Shilin Lu, Wenbo Zhou 0004, Jiaming Chu, Jie Zhang 0073, Mengxi Jia, Bang Zhang, Zhaoxin Fan, Weiming Zhang 0001 |
ICML | 6 |
| 2025 | SAFA: Lifelong Person Re-Identification learning by statistics-aware feature alignment
Qiankun Gao, Mengxi Jia, Jie Chen 0001, Jian Zhang 0018 |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Towards Close-to-Zero Runtime Collection Overhead: Raft-Based Anomaly Diagnosis on System Faults for Distributed Storage SystemabstractDistributed storage systems are fundamental infrastructures of today’s large-scale software systems such as cloud systems. Diagnosing anomalies in distributed storage systems is essential for maintaining software availability. Existing anomaly diagnosis approaches mainly rely on the run-time data including monitoring data and application logs. However, collecting and analyzing the run-time data requires huge computing, storage, and management costs. Typically, more fine-grained run-time data can reveal more symptoms of anomalies, but on the contrary, requires more computing, storage, and management costs. As a result, solving the anomaly diagnosis problem is a balancing between the quality of run-time data and system overhead or cost. In this paper, we take into account both data quality and system overhead or cost by introducing a new type of run-time data-Raft logs. Raft logs are naturally produced by distributed storage systems and collecting raft logs will not bring any extra system overhead. To verify the ability of Raft logs in reflecting anomalies, we conduct a comprehensive study on the interconnection between the anomalies and Raft logs. Based on the study, we propose an effectiveRaft-BasedAnomalyDiagnosis approach namedRBAD. For evaluation, we expose the first open-sourced comprehensive dataset with multiple runtime data containing both Raft logs, application logs and monitoring data. Experiments based on this dataset demonstrate RBAD’s superiority, outperforming monitoring-based methods by 15.38% and log-based methods by 53.10%. Lingzhe Zhang, Mengxi Jia, Yong Yang 0011, Zhonghai Wu, Ying Li 0012 |
IEEE Trans. Serv. Comput. | 3 |
| 2025 | E-Log: Fine-Grained Elastic Log-Based Anomaly Detection and Diagnosis for DatabasesabstractDatabase Management Systems (DBMS) form the backbone of modern large-scale software systems, where reliable anomaly detection and diagnosis are essential for ensuring system availability. However, existing log-based methods often impose significant performance overhead by collecting large volumes of logs, which is impractical for DBMS requiring high read/write throughput. This paper addresses a critical yet underexplored challenge: how to balance logging granularity with runtime efficiency for effective anomaly management in databases. We presentE-Log, a novel fine-grained elastic log-based framework for anomaly detection and diagnosis. E-Log intelligently adjusts the amount and detail of logging based on system state—maintaining lightweight logging during normal operation for efficient anomaly detection, and triggering rich, informative logging only upon anomaly suspicion for accurate diagnosis. This adaptive strategy significantly reduces runtime overhead while preserving diagnostic precision. We implement E-Log on Apache IoTDB and evaluate it using benchmarks including TSBS, TPCx-IoT, and IoT-Bench. Experimental results show that E-Log improves anomaly detection accuracy by 3.15% and diagnosis performance by 9.32% compared to state-of-the-art methods. Moreover, it reduces log storage size by 43.53% and increases average write throughput by 26.22%. These results highlight E-Log's potential to enable efficient, accurate, and scalable anomaly management in high-performance database systems. Lingzhe Zhang, Mengxi Jia, Zhonghai Wu, Ying Li 0012 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Reducing Events to Augment Log-based Anomaly Detection Models: An Empirical StudyabstractAs software systems grow increasingly intricate, the precise detection of anomalies have become both essential and challenging. Current log-based anomaly detection methods depend heavily on vast amounts of log data leading to inefficient inference and potential misguidance by noise logs. However, the quantitative effects of log reduction on the effectiveness of anomaly detection remain unexplored. Therefore, we first conduct a comprehensive study on six distinct models spanning three datasets. Through the study, the impact of log quantity and their effectiveness in representing anomalies is qualifies, uncovering three distinctive log event types that differently influence model performance. Drawing from these insights, we propose LogCleaner: an efficient methodology for the automatic reduction of log events in the context of anomaly detection. Serving as middleware between software systems and models, LogCleaner continuously updates and filters anti-events and duplicative-events in the raw generated logs. Experimental outcomes highlight LogCleaner’s capability to reduce over 70% of log events in anomaly detection, accelerating the model’s inference speed by approximately 300%, and universally improving the performance of models for anomaly detection. Lingzhe Zhang, Kangjin Wang, Mengxi Jia, Yong Yang 0011, Ying Li 0012 |
ESEM | 4 |
| 2024 | Multivariate Log-based Anomaly Detection for Distributed DatabaseabstractDistributed databases are fundamental infrastructures of today's large-scale software systems such as cloud systems. Detecting anomalies in distributed databases is essential for maintaining software availability. Existing approaches, predominantly developed using Loghub-a comprehensive collection of log datasets from various systems-lack datasets specifically tailored to distributed databases, which exhibit unique anomalies. Additionally, there's a notable absence of datasets encompassing multi-anomaly, multi-node logs. Consequently, models built upon these datasets, primarily designed for standalone systems, are inadequate for distributed databases, and the prevalent method of deeming an entire cluster anomalous based on irregularities in a single node leads to a high false-positive rate. This paper addresses the unique anomalies and multivariate nature of logs in distributed databases. We expose the first open-sourced, comprehensive dataset with multivariate logs from distributed databases. Utilizing this dataset, we conduct an extensive study to identify multiple database anomalies and to assess the effectiveness of state-of-the-art anomaly detection using multivariate log data. Our findings reveal that relying solely on logs from a single node is insufficient for accurate anomaly detection on distributed database. Leveraging these insights, we propose MultiLog, an innovative multivariate log-based anomaly detection approach tailored for distributed databases. Our experiments, based on this novel dataset, demonstrate MultiLog's superiority, outperforming existing state-of-the-art methods by approximately 12%. Lingzhe Zhang, Mengxi Jia, Ying Li 0012, Yong Yang 0011, Zhonghai Wu |
KDD | 3 |
| 2024 | UAC-AD: Unsupervised Adversarial Contrastive Learning for Anomaly Detection on Multi-Modal Data in Microservice SystemsabstractTo ensure the stability and reliability of microservice systems, timely and accurate anomaly detection is of utmost importance. Recently, considering the lack of labels in real-world scenarios and the collaborative and complementary relationships of multi-modal data in reflecting system anomalies, unsupervised multi-modal anomaly methods have been proposed. However, existing methods face challenges in effectively distinguishing normal hard samples (they are normal but hard to classify correctly) from anomalies. This is mainly caused by two aspects. First, the hard sample patterns are complex. Second, the convergence speed is inconsistent between hard and simple samples. To overcome these issues, we propose an unsupervised adversarial contrastive multi-modal anomaly detection method (UAC-AD). We utilize contrastive learning to help learn the complex patterns of hard samples and enlarge the distance between hard and anomaly samples. Meanwhile, the adversarial framework automatically identifies hard samples and fine-grained adjusts the training weights to each modality part of these hard samples. In this case, The hard sample problems of two aspects can be alleviated. We extensively evaluate UAC-AD on two open-source simulated datasets and a real industrial dataset from a large communication company. Extensive experimental results demonstrate the effectiveness of our approach in anomaly detection. We also release the code and dataset for replication and future research. Xiaosong Huang, Mengxi Jia, Zhonghai Wu, Ying Li 0012 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Semi-attention Partition for Occluded Person Re-identificationabstractThis paper proposes a Semi-Attention Partition (SAP) method to learn well-aligned part features for occluded person re-identification (re-ID). Currently, the mainstream methods employ either external semantic partition or attention-based partition, and the latter manner is usually better than the former one. Under this background, this paper explores a potential that the weak semantic partition can be a good teacher for the strong attention-based partition. In other words, the attention-based student can substantially surpass its noisy semantic-based teacher, contradicting the common sense that the student usually achieves inferior (or comparable) accuracy. A key to this effect is: the proposed SAP encourages the attention-based partition of the (transformer) student to be partially consistent with the semantic-based teacher partition through knowledge distillation, yielding the so-called semi-attention. Such partial consistency allows the student to have both consistency and reasonable conflict with the noisy teacher. More specifically, on the one hand, the attention is guided by the semantic partition from the teacher. On the other hand, the attention mechanism itself still has some degree of freedom to comply with the inherent similarity between different patches, thus gaining resistance against noisy supervision. Moreover, we integrate a battery of well-engineered designs into SAP to reinforce their cooperation (e.g., multiple forms of teacher-student consistency), as well as to promote reasonable conflict (e.g., mutual absorbing partition refinement and a supervision signal dropout strategy). Experimental results confirm that the transformer student achieves substantial improvement after this semi-attention learning scheme, and produces new state-of-the-art accuracy on several standard re-ID benchmarks. Mengxi Jia, Yifan Sun 0003, Yunpeng Zhai, Xinhua Cheng, Yi Yang 0001, Ying Li 0012 |
AAAI | 1 |
| 2023 | Panoptic Compositional Feature Field for Editable Scene Rendering with Network-Inferred Labels via Metric LearningabstractDespite neural implicit representations demonstrating impressive high-quality view synthesis capacity, decom-posing such representations into objects for instance-level editing is still challenging. Recent works learn object-compositional representations supervised by ground truth instance annotations and produce promising scene editing results. However, ground truth annotations are manually labeled and expensive in practice, which limits their usage in real-world scenes. In this work, we attempt to learn an object-compositional neural implicit representation for editable scene rendering by leveraging labels inferred from the off-the-shelf 2D panoptic segmentation networks instead of the ground truth annotations. We propose a novel framework named Panoptic Compositional Feature Field (PCFF), which introduces an instance quadruplet metric learning to build a discriminating panoptic feature space for reliable scene editing. In addition, we propose semantic-related strategies to further exploit the correlations between semantic and appearance attributes for achieving better rendering results. Experiments on multiple scene datasets including ScanNet, Replica, and ToyDesk demonstrate that our proposed method achieves superior performance for novel view synthesis and produces convincing real-world scene editing results. Xinhua Cheng, Yanmin Wu, Mengxi Jia, Jian Zhang 0018 |
CVPR | 3 |
| 2023 | Population-Based Evolutionary Gaming for Unsupervised Person Re-identification
Yunpeng Zhai, Peixi Peng, Mengxi Jia, Xuesong Gao, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | Learning Disentangled Representation Implicitly Via Transformer for Occluded Person Re-IdentificationabstractPerson re-IDentification (re-ID) under various occlusions has been a long-standing challenge as person images with different types of occlusions often suffer from misalignment in image matching and ranking. Most existing methods tackle this challenge by aligning spatial features of body parts according to external semantic cues or feature similarities but this alignment approach is complicated and sensitive to noises. We design DRL-Net, a disentangled representation learning network that handles occluded re-ID without requiring strict person image alignment or any additional supervision. Leveraging transformer architectures, DRL-Net achieves alignment-free re-ID via global reasoning of local features of occluded person images. It measures image similarity by automatically disentangling the representation of undefined semantic components, e.g., human body parts or obstacles, under the guidance of semantic preference object queries in the transformer. In addition, we design a decorrelation constraint in the transformer decoder and impose it over object queries for better focus on different semantic components. To better eliminate interference from occlusions, we design a contrast feature learning technique (CFL) for better separation of occlusion features and discriminative ID features. Extensive experiments over occluded and holistic re-ID benchmarks show that the DRL-Net achieves superior re-ID performance consistently and outperforms the state-offi-the-art by large margins for occluded re-ID dataset. Mengxi Jia, Xinhua Cheng, Shijian Lu, Jian Zhang 0018 |
IEEE Trans. Multim. | 1 |
| 2022 | More is better: Multi-source Dynamic Parsing Attention for Occluded Person Re-identificationabstractOccluded person re-identification (re-ID) has been a long-standing challenge in surveillance systems. Most existing methods tackle this challenge by aligning spatial features of human parts according to external semantic cues, which are inferred from the off-the-shelf semantic models (e.g. human parsing and pose estimation). However, there is a significant domain gap between the images in re-ID datasets and the images used for training the semantic models, such that inevitably making those semantic cues unreliable and deteriorating the re-ID performance. Multi-source knowledge ensemble has been proved to be effective for domain adaptation. Inspired by this, we propose a multi-source dynamic parsing attention (MSDPA) mechanism that leverages knowledge learned from different source datasets to generate reliable semantic cues and dynamically integrate and adapt them in a self-supervised manner by attention mechanism. Specifically, we first design a parsing embedding module (PEM) to integrate and embed the multi-source semantic cues into the patch tokens through a voting procedure. To further exploit correlations among body parts with similar semantics, we design a dynamic parsing attention block (DPAB) to guide the patch sequences aggregation by prior attentions which are dynamically generated from human parsing results. Extensive experiments over occluded, partial, and holistic re-ID datasets show that the MSDPA achieves superior re-ID performance consistently and outperforms the state-of-the-art methods by large margins on occluded datasets. Xinhua Cheng, Mengxi Jia, Jian Zhang 0018 |
ACM Multimedia | 2 |
| 2022 | A Simple Visual-Textual Baseline for Pedestrian Attribute RecognitionabstractPedestrian attribute recognition (PAR), which aims to identify attributes of the pedestrians captured in video surveillance, is a challenging task due to the poor quality of images and diverse spatial distribution among attributes. Existing methods usually model PAR as a multi-label classification problem and manually map attributes to an ordered list corresponding to the outputs of classifiers or sequential models. However, the inherent textual information among attribute annotations is largely neglected in these visual-only methods. In this paper, we first alleviate this issue by proposing a novel visual-textual baseline (VTB) for PAR which introduces an additional textual modality to explore the textual semantic correlations from attribute annotations by pre-trained textual encoders instead of human definitions. VTB encodes pedestrian images and attribute annotations into visual and textual features respectively, interacts with information across modalities, and predicts recognition results independently to remove the influence of attribute orders. Furthermore, we introduce transformer encoder as the cross-modal fusion module in VTB for sufficient intra-modal and cross-modal correlations exploration. Our method achieves superior performance over most existing visual-only methods on two widely used datasets including RAP and PA-100K, demonstrating the effectiveness of utilizing textual modality to PAR. Our method is expected to serve as a multimodal PAR baseline and inspire new insights for multimodal fusion in future PAR research. Our code is available athttps://github.com/cxh0519/VTB. Xinhua Cheng, Mengxi Jia, Jian Zhang 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Matching on Sets: Conquer Occluded Person Re-identification Without AlignmentabstractOccluded person re-identification (re-ID) is a challenging task as different human parts may become invisible in cluttered scenes, making it hard to match person images of different identities. Most existing methods address this challenge by aligning spatial features of body parts according to semantic information (e.g. human poses) or feature similarities but this approach is complicated and sensitive to noises. This paper presents Matching on Sets (MoS), a novel method that positions occluded person re-ID as a set matching task without requiring spatial alignment. MoS encodes a person image by a pattern set as represented by a `global vector’ with each element capturing one specific visual pattern, and it introduces Jaccard distance as a metric to compute the distance between pattern sets and measure image similarity. To enable Jaccard distance over continuous real numbers, we employ minimization and maximization to approximate the operations of intersection and union, respectively. In addition, we design a Jaccard triplet loss that enhances the pattern discrimination and allows to embed set matching into deep neural networks for end-to-end training. In the inference stage, we introduce a conflict penalty mechanism that detects mutually exclusive patterns in the pattern union of image pairs and decreases their similarities accordingly. Extensive experiments over three widely used datasets (Market1501, DukeMTMC and Occluded-DukeMTMC) show that MoS achieves superior re-ID performance. Additionally, it is tolerant of occlusions and outperforms the state-of-the-art by large margins for Occluded-DukeMTMC. Mengxi Jia, Xinhua Cheng, Yunpeng Zhai, Shijian Lu, Siwei Ma 0001, Yonghong Tian 0001, Jian Zhang 0018 |
AAAI | 1 |
| 2020 | Multiple Expert Brainstorming for Domain Adaptive Person Re-Identification
Yunpeng Zhai, Qixiang Ye, Shijian Lu, Mengxi Jia, Rongrong Ji, Yonghong Tian 0001 |
ECCV (7) | 4 |
| 2020 | A Similarity Inference Metric for RGB-Infrared Cross-Modality Person Re-identificationabstractRGB-Infrared (IR) cross-modality person re-identification (re-ID), which aims to search an IR image in RGB gallery or vice versa, is a challenging task due to the large discrepancy between IR and RGB modalities. Existing methods address this challenge typically by aligning feature distributions or image styles across modalities, whereas the very useful similarities among gallery samples of the same modality (i.e. intra-modality sample similarities) are largely neglected. This paper presents a novel similarity inference metric (SIM) that exploits the intra-modality sample similarities to circumvent the cross-modality discrepancy targeting optimal cross-modality image matching. SIM works by successive similarity graph reasoning and mutual nearest-neighbor reasoning that mine cross-modality sample similarities by leveraging intra-modality sample similarities from two different perspectives. Extensive experiments over two cross-modality re-ID datasets (SYSU-MM01 and RegDB) show that SIM achieves significant accuracy improvement but with little extra training as compared with the state-of-the-art. Mengxi Jia, Yunpeng Zhai, Shijian Lu, Siwei Ma 0001, Jian Zhang 0018 |
IJCAI | 1 |