VLDB 2026 Research / reviewers in the wild / expert
Mingfei Cheng
dblp:241/6053
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-8982-1483ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causality-Aware Safety Testing for Autonomous Driving SystemsabstractSimulation-based testing is essential for evaluating the safety of Autonomous Driving Systems (ADSs). Comprehensive evaluation requires testing across diverse scenarios that can trigger various types of violations under different conditions. While existing methods typically focus on individual diversity metrics, such as input scenarios, ADS-generated motion commands, and system violations, they often fail to capture the complex interrelationships among these elements. For instance, identical motion commands can produce different collision risks in varying scenes, and the same collision may result from different commands under different scenarios. This oversight leads to gaps in testing coverage, potentially missing critical issues in the ADS under evaluation. In this paper, we proposeCausal-Fuzzer, the first causality-aware fuzzing technique that enables efficient and comprehensive testing of ADSs by constructing causal graphs to model the interrelationships among scenarios, actions, and violations. Unlike existing methods that treat diversity metrics independently, we recognize these elements are causally interconnected and use their relationships to identify more diverse violations triggered by fundamentally different causal mechanisms. Specifically,Causal-Fuzzerproposes (1) a causality-based feedback mechanism that quantifies the combined diversity of test scenarios by assessing whether they activate new causal relationships, and (2) a causality-driven mutation strategy that prioritizes mutations on input scenario elements with higher causal impact on ego action changes and violation occurrence to enable interpretable and efficient test generation. We evaluatedCausal-Fuzzeron an industry-grade ADS Apollo, with a high-fidelity simulator LGSVL. Our empirical results demonstrate thatCausal-Fuzzersignificantly outperforms existing methods in (1) identifying a greater diversity of violations (96.5 violations on average, compared to 66.9 for the best baseline method), (2) providing enhanced testing sufficiency with improved coverage of causal relationships (13.6 unique sceneaction- violation patterns on average, compared to 8.6 for the best baseline method), and (3) achieving greater efficiency in detecting critical scenarios, strong robustness under noise conditions, and good generalizability across varying scenario complexities and violation types. Our source code and experimental results are available athttps://sites.google.com/view/causal-fuzzer. Wenbing Tang 0001, Mingfei Cheng, Yuan Zhou 0005, Yang Liu 0003, Zuohua Ding |
IEEE Trans. Software Eng. | 2 |
| 2025 | Decictor: Towards Evaluating the Robustness of Decision-Making in Autonomous Driving SystemsabstractAutonomous Driving System (ADS) testing is crucial in ADS development, with the current primary focus being on safety. However, the evaluation of non-safety-critical performance, particularly the ADS's ability to make optimal decisions and produce optimal paths for autonomous vehicles (AVs), is also vital to ensure the intelligence and reduce risks of AVs. Currently, there is little work dedicated to assessing the robustness of ADSs' path-planning decisions (PPDs), i.e., whether an ADS can maintain the optimal PPD after an insignificant change in the environment. The key challenges include the lack of clear oracles for assessing PPD optimality and the difficulty in searching for scenarios that lead to non-optimal PPDs. To fill this gap, in this paper, we focus on evaluating the robustness of ADSs' PPDs and propose the first method, Decictor, for generating nonoptimal decision scenarios (NoDSs), where the ADS does not plan optimal paths for AVs. Decictor comprises three main components: Non-invasive Mutation, Consistency Check, and Feedback. To overcome the oracle challenge, Non-invasive Mutation is devised to implement conservative modifications, ensuring the preservation of the original optimal path in the mutated scenarios. Subsequently, the Consistency Check is applied to determine the presence of nonoptimal PPDs by comparing the driving paths in the original and mutated scenarios. To deal with the challenge of large environment space, we design Feedback metrics that integrate spatial and temporal dimensions of the AV's movement. These metrics are crucial for effectively steering the generation of NoDSs. Therefore, Decictor can generate NoDSs by generating new scenarios and then identifying NoDSs in the new scenarios. We evaluate Decictor on Baidu Apollo, an open-source and production-grade ADS. The experimental results validate the effectiveness of Decictor in detecting non-optimal PPDs of ADSs. It generates 63.9 NoDSs in total, while the best-performing baseline only detects 35.4 NoDSs. Mingfei Cheng, Xiaofei Xie, Yuan Zhou 0005, Junjie Wang 0007, Guozhu Meng, Kairui Yang |
ICSE | 1 |
| 2025 | ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case PairsabstractAutomated Program Repair (APR) aims to automatically generate patches for rectifying software bugs. Recent strides in Large Language Models (LLM), such as ChatGPT, have yielded encouraging outcomes in APR, especially within the conversation-driven APR framework. Nevertheless, the efficacy of conversation-driven APR is contingent on the quality of the feedback information. In this article, we propose ContrastRepair , a novel conversation-based APR approach that augments conversation-driven APR by providing LLMs with contrastive test pairs. A test pair consists of a failing test and a passing test, which offer contrastive feedback to the LLM. Our key insight is to minimize the difference between the generated passing test and the given failing test, which can better isolate the root causes of bugs. By providing such informative feedback, ContrastRepair enables the LLM to produce effective bug fixes. The implementation of ContrastRepair is based on the state-of-the-art LLM, ChatGPT, and it iteratively interacts with ChatGPT until plausible patches are generated. We evaluate ContrastRepair on multiple benchmark datasets, including Defects4J, QuixBugs, and HumanEval-Java. The results demonstrate that ContrastRepair significantly outperforms existing methods, achieving a new state-of-the-art in program repair. For instance, among Defects4J 1.2 and 2.0, ContrastRepair correctly repairs 143 out of all 337 bug cases, while the best-performing baseline fixes 124 bugs. Jiaolong Kong, Xiaofei Xie, Mingfei Cheng, Shangqing Liu, Xiaoning Du 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2024 | CofiPara: A Coarse-to-fine Paradigm for Multimodal Sarcasm Target Identification with Large Multimodal ModelsabstractSocial media abounds with multimodal sarcasm, and identifying sarcasm targets is particularly challenging due to the implicit incongruity not directly evident in the text and image modalities. Current methods for Multimodal Sarcasm Target Identification (MSTI) predominantly focus on superficial indicators in an end-to-end manner, overlooking the nuanced understanding of multimodal sarcasm conveyed through both the text and image. This paper proposes a versatile MSTI framework with a coarse-to-fine paradigm, by augmenting sarcasm explainability with reasoning and pre-training knowledge. Inspired by the powerful capacity of Large Multimodal Models (LMMs) on multimodal reasoning, we first engage LMMs to generate competing rationales for coarser-grained pre-training of a small language model on multimodal sarcasm detection. We then propose fine-tuning the model for finer-grained sarcasm target identification. Our framework is thus empowered to adeptly unveil the intricate targets within multimodal sarcasm and mitigate the negative impact posed by potential noise inherently in LMMs. Experimental results demonstrate that our model far outperforms state-of-the-art MSTI methods, and markedly exhibits explainability in deciphering sarcasm as well. Zixin Chen, Hongzhan Lin 0001, Mingfei Cheng, Jing Ma 0004, Guang Chen 0003 |
ACL (1) | 4 |
| 2024 | Towards low-resource rumor detection: Unified contrastive transfer with propagation structure
Hongzhan Lin 0001, Jing Ma 0004, Ruichao Yang, Zhiwei Yang 0005, Mingfei Cheng |
Neurocomputing | 5 |
| 2024 | Historical Embedding-Guided Efficient Large-Scale Federated Graph LearningabstractGraph convolutional networks (GCNs) are promising for graph learning tasks. For privacy-preserving graph learning tasks involving distributed graph datasets, federated learning (FL)-based GCN (FedGCN) training is required. An important open challenge for FedGCN is scaling to large graphs, which typically incurs 1) high computation overhead for handling the explosively-increasing number of neighbors, and 2) high communication overhead of training GCNs involving multiple FL clients. Thus, neighbor sampling is being studied to enhance the scalability of FedGCNs. Existing FedGCN training techniques with neighbor sampling often produce extremely large communication and computation overhead and inaccurate node embeddings, leading to poor model performance. To bridge this gap, we propose the Federated Adaptive Attention-based Sampling (FedAAS) approach. It achieves substantial cost savings by efficiently leveraging historical embedding estimators and focusing the limited communication resources on transmitting the most influential neighbor node embeddings across FL clients. We further design an adaptive embedding synchronization scheme to optimize the efficiency and accuracy of FedAAS on large-scale datasets. Theoretical analysis shows that the approximation error induced by the staleness of historical embedding is upper bounded, and the model is guaranteed to converge in an efficient manner. Extensive experimental evaluation against four state-of-the-art baselines on six real-world graph datasets show that FedAAS achieves up to 5.12% higher test accuracy, while saving communication and computation costs by 95.11% and 94.76%, respectively. Anran Li 0001, Yuanyuan Chen 0012, Jian Zhang 0087, Mingfei Cheng, Yihao Huang 0001, Yueming Wu 0001, Anh Tuan Luu, Han Yu 0001 |
Proc. ACM Manag. Data | 4 |
| 2023 | BehAVExplor: Behavior Diversity Guided Testing for Autonomous Driving SystemsabstractTesting Autonomous Driving Systems (ADSs) is a critical task for ensuring the reliability and safety of autonomous vehicles. Existing methods mainly focus on searching for safety violations while the diversity of the generated test cases is ignored, which may generate many redundant test cases and failures. Such redundant failures can reduce testing performance and increase failure analysis costs. In this paper, we present a novel behavior-guided fuzzing technique (BehAVExplor) to explore the different behaviors of the ego vehi- cle (i.e., the vehicle controlled by the ADS under test) and detect diverse violations. Specifically, we design an efficient unsupervised model, called BehaviorMiner, to characterize the behavior of the ego vehicle. BehaviorMiner extracts the temporal features from the given scenarios and performs a clustering-based abstraction to group behaviors with similar features into abstract states. A new test case will be added to the seed corpus if it triggers new behav- iors (e.g., cover new abstract states). Due to the potential conflict between the behavior diversity and the general violation feedback, we further propose an energy mechanism to guide the seed selec- tion and the mutation. The energy of a seed quantifies how good it is. We evaluated BehAVExplor on Apollo, an industrial-level ADS, and LGSVL simulation environment. Empirical evaluation results show that BehAVExplor can effectively find more diverse violations than the state-of-the-art. Mingfei Cheng, Yuan Zhou 0005, Xiaofei Xie |
ISSTA | 1 |
| 2023 | Generative Model-Based Testing on Decision-Making PoliciesabstractThe reliability of decision-making policies is urgently important today as they have established the fundamentals of many critical applications, such as autonomous driving and robotics. To ensure reliability, there have been a number of research efforts on testing decision-making policies that solve Markov decision processes (MDPs). However, due to the deep neural network (DNN)-based inherit and infinite state space, developing scalable and effective testing frameworks for decision-making policies still remains open and challenging. In this paper, we present an effective testing framework for decision-making policies. The framework adopts a generative diffusion model-based test case generator that can easily adapt to different search spaces, ensuring the practicality and validity of test cases. Then, we propose a termination state novelty-based guidance to diversify agent behaviors and improve the test effectiveness. Finally, we evaluate the framework on five widely used benchmarks, including autonomous driving, aircraft collision avoidance, and gaming scenarios. The results demonstrate that our approach identifies more diverse and influential failure-triggering test cases compared to current state-of-the-art techniques. Moreover, we employ the detected failure cases to repair the evaluated models, achieving better robustness enhancement compared to the baseline method. Zhuo Li 0021, Xiongfei Wu, Derui Zhu, Mingfei Cheng, Fuyuan Zhang, Xiaofei Xie, Lei Ma 0003, Jianjun Zhao 0001 |
ASE | 4 |
| 2022 | A Joint Framework Towards Class-aware and Class-agnostic Alignment for Few-shot Segmentation
Mingfei Cheng, Bochen Wang, Ye Xi, Feigege Wang |
ACCV (7) | 2 |
| 2021 | Rumor Detection on Twitter with Claim-Guided Hierarchical Graph Attention NetworksabstractRumors are rampant in the era of social media.Conversation structures provide valuable clues to differentiate between real and fake claims.However, existing rumor detection methods are either limited to the strict relation of user responses or oversimplify the conversation structure.In this study, to substantially reinforces the interaction of user opinions while alleviating the negative impact imposed by irrelevant posts, we first represent the conversation thread as an undirected interaction graph.We then present a Claim-guided Hierarchical Graph Attention Network for rumor classification, which enhances the representation learning for responsive posts considering the entire social contexts and attends over the posts that can semantically infer the target claim.Extensive experiments on three Twitter datasets demonstrate that our rumor detection method achieves much better performance than stateof-the-art methods and exhibits a superior capacity for detecting rumors at early stages. Hongzhan Lin 0001, Jing Ma 0004, Mingfei Cheng, Zhiwei Yang 0005, Guang Chen 0003 |
EMNLP (1) | 3 |
| 2021 | Joint Topology-preserving and Feature-refinement Network for Curvilinear Structure SegmentationabstractCurvilinear structure segmentation (CSS) is under semantic segmentation, whose applications include crack detection, aerial road extraction, and biomedical image segmentation. In general, geometric topology and pixel-wise features are two critical aspects of CSS. However, most semantic segmentation methods only focus on enhancing feature representations while existing CSS techniques emphasize preserving topology alone. In this paper, we present a Joint Topology-preserving and Feature-refinement Network (JTFN) that jointly models global topology and refined features based on an iterative feedback learning strategy. Specifically, we explore the structure of objects to help preserve corresponding topologies of predicted masks, thus design a reciprocative two-stream module for CSS and boundary detection. In addition, we introduce such topology-aware predictions as feedback guidance that refines attentive features by supplementing and enhancing saliencies. To the best of our knowledge, this is the first work that jointly addresses topology preserving and feature refinement for CSS. We evaluate JTFN on four datasets of diverse applications: Crack500, CrackTree200, Roads, and DRIVE. Results show that JTFN performs best in comparison with alternative methods. Code is available.1 Mingfei Cheng, Kaili Zhao, Xuhong Guo, Jun Guo 0002 |
ICCV | 1 |
| 2019 | Cyclone Intensity Estimate with Context-Aware CycleganabstractDeep learning approaches to cyclone intensity estimation have recently shown promising results. However, suffering from the extreme scarcity of cyclone data on specific intensity, most existing deep learning methods fail to achieve satisfactory performance on cyclone intensity estimation, especially on classes with few instances. To avoid the degradation of recognition performance caused by scarce samples, we propose a context-aware CycleGAN which learns the latent evolution features from adjacent cyclone intensity and synthesizes CNN features of classes lacking samples from unpaired source classes. Specifically, our approach synthesizes features conditioned on the learned evolution features, while the extra information is not required. Experimental results of several evaluation methods show the effectiveness of our approach, even can predicting unseen classes. Haitao Yang 0008, Mingfei Cheng, Si Li 0001 |
ICIP | 3 |