VLDB 2026 Research / reviewers in the wild / expert
Xing Cui
dblp:210/4637
· DBLP profile ↗
14ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VFCionX: Bridging Large and Small Models for Robust Vulnerability-Fixing Commit IdentificationabstractVulnerability-Fixing Commit Identification(VFCI) is a critical task in software security maintenance that aims to automatically identify code commits that patch security vulnerabilities. However, existing approaches face challenges in handling low-quality commit messages and entangled commits, which limit their identification performance. To address these issues, we propose VFCionX, a novel VFCI framework that integrates large and small language models in a collaborative architecture. VFCionX consists of three core modules: Message Classifier, Patch Classifier, and Ensemble Classifier. The Message Classifier employs a multi-source contextual augmentation strategy to enhance the quality of commit messages and fine-tunes the Qwen2.5-1.5B model, significantly improving classification performance in the textual modality. The Patch Classifier combines heuristic rules with a Qwen2.5-Coder-7B-driven file selector to filter noise from entangled commits, and incorporates a line-level feature extractor based on CodeBERT and CNN to capture local pattern differences between added and deleted code lines. The Ensemble Classifier integrates predictions from both channels using the AdaBoost algorithm, enhancing model robustness and generalization. Experimental results on five popular C/C++ repositories comprising 24,630 commits show that VFCionX achieves an F1-score of 81.47%, outperforming the best baseline by 9.42%. Ablation studies validate the effectiveness of each component, while sensitivity analysis reveals optimal parameter settings for balancing performance and noise resilience. This work provides a new and effective solution for robust vulnerability patch identification. Xing Cui, JingZheng Wu, Wenxiang Ou, Tianyue Luo, Xiang Ling 0001 |
AAAI | 1 |
| 2026 | T2Agent: A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree SearchabstractReal-world multimodal misinformation often arises from mixed forgery sources, requiring dynamic reasoning and adaptive verification. However, existing methods mainly rely on static pipelines and limited tool usage, limiting their ability to handle such complexity and diversity. To address this challenge, we propose T2Agent, a novel misinformation detection agent that incorporates an extensible toolkit with Monte Carlo Tree Search (MCTS). The toolkit consists of modular tools such as web search, forgery detection, and consistency analysis. Each tool is described using standardized templates, enabling seamless integration and future expansion. To avoid inefficiency from using all tools simultaneously, a greedy search-based selector is proposed to identify a task-relevant subset. This subset then serves as the action space for MCTS to dynamically collect evidence and perform multi-source verification. To better align MCTS with the multi-source nature of misinformation detection, T2Agent extends traditional MCTS with multi-source verification, which decomposes the task into coordinated subtasks targeting different forgery sources. A dual reward mechanism containing a reasoning trajectory score and a confidence score is further proposed to encourage a balance between exploration across mixed forgery sources and exploitation for more reliable evidence. We conduct ablation studies to confirm the effectiveness of the tree search mechanism and tool usage. Extensive experiments further show that T2Agent consistently outperforms existing baselines on challenging mixed-source multimodal misinformation benchmarks, demonstrating its strong potential as a training-free detector. Xing Cui, Yueying Zou, Zekun Li 0001, Peipei Li 0002, Xuannan Liu, Huaibo Huang |
AAAI | 1 |
| 2025 | We Know What You're Looking For: Recommendation for Large-Scale Open Source SoftwareabstractBackground: In recent years, with the advancement of software engineering technologies and industry, Open Source Software (OSS) has become a mainstream model for software development and innovation. Increasingly, organizations and developers are adopting and customizing existing OSS to simplify and accelerate development processes. During OSS adoption, recommending suitable software based on user needs is crucial for enhancing development efficiency and addressing diverse requirements. However, the vast number and diversity of OSS make the recommendation task highly challenging. Despite progress in previous research, several issues remain, such as neglect of key software attributes, complexity in extracting multilingual features, and challenges of cold start and data sparsity. Aims: This paper presents AthenaRec, a large-scale OSS recommendation system comprising three core modules: Delphi, Argus, and Hestia. AthenaRec aims to recommend relevant and suitable software from a vast OSS based on user needs. Method: Specifically, Delphi first analyzes user queries to identify intention; Argus employs a heterogeneous ensemble recall approach to retrieve a large set of candidate software relevant to the identified intention; finally, Hestia adopts a two-stage deep ranking strategy. It performs coarse ranking by integrating multilingual modeling with contrastive learning, followed by fine ranking with a large language model, augmented by retrieval-augmented generation to incorporate external evidence. To evaluate the effectiveness of AthenaRec, we use a query dataset from real application scenarios. Results: Experimental results demonstrate that, on the test set of 7,500 queries, AthenaRec achieves superior recommendation performance, with Hits@20, MAP@20, NDCG@20, and MRR scores of$98.27 \%, 95.60 \%, 95.05 {\%}$, and 92.92%, respectively. On average, AthenaRec outperforms other top methods by 10.9% across all evaluation metrics. Conclusions: Additionally, we develop a Visual Studio Code (VSCode) plugin based on AthenaRec, which can be accessed via URL. We intend for this research to provide a reference for software developers, advancing the efficiency and accuracy of OSS recommendation. Xing Cui, JingZheng Wu, Xiang Ling 0001, Tianyue Luo |
ESEM | 1 |
| 2025 | MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMsabstractCurrent multimodal misinformation detection (MMD) methods often assume a single source and type of forgery for each sample, which is insufficient for real-world scenarios where multiple forgery sources coexist. The lack of a benchmark for mixed-source misinformation has hindered progress in this field. To address this, we introduce MMFakeBench, the first comprehensive benchmark for mixed-source MMD. MMFakeBench includes 3 critical sources: textual veracity distortion, visual veracity distortion, and cross-modal consistency distortion, along with 12 sub-categories of misinformation forgery types. We further conduct an extensive evaluation of 6 prevalent detection methods and 15 Large Vision-Language Models (LVLMs) on MMFakeBench under a zero-shot setting. The results indicate that current methods struggle under this challenging and realistic mixed-source MMD setting. Additionally, we propose MMD-Agent, a novel approach to integrate the reasoning, action, and tool-use capabilities of LVLM agents, significantly enhancing accuracy and generalization. We believe this study will catalyze future research into more realistic mixed-source multimodal misinformation and provide a fair evaluation of misinformation detection methods. Xuannan Liu, Zekun Li 0001, Peipei Li 0002, Huaibo Huang, Shuhan Xia, Xing Cui, Linzhi Huang, Weihong Deng, Zhaofeng He 0001 |
ICLR | 6 |
| 2025 | RMGenie: An LLM-Based Agent Framework for Open Source Software README GenerationabstractOpen Source Software (OSS) plays a vital role in modern software ecosystems, with README files providing essential information on functionality, configuration, and usage. However, approximately 9.03% of OSS projects lack adequate README documentation, impacting developer efficiency and software maintainability. Recent advances in natural language processing and large language models (LLMs) have automated README generation to reduce manual effort. Yet, existing methods struggle with integrating real-time external knowledge, parsing complex software structures, and adapting to projectspecific requirements. To address these challenges, this paper introduces RMGenie, a framework that leverages LLM-based agents and external tool invocation for automated README generation. RMGenie constructs an agent-driven workflow enabling multi-round interactions with external tools to dynamically extract critical code insights, particularly for complex software structures. It employs a Tree of Actions model with an entropy-based scoring mechanism to optimize decision paths for README generation. Additionally, a reflexion Mechanism enhances accuracy by mitigating decision biases and tool invocation errors. Experimental results confirm that RMGenie significantly outperforms baseline methods in content completeness, instruction adherence, and factual accuracy. Furthermore, we develop a plugin that integrates RMGenie to automatically generate README files from GitHub repositories, improving developer productivity and standardizing OSS documentation. The plugin is available at VSCode Marketplace. Xing Cui, JingZheng Wu, Tianyue Luo, Xiang Ling 0001 |
ICSME | 1 |
| 2025 | SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology AnalysisabstractWith the increasing integration of Multimodal Large Language Models (MLLMs) into the medical field, comprehensive evaluation of their performance in various medical domains becomes critical. However, existing benchmarks primarily assess general medical tasks, inadequately capturing performance in nuanced areas like the spine, which relies heavily on visual input. To address this, we introduce SpineBench, a comprehensive Visual Question Answering (VQA) benchmark designed for fine-grained analysis and evaluation of MLLMs in the spinal domain. SpineBench comprises 64,878 QA pairs from 40,263 spine images, covering 11 spinal diseases through two critical clinical tasks: spinal disease diagnosis and spinal lesion localization, both in multiple-choice format. SpineBench is built by integrating and standardizing image-label pairs from open-source spinal disease datasets, and samples challenging hard negative options for each VQA pair based on visual similarity (similar but not the same disease), simulating real-world challenging scenarios. We evaluate 12 leading MLLMs on SpineBench. The results reveal that these models exhibit poor performance in spinal tasks, highlighting limitations of current MLLM in the spine domain and guiding future improvements in spinal medicine applications. SpineBench is publicly available at https://zhangchenghanyu.github.io/SpineBench.github.io/. Chenghanyu Zhang, Zekun Li 0001, Peipei Li 0002, Xing Cui, Shuhan Xia, Weixiang Yan, Yiqiao Zhang, Qianyu Zhuang |
ACM Multimedia | 4 |
| 2025 | Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMsabstractThe increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily focus on model vulnerabilities exposed by static image inputs, ignoring the temporal dynamics of video that may induce distinct safety risks. To bridge this gap, we introduce Video-SafetyBench, the first comprehensive benchmark designed to evaluate the safety of LVLMs under video-text attacks. It comprises 2,264 video-text pairs spanning 48 fine-grained unsafe categories, each pairing a synthesized video with either a harmful query, which contains explicit malice, or a benign query, which appears harmless but triggers harmful behavior when interpreted alongside the video. To generate semantically accurate videos for safety evaluation, we design a controllable pipeline that decomposes video semantics into subject images (what is shown) and motion text (how it moves), which jointly guide the synthesis of query-relevant videos. To effectively evaluate uncertain or borderline harmful outputs, we propose RJScore, a novel LLM-based metric that incorporates the confidence of judge models and human-aligned decision threshold calibration. Extensive experiments show that benign-query video composition achieves average attack success rates of 67.2%, revealing consistent vulnerabilities to video-induced attacks. We believe Video-SafetyBench will catalyze future research into video-based safety evaluation and defense strategies. Xuannan Liu, Zekun Li 0001, Zheqi He, Peipei Li 0002, Shuhan Xia, Xing Cui, Huaibo Huang, Xi Yang 0023, Ran He 0001 |
NeurIPS | 6 |
| 2025 | AdvCloak: Customized adversarial cloak for privacy protection
Xuannan Liu, Yaoyao Zhong, Xing Cui, Yuhang Zhang 0016, Peipei Li 0002, Weihong Deng |
Pattern Recognit. | 3 |
| 2024 | INSTASTYLE: Inversion Noise of a Stylized Image is Secretly a Style Adviser
Xing Cui, Zekun Li 0001, Peipei Li 0002, Huaibo Huang, Xuannan Liu, Zhaofeng He 0001 |
ECCV (51) | 1 |
| 2024 | Exploring 3D-aware Lifespan Face Aging via Disentangled Shape-Texture RepresentationsabstractExisting face aging methods often focus on modeling either texture aging or using an entangled shape-texture representation to achieve face aging. However, shape and texture are two distinct factors that mutually affect the human face aging process. In this paper, we propose 3D-STD, a novel 3D-aware Shape-Texture Disentangled face aging network that explicitly disentangles the facial image into shape and texture representations using 3D face reconstruction. Additionally, to facilitate high-fidelity texture synthesis, we propose a novel texture generation method based on Empirical Mode Decomposition (EMD). Extensive qualitative and quantitative experiments show that our method achieves state-of-the-art performance in terms of shape and texture transformation. Moreover, our method supports producing plausible 3D face aging results, which is rarely accomplished by current methods. Qianrui Teng, Rui Wang 0124, Xing Cui, Peipei Li 0002, Zhaofeng He 0001 |
ICME | 3 |
| 2024 | FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMsabstractThe massive generation of multimodal fake news involving both text and images exhibits substantial distribution discrepancies, prompting the need for generalized detectors. However, the insulated nature of training restricts the capability of classical detectors to obtain open-world facts. While Large Vision-Language Models (LVLMs) have encoded rich world knowledge, they are not inherently tailored for combating fake news and struggle to comprehend local forgery details. In this paper, we propose FKA-Owl, a novel framework that leverages forgery-specific knowledge to augment LVLMs, enabling them to reason about manipulations effectively. The augmented forgery-specific knowledge includes semantic correlation between text and images, and artifact trace in image manipulation. To inject these two kinds of knowledge into the LVLM, we design two specialized modules to establish their representations, respectively. The encoded knowledge embeddings are then incorporated into LVLMs. Extensive experiments on the public benchmark demonstrate that FKA-Owl achieves superior cross-domain performance compared to previous methods. Code is publicly available at https://liuxuannan.github.io/FKA_Owl.github.io/. Xuannan Liu, Peipei Li 0002, Huaibo Huang, Zekun Li 0001, Xing Cui, Lixiong Qin, Weihong Deng, Zhaofeng He 0001 |
ACM Multimedia | 5 |
| 2024 | Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention ReasonerabstractFlexible and accurate drag-based editing is a challenging task that has recently garnered significant attention. Current methods typically model this problem as automatically learning "how to drag" through point dragging and often produce one deterministic estimation, which presents two key limitations: 1) Overlooking the inherently ill-posed nature of drag-based editing, where multiple results may correspond to a given input, as illustrated in Fig.1; 2) Ignoring the constraint of image quality, which may lead to unexpected distortion.
To alleviate this, we propose LucidDrag, which shifts the focus from "how to drag" to "what-then-how" paradigm. LucidDrag comprises an intention reasoner and a collaborative guidance sampling mechanism. The former infers several optimal editing strategies, identifying what content and what semantic direction to be edited. Based on the former, the latter addresses "how to drag" by collaboratively integrating existing editing guidance with the newly proposed semantic guidance and quality guidance.
Specifically, semantic guidance is derived by establishing a semantic editing direction based on reasoned intentions, while quality guidance is achieved through classifier guidance using an image fidelity discriminator.
Both qualitative and quantitative comparisons demonstrate the superiority of LucidDrag over previous methods. Xing Cui, Peipei Li 0002, Zekun Li 0001, Xuannan Liu, Yueying Zou, Zhaofeng He 0001 |
NeurIPS | 1 |
| 2024 | Bidirectional Knowledge Reconfiguration for Lightweight Point Cloud AnalysisabstractPoint cloud analysis faces computational system overhead, limiting its application on mobile or edge devices. Directly employing small models may result in a significant drop in performance since it is difficult for a small model to adequately capture local structure and global shape information simultaneously, which are essential clues for point cloud analysis. This paper explores feature distillation for lightweight point cloud models. To mitigate the semantic gap between the lightweight student and the cumbersome teacher, we propose bidirectional knowledge reconfiguration (BKR) to distill informative contextual knowledge from the teacher to the student. Specifically, a top-down knowledge reconfiguration and a bottom-up knowledge reconfiguration are developed to inherit diverse local structure information and consistent global shape knowledge from the teacher, respectively. However, due to the farthest point sampling in most point cloud models, the intermediate features between teacher and student are misaligned, deteriorating the feature distillation performance. To eliminate it, we propose a feature mover's distance (FMD) loss based on optimal transportation, which can measure the distance between unordered point cloud features effectively. Extensive experiments conducted on shape classification, part segmentation, and semantic segmentation benchmarks demonstrate the universality and superiority of our method. Peipei Li 0002, Xing Cui, Yibo Hu 0001, Man Zhang 0005, Ting Yao 0003, Tao Mei 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | ChatEdit: Towards Multi-turn Interactive Facial Image Editing via DialogueabstractSingle-turn Multi-turn Input Lipstick Pale skin Smiling Black Hair Figure 2: Comparison of previous repeated singleturn editing approaches and our proposed multiturn editing approach.The cascaded errors in the single-turn approach lead to unintended changes in gender and eye makeup. Xing Cui, Zekun Li 0001, Yibo Hu 0001, Hailin Shi, Chunshui Cao, Zhaofeng He 0001 |
EMNLP | 1 |