VLDB 2026 Research / reviewers in the wild / expert
Xuankun Rong
dblp:400/6453
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 32% Vision and language · 26% Learning paradigms · 21% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 100% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
1.7 | 2 | 2025 | MOTION: Multi-Sculpt Evolutionary Coarsening for Federated Continual Graph Learning · NeurIPS 2025 CAN: Leveraging Clients As Navigators for Generative Replay in Federated Continual Learning · ICML 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.3 | 2 | 2026 | PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language Models · AAAI 2026 Probing Semantic Insensitivity for Inference-Time Backdoor Defense in Multimodal Large Language Model · AAAI 2026 |
Security and privacy of machine learning › adversarial attack
backdoor attack |
1.3 | 2 | 2026 | Probing Semantic Insensitivity for Inference-Time Backdoor Defense in Multimodal Large Language Model · AAAI 2026 PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language Models · AAAI 2026 |
Security and privacy of machine learning › adversarial attack › backdoor attack
backdoor defense |
1.0 | 1 | 2026 | PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language Models · AAAI 2026 |
Security and privacy of machine learning › adversarial attack › backdoor attack › backdoor defense
backdoor detection |
1.0 | 1 | 2026 | Probing Semantic Insensitivity for Inference-Time Backdoor Defense in Multimodal Large Language Model · AAAI 2026 |
Machine learning › Learning paradigms
continual learning |
0.9 | 1 | 2025 | MOTION: Multi-Sculpt Evolutionary Coarsening for Federated Continual Graph Learning · NeurIPS 2025 |
Machine learning › Efficient and distributed learning › federated learning
federated continual learning |
0.9 | 1 | 2025 | CAN: Leveraging Clients As Navigators for Generative Replay in Federated Continual Learning · ICML 2025 |
Machine learning › Learning paradigms › continual learning › rehearsal-based continual learning
generative replay |
0.9 | 1 | 2025 | CAN: Leveraging Clients As Navigators for Generative Replay in Federated Continual Learning · ICML 2025 |
Machine learning › Graph learning
graph neural network |
0.9 | 1 | 2025 | MOTION: Multi-Sculpt Evolutionary Coarsening for Federated Continual Graph Learning · NeurIPS 2025 |
Natural language and speech › Question answering and dialogue systems › dialogue evaluation
multi-turn dialogue evaluation |
0.9 | 1 | 2025 | MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models · ICCV 2025 |
Computer vision › Vision and language › vision-language model
vision-language model evaluation |
0.9 | 1 | 2025 | MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models · ICCV 2025 |
Methods — techniques the papers use, named apart from their topics
visual token purification · 2.0semantic perturbation · 2.0confidence drift analysis · 2.0attention-guided token filtering · 2.0parameter aggregation · 0.9graph coarsening · 0.9generative replay · 0.9client-centric data synthesis · 0.9checklist-based evaluation · 0.9GPT-4o as evaluator · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language ModelsabstractDownstream fine-tuning of Multimodal Large Language Models (MLLMs) is advancing rapidly, allowing general models to achieve superior performance on domain-specific tasks. Yet most prior research focuses on performance gains and overlooks the vulnerability of the fine-tuning pipeline: attackers can easily poison the dataset to implant backdoors into MLLMs. We conduct an in-depth investigation of backdoor attacks on MLLMs and reveal the phenomenon of Attention Hijacking and its Hierarchical Mechanism. Guided by this insight, we propose PurMM, a test-time backdoor purification framework that removes visual tokens exhibiting anomalous attention, thereby avoiding targeted outputs while restoring correct answers. PurMM contains three stages: (1) locating tokens with abnormal attention, (2) filtering them using deep-layer cues, and (3) zeroing out their corresponding components in the visual embeddings. Unlike existing defences, PurMM dispenses with retraining and training-process modifications, operating at test-time to restore model performance while eliminating the backdoor. Extensive experiments across multiple MLLMs and datasets show that PurMM maintains normal performance, sharply reduces attack success rates, and consistently converts backdoor outputs to benign ones, offering a new perspective for safeguarding MLLMs. Wenzheng Jiang, Ke Liang 0006, Xuankun Rong, Jingxuan Zhou, Zhengyi Zhong, Guancheng Wan, Ji Wang 0002 |
AAAI | 3 |
| 2026 | Probing Semantic Insensitivity for Inference-Time Backdoor Defense in Multimodal Large Language ModelabstractThe massive scale of data and computation required for training Multimodal Large Language Models (MLLMs) has fueled the rise of Fine-Tuning as a Service (FTaaS), enabling users to rapidly customize models for diverse real-world tasks. While FTaaS democratizes access to advanced multimodal intelligence, it also introduces serious security concerns, particularly backdoor attacks. In this work, we systematically analyze backdoor vulnerabilities in MLLMs under the FTaaS paradigm, revealing two key phenomena: (1) markedly reduced sensitivity to textual variations when a visual trigger is present, and (2) abnormally stable model confidence even under strong semantic perturbations. Building on these insights, we propose Trap on Text (ToT), a novel inference-time backdoor detection framework. ToT applies controlled semantic perturbations to textual prompts and jointly analyzes the semantic consistency and confidence drift of the model’s responses, enabling robust detection of backdoor activations without requiring model parameters, architectures or clean reference data. Extensive experiments across architectures and datasets show that ToT achieves strong attack mitigation and preserves clean accuracy, offering a practical solution for safeguarding FTaaS workflows. Xuankun Rong, Wenke Huang 0003, Wenzheng Jiang, Yiming Li 0004, Wenxuan Wang 0001, Mang Ye |
AAAI | 1 |
| 2025 | MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language ModelsabstractVision-and-Language Models (VLMs) have shown impressive capabilities on single-turn benchmarks, yet real-world applications often demand more intricate multi-turn dialogues. Existing multi-turn datasets (e.g, MMDU, ConvBench) only partially capture the breadth and depth of conversational scenarios encountered by users. In this work, we introduce MultiVerse, a novel multi-turn conversation benchmark featuring 647 dialogues - each averaging four turns - derived from a diverse set of 12 popular VLM evaluation benchmarks. With 484 tasks and 484 interaction goals, MultiVerse covers a wide range of topics, from factual knowledge and perception to advanced reasoning tasks such as mathematics and coding. To facilitate robust assessment, we propose a checklist-based evaluation method that leverages GPT-4o as the automated evaluator, measuring performance across 37 key aspects, including perceptual accuracy, linguistic clarity, and factual correctness. We evaluate 18 VLMs on MultiVerse, revealing that even the strongest models (e.g., GPT-4o) achieve only a 50% success rate in complex multi-turn conversations, highlighting the dataset's challenging nature. Notably, we find that providing full dialogue context significantly enhances performance for smaller or weaker models, emphasizing the importance of in-context learning. We believe MultiVerse is a landscape of evaluating multi-turn interaction abilities for VLMs. Young-Jun Lee, Yechan Hwang, Byungsoo Ko, Han-Gyu Kim, Dongyu Yao, Xuankun Rong, Eojin Joo, Seung-Ho Han 0001, Bowon Ko, Ho-Jin Choi |
ICCV | 8 |
| 2025 | CAN: Leveraging Clients As Navigators for Generative Replay in Federated Continual LearningabstractGenerative replay (GR) has been extensively validated in continual learning as a mechanism to synthesize data and replay past knowledge to mitigate forgetting.
By leveraging synthetic rather than real data for the replay, GR has been adopted in some federated continual learning (FCL) approaches to ensure the privacy of client-side data.
While existing GR-based FCL approaches have introduced improvements, none of their enhancements specifically take into account the unique characteristics of federated learning settings.
Beyond privacy constraints, what other fundamental aspects of federated learning should be explored in the context of FCL?
In this work, we explore the potential benefits that come from emphasizing the role of clients throughout the process.
We begin by highlighting two key observations: (a) Client Expertise Superiority, where clients, rather than the server, act as domain experts, and (b) Client Forgetting Variance, where heterogeneous data distributions across clients lead to varying levels of forgetting.
Building on these insights, we propose CAN (Clients As Navigators), highlighting the pivotal role of clients in both data synthesis and data replay.
Extensive evaluations demonstrate that this client-centric approach achieves state-of-the-art performance. Notably, it requires a smaller buffer size, reducing storage overhead and enhancing computational efficiency. Xuankun Rong, Jianshu Zhang 0003, Mang Ye |
ICML | 1 |
| 2025 | Backdoor Cleaning without External Guidance in MLLM Fine-tuningabstractMultimodal Large Language Models (MLLMs) are increasingly deployed in fine-tuning-as-a-service (FTaaS) settings, where user-submitted datasets adapt general-purpose models to downstream tasks. This flexibility, however, introduces serious security risks, as malicious fine-tuning can implant backdoors into MLLMs with minimal effort. In this paper, we observe that backdoor triggers systematically disrupt cross-modal processing by causing abnormal attention concentration on non-semantic regions—a phenomenon we term **attention collapse**. Based on this insight, we propose **Believe Your Eyes (BYE)**, a data filtering framework that leverages attention entropy patterns as self-supervised signals to identify and filter backdoor samples. BYE operates via a three-stage pipeline: (1) extracting attention maps using the fine-tuned model, (2) computing entropy scores and profiling sensitive layers via bimodal separation, and (3) performing unsupervised clustering to remove suspicious samples. Unlike prior defenses, BYE equires no clean supervision, auxiliary labels, or model modifications. Extensive experiments across various datasets, models, and diverse trigger types validate BYE's effectiveness: it achieves near-zero attack success rates while maintaining clean-task performance, offering a robust and generalizable solution against backdoor threats in MLLMs. Xuankun Rong, Wenke Huang 0003, Jian Liang 0003, Jinhe Bi, Xun Xiao, Yiming Li 0004, Bo Du 0001, Mang Ye |
NeurIPS | 1 |
| 2025 | MOTION: Multi-Sculpt Evolutionary Coarsening for Federated Continual Graph LearningabstractGraph neural networks (GNNs) have achieved remarkable success in various domains but typically rely on centralized, static graphs, which limits their applicability in distributed, evolving environments. To address this limitation, we define the task of Federated Continual Graph Learning (FCGL), a paradigm for incremental learning on dynamic graphs distributed across decentralized clients. Existing methods, however, neither preserve graph topology during task transitions nor mitigate parameter conflicts in server‐side aggregation. To overcome these challenges, we introduce **MOTION**, a generalizable FCGL framework that integrates two complementary modules: the Graph Topology‐preserving Multi‐Sculpt Coarsening (G‐TMSC) module, which maintains the structural integrity of past graphs through a multi‐expert, similarity‐guided fusion process, and the Graph‐Aware Evolving Parameter Adaptive Engine (G‐EPAE) module, which refines global model updates by leveraging a topology‐sensitive compatibility matrix. Extensive experiments on real‐world datasets show that our approach improves average accuracy (AA) by an average of 30\% $\uparrow$ over the FedAvg baseline across five datasets while maintaining a negative $\downarrow$ average forgetting (AF) rate, significantly enhancing generalization and robustness under FCGL settings. The code is available for anonymous access at https://anonymous.4open.science/r/MOTION. Frank Wan, Fengyuan Ran, Wenke Huang 0003, Xuankun Rong, Guibin Zhang, Bo Du 0001, Mang Ye |
NeurIPS | 5 |