EDBT 2026 Demo / reviewers in the wild / expert
Rui Miao 0003
dblp:62/675-3
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-2917-2311ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BlindGuard: Safeguarding LLM-based Multi-Agent Systems under Unknown AttacksabstractRui Miao, Yixin Liu, Yili Wang, Xu Shen, Yue Tan, Yiwei Dai, Shirui Pan, Xin Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rui Miao 0003, Yixin Liu 0001, Yili Wang 0004, Xu Shen 0002, Yiwei Dai, Shirui Pan, Xin Wang 0035 |
ACL (1) | 1 |
| 2026 | Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly DetectionabstractJunjun Pan, Yixin Liu, Rui Miao, Kaize Ding, Yu Zheng, Quoc Viet Hung Nguyen, Alan Wee-Chung Liew, Shirui Pan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. JunJun Pan, Yixin Liu 0001, Rui Miao 0003, Kaize Ding, Yu Zheng 0013, Nguyen Quoc Viet Hung, Alan Wee-Chung Liew, Shirui Pan |
ACL (1) | 3 |
| 2026 | Graph Defense Diffusion ModelabstractGraph Neural Networks (GNNs) are highly vulnerable to adversarial attacks, which can greatly degrade their performance. Existing graph purification methods attempt to address this issue by filtering attacked graphs. However, they struggle to defend effectively against multiple types of adversarial attacks (e.g., targeted attacks and non-targeted attacks) simultaneously due to limited flexibility. Additionally, these methods lack comprehensive modeling of graph data, relying heavily on heuristic prior knowledge. To overcome these challenges, we introduce the Graph Defense Diffusion Model (GDDM), a flexible purification method that leverages the denoising and modeling capabilities of diffusion models. The iterative nature of diffusion models aligns well with the stepwise process of adversarial attacks, making them particularly suitable for defense. By iteratively adding and removing noises (edges), GDDM effectively purifies attacked graphs, restoring their original structures and features. Our GDDM consists of two key components: (1) Graph Structure-Driven Refiner, which preserves the basic fidelity of the graph during the denoising process, and ensures that the generated graph remains consistent with the original scope; and (2) Node Feature-Constrained Regularizer, which removes residual impurities from the denoised graph, further enhancing the purification effect. By designing tailored denoising strategies to handle different types of adversarial attacks, we improve the GDDM's adaptability to various attack scenarios. Furthermore, GDDM demonstrates strong scalability, leveraging its structural properties to seamlessly transfer across similar datasets without retraining. Extensive experiments on three real-world datasets demonstrate that GDDM outperforms state-of-the-art methods in defending against various adversarial attacks, showcasing its robustness and effectiveness. Xin He 0003, Wenqi Fan, Yili Wang 0004, Chengyi Liu 0001, Rui Miao 0003, Xin Juan, Xin Wang 0035 |
KDD (1) | 5 |
| 2025 | Understanding the Information Propagation Effects of Communication Topologies in LLM-based Multi-Agent SystemsabstractThe communication topology in large language model-based multi-agent systems fundamentally governs inter-agent collaboration patterns, critically shaping both the efficiency and effectiveness of collective decision-making.While recent studies for communication topology automated design tend to construct sparse structures for efficiency, they often overlook why and when sparse and dense topologies help or hinder collaboration.In this paper, we present a causal framework to analyze how agent outputs, whether correct or erroneous, propagate under topologies with varying sparsity.Our empirical studies reveal that moderately sparse topologies, which effectively suppress error propagation while preserving beneficial information diffusion, typically achieve optimal task performance.Guided by this insight, we propose a novel topology design approach, EIB-LEARNER, that balances error suppression and beneficial information propagation by fusing connectivity patterns from both dense and sparse graphs.Extensive experiments show the superior effectiveness, communication cost, and robustness of EIB-LEARNER.The code is in Xu Shen 0002, Yixin Liu 0001, Yiwei Dai, Yili Wang 0004, Rui Miao 0003, Shirui Pan, Xin Wang 0035 |
EMNLP | 5 |
| 2025 | Unifying Unsupervised Graph-Level Anomaly Detection and Out-of-Distribution Detection: A BenchmarkabstractTo build safe and reliable graph machine learning systems, unsupervised graph-level anomaly detection (GLAD) and unsupervised graph-level out-of-distribution (OOD) detection (GLOD) have received significant attention in recent years. Though these two lines of research share the same objective, they have been studied independently in the community due to distinct evaluation setups, creating a gap that hinders the application and evaluation of methods from one to the other. To bridge the gap, in this work, we present a Unified Benchmark for unsupervised Graph-level OOD and anomaly Detection (UB-GOLD), a comprehensive evaluation framework that unifies GLAD and GLOD under the concept of generalized graph-level OOD detection. Our benchmark encompasses 35 datasets spanning four practical anomaly and OOD detection scenarios, facilitating the comparison of 18 representative GLAD/GLOD methods. We conduct multi-dimensional analyses to explore the effectiveness, generalizability, robustness, and efficiency of existing methods, shedding light on their strengths and limitations. Furthermore, we provide an open-source codebase of UB-GOLD to foster reproducible research and outline potential directions for future investigations based on our insights. Yili Wang 0004, Yixin Liu 0001, Xu Shen 0002, Rui Miao 0003, Kaize Ding, Ying Wang 0009, Shirui Pan, Xin Wang 0035 |
ICLR | 5 |
| 2025 | Mamba-Based Graph Convolutional Networks: Tackling Over-smoothing with Selective State SpaceabstractGraph Neural Networks (GNNs) have shown great success in various graph-based learning tasks. However, it often faces the issue of over-smoothing as the model depth increases, which causes all node representations to converge to a single value and become indistinguishable. This issue stems from the inherent limitations of GNNs, which struggle to distinguish the importance of information from different neighborhoods. In this paper, we introduce MbaGCN, a novel graph convolutional architecture that draws inspiration from the Mamba paradigm—originally designed for sequence modeling. MbaGCN presents a new backbone for GNNs, consisting of three key components: the Message Aggregation Layer, the Selective State Space Transition Layer, and the Node State Prediction Layer. These components work in tandem to adaptively aggregate neighborhood information, providing greater flexibility and scalability for deep GNN models. While MbaGCN may not consistently outperform all existing methods on each dataset, it provides a foundational framework that demonstrates the effective integration of the Mamba paradigm into graph representation learning. Through extensive experiments on benchmark datasets, we demonstrate that MbaGCN paves the way for future advancements in graph neural network research. Our code is in https://github.com/hexin5515/MbaGCN. Xin He 0003, Yili Wang 0004, Wenqi Fan, Xu Shen 0002, Xin Juan, Rui Miao 0003, Xin Wang 0035 |
IJCAI | 6 |
| 2025 | AdaGCL+: An Adaptive Subgraph Contrastive Learning Toward Tackling Topological BiasabstractLarge-scale graph data poses a training scalability challenge, which is generally treated by employing batch sampling methods to divide the graph into smaller subgraphs and train them in batches. However, such an approach introduces a topological bias in the local batches compared with the complete graph structure, missing either node features or edges. This topological bias is empirically shown to affect the generalization capabilities of graph neural networks (GNNs). To address this issue, we propose adaptive subgraph contrastive learning (AdaGCL) that bridges the gap between large-scale batch sampling and its generalization poorness. Specifically, AdaGCL augments graphs depending on the sampled batches and leverages a subgraph-granularity contrastive loss to learn the node embeddings invariant among the augmented imperfect graphs. To optimize the augmentation strategy for each downstream application, we introduce a node-centric information bottleneck (Node-IB) to control the trade-off regarding the similarity and diversity between the original and augmented graphs. This enhanced version of AdaGCL referred to as AdaGCL+, automates the graph augmentation process by dynamically adjusting graph perturbation parameters (e.g., edge dropping rate) to minimize the downstream loss. Extensive experimental results showcase the scalability of AdaGCL+ to graphs with millions of nodes using batch sampling methods. AdaGCL+ consistently outperforms existing methods on numerous benchmark datasets in terms of node classification accuracy and runtime efficiency. Yili Wang 0004, Ninghao Liu 0001, Rui Miao 0003, Ying Wang 0009, Xin Wang 0035 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Rethinking Independent Cross-Entropy Loss For Graph-Structured DataabstractGraph neural networks (GNNs) have exhibited prominent performance in learning graph-structured data. Considering node classification task, based on the i.i.d assumption among node labels, the traditional supervised learning simply sums up cross-entropy losses of the independent training nodes and applies the average loss to optimize GNNs’ weights. But different from other data formats, the nodes are naturally connected. It is found that the independent distribution modeling of node labels restricts GNNs’ capability to generalize over the entire graph and defend adversarial attacks. In this work, we propose a new framework, termed joint-cluster supervised learning, to model the joint distribution of each node with its corresponding cluster. We learn the joint distribution of node and cluster labels conditioned on their representations, and train GNNs with the obtained joint loss. In this way, the data-label reference signals extracted from the local cluster explicitly strengthen the discrimination ability on the target node. The extensive experiments demonstrate that our joint-cluster supervised learning can effectively bolster GNNs’ node classification accuracy. Furthermore, being benefited from the reference signals which may be free from spiteful interference, our learning paradigm significantly protects the node classification from being affected by the adversarial attack. Rui Miao 0003, Kaixiong Zhou, Yili Wang 0004, Ninghao Liu 0001, Ying Wang 0009, Xin Wang 0035 |
ICML | 1 |
| 2022 | AdaGCL: Adaptive Subgraph Contrastive Learning to Generalize Large-scale Graph TrainingabstractTraining graph neural networks (GNNs) with good generalizability on large-scale graphs is a challenging problem. Existing methods mainly divide the input graph into multiple subgraphs and train them in different batches to improve training scalability. However, the local batches obtained by such a strategy could contain topological bias compared with the complete graph structure. It has been studied that the topological bias results in more significant gaps between training and testing performances, or worse generalization robustness. A straightforward solution is to utilize contrastive learning, and train node embeddings to be robust and invariant among the augmented imperfect graphs. However, most of the existing work are inefficient by contrasting extensive node pairs at the large-scale graph. With random data augmentation, they may deteriorate the embedding process by transforming well-sampled batches into meaningless graph structures. Yili Wang 0004, Kaixiong Zhou, Rui Miao 0003, Ninghao Liu 0001, Xin Wang 0035 |
CIKM | 3 |
| 2022 | Contrastive Graph Convolutional Networks with adaptive augmentation for text classification
Yintao Yang, Rui Miao 0003, Yili Wang 0004, Xin Wang 0035 |
Inf. Process. Manag. | 2 |
| 2022 | Negative samples selecting strategy for graph contrastive learningabstractGraph neural networks (GNNs) have emerged as a successful method on graph structured data. Limited by expensive labeled data, contrastive learning has been adopted to the graph domain. In most existing node-level graph contrastive learning methods, when applying contrastive learning to a certain unlabeled node (the center node), its corresponding “similar” node (positive sample) is usually generated by data augmentation. Other nodes in the graph are served as the “dissimilar” nodes (negative samples), which leads to two major problems. First, the computational cost can be prohibitively expensive, especially when the graph is large. Second, utilizing some nodes which share the same label with the center node as the negative samples will damage the learning process. Hence, to address these issues, we explore the feasibility of only sampling a part of nodes for graph contrastive learning process. And unlike the previous self-supervised contrastive methods, we use joint training to exploit supervised signals as much as possible in contrastive learning. Hence, we propose a Negative Samples Selecting Strategy to utilize the classification prediction to guide the selection of the negative samples for sampled nodes. Then, we further incorporate this strategy for performing contrastive learning on graphs and propose a framework named Graph Contrastive Learning with Negative Samples Selecting Strategy (GCNSS). We demonstrate that GCNSS can be trained much faster with much less computation memory than graph contrastive learning baselines, and GCNSS can effectively boost the performance of existing GNN models on semi-supervised node classification tasks across many different datasets. The code is in: https://github.com/MR9812/GCNSS. Rui Miao 0003, Yintao Yang, Yao Ma 0001, Xin Juan, Haotian Xue 0001, Jiliang Tang, Ying Wang 0009, Xin Wang 0035 |
Inf. Sci. | 1 |