VLDB 2026 Research / reviewers in the wild / expert
Yilin Xiao 0002
dblp:256/8598-2
· DBLP profile ↗
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-9827-6149ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | You Don't Need Pre-Built Graphs for RAG: Retrieval Augmented Generation with Adaptive Reasoning StructuresabstractLarge language models (LLMs) often suffer from hallucination, generating factually incorrect statements when handling questions beyond their knowledge and perception. Retrieval-augmented generation (RAG) addresses this by retrieving query-relevant contexts from knowledge bases to support LLM reasoning. Recent advances leverage pre-constructed graphs to capture the relational connections among distributed documents, showing remarkable performance in complex tasks. However, existing Graph-based RAG (GraphRAG) methods rely on a costly process to transform the corpus into a graph, introducing overwhelming token cost and update latency. Moreover, real-world queries vary in type and complexity, requiring different logic structures for accurate reasoning. The pre-built graph may not align with these required structures, resulting in ineffective knowledge retrieval. To this end, we propose a Logic-aware Retrieval Augmented Generation framework (LogicRAG) that dynamically extracts reasoning structures at inference time to guide adaptive retrieval without any pre-built graph. LogicRAG begins by decomposing the input query into a set of subproblems and constructing a directed acyclic graph (DAG) to model the logical dependencies among them. To support coherent multi-step reasoning, LogicRAG then linearizes the graph using topological sort, so that subproblems can be addressed in a logically consistent order. Besides, LogicRAG applies graph pruning to reduce redundant retrieval and uses context pruning to filter irrelevant context, significantly reducing the overall token cost. Extensive experiments demonstrate that LogicRAG achieves both superior performance and efficiency compared to state-of-the-art baselines. Shengyuan Chen, Chuang Zhou 0002, Zheng Yuan 0013, Qinggang Zhang, Zeyang Cui, Hao Chen 0062, Yilin Xiao 0002, Jiannong Cao 0001, Xiao Huang 0001 |
AAAI | 7 |
| 2026 | Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables QuestionsabstractRecent studies have raised significant concerns regarding the reliability of current mathematical benchmarks, highlighting key limitations such as simplistic design and potential data contamination that undermine evaluation accuracy. Consequently, developing a reliable benchmark that effectively evaluates large language models' (LLMs) genuine capabilities in mathematical reasoning remains a critical challenge. To address these concerns, we propose RV-Bench, a novel evaluation methodology for Benchmarking LLMs with Random Variables in mathematical reasoning. Specifically, we develop question-generating functions to produce random variable questions (RVQs), whose background content mirrors the original benchmark problems, but with randomized variable combinations, rendering them "unseen" to LLMs. Models must completely understand the inherent question pattern to correctly answer RVQs with diverse variable combinations. Thus, an LLMs' genuine reasoning capability is reflected through its accuracy and robustness on RV-Bench. We conducted extensive experiments on over 30 representative LLMs across more than 1,000 RVQs. Our findings reveal that LLMs exhibit a proficiency imbalance between encountered and "unseen" data distributions. Furthermore, RV-Bench reveals that proficiency generalization across similar mathematical reasoning tasks is limited, but we verified that it can still be effectively elicited through test-time scaling. Zijin Hong, Hao Wu 0070, Su Dong 0002, Junnan Dong, Yilin Xiao 0002, Yujing Zhang 0001, Zhu Wang 0016, Feiran Huang, Hongxia Yang, Xiao Huang 0001 |
AAAI | 5 |
| 2026 | LogicPoison: Logical Attacks on Graph Retrieval-Augmented GenerationabstractYilin Xiao, Jin Chen, Qinggang Zhang, Yujing Zhang, Chuang Zhou, Longhao Yang, Lingfei Ren, Xin Yang, Xiao Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yilin Xiao 0002, Qinggang Zhang, Yujing Zhang 0001, Chuang Zhou 0002, Longhao Yang, Lingfei Ren, Xiao Huang 0001 |
ACL (1) | 1 |
| 2026 | Query-Aware Knowledge Retrieval via Hyperbolic StructuringabstractChuang Zhou, Junnan Dong, Yilin Xiao, Shengyuan Chen, Su Dong, di Yin, Xing Sun, Zhaozhuo Xu, Xiao Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chuang Zhou 0002, Junnan Dong, Yilin Xiao 0002, Shengyuan Chen, Su Dong 0002, Xing Sun 0001, Zhaozhuo Xu, Xiao Huang 0001 |
ACL (1) | 3 |
| 2026 | Collision to Cognition: Hash-Driven Graph Construction for Efficient RAGabstractChuang Zhou, Zheng Yuan, Linhao Luo, Zhaozhuo Xu, Yilin Xiao, Junnan Dong, Siyu An, di Yin, Xing Sun, Xiao Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chuang Zhou 0002, Zheng Yuan 0013, Linhao Luo, Zhaozhuo Xu, Yilin Xiao 0002, Junnan Dong, Siyu An, Xing Sun 0001, Xiao Huang 0001 |
ACL (1) | 5 |
| 2026 | Reliable Reasoning Path: Distilling Effective Guidance for LLM Reasoning With Knowledge Graphs
Yilin Xiao 0002, Chuang Zhou 0002, Qinggang Zhang, Bo Li 0037, Qing Li 0001, Xiao Huang 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | FaithfulRAG: Fact-Level Conflict Modeling for Context-Faithful Retrieval-Augmented GenerationabstractQinggang Zhang, Zhishang Xiang, Yilin Xiao, Le Wang, Junhui Li, Xinrun Wang, Jinsong Su. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Qinggang Zhang, Zhishang Xiang, Yilin Xiao 0002, Xinrun Wang, Jinsong Su |
ACL (1) | 3 |
| 2025 | Power on graph: Mining power relationship via user interaction correlation
Yilong Zang, Lingfei Ren, Junhang Wu, Yilin Xiao 0002, Ruimin Hu |
Expert Syst. Appl. | 4 |
| 2025 | DQFormer: Toward Unified LiDAR Panoptic Segmentation With Decoupled Queries for Large-Scale Outdoor ScenesabstractLiDAR panoptic segmentation (LPS) performs semantic and instance segmentation for things (foreground objects) and stuff (background elements), essential for scene perception and remote sensing. While most existing methods separate these tasks using distinct branches (i.e., semantic and instance), recent approaches have unified LPS through a query-based paradigm. However, the distinct spatial distributions of foreground objects and background elements in large-scale outdoor scenes pose challenges. This paper presents DQFormer, a novel framework for unified LPS that employs a decoupled query workflow to adapt to the characteristics of things and stuff in outdoor scenes. It first utilizes a feature encoder to extract multi-scale voxel-wise, point-wise, and BEV features. Then, a decoupled query generator proposes informative queries by localizing things/stuff positions and fusing multi-level BEV embeddings. A query-oriented mask decoder uses masked cross-attention to decode segmentation masks, which are combined with query semantics to produce panoptic results. Extensive experiments on large-scale outdoor scenes, including the vehicular datasets nuScenes and SemanticKITTI, as well as the aerial point cloud dataset DALES, show that DQFormer outperforms superior methods by +1.8%, +0.9%, and +3.5% in panoptic quality (PQ), respectively. Code is available at https://github.com/yuyang-cloud/DQFormer. Yu Yang 0001, Jianbiao Mei, Siliang Du, Yilin Xiao 0002, Huifeng Wu, Yong Liu 0007 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Heterophilic Graph Invariant Learning for Out-of-Distribution of Fraud DetectionabstractGraph-based fraud detection (GFD) has garnered increasing attention due to its effectiveness in identifying fraudsters within multimedia data such as online transactions, product reviews, or telephone voices. However, the prevalent in-distribution (ID) assumption significantly impedes the generalization of GFD approaches to out-of-distribution (OOD) scenarios, which is a pervasive challenge considering the dynamic nature of fraudulent activities. In this paper, we introduce the Heterophilic Graph Invariant Learning Framework (HGIF), a novel approach to bolster the OOD generalization of GFD. HGIF addresses two pivotal challenges: creating diverse virtual training environments and adapting to varying target distributions. Leveraging edge-aware augmentation, HGIF efficiently generates multiple virtual training environments characterized by generalized heterophily distributions, thereby facilitating robust generalization against fraud graphs with diverse heterophily degrees. Moreover, HGIF employs a shared dual-channel encoder with heterophilic graph contrastive learning, enabling the model to acquire stable high-pass and low-pass node representations during training. During the Test-time Training phase, the shared dual-channel encoder is flexibly fine-tuned to adapt to the test distribution through graph contrastive learning. Extensive experiments showcase HGIF's superior performance over existing methods in OOD generalization, setting a new benchmark for GFD in OOD scenarios. Lingfei Ren, Ruimin Hu, Zheng Wang 0007, Yilin Xiao 0002, Dengshi Li, Junhang Wu, Yilong Zang, Jinzhang Hu |
ACM Multimedia | 4 |
| 2024 | Dual-attention-transformer-based semantic reranking for large-scale image localization
Yilin Xiao 0002, Siliang Du, Mingzhong Liu |
Appl. Intell. | 1 |
| 2023 | Where Have You Gone: Category-aware Multigraph Embedding for Missing Point-of-Interest Identification
Junhang Wu, Ruimin Hu, Dengshi Li, Yilin Xiao 0002, Lingfei Ren, Wenyi Hu |
Neural Process. Lett. | 4 |
| 2022 | User Alignment Across Social Networks Based On ego-Network EmbeddingabstractCross-social network user alignment is to find users with the same identity in multiple social networks. It has important applications in natural and scientific fields, such as link prediction and personality recommendation, and has certain research value in the field of data mining. Most current approaches embed social networks in a low-dimensional vector space and then align users in the low-dimensional space. However, because the social network is extremely complex and large, it is easy to be affected by error propagation and noise of different neighbors in the process of network embedding. Therefore, to obtain better embedding, we first form the user's EGO network, then use the random walk to extract the user node sequence, then use the framework of the natural language model to learn the low-dimensional vector representation of the user, and finally train a matrix to map the two social networks into the same feature space for alignment. Our experiments on real-world data set Foursquare-Twitter and Livejournal-myspace show some improvement over several baseline results. Yu Zhen, Ruimin Hu, Dengshi Li, Yilin Xiao 0002 |
IJCNN | 4 |
| 2022 | Where have you been: Dual spatiotemporal-aware user mobility modeling for missing check-in POI identification
Junhang Wu, Ruimin Hu, Dengshi Li, Lingfei Ren, Wenyi Hu, Yilin Xiao 0002 |
Inf. Process. Manag. | 6 |
| 2021 | Multi-level Graph Attention Network based Unsupervised Network AlignmentabstractNetwork alignment is the matching of two networks with corresponding nodes that belong to the same user or entity. The most common application is to analyze which accounts belong to the same user in two social networks. Most of existing techniques rely on matrix factorization so that they cannot be scaled to large-scale networks, are constrained by strict constraints, and cannot learn node embedding without a training set. In this paper, we propose an unsupervised network alignment model based on multi-level graph attention networks. The model uses multi-level graph attention network to learn the embedded representation of nodes, satisfying attribute and structure constraints of alignment. Augmented learning process is proposed to simulate attribute noise and structural noise to improve adaptability of the model. Extensive experiments on real datasets show that the proposed model performs better than the state-of-the-art network alignment model. We also demonstrate the robustness of the proposed model. Yilin Xiao 0002, Ruimin Hu, Dengshi Li, Junhang Wu, Yu Zhen, Lingfei Ren |
LCN | 1 |