VLDB 2026 Research / reviewers in the wild / expert
Xiaqiang Tang
dblp:353/6066
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0004-5827-0297ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 53% Deep learning architectures and training · 19% Efficient and distributed learning · 9% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 50% Parallel and multicore computing · 50% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.9 | 2 | 2026 | REFO: Reinforced Evolutionary Faithfulness Optimization for Large Language Models · AAAI 2026 Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs · AAAI 2025 |
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
faithfulness |
1.0 | 1 | 2026 | REFO: Reinforced Evolutionary Faithfulness Optimization for Large Language Models · AAAI 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
faithfulness evaluation |
0.9 | 1 | 2025 | CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models · ACL (1) 2025 |
Machine learning › Trustworthy machine learning
hallucination |
0.9 | 1 | 2025 | CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models · ACL (1) 2025 |
Natural language and speech › Language models and text generation
hallucination detection |
0.9 | 1 | 2025 | CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models · ACL (1) 2025 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
0.9 | 1 | 2025 | Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs · AAAI 2025 |
Machine learning › Efficient and distributed learning › efficient training
long-context training |
0.9 | 1 | 2025 | Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large Clusters · AAAI 2025 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.9 | 1 | 2025 | Improving Bilinear RNN with Closed-loop Control · NeurIPS 2025 |
Information retrieval › interactive information retrieval
adaptive retrieval |
0.9 | 1 | 2025 | Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs · AAAI 2025 |
Information retrieval › retrieval models
retrieval model selection |
0.9 | 1 | 2025 | Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs · AAAI 2025 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.3 | 1 | 2026 | REFO: Reinforced Evolutionary Faithfulness Optimization for Large Language Models · AAAI 2026 |
Distributed systems › distributed machine learning
distributed training |
0.3 | 1 | 2025 | Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large Clusters · AAAI 2025 |
Parallel and multicore computing
pipeline parallelism |
0.3 | 1 | 2025 | Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large Clusters · AAAI 2025 |
Mathematical optimization
control theory |
0.3 | 1 | 2025 | Improving Bilinear RNN with Closed-loop Control · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 2.7sequence accumulation · 1.7pipeline parallelism · 1.7multi-armed bandit · 1.7delta learning rule · 1.7chunk-wise parallel kernel · 1.7self-evolution · 1.0attention-based loss · 1.0state feedback · 0.9output feedback · 0.9hallucination detection · 0.9automatic annotation pipeline · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REFO: Reinforced Evolutionary Faithfulness Optimization for Large Language ModelsabstractDespite its success in enriching LLMs with external knowledge, RAG remains plagued by faithfulness hallucinations, where generated text contradicts the retrieved source information. Previous research on faithfulness hallucination in LLMs is frequently hindered by prohibitive manual annotation costs and a dependency on static datasets, which caps their performance and adaptability. Furthermore, these models lack a clear training mechanism to explicitly promote contextual focus. In this work, we propose a novel iterative self-evolution framework to enhance model faithfulness. This framework autonomously generates high-quality data and leverages it for the continuous self-optimization of the model, leading to significant improvements in faithfulness. Our experimental analysis reveals that improving model faithfulness encourages a closer alignment of the attention distribution with the given context. Based on this finding, we design an attention-based loss function to further promote this process. Experimental results show that our model achieves state-of-the-art faithfulness on a range of context-based question-answering datasets, marking a significant advancement over previous approaches. Xiaqiang Tang, Keyu Hu, Haojie Lu, Sihong Xie |
AAAI | 2 |
| 2025 | Sequence Accumulation and Beyond: Infinite Context Length on Single GPU and Large ClustersabstractLinear sequence modeling methods, such as linear attention, state space modeling, and linear RNNs, have recently been recognized as potential alternatives to softmax attention thanks to their linear complexity and competitive performance. However, although their linear-memory advantage during training enables dealing with long sequences, it is still hard to handle extremely long sequences with very limited computational resources. In this paper, we propose Sequence Accumulation (SA) which leverages the common recurrence feature of linear sequence modeling methods to manage infinite context length even on a single GPU. Specifically, SA divides long input sequences into fixed-length sub-sequences and accumulates intermediate states sequentially, which reaches only constant-memory consumption. Additionally, we further propose Sequence Accumulation with Pipeline Parallelism (SAPP), to train large models with infinite context length, without incurring any additional synchronization costs in the sequence dimension. Extensive experiments with a wide range of context lengths are conducted to validate the effectiveness of SA and SAPP on both single and multiple GPUs. Results show that SA and SAPP enable the training of infinite context length on even very limited resources, and are well compatible with the out-of-the-box distributed training techniques. Weigao Sun, Yongtuo Liu, Xiaqiang Tang, Xiaoyu Mo |
AAAI | 3 |
| 2025 | Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge GraphsabstractDespite the superior performance of Large language models on many NLP tasks, they still face significant limitations in memorizing extensive world knowledge. Recent studies have demonstrated that leveraging the Retrieval-Augmented Generation (RAG) framework, combined with Knowledge Graphs that encapsulate extensive factual data in a structured format, robustly enhances the reasoning capabilities of LLMs. However, deploying such systems in real-world scenarios presents challenges: the continuous evolution of non-stationary environments may lead to performance degradation and user satisfaction requires a careful balance of performance and responsiveness. To address these challenges, we introduce a Multi-objective Multi-Armed Bandit enhanced RAG framework, supported by multiple retrieval methods with diverse capabilities under rich and evolving retrieval contexts in practice. Within this framework, each retrieval method is treated as a distinct "arm''. The system utilizes real-time user feedback to adapt to dynamic environments, by selecting the appropriate retrieval method based on input queries and the historical multi-objective performance of each arm. Extensive experiments conducted on two benchmark KGQA datasets demonstrate that our method significantly outperforms baseline methods in non-stationary settings while achieving state-of-the-art performance in station environments. Xiaqiang Tang, Sihong Xie |
AAAI | 1 |
| 2025 | CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language ModelsabstractFaithfulness hallucinations are claims generated by a Large Language Model (LLM) not supported by contexts provided to the LLM. Lacking assessment standards, existing benchmarks focus on “factual statements” that rephrase source materials while overlooking “cognitive statements” that involve making inferences from the given context. Consequently, evaluating and detecting the hallucination of cognitive statements remains challenging. Inspired by how evidence is assessed in the legal domain, we design a rigorous framework to assess different levels of faithfulness of cognitive statements and introduce the CogniBench dataset where we reveal insightful statistics. To keep pace with rapidly evolving LLMs, we further develop an automatic annotation pipeline that scales easily across different models. This results in a large-scale CogniBench-L dataset, which facilitates training accurate detectors for both factual and cognitive hallucinations. We release our model and datasets at: https://github.com/FUTUREEEEEE/CogniBench Xiaqiang Tang, Keyu Hu, Xi Zhang 0008, Weigao Sun, Sihong Xie |
ACL (1) | 1 |
| 2025 | MBA-RAG: a Bandit Approach for Adaptive Retrieval-Augmented Generation through Question ComplexityabstractRetrieval Augmented Generation (RAG) has proven to be highly effective in boosting the generative performance of language model in knowledge-intensive tasks. However, existing RAG framework either indiscriminately perform retrieval or rely on rigid single-label classifiers to select retrieval methods, leading to inefficiencies and suboptimal performance across queries of varying complexity. To address these challenges, we propose a reinforcement learning-based framework that dynamically selects the most suitable retrieval strategy based on query complexity. To address these challenges, we propose a reinforcement learning-based framework that dynamically selects the most suitable retrieval strategy based on query complexity. Our approach leverages a multi-armed bandit algorithm, which treats each retrieval method as a distinct “arm” and adapts the selection process by balancing exploration and exploitation. Additionally, we introduce a dynamic reward function that balances accuracy and efficiency, penalizing methods that require more retrieval steps, even if they lead to a correct result. Our method achieves new state of the art results on multiple single-hop and multi-hop datasets while reducing retrieval costs. Our code are available at https://github.com/FUTUREEEEEE/MBA. Xiaqiang Tang, Sihong Xie |
COLING | 1 |
| 2025 | Improving Bilinear RNN with Closed-loop ControlabstractRecent efficient sequence modeling methods, such as Gated DeltaNet, TTT, and RWKV-7, have achieved performance improvements by supervising the recurrent memory management through the Delta learning rule. Unlike previous state-space models (e.g., Mamba) and gated linear attentions (e.g., GLA), these models introduce interactions between the recurrent state and the key vector, resulting in a bilinear recursive structure. In this paper, we first introduce the concept of Bilinear RNNs with a comprehensive analysis on the advantages and limitations of these models. Then based on the closed-loop control theory, we propose a novel Bilinear RNN variant named Comba, which adopts a scalar-plus-low-rank state transition, with both state feedback and output feedback corrections. We also implement a hardware-efficient chunk-wise parallel kernel in Triton and train models with 340M/1.3B parameters on a large-scale corpus. Comba demonstrates its superior performance and computation efficiency on both language modeling and vision tasks. Jiaxi Hu, Yongqi Pan, Jusen Du, Disen Lan, Xiaqiang Tang, Qingsong Wen, Yuxuan Liang 0002, Weigao Sun |
NeurIPS | 5 |
| 2023 | FGNet: A Graph-Based Motion Forecasting Method From a Future PerspectiveabstractOne essential task for autonomous driving is to accurately predict the future motions of surrounding traffic agents. Recently, Graph Neural Network (GNN) approaches have shown potential in motion prediction due to the fact that information in the traffic scenario can be inherently formed into a graph structure. However, existing approaches are limited to using a static graph representing the states of the current timestamp, ignoring the prediction scenario’s derivation. In this work, we propose the Future Graph Network (FGNet) a two-stage GNN-based model for accurate and real-time multi-agent motion prediction. We design a future graph decoder that takes first-stage F mode multi-agent prediction result as input and reconstructs a graph to represent the derivation result of the scenario. The decoder enhances the first-stage result by (1) Constraining the agent’s future trajectory in the same modal to be harmonious and (2) Introducing the subsequent road information to guide the long-term prediction. Meanwhile, we incorporate statistical characteristics of the vehicles’ trajectory into the design of the loss function, which further boosts the performance of our model. Experiments show that FGNet achieves competitive performance on the Argoverse motion forecasting benchmark with real-time inference. Xiaqiang Tang, Yafeng Guo, Jun Wang 0025 |
IV | 1 |