Kai Xiong 0002

dblp:38/6410-2 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-5909-3075ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 4 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired Benchmark
abstract
Translating brain signals into text could restore communication for people with severe paralysis, yet practically usable systems to date rely on invasive electrocorticography (ECoG).Electroencephalography (EEG) offers a non-invasive alternative, and EEG-to-text (EEG2Text) has been widely explored.Interestingly, however, EEG2Text models generally rely on teacher-forcing evaluation; without it, they fail to generate meaningful decoding.This reliance prevents EEG2Text from being applied in real-world, non-academic settings.This has fueled numerous debates about whether EEG2Text is a meaningful direction, by extension, and whether EEG truly contains decodable linguistic information.Here, using a neuropsychology-informed paradigm, we find that existing EEG2Text benchmarks have neglected EEG instability, a flaw that has confounded inference and sparked debate.Our experiments furnish key evidence for the feasibility of teacher-forcing-free EEG2Text decoding.Accordingly, we assemble the Corpus OF Eeg-To-Text (COFETT) using a 128-channel highdensity EEG cap, providing a benchmark dedicated to evaluating EEG2Text models.In comparisons with multiple existing benchmarks, COFETT achieves SOTA ability to distinguish among model performances and enables robust, teacher-forcing-free evaluation, thereby opening a path toward practical EEG2Text applications.COFETT is open sourced in https: //github.com/baoyudu/COFETT.
Tianyi Jiang, Kai Xiong 0002
ACL (1)5
2026 Consolidation or Adaptation? PRISM: Disentangling SFT and RL Data via Gradient Concentration
abstract
Yang Zhao, Yangou Ouyang, Xiao Ding, Hepeng Wang, Bibo Cai, Kai Xiong, Jinglong Gao, Zhouhao Sun, Li Du, Bing Qin, Ting Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yang Zhao 0023, Yangou Ouyang, Hepeng Wang, Bibo Cai, Kai Xiong 0002, Jinglong Gao, Zhouhao Sun, Bing Qin 0001, Ting Liu 0001
ACL (1)6
2026 MAESTRO: Meta-learning Adaptive Estimation of Scalarization Trade-offs for Reward Optimization
abstract
Yang Zhao, Hepeng Wang, Xiao Ding, Yangou Ouyang, Bibo Cai, Kai Xiong, Jinglong Gao, Zhouhao Sun, Li Du, Bing Qin, Ting Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yang Zhao 0023, Hepeng Wang, Yangou Ouyang, Bibo Cai, Kai Xiong 0002, Jinglong Gao, Zhouhao Sun, Bing Qin 0001, Ting Liu 0001
ACL (1)6
2026 Necessary and sufficient knowledge enhanced collaborative logical reasoning in LLMs
Kai Xiong 0002, Bing Qin 0001, Ting Liu 0001
Neural Networks3
2025 Com² : A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models
abstract
Large language models (LLMs) have mastered abundant simple and explicit commonsense knowledge through pre-training, enabling them to achieve human-like performance in simple commonsense reasoning. Nevertheless, LLMs struggle to reason with complex and implicit commonsense knowledge that is derived from simple ones (such as understanding the long-term effects of certain events), an aspect humans tend to focus on more. Existing works focus on complex tasks like math and code, while complex commonsense reasoning remains underexplored due to its uncertainty and lack of structure. To fill this gap and align with real-world concerns, we propose a benchmark Com^2 focusing on complex commonsense reasoning. We first incorporate causal event graphs to serve as structured complex commonsense. Then we adopt causal theory (e.g., intervention) to modify the causal event graphs and obtain different scenarios that meet human concerns. Finally, an LLM is employed to synthesize examples with slow thinking, which is guided by the logical relationships in the modified causal graphs. Furthermore, we use detective stories to construct a more challenging subset. Experiments show that LLMs struggle in reasoning depth and breadth, while post-training and slow thinking can alleviate this. The code and data are available at https://github.com/Waste-Wood/Com2.
Kai Xiong 0002, Yixin Cao 0002, Yuxiong Yan, Jinglong Gao, Jiaqian Liu, Bing Qin 0001, Ting Liu 0001
ACL (1)1
2025 Analyzing the Rapid Generalization of SFT via the Perspective of Attention Head Activation Patterns
abstract
LLMs’ performance on complex tasks is still unsatisfactory. A key issue is that presently LLMs learn in a data-driven schema, while the instructions about these complex tasks are both scarce and hard to collect or construct. On the contrary, a prominent phenomenon is that LLMs can learn rather fast on simpler tasks with adequate prior knowledge captured during pretraining stage. Thus, if the prerequisite and mechanism of such rapid generalization could be elucidated, it could enhance the efficiency and effectiveness of the LLM’s ability to learn complex tasks. Thus, in this paper, we employ a gradient-based method, to dissect the process that the SFT process adapts LLMs to downstream tasks via the perspective of attention patterns. We find that: (1) LLMs selectively activate task-specific attention heads during SFT; (2) activation patterns for complex tasks are combinations of basic task patterns; and (3) changes in a few parameters can significantly impact activation patterns after SFT on a small number of samples.Based on these insights, experiments are conducted to actually enhance the efficiency and effectiveness of SFT.
Yang Zhao 0023, Kai Xiong 0002, Ting Liu 0001, Bing Qin 0001
ACL (1)4
2025 Beyond Similarity: A Gradient-based Graph Method for Instruction Tuning Data Selection
abstract
Yang Zhao, Li Du, Xiao Ding, Yangou Ouyang, Hepeng Wang, Kai Xiong, Jinglong Gao, Zhouhao Sun, Dongliang Xu, Qing Yang, Dongchen Li, Bing Qin, Ting Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yang Zhao 0023, Yangou Ouyang, Hepeng Wang, Kai Xiong 0002, Jinglong Gao, Zhouhao Sun, Dongliang Xu, Qing Yang 0033, Bing Qin 0001, Ting Liu 0001
ACL (1)6
2025 UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection
abstract
A primary impediment to scaling reinforcement learning (RL) for large language model (LLM) training is the substantial computational cost, predominantly arising from the necessity of multi-sampling for policy optimization and evaluation. This underscores the critical yet challenging nature of efficient training data selection. Drawing inspiration from the Zone of Proximal Development (ZPD) theory, which posits that learners acquire knowledge more effectively from tasks of intermediate difficulty, we hypothesize that LLMs exhibit optimal learning from data they have not yet mastered but demonstrate the potential to comprehend. Conventional methodologies for assessing data difficulty or informativeness typically rely on computationally intensive multi-sampling or iterative procedures. To address this limitation, we introduce UFO-RL (**U**ncertainty-**F**ocused **O**ptimization for **R**einforcement **L**earning), a novel framework that employs a computationally efficient single-pass uncertainty estimation technique to identify informative training instances. This method, requiring only a single forward pass and obviating the need for iterative next-token computation, achieves a significant acceleration (up to 185$\times$) in data evaluation compared to multi-sampling approaches. UFO-RL leverages this efficient metric to select data within the model's estimated ZPD for training. Extensive experimentation across diverse LLMs and mathematical benchmarks demonstrates that training with a mere 10\% of the data, carefully selected by UFO-RL, yields performance comparable to or even surpassing that of full-data training. Furthermore, this targeted data selection results in up to a 16$\times$ reduction in overall training time, concurrently enhancing training stability and improving generalization capabilities. Thus, UFO-RL presents a practical and highly efficient strategy for scaling RL fine-tuning of LLMs by focusing learning efforts on the most informative and valuable data, thereby mitigating the computational bottlenecks associated with traditional RL training.
Yang Zhao 0023, Kai Xiong 0002, Yangou Ouyang, Zhouhao Sun, Jiannan Guan, Bing Qin 0001, Ting Liu 0001
NeurIPS2
2025 Improving cross-task generalization with step-by-step instructions
Yang Wu 0010, Bing Qin 0001, Kai Xiong 0002
Sci. China Inf. Sci.5
2025 Think straight or think again? Continual joint learning of deduction, abduction and induction
Kai Xiong 0002, Yixin Cao 0002, Yang Zhao 0023, Ting Liu 0001, Bing Qin 0001
Neural Networks1
2024 Intuitive or Dependent? Investigating LLMs' Behavior Style to Conflicting Prompts
abstract
This study investigates the behaviors of Large Language Models (LLMs) when faced with conflicting prompts versus their internal memory.This will not only help to understand LLMs' decision mechanism but also benefit real-world applications, such as retrievalaugmented generation (RAG).Drawing on cognitive theory, we target the first scenario of decision-making styles where there is no superiority in the conflict and categorize LLMs' preference into dependent, intuitive, and rational/irrational styles.Another scenario of factual robustness considers the correctness of prompt and memory in knowledge-intensive tasks, which can also distinguish if LLMs behave rationally or irrationally in the first scenario.To quantify them, we establish a complete benchmarking framework including a dataset, a robustness evaluation pipeline, and corresponding metrics.Extensive experiments with seven LLMs reveal their varying behaviors.And, with role play intervention, we can change the styles, but different models present distinct adaptivity and upper-bound.One of our key takeaways is to optimize models or the prompts according to the identified style.For instance, RAG models with high role play adaptability may dynamically adjust the interventions according to the quality of retrieval results -being dependent to better leverage informative context; and, being intuitive when the external prompt is noisy.Our dataset can be found at https://github.com/yingjiahao14/KRE.
Jiahao Ying, Yixin Cao 0002, Kai Xiong 0002, Long Cui, Yidong He
ACL (1)3
2024 Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact Guidance
abstract
Large language models (LLMs) have developed impressive performance and strong explainability across various reasoning scenarios, marking a significant stride towards mimicking human-like intelligence. Despite this, when tasked with several simple questions supported by a generic fact, LLMs often struggle to abstract and apply the generic fact to provide consistent and precise answers, revealing a deficiency in abstract reasoning abilities. This has sparked a vigorous debate about whether LLMs are genuinely reasoning or merely memorizing. In light of this, we design a preliminary study to quantify and delve into the abstract reasoning abilities of existing LLMs. Our findings reveal a substantial discrepancy between their general reasoning and abstract reasoning performances. To relieve this problem, we tailor an abstract reasoning dataset (AbsR) together with a meaningful learning paradigm to teach LLMs how to leverage generic facts for reasoning purposes. The results show that our approach not only boosts the general reasoning performance of LLMs but also makes considerable strides towards their capacity for abstract reasoning, moving beyond simple memorization or imitation to a more nuanced understanding and application of generic facts. The code is available at https://github.com/Waste-Wood/MeanLearn.
Kai Xiong 0002, Ting Liu 0001, Bing Qin 0001, Dongliang Xu, Qing Yang 0033, Hongtao Liu 0008, Yixin Cao 0002
NeurIPS1
2022 e-CARE: a New Dataset for Exploring Explainable Causal Reasoning
abstract
Understanding causality has vital importance for various Natural Language Processing (NLP) applications.Beyond the labeled instances, conceptual explanations of the causality can provide deep understanding of the causal facts to facilitate the causal reasoning process.However, such explanation information still remains absent in existing causal reasoning resources.In this paper, we fill this gap by presenting a human-annotated explainable CAusal REasoning dataset (e-CARE), which contains over 21K causal reasoning questions, together with natural language formed explanations of the causal questions.Experimental results show that generating valid explanations for causal facts still remains especially challenging for the state-of-the-art models, and the explanation information can be helpful for promoting the accuracy and stability of causal reasoning models.
Kai Xiong 0002, Ting Liu 0001, Bing Qin 0001
ACL (1)3
2022 ReCo: Reliable Causal Chain Reasoning via Structural Causal Recurrent Neural Networks
abstract
Causal chain reasoning (CCR) is an essential ability for many decision-making AI systems, which requires the model to build reliable causal chains by connecting causal pairs.However, CCR suffers from two main transitive problems: threshold effect and scene drift.In other words, the causal pairs to be spliced may have a conflicting threshold boundary or scenario.To address these issues, we propose a novel Reliable Causal chain reasoning framework (ReCo), which introduces exogenous variables to represent the threshold and scene factors of each causal pair within the causal chain, and estimates the threshold and scene contradictions across exogenous variables via structural causal recurrent neural networks (SRNN).Experiments show that ReCo outperforms a series of strong baselines on both Chinese and English CCR datasets.Moreover, by injecting reliable causal chain knowledge distilled by ReCo, BERT can achieve better performances on four downstream causal-related tasks than BERT models enhanced by other kinds of knowledge.
Kai Xiong 0002, Ting Liu 0001, Bing Qin 0001, Baoxing Huai
EMNLP1
2022 Enhancing pretrained language models with structured commonsense knowledge for textual inference
Kai Xiong 0002, Ting Liu 0001, Bing Qin 0001
Knowl. Based Syst.3
2021 ExCAR: Event Graph Knowledge Enhanced Explainable Causal Reasoning
abstract
Li Du, Xiao Ding, Kai Xiong, Ting Liu, Bing Qin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Kai Xiong 0002, Ting Liu 0001, Bing Qin 0001
ACL/IJCNLP (1)3