Amy Xin

dblp:349/5224 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0001-2404-0475ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 55% Knowledge representation and reasoning · 22% Question answering and dialogue systems · 15%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
1.622025
AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios · NeurIPS 2025
KoLA: Carefully Benchmarking World Knowledge of Large Language Models · ICLR 2024
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Natural language and speech › Language models and text generation
instruction following
0.912025
AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios · NeurIPS 2025
Natural language and speech › Language models and text generation › instruction following
instruction-following agents
0.912025
AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
How do Transformers Learn Implicit Reasoning? · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
knowledge graph querying
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Natural language and speech › Question answering and dialogue systems
knowledge-intensive question answering
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Natural language and speech › Language models and text generation
LLM agents
0.312025
AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 0.9representational analysis · 0.9large language model · 0.9cross-query semantic patching · 0.9controlled symbolic training · 0.9code-based evaluation · 0.9chain-of-thought · 0.9LLM-based evaluation · 0.9self-contrast metric · 0.8contrastive evaluation · 0.8
YearPublicationVenuePosition
2025 LLMAEL: Large Language Models are Good Context Augmenters for Entity Linking
abstract
Specialized entity linking (EL) models are well-trained at mapping mentions to unique knowledge base (KB) entities according to a given context. However, specialized EL models struggle to disambiguate long-tail entities due to their limited training data. Meanwhile, extensively pre-trained large language models (LLMs) possess broader knowledge of uncommon entities. Yet, with a lack of specialized EL training, LLMs frequently fail to generate accurate KB entity names, limiting their standalone effectiveness in EL. With the observation that LLMs are more adept at context generation instead of EL execution, we introduce LLM-Augmented Entity Linking (LLMAEL), the first framework to enhance specialized EL models with LLM data augmentation. LLMAEL leverages off-the-shelf, tuning-free LLMs as context augmenters, generating entity descriptions to serve as additional input for specialized EL models. Experiments show that LLMAEL sets new state-of-the-art results across 6 widely adopted EL benchmarks: compared to prior methods that integrate tuning-free LLMs into EL, LLMAEL achieves an absolute 8.9% gain in EL accuracy. We release our code and datasets.
Amy Xin, Yunjia Qi, Zijun Yao 0002, Fangwei Zhu, Kaisheng Zeng, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li
CIKM1
2025 AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning
abstract
Despite the outstanding capabilities of large language models (LLMs), knowledge-intensive reasoning still remains a challenging task due to LLMs' limitations in compositional reasoning and the hallucination problem. A prevalent solution is to employ chain-of-thought (CoT) with retrieval-augmented generation (RAG), which first formulates a reasoning plan by decomposing complex questions into simpler sub-questions, and then applies iterative RAG at each sub-question. However, prior works exhibit two crucial problems: inadequate reasoning planning and poor incorporation of heterogeneous knowledge. In this paper, we introduce AtomR, a framework for LLMs to conduct accurate heterogeneous knowledge reasoning at the atomic level. Inspired by how knowledge graph query languages model compositional reasoning through combining predefined operations, we propose three atomic knowledge operators, a unified set of operators for LLMs to retrieve and manipulate knowledge from heterogeneous sources. First, in the reasoning planning stage, AtomR decomposes a complex question into a reasoning tree where each leaf node corresponds to an atomic knowledge operator, achieving question decomposition that is highly fine-grained and orthogonal. Subsequently, in the reasoning execution stage, AtomR executes each atomic knowledge operator, which flexibly selects, retrieves, and operates atomic level knowledge from heterogeneous sources. We also introduce BlendQA, a challenging benchmark specially tailored for heterogeneous knowledge reasoning. Experiments on three single-source and two multi-source datasets show that AtomR outperforms state-of-the-art baselines by a large margin, with absolute F1 score improvements of 9.4% on 2WikiMultihop and 9.5% on BlendQA. We release our code and data https://github.com/THU-KEG/AtomR.git.
Amy Xin, Jinxin Liu 0002, Zijun Yao 0002, Zhicheng Lee, Shulin Cao, Lei Hou 0001, Juan-Zi Li
KDD (2)1
2025 AGENTIF: Benchmarking Large Language Models Instruction Following Ability in Agentic Scenarios
abstract
Large Language Models (LLMs) have demonstrated advanced capabilities in real-world agentic applications. Growing research efforts aim to develop LLM-based agents to address practical demands, introducing a new challenge: agentic scenarios often involve lengthy instructions with complex constraints, such as extended system prompts and detailed tool specifications. While adherence to such instructions is crucial for agentic applications, whether LLMs can reliably follow them remains underexplored. In this paper, we introduce AgentIF, the first benchmark for systematically evaluating LLM instruction following ability in agentic scenarios. AgentIF features three key characteristics: (1) Realistic, constructed from $50$ real-world agentic applications. (2) Long, averaging $1,723$ words with a maximum of $15,630$ words. (3) Complex, averaging $11.9$ constraints per instruction, covering diverse constraint types, such as tool specifications and condition constraints.To construct AgentIF, we collect $707$ human-annotated instructions across $50$ agentic tasks from industrial application agents and open-source agentic systems. For each instruction, we annotate the associated constraints and corresponding evaluation metrics, including code-based evaluation, LLM-based evaluation, and hybrid code-LLM evaluation.We use AgentIF to systematically evaluate existing advanced LLMs. We observe that current models generally perform poorly, especially in handling complex constraint structures and tool specifications. We further conduct error analysis and analytical experiments on instruction length and meta constraints, providing some findings about the failure modes of existing LLMs. We have released the code and data to facilitate future research.
Yunjia Qi, Hao Peng 0015, Xiaozhi Wang, Amy Xin, Youfeng Liu, Bin Xu 0001, Lei Hou 0001, Juan-Zi Li
NeurIPS4
2025 How do Transformers Learn Implicit Reasoning?
abstract
Recent work suggests that large language models (LLMs) can perform multi-hop reasoning implicitly---producing correct answers without explicitly verbalizing intermediate steps---but the underlying mechanisms remain poorly understood. In this paper, we study how such implicit reasoning emerges by training transformers from scratch in a controlled symbolic environment. Our analysis reveals a three-stage developmental trajectory: early memorization, followed by in-distribution generalization, and eventually cross-distribution generalization. We find that training with atomic triples is not necessary but accelerates learning, and that second-hop generalization relies on query-level exposure to specific compositional structures. To interpret these behaviors, we introduce two diagnostic tools: cross-query semantic patching, which identifies semantically reusable intermediate representations, and a cosine-based representational lens, which reveals that successful reasoning correlates with the cosine-base clustering in hidden space. This clustering phenomenon in turn provides a coherent explanation for the behavioral dynamics observed across training, linking representational structure to reasoning capability. These findings provide new insights into the interpretability of implicit multi-hop reasoning in LLMs, helping to clarify how complex reasoning processes unfold internally and offering pathways to enhance the transparency of such models.
Jiaran Ye, Zijun Yao 0002, Zhidian Huang, Liangming Pan, Jinxin Liu 0002, Yushi Bai, Amy Xin, Weichuan Liu, Xiaoyin Che, Lei Hou 0001, Juan-Zi Li
NeurIPS7
2024 DiaKoP: Dialogue-based Knowledge-oriented Programming for Neural-symbolic Knowledge Base Question Answering
Zhicheng Lee, Zhidian Huang, Zijun Yao 0002, Jinxin Liu 0002, Amy Xin, Lei Hou 0001, Juan-Zi Li
CIKM5
2024 KoLA: Carefully Benchmarking World Knowledge of Large Language Models
abstract
The unprecedented performance of large language models (LLMs) necessitates improvements in evaluations. Rather than merely exploring the breadth of LLM abilities, we believe meticulous and thoughtful designs are essential to thorough, unbiased, and applicable evaluations. Given the importance of world knowledge to LLMs, we construct a Knowledge-oriented LLM Assessment benchmark (KoLA), in which we carefully design three crucial factors: (1) For ability modeling, we mimic human cognition to form a four-level taxonomy of knowledge-related abilities, covering 19 tasks. (2) For data, to ensure fair comparisons, we use both Wikipedia, a corpus prevalently pre-trained by LLMs, along with continuously collected emerging corpora, aiming to evaluate the capacity to handle unseen data and evolving knowledge. (3) For evaluation criteria, we adopt a contrastive system, including overall standard scores for better numerical comparability across tasks and models, and a unique self-contrast metric for automatically evaluating knowledge-creating ability. We evaluate 21 open-source and commercial LLMs and obtain some intriguing findings. The KoLA dataset will be updated every three months to provide timely references for developing LLMs and knowledge-related systems.
Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, Hao Peng 0015, Zijun Yao 0002, Hanming Li, Zheyuan Zhang 0002, Yushi Bai, Yantao Liu, Amy Xin, Kaifeng Yun, Linlu Gong, Nianyi Lin, Zhi-Li Wu, Yunjia Qi, Weikai Li 0002, Kaisheng Zeng, Ji Qi 0003, Hailong Jin, Jinxin Liu 0002, Yu Gu 0029, Yuan Yao 0011, Ning Ding 0002, Lei Hou 0001, Zhiyuan Liu 0001, Bin Xu 0001, Jie Tang 0001, Juan-Zi Li
ICLR15