Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jinxin Liu 0002

dblp:20/6480-2 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0009-4673-9824ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 36% Knowledge representation and reasoning · 25% Question answering and dialogue systems · 17%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model evaluation
1.022024
KoLA: Carefully Benchmarking World Knowledge of Large Language Models · ICLR 2024
Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction · EMNLP 2023
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
How do Transformers Learn Implicit Reasoning? · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
knowledge graph querying
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Natural language and speech › Question answering and dialogue systems
knowledge-intensive question answering
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning · KDD (2) 2025
Natural language and speech › Information extraction and text analysis
open information extraction
0.712023
Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction · EMNLP 2023
Machine learning › Trustworthy machine learning
robustness evaluation
0.712023
Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 0.9representational analysis · 0.9large language model · 0.9cross-query semantic patching · 0.9controlled symbolic training · 0.9chain-of-thought · 0.9self-contrast metric · 0.8contrastive evaluation · 0.8knowledge-invariant clique benchmark · 0.7
YearPublicationVenuePosition
2025 AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning
abstract
Despite the outstanding capabilities of large language models (LLMs), knowledge-intensive reasoning still remains a challenging task due to LLMs' limitations in compositional reasoning and the hallucination problem. A prevalent solution is to employ chain-of-thought (CoT) with retrieval-augmented generation (RAG), which first formulates a reasoning plan by decomposing complex questions into simpler sub-questions, and then applies iterative RAG at each sub-question. However, prior works exhibit two crucial problems: inadequate reasoning planning and poor incorporation of heterogeneous knowledge. In this paper, we introduce AtomR, a framework for LLMs to conduct accurate heterogeneous knowledge reasoning at the atomic level. Inspired by how knowledge graph query languages model compositional reasoning through combining predefined operations, we propose three atomic knowledge operators, a unified set of operators for LLMs to retrieve and manipulate knowledge from heterogeneous sources. First, in the reasoning planning stage, AtomR decomposes a complex question into a reasoning tree where each leaf node corresponds to an atomic knowledge operator, achieving question decomposition that is highly fine-grained and orthogonal. Subsequently, in the reasoning execution stage, AtomR executes each atomic knowledge operator, which flexibly selects, retrieves, and operates atomic level knowledge from heterogeneous sources. We also introduce BlendQA, a challenging benchmark specially tailored for heterogeneous knowledge reasoning. Experiments on three single-source and two multi-source datasets show that AtomR outperforms state-of-the-art baselines by a large margin, with absolute F1 score improvements of 9.4% on 2WikiMultihop and 9.5% on BlendQA. We release our code and data https://github.com/THU-KEG/AtomR.git.
Amy Xin, Jinxin Liu 0002, Zijun Yao 0002, Zhicheng Lee, Shulin Cao, Lei Hou 0001, Juan-Zi Li
KDD (2)2
2025 How do Transformers Learn Implicit Reasoning?
abstract
Recent work suggests that large language models (LLMs) can perform multi-hop reasoning implicitly---producing correct answers without explicitly verbalizing intermediate steps---but the underlying mechanisms remain poorly understood. In this paper, we study how such implicit reasoning emerges by training transformers from scratch in a controlled symbolic environment. Our analysis reveals a three-stage developmental trajectory: early memorization, followed by in-distribution generalization, and eventually cross-distribution generalization. We find that training with atomic triples is not necessary but accelerates learning, and that second-hop generalization relies on query-level exposure to specific compositional structures. To interpret these behaviors, we introduce two diagnostic tools: cross-query semantic patching, which identifies semantically reusable intermediate representations, and a cosine-based representational lens, which reveals that successful reasoning correlates with the cosine-base clustering in hidden space. This clustering phenomenon in turn provides a coherent explanation for the behavioral dynamics observed across training, linking representational structure to reasoning capability. These findings provide new insights into the interpretability of implicit multi-hop reasoning in LLMs, helping to clarify how complex reasoning processes unfold internally and offering pathways to enhance the transparency of such models.
Jiaran Ye, Zijun Yao 0002, Zhidian Huang, Liangming Pan, Jinxin Liu 0002, Yushi Bai, Amy Xin, Weichuan Liu, Xiaoyin Che, Lei Hou 0001, Juan-Zi Li
NeurIPS5
2025 Dynamic multi teacher knowledge distillation for semantic parsing in KBQA
Ao Zou, Shulin Cao, Jinxin Liu 0002, Lei Hou 0001
Expert Syst. Appl.5
2024 DiaKoP: Dialogue-based Knowledge-oriented Programming for Neural-symbolic Knowledge Base Question Answering
Zhicheng Lee, Zhidian Huang, Zijun Yao 0002, Jinxin Liu 0002, Amy Xin, Lei Hou 0001, Juan-Zi Li
CIKM4
2024 KoLA: Carefully Benchmarking World Knowledge of Large Language Models
abstract
The unprecedented performance of large language models (LLMs) necessitates improvements in evaluations. Rather than merely exploring the breadth of LLM abilities, we believe meticulous and thoughtful designs are essential to thorough, unbiased, and applicable evaluations. Given the importance of world knowledge to LLMs, we construct a Knowledge-oriented LLM Assessment benchmark (KoLA), in which we carefully design three crucial factors: (1) For ability modeling, we mimic human cognition to form a four-level taxonomy of knowledge-related abilities, covering 19 tasks. (2) For data, to ensure fair comparisons, we use both Wikipedia, a corpus prevalently pre-trained by LLMs, along with continuously collected emerging corpora, aiming to evaluate the capacity to handle unseen data and evolving knowledge. (3) For evaluation criteria, we adopt a contrastive system, including overall standard scores for better numerical comparability across tasks and models, and a unique self-contrast metric for automatically evaluating knowledge-creating ability. We evaluate 21 open-source and commercial LLMs and obtain some intriguing findings. The KoLA dataset will be updated every three months to provide timely references for developing LLMs and knowledge-related systems.
Jifan Yu, Xiaozhi Wang, Shangqing Tu, Shulin Cao, Daniel Zhang-Li, Hao Peng 0015, Zijun Yao 0002, Hanming Li, Zheyuan Zhang 0002, Yushi Bai, Yantao Liu, Amy Xin, Kaifeng Yun, Linlu Gong, Nianyi Lin, Zhi-Li Wu, Yunjia Qi, Weikai Li 0002, Kaisheng Zeng, Ji Qi 0003, Hailong Jin, Jinxin Liu 0002, Yu Gu 0029, Yuan Yao 0011, Ning Ding 0002, Lei Hou 0001, Zhiyuan Liu 0001, Bin Xu 0001, Jie Tang 0001, Juan-Zi Li
ICLR27
2023 Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction
abstract
The robustness to distribution changes ensures that NLP models can be successfully applied in the realistic world, especially for information extraction tasks.However, most prior evaluation benchmarks have been devoted to validating pairwise matching correctness, ignoring the crucial validation of robustness.In this paper, we present the first benchmark that simulates the evaluation of open information extraction models in the real world, where the syntactic and expressive distributions under the same knowledge meaning may drift variously.We design and annotate a large-scale testbed in which each example is a knowledge-invariant clique that consists of sentences with structured knowledge of the same meaning but with different syntactic and expressive forms.By further elaborating the robustness metric, a model is judged to be robust if its performance is consistently accurate on the overall cliques.We perform experiments on typical models published in the last decade as well as a representative large language model, and the results show that the existing successful models exhibit a frustrating degradation, with a maximum drop of 23.43 F 1 score.Our resources and code are available at https://github.com/qijimrc/ROBUST.
Ji Qi 0003, Chuchun Zhang, Xiaozhi Wang, Kaisheng Zeng, Jifan Yu, Jinxin Liu 0002, Lei Hou 0001, Juan-Zi Li, Xu Bin
EMNLP6