Tao He 0014

dblp:94/5035-14 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2026
0000-0002-6052-4573ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Question answering and dialogue systems · 37% Reinforcement learning · 20% Trustworthy machine learning · 15%
Databases, data mining, and information retrieval
2 papers
Knowledge graphs · 100%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › dialogue management
dialogue policy planning
1.722025
Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning · SIGIR 2025
Simulation-Free Hierarchical Latent Policy Planning for Proactive Dialogues · AAAI 2025
Natural language and speech › Question answering and dialogue systems
proactive dialogue
1.622025
Simulation-Free Hierarchical Latent Policy Planning for Proactive Dialogues · AAAI 2025
Planning Like Human: A Dual-process Framework for Dialogue Planning · ACL (1) 2024
Machine learning › Graph learning
graph reasoning
1.012026
Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning
large language model trustworthiness
1.012026
From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty Calibration · AAAI 2026
Machine learning › Reinforcement learning
reinforcement learning for reasoning
1.012026
Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models · ACL (1) 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning
structured reasoning
1.012026
Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models · ACL (1) 2026
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty calibration
1.012026
From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty Calibration · AAAI 2026
Knowledge graphs
knowledge graph reasoning
1.012026
Subgraph-Centric Multi-Agent Reinforcement Learning for Multi-Hop Knowledge Graph Reasoning · IEEE Trans. Knowl. Data Eng. 2026
Knowledge graphs › knowledge graph reasoning
multi-hop reasoning
1.012026
Subgraph-Centric Multi-Agent Reinforcement Learning for Multi-Hop Knowledge Graph Reasoning · IEEE Trans. Knowl. Data Eng. 2026
Knowledge graphs › knowledge graph reasoning
reinforcement learning-based reasoning
1.012026
Subgraph-Centric Multi-Agent Reinforcement Learning for Multi-Hop Knowledge Graph Reasoning · IEEE Trans. Knowl. Data Eng. 2026
Knowledge graphs › knowledge graph reasoning
subgraph reasoning
1.012026
Subgraph-Centric Multi-Agent Reinforcement Learning for Multi-Hop Knowledge Graph Reasoning · IEEE Trans. Knowl. Data Eng. 2026
Machine learning › Reinforcement learning › hierarchical reinforcement learning
hierarchical policy learning
0.912025
Simulation-Free Hierarchical Latent Policy Planning for Proactive Dialogues · AAAI 2025
Machine learning › Reinforcement learning
offline reinforcement learning
0.912025
Simulation-Free Hierarchical Latent Policy Planning for Proactive Dialogues · AAAI 2025
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
user simulation
0.912025
Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning · SIGIR 2025
Knowledge graphs › knowledge graph alignment
entity alignment
0.912025
How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Future · EMNLP 2025
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.812024
Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future · ACL (1) 2024
Natural language and speech › Question answering and dialogue systems › dialogue management
dialogue planning
0.812024
Planning Like Human: A Dual-process Framework for Dialogue Planning · ACL (1) 2024
Natural language and speech › Language models and text generation
large language model reasoning
0.312026
Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models · ACL (1) 2026
Machine learning › Efficient and distributed learning
active learning
0.312025
Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning · SIGIR 2025
Knowledge graphs
knowledge base integration
0.312025
How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Future · EMNLP 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.212024
Planning Like Human: A Dual-process Framework for Dialogue Planning · ACL (1) 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.8topology-aware reasoning · 1.0symbolic reasoning · 1.0subgraph retriever · 1.0multi-agent reinforcement learning · 1.0markov decision process · 1.0cognitive confidence modeling · 1.0alignment framework · 1.0variational autoencoder · 0.9survey · 0.9latent policy discovery · 0.9large language model · 0.9diffusion model · 0.9brownian bridge · 0.9active learning · 0.9
YearPublicationVenuePosition
2026 From Sampling to Cognition: Modeling Internal Cognitive Confidence in Language Models for Robust Uncertainty Calibration
abstract
Large Language Models (LLMs) have demonstrated remarkable performance across a wide range of tasks, yet they generally lack self-awareness, often displaying overconfidence when confronted with questions beyond their knowledge boundaries. This limitation severely hinders their trustworthiness in high-stakes scenarios. Existing calibration methods typically rely on sampling accuracy, derived from multiple outputs, as a proxy for model confidence. However, this coarse-grained metric fails to capture the model’s internal cognitive states, such as confusion, hallucination, or persistent belief in false knowledge. To address this, we propose CogConf (Cognitive Confidence), a cognitively grounded uncertainty signal that extends sampling accuracy by incorporating the semantic diversity of incorrect answers and the model’s abstention behaviors. By shifting the focus from sampling-based to cognition-oriented uncertainty modeling, CogConf offers a more faithful reflection of the model's internal beliefs. Building on this signal, we introduce CogAlign, a simple yet effective alignment framework that explicitly aligns the model’s verbalized confidence with CogConf, thereby producing uncertainty estimates that better reflect the model’s internal cognition. Experimental results on six knowledge-intensive in-domain and out-of-domain QA datasets demonstrate that CogConf robustly characterizes the model's internal uncertainty. Building on this foundation, CogAlign guides the model's expression to significantly enhance the trustworthiness and utility of its uncertainty calibration without compromising its underlying QA capabilities, while also demonstrating strong cross-task generalization and output stability. Offering a new pathway toward building more trustworthy LLMs.
Tao He 0014, Jiafeng Liang, Ming Liu 0004
AAAI2
2026 Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models
abstract
Runxuan Liu, Xianhao Ou, Xinyan Ma, Jiyuan Wang, Jiafeng Liang, Jiaqi Li, Tao He, Zheng Chu, Rongchuan Mu, Zekun Wang, Baoxin Wang, Dayong Wu, Ming Liu, Shijin Wang, Guoping Hu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Runxuan Liu, Xianhao Ou, Xinyan Ma, Jiafeng Liang, Jiaqi Li 0004, Tao He 0014, Rongchuan Mu, Zekun Wang 0001, Baoxin Wang, Dayong Wu, Ming Liu 0004, Shijin Wang 0001, Bing Qin 0001
ACL (1)7
2026 Subgraph-Centric Multi-Agent Reinforcement Learning for Multi-Hop Knowledge Graph Reasoning
abstract
Multi-hop Knowledge Graph Reasoning (KGR) seeks to identify accurate answers within Knowledge Graphs (KGs) via multi-step reasoning, predominantly utilizing reinforcement learning (RL) to enhance the efficiency of the reasoning process. Unlike traditional Knowledge Graph Embedding (KGE) methods, RL-based approaches offer superior interpretability. However, these methods often underperform due to two critical limitations: (1) their over-reliance on Horn rules for reasoning paths, which restricts their expressive power; and (2) inadequate utilization of reasoning states during the process. To address these issues, we propose a novel RL-based framework, RAR, which shifts focus from individual paths to subgraph structures for more robust predictions. RAR frames the retrieval of reasoning subgraphs from the KG as a Markov Decision Process (MDP) and incorporates a subgraph retriever. To efficiently explore the extensive subgraph space, we integrate multi-agent RL to enhance the retriever's capabilities. Additionally, RAR features an advanced analyst module that meticulously examines reasoning states. These modules function iteratively: the retriever expands the subgraph, followed by the analyst module's in-depth analysis. The insights gained are then used to inform subsequent retrieval steps. Ultimately, the predicted scores from both modules are synthesized to produce more precise posterior scores. Experimental results across multiple datasets demonstrate RAR's efficacy, showcasing a notable improvement over existing state-of-the-art RL-based KGR methods.
Tao He 0014, Zerui Chen, Lizi Liao, Yixin Cao 0002, Yuanxing Liu 0001, Wei Tang 0015, Xun Mao, Ming Liu 0004, Bing Qin 0001
IEEE Trans. Knowl. Data Eng.1
2025 Simulation-Free Hierarchical Latent Policy Planning for Proactive Dialogues
abstract
Recent advancements in proactive dialogues have garnered significant attention, particularly for more complex objectives (e.g. emotion support and persuasion). Unlike traditional task-oriented dialogues, proactive dialogues demand advanced policy planning and adaptability, requiring rich scenarios and comprehensive policy repositories to develop such systems. However, existing approaches tend to rely on Large Language Models (LLMs) for user simulation and online learning, leading to biases that diverge from realistic scenarios and result in suboptimal efficiency. Moreover, these methods depend on manually defined, context-independent, coarse-grained policies, which not only incur high expert costs but also raise concerns regarding their completeness. In our work, we highlight the potential for automatically discovering policies directly from raw, real-world dialogue records. To this end, we introduce a novel dialogue policy planning framework, LDPP. It fully automates the process from mining policies in dialogue records to learning policy planning. Specifically, we employ a variant of the Variational Autoencoder to discover fine-grained policies represented as latent vectors. After automatically annotating the data with these latent policy labels, we propose an Offline Hierarchical Reinforcement Learning (RL) algorithm in the latent space to develop effective policy planning capabilities. Our experiments demonstrate that LDPP outperforms existing methods on two proactive scenarios, even surpassing ChatGPT with only a 1.8-billion-parameter LLM.
Tao He 0014, Lizi Liao, Yixin Cao 0002, Yuanxing Liu 0001, Yiheng Sun, Zerui Chen, Ming Liu 0004, Bing Qin 0001
AAAI1
2025 How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Future
abstract
Entity alignment (EA), critical for knowledge graph (KG) integration, identifies equivalent entities across different KGs.Traditional methods often face challenges in semantic understanding and scalability.The rise of language models (LMs), particularly large language models (LLMs), has provided powerful new strategies.This paper systematically reviews LM-driven EA methods, proposing a novel taxonomy that categorizes methods in three key stages: data preparation, feature embedding, and alignment.We further summarize key benchmarks, evaluation metrics, and discuss future directions.This paper aims to provide researchers and practitioners with a clear and comprehensive understanding of how language models reshape the field of entity alignment.* These authors contributed equally.
Zerui Chen, Huiming Fan, Tao He 0014, Ming Liu 0004, Heng Chang, Weijiang Yu, Bing Qin 0001
EMNLP4
2025 Simulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy Planning
abstract
Recent advancements in dialogue policy planning have focused on optimizing system agent policies to achieve predefined goals, emphasizing strategy design, trajectory acquisition, and training efficiency.However, these approaches often overlook the critical role of user characteristics, which are essential in real-world scenarios like conversational search and recommendation, where interactions must adapt to individual user traits such as personality, preferences, and goals.To address this gap, we conduct a comprehensive study using task-specific user personas to evaluate dialogue policy planning under diverse user behaviors.Our analysis, based on these user profiles, reveals significant shortcomings in existing approaches, underscoring the necessity for user-tailored dialogue policies.Building on these insights, we propose the User-Tailored Dialogue Policy Planning (UDP) framework, which integrates an Intrinsic User World Model to capture user traits and feedback.UDP operates in three stages: (1) User Persona Portraying, employing a diffusion model to dynamically infer user profiles; (2) User Feedback Anticipating, using a Brownian Bridge-inspired mechanism to predict user reactions; and (3) User-Tailored Policy Planning, synthesizing these elements to optimize response strategies.To enhance robustness, we introduce an active learning approach that prioritizes challenging user personas during training.Extensive experiments across benchmarks, including both collaborative and non-collaborative settings, demonstrate UDP's effectiveness in learning user-specific dialogue strategies.Results confirm the framework's utility, highlighting its robustness, adaptability, and potential to advance user-centric dialogue systems.
Tao He 0014, Lizi Liao, Ming Liu 0004, Bing Qin 0001
SIGIR1
2025 Exploring & exploiting high-order graph structure for sparse knowledge graph completion
Tao He 0014, Ming Liu 0004, Yixin Cao 0002, Zekun Wang 0001, Bing Qin 0001
Frontiers Comput. Sci.1
2024 Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future
abstract
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, Ting Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He 0014, Haotian Wang 0007, Weihua Peng, Ming Liu 0004, Bing Qin 0001, Ting Liu 0001
ACL (1)5
2024 Planning Like Human: A Dual-process Framework for Dialogue Planning
abstract
In proactive dialogue, the challenge lies not just in generating responses but in steering conversations toward predetermined goals, a task where Large Language Models (LLMs) typically struggle due to their reactive nature.Traditional approaches to enhance dialogue planning in LLMs, ranging from elaborate prompt engineering to the integration of policy networks, either face efficiency issues or deliver suboptimal performance.Inspired by the dualprocess theory in psychology, which identifies two distinct modes of thinking-intuitive (fast) and analytical (slow), we propose the Dual-Process Dialogue Planning (DPDP) framework.DPDP embodies this theory through two complementary planning systems: an instinctive policy model for familiar contexts and a deliberative Monte Carlo Tree Search (MCTS) mechanism for complex, novel scenarios.This dual strategy is further coupled with a novel two-stage training regimen: offline Reinforcement Learning for robust initial policy model formation followed by MCTS-enhanced on-thefly learning, which ensures a dynamic balance between efficiency and strategic depth.Our empirical evaluations across diverse dialogue tasks affirm DPDP's superiority in achieving both high-quality dialogues and operational efficiency, outpacing existing methods. 1
Tao He 0014, Lizi Liao, Yixin Cao 0002, Yuanxing Liu 0001, Ming Liu 0004, Zerui Chen, Bing Qin 0001
ACL (1)1
2024 Relational Graph-Bridged Image-Text Interaction: A Novel Method for Multi-Modal Relation Extraction
abstract
Multi-modal relation extraction (MRE) requires the integration of multi-modal information to identify relationships between entities. Although fine-grained correlations between visual objects and textual words have the potential to improve cross-modal interaction, they are typically modeled implicitly and hindered by the modality gap. This paper introduces a novel method called relational Graph-Bridged cross-modal InTeraction (GBIT). GBIT aims to model fine-grained cross-modal correlations into the interaction process explicitly. This is achieved by constructing a fine-grained cross-modal relational graph, which acts as a bridge for effective cross-modal interaction in multiple layers. Within GBIT, a gated interaction strategy and an adaptive integration module are proposed for irrelevance-filtered information exchange and final information collation. Through extensive experiments on the benchmark MRE, we demonstrate the superiority of our proposed method for MRE.
Tao He 0014, Ming Liu 0004, Zhongyuan Wang 0006, Ruiji Fu, Bing Qin 0001
ICASSP2
2024 VEM2L: an easy but effective framework for fusing text and structure knowledge on sparse knowledge graph completion
Tao He 0014, Ming Liu 0004, Yixin Cao 0002, Meng Qu, Bing Qin 0001
Data Min. Knowl. Discov.1