Linhai Zhang

dblp:256/3827 · DBLP profile ↗
← Back
20ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Static Artifacts: An Evolutionary Framework for Synthetic Claim Generation
abstract
Yeqing Teng, Jiasheng Si, Shuxia Lin, Linhai Zhang, Weiyu Zhang, Wenpeng Lu, Deyu Zhou, Xiaoming Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yeqing Teng, Jiasheng Si, Shuxia Lin, Linhai Zhang, Weiyu Zhang 0001, Wenpeng Lu
ACL (1)4
2026 Beyond Meta-Reasoning: Metacognitive Consolidation for Self-Improving LLM Reasoning
abstract
Large language models (LLMs) have demonstrated strong reasoning capabilities, and as existing approaches for enhancing LLM reasoning continue to mature, increasing attention has shifted toward meta-reasoning as a promising direction for further improvement.However, most existing meta-reasoning methods remain episodic: they focus on executing complex meta-reasoning routines within individual instances, but ignore the accumulation of reusable meta-reasoning skills across instances, leading to recurring failure modes and repeatedly high metacognitive effort.In this paper, we introduce Metacognitive Consolidation, a novel framework in which a model consolidates metacognitive experience from past reasoning episodes into reusable knowledge that improves future meta-reasoning.We instantiate this framework by structuring instance-level problem solving into distinct roles for reasoning, monitoring, and control to generate rich, attributable meta-level traces.These traces are then consolidated through a hierarchical, multi-timescale update mechanism that gradually forms evolving metaknowledge.Experimental results demonstrate consistent performance gains across benchmarks and backbone models, and show that performance improves as metacognitive experience accumulates over time.
Ziqing Zhuang, Linhai Zhang, Jiasheng Si, Yulan He 0001
ACL (1)2
2026 From Personal to Clinical: Personalisation and Depersonalisation for Explainable Depression Detection
Ziyang Gao, Linhai Zhang, Yulan He 0001
DASFAA (3)2
2025 Causal Prompting: Debiasing Large Language Model Prompting Based on Front-Door Adjustment
abstract
Despite the notable advancements of existing prompting methods, such as In-Context Learning and Chain-of-Thought for Large Language Models (LLMs), they still face challenges related to various biases. Traditional debiasing methods primarily focus on the model training stage, including approaches based on data augmentation and reweighting, yet they struggle with the complex biases inherent in LLMs. To address such limitations, the causal relationship behind the prompting methods is uncovered using a structural causal model, and a novel causal prompting method based on front-door adjustment is proposed to effectively mitigate LLMs biases. In specific, causal intervention is achieved by designing the prompts without accessing the parameters and logits of LLMs. The chain-of-thought generated by LLM is employed as the mediator variable and the causal effect between input prompts and output answers is calculated through front-door adjustment to mitigate model biases. Moreover, to accurately represent the chain-of-thoughts and estimate the causal effects, contrastive learning is used to fine-tune the encoder of chain-of-thought by aligning its space with that of the LLM. Experimental results show that the proposed causal prompting approach achieves excellent performance across seven natural language processing datasets on both open-source and closed-source LLMs.
Congzhi Zhang, Linhai Zhang, Jialong Wu 0007, Yulan He 0001
AAAI2
2025 SCOPE: Optimizing Key-Value Cache Compression in Long-context Generation
abstract
Key-Value (KV) cache has become a bottleneck of LLMs for long-context generation.Despite the numerous efforts in this area, the optimization for the decoding phase is generally ignored.However, we believe such optimization is crucial, especially for long-output generation tasks based on the following two observations: (i) Excessive compression during the prefill phase which requires specific full context, impairs the comprehension of the reasoning task; (ii) Deviation of heavy hitters 1 occurs in the reasoning tasks with long outputs.Therefore, SCOPE, a simple yet efficient framework that separately performs KV cache optimization during the prefill and decoding phases, is introduced.Specifically, the KV cache during the prefill phase is preserved to maintain the essential information, while a novel strategy based on sliding is proposed to select essential heavy hitters for the decoding phase.Memory usage and memory transfer are further optimized using adaptive and discontinuous strategies.Extensive experiments on LONGGENBENCH show the effectiveness and generalization of SCOPE and its compatibility as a plug-in to other prefill-only KV compression methods. 2
Jialong Wu 0007, Zhenglin Wang, Linhai Zhang, Yilong Lai, Yulan He 0001
ACL (1)3
2025 WebWalker: Benchmarking LLMs in Web Traversal
abstract
Retrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. However, traditional search engines may retrieve shallow content, limiting the ability of LLMs to handle complex, multi-layered information. To address this, we introduce WebWalkerQA, a benchmark designed to assess the ability of LLMs to perform web traversal. It evaluates the capacity of LLMs to traverse a website’s subpages to extract high-quality data systematically. We propose WebWalker, which is a multi-agent framework that mimics human-like web navigation through an explore-critic paradigm. Extensive experimental results show that WebWalkerQA is challenging and demonstrates the effectiveness of RAG combined with WebWalker, through this horizontal and vertical integration in real-world scenarios.
Jialong Wu 0007, Wenbiao Yin, Yong Jiang 0005, Zhenglin Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He 0001, Pengjun Xie, Fei Huang 0002
ACL (1)7
2025 PROPER: A Progressive Learning Framework for Personalized Large Language Models with Group-Level Adaptation
abstract
Personalized large language models (LLMs) aim to tailor their outputs to user preferences.Recent advances in parameter-efficient finetuning (PEFT) methods have highlighted the effectiveness of adapting population-level LLMs to personalized LLMs by fine-tuning userspecific parameters with user history.However, user data is typically sparse, making it challenging to adapt LLMs to specific user patterns.To address this challenge, we propose PROgressive PERsonalization (PROPER), a novel progressive learning framework inspired by meso-level theory in social science.PROPER bridges population-level and user-level models by grouping users based on preferences and adapting LLMs in stages.It combines a Mixture-of-Experts (MoE) structure with Low Ranked Adaptation (LoRA), using a user-aware router to assign users to appropriate groups automatically.Additionally, a LoRA-aware router is proposed to facilitate the integration of individual user LoRAs with group-level LoRAs.Experimental results show that PROPER significantly outperforms SOTA models across multiple tasks, demonstrating the effectiveness of our approach.Our code is available at https://github.com/callanwu/PROPER.
Linhai Zhang, Jialong Wu 0007, Yulan He 0001
ACL (1)1
2025 CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
abstract
Chain-of-Thought (CoT) reasoning enhances Large Language Models (LLMs) by encouraging step-by-step reasoning in natural language.However, leveraging a latent continuous space for reasoning may offer benefits in terms of both efficiency and robustness.Prior implicit CoT methods attempt to bypass language completely by reasoning in continuous space but have consistently underperformed compared to the standard explicit CoT approach.We introduce CODI (Continuous Chain-of-Thought via Self-Distillation), a novel training framework that effectively compresses natural language CoT into continuous space.CODI jointly trains a teacher task (Explicit CoT) and a student task (Implicit CoT), distilling the reasoning ability from language into continuous space by aligning the hidden states of a designated token.Our experiments show that CODI is the first implicit CoT approach to match the performance of explicit CoT on GSM8k at the GPT-2 scale, achieving a 3.1x compression rate and outperforming the previous stateof-the-art by 28.2% in accuracy.CODI also demonstrates robustness, generalizable to complex datasets, and interpretability.These results validate that LLMs can reason effectively not only in natural language, but also in a latent continuous space.Code is available at https://github.com/zhenyi4/codi.
Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du 0001, Yulan He 0001
EMNLP3
2025 Relevant emotion ranking with topic-enhanced emotion transition
Linhai Zhang, Xin Zhang 0108
Knowl. Based Syst.1
2024 Causal Walk: Debiasing Multi-Hop Fact Verification with Front-Door Adjustment
abstract
Multi-hop fact verification aims to detect the veracity of the given claim by integrating and reasoning over multiple pieces of evidence. Conventional multi-hop fact verification models are prone to rely on spurious correlations from the annotation artifacts, leading to an obvious performance decline on unbiased datasets. Among the various debiasing works, the causal inference-based methods become popular by performing theoretically guaranteed debiasing such as casual intervention or counterfactual reasoning. However, existing causal inference-based debiasing methods, which mainly formulate fact verification as a single-hop reasoning task to tackle shallow bias patterns, cannot deal with the complicated bias patterns hidden in multiple hops of evidence. To address the challenge, we propose Causal Walk, a novel method for debiasing multi-hop fact verification from a causal perspective with front-door adjustment. Specifically, in the structural causal model, the reasoning path between the treatment (the input claim-evidence graph) and the outcome (the veracity label) is introduced as the mediator to block the confounder. With the front-door adjustment, the causal effect between the treatment and the outcome is decomposed into the causal effect between the treatment and the mediator, which is estimated by applying the idea of random walk, and the causal effect between the mediator and the outcome, which is estimated with normalized weighted geometric mean approximation. To investigate the effectiveness of the proposed method, an adversarial multi-hop fact verification dataset and a symmetric multi-hop fact verification dataset are proposed with the help of the large language model. Experimental results show that Causal Walk outperforms some previous debiasing methods on both existing datasets and the newly constructed datasets. Code and data will be released at https://github.com/zcccccz/CausalWalk.
Congzhi Zhang, Linhai Zhang
AAAI2
2024 TECA: A Two-stage Approach with Controllable Attention Soft Prompt for Few-shot Nested Named Entity Recognition
abstract
Few-shot nested named entity recognition (NER), identifying named entities that are nested with a small number of labeled data, has attracted much attention. Recently, a span-based method based on three stages ( focusing, bridging and prompting) has been proposed for few-shot nested NER. However, such a span-based approach for few-shot nested NER suffers from two challenges: 1) error propagation because of its 3-stage pipeline-based framework; 2) ignoring the relationship between inner and outer entities, which is crucial for few-shot nested NER. Therefore, in this work, we propose a two-stage approach with a controllable attention soft prompt for few-shot nested named entity recognition (TECA). It consists of two components: span part identification and entity mention recognition. The span part identification provides possible entity mentions without an extra filtering module. The entity mention recognition pays fine-grained attention to the inner and outer entities and the corresponding adjacent context through the controllable attention soft prompt to classify the candidate entity mentions. Experimental results show that the TECA approach achieves state-of-the-art performance consistently on the four benchmark datasets (ACE2004, ACE2005, GENIA, and KBP2017) and outperforms several competing baseline models on F1-score by 5.62% on ACE04, 5.11% on ACE05, 3.41% on KBP2017 and 0.7% on GENIA on the 10-shot setting.
Linhai Zhang
LREC/COLING2
2023 Sentiment Analysis on Streaming User Reviews via Dual-Channel Dynamic Graph Neural Network
abstract
Sentiment analysis on user reviews has achieved great success thanks to the rapid growth of deep learning techniques.The large number of online streaming reviews also provides the opportunity to model temporal dynamics for users and products on the timeline.However, existing methods model users and products in the real world based on a static assumption and neglect their time-varying characteristics.In this paper, we present DC-DGNN, a dual-channel framework based on a dynamic graph neural network that models temporal user and product dynamics for sentiment analysis.Specifically, a dual-channel text encoder is employed to extract current local and global contexts from review documents for users and products.Moreover, user review streams are integrated into the dynamic graph neural network by treating users and products as nodes and reviews as new edges.Node representations are dynamically updated along with the evolution of the dynamic graph and used for the final prediction.Experimental results on five real-world datasets demonstrate the superiority of the proposed method.
Xin Zhang 0108, Linhai Zhang
EMNLP2
2022 Pre-training and Fine-tuning Neural Topic Model: A Simple yet Effective Approach to Incorporating External Knowledge
abstract
Recent years have witnessed growing interests in incorporating external knowledge such as pre-trained word embeddings (PWEs) or pretrained language models (PLMs) into neural topic modeling.However, we found that employing PWEs and PLMs for topic modeling only achieved limited performance improvements but with huge computational overhead.In this paper, we propose a novel strategy to incorporate external knowledge into neural topic modeling where the neural topic model is pretrained on a large corpus and then fine-tuned on the target dataset.Experiments have been conducted on three datasets and results show that the proposed approach significantly outperforms both current state-of-the-art neural topic models and some topic modeling approaches enhanced with PWEs or PLMs.Moreover, further study shows that the proposed approach greatly reduces the need for the huge size of training data.
Linhai Zhang, Xuemeng Hu, Qian-Wen Zhang, Yunbo Cao
ACL (1)1
2022 SEE-Few: Seed, Expand and Entail for Few-shot Named Entity Recognition
abstract
Few-shot named entity recognition (NER) aims at identifying named entities based on only few labeled instances. Current few-shot NER methods focus on leveraging existing datasets in the rich-resource domains which might fail in a training-from-scratch setting where no source-domain data is used. To tackle training-from-scratch setting, it is crucial to make full use of the annotation information (the boundaries and entity types). Therefore, in this paper, we propose a novel multi-task (Seed, Expand and Entail) learning framework, SEE-Few, for Few-shot NER without using source domain data. The seeding and expanding modules are responsible for providing as accurate candidate spans as possible for the entailing module. The entailing module reformulates span classification as a textual entailment task, leveraging both the contextual clues and entity type information. All the three modules share the same text encoder and are jointly learned. Experimental results on several benchmark datasets under the training-from-scratch setting show that the proposed method outperformed several state-of-the-art few-shot NER methods with a large margin. Our code is available at https://github.com/unveiled-the-red-hat/SEE-Few.
Zeng Yang, Linhai Zhang
COLING2
2022 Temporal Knowledge Graph Completion with Approximated Gaussian Process Embedding
abstract
Knowledge Graphs (KGs) stores world knowledge that benefits various reasoning-based applications. Due to their incompleteness, a fundamental task for KGs, which is known as Knowledge Graph Completion (KGC), is to perform link prediction and infer new facts based on the known facts. Recently, link prediction on the temporal KGs becomes an active research topic. Numerous Temporal Knowledge Graph Completion (TKGC) methods have been proposed by mapping the entities and relations in TKG to the high-dimensional representations. However, most existing TKGC methods are mainly based on deterministic vector embeddings, which are not flexible and expressive enough. In this paper, we propose a novel TKGC method, TKGC-AGP, by mapping the entities and relations in TKG to the approximations of multivariate Gaussian processes (MGPs). Equipped with the flexibility and capacity of MGP, the global trends as well as the local fluctuations in the TKGs can be simultaneously modeled. Moreover, the temporal uncertainties can be also captured with the kernel function and the covariance matrix of MGP. Moreover, a first-order Markov assumption-based training algorithm is proposed to effective optimize the proposed method. Experimental results show the effectiveness of the proposed approach on two real-world benchmark datasets compared with some state-of-the-art TKGC methods.
Linhai Zhang
COLING1
2021 MERL: Multimodal Event Representation Learning in Heterogeneous Embedding Spaces
abstract
Previous work has shown the effectiveness of using event representations for tasks such as script event prediction and stock market prediction. It is however still challenging to learn the subtle semantic differences between events based solely on textual descriptions of events often represented as (subject, predicate, object) triples. As an alternative, images offer a more intuitive way of understanding event semantics. We observe that event described in text and in images show different abstraction levels and therefore should be projected onto heterogeneous embedding spaces, as opposed to what have been done in previous approaches which project signals from different modalities onto a homogeneous space. In this paper, we propose a Multimodal Event Representation Learning framework (MERL) to learn event representations based on both text and image modalities simultaneously. Event textual triples are projected as Gaussian density embeddings by a dual-path Gaussian triple encoder, while event images are projected as point embeddings by a visual event component-aware image encoder. Moreover, a novel score function motivated by statistical hypothesis testing is introduced to coordinate two embedding spaces. Experiments are conducted on various multimodal event-related tasks and results show that MERL outperforms a number of unimodal and multimodal baselines, demonstrating the effectiveness of the proposed framework.
Linhai Zhang, Yulan He 0001, Zeng Yang
AAAI1
2021 A Neural Group-wise Sentiment Analysis Model with Data Sparsity Awareness
abstract
Sentiment analysis on user-generated content has achieved notable progress by introducing user information to consider each individual’s preference and language usage. However, most existing approaches ignore the data sparsity problem, where the content of some users is limited and the model fails to capture discriminative features of users. To address this issue, we hypothesize that users could be grouped together based on their rating biases as well as degree of rating consistency and the knowledge learned from groups could be employed to analyze the users with limited data. Therefore, in this paper, a neural group-wise sentiment analysis model with data sparsity awareness is proposed. The user-centred document representations are generated by incorporating a group-based user encoder. Furthermore, a multi-task learning framework is employed to jointly modelusers’ rating biases and their degree of rating consistency. One task is vanilla populationlevel sentiment analysis and the other is groupwise sentiment analysis. Experimental results on three real-world datasets show that the proposed approach outperforms some state-of the-art methods. Moreover, model analysis and case study demonstrate its effectiveness of modeling user rating biases and variances.
Linhai Zhang, Yulan He 0001
AAAI3
2021 Beyond Text: Incorporating Metadata and Label Structure for Multi-Label Document Classification using Heterogeneous Graphs
abstract
Multi-label document classification, associating one document instance with a set of relevant labels, is attracting more and more research attention.Existing methods explore the incorporation of information beyond text, such as document metadata or label structure.These approaches however either simply utilize the semantic information of metadata or employ the predefined parent-child label hierarchy, ignoring the heterogeneous graphical structures of metadata and labels, which we believe are crucial for accurate multi-label document classification.Therefore, in this paper, we propose a novel neural network based approach for multi-label document classification, in which two heterogeneous graphs are constructed and learned using heterogeneous graph transformers.One is metadata heterogeneous graph, which models various types of metadata and their topological relations.The other is label heterogeneous graph, which is constructed based on both the labels' hierarchy and their statistical dependencies.Experimental results on two benchmark datasets show the proposed approach outperforms several stateof-the-art baselines.
Chenchen Ye 0002, Linhai Zhang, Yulan He 0001
EMNLP (1)2
2021 Implicit Sentiment Analysis with Event-centered Text Representation
abstract
Implicit sentiment analysis, aiming at detecting the sentiment of a sentence without sentiment words, has become an attractive research topic in recent years.In this paper, we focus on event-centric implicit sentiment analysis that utilizes the sentiment-aware event contained in a sentence to infer its sentiment polarity.Most existing methods in implicit sentiment analysis simply view noun phrases or entities in text as events or indirectly model events with sophisticated models.Since events often trigger sentiments in sentences, we argue that this task would benefit from explicit modeling of events and event representation learning.To this end, we represent an event as the combination of its event type and the event triplet .Based on such event representation, we further propose a novel model with hierarchical tensor-based composition mechanism to detect sentiment in text.In addition, we present a dataset 1 for event-centric implicit sentiment analysis where each sentence is labeled with the event representation described above.Experimental results on our constructed dataset and an existing benchmark dataset show the effectiveness of the proposed approach.
Linhai Zhang, Yulan He 0001
EMNLP (1)3
2021 A Bayesian end-to-end model with estimated uncertainties for simple question answering over knowledge bases
Linhai Zhang, Yulan He 0001
Comput. Speech Lang.1