EDBT 2026 Demo / reviewers in the wild / expert
Hanqi Yan
dblp:254/8174
· DBLP profile ↗
18ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0001-5034-4520ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection ScoreabstractLarge Language Models (LLMs) have shown improved generation performance through retrieval-augmented generation (RAG) following the retriever-reader paradigm, which supplements model inputs with externally retrieved knowledge. However, prior work often evaluates RAG holistically, assessing the retriever and reader jointly, making it difficult to isolate the true contribution of retrieval, particularly given the prompt sensitivity of LLMs used as readers. We move beyond perplexity and introduce Spectrum Projection Score (SPS), a lightweight and supervision-free metric that allows the reader to gauge the semantic alignment of a retrieved summary with its hidden representation by comparing the area formed by generated tokens from the summary, and the principal directions of subspace in the reader and to measure the relevance. Building on SPS we present xCompress, an inference‑time controller framework that dynamically samples, ranks, and compresses retrieval summary candidates. Extensive experiments on five QA benchmarks with four open-sourced LLMs show that SPS not only enhances performance across a range of tasks but also provides a principled perspective on the interaction between retrieval and generation. Zhanghao Hu, Qinglin Zhu, Siya Qi, Yulan He 0001, Hanqi Yan, Lin Gui 0003 |
AAAI | 5 |
| 2025 | Drift: Enhancing LLM Faithfulness in Rationale Generation via Dual-Reward Probabilistic InferenceabstractAs Large Language Models (LLMs) are increasingly applied to complex reasoning tasks, achieving both accurate task performance and faithful explanations becomes crucial.However, LLMs often generate unfaithful explanations, partly because they do not consistently adhere closely to the provided context.Existing approaches to this problem either rely on superficial calibration methods, such as decomposed Chain-of-Thought prompting, or require costly retraining to improve model faithfulness.In this work, we propose a probabilistic inference paradigm that leverages taskspecific and lookahead rewards to ensure that LLM-generated rationales are more faithful to model decisions and align better with input context.These rewards are derived from a domainspecific proposal distribution, allowing for optimized sequential Monte Carlo approximations.Our evaluations across three different reasoning tasks show that this method, which allows for controllable generation during inference, improves both accuracy and faithfulness of LLMs.This method offers a promising path towards making LLMs more reliable for reasoning tasks without sacrificing performance. Jiazheng Li 0002, Hanqi Yan, Yulan He 0001 |
ACL (1) | 2 |
| 2025 | Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question AnsweringabstractLarge language models (LLMs) have recently pushed open-domain question answering (ODQA) to new frontiers.However, prevailing retriever-reader pipelines often depend on multiple rounds of prompt-level instructions, leading to high computational overhead, instability, and suboptimal retrieval coverage.In this paper, we propose EmbQA, an embedding-level framework that alleviates these shortcomings by enhancing both the retriever and the reader.Specifically, we refine query representations via lightweight linear layers under an unsupervised contrastive learning objective, thereby reordering retrieved passages to highlight those most likely to contain correct answers.Additionally, we introduce an exploratory embedding that broadens the model's latent semantic space to diversify candidate generation and employs an entropy-based selection mechanism to choose the most confident answer automatically.Extensive experiments across three opensource LLMs, three retrieval methods, and four ODQA benchmarks demonstrate that EmbQA substantially outperforms recent baselines in both accuracy and efficiency. Zhanghao Hu, Hanqi Yan, Qinglin Zhu, Zhenyi Shen, Yulan He 0001, Lin Gui 0003 |
ACL (1) | 2 |
| 2025 | CODI: Compressing Chain-of-Thought into Continuous Space via Self-DistillationabstractChain-of-Thought (CoT) reasoning enhances Large Language Models (LLMs) by encouraging step-by-step reasoning in natural language.However, leveraging a latent continuous space for reasoning may offer benefits in terms of both efficiency and robustness.Prior implicit CoT methods attempt to bypass language completely by reasoning in continuous space but have consistently underperformed compared to the standard explicit CoT approach.We introduce CODI (Continuous Chain-of-Thought via Self-Distillation), a novel training framework that effectively compresses natural language CoT into continuous space.CODI jointly trains a teacher task (Explicit CoT) and a student task (Implicit CoT), distilling the reasoning ability from language into continuous space by aligning the hidden states of a designated token.Our experiments show that CODI is the first implicit CoT approach to match the performance of explicit CoT on GSM8k at the GPT-2 scale, achieving a 3.1x compression rate and outperforming the previous stateof-the-art by 28.2% in accuracy.CODI also demonstrates robustness, generalizable to complex datasets, and interpretability.These results validate that LLMs can reason effectively not only in natural language, but also in a latent continuous space.Code is available at https://github.com/zhenyi4/codi. Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du 0001, Yulan He 0001 |
EMNLP | 2 |
| 2025 | Constrain Alignment with Sparse AutoencodersabstractThe alignment of large language models (LLMs) with human preferences remains a key challenge. While post-training techniques like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) have achieved notable success, they often experience computational inefficiencies and training instability. In this paper, we propose Feature-level constrained Preference Optimization (FPO), a novel method designed to simplify the alignment process while ensuring stability. FPO leverages pre-trained Sparse Autoencoders (SAEs) and introduces feature-level constraints, allowing for efficient, sparsity-enforced alignment. Our approach enjoys efficiency by using sparse features activated in a well-trained sparse autoencoder and the quality of sequential KL divergence by using the feature-level offline reference. Experimental results on benchmark datasets demonstrate that FPO achieves an above 5% absolute improvement in win rate with much lower computational cost compared to state-of-the-art baselines, making it a promising solution for efficient and controllable LLM alignments. Qingyu Yin, Chak Tou Leong, Minjun Zhu, Hanqi Yan, Qiang Zhang 0026, Yulan He 0001, Wenjie Li 0002, Jun Wang 0012, Yue Zhang 0004, Linyi Yang |
ICML | 5 |
| 2025 | Soft Reasoning: Navigating Solution Spaces in Large Language Models through Controlled Embedding ExplorationabstractLarge Language Models (LLMs) struggle with complex reasoning due to limited diversity and inefficient search. We propose Soft Reasoning, an embedding-based search framework that optimises the embedding of the first token to guide generation. It combines (1) embedding perturbation for controlled exploration and (2) Bayesian optimisation to refine embeddings via a verifier-guided objective, balancing exploration and exploitation. This approach improves reasoning accuracy and coherence while avoiding reliance on heuristic search. Experiments demonstrate superior correctness with minimal computation, making it a scalable, model-agnostic solution. Qinglin Zhu, Runcong Zhao, Hanqi Yan, Yulan He 0001, Lin Gui 0003 |
ICML | 3 |
| 2024 | Mirror: Multiple-perspective Self-Reflection Method for Knowledge-rich ReasoningabstractWhile Large language models (LLMs) have the capability to iteratively reflect on their own outputs, recent studies have observed their struggles with knowledge-rich problems without access to external resources.In addition to the inefficiency of LLMs in self-assessment, we also observe that LLMs struggle to revisit their predictions despite receiving explicit negative feedback.Therefore, We propose Mirror, a Multiple-perspective self-reflection method for knowledge-rich reasoning, to avoid getting stuck at a particular reflection iteration.Mirror enables LLMs to reflect from multipleperspective clues, achieved through a heuristic interaction between a Navigator and a Reasoner.It guides agents toward diverse yet plausibly reliable reasoning trajectory without access to ground truth by encouraging (1) diversity of directions generated by Navigator and (2) agreement among strategically induced perturbations in responses generated by the Reasoner.The experiments on five reasoning datasets demonstrate that Mirror's superiority over several contemporary self-reflection approaches.Additionally, the ablation study studies clearly indicate that our strategies alleviate the aforementioned challenges.The code is released at https://github.com/hanqi-qi/Mirror.git. Hanqi Yan, Qinglin Zhu, Xinyu Wang 0062, Lin Gui 0003, Yulan He 0001 |
ACL (1) | 1 |
| 2024 | Weak Reward Model Transforms Generative Models into Robust Causal Event Extraction SystemsabstractThe inherent ambiguity of cause and effect boundaries poses a challenge in evaluating causal event extraction tasks.Traditional metrics like Exact Match and BertScore poorly reflect model performance, so we trained evaluation models to approximate human evaluation, achieving high agreement.We used them to perform Reinforcement Learning with extraction models to align them with human preference, prioritising semantic understanding.We successfully explored our approach through multiple datasets, including transferring an evaluator trained on one dataset to another as a way to decrease the reliance on human-annotated data.In that vein, we also propose a weak-to-strong supervision method that uses a fraction of the annotated data to train an evaluation model while still achieving high performance in training an RL model. 1 Italo Luis da Silva, Hanqi Yan, Lin Gui 0003, Yulan He 0001 |
EMNLP | 2 |
| 2024 | Encourage or Inhibit Monosemanticity? Revisit Monosemanticity from a Feature Decorrelation PerspectiveabstractTo better interpret the intrinsic mechanism of large language models (LLMs), recent studies focus on monosemanticity on its basic units.A monosemantic neuron is dedicated to a single and specific concept, which forms a oneto-one correlation between neurons and concepts.Despite extensive research in monosemanticity probing, it remains unclear whether monosemanticity is beneficial or harmful to model capacity.To explore this question, we revisit monosemanticity from the feature decorrelation perspective and advocate for its encouragement.We experimentally observe that the current conclusion by Wang et al. ( 2024), which suggests that decreasing monosemanticity enhances model performance, does not hold when the model changes.Instead, we demonstrate that monosemanticity consistently exhibits a positive correlation with model capacity, in the preference alignment process.Consequently, we apply feature correlation as a proxy for monosemanticity and incorporate a feature decorrelation regularizer into the dynamic preference optimization process.The experiments show that our method not only enhances representation diversity but also improves preference alignment performance 1 . Hanqi Yan, Yanzheng Xiang, Guangyi Chen 0002, Lin Gui 0003, Yulan He 0001 |
EMNLP | 1 |
| 2024 | The Mystery of In-Context Learning: A Comprehensive Survey on Interpretation and AnalysisabstractUnderstanding in-context learning (ICL) capability that enables large language models (LLMs) to excel in proficiency through demonstration examples is of utmost importance.This importance stems not only from the better utilization of this capability across various tasks, but also from the proactive identification and mitigation of potential risks, including concerns regarding truthfulness, bias, and toxicity, that may arise alongside the capability.In this paper, we present a thorough survey on the interpretation and analysis of in-context learning.First, we provide a concise introduction to the background and definition of in-context learning.Then, we give an overview of advancements from two perspectives: 1) the theoretical perspective, emphasizing studies on mechanistic interpretability and delving into the mathematical foundations behind ICL; and 2) the empirical perspective, concerning studies that empirically analyze factors associated with ICL.We conclude by discussing open questions and the challenges encountered and, by suggesting potential avenues for future research.We believe that our work establishes the basis for further exploration into the interpretation of incontext learning.To aid this effort, we have created a repository 1 containing resources that will be continually updated.1 https://github.com/zyxnlp Jiazheng Li 0002, Yanzheng Xiang, Hanqi Yan, Lin Gui 0003, Yulan He 0001 |
EMNLP | 4 |
| 2024 | Explainable Recommender With Geometric Information BottleneckabstractExplainable recommender systems can explain their recommendation decisions, enhancing user trust in the systems. Most explainable recommender systems either rely on human-annotated rationales to train models for explanation generation or leverage the attention mechanism to extract important text spans from reviews as explanations. The extracted rationales are often confined to an individual review and may fail to identify the implicit features beyond the review text. To avoid the expensive human annotation process and to generate explanations beyond individual reviews, we propose to incorporate a geometric prior learnt from user-item interactions into a variational network which infers latent factors from user-item reviews. The latent factors from an individual user-item pair can be used for both recommendation and explanation generation, which naturally inherit the global characteristics encoded in the prior knowledge. Experimental results on three e-commerce datasets show that our model significantly improves the interpretability of a variational recommender using the Wasserstein distance while achieving performance comparable to existing content-based recommender systems in terms of recommendation behaviours. Hanqi Yan, Lin Gui 0003, Kun Zhang 0001, Yulan He 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Counterfactual Generation with Identifiability GuaranteesabstractCounterfactual generation lies at the core of various machine learning tasks, including image translation and controllable text generation. This generation process usually requires the identification of the disentangled latent representations, such as content and style, that underlie the observed data. However, it becomes more challenging when faced with a scarcity of paired data and labelling information. Existing disentangled methods crucially rely on oversimplified assumptions, such as assuming independent content and style variables, to identify the latent variables, even though such assumptions may not hold for complex data distributions. For instance, food reviews tend to involve words like “tasty”, whereas movie reviews commonly contain words such as “thrilling” for the same positive sentiment. This problem is exacerbated when data are sampled from multiple domains since the dependence between content and style may vary significantly over domains. In this work, we tackle the domain-varying dependence between the content and the style variables inherent in the counterfactual generation task. We provide identification guarantees for such latent-variable models by leveraging the relative sparsity of the influences from different latent variables. Our theoretical insights enable the development of a doMain AdapTive counTerfactual gEneration model, called (MATTE). Our theoretically grounded framework achieves state-of-the-art performance in unsupervised style transfer tasks, where neither paired data nor style labels are utilized, across four large-scale datasets. Hanqi Yan, Lin Gui 0003, Yuejie Chi, Eric P. Xing, Yulan He 0001, Kun Zhang 0001 |
NeurIPS | 1 |
| 2023 | Tracking Brand-Associated Polarity-Bearing Topics in User ReviewsabstractAbstract Monitoring online customer reviews is important for business organizations to measure customer satisfaction and better manage their reputations. In this paper, we propose a novel dynamic Brand-Topic Model (dBTM) which is able to automatically detect and track brand-associated sentiment scores and polarity-bearing topics from product reviews organized in temporally ordered time intervals. dBTM models the evolution of the latent brand polarity scores and the topic-word distributions over time by Gaussian state space models. It also incorporates a meta learning strategy to control the update of the topic-word distribution in each time interval in order to ensure smooth topic transitions and better brand score predictions. It has been evaluated on a dataset constructed from MakeupAlley reviews and a hotel review dataset. Experimental results show that dBTM outperforms a number of competitive baselines in brand ranking, achieving a good balance of topic coherence and uniqueness, and extracting well-separated polarity-bearing topics across time intervals.1 Runcong Zhao, Lin Gui 0003, Hanqi Yan, Yulan He 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Addressing token uniformity in transformers via singular value transformationabstractToken uniformity is commonly observed in transformer-based models, in which different tokens share a large proportion of similar information after going through stacked multiple self-attention layers in a transformer. In this paper, we propose to use the distribution of singular values of outputs of each transformer layer to characterise the phenomenon of token uniformity and empirically illustrate that a less skewed singular value distribution can alleviate the token uniformity problem. Base on our observations, we define several desirable properties of singular value distributions and propose a novel transformation function for updating the singular values. We show that apart from alleviating token uniformity, the transformation function should preserve the local neighbourhood structure in the original embedding space. Our proposed singular value transformation function is applied to a range of transformer-based language models such as BERT, ALBERT, RoBERTa and DistilBERT, and improved performance is observed in semantic textual similarity evaluation and a range of GLUE tasks. Hanqi Yan, Lin Gui 0003, Wenjie Li 0002, Yulan He 0001 |
UAI | 1 |
| 2022 | Hierarchical Interpretation of Neural Text ClassificationabstractAbstract Recent years have witnessed increasing interest in developing interpretable models in Natural Language Processing (NLP). Most existing models aim at identifying input features such as words or phrases important for model predictions. Neural models developed in NLP, however, often compose word semantics in a hierarchical manner. As such, interpretation by words or phrases only cannot faithfully explain model decisions in text classification. This article proposes a novel Hierarchical Interpretable Neural Text classifier, called HINT, which can automatically generate explanations of model predictions in the form of label-associated topics in a hierarchical manner. Model interpretation is no longer at the word level, but built on topics as the basic semantic unit. Experimental results on both review datasets and news datasets show that our proposed approach achieves text classification results on par with existing state-of-the-art text classifiers, and generates interpretations more faithful to model predictions and better understood by humans than other interpretable neural text classifiers.1 Hanqi Yan, Lin Gui 0003, Yulan He 0001 |
Comput. Linguistics | 1 |
| 2021 | Position Bias Mitigation: A Knowledge-Aware Graph Model for Emotion Cause ExtractionabstractHanqi Yan, Lin Gui, Gabriele Pergola, Yulan He. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hanqi Yan, Lin Gui 0003, Gabriele Pergola, Yulan He 0001 |
ACL/IJCNLP (1) | 1 |
| 2020 | Reinforcement Learning based Recommendation with Graph Convolutional Q-networkabstractReinforcement learning (RL) has been successfully applied to recommender systems. However, the existing RL-based recommendation methods are limited by their unstructured state/action representations. To address this limitation, we propose a novel way that builds high-quality graph-structured states/actions according to the user-item bipartite graph. More specifically, we develop an end-to-end RL agent, termed Graph Convolutional Q-network (GCQN), which is able to learn effective recommendation policies based on the inputs of the proposed graph-structured representations. We show that GCQN achieves significant performance margins over the existing methods, across different datasets and task settings. Yu Lei 0004, Hongbin Pei, Hanqi Yan, Wenjie Li 0002 |
SIGIR | 3 |
| 2019 | LexicalAT: Lexical-Based Adversarial Reinforcement Training for Robust Sentiment ClassificationabstractJingjing Xu, Liang Zhao, Hanqi Yan, Qi Zeng, Yun Liang, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jingjing Xu 0001, Hanqi Yan, Qi Zeng 0001, Yun Liang 0001, Xu Sun 0001 |
EMNLP/IJCNLP (1) | 3 |