EDBT 2026 Demo / reviewers in the wild / expert
Dongfang Li 0002
dblp:98/6118-2
· DBLP profile ↗
22ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 8 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improving Value-based Process Verifier via Low-Cost Variance ReductionabstractLarge language models (LLMs) have achieved remarkable success in a wide range of tasks. However, their reasoning capabilities, particularly in complex domains like mathematics, remain a significant challenge. Value-based process verifiers, which estimate the probability of a partial reasoning chain leading to a correct solution, are a promising approach for improving reasoning. Nevertheless, their effectiveness is often hindered by estimation error in their training annotations, a consequence of the limited number of Monte Carlo (MC) samples feasible due to the high cost of LLM inference. In this paper, we identify that the estimation error primarily arises from high variance rather than bias, and the MC estimator is a Minimum Variance Unbiased Estimator (MVUE). To address the problem, we propose the Compound Monte Carlo Sampling (ComMCS) method, which constructs an unbiased estimator by linearly combining the MC estimators from the current and subsequent steps. Theoretically, we show that our method leads to a predictable reduction in variance, while maintaining an unbiased estimation without additional LLM inference cost. We also perform empirical experiments on the MATH-500 and GSM8K benchmarks to demonstrate the effectiveness of our method. Notably, ComMCS outperforms regression-based optimization method by 2.8 points, the non-variance-reduced baseline by 2.2 points on MATH-500 on Best-of-32 sampling experiment. Zetian Sun, Dongfang Li 0002, Baotian Hu, Min Zhang 0005 |
AAAI | 2 |
| 2026 | Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement LearningabstractLarge Language Models (LLMs) face severe challenges in long-context processing, including quadratic computational costs, information forgetting, and the context fragmentation inherent in Retrieval-Augmented Generation (RAG).We introduce LycheeMemory, a cognitively inspired framework that enables efficient long-context inference via chunk-wise compression and selective memory recall, rather than processing all raw tokens.LycheeMemory segments the input into chunks and encodes each into compressed KV-cache-style representations using a Compressor.A Gate then dynamically selects relevant memory blocks, which a Reasoner iteratively processes with an evolving working memory to solve downstream tasks.The Compressor and Reasoner are jointly optimized via end-to-end reinforcement learning, while the Gate is trained separately as a classifier.Experimental results demonstrate that LycheeMemory achieves competitive accuracy (up to 82% in ablation variants) on multi-hop reasoning benchmarks (e.g., RULER-HQA), successfully extrapolates context length from 7K to 1.75M, and provides a favorable accuracy-efficiency trade-off against strong long-context baselines.Notably, compared to MemAgent, LycheeMemory achieves an average 2× reduction in peak GPU memory usage and a 6× speedup during inference. Zhuoen Chen, Dongfang Li 0002, Meishan Zhang, Baotian Hu, Min Zhang 0005 |
ACL (1) | 2 |
| 2026 | Structured Episodic Event MemoryabstractCurrent approaches to memory in Large Language Models (LLMs) predominantly rely on static Retrieval-Augmented Generation (RAG), which often results in scattered retrieval and fails to capture the structural dependencies required for complex reasoning.For autonomous agents, these passive and flat architectures lack the cognitive organization necessary to model the dynamic and associative nature of longterm interaction.To address this, we propose Structured Episodic Event Memory (SEEM), a hierarchical framework that synergizes a graph memory layer for relational facts with a dynamic episodic memory layer for narrative progression.Grounded in cognitive frame theory, SEEM transforms interaction streams into structured Episodic Event Frames (EEFs) anchored by precise provenance pointers.Furthermore, we introduce an agentic associative fusion and Reverse Provenance Expansion (RPE) mechanism to reconstruct coherent narrative contexts from fragmented evidence.Experimental results on the LoCoMo and Long-MemEval benchmarks demonstrate that SEEM significantly outperforms baselines, enabling agents to maintain superior narrative coherence and logical consistency. Zhengxuan Lu, Dongfang Li 0002, Yukun Shi, Beilun Wang, Longyue Wang, Baotian Hu |
ACL (1) | 2 |
| 2025 | CMT: A Memory Compression Method for Continual Knowledge Learning of Large Language ModelsabstractLarge Language Models (LLMs) need to adapt to the continuous changes in data, tasks, and user preferences. Due to their massive size and the high costs associated with training, LLMs are not suitable for frequent retraining. However, updates are necessary to keep them in sync with rapidly evolving human knowledge. To address these challenges, this paper proposes the Compression Memory Training (CMT) method, an efficient and effective online adaptation framework for LLMs that features robust knowledge retention capabilities. Inspired by human memory mechanisms, CMT compresses and extracts information from new documents to be stored in a memory bank. When answering to queries related to these new documents, the model aggregates these document memories from the memory bank to better answer user questions. The parameters of the LLM itself do not change during training and inference, reducing the risk of catastrophic forgetting. To enhance the encoding, retrieval, and aggregation of memory, we further propose three new general and flexible techniques, including memory-aware objective, self-matching and top-k aggregation. Extensive experiments conducted on three continual learning datasets (i.e., StreamingQA, SQuAD and ArchivalQA) demonstrate that the proposed method improves model adaptability and robustness across multiple base LLMs (e.g., +4.07 EM & +4.19 F1 in StreamingQA with Llama-2-7b). Dongfang Li 0002, Zetian Sun, Xinshuo Hu, Baotian Hu, Min Zhang 0005 |
AAAI | 1 |
| 2024 | Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module OperationabstractLarge language models (LLMs) have been widely used in various applications but are known to suffer from issues related to untruthfulness and toxicity. While parameter-efficient modules (PEMs) have demonstrated their effectiveness in equipping models with new skills, leveraging PEMs for deficiency unlearning remains underexplored. In this work, we propose a PEMs operation approach, namely Extraction-before-Subtraction (Ext-Sub), to enhance the truthfulness and detoxification of LLMs through the integration of ``expert'' PEM and ``anti-expert'' PEM. Remarkably, even anti-expert PEM possess valuable capabilities due to their proficiency in generating fabricated content, which necessitates language modeling and logical narrative competence. Rather than merely negating the parameters, our approach involves extracting and eliminating solely the deficiency capability within anti-expert PEM while preserving the general capabilities. To evaluate the effectiveness of our approach in terms of truthfulness and detoxification, we conduct extensive experiments on LLMs, encompassing additional abilities such as language modelling and mathematical reasoning. Our empirical results demonstrate that our approach effectively improves truthfulness and detoxification, while largely preserving the fundamental abilities of LLMs. Xinshuo Hu, Dongfang Li 0002, Baotian Hu, Min Zhang 0005 |
AAAI | 2 |
| 2024 | Temporal Knowledge Question Answering via Abstract Reasoning InductionabstractIn this study, we address the challenge of enhancing temporal knowledge reasoning in Large Language Models (LLMs).LLMs often struggle with this task, leading to the generation of inaccurate or misleading responses.This issue mainly arises from their limited ability to handle evolving factual knowledge and complex temporal logic.To overcome these limitations, we propose Abstract Reasoning Induction (ARI) framework, which divides temporal reasoning into two distinct phases: Knowledgeagnostic and Knowledge-based.This framework offers factual knowledge support to LLMs while minimizing the incorporation of extraneous noisy data.Concurrently, informed by the principles of constructivism, ARI provides LLMs the capability to engage in proactive, self-directed learning from both correct and incorrect historical reasoning samples.By teaching LLMs to actively construct knowledge and methods, it can significantly boosting their temporal reasoning abilities.Our approach achieves significant improvements, with relative gains of 29.7% and 9.27% on two temporal QA datasets, underscoring its efficacy in advancing temporal reasoning in LLMs.The code can be found at https: //github.com/czy1999/ARI-QA. Dongfang Li 0002, Xiang Zhao 0002, Baotian Hu, Min Zhang 0005 |
ACL (1) | 2 |
| 2024 | Does the Generator Mind Its Contexts? An Analysis of Generative Model Faithfulness under Context Transferabstracthe present study introduces the knowledge-augmented generator, which is specifically designed to produce information that remains grounded in contextual knowledge, regardless of alterations in the context. Previous research has predominantly focused on examining hallucinations stemming from static input, such as in the domains of summarization or machine translation. However, our investigation delves into the faithfulness of generative question answering in the presence of dynamic knowledge. Our objective is to explore the existence of hallucinations arising from parametric memory when contextual knowledge undergoes changes, while also analyzing the underlying causes for their occurrence. In order to efficiently address this issue, we propose a straightforward yet effective measure for detecting such hallucinations. Intriguingly, our investigation uncovers that all models exhibit a tendency to generate previous answers as hallucinations. To gain deeper insights into the underlying causes of this phenomenon, we conduct a series of experiments that verify the critical role played by context in hallucination, both during training and testing, from various perspectives. Xinshuo Hu, Dongfang Li 0002, Yuxiang Wu, Lifeng Shang, Baotian Hu |
LREC/COLING | 2 |
| 2024 | Take Off the Training Wheels! Progressive In-Context Learning for Effective AlignmentabstractRecent studies have explored the working mechanisms of In-Context Learning (ICL).However, they mainly focus on classification and simple generation tasks, limiting their broader application to more complex generation tasks in practice.To address this gap, we investigate the impact of demonstrations on token representations within the practical alignment tasks.We find that the transformer embeds the task function learned from demonstrations into the separator token representation, which plays an important role in the generation of prior response tokens.Once the prior response tokens are determined, the demonstrations become redundant.Motivated by this finding, we propose an efficient Progressive In-Context Alignment (PICA) method consisting of two stages.In the first few-shot stage, the model generates several prior response tokens via standard ICL while concurrently extracting the ICL vector that stores the task function from the separator token representation.In the following zero-shot stage, this ICL vector guides the model to generate responses without further demonstrations.Extensive experiments demonstrate that our PICA not only surpasses vanilla ICL but also achieves comparable performance to other alignment tuning methods.The proposed training-free method reduces the time cost (e.g., 5.45×) with improved alignment performance (e.g., 6.57+).Consequently, our work highlights the application of ICL for alignment and calls for a deeper understanding of ICL for complex generations. Dongfang Li 0002, Xinshuo Hu, Xinping Zhao, Yibin Chen, Baotian Hu, Min Zhang 0005 |
EMNLP | 2 |
| 2024 | SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented GenerationabstractRecent studies in Retrieval-Augmented Generation (RAG) have investigated extracting evidence from retrieved passages to reduce computational costs and enhance the final RAG performance, yet it remains challenging.Existing methods heavily rely on heuristic-based augmentation, encountering several issues: (1) Poor generalization due to hand-crafted context filtering; (2) Semantics deficiency due to rulebased context chunking; (3) Skewed length due to sentence-wise filter learning.To address these issues, we propose a model-based evidence extraction learning framework, SEER, optimizing a vanilla model as an evidence extractor with desired properties through selfaligned learning.Extensive experiments show that our method largely improves the final RAG performance, enhances the faithfulness, helpfulness, and conciseness of the extracted evidence, and reduces the evidence length by 9.25 times.The code will be available at https://github.com/HITsz-TMG/SEER. Xinping Zhao, Dongfang Li 0002, Boren Hu, Yibin Chen, Baotian Hu, Min Zhang 0005 |
EMNLP | 2 |
| 2024 | In-Context Learning State Vector with Inner and Momentum OptimizationabstractLarge Language Models (LLMs) have exhibited an impressive ability to perform In-Context Learning (ICL) from only a few examples. Recent works have indicated that the functions learned by ICL can be represented through compressed vectors derived from the transformer. However, the working mechanisms and optimization of these vectors are yet to be thoroughly explored. In this paper, we address this gap by presenting a comprehensive analysis of these compressed vectors, drawing parallels to the parameters trained with gradient descent, and introducing the concept of state vector. Inspired by the works on model soup and momentum-based gradient descent, we propose inner and momentum optimization methods that are applied to refine the state vector progressively as test-time adaptation. Moreover, we simulate state vector aggregation in the multiple example setting, where demonstrations comprising numerous examples are usually too lengthy for regular ICL, and further propose a divide-and-conquer aggregation method to address this challenge. We conduct extensive experiments using Llama-2 and GPT-J in both zero-shot setting and few-shot setting. The experimental results show that our optimization method effectively enhances the state vector and achieves the state-of-the-art performance on diverse tasks. Dongfang Li 0002, Xinshuo Hu, Zetian Sun, Baotian Hu, Min Zhang 0005 |
NeurIPS | 1 |
| 2024 | SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-ReflectionabstractInstruction tuning (IT) is crucial to tailoring large language models (LLMs) towards human-centric interactions. Recent advancements have shown that the careful selection of a small, high-quality subset of IT data can significantly enhance the performance of LLMs. Despite this, common approaches often rely on additional models or data, which increases costs and limits widespread adoption. In this work, we propose a novel approach, termed $\textit{SelectIT}$, that capitalizes on the foundational capabilities of the LLM itself. Specifically, we exploit the intrinsic uncertainty present in LLMs to more effectively select high-quality IT data, without the need for extra resources. Furthermore, we introduce a curated IT dataset, the $\textit{Selective Alpaca}$, created by applying SelectIT to the Alpaca-GPT4 dataset. Empirical results demonstrate that IT using Selective Alpaca leads to substantial model ability enhancement. The robustness of SelectIT has also been corroborated in various foundation models and domain-specific tasks. Our findings suggest that longer and more computationally intensive IT data may serve as superior sources of IT, offering valuable insights for future research in this area. Data, code, and scripts are freely available at https://github.com/Blue-Raincoat/SelectIT. Liangxin Liu, Xuebo Liu 0002, Derek F. Wong, Dongfang Li 0002, Baotian Hu, Min Zhang 0005 |
NeurIPS | 4 |
| 2022 | Diaformer: Automatic Diagnosis via Symptoms Sequence GenerationabstractAutomatic diagnosis has attracted increasing attention but remains challenging due to multi-step reasoning. Recent works usually address it by reinforcement learning methods. However, these methods show low efficiency and require task-specific reward functions. Considering the conversation between doctor and patient allows doctors to probe for symptoms and make diagnoses, the diagnosis process can be naturally seen as the generation of a sequence including symptoms and diagnoses. Inspired by this, we reformulate automatic diagnosis as a symptoms Sequence Generation (SG) task and propose a simple but effective automatic Diagnosis model based on Transformer (Diaformer). We firstly design the symptom attention framework to learn the generation of symptom inquiry and the disease diagnosis. To alleviate the discrepancy between sequential generation and disorder of implicit symptoms, we further design three orderless training mechanisms. Experiments on three public datasets show that our model outperforms baselines on disease diagnosis by 1%, 6% and 11.5% with the highest training efficiency. Detailed analysis on symptom inquiry prediction demonstrates that the potential of applying symptoms sequence generation for automatic diagnosis. Dongfang Li 0002, Qingcai Chen, Wenxiu Zhou, Xin Liu 0054 |
AAAI | 2 |
| 2022 | Unifying Model Explainability and Robustness for Joint Text Classification and Rationale ExtractionabstractRecent works have shown explainability and robustness are two crucial ingredients of trustworthy and reliable text classification. However, previous works usually address one of two aspects: i) how to extract accurate rationales for explainability while being beneficial to prediction; ii) how to make the predictive model robust to different types of adversarial attacks. Intuitively, a model that produces helpful explanations should be more robust against adversarial attacks, because we cannot trust the model that outputs explanations but changes its prediction under small perturbations. To this end, we propose a joint classification and rationale extraction model named AT-BMC. It includes two key mechanisms: mixed Adversarial Training (AT) is designed to use various perturbations in discrete and embedding space to improve the model’s robustness, and Boundary Match Constraint (BMC) helps to locate rationales more precisely with the guidance of boundary information. Performances on benchmark datasets demonstrate that the proposed AT-BMC outperforms baselines on both classification and rationale extraction by a large margin. Robustness analysis shows that the proposed AT-BMC decreases the attack success rate effectively by up to 69%. The results indicate that there are connections between robust models and better explanations. Dongfang Li 0002, Baotian Hu, Qingcai Chen, Tujie Xu, Jingcong Tao, Yunan Zhang 0003 |
AAAI | 1 |
| 2022 | Prompt-based Text Entailment for Low-Resource Named Entity RecognitionabstractPre-trained Language Models (PLMs) have been applied in NLP tasks and achieve promising results. Nevertheless, the fine-tuning procedure needs labeled data of the target domain, making it difficult to learn in low-resource and non-trivial labeled scenarios. To address these challenges, we propose Prompt-based Text Entailment (PTE) for low-resource named entity recognition, which better leverages knowledge in the PLMs. We first reformulate named entity recognition as the text entailment task. The original sentence with entity type-specific prompts is fed into PLMs to get entailment scores for each candidate. The entity type with the top score is then selected as final label. Then, we inject tagging labels into prompts and treat words as basic units instead of n-gram spans to reduce time complexity in generating candidates by n-grams enumeration. Experimental results demonstrate that the proposed method PTE achieves competitive performance on the CoNLL03 dataset, and better than fine-tuned counterparts on the MIT Movie and Few-NERD dataset in low-resource settings. Dongfang Li 0002, Baotian Hu, Qingcai Chen |
COLING | 1 |
| 2022 | Calibration Meets Explanation: A Simple and Effective Approach for Model Confidence EstimatesabstractCalibration strengthens the trustworthiness of black-box models by producing better accurate confidence estimates on given examples.However, little is known about if model explanations can help confidence calibration.Intuitively, humans look at important features attributions and decide whether the model is trustworthy.Similarly, the explanations can tell us when the model may or may not know.Inspired by this, we propose a method named CME that leverages model explanations to make the model less confident with non-inductive attributions.The idea is that when the model is not highly confident, it is difficult to identify strong indications of any class, and the tokens accordingly do not have high attribution scores for any class and vice versa.We conduct extensive experiments on six datasets with two popular pre-trained language models in the in-domain and out-of-domain settings.The results show that CME improves calibration performance in all settings.The expected calibration errors are further reduced when combined with temperature scaling.Our findings highlight that model explanations can help calibrate posterior estimates. Dongfang Li 0002, Baotian Hu, Qingcai Chen |
EMNLP | 1 |
| 2021 | FHTC: Few-Shot Hierarchical Text Classification in Financial Domain
Qingcai Chen, Dongfang Li 0002 |
ICONIP (2) | 3 |
| 2021 | MSDF: A General Open-Domain Multi-skill Dialog Framework
Yu Zhao 0043, Xinshuo Hu, Yunxin Li, Baotian Hu, Dongfang Li 0002, Sichao Chen, Xiaolong Wang 0001 |
NLPCC (2) | 5 |
| 2021 | Attentive capsule network for click-through rate and conversion rate prediction in online advertising
Dongfang Li 0002, Baotian Hu, Qingcai Chen, Quanchang Qi, Liubin Wang, Haishan Liu |
Knowl. Based Syst. | 1 |
| 2020 | Towards Medical Machine Reading Comprehension with Structural Knowledge and Plain TextabstractMachine reading comprehension (MRC) has achieved significant progress on the open domain in recent years, mainly due to large-scale pre-trained language models.However, it performs much worse in specific domains such as the medical field due to the lack of extensive training data and professional structural knowledge neglect.As an effort, we first collect a large scale medical multi-choice question dataset (more than 21k instances) for the National Licensed Pharmacist Examination in China.It is a challenging medical examination with a passing rate of less than 14.2% in 2018.Then we propose a novel reading comprehension model KMQA, which can fully exploit the structural medical knowledge (i.e., medical knowledge graph) and the reference medical plain text (i.e., text snippets retrieved from reference books).The experimental results indicate that the KMQA outperforms existing competitive models with a large margin and passes the exam with 61.8% accuracy rate on the test set. Dongfang Li 0002, Baotian Hu, Qingcai Chen, Weihua Peng |
EMNLP (1) | 1 |
| 2020 | Gated Semantic Difference Based Sentence Semantic Equivalence IdentificationabstractThis article proposes a novel sentence semantic equivalence identification (SSEI) method by using the semantic difference features between sentences. The lexical differences of a sentence pair are first extracted, and the bidirectional long short term memory (BiLSTM) network is then applied on them to generate the semantic difference representations. Finally, an efficient gate mechanism is proposed to integrate the semantic differences with existing models (called base model) to enhance their encoding capability in the SSEI task. Exhaustive experiments conducted on the standard Quora corpus, and the Large-scale Chinese Question Matching Corpus (LCQMC) show that the proposed gated semantic difference (GSD) method brings significant improvement for different existing state-of-the-art models. When the bidirectional encoder representations from transformers model (BERT) is used as the base model, the accuracy for SSEI on Quora is improved from 90.63% to 91.98%, and the F1 score on the LCQMC is improved from 87.0% to 87.7%, which outperforms the best-published results. Xin Liu 0054, Qingcai Chen, Xiangping Wu 0001, Yang Hua 0004, Dongfang Li 0002, Buzhou Tang, Xiaolong Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2019 | Multi-strategies Method for Cold-Start Stage Question Matching of rQA Task
Dongfang Li 0002, Qingcai Chen, Songjian Chen, Xin Liu 0054, Buzhou Tang, Ben Tan |
NLPCC (1) | 1 |
| 2018 | LCQMC: A Large-scale Chinese Question Matching CorpusabstractThe lack of large-scale question matching corpora greatly limits the development of matching methods in question answering (QA) system, especially for non-English languages. To ameliorate this situation, in this paper, we introduce a large-scale Chinese question matching corpus (named LCQMC), which is released to the public1. LCQMC is more general than paraphrase corpus as it focuses on intent matching rather than paraphrase. How to collect a large number of question pairs in variant linguistic forms, which may present the same intent, is the key point for such corpus construction. In this paper, we first use a search engine to collect large-scale question pairs related to high-frequency words from various domains, then filter irrelevant pairs by the Wasserstein distance, and finally recruit three annotators to manually check the left pairs. After this process, a question matching corpus that contains 260,068 question pairs is constructed. In order to verify the LCQMC corpus, we split it into three parts, i.e., a training set containing 238,766 question pairs, a development set with 8,802 question pairs, and a test set with 12,500 question pairs, and test several well-known sentence matching methods on it. The experimental results not only demonstrate the good quality of LCQMC but also provide solid baseline performance for further researches on this corpus. Xin Liu 0054, Qingcai Chen, Chong Deng, Hua-Jun Zeng, Dongfang Li 0002, Buzhou Tang |
COLING | 6 |