EDBT 2026 Demo / reviewers in the wild / expert
Xueliang Zhao
dblp:14/7742
· DBLP profile ↗
22ranked-venue papers
9as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Retrieval-Augmented Generation for Biomedical Question Answering Based on Hierarchical Knowledge TreeabstractBiomedical literature is characterized by vast, structurally complex, and rapidly evolving knowledge that poses significant challenges for conventional retrieval-augmented generation (RAG) methods. Existing approaches typically rely on fixed-length chunking strategies that disrupt semantic coherence and fail to capture hierarchical relationships in biomedical texts, limiting their effectiveness in multi-hop reasoning tasks. To address these limitations, we propose Bio-RAPTOR, a hierarchical RAG framework specifically designed to accommodate the distinct structure and semantics of biomedical texts. Our approach integrates semantic chunking, recursive clustering with HDBSCAN, and hierarchical summarization to construct hierarchically organized semantic structures from biomedical documents. During inference, we employ HyDE-based query expansion and hybrid retrieval with reranking to retrieve relevant information across different abstraction levels. Experimental results on authoritative biomedical QA datasets demonstrate that Bio-RAPTOR significantly outperforms both standard RAG and the original RAPTOR framework in the biomedical question answering. Xueliang Zhao, Xiaoguang Lin, Qilong Sun |
BIBM | 2 |
| 2025 | Forewarned is Forearmed: Harnessing LLMs for Data Synthesis via Failure-induced ExplorationabstractLarge language models (LLMs) have significantly benefited from training on diverse, high-quality task-specific data, leading to impressive performance across a range of downstream applications. Current methods often rely on human-annotated data or predefined task templates to direct powerful LLMs in synthesizing task-relevant data for effective model training. However, this dependence on manually designed components may constrain the scope of generated data, potentially overlooking critical edge cases or novel scenarios that could challenge the model. In this paper, we present a novel approach, ReverseGen, designed to automatically generate effective training samples that expose the weaknesses of LLMs. Specifically, we introduce a dedicated proposer trained to produce queries that lead target models to generate unsatisfactory responses. These failure-inducing queries are then used to construct training data, helping to address the models' shortcomings and improve overall performance. Our approach is flexible and can be applied to models of various scales (3B, 7B, and 8B). We evaluate ReverseGen on three key applications—safety, honesty, and math—demonstrating that our generated data is both highly effective and diverse. Models fine-tuned with ReverseGen-generated data consistently outperform those trained on human-annotated or general model-generated data, offering a new perspective on data synthesis for task-specific LLM enhancement. Qintong Li, Jiahui Gao 0002, Renjie Pi, Xueliang Zhao, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong |
ICLR | 5 |
| 2025 | DynaAct: Large Language Model Reasoning with Dynamic Action SpacesabstractIn modern sequential decision-making systems, the construction of an optimal candidate action space is critical to efficient inference. However, existing approaches either rely on manually defined action spaces that lack scalability or utilize unstructured spaces that render exhaustive search computationally prohibitive. In this paper, we propose a novel framework named \textsc{DynaAct} for automatically constructing a compact action space to enhance sequential reasoning in complex problem-solving scenarios. Our method first estimates a proxy for the complete action space by extracting general sketches observed in a corpus covering diverse complex reasoning problems using large language models. We then formulate a submodular function that jointly evaluates candidate actions based on their utility to the current state and their diversity, and employ a greedy algorithm to select an optimal candidate set.
Extensive experiments on six diverse standard benchmarks demonstrate that our approach significantly improves overall performance, while maintaining efficient inference without introducing substantial latency. The implementation is available at \url{https://github.com/zhaoxlpku/DynaAct}. Xueliang Zhao, Wei Wu 0014, Jian Guan 0002, Qintong Li, Lingpeng Kong |
NeurIPS | 1 |
| 2024 | GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem SolversabstractLarge language models (LLMs) have achieved impressive performance across various mathematical reasoning benchmarks.However, there are increasing debates regarding whether these models truly understand and apply mathematical knowledge or merely rely on shortcuts for mathematical reasoning.One essential and frequently occurring evidence is that when the math questions are slightly changed, LLMs can behave incorrectly.This motivates us to evaluate the robustness of LLMs' math reasoning capability by testing a wide range of question variations.We introduce the adversarial grade school math (GSM-PLUS) dataset, an extension of GSM8K augmented with various mathematical perturbations.Our experiments on 25 LLMs and 4 prompting techniques show that while LLMs exhibit different levels of math reasoning abilities, their performances are far from robust.In particular, even for problems that have been solved in GSM8K, LLMs can make mistakes when new statements are added or the question targets are altered.We also explore whether more robust performance can be achieved by composing existing prompting methods, in which we try an iterative method that generates and verifies each intermediate thought based on its reasoning goal and calculation result. Qintong Li, Leyang Cui, Xueliang Zhao, Lingpeng Kong, Wei Bi |
ACL (1) | 3 |
| 2024 | SEGO: Sequential Subgoal Optimization for Mathematical Problem-SolvingabstractLarge Language Models (LLMs) have driven substantial progress in artificial intelligence in recent years, exhibiting impressive capabilities across a wide range of tasks, including mathematical problem-solving.Inspired by the success of subgoal-based methods, we propose a novel framework called SEquential subGoal Optimization (SEGO) to enhance LLMs' ability to solve mathematical problems.By establishing a connection between the subgoal breakdown process and the probability of solving problems, SEGO aims to identify better subgoals with theoretical guarantees.Addressing the challenge of identifying suitable subgoals in a large solution space, our framework generates problem-specific subgoals and adjusts them according to carefully designed criteria.Incorporating these optimized subgoals into the policy model training leads to significant improvements in problem-solving performance.We validate SEGO's efficacy through experiments on two benchmarks, GSM8K and MATH, where our approach outperforms existing methods, highlighting the potential of SEGO in AI-driven mathematical problemsolving. * This work Xueliang Zhao, Xinting Huang, Wei Bi, Lingpeng Kong |
ACL (1) | 1 |
| 2024 | Subgoal-based Demonstration Learning for Formal Theorem ProvingabstractLarge language models (LLMs) present a promising pathway for advancing the domain of formal theorem proving. In this paper, we aim to improve the performance of LLMs in formal theorem proving by thoroughly examining the structure and organization of demonstrative in-context examples. We introduce a subgoal-based demonstration learning framework, specifically designed to enhance the efficiency of proof search in LLMs. First, drawing upon the insights of subgoal learning from reinforcement learning and robotics, we propose the construction of distinct subgoals for each demonstration example and refine these subgoals in accordance with the pertinent theories of subgoal learning. Second, we build upon recent advances in diffusion models to predict the optimal organization, simultaneously addressing two intricate issues that persist within the domain of demonstration organization: subset selection and order determination. Our integration of subgoal-based learning has notably increased proof accuracy from 38.9% to 44.1% on the miniF2F benchmark. Furthermore, the adoption of diffusion models for demonstration organization can lead to an additional enhancement in accuracy to 45.5%, or a $5\times$ improvement in sampling efficiency compared to previously established methods. Xueliang Zhao, Wenda Li 0001, Lingpeng Kong |
ICML | 1 |
| 2024 | An overlap estimation guided feature metric approach for real point cloud registration
Fukai Zhang, Tiancheng He, Yiran Sun, Shan Zhao 0009, Xueliang Zhao, Weiye Zhao |
Comput. Graph. | 7 |
| 2023 | On the Compositional Generalization in Versatile Open-domain DialogueabstractPrevious research has demonstrated the potential of multi-task learning to foster a conversational agent's ability to acquire a variety of skills.However, these approaches either suffer from interference among different datasets (also known as negative transfer), or fail to effectively reuse knowledge and skills learned from other datasets.In contrast to previous works, we develop a sparsely activated modular network: (1) We propose a wellrounded set of operators and instantiate each operator with an independent module; (2) We formulate dialogue generation as the execution of a generated programme which recursively composes and assembles modules.Extensive experiments on 9 datasets verify the efficacy of our methods through automatic evaluation and human evaluation.Notably, our model outperforms state-of-the-art supervised approaches on 4 datasets with only 10% training data thanks to the modular architecture and multi-task learning.1 † Tingchen Fu and Xueliang Zhao contribute equally to this work. Tingchen Fu, Xueliang Zhao, Lemao Liu, Rui Yan 0001 |
ACL (1) | 2 |
| 2023 | VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic TransitionsabstractYeah, we're pretty lousy with pens around here, so knock yourself out.Really? Thanks.Here are a few notes on Yuxuan Wang 0004, Zilong Zheng, Xueliang Zhao, Jinpeng Li 0003, Yueqian Wang, Dongyan Zhao 0001 |
ACL (1) | 3 |
| 2023 | Delving into Global Dialogue Structures: Structure Planning Augmented Response Selection for Multi-turn ConversationsabstractRetrieval-based dialogue systems are a crucial component of natural language processing, employing information retrieval techniques to select responses from a predefined pool of candidates. The advent of pre-trained language models (PLMs) has significantly advanced the field, with a prevailing paradigm that involves post-training PLMs on specific dialogue corpora, followed by fine-tuning for the response selection (RS) task. This post-training process aims to capture dialogue-specific features, as most PLMs are originally trained on plain text. However, prior approaches predominantly rely on self-supervised tasks or session-level graph neural networks during post-training, focusing on capturing underlying patterns of coherent dialogues without explicitly refining the global pattern across the entire dialogue corpus. Consequently, the learned knowledge for organizing coherent dialogues remains isolated, heavily reliant on specific contexts. Additionally, interpreting or visualizing the implicit knowledge acquired through self-supervised tasks proves challenging. In this study, we address these limitations by explicitly refining the knowledge required for response selection and structuring it into a coherent global flow, known as "dialogue structure." This structure captures the inter-dependency of utterances and topic shifts, thereby enhancing the response selection task. To achieve this, we propose a novel structure model comprising a state recognizer and a structure planner. This model effectively captures the flow within the utterance history and plans the trajectory of future utterances. Importantly, the structure model operates orthogonally to the retrieval model, enabling seamless integration with existing retrieval models and facilitating collaborative training. Extensive experiments conducted on three benchmark datasets demonstrate the superior performance of our method over a wide range of competitive baselines, establishing a new state-of-the-art in the field. Tingchen Fu, Xueliang Zhao, Rui Yan 0001 |
KDD | 2 |
| 2023 | Learning Multi-turn Response Selection in Grounded Dialogues with Reinforced Knowledge and Context DistillationabstractRecently, knowledge-grounded dialogue systems have gained increasing attention. Great efforts have been made to build response matching models where all dialogue content and knowledge sentences are leveraged. However, knowledge redundancy and distraction of irrelevant dialogue content often exist in knowledge-grounded conversations, which may affect the matching process and lead to inferior performance. In addition, irrelevant dialogue history and excessive knowledge also hinder the exploitation of popular pre-trained language models (PLMs) due to the limitation of input length. To address these challenges, we propose a new knowledge-grounded dialogue model based on PLMs, where a knowledge selector and a context selector are designed for filtering out irrelevant knowledge sentences and redundant dialogue history, respectively. Considering the lack of labeled data for the learning of two selectors, we pre-train them with weakly-supervised tasks and then jointly conduct the optimization of knowledge and context selection and fine-tuning of PLMs for response ranking with reinforcement learning (RL). By this means, the dialogue model can distill more accurate and concise knowledge and dialogue content for subsequent response ranking module, and the overall model can converge and perform better. We conduct experiments on two benchmarks and evaluation results indicate that our model can significantly outperform the state-of-the-art methods. Jiazhan Feng, Chongyang Tao, Xueliang Zhao, Dongyan Zhao 0001 |
ACM Trans. Inf. Syst. | 3 |
| 2022 | There Are a Thousand Hamlets in a Thousand People's Eyes: Enhancing Knowledge-grounded Dialogue with Personal MemoryabstractKnowledge-grounded conversation (KGC) shows great potential in building an engaging and knowledgeable chatbot, and knowledge selection is a key ingredient in it.However, previous methods for knowledge selection only concentrate on the relevance between knowledge and dialogue context, ignoring the fact that age, hobby, education and life experience of an interlocutor have a major effect on his or her personal preference over external knowledge.Without taking the personalization issue into account, it is difficult to select the proper knowledge and generate persona-consistent responses.In this work, we introduce personal memory into knowledge selection in KGC to address the personalization issue.We propose a variational method to model the underlying relationship between one's personal memory and his or her selection of knowledge, and devise a learning scheme in which the forward mapping from personal memory to knowledge and its inverse mapping is included in a closed loop so that they could teach each other.Experiment results show that our method outperforms existing KGC methods significantly on both automatic evaluation and human evaluation. Tingchen Fu, Xueliang Zhao, Chongyang Tao, Ji-Rong Wen, Rui Yan 0001 |
ACL (1) | 2 |
| 2022 | There Is No Standard Answer: Knowledge-Grounded Dialogue Generation with Adversarial Activated Multi-Reference LearningabstractKnowledge-grounded conversation (KGC) shows excellent potential to deliver an engaging and informative response.However, existing approaches emphasize selecting one golden knowledge given a particular dialogue context, overlooking the one-to-many phenomenon in dialogue.As a result, the existing paradigm limits the diversity of knowledge selection and generation.To this end, we establish a multireference KGC dataset and propose a series of metrics to systematically assess the one-tomany efficacy of existing KGC models.Furthermore, to extend the hypothesis space of knowledge selection to enhance the mapping relationship between multiple knowledge and multiple responses, we devise a span-based variational model and optimize the model in a wake-sleep style with an ameliorated evidence lower bound objective to learn the oneto-many generalization.Both automatic and human evaluations demonstrate the efficacy of our approach. Xueliang Zhao, Tingchen Fu, Chongyang Tao, Rui Yan 0001 |
EMNLP | 1 |
| 2022 | Towards Efficient Dialogue Pre-training with Transferable and Interpretable Latent StructureabstractWith the availability of massive generaldomain dialogue data, pre-trained dialogue generation appears to be super appealing to transfer knowledge from the general domain to downstream applications.In most existing work, such transferable ability is mainly obtained by fitting a large model with hundreds of millions of parameters on massive data in an exhaustive way, leading to inefficient running and poor interpretability.This paper proposes a novel dialogue generation model with a latent structure that is easily transferable from the general domain to downstream tasks in a lightweight and transparent way.Experiments on two benchmarks validate the effectiveness of the proposed model.Thanks to the transferable latent structure, our model is able to yield better dialogue responses than four strong baselines in terms of both automatic and human evaluations, and our model with about 22% parameters particularly delivers a 5x speedup in running time compared with the strongest baseline.Moreover, the proposed model is explainable by interpreting the discrete latent variables. Xueliang Zhao, Lemao Liu, Tingchen Fu, Shuming Shi 0001, Dongyan Zhao 0001, Rui Yan 0001 |
EMNLP | 1 |
| 2022 | Learning to Express in Knowledge-Grounded ConversationabstractXueliang Zhao, Tingchen Fu, Chongyang Tao, Wei Wu, Dongyan Zhao, Rui Yan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Xueliang Zhao, Tingchen Fu, Chongyang Tao, Wei Wu 0014, Dongyan Zhao 0001, Rui Yan 0001 |
NAACL-HLT | 1 |
| 2022 | Overview of the NLPCC 2022 Shared Task: Multi-modal Dialogue Understanding and Generation
Yuxuan Wang 0004, Xueliang Zhao, Dongyan Zhao 0001 |
NLPCC (2) | 2 |
| 2021 | Learning an Effective Context-Response Matching Model with Self-Supervised Tasks for Retrieval-based DialoguesabstractBuilding an intelligent dialogue system with the ability to select a proper response according to a multi-turn context is a great challenging task. Existing studies focus on building a context-response matching model with various neural architectures or pretrained language models (PLMs) and typically learning with a single response prediction task. These approaches overlook many potential training signals contained in dialogue data, which might be beneficial for context understanding and produce better features for response prediction. Besides, the response retrieved from existing dialogue systems supervised by the conventional way still faces some critical challenges, including incoherence and inconsistency. To address these issues, in this paper, we propose learning a context-response matching model with auxiliary self-supervised tasks designed for the dialogue data based on pre-trained language models. Specifically, we introduce four self-supervised tasks including next session prediction, utterance restoration, incoherence detection and consistency discrimination, and jointly train the PLM-based response selection model with these auxiliary tasks in a multi-task manner. By this means, the auxiliary tasks can guide the learning of the matching model to achieve a better local optimum and select a more proper response. Experiment results on two benchmarks indicate that the proposed auxiliary self-supervised tasks bring significant improvement for multi-turn response selection in retrieval-based dialogues, and our model achieves new state-of-the-art results on both datasets. Ruijian Xu, Chongyang Tao, Daxin Jiang, Xueliang Zhao, Dongyan Zhao 0001, Rui Yan 0001 |
AAAI | 4 |
| 2020 | Knowledge-Grounded Dialogue Generation with Pre-trained Language ModelsabstractWe study knowledge-grounded dialogue generation with pre-trained language models. To leverage the redundant external knowledge under capacity constraint, we propose equipping response generation defined by a pre-trained language model with a knowledge selection module, and an unsupervised approach to jointly optimizing knowledge selection and response generation with unlabeled dialogues. Empirical results on two benchmarks indicate that our model can significantly outperform state-of-the-art methods in both automatic evaluation and human judgment. Xueliang Zhao, Wei Wu 0014, Can Xu 0002, Chongyang Tao, Dongyan Zhao 0001, Rui Yan 0001 |
EMNLP (1) | 1 |
| 2020 | Low-Resource Knowledge-Grounded Dialogue Generation
Xueliang Zhao, Wei Wu 0014, Chongyang Tao, Can Xu 0002, Dongyan Zhao 0001, Rui Yan 0001 |
ICLR | 1 |
| 2020 | Zero-Resource Knowledge-Grounded Dialogue GenerationabstractWhile neural conversation models have shown great potentials towards generating informative and engaging responses via introducing external knowledge, learning such a model often requires knowledge-grounded dialogues that are difficult to obtain. To overcome the data challenge and reduce the cost of building a knowledge-grounded dialogue system, we explore the problem under a zero-resource setting by assuming no context-knowledge-response triples are needed for training. To this end, we propose representing the knowledge that bridges a context and a response and the way that the knowledge is expressed as latent variables, and devise a variational approach that can effectively estimate a generation model from independent dialogue corpora and knowledge corpora. Evaluation results on three benchmarks of knowledge-grounded dialogue generation indicate that our model can achieve comparable performance with state-of-the-art methods that rely on knowledge-grounded dialogues for training, and exhibits a good generalization ability over different datasets. Can Xu 0002, Wei Wu 0014, Yufan Zhao, Xueliang Zhao, Chongyang Tao |
NeurIPS | 5 |
| 2019 | A Document-grounded Matching Network for Response Selection in Retrieval-based ChatbotsabstractWe present a document-grounded matching network (DGMN) for response selection that can power a knowledge-aware retrieval-based chatbot system. The challenges of building such a model lie in how to ground conversation contexts with background documents and how to recognize important information in the documents for matching. To overcome the challenges, DGMN fuses information in a document and a context into representations of each other, and dynamically determines if grounding is necessary and importance of different parts of the document and the context through hierarchical interaction with a response at the matching step. Empirical studies on two public data sets indicate that DGMN can significantly improve upon state-of-the-art methods and at the same time enjoys good interpretability. Xueliang Zhao, Chongyang Tao, Wei Wu 0014, Can Xu 0002, Dongyan Zhao 0001, Rui Yan 0001 |
IJCAI | 1 |
| 2016 | Reliability and Performance Evaluation of Joint Redundancy and Inspection-Based Maintenance Strategy in Virtualized SystemabstractIn virtualized system, both redundancy and inspection-based maintenance has been used to maintain reliability. Analysis is often conducted on a single strategy, and the overall impact of a joint mechanism has not been analyzed in detail. So, a method is presented to analyze the reliability and performance of virtualized system with the joint mechanism. The reliability and performance indicators are based on the Markov chain constructed from the state transition diagram for the joint mechanism. Sensitivity analysis is conducted to analyze the impact of system configuration parameters change. Empirical studies show the process of evaluation model and sensitivity analysis. Changing the value of redundancy and inspection rate value, the system reliability and performance could be calculated through the analysis method. The increase of redundancy leads to the increase of reliability and performance rate. Whereas, the increase in inspection rate could only improve the performance to an extent. The influence of both parameters declines rapidly as the value increases. Pan He, Xiaoguang Lin, Xueliang Zhao |
ISPDC | 4 |