Yusheng Liao

dblp:37/4774 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-7549-3944ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MedS³: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process Supervision
abstract
Medical language models face critical barriers to real-world clinical reasoning applications. However, mainstream efforts, which fall short in task coverage, lack fine-grained supervision for intermediate reasoning steps, and rely on proprietary systems, are still far from a versatile, credible and efficient language model for clinical reasoning usage. To this end, we propose MedS3, a self-evolving framework that imparts robust reasoning capabilities to small, deployable models. Starting with 8,000 curated instances sampled via a curriculum strategy across five medical domains and 16 datasets, we use a small base policy model to conduct Monte Carlo Tree Search (MCTS) for constructing rule-verifiable reasoning trajectories. Self-explored reasoning trajectories ranked by node values are used to bootstrap the policy model via reinforcement fine-tuning and preference learning. Moreover, we introduce a soft dual process reward model that incorporates value dynamics: steps that degrade node value are penalized, enabling fine-grained identification of reasoning errors even when the final answer is correct. Experiments on eleven benchmarks show that MedS3 outperforms the previous state-of-the-art medical model by +6.45 accuracy points and surpasses 32B-scale general-purpose reasoning models by +8.57 points. Additional empirical analysis further demonstrates that MedS3 achieves robust and faithful reasoning behavior.
Shuyang Jiang, Yusheng Liao, Zhe Chen 0024, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027
AAAI2
2026 SLoRA: Balancing Plasticity and Forgetting in Large Language Models for Continual Learning
abstract
Large language models (LLMs) have achieved remarkable success across diverse tasks through large-scale pretraining.However, they remain prone to catastrophic forgetting in continual learning.To the best of our knowledge, this is the first work to identify noise accumulation in LoRA updates as a key cause of forgetting in continual learning.A preliminary two-task experiment demonstrates that removing less important components of the second task's LoRA parameters improves performance on the first task, suggesting that later updates introduce noisy interference.Building on this insight, we propose Subspace-Denoised Low-Rank Adaptation (SLoRA), a simple and effective framework that filters noisy components from LoRA updates via subspace similarity with the base model.SLoRA is a regularizationfree method without accessing data or gradients from previous tasks or modifying the training process.It offers two variants, SLoRA-Pre and SLoRA-Post, for online and offline continual learning, respectively.Extensive experiments across tasks and models validate the effectiveness of SLoRA.It improves final accuracy by up to 12%, reduces forgetting by 29%, and filters out over 30% of LoRA parameters identified as noisy.Our code is available at https://github.com/alina1031/SLoRA.
Yusheng Liao, Yanfeng Wang 0001, Yu Wang 0027
ACL (1)2
2025 Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications
abstract
Large language models hold promise for addressing medical challenges, such as medical diagnosis reasoning, research knowledge acquisition, clinical decision-making, and consumer health inquiry support. However, they often generate hallucinations due to limited medical knowledge. Incorporating external knowledge is therefore critical, which necessitates multi-source knowledge acquisition. We address this challenge by framing it as a source planning problem, which is to formulate context-appropriate queries tailored to the attributes of diverse sources. Existing approaches either overlook source planning or fail to achieve it effectively due to misalignment between the model’s expectation of the sources and their actual content. To bridge this gap, we present MedOmniKB, a repository comprising multigenre and multi-structured medical knowledge sources. Leveraging these sources, we propose the Source Planning Optimisation method, which enhances multi-source utilisation. Our approach involves enabling an expert model to explore and evaluate potential plans while training a smaller model to learn source alignment. Experimental results demonstrate that our method substantially improves multi-source planning performance, enabling the optimised small model to achieve state-of-the-art results in leveraging diverse medical knowledge sources.
Zhe Chen 0024, Yusheng Liao, Shuyang Jiang, Pingjie Wang, Yiqiu Guo, Yanfeng Wang 0001, Yu Wang 0027
ACL (1)2
2025 ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents
abstract
Large Language Models (LLMs) have shown promising potential in the medical domain, assisting with tasks like clinical note generation and patient communication.However, current LLMs are limited to text-based communication, hindering their ability to interact with diverse forms of information in clinical environments.Despite clinical agents succeeding in diverse signal interaction, they are oriented to a single clinical scenario and hence fail for broader applications.To evaluate clinical agents holistically, we propose ClinicalAgent Bench (CAB), a comprehensive medical agent benchmark consisting of 18 tasks across five key realistic clinical dimensions.Building on this, we introduce REFLECTOOL, a novel framework that excels at utilizing domain-specific tools within two stages.The first optimization stage progressively enlarges a long-term memory by saving successful solving processes and toolwise experience of agents in a tiny pre-defined training set.In the following inference stage, REFLECTOOL can search for supportive successful demonstrations from already built longterm memory to guide the tool selection strategy, and a verifier improves the tool usage according to the tool-wise experience with two verification methods-iterative refinement and candidate selection.Extensive experiments on CAB demonstrate that REFLECTOOL surpasses the pure LLMs with more than 10 points and the well-established agent-based methods with 3 points, highlighting its adaptability and effectiveness in solving complex clinical tasks.Our code and datasets are available at https: //github.com/BlueZeros/ReflecTool.
Yusheng Liao, Shuyang Jiang, Yanfeng Wang 0001, Yu Wang 0027
ACL (1)1
2025 EvolveBench: A Comprehensive Benchmark for Assessing Temporal Awareness in LLMs on Evolving Knowledge
abstract
Large language models (LLMs) are trained on extensive historical corpora, but their ability to understand time and maintain temporal awareness of time-evolving factual knowledge remains limited. Previous studies often neglect the critical aspect of utilizing knowledge from various sources. To address this gap, we introduce EvolveBench, a comprehensive benchmark that evaluates temporal competence along five key dimensions: Cognition, which examines the ability to recall and contextualize historical facts. Awareness, which tests LLMs’ awareness of temporal misalignment between external inputs and the temporal context of a query. Trustworthiness, which assesses whether models can identify and appropriately refuse queries based on invalid timestamps. Understanding, which focuses on interpreting both explicit dates and implicit historical markers. Finally, reasoning evaluates the capacity to analyze temporal relationships and draw accurate inferences. Evaluating 15 widely used LLMs on EvolveBench shows that GPT-4o achieves the highest average EM score of 79.36, while the open-source Llama3.1-70B demonstrates notable strength in handling temporally misaligned contexts with an average score of 72.47. Despite these advances, all models still struggle with handling temporal misaligned context. Our code and dataset are available at https://github.com/zzysjtuiwct/EvolveBench.
Yusheng Liao, Zhe Chen 0024, Yunfeng Guan 0001, Yanfeng Wang 0001, Yu Wang 0027
ACL (1)2
2025 DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models
abstract
The reliability of large language models remains a critical challenge, particularly due to their susceptibility to hallucinations and factual inaccuracies during text generation.Existing solutions either underutilize models' selfcorrection with preemptive strategies or use costly post-hoc verification.To further explore the potential of real-time self-verification and correction, we present Dynamic Self-Verify Decoding (DSVD), a novel decoding framework that enhances generation reliability through real-time hallucination detection and efficient error correction.DSVD integrates two key components: (1) parallel self-verification architecture for continuous quality assessment, (2) dynamic rollback mechanism for targeted error recovery.Extensive experiments across five benchmarks demonstrate DSVD's effectiveness, achieving significant improvement in truthfulness (Quesetion-Answering) and factual accuracy (FActScore).Results show the DSVD can be further incorporated with existing faithful decoding methods to achieve stronger performance.Our work establishes that real-time self-verification during generation offers a viable path toward more faithful language models without sacrificing practical deployability.
YiQiu Guo, Zhe Chen 0024, Pingjie Wang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027
EMNLP5
2025 DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
abstract
When performing reasoning tasks with userspecific requirements, such as strict output formats, large language models (LLMs) often prioritize reasoning over adherence to detailed instructions.Fine-tuning LLMs on supervised datasets to address this is impractical due to high computational costs and limited parameter access.To tackle this, we propose DICE, a lightweight framework that guides small language models (SLMs) to refine LLMs' outputs through chain-of-thought (CoT) correction.DICE decouples the process by first prompting LLMs to generate natural language responses, then using trained SLMs to analyze and refine these outputs to meet structured output specifications.This framework preserves LLMs' broad knowledge and reasoning capabilities while ensuring the outputs conform to user demands.Specifically, DICE first constructs structured CoT adaptation datasets via a two-stage method and subsequently applies a dual-tuning strategy to fine-tune SLMs for generating structured outputs in an analyze-thenanswer pattern. 1 Experiments demonstrate that DICE improves the average format accuracy and content correctness of LLM outputs by 35.4% and 29.4%, respectively, achieving stateof-the-art (SOTA) performance over other competitive baselines.
Yusheng Liao, Zhe Chen 0024, Yanfeng Wang 0001, Yu Wang 0027
EMNLP2
2025 Fine-tuning with Reserved Majority for Noise Reduction
abstract
Parameter-efficient fine-tuning (PEFT) has revolutionized supervised fine-tuning, where LoRA and its variants gain the most popularity due to their low training costs and zero inference latency. However, LoRA tuning not only injects knowledgeable features but also noisy hallucination during fine-tuning, which hinders the utilization of tunable parameters with the increasing LoRA rank. In this work, we first investigate in-depth the redundancies among LoRA parameters with substantial empirical studies. Aiming to resemble the learning capacity of high ranks from the findings, we set up a new fine-tuning framework, \textbf{P}arameter-\textbf{Re}dundant \textbf{F}ine-\textbf{T}uning (\preft), which follows the vanilla LoRA tuning process but is required to reduce redundancies before merging LoRA parameters back to pre-trained models. Based on this framework, we propose \textbf{No}ise reduction with \textbf{R}eserved \textbf{M}ajority~(\norm), which decomposes the LoRA parameters into majority parts and redundant parts with random singular value decomposition. The major components are determined by the proposed \search method, specifically employing subspace similarity to confirm the parameter groups that share the highest similarity with the base weight. By employing \norm, we enhance both the learning capacity and benefits from larger ranks, which consistently outperforms both LoRA and other \preft-based methods on various downstream tasks, such as general instruction tuning, math reasoning and code generation. Code is available at \url{https://github.com/pixas/NoRM}.
Shuyang Jiang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027
ICLR2
2024 MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
abstract
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in visual perception and understanding.However, these models also suffer from hallucinations, which limit their reliability as AI systems.We believe that these hallucinations are partially due to the models' struggle with understanding what they can and cannot perceive from images, a capability we refer to as self-awareness in perception.Despite its importance, this aspect of MLLMs has been overlooked in prior studies.In this paper, we aim to define and evaluate the selfawareness of MLLMs in perception.To do this, we first introduce the knowledge quadrant in perception, which helps define what MLLMs know and do not know about images.Using this framework, we propose a novel benchmark, the Self-Awareness in Perception for MLLMs (MM-SAP), specifically designed to assess this capability.We apply MM-SAP to a variety of popular MLLMs, offering a comprehensive analysis of their self-awareness and providing detailed insights.The experiment results reveal that current MLLMs possess limited selfawareness capabilities, pointing to a crucial area for future advancement in the development of trustworthy MLLMs.
Yusheng Liao, Heyang Liu, Yanfeng Wang 0001, Yu Wang 0027
ACL (1)2
2024 RA2FD: Distilling Faithfulness into Efficient Dialogue Systems
abstract
Generating faithful and fast responses is crucial in the knowledge-grounded dialogue.Retrieval Augmented Generation (RAG) strategies are effective but are inference inefficient, while previous Retrieval Free Generations (RFG) are more efficient but sacrifice faithfulness.To solve this faithfulness-efficiency trade-off dilemma, we propose a novel retrieval-free model training scheme named Retrieval Augmented to Retrieval Free Distillation (RA2FD) to build a retrieval-free model that achieves higher faithfulness than the previous RFG method while maintaining inference efficiency.The core idea of RA2FD is to use a teacher-student framework to distill the faithfulness capacity of a teacher, which is an oracle RAG model that generates multiple knowledge-infused responses.The student retrieval-free model learns how to generate faithful responses from these teacher labels through sequence-level distillation and contrastive learning.Experiment results show that RA2FD let the faithfulness performance of an RFG model surpass the previous SOTA RFG baseline on three knowledge-grounded dialogue datasets by an average of 33% and even matching an RAG model's performance while significantly improving inference efficiency.Our code is available at https:// github.com/zzysjtuiwct/RA2FD.
Yusheng Liao, Chenxin Xu, Yunfeng Guan 0001, Yanfeng Wang 0001, Yu Wang 0027
EMNLP2
2024 TAIA: Large Language Models are Out-of-Distribution Data Learners
abstract
Fine-tuning on task-specific question-answer pairs is a predominant method for enhancing the performance of instruction-tuned large language models (LLMs) on downstream tasks. However, in certain specialized domains, such as healthcare or harmless content generation, it is nearly impossible to obtain a large volume of high-quality data that matches the downstream distribution. To improve the performance of LLMs in data-scarce domains with domain-mismatched data, we re-evaluated the Transformer architecture and discovered that not all parameter updates during fine-tuning contribute positively to downstream performance. Our analysis reveals that within the self-attention and feed-forward networks, only the fine-tuned attention parameters are particularly beneficial when the training set's distribution does not fully align with the test set. Based on this insight, we propose an effective inference-time intervention method: \uline{T}raining \uline{A}ll parameters but \uline{I}nferring with only \uline{A}ttention (TAIA). We empirically validate TAIA using two general instruction-tuning datasets and evaluate it on seven downstream tasks involving math, reasoning, and knowledge understanding across LLMs of different parameter sizes and fine-tuning techniques. Our comprehensive experiments demonstrate that TAIA achieves superior improvements compared to both the fully fine-tuned model and the base model in most scenarios, with significant performance gains. The high tolerance of TAIA to data mismatches makes it resistant to jailbreaking tuning and enhances specialized tasks using general data. Code is available in \url{https://github.com/pixas/TAIA_LLM}.
Shuyang Jiang, Yusheng Liao, Ya Zhang 0002, Yanfeng Wang 0001, Yu Wang 0027
NeurIPS2
2024 Leveraging Diverse Modeling Contexts With Collaborating Learning for Neural Machine Translation
abstract
Autoregressive (AR) and Non-autoregressive (NAR) models are two types of generative models for Neural Machine Translation (NMT). AR models predict tokens in a word-byword manner and can effectively capture the distribution of real translations. NAR models predict tokens by extracting bidirectional contextual information which can improve the inference speed but they suffer from performance degradation. Previous works utilized AR models to enhance NAR models by reducing the training data's complexity or incorporating the global information into AR models by virtue of NAR models. However, those investigated methods only take advantage of the contextual information of a single type of model while neglecting the diversity in the contextual information that can be provided by different types of models. In this paper, we propose a novel generic collaborative learning method, DCMCL, where AR and NAR models are treated as collaborators instead of teachers and students. To hierarchically leverage the bilateral contextual information, token-level mutual learning and sequence-level contrastive learning are adopted between AR and NAR models. Extensive experiments on four widely used benchmarks show that the proposed DCMCL method can simultaneously improve both AR and NAR models with up to 1.38 and 2.98 BLEU scores, respectively, and can also outperform the current best-unified model with up to 0.97 BLEU scores for both AR and NAR decoding.
Yusheng Liao, Yanfeng Wang 0001, Yu Wang 0027
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Self-Improvement of Non-autoregressive Model via Sequence-Level Distillation
abstract
Although Non-autoregressive Transformer (NAT) models have achieved great success in terms of fast inference speed, this speedup comes with a performance drop due to the inherent multi-modality problem of the NAT model.Previous works commonly alleviate this problem by replacing the target side of the raw data with distilled data generated by Autoregressive Transformer (AT) models.However, the multimodality problem in the distilled data is still significant and thus limits further improvement of the NAT models.In this paper, we propose a method called Sequence-Level Self-Distillation (SLSD), which aims to generate distilled data by the NAT model itself, eliminating the need for additional teacher networks.Furthermore, SLSD can adapt to different NAT models without precise adjustments since the self-distilled data is generated from the same types of NAT models.We conduct extensive experiments on WMT14 EN↔DE and WMT16 EN↔RO and choose five classic NAT models as the backbones to validate the generality and effectiveness of SLSD.The results show that our approach can consistently improve all models on both raw data and distilled data without sacrificing the inference speed.
Yusheng Liao, Shuyang Jiang, Yu Wang 0027, Yanfeng Wang 0001
EMNLP1
2023 Contrastive Learning Based ASR Robust Knowledge Selection For Spoken Dialogue System
Yusheng Liao, Yu Wang 0027, Yunfeng Guan 0001
INTERSPEECH2
2021 Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopy
abstract
The Endoscopy Computer Vision Challenge (EndoCV) is a crowd-sourcing initiative to address eminent problems in developing reliable computer aided detection and diagnosis endoscopy systems and suggest a pathway for clinical translation of technologies. Whilst endoscopy is a widely used diagnostic and treatment tool for hollow-organs, there are several core challenges often faced by endoscopists, mainly: 1) presence of multi-class artefacts that hinder their visual interpretation, and 2) difficulty in identifying subtle precancerous precursors and cancer abnormalities. Artefacts often affect the robustness of deep learning methods applied to the gastrointestinal tract organs as they can be confused with tissue of interest. EndoCV2020 challenges are designed to address research questions in these remits. In this paper, we present a summary of methods developed by the top 17 teams and provide an objective comparison of state-of-the-art methods and methods designed by the participants for two sub-challenges: i) artefact detection and segmentation (EAD2020), and ii) disease detection and segmentation (EDD2020). Multi-center, multi-organ, multi-class, and multi-modal clinical endoscopy datasets were compiled for both EAD2020 and EDD2020 sub-challenges. The out-of-sample generalization ability of detection algorithms was also evaluated. Whilst most teams focused on accuracy improvements, only a few methods hold credibility for clinical usability. The best performing teams provided solutions to tackle class imbalance, and variabilities in size, origin, modality and occurrences by exploring data augmentation, data fusion, and optimal class thresholding techniques.
Sharib Ali, Mariia Dmitrieva, Noha M. Ghatwary, Sophia Bano, Gorkem Polat, Alptekin Temizel, Adrian Krenzer, Amar Hekalo, Bogdan J. Matuszewski, Mourad Gridach, Irina Voiculescu, Vishnusai Yoganand, Arnav Chavan, Aryan Raj, Nhan T. Nguyen, Dat Q. Tran, Lê Duy Huynh, Nicolas Boutry, Shahadate Rezvy, Haijian Chen, Yoon Ho Choi, Anand Subramanian 0004, Velmurugan Balasubramanian, Xiaohong W. Gao, Hongyu Hu, Yusheng Liao, Danail Stoyanov, Christian Daul, Stefano Realdon, Renato Cannizzaro, Dominique Lamarque, Terry Tran-Nguyen, Adam Bailey, Barbara Braden, James E. East, Jens Rittscher
Medical Image Anal.27