Kaishuai Xu

dblp:295/3979 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 72% Vision and language · 15% Speech recognition and synthesis · 8%
Human-computer interaction and pervasive computing
1 paper
Health and well-being technologies · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language › medical report generation
radiology report generation
1.522025
RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection · ACL (1) 2025
ORGAN: Observation-Guided Radiology Report Generation via Tree Reasoning · ACL (1) 2023
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.912025
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling · NeurIPS 2025
Natural language and speech › Language models and text generation
decoding
0.912025
Integrative Decoding: Improving Factuality via Implicit Self-consistency · ICLR 2025
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability › factuality
large language model factuality
0.912025
Integrative Decoding: Improving Factuality via Implicit Self-consistency · ICLR 2025
Natural language and speech › Language models and text generation › large language model
reasoning model
0.912025
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling · NeurIPS 2025
Natural language and speech › Language models and text generation › decoding
self-consistency decoding
0.912025
Integrative Decoding: Improving Factuality via Implicit Self-consistency · ICLR 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling · NeurIPS 2025
Medical and health informatics › clinical text processing
clinical text generation
0.912025
RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection · ACL (1) 2025
Health and well-being technologies › mental health detection
depression detection
0.812024
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection · EMNLP 2024
Health and well-being technologies
mental health detection
0.812024
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection · EMNLP 2024
Natural language and speech › Language models and text generation › text generation
content planning
0.712023
ORGAN: Observation-Guided Radiology Report Generation via Tree Reasoning · ACL (1) 2023
Machine learning › Efficient and distributed learning
inference efficiency
0.312025
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling · NeurIPS 2025
Natural language and speech › Language models and text generation › text generation
open-ended text generation
0.312025
Integrative Decoding: Improving Factuality via Implicit Self-consistency · ICLR 2025
Machine learning › Reinforcement learning
preference learning
0.312025
Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing · ACL (1) 2025
Information retrieval › retrieval-augmented generation
knowledge retrieval
0.312025
RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model
0.212024
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

multimodal learning · 2.6large language model · 2.6knowledge retrieval · 2.6self-editing · 0.9self-consistency · 0.9sampling · 0.9preference learning · 0.9perplexity-based importance refinement · 0.9error injection · 0.9chain-of-thought distillation · 0.9multimodal fusion · 0.8acoustic landmark · 0.8
YearPublicationVenuePosition
2025 RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in various domains, including radiology report generation.Previous approaches have attempted to utilize multimodal LLMs for this task, enhancing their performance through the integration of domainspecific knowledge retrieval.However, these approaches often overlook the knowledge already embedded within the LLMs, leading to redundant information integration.To address this limitation, we propose RADAR, a framework for enhancing radiology report generation with supplementary knowledge injection.RADAR improves report generation by systematically leveraging both the internal knowledge of an LLM and externally retrieved information.Specifically, it first extracts the model's acquired knowledge that aligns with expert imagebased classification outputs.It then retrieves relevant supplementary knowledge to further enrich this information.Finally, by aggregating both sources, RADAR generates more accurate and informative radiology reports.Extensive experiments on MIMIC-CXR, CHEXPERT-PLUS, and IU X-RAY demonstrate that our model outperforms state-of-the-art LLMs in both language quality and clinical accuracy 1 .
Kaishuai Xu, Heng Li 0010, Jiang Liu 0001
ACL (1)3
2025 Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
abstract
Kaishuai Xu, Tiezheng Yu, Wenjun Hou, Yi Cheng, Chak Tou Leong, Liangyou Li, Xin Jiang, Lifeng Shang, Qun Liu, Wenjie Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kaishuai Xu, Tiezheng Yu, Chak Tou Leong, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001, Wenjie Li 0002
ACL (1)1
2025 Integrative Decoding: Improving Factuality via Implicit Self-consistency
abstract
Self-consistency-based approaches, which involve repeatedly sampling multiple outputs and selecting the most consistent one as the final response, prove to be remarkably effective in improving the factual accuracy of large language models. Nonetheless, existing methods usually have strict constraints on the task format, largely limiting their applicability. In this paper, we present Integrative Decoding (ID), to unlock the potential of self-consistency in open-ended generation tasks. ID operates by constructing a set of inputs, each prepended with a previously sampled response, and then processes them concurrently, with the next token being selected by aggregating of all their corresponding predictions at each decoding step. In essence, this simple approach implicitly incorporates self-consistency in the decoding objective. Extensive evaluation shows that ID consistently enhances factuality over a wide range of language models, with substantial improvements on the TruthfulQA (+11.2%), Biographies (+15.4%) and LongFact (+8.5%) benchmarks. The performance gains amplify progressively as the number of sampled responses increases, indicating the potential of ID to scale up with repeated sampling.
Yeyun Gong, Yuji Zhang 0002, Kaishuai Xu, Wenge Liu, Wenjie Li 0002, Jian Jiao 0007, Qi Chen 0009, Peng Cheng 0005, Wayne Xiong
ICLR8
2025 LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
abstract
Large language models (LLMs) have demonstrated remarkable reasoning capabilities through test-time scaling approaches, particularly when fine-tuned with chain-of-thought (CoT) data distilled from more powerful large reasoning models (LRMs). However, these reasoning chains often contain verbose elements that mirror human problem-solving, categorized as progressive reasoning (the essential solution development path) and functional elements (verification processes, alternative solution approaches, and error corrections). While progressive reasoning is crucial, the functional elements significantly increase computational demands during test-time inference. We introduce PIR (Perplexity-based Importance Refinement), a principled framework that quantitatively evaluates the importance of each reasoning step based on its impact on answer prediction confidence. PIR systematically identifies and selectively prunes only low-importance functional steps while preserving all progressive reasoning components, creating optimized training data that maintains the integrity of the core solution path while reducing verbosity. Models fine-tuned on PIR-optimized data exhibit superior test-time scaling properties, generating more concise reasoning chains while achieving improved accuracy (+0.9\% to +6.6\%) with significantly reduced token usage (-3\% to -41\%) across challenging reasoning benchmarks (AIME, AMC, and GPQA Diamond). Our approach demonstrates strong generalizability across different model sizes, data sources, and token budgets, offering a practical solution for deploying reasoning-capable LLMs in scenarios where efficient test-time scaling, response time, and computational efficiency are valuable constraints. Code and dataset are available at the [LIMOPro GitHub repository.](https://github.com/GAIR-NLP/LIMOPro)
Jiashuo Wang, Ruifeng Yuan, Chunpu Xu, Kaishuai Xu, Wenjie Li 0002, Pengfei Liu 0003
NeurIPS5
2024 When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
abstract
Depression is a critical concern in global mental health, prompting extensive research into AIbased detection methods.Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications.However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities.Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped.In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection.We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks.By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts.This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals.Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines.In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals.
Xiangyu Zhang 0005, Hexin Liu, Kaishuai Xu, Qiquan Zhang, Daijiao Liu, Beena Ahmed, Julien Epps
EMNLP3
2023 ORGAN: Observation-Guided Radiology Report Generation via Tree Reasoning
abstract
This paper explores the task of radiology report generation, which aims at generating free-text descriptions for a set of radiographs.One significant challenge of this task is how to correctly maintain the consistency between the images and the lengthy report.Previous research explored solving this issue through planningbased methods, which generate reports only based on high-level plans.However, these plans usually only contain the major observations from the radiographs (e.g., lung opacity), lacking much necessary information, such as the observation characteristics and preliminary clinical diagnoses.To address this problem, the system should also take the image information into account together with the textual plan and perform stronger reasoning during the generation process.In this paper, we propose an Observation-guided radiology Report GenerAtioN framework (ORGAN).It first produces an observation plan and then feeds both the plan and radiographs for report generation, where an observation graph and a tree reasoning mechanism are adopted to precisely enrich the plan information by capturing the multiformats of each observation.Experimental results demonstrate that our framework outperforms previous state-of-the-art methods regarding text quality and clinical efficacy.1
Kaishuai Xu
ACL (1)2