Sirry Chen

dblp:377/5735 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Question answering and dialogue systems · 50% Speech recognition and synthesis · 50%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
speech language model
1.012026
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems
1.012026
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation · ACL (1) 2026
Medical and health informatics › clinical text processing
clinical dialogue
0.312026
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

two-stage training · 2.0modality re-alignment · 2.0knowledge injection · 2.0
YearPublicationVenuePosition
2026 SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
abstract
Medical consultations are intrinsically speechcentric.However, most prior works focus on long-text-based interactions, which are cumbersome and patient-unfriendly.Recent advances in speech language models (SpeechLMs) have enabled more natural speech-based interaction, yet the scarcity of medical speech data and the inefficiency of directly fine-tuning on speech data jointly hinder the adoption of SpeechLMs in medical consultation.In this paper, we propose SpeechMedAssist, a SpeechLM natively capable of conducting speech-based multi-turn interactions with patients.By exploiting the architectural properties of SpeechLMs, we decouple the conventional one-stage training into a two-stage paradigm consisting of (1) Knowledge & Capability Injection via Text and (2) Modality Re-alignment with Limited Speech Data, thereby reducing the requirement for medical speech data to only 10k synthesized samples.To evaluate SpeechLMs for medical consultation scenarios, we design a benchmark comprising both single-turn question answering and multi-turn simulated interactions.Experimental results show that our model outperforms all baselines in both effectiveness and robustness in most evaluation settings.
Sirry Chen, Jieyi Wang, Wei Chen 0088, Zhongyu Wei
ACL (1)1