Jieyi Wang

dblp:205/7406 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Question answering and dialogue systems · 44% Language models and text generation · 32% Speech recognition and synthesis · 16%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
knowledge editing
1.012026
MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models · AAAI 2026
Natural language and speech › Language models and text generation › knowledge editing
multimodal knowledge editing
1.012026
MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models · AAAI 2026
Natural language and speech › Speech recognition and synthesis
speech language model
1.012026
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
spoken dialogue systems
1.012026
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation · ACL (1) 2026
Medical and health informatics › clinical informatics › clinical AI › clinical machine learning
medical multimodal large language model
1.012026
MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models · AAAI 2026
Natural language and speech › Question answering and dialogue systems › conversational agents
empathetic dialogue systems
0.912025
STAMPsy: Towards SpatioTemporal-Aware Mixed-Type Dialogues for Psychological Counseling · AAAI 2025
Natural language and speech › Question answering and dialogue systems › medical dialogue systems
psychological counseling dialogue
0.912025
STAMPsy: Towards SpatioTemporal-Aware Mixed-Type Dialogues for Psychological Counseling · AAAI 2025
Computer vision › Vision and language › visual question answering
medical visual question answering
0.312026
MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models · AAAI 2026
Medical and health informatics › clinical text processing
clinical dialogue
0.312026
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

two-stage training · 2.0modality re-alignment · 2.0knowledge injection · 2.0benchmark construction · 2.0knowledge-grounded dialogue · 0.9iterative self-feedback · 0.9
YearPublicationVenuePosition
2026 MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models
abstract
Recent advances in multimodal large language models (MLLMs) have significantly improved medical AI, enabling it to unify the understanding of visual and textual information. However, as medical knowledge continues to evolve, it is critical to allow these models to efficiently update outdated or incorrect information without retraining from scratch. Although textual knowledge editing has been widely studied, there is still a lack of systematic benchmarks for multimodal medical knowledge editing involving image and text modalities. To fill this gap, we present MedMKEB, the first comprehensive benchmark designed to evaluate the reliability, generality, locality, portability, and robustness of knowledge editing in medical multimodal large language models. MedMKEB is built on a high-quality medical visual question-answering dataset and enriched with carefully constructed editing tasks, including counterfactual correction, semantic generalization, knowledge transfer, and adversarial robustness. We incorporate human expert validation to ensure the accuracy and reliability of the benchmark. Extensive experiments on state-of-the-art general and medical MLLMs demonstrate the limitations of existing knowledge editing methods in the medical domain, highlighting the need to develop specialized editing strategies.
Dexuan Xu, Jieyi Wang, Zhongyan Chai, Yongzhi Cao, Hanpin Wang, Huamin Zhang, Yu Huang 0004
AAAI2
2026 SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation
abstract
Medical consultations are intrinsically speechcentric.However, most prior works focus on long-text-based interactions, which are cumbersome and patient-unfriendly.Recent advances in speech language models (SpeechLMs) have enabled more natural speech-based interaction, yet the scarcity of medical speech data and the inefficiency of directly fine-tuning on speech data jointly hinder the adoption of SpeechLMs in medical consultation.In this paper, we propose SpeechMedAssist, a SpeechLM natively capable of conducting speech-based multi-turn interactions with patients.By exploiting the architectural properties of SpeechLMs, we decouple the conventional one-stage training into a two-stage paradigm consisting of (1) Knowledge & Capability Injection via Text and (2) Modality Re-alignment with Limited Speech Data, thereby reducing the requirement for medical speech data to only 10k synthesized samples.To evaluate SpeechLMs for medical consultation scenarios, we design a benchmark comprising both single-turn question answering and multi-turn simulated interactions.Experimental results show that our model outperforms all baselines in both effectiveness and robustness in most evaluation settings.
Sirry Chen, Jieyi Wang, Wei Chen 0088, Zhongyu Wei
ACL (1)2
2025 STAMPsy: Towards SpatioTemporal-Aware Mixed-Type Dialogues for Psychological Counseling
abstract
Online psychological counseling dialogue systems are trending, offering a convenient and accessible alternative to traditional in-person therapy. However, existing psychological counseling dialogue systems mainly focus on basic empathetic dialogue or QA with minimal professional knowledge and without goal guidance. In many real-world counseling scenarios, clients often seek multi-type help, such as diagnosis, consultation, therapy, console, and common questions, but existing dialogue systems struggle to combine different dialogue types naturally. In this paper, we identify this challenge as how to construct mixed-type dialogue systems for psychological counseling that enable clients to clarify their goals before proceeding with counseling. To mitigate the challenge, we collect a mixed-type counseling dialogues corpus termed STAMPsy, covering five dialogue types, task-oriented dialogue for diagnosis, knowledge-grounded dialogue, conversational recommendation, empathetic dialogue, and question answering, over 5,000 conversations. Moreover, spatiotemporal-aware knowledge enables systems to have world awareness and has been proven to affect one's mental health. Therefore, we link dialogues in STAMPsy to spatiotemporal state and propose a spatiotemporal-aware mixed-type psychological counseling dataset. Additionally, we build baselines on STAMPsy and develop an iterative self-feedback psychological dialogue generation framework, named Self-STAMPsy. Results indicate that clarifying dialogue goals in advance and utilizing spatiotemporal states are effective.
Jieyi Wang, Zeming Liu, Dexuan Xu, Chuan Wang 0002, Ruiyuan Guan, Weihua Yue, Yu Huang 0004
AAAI1