Wen-Wai Yim

dblp:152/2666 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
10since 2021 · last 2026
0000-0001-9011-0817ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 7 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 ImageCLEF 2026: Multimodal Challenges in Medicine, Science, Agritech, and Security
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandra Baicoianu, Ana Neacsu, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Lea Reinartz, Benjamin Lecouteux, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Corneliu Florea, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Hendrik Damm, Henning Schäfer, Ivan Koychev, Josiane Mothe, Liviu-Daniel Stefan, Maja J. Hjuler, Mehmet Kurt, Meliha Yetisgen, Michael Riegler 0001, Mihai Dogariu, Mihai Ivanovici, Ming Shan Hee, Mohammad El Sakka, Momina Ahsan, Obioma Pelka, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Bahadir Eryilmaz, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Yuri Prokopchuk, Zhuohan Xie
ECIR (4)47
2026 MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
abstract
Evaluating natural language generation (NLG) systems in the medical domain presents unique challenges due to the critical demands for accuracy, relevance, and domain-specific expertise. Traditional automatic evaluation metrics, such as BLEU, ROUGE, and BERTScore, often fall short in distinguishing between high-quality outputs, especially given the open-ended nature of medical question answering (QA) tasks where multiple valid responses may exist. In this work, we introduce MORQA (Medical Open-Response QA), a new multilingual benchmark designed to assess the effectiveness of NLG evaluation metrics across three medical visual and text-based QA datasets in English and Chinese. Unlike prior resources, our datasets feature 2-4+ gold-standard answers authored by medical professionals, along with expert human ratings for three English and Chinese subsets. We benchmark both traditional metrics and large language model (LLM)-based evaluators, such as GPT-4 and Gemini, finding that LLM-based approaches significantly outperform traditional metrics in correlating with expert judgments. We further analyze factors driving this improvement, including LLMs' sensitivity to semantic nuances and robustness to variability among reference answers. Our results provide the first comprehensive, multilingual qualitative study of NLG evaluation in the medical domain, highlighting the need for human-aligned evaluation methods. All datasets and annotations will be publicly released to support future research.
Wen-Wai Yim, Asma Ben Abacha, Zixuan Yu, Robert Doerning, Fei Xia 0004, Meliha Yetisgen
LREC1
2025 ImageCLEF 2025: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Dan-Cristian Stanciu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Andrea M. Storås, Asma Ben Abacha, Benjamin Bracke, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Diandra Fabre, Didier Schwab, Dimitar Dimitrov 0003, Emmanuelle Esperança-Rodier, Mihai Gabriel Constantin, Helmut Becker, Hendrik Damm, Henning Schäfer, Ivan Rodkin, Ivan Koychev, Johannes Kiesel, Johannes Rückert, Josep Malvehy, Liviu-Daniel Stefan, Louise Bloch, Martin Potthast, Maximilian Heinrich, Michael Riegler 0001, Mihai Dogariu, Noel Codella, Pål Halvorsen, Preslav Nakov, Raphael Brüngel, Roberto A. Novoa, Rocktim Jyoti Das, Steven Alexander Hicks, Sushant Gautam, Tabea Margareta Grace Pakull, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim, Zhuohan Xie
ECIR (5)46
2025 A scoping review of natural language processing in addressing medically inaccurate information: Errors, misinformation, and hallucination
Zhaoyi Sun, Wen-Wai Yim, Özlem Uzuner, Fei Xia 0004, Meliha Yetisgen
J. Biomed. Informatics2
2025 WoundcareVQA: A multilingual visual question answering benchmark dataset for wound care
abstract
OBJECTIVE: Introduce the task of wound care multimodal multilingual visual question answering, provide baseline performances, and identify areas of future study. METHODS: A dataset of wound care multimodal multilingual visual question answering (VQA) was created using consumer health questions asked online. Practicing US medical doctors were tasked with providing metadata and expert responses labels. Several instruct-enabled, multilingual visual question answering models (GPT-4o, Gemini-1.5-Pro, and Qwen-VL) were tested to benchmark performances. Finally, automatic evaluations were tested against domain expert response ratings. RESULTS: A multilingual dataset of 477 wound care cases, 768 responses, 748 images, 3k structured data labels, 1362 translation instances, and 10k judgments was constructed (https://osf.io/xsj5u/). Metadata scores ranged from 0.32-0.78 accuracy depending on classification type; response generation performances 0.06 BLEU, 0.66 BERTScore, 0.45 ROUGE-L in English and 0.12 BLEU, 0.69 BERTScore, and 0.50 ROUGE-L in Chinese. CONCLUSION: We construct and explore the tasks of multimodal, multilingual VQA. We hope the work here can inspire further research in wound care metadata classification, VQA response generation, and open response automatic evaluation.
Wen-Wai Yim, Asma Ben Abacha, Robert Doerning, Chia-Yu Chen, Jiaying Xu, Anita Subbarao, Zixuan Yu, Fei Xia 0004, M. Kennedy Hall, Meliha Yetisgen
J. Biomed. Informatics1
2024 To Err Is Human, How about Medical Large Language Models? Comparing Pre-trained Language Models for Medical Assessment Errors and Reliability
abstract
Unpredictability, especially unpredictability with unknown error characteristics, is a highly undesirable trait, particularly in medical patient care applications. Although large pre-trained language models (LLM) have been applied to a variety of unseen tasks with highly competitive and successful results, their sensitivity to language inputs and resulting performance variability is not well-studied. In this work, we test state-of-the-art pre-trained language models from a variety of families to characterize their error generation and reliability in medical assessment ability. Particularly, we experiment with general medical assessment multiple choice tests, as well as their open-ended and true-false alternatives. We also profile model consistency, error agreements with each other and to humans; and finally, quantify their ability to recover and explain errors. The findings in this work can be used to give further information about medical models so that modelers can make better-informed decisions rather than relying on standalone performance metrics alone.
Wen-Wai Yim, Yujuan Fu, Asma Ben Abacha, Meliha Yetisgen
LREC/COLING1
2024 Advancing Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications with ImageCLEF 2024
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Ahmad Idrissi-Yaghir, Ahmedkhan Radzhabov, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Benjamin Lecouteux, Benno Stein 0001, Cécile Macaire, Christoph M. Friedrich, Cynthia Sabrina Schmidt, Didier Schwab, Emmanuelle Esperança-Rodier, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ioan Coman, Johanna Schöler, Johannes Kiesel, Johannes Rückert, Louise Bloch, Martin Potthast, Maximilian Heinrich, Meliha Yetisgen, Michael Riegler 0001, Neal Snider, Pål Halvorsen, Raphael Brüngel, Steven Alexander Hicks, Vajira Thambawita, Vassili Kovalev, Yuri Prokopchuk, Wen-Wai Yim
ECIR (6)38
2024 DermaVQA: A Multilingual Visual Question Answering Dataset for Dermatology
Wen-Wai Yim, Yujuan Fu, Zhaoyi Sun, Asma Ben Abacha, Meliha Yetisgen, Fei Xia 0004
MICCAI (5)1
2023 An Empirical Study of Clinical Note Generation from Doctor-Patient Encounters
abstract
Medical doctors spend on average 52 to 102 minutes per day writing clinical notes from their patient encounters (Hripcsak et al., 2011).Reducing this workload calls for relevant and efficient summarization methods.In this paper, we introduce new resources and empirical investigations for the automatic summarization of doctor-patient conversations in a clinical setting.In particular, we introduce the MTS-DIALOG dataset; a new collection of 1,700 doctor-patient dialogues and corresponding clinical notes.We use this new dataset to investigate the feasibility of this task and the relevance of existing language models, data augmentation, and guided summarization techniques.We compare standard evaluation metrics based on n-gram matching, contextual embeddings, and Fact Extraction to assess the accuracy and the factual consistency of the generated summaries.To ground these results, we perform an expert-based evaluation using relevant natural language generation criteria and task-specific criteria such as critical omissions, and study the correlation between the automatic metrics and expert judgments.To the best of our knowledge, this study is the first attempt to introduce an open dataset of doctor-patient conversations and clinical notes, with detailed automated and manual evaluations of clinical note generation.
Asma Ben Abacha, Wen-Wai Yim, Yadan Fan, Thomas Lin
EACL2
2023 ImageCLEF 2023 Highlight: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications
Bogdan Ionescu, Henning Müller, Ana-Maria Claudia Dragulinescu, Adrian Popescu 0001, Ahmad Idrissi-Yaghir, Alba Garcia Seco de Herrera, Alexandra-Georgiana Andrei, Alexandru Stan, Andrea M. Storås, Asma Ben Abacha, Christoph M. Friedrich, George Ioannidis, Griffin Adams, Henning Schäfer, Hugo Manguinhas, Ihar Filipovich, Ioan Coman, Jérôme Deshayes-Chossart, Johanna Schöler, Johannes Rückert, Liviu-Daniel Stefan, Louise Bloch, Meliha Yetisgen, Michael Riegler 0001, Mihai Dogariu, Mihai Gabriel Constantin, Neal Snider, Nikolaos Papachrysos, Pål Halvorsen, Raphael Brüngel, Serge Kozlovski, Steven Alexander Hicks, Thomas de Lange, Vajira Thambawita, Vassili Kovalev, Wen-Wai Yim
ECIR (3)36
2020 Alignment Annotation for Clinic Visit Dialogue to Clinical Note Sentence Language Generation
abstract
For every patient’s visit to a clinician, a clinical note is generated documenting their medical conversation, including complaints discussed, treatments, and medical plans. Despite advances in natural language processing, automating clinical note generation from a clinic visit conversation is a largely unexplored area of research. Due to the idiosyncrasies of the task, traditional methods of corpus creation are not effective enough approaches for this problem. In this paper, we present an annotation methodology that is content- and technique- agnostic while associating note sentences to sets of dialogue sentences. The sets can further be grouped with higher order tags to mark sets with related information. This direct linkage from input to output decouples the annotation from specific language understanding or generation strategies. Here we provide data statistics and qualitative analysis describing the unique annotation challenges. Given enough annotated data, such a resource would support multiple modeling methods including information extraction with template language generation, information retrieval type language generation, or sequence to sequence modeling.
Wen-Wai Yim, Meliha Yetisgen, Jenny Huang, Micah Grossman
LREC1
2017 Mining Electronic Health Records to Extract Patient-Centered Outcomes Following Prostate cancer Treatment
Tina Hernandez-Boussard, Panayotis Kourdis, Wen-Wai Yim, Rajendra Dulal, Douglas W. Blayney, James D. Brooks
AMIA3
2017 Classification of hepatocellular carcinoma stages from free-text clinical and radiology reports
Wen-Wai Yim, Sharon W. Kwan, Guy Johnson, Meliha Yetisgen
AMIA1
2017 Classifying tumor event attributes in radiology reports
abstract
Radiology reports contain vital diagnostic information that characterizes patient disease progression. However, information from reports is represented in free text, which is difficult to query against for secondary use. Automatic extraction of important information, such as tumor events using natural language processing, offers possibilities in improved clinical decision support, cohort identification, and retrospective evidence‐based research for cancer patients. The goal of this work was to classify tumor event attributes: negation, temporality, and malignancy, using biomedical ontology and linguistically enriched features. We report our results on an annotated corpus of 101 hepatocellular carcinoma patient radiology reports, and show that the improved classification improves overall template structuring. Classification performances for negation identification, past temporality classification, and malignancy classification were at 0.94, 0.62, and 0.77 F1, respectively. Incorporating the attributes into full templates led to an improvement of 0.72 F1 for tumor‐related events over a baseline of 0.65 F1. Improvement of negation, malignancy, and temporality classifications led to significant improvements in template extraction for the majority of categories. We present our machine‐learning approach to identifying these several tumor event attributes from radiology reports, as well as highlight challenges and areas for improvement.
Wen-Wai Yim, Sharon W. Kwan, Meliha Yetisgen
J. Assoc. Inf. Sci. Technol.1
2016 Annotation of Tumor Reference Resolution and Tumor Characteristics for Cancer Liver Stage Prediction
Wen-Wai Yim, Tyler Denman, Sharon W. Kwan, Meliha Yetisgen
AMIA1
2016 Tumor reference resolution and characteristic extraction in radiology reports for liver cancer stage prediction
Wen-Wai Yim, Sharon W. Kwan, Meliha Yetisgen
J. Biomed. Informatics1
2015 Annotation of Disease Characteristics for Cancer Liver Stage Prediction
Wen-Wai Yim, Sharon W. Kwan, Guy Johnson, Meliha Yetisgen
AMIA1
2014 A New Corpus for Structured Microbiology Results
Wen-Wai Yim, Xavier Engle, Heather L. Evans, Meliha Yetisgen
AMIA1