VLDB 2026 Research / reviewers in the wild / expert
Douglas Teodoro
dblp:01/7332 · also Douglas Theodoro
· DBLP profile ↗
14ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0001-6238-4503ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 9 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GLiNER-BioMed: a suite of efficient models for open biomedical named entity recognitionabstractMOTIVATION: Biomedical named entity recognition (NER) presents unique challenges due to specialized vocabularies, the sheer volume of entities, and the continuous emergence of novel entities. Traditional NER models, constrained by fixed taxonomies and human annotations, struggle to generalize beyond predefined entity types. RESULTS: To address these issues, we introduce GLiNER-BioMed, a domain-adapted suite of GLiNER models for biomedicine. Our approach first distills the annotation capabilities of large language models (LLMs) into a smaller, more efficient model, enabling the generation of high-coverage biomedical NER data. We subsequently train two GLiNER architectures, uni- and bi-encoder, at multiple scales to balance computational efficiency and performance. Experiments on eight biomedical datasets demonstrate that GLiNER-BioMed achieved state-of-the-art zero-shot performance (micro-F1 59.77%), exceeding the strongest baseline by 5.96 points (P < .001). In few-shot learning, the bi-encoder variant reached 70.39% (10-shot), consistently outperforming the strongest baseline across all settings (P < .05). Our findings show that the uni-encoder GLiNER-BioMed achieves the strongest zero-shot performance, while the bi-encoder offers superior few-shot gains and substantially higher inference throughput (+39%-568%), making it well-suited to annotation-limited, latency-sensitive, or large-label-space settings. Ablation studies further indicate that combining synthetic biomedical pre-training with general-domain post-training is essential for capturing domain-specific knowledge while maintaining precision-recall balance. AVAILABILITY AND IMPLEMENTATION: The source code, datasets, and models are publicly available at https://github.com/ds4dh/GLiNER-biomed. Anthony Yazdani, Ihor Stepanov, Douglas Teodoro |
Bioinform. | 3 |
| 2026 | The detectability paradox: bilingual medical report generation with open-weight models and the limits of human oversightabstractOBJECTIVES: The automation of medical report generation using large language models (LLMs) could significantly reduce physicians' documentation burden while enhancing healthcare efficiency. However, the misuse of generative artificial intelligence in medical reporting can lead to important safety risks for patients. We addressed 2 questions: (1) What is the quality of medical reports generated by LLMs in English and French? and (2) Can we distinguish between human-written and LLM-generated medical reports? MATERIALS AND METHODS: We evaluated the quality of reports generated by several multilingual, open-weight LLMs using text similarity metrics on 4212 medical reports in English and French across multiple specialties. A bilingual expert panel of certified physicians (n = 4) and medical residents (n = 5) scored accuracy, fluency, and completeness of generated reports using a 1-5 Likert scale. Experts also completed a Turing-like test, blindly identifying reports as human or machine-generated. RESULTS: Phi-4 achieved the best overall performance (ROUGE-1: 0.70, BERTScore: 0.83). Expert evaluation confirmed high-quality reports in both languages (overall 4.6/5.0). Medical experts performed better than chance but struggled to differentiate human versus machine reports (accuracy: 0.60). Automatic classifiers showed strong performance (accuracy: 0.98). DISCUSSION: The high quality of LLM-generated reports supports their potential to enhance healthcare efficiency in multilingual settings. However, the discrepancy between human detection difficulty and automated detection success reveals inherent limitations in relying solely on human oversight for quality assurance and misuse prevention. CONCLUSIONS: Deployment of LLMs for medical reporting requires combining automated detection tools with human expertise to ensure patient safety. Dataset and code: https://github.com/ds4dh/medical_report_generation. Hossein Rouhizadeh, Abiram Sandralegar, Anthony Yazdani, Weibo Feng, Oren Schreier, Yonnou Ahn-Kim, Assiya Sirbal, Valentino Pirelli, Rui Yang 0016, Lukas Sveikata, Elena Tessitore, Nan Liu 0003, Philippe Bijlenga, Douglas Teodoro |
J. Am. Medical Informatics Assoc. | 14 |
| 2025 | STM-GNN: Space-Time-and-Memory Graph Neural Networks for Predicting Multi-Drug Resistance Risks in Dynamic Patient Networks
Damien Geissbuhler, Alban Bornet, Catarina Marques, André Anjos, Sónia Pereira, Douglas Teodoro |
AIME (1) | 6 |
| 2025 | ICU-TSB: A Benchmark for Temporal Patient Representation Learning for Unsupervised Stratification Into Patient CohortsabstractPatient stratification-identifying clinically meaningful sub-groups-is essential for advancing personalized medicine through improved diagnostics and treatment strategies. Electronic health records (EHRs), particularly those from intensive care units (ICUs), contain rich temporal clinical data that can be leveraged for this purpose. In this work, we introduce ICU-TSB (Temporal Stratification Benchmark), the first comprehensive benchmark for evaluating patient stratification based on temporal patient representation learning using three publicly available ICU EHR datasets. A key contribution of our benchmark is a novel hierarchical evaluation framework utilizing disease taxonomies to measure the alignment of discovered clusters with clinically validated disease groupings. In our experiments with ICU-TSB, we compared statistical methods and several recurrent neural networks, including LSTM and GRU, for their ability to generate effective patient representations for subsequent clustering of patient trajectories. Our results demonstrate that temporal representation learning can rediscover clinically meaningful patient cohorts; nevertheless, it remains a challenging task, with v-measuring varying from up to 0.46 at the top level of the taxonomy to up to 0.40 at the lowest level. To further enhance the practical utility of our findings, we also evaluate multiple strategies for assigning interpretable labels to the identified clusters. The experiments and benchmark are fully reproducible and available at https://github.com/ds4dh/CBMS2025stratification. Dimitrios Proios, Alban Bornet, Anthony Yazdani, Jose F. Rodrigues, Douglas Teodoro |
CBMS | 5 |
| 2025 | MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model EvaluationabstractWeihao Xuan, Rui Yang, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing, Junjue Wang, Fan Gao, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen, Douglas Teodoro, Nan Liu, Randy Goebel, Lei Ma, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Weihao Xuan, Rui Yang 0016, Heli Qi, Qingcheng Zeng, Yunze Xiao, Aosong Feng, Dairui Liu, Yun Xing 0001, Jinghui Lu, Yuang Jiang, Huitao Li, Xin Li 0079, Kunyu Yu, Ruihai Dong, Shangding Gu, Yuekang Li, Xiaofei Xie, Felix Juefei-Xu, Foutse Khomh, Osamu Yoshie, Qingyu Chen 0001, Douglas Teodoro, Nan Liu 0003, Randy Goebel, Lei Ma 0003, Edison Marrese-Taylor, Shijian Lu, Yusuke Iwasawa, Yutaka Matsuo, Irene Li |
EMNLP | 24 |
| 2025 | Comparing neural language models for medical concept representation and patient trajectory predictionabstractEffective representation of medical concepts is crucial for secondary analyses of electronic health records. Neural language models have shown promise in automatically deriving medical concept representations from clinical data. However, the comparative performance of different language models for creating these empirical representations, and the extent to which they encode medical semantics, has not been extensively studied. This study aims to address this gap by evaluating the effectiveness of three popular language models - word2vec, fastText, and GloVe - in creating medical concept embeddings that capture their semantic meaning. By using a large dataset of digital health records, we created patient trajectories and used them to train the language models. We then assessed the ability of the learned embeddings to encode semantics through an explicit comparison with biomedical terminologies, and implicitly by predicting patient outcomes and trajectories with different levels of available information. Our qualitative analysis shows that empirical clusters of embeddings learned by fastText exhibit the highest similarity with theoretical clustering patterns obtained from biomedical terminologies, with a similarity score between empirical and theoretical clusters of 0.88, 0.80, and 0.92 for diagnosis, procedure, and medication codes, respectively. Conversely, for outcome prediction, word2vec and GloVe tend to outperform fastText, with the former achieving AUROC as high as 0.78, 0.62, and 0.85 for length-of-stay, readmission, and mortality prediction, respectively. In predicting medical codes in patient trajectories, GloVe achieves the highest performance for diagnosis and medication codes (AUPRC of 0.45 and of 0.81, respectively) at the highest level of the semantic hierarchy, while fastText outperforms the other models for procedure codes (AUPRC of 0.66). Our study demonstrates that subword information is crucial for learning medical concept representations, but global embedding vectors are better suited for more high-level downstream tasks, such as trajectory prediction. Thus, these models can be harnessed to learn representations that convey clinical meaning, and our insights highlight the potential of using machine learning techniques to semantically encode medical data. Alban Bornet, Dimitrios Proios, Anthony Yazdani, Fernando Jaume-Santero, Guy Haller, Edward Choi 0003, Douglas Teodoro |
Artif. Intell. Medicine | 7 |
| 2025 | Analysis of eligibility criteria clusters based on large language models for clinical trial designabstractOBJECTIVES: Clinical trials (CTs) are essential for improving patient care by evaluating new treatments' safety and efficacy. A key component in CT protocols is the study population defined by the eligibility criteria. This study aims to evaluate the effectiveness of large language models (LLMs) in encoding eligibility criterion information to support CT-protocol design. MATERIALS AND METHODS: We extracted eligibility criterion sections, phases, conditions, and interventions from CT protocols available in the ClinicalTrials.gov registry. Eligibility sections were split into individual rules using a criterion tokenizer and embedded using LLMs. The obtained representations were clustered. The quality and relevance of the clusters for protocol design was evaluated through 3 experiments: intrinsic alignment with protocol information and human expert cluster coherence assessment, extrinsic evaluation through CT-level classification tasks, and eligibility section generation. RESULTS: Sentence embeddings fine-tuned using biomedical corpora produce clusters with the highest alignment to CT-level information. Human expert evaluation confirms that clusters are well structured and coherent. Despite the high information compression, clusters retain significant CT information, up to 97% of the classification performance obtained with raw embeddings. Finally, eligibility sections automatically generated using clusters achieve 95% of the ROUGE scores obtained with a generative LLM prompted with CT-protocol details, suggesting that clusters encapsulate information useful to CT-protocol design. DISCUSSION: Clusters derived from sentence-level LLM embeddings effectively summarize complex eligibility criterion data while retaining relevant CT-protocol details. Clustering-based approaches provide a scalable enhancement in CT design that balances information compression with accuracy. CONCLUSIONS: Clustering eligibility criteria using LLM embeddings provides a practical and efficient method to summarize critical protocol information. We provide an interactive visualization of the pipeline here. Alban Bornet, Philipp Khlebnikov, Florian Meer, Quentin Haas, Anthony Yazdani, Poorya Amini, Douglas Teodoro |
J. Am. Medical Informatics Assoc. | 8 |
| 2025 | A machine learning approach for automating review of a RxNorm medication mapping pipeline output
Matthias Hüser, John E. Doole, Vinicius Pinho, Hossein Rouhizadeh, Douglas Teodoro, Ahson Saiyed, Matvey Palchuk |
J. Biomed. Informatics | 5 |
| 2024 | PRIMIS: Privacy-preserving medical image sharing via deep sparsifying transform learning with obfuscationabstractOBJECTIVE: The primary objective of our study is to address the challenge of confidentially sharing medical images across different centers. This is often a critical necessity in both clinical and research environments, yet restrictions typically exist due to privacy concerns. Our aim is to design a privacy-preserving data-sharing mechanism that allows medical images to be stored as encoded and obfuscated representations in the public domain without revealing any useful or recoverable content from the images. In tandem, we aim to provide authorized users with compact private keys that could be used to reconstruct the corresponding images. METHOD: Our approach involves utilizing a neural auto-encoder. The convolutional filter outputs are passed through sparsifying transformations to produce multiple compact codes. Each code is responsible for reconstructing different attributes of the image. The key privacy-preserving element in this process is obfuscation through the use of specific pseudo-random noise. When applied to the codes, it becomes computationally infeasible for an attacker to guess the correct representation for all the codes, thereby preserving the privacy of the images. RESULTS: The proposed framework was implemented and evaluated using chest X-ray images for different medical image analysis tasks, including classification, segmentation, and texture analysis. Additionally, we thoroughly assessed the robustness of our method against various attacks using both supervised and unsupervised algorithms. CONCLUSION: This study provides a novel, optimized, and privacy-assured data-sharing mechanism for medical images, enabling multi-party sharing in a secure manner. While we have demonstrated its effectiveness with chest X-ray images, the mechanism can be utilized in other medical images modalities as well. Isaac Shiri, Behrooz Razeghi, Sohrab Ferdowsi, Yazdan Salimi, Deniz Gündüz, Douglas Teodoro, Sviatoslav Voloshynovskiy, Habib Zaidi |
J. Biomed. Informatics | 6 |
| 2023 | CardioBERTpt: Transformer-based Models for Cardiology Language Representation in PortugueseabstractContextual word embeddings and the Transformers architecture have reached state-of-the-art results in many natural language processing (NLP) tasks and improved the adaptation of models for multiple domains. Despite the improvement in the reuse and construction of models, few resources are still developed for the Portuguese language, especially in the health domain. Furthermore, the clinical models available for the language are not representative enough for all medical specialties. This work explores deep contextual embedding models for the Portuguese language to support clinical NLP tasks. We transferred learned information from electronic health records of a Brazilian tertiary hospital specialized in cardiology diseases and pre-trained multiple clinical BERT-based models. We evaluated the performance of these models in named entity recognition experiments, fine-tuning them in two annotated corpora containing clinical narratives. Our pre-trained models outperformed previous multilingual and Portuguese BERT-based models for cardiology and multi-specialty environments, reaching the state-of-the-art for analyzed corpora, with 5.5% F1 score improvement in TempClinBr (all entities) and 1.7% in SemClinBr (Disorder entity) corpora. Hence, we demonstrate that data representativeness and a high volume of training data can improve the results for clinical tasks, aligned with results for other languages. Elisa Terumi Rubel Schneider, Yohan Bonescki Gumiel, João Vitor Andrioli de Souza, Lilian Mie Mukai Cintho, Lucas Emanuel Silva e Oliveira, Marina de Sá Rebelo, Marco A. Gutierrez 0001, José Eduardo Krieger, Douglas Teodoro, Claudia Maria Cabral Moro Barra, Emerson Cabrera Paraiso |
CBMS | 9 |
| 2023 | Leveraging patient similarities via graph neural networks to predict phenotypes from temporal dataabstractSeveral machine learning approaches have been proposed to automatically derive clinical phenotypes from patient data. Nevertheless, methods leveraging similarity-based patient networks remain underexplored for temporal data. In this work, we propose a graph neural network (GNN) model that learns patient representation using different network configurations and feature modes. To explore the sequential nature of time series, features were extracted using a recurrent neural network (RNN) and embedded using information from the network structure via the GNN. Our method improves upon statistical and RNN baselines, with performance boosts up to 1% and 22% accuracy in the inductive and transductive settings, respectively. We also show that network configurations significantly impact performance in the transductive learning setting. Thus, automated phenotyping models based on GNNs could be used to support phenotype-based clinical research and ultimately for personalized clinical decision support.Data and Code Availability: This paper uses the MIMIC-III dataset [1], which is available on the PhysioNet repository [2]. The experiments are based on the public open source phenotyping benchmark of Harutyunyan et al. [3]. All our source code is publicly available at https://github.com/ds4dh/mimic3-benchmarks-GraDSCI23. Dimitrios Proios, Anthony Yazdani, Alban Bornet, Julien Ehrsam, Islem Rekik, Douglas Teodoro |
DSAA | 6 |
| 2022 | On Graph Construction for Classification of Clinical Trials Protocols Using Graph Neural Networks
Sohrab Ferdowsi, Jenny Copara, Racha Gouareb, Nikolay Borissov, Fernando Jaume-Santero, Poorya Amini, Douglas Teodoro |
AIME | 7 |
| 2021 | Classification of hierarchical text using geometric deep learning: the case of clinical trials corpusabstractWe consider the hierarchical representation of documents as graphs and use geometric deep learning to classify them into different categories.While graph neural networks can efficiently handle the variable structure of hierarchical documents using the permutation invariant message passing operations, we show that we can gain extra performance improvements using our proposed selective graph pooling operation that arises from the fact that some parts of the hierarchy are invariable across different documents.We applied our model to classify clinical trial (CT) protocols into completed and terminated categories.We use bag-of-words based, as well as pre-trained transformer-based embeddings to featurize the graph nodes, achieving f1-scores 0.85 on a publicly available large scale CT registry of around 360K protocols.We further demonstrate how the selective pooling can add insights into the CT termination status prediction.We make the source code and dataset splits accessible. Sohrab Ferdowsi, Nikolay Borissov, Julien Knafou, Poorya Amini, Douglas Teodoro |
EMNLP (1) | 5 |
| 2009 | QA-driven Guidelines Generation for Bacteriotherapy
Emilie Pasche, Douglas Teodoro, Julien Gobeill, Patrick Ruch, Christian Lovis |
AMIA | 2 |