EDBT 2026 Demo / reviewers in the wild / expert
Markus Kreuzthaler
dblp:55/8726
· DBLP profile ↗
13ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-9824-9004ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ontology-Aligned Clinical NER in Emergency Medicine Using Weak Supervision
Amila Kugic, Adrian Schiegl, Alistair Tiefenbacher, Lukas Seper, Stefan Schulz 0001, Markus Kreuzthaler |
AIME (2) | 6 |
| 2025 | Zero- and few-shot Named Entity Recognition and Text Expansion in medication prescriptions using large language modelsabstractMedication prescriptions in electronic health records (EHR) are often in free-text and may include a mix of languages, local brand names, and a wide range of idiosyncratic formats and abbreviations. Large language models (LLMs) have shown a promising ability to generate text in response to input prompts. We use ChatGPT3.5 to automatically structure and expand medication statements in discharge summaries and thus make them easier to interpret for people and machines. Named Entity Recognition (NER) and Text Expansion (EX) are used with different prompt strategies in a zero- and few-shot setting. 100 medication statements were manually annotated and curated. NER performance was measured by using strict and partial matching. For the EX task, two experts interpreted the results by assessing semantic equivalence between original and expanded statements. The model performance was measured by precision, recall, and F1 score. For NER, the best-performing prompt reached an average F1 score of 0.94 in the test set. For EX, the few-shot prompt showed superior performance among other prompts, with an average F1 score of 0.87. Our study demonstrates good performance for NER and EX tasks in free-text medication statements using ChatGPT3.5. Compared to a zero-shot baseline, a few-shot approach prevented the system from hallucinating, which is essential when processing safety-relevant medication data. We tested ChatGPT3.5-tuned prompts on other LLMs, including ChatGPT4o, Gemini 2.0 Flash, MedLM-1.5-Large, and DeepSeekV3. The findings showed most models outperformed ChatGPT3.5 in NER and EX tasks. Natthanaphop Isaradech, Andrea Riedel, Wachiranun Sirikul, Markus Kreuzthaler, Stefan Schulz 0001 |
Artif. Intell. Medicine | 4 |
| 2024 | Smoking Status Classification: A Comparative Analysis of Machine Learning Techniques with Clinical Real World Data
Amila Kugic, Akhila Abdulnazar, Anto Knezovic, Stefan Schulz 0001, Markus Kreuzthaler |
AIME (1) | 5 |
| 2024 | Sequence-Model-Based Medication Extraction from Clinical Narratives in German
Vishakha Sharma 0004, Andreas Thalhammer 0001, Amila Kugic, Stefan Schulz 0001, Markus Kreuzthaler |
AIME (1) | 5 |
| 2024 | Federated unsupervised random forest for privacy-preserving patient stratificationabstractMOTIVATION: In the realm of precision medicine, effective patient stratification and disease subtyping demand innovative methodologies tailored for multi-omics data. Clustering techniques applied to multi-omics data have become instrumental in identifying distinct subgroups of patients, enabling a finer-grained understanding of disease variability. Meanwhile, clinical datasets are often small and must be aggregated from multiple hospitals. Online data sharing, however, is seen as a significant challenge due to privacy concerns, potentially impeding big data's role in medical advancements using machine learning. This work establishes a powerful framework for advancing precision medicine through unsupervised random forest-based clustering in combination with federated computing. RESULTS: We introduce a novel multi-omics clustering approach utilizing unsupervised random forests. The unsupervised nature of the random forest enables the determination of cluster-specific feature importance, unraveling key molecular contributors to distinct patient groups. Our methodology is designed for federated execution, a crucial aspect in the medical domain where privacy concerns are paramount. We have validated our approach on machine learning benchmark datasets as well as on cancer data from The Cancer Genome Atlas. Our method is competitive with the state-of-the-art in terms of disease subtyping, but at the same time substantially improves the cluster interpretability. Experiments indicate that local clustering performance can be improved through federated computing. AVAILABILITY AND IMPLEMENTATION: The proposed methods are available as an R-package (https://github.com/pievos101/uRF). Bastian Pfeifer, Christel Sirocchi, Marcus D. Bloice, Markus Kreuzthaler, Martin Urschler |
Bioinform. | 4 |
| 2024 | Disambiguation of acronyms in clinical narratives with large language modelsabstractOBJECTIVE: To assess the performance of large language models (LLMs) for zero-shot disambiguation of acronyms in clinical narratives. MATERIALS AND METHODS: Clinical narratives in English, German, and Portuguese were applied for testing the performance of four LLMs: GPT-3.5, GPT-4, Llama-2-7b-chat, and Llama-2-70b-chat. For English, the anonymized Clinical Abbreviation Sense Inventory (CASI, University of Minnesota) was used. For German and Portuguese, at least 500 text spans were processed. The output of LLM models, prompted with contextual information, was analyzed to compare their acronym disambiguation capability, grouped by document-level metadata, the source language, and the LLM. RESULTS: On CASI, GPT-3.5 achieved 0.91 in accuracy. GPT-4 outperformed GPT-3.5 across all datasets, reaching 0.98 in accuracy for CASI, 0.86 and 0.65 for two German datasets, and 0.88 for Portuguese. Llama models only reached 0.73 for CASI and failed severely for German and Portuguese. Across LLMs, performance decreased from English to German and Portuguese processing languages. There was no evidence that additional document-level metadata had a significant effect. CONCLUSION: For English clinical narratives, acronym resolution by GPT-4 can be recommended to improve readability of clinical text by patients and professionals. For German and Portuguese, better models are needed. Llama models, which are particularly interesting for processing sensitive content on premise, cannot yet be recommended for acronym resolution. Amila Kugic, Stefan Schulz 0001, Markus Kreuzthaler |
J. Am. Medical Informatics Assoc. | 3 |
| 2023 | Identification of Non-Lexical Content in Croatian Health Forum EntriesabstractMedical texts often contain expressions that are not listed in biomedical dictionaries and terminology systems. In this investigation, the detection of such entities is examined with online health forum entries in the Croatian language. Emphasis is put on short-form content, lexical variations, brand names, and proper names. By leveraging Transformer architectures on token and entity level, noteworthy results (>90% in F1-measure) could be achieved, which is on par with state-of-the-art named entity recognition in high-resource languages. Additionally, this investigation showcases a way to recognize non-lexicalized tokens from texts, which can be of use for further research questions in combination with word sense disambiguation, word-embeddings, and terminology expansion workflows. Amila Kugic, Stefan Schulz 0001, Markus Kreuzthaler |
BIBM | 3 |
| 2023 | Embedding-based terminology expansion via secondary use of large clinical real-world datasetsabstractA log-likelihood based co-occurrence analysis of ∼1.9 million de-identified ICD-10 codes and related short textual problem list entries generated possible term candidates at a significance level of p<0.01. These top 10 term candidates, consisting of 1 to 5-grams, were used as seed terms for an embedding based nearest neighbor approach to fetch additional synonyms, hypernyms and hyponyms in the respective n-gram embedding spaces by leveraging two different language models. This was done to analyze the lexicality of the resulting term candidates and to compare the term classifications of both models. We found no difference in system performance during the processing of lexical and non-lexical content, i.e. abbreviations, acronyms, etc. Additionally, an application-oriented analysis of the SapBERT (Self-Alignment Pretraining for Biomedical Entity Representations) language model indicates suitable performance for the extraction of all term classifications such as synonyms, hypernyms, and hyponyms. Amila Kugic, Bastian Pfeifer, Stefan Schulz 0001, Markus Kreuzthaler |
J. Biomed. Informatics | 4 |
| 2019 | Evaluating shallow and deep learning strategies for the 2018 n2c2 shared task on clinical text classificationabstractOBJECTIVE: Automated clinical phenotyping is challenging because word-based features quickly turn it into a high-dimensional problem, in which the small, privacy-restricted, training datasets might lead to overfitting. Pretrained embeddings might solve this issue by reusing input representation schemes trained on a larger dataset. We sought to evaluate shallow and deep learning text classifiers and the impact of pretrained embeddings in a small clinical dataset. MATERIALS AND METHODS: We participated in the 2018 National NLP Clinical Challenges (n2c2) Shared Task on cohort selection and received an annotated dataset with medical narratives of 202 patients for multilabel binary text classification. We set our baseline to a majority classifier, to which we compared a rule-based classifier and orthogonal machine learning strategies: support vector machines, logistic regression, and long short-term memory neural networks. We evaluated logistic regression and long short-term memory using both self-trained and pretrained BioWordVec word embeddings as input representation schemes. RESULTS: Rule-based classifier showed the highest overall micro F1 score (0.9100), with which we finished first in the challenge. Shallow machine learning strategies showed lower overall micro F1 scores, but still higher than deep learning strategies and the baseline. We could not show a difference in classification efficiency between self-trained and pretrained embeddings. DISCUSSION: Clinical context, negation, and value-based criteria hindered shallow machine learning approaches, while deep learning strategies could not capture the term diversity due to the small training dataset. CONCLUSION: Shallow methods for clinical phenotyping can still outperform deep learning methods in small imbalanced data, even when supported by pretrained embeddings. Michel Oleynik, Amila Kugic, Zdenko Kasác, Markus Kreuzthaler |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | SEMCARE - Semantic Data Platform for Healthcare
Philipp Daumke, Claudia Riede, Thomas Fassbender, Angel Honrado, Markus Kreuzthaler, Pablo López-García, Stefan Schulz 0001, Erik M. van Mulligen, Herman van Haagen, Jan A. Kors, Hanney Gonna, Xinkai Wang 0002, Elijah Behr |
AMIA | 5 |
| 2015 | Knowledge Extraction from MEDLINE by Combining Clustering with Natural Language Processing
José Antonio Miñarro-Giménez, Markus Kreuzthaler, Stefan Schulz 0001 |
AMIA | 2 |
| 2015 | Secondary use of electronic health records for building cohort studies through top-down information extraction
Markus Kreuzthaler, Stefan Schulz 0001, Andrea Berghold |
J. Biomed. Informatics | 1 |
| 2012 | Metonymies in Medical Terminologies. A SNOMED CT Case Study
Markus Kreuzthaler, Stefan Schulz 0001 |
AMIA | 1 |