Florian Borchert

dblp:270/1909 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-1079-6500ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Electronic Health Records-based identification of newly diagnosed Crohn's Disease cases
Susanne Ibing, Julian Hugo, Florian Borchert, Linea Schmidt, Caroline Benson, Allison A. Marshall, Colleen Chasteau, Ujunwa Korie, Diana Paguay, Jan-Philipp Sachs, Bernhard Y. Renard, Judy H. Cho, Erwin P. Bottinger, Ryan C. Ungaro
Artif. Intell. Medicine3
2024 medBERT.de: A comprehensive German BERT model for the medical domain
abstract
This paper presents medBERT.de , a pre-trained German BERT model specifically designed for the German medical domain. The model has been trained on a large corpus of 4.7 Million German medical documents and has been shown to achieve new state-of-the-art performance on eight different medical benchmarks covering a wide range of disciplines and medical document types. In addition to evaluating the overall performance of the model, this paper also conducts a more in-depth analysis of its capabilities. We investigate the impact of data deduplication on the model's performance, as well as the potential benefits of using more efficient tokenization methods. Our results indicate that domain-specific models such as medBERT.de are particularly useful for longer texts, and that deduplication of training data does not necessarily lead to improved performance. Furthermore, we found that efficient tokenization plays only a minor role in improving model performance, and attribute most of the improved performance to the large amount of training data. To encourage further research, the pre-trained model weights and new benchmarks based on radiological data are made publicly available for use by the scientific community.
Keno Bressem, Jens-Michalis Papaioannou, Paul Grundmann, Florian Borchert, Lisa Adams, Leonhard Liu, Felix Busch, Jan P. Loyen, Stefan Markus Niehues, Moritz Augustin, Lennart Grosser, Marcus R. Makowski, Hugo J. W. L. Aerts, Alexander Löser
Expert Syst. Appl.4
2023 Machine Learning Based Prediction of Incident Cases of Crohn's Disease Using Electronic Health Records from a Large Integrated Health System
Julian Hugo, Susanne Ibing, Florian Borchert, Jan-Philipp Sachs, Judy Cho, Ryan C. Ungaro, Erwin P. Bottinger
AIME3
2023 GGTWEAK: Gene Tagging with Weak Supervision for German Clinical Text
Sandro Steinwand, Florian Borchert, Silvia Winkler, Matthieu-P. Schapranow
AIME2
2022 GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers
abstract
Despite remarkable advances in the development of language resources over the recent years, there is still a shortage of annotated, publicly available corpora covering (German) medical language. With the initial release of the German Guideline Program in Oncology NLP Corpus (GGPONC), we have demonstrated how such corpora can be built upon clinical guidelines, a widely available resource in many natural languages with a reasonable coverage of medical terminology. In this work, we describe a major new release for GGPONC. The corpus has been substantially extended in size and re-annotated with a new annotation scheme based on SNOMED CT top level hierarchies, reaching high inter-annotator agreement (γ=.94). Moreover, we annotated elliptical coordinated noun phrases and their resolutions, a common language phenomenon in (not only German) scientific documents. We also trained BERT-based named entity recognition models on this new data set, which achieve high performance on short, coarse-grained entity spans (F1=.89), while the rate of boundary errors increases for long entity spans. GGPONC is freely available through a data use agreement. The trained named entity recognition models, as well as the detailed annotation guide, are also made publicly available.
Florian Borchert, Christina Lohr, Luise Modersohn, Jonas Witt, Thomas Langer, Markus Follmann, Matthias Gietzelt, Bert Arnrich, Udo Hahn, Matthieu-P. Schapranow
LREC1
2021 Controversial Trials First: Identifying Disagreement Between Clinical Guidelines and New Evidence
Florian Borchert, Laura Meister, Thomas Langer, Markus Follmann, Bert Arnrich, Matthieu-P. Schapranow
AMIA1
2021 A Comparison of Concept Embeddings for German Clinical Corpora
abstract
Clinical concept embeddings enable unsupervised learning of relationships among medical concepts. A range of benchmarks quantifies the degree to which learned representations capture medical semantics. However, training and evaluation of embeddings require a large amount of data. In addition, embeddings’ benchmark score varies in different languages because it differs with the size of the available corpora. Multi-modal data increases the corpus size, but data protection regulations limit access to clinical multi-modal data. We present an extendable pipeline for training clinical concept embeddings on various text corpora and evaluating the quality of trained embeddings on selected benchmark tasks. Our work provides different ways to identify clinical concepts in textual corpora. We train embeddings on selected German clinical text corpora and evaluate them on various benchmark scores. Our work can be extended to train embeddings in other languages in which a large multi-modal dataset is not available.
Aadil Rasheed, Florian Borchert, Lasse Kohlmeyer, Richard Henkenjohann, Matthieu-P. Schapranow
BIBM2
2021 Knowledge bases and software support for variant interpretation in precision oncology
abstract
[14 June 2021]: notice amended to state that the funding information has been corrected to remove a duplicate reference to two funding organizations introduced in error during the initial correction. In the originally published version of this manuscript, requested author amendments to the Funding section were inadvertently omitted prior to publishing. The section should read: ``German Federal Ministry of Research and Education (01ZZ1802); Physician-Scientist Program of the University of Heidelberg, Faculty of Medicine, DKTK (German Cancer Consortium) School of Oncology and the Cancer Core Europe TRYTRAC program (to A.M.); MTB-Report project (VolkswagenStiftung ZN3424) (to J.H.).'' instead of ``German Federal Ministry of Research and Education (01ZZ1802); University of Heidelberg, Faculty of Medicine (to A.M.); Volkswagen Foundation (to J.H.).'' This has now been corrected.
Florian Borchert, Andreas Mock, Aurelie Tomczak, Jonas Hügel, Samer Alkarkoukly, Alexander Knurr, Anna-Lena Volckmar, Albrecht Stenzinger, Peter Schirmacher, Jürgen Debus, Dirk Jäger, Thomas Longerich, Stefan Fröhling, Roland Eils, Nina Bougatf, Ulrich Sax, Matthieu-P. Schapranow
Briefings Bioinform.1
2021 Knowledge bases and software support for variant interpretation in precision oncology
abstract
Precision oncology is a rapidly evolving interdisciplinary medical specialty. Comprehensive cancer panels are becoming increasingly available at pathology departments worldwide, creating the urgent need for scalable cancer variant annotation and molecularly informed treatment recommendations. A wealth of mainly academia-driven knowledge bases calls for software tools supporting the multi-step diagnostic process. We derive a comprehensive list of knowledge bases relevant for variant interpretation by a review of existing literature followed by a survey among medical experts from university hospitals in Germany. In addition, we review cancer variant interpretation tools, which integrate multiple knowledge bases. We categorize the knowledge bases along the diagnostic process in precision oncology and analyze programmatic access options as well as the integration of knowledge bases into software tools. The most commonly used knowledge bases provide good programmatic access options and have been integrated into a range of software tools. For the wider set of knowledge bases, access options vary across different parts of the diagnostic process. Programmatic access is limited for information regarding clinical classifications of variants and for therapy recommendations. The main issue for databases used for biological classification of pathogenic variants and pathway context information is the lack of standardized interfaces. There is no single cancer variant interpretation tool that integrates all identified knowledge bases. Specialized tools are available and need to be further developed for different steps in the diagnostic process.
Florian Borchert, Andreas Mock, Aurelie Tomczak, Jonas Hügel, Samer Alkarkoukly, Alexander Knurr, Anna-Lena Volckmar, Albrecht Stenzinger, Peter Schirmacher, Jürgen Debus, Dirk Jäger, Thomas Longerich, Stefan Fröhling, Roland Eils, Nina Bougatf, Ulrich Sax, Matthieu-P. Schapranow
Briefings Bioinform.1