Matthieu-P. Schapranow

dblp:92/7637 · also Matthieu-Patrick Schapranow · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
7since 2021 · last 2023
0000-0001-6601-2942ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorComputer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2023 GGTWEAK: Gene Tagging with Weak Supervision for German Clinical Text
Sandro Steinwand, Florian Borchert, Silvia Winkler, Matthieu-P. Schapranow
AIME4
2022 GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers
abstract
Despite remarkable advances in the development of language resources over the recent years, there is still a shortage of annotated, publicly available corpora covering (German) medical language. With the initial release of the German Guideline Program in Oncology NLP Corpus (GGPONC), we have demonstrated how such corpora can be built upon clinical guidelines, a widely available resource in many natural languages with a reasonable coverage of medical terminology. In this work, we describe a major new release for GGPONC. The corpus has been substantially extended in size and re-annotated with a new annotation scheme based on SNOMED CT top level hierarchies, reaching high inter-annotator agreement (γ=.94). Moreover, we annotated elliptical coordinated noun phrases and their resolutions, a common language phenomenon in (not only German) scientific documents. We also trained BERT-based named entity recognition models on this new data set, which achieve high performance on short, coarse-grained entity spans (F1=.89), while the rate of boundary errors increases for long entity spans. GGPONC is freely available through a data use agreement. The trained named entity recognition models, as well as the detailed annotation guide, are also made publicly available.
Florian Borchert, Christina Lohr, Luise Modersohn, Jonas Witt, Thomas Langer, Markus Follmann, Matthias Gietzelt, Bert Arnrich, Udo Hahn, Matthieu-P. Schapranow
LREC10
2021 Controversial Trials First: Identifying Disagreement Between Clinical Guidelines and New Evidence
Florian Borchert, Laura Meister, Thomas Langer, Markus Follmann, Bert Arnrich, Matthieu-P. Schapranow
AMIA6
2021 A Comparison of Concept Embeddings for German Clinical Corpora
abstract
Clinical concept embeddings enable unsupervised learning of relationships among medical concepts. A range of benchmarks quantifies the degree to which learned representations capture medical semantics. However, training and evaluation of embeddings require a large amount of data. In addition, embeddings’ benchmark score varies in different languages because it differs with the size of the available corpora. Multi-modal data increases the corpus size, but data protection regulations limit access to clinical multi-modal data. We present an extendable pipeline for training clinical concept embeddings on various text corpora and evaluating the quality of trained embeddings on selected benchmark tasks. Our work provides different ways to identify clinical concepts in textual corpora. We train embeddings on selected German clinical text corpora and evaluate them on various benchmark scores. Our work can be extended to train embeddings in other languages in which a large multi-modal dataset is not available.
Aadil Rasheed, Florian Borchert, Lasse Kohlmeyer, Richard Henkenjohann, Matthieu-P. Schapranow
BIBM5
2021 Using interpretability approaches to update "black-box" clinical prediction models: an external validation study in nephrology
Harry Freitas Da Cruz, Boris Pfahringer, Tom Martensen, Frederic Schneider, Alexander Meyer, Erwin P. Bottinger, Matthieu-P. Schapranow
Artif. Intell. Medicine7
2021 Knowledge bases and software support for variant interpretation in precision oncology
abstract
[14 June 2021]: notice amended to state that the funding information has been corrected to remove a duplicate reference to two funding organizations introduced in error during the initial correction. In the originally published version of this manuscript, requested author amendments to the Funding section were inadvertently omitted prior to publishing. The section should read: ``German Federal Ministry of Research and Education (01ZZ1802); Physician-Scientist Program of the University of Heidelberg, Faculty of Medicine, DKTK (German Cancer Consortium) School of Oncology and the Cancer Core Europe TRYTRAC program (to A.M.); MTB-Report project (VolkswagenStiftung ZN3424) (to J.H.).'' instead of ``German Federal Ministry of Research and Education (01ZZ1802); University of Heidelberg, Faculty of Medicine (to A.M.); Volkswagen Foundation (to J.H.).'' This has now been corrected.
Florian Borchert, Andreas Mock, Aurelie Tomczak, Jonas Hügel, Samer Alkarkoukly, Alexander Knurr, Anna-Lena Volckmar, Albrecht Stenzinger, Peter Schirmacher, Jürgen Debus, Dirk Jäger, Thomas Longerich, Stefan Fröhling, Roland Eils, Nina Bougatf, Ulrich Sax, Matthieu-P. Schapranow
Briefings Bioinform.17
2021 Knowledge bases and software support for variant interpretation in precision oncology
abstract
Precision oncology is a rapidly evolving interdisciplinary medical specialty. Comprehensive cancer panels are becoming increasingly available at pathology departments worldwide, creating the urgent need for scalable cancer variant annotation and molecularly informed treatment recommendations. A wealth of mainly academia-driven knowledge bases calls for software tools supporting the multi-step diagnostic process. We derive a comprehensive list of knowledge bases relevant for variant interpretation by a review of existing literature followed by a survey among medical experts from university hospitals in Germany. In addition, we review cancer variant interpretation tools, which integrate multiple knowledge bases. We categorize the knowledge bases along the diagnostic process in precision oncology and analyze programmatic access options as well as the integration of knowledge bases into software tools. The most commonly used knowledge bases provide good programmatic access options and have been integrated into a range of software tools. For the wider set of knowledge bases, access options vary across different parts of the diagnostic process. Programmatic access is limited for information regarding clinical classifications of variants and for therapy recommendations. The main issue for databases used for biological classification of pathogenic variants and pathway context information is the lack of standardized interfaces. There is no single cancer variant interpretation tool that integrates all identified knowledge bases. Specialized tools are available and need to be further developed for different steps in the diagnostic process.
Florian Borchert, Andreas Mock, Aurelie Tomczak, Jonas Hügel, Samer Alkarkoukly, Alexander Knurr, Anna-Lena Volckmar, Albrecht Stenzinger, Peter Schirmacher, Jürgen Debus, Dirk Jäger, Thomas Longerich, Stefan Fröhling, Roland Eils, Nina Bougatf, Ulrich Sax, Matthieu-P. Schapranow
Briefings Bioinform.17
2019 External Validation of a "Black-Box" Clinical Predictive Model in Nephrology: Can Interpretability Methods Help Illuminate Performance Differences?
Harry Freitas Da Cruz, Boris Pfahringer, Frederic Schneider, Alexander Meyer, Matthieu-P. Schapranow
AIME5
2019 MORPHER - A Platform to Support Modeling of Outcome and Risk Prediction in Health Research
abstract
Machine learning is rapidly becoming a mainstay in research and industry. Particularly for clinical predictive modeling, these approaches are being increasingly applied, as evidenced by the growth in the number of related publications. While different computer tools exist that support rapid prototyping, we observe that the state of the art is lacking in the extent to which the needs of research clinicians are addressed. This leads to an increase in the time needed for development and validation of such models. In this paper, we outline the requirements and challenges inherent to this domain and present a platform for rapid prototyping tailored to the specific needs of clinical modeling for outcome and risk prediction. We argue that a move towards hybrid solutions, i.e., a mix of cloud and on-premise infrastructure, constitutes a viable way to reduce the time needed to develop and validate clinical predictive models in a standardized, reproducible fashion.
Harry Freitas Da Cruz, Benjamin Bergner, Orhan Konak, Frederic Schneider, Philipp Bode, Conrad Lempert, Matthieu-P. Schapranow
BIBE7
2019 An Information and Communication Platform Supporting Analytics for Elderly Care
Orhan Konak, Harry Freitas Da Cruz, Marvin Thiele, David Golla, Matthieu-P. Schapranow
ICT4AWE5
2016 IMDBfs: Bridging the gap between in-memory database technology and file-based tools for life sciences
abstract
Many established processing and analysis tools in life sciences are still operating on files, e.g. alignment of genome data. Their integration into optimized workflows requires time-consuming transformation, import and export of data. In the given contribution, we introduce IMDBfs: a shared file system operating on top of an in-memory database system. Our IMDBfs provides transparent access to database resources via traditional file system operations. For the first time, it allows seamless integration of file-based life science tools into processes optimized for latest IMDB technology. As a result, existing tools neither need to be ported nor modified whilst transparent IMDB data access is provided.
Matthieu-P. Schapranow, Milena Kraus, Marius Danner, Hasso Plattner
BIBM1
2016 Towards an integrated health research process: A cloud-based approach
abstract
Today, health research and health care generate a steadily increasing amount of data. Making these available for secondary use cases is essential for efficiency gains in health research, e.g. by reducing time-and costs-intensive acquisition of data. In this contribution, we introduce our SAHRA software platform enabling reproducible research, e.g. by combining multiple data sources, performing data de-identification, and content filtering. We define an innovative research process combining retrospective and prospective research for the first time. Thus, authorized users, e.g. clinical researchers, are able to gain access through our system relevant research data and to perform interactive analyses. As a result, existing sensitive health data is securely transformed into de-identified research data, which can be used to improve future health research.
Matthieu-P. Schapranow, Matthias Uflacker, Murat Sariyar, Sebastian Claudius Semler, Johannes Klaus Fichte, Dietmar Schielke, Kismet Ekinci, Thomas Zahn
IEEE BigData1
2015 The Medical Knowledge Cockpit: Real-time analysis of big medical data enabling precision medicine
abstract
Significant medical knowledge has been generated over past decades, but is stored in databases distributed all over the globe using individual data formats. Detailed diagnostic tests result in steadily growing patient data. Consequently, medical experts are facing challenges outside of their field of expertise, e.g. analyzing, interpreting and linking medical data. In this contribution, we share details about our Medical Knowledge Cockpit, an application designed in an interdisciplinary cooperation with medical experts to improve the identification of relevant knowledge. Built upon latest in-memory database technology, it offers medical experts a unique starting point to link and analyze medical data on their own in real time. Result sets providing links back to primary data sources are assembled using patient specifics.
Matthieu-P. Schapranow, Milena Kraus, Cindy Perscheid, Cornelius Bock, Franz Liedke, Hasso Plattner
BIBM1
2014 Towards integrating the detection of genetic variants into an in-memory database
abstract
Next-generation sequencing enables whole genome sequencing within a few hours at a minimum of cost, entailing advanced medical applications such as personalized treatments. However, this recent technology imposes new challenges to alignment and variant calling as subsequent analysis steps. Compared to former sequencing, both must deal with an increasing amount of data to process at a significantly lower data quality - and are currently not capable of that. In this work, we focus on addressing these challenges for identifying Single Nucleotide Polymorphisms, i.e. SNP calling, in genome data as one subtask of variant calling. We propose the application of a column-store in-memory database for efficient data processing and apply the statistical model that is provided by the Genome Analysis Toolkit's UnifiedGenotyper. Comparisons with the UnifiedGenotyper show that our approach can exploit all computational resources available and accelerates SNP calling up to a factor of 22x.
Cindy Fähnrich, Matthieu-P. Schapranow, Hasso Plattner
IEEE BigData2
2014 In-memory technology enables interactive drug response analysis
abstract
Latest medical diagnostics generate increasing amounts of big medical data. Specific software tools optimized for the use by healthcare experts and researchers as well as systematic processes for data processing and analysis in clinical and research environments are still missing. Our work focuses on the integration of high-throughput next-generation sequencing data and its systematic processing and its instantaneous analysis to use them in the course of precision medicine. We share our research results on designing a generic research process for drug response analysis including specific software tools built on top of our distributed in-memory computing platform for processing of big medical data. Furthermore, we present our technical foundations as well as process aspects of integrating and combining heterogeneous data sources, such as genome, patient, and experimental data.
Matthieu-P. Schapranow, Konrad Klinghammer, Cindy Fähnrich, Hasso Plattner
Healthcom1
2013 HIG - An in-memory database platform enabling real-time analyses of genome data
abstract
Costs and time required for sequencing of DNA and RNA declined through use of next generation sequencing technology, e.g. up to 30-times coverage reads are generated in less than two days. However, its interpretation and analysis is still a time-consuming process potentially taking weeks. In this work, we present a completely new architecture for processing and analyzing genome data. It builds on the in-memory database technology to eliminate time-consuming file-based data operations and to enable real-time data analysis. We found out that the use of in-memory technology as an integral component for genome data processing and its analysis significantly reduces time and costs to obtain relevant results, e.g. in the course of personalized medicine.
Matthieu-P. Schapranow, Hasso Plattner
IEEE BigData1
2012 Costs of authentic pharmaceuticals: research on qualitative and quantitative aspects of enabling anti-counterfeiting in RFID-aided supply chains
Matthieu-P. Schapranow, Jürgen Müller 0003, Alexander Zeier, Hasso Plattner
Pers. Ubiquitous Comput.1
2011 Security Extensions for Improving Data Security of Event Repositories in EPCglobal Networks
abstract
Location-based event data is captured in RFID-aided supply chains for tacking individual goods. They are stored in distributed event repositories by involved supply chain parties. Performing anti-counterfeiting checks involves exchange of event data without exposure of sensitive business secrets. Current EPC global standards leave the definition of security strategies open for concrete implementation. We consider data security as the major aspect that needs to be clarified before wide adaption of EPC global standards will be considered by industries. The given work contributes by defining security extensions for EPC global networks and sharing implementation details of our research prototype. We show that incorporating in-memory technology enables history-based access control while keeping response times fast.
Matthieu-P. Schapranow, Alexander Zeier, Hasso Plattner
EUC1
2010 Enabling real-time charging for smart grid infrastructures using in-memory databases
abstract
The emerging trend towards smart grids defines new requirements for designing enterprise applications for the energy market. Current solutions were built to process single billing runs as time-consuming batch jobs. Rather than processing some readings per year and household, a constant stream of meter readings has to be processed in context of a smart power grid. Additionally, consumers demand for convenient ways to monitor power consumption while getting real-time charging information on a daily basis. In this paper, we share our experiences of integrating meter readings in an industry-specific enterprise application for utilities to support real-time billing. Our prototype enables customers to track consumptions and to view billing details online. Besides, we discuss new business scenarios enabled by this real-timeliness.
Matthieu-P. Schapranow, Ralph Kühne, Alexander Zeier, Hasso Plattner
LCN1
2010 A Dynamic Mutual RFID Authentication Model Preventing Unauthorized Third Party Access
abstract
RFID implementations leverage competitive business advantages in processing, tracking and tracing of fast-moving goods. Most of them suffer from security threats and the resulting privacy risks as RFID technology was not designed for exchange of sensible data. Emerging global RFID-aided supply chains require open interfaces for data exchange of confidential business data between business partners. We present a mutual authentication model based on one-time passwords preventing tag access by unauthorized third parties. Compared to models using complex on-tag encryption methods our implementation focuses on reducing tag-manufacturing costs while increasing customers' acceptance for RFID technology.
Matthieu-P. Schapranow, Alexander Zeier, Hasso Plattner
NSS1