EDBT 2026 Demo / reviewers in the wild / expert
Stefan Schulz 0001
dblp:91/3020-1
· DBLP profile ↗
86ranked-venue papers
28as first author
8since 2021 · last 2026
0000-0001-7222-3287ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 57 · 16 first-author · 6 since 2021Artificial intelligence and machine learning · 28 · 12 first-author · 2 since 2021Theory of computation · 8 · 7 first-authorDatabases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ontology-Aligned Clinical NER in Emergency Medicine Using Weak Supervision
Amila Kugic, Adrian Schiegl, Alistair Tiefenbacher, Lukas Seper, Stefan Schulz 0001, Markus Kreuzthaler |
AIME (2) | 5 |
| 2026 | Developing the German Medical Text Corpus (GeMTeX): Legal Compliance and Semantic Enrichment
Justin Hofenbitzer, Christina Lohr, Andrea Riedel, Rebekka Kiser, Aliaksandra Shutsko, Abanoub Abdelmalak, Peter Klügl, Jutta Romberg, Sarah Riepenhausen, Miriam Schechner, Jakob Faller, Frank A. Meineke, Luise Modersohn, Markus Löffler, Juliane Fluck, Udo Hahn, Stefan Schulz 0001, Martin Boeker |
LREC | 17 |
| 2025 | Zero- and few-shot Named Entity Recognition and Text Expansion in medication prescriptions using large language modelsabstractMedication prescriptions in electronic health records (EHR) are often in free-text and may include a mix of languages, local brand names, and a wide range of idiosyncratic formats and abbreviations. Large language models (LLMs) have shown a promising ability to generate text in response to input prompts. We use ChatGPT3.5 to automatically structure and expand medication statements in discharge summaries and thus make them easier to interpret for people and machines. Named Entity Recognition (NER) and Text Expansion (EX) are used with different prompt strategies in a zero- and few-shot setting. 100 medication statements were manually annotated and curated. NER performance was measured by using strict and partial matching. For the EX task, two experts interpreted the results by assessing semantic equivalence between original and expanded statements. The model performance was measured by precision, recall, and F1 score. For NER, the best-performing prompt reached an average F1 score of 0.94 in the test set. For EX, the few-shot prompt showed superior performance among other prompts, with an average F1 score of 0.87. Our study demonstrates good performance for NER and EX tasks in free-text medication statements using ChatGPT3.5. Compared to a zero-shot baseline, a few-shot approach prevented the system from hallucinating, which is essential when processing safety-relevant medication data. We tested ChatGPT3.5-tuned prompts on other LLMs, including ChatGPT4o, Gemini 2.0 Flash, MedLM-1.5-Large, and DeepSeekV3. The findings showed most models outperformed ChatGPT3.5 in NER and EX tasks. Natthanaphop Isaradech, Andrea Riedel, Wachiranun Sirikul, Markus Kreuzthaler, Stefan Schulz 0001 |
Artif. Intell. Medicine | 5 |
| 2024 | Smoking Status Classification: A Comparative Analysis of Machine Learning Techniques with Clinical Real World Data
Amila Kugic, Akhila Abdulnazar, Anto Knezovic, Stefan Schulz 0001, Markus Kreuzthaler |
AIME (1) | 4 |
| 2024 | Sequence-Model-Based Medication Extraction from Clinical Narratives in German
Vishakha Sharma 0004, Andreas Thalhammer 0001, Amila Kugic, Stefan Schulz 0001, Markus Kreuzthaler |
AIME (1) | 4 |
| 2024 | Disambiguation of acronyms in clinical narratives with large language modelsabstractOBJECTIVE: To assess the performance of large language models (LLMs) for zero-shot disambiguation of acronyms in clinical narratives. MATERIALS AND METHODS: Clinical narratives in English, German, and Portuguese were applied for testing the performance of four LLMs: GPT-3.5, GPT-4, Llama-2-7b-chat, and Llama-2-70b-chat. For English, the anonymized Clinical Abbreviation Sense Inventory (CASI, University of Minnesota) was used. For German and Portuguese, at least 500 text spans were processed. The output of LLM models, prompted with contextual information, was analyzed to compare their acronym disambiguation capability, grouped by document-level metadata, the source language, and the LLM. RESULTS: On CASI, GPT-3.5 achieved 0.91 in accuracy. GPT-4 outperformed GPT-3.5 across all datasets, reaching 0.98 in accuracy for CASI, 0.86 and 0.65 for two German datasets, and 0.88 for Portuguese. Llama models only reached 0.73 for CASI and failed severely for German and Portuguese. Across LLMs, performance decreased from English to German and Portuguese processing languages. There was no evidence that additional document-level metadata had a significant effect. CONCLUSION: For English clinical narratives, acronym resolution by GPT-4 can be recommended to improve readability of clinical text by patients and professionals. For German and Portuguese, better models are needed. Llama models, which are particularly interesting for processing sensitive content on premise, cannot yet be recommended for acronym resolution. Amila Kugic, Stefan Schulz 0001, Markus Kreuzthaler |
J. Am. Medical Informatics Assoc. | 2 |
| 2023 | Identification of Non-Lexical Content in Croatian Health Forum EntriesabstractMedical texts often contain expressions that are not listed in biomedical dictionaries and terminology systems. In this investigation, the detection of such entities is examined with online health forum entries in the Croatian language. Emphasis is put on short-form content, lexical variations, brand names, and proper names. By leveraging Transformer architectures on token and entity level, noteworthy results (>90% in F1-measure) could be achieved, which is on par with state-of-the-art named entity recognition in high-resource languages. Additionally, this investigation showcases a way to recognize non-lexicalized tokens from texts, which can be of use for further research questions in combination with word sense disambiguation, word-embeddings, and terminology expansion workflows. Amila Kugic, Stefan Schulz 0001, Markus Kreuzthaler |
BIBM | 2 |
| 2023 | Embedding-based terminology expansion via secondary use of large clinical real-world datasetsabstractA log-likelihood based co-occurrence analysis of ∼1.9 million de-identified ICD-10 codes and related short textual problem list entries generated possible term candidates at a significance level of p<0.01. These top 10 term candidates, consisting of 1 to 5-grams, were used as seed terms for an embedding based nearest neighbor approach to fetch additional synonyms, hypernyms and hyponyms in the respective n-gram embedding spaces by leveraging two different language models. This was done to analyze the lexicality of the resulting term candidates and to compare the term classifications of both models. We found no difference in system performance during the processing of lexical and non-lexical content, i.e. abbreviations, acronyms, etc. Additionally, an application-oriented analysis of the SapBERT (Self-Alignment Pretraining for Biomedical Entity Representations) language model indicates suitable performance for the extraction of all term classifications such as synonyms, hypernyms, and hyponyms. Amila Kugic, Bastian Pfeifer, Stefan Schulz 0001, Markus Kreuzthaler |
J. Biomed. Informatics | 3 |
| 2020 | Risk prediction of delirium in hospitalized patients using machine learning: An implementation and prospective evaluation studyabstractOBJECTIVE: Machine learning models trained on electronic health records have achieved high prognostic accuracy in test datasets, but little is known about their embedding into clinical workflows. We implemented a random forest-based algorithm to identify hospitalized patients at high risk for delirium, and evaluated its performance in a clinical setting. MATERIALS AND METHODS: Delirium was predicted at admission and recalculated on the evening of admission. The defined prediction outcome was a delirium coded for the recent hospital stay. During 7 months of prospective evaluation, 5530 predictions were analyzed. In addition, 119 predictions for internal medicine patients were compared with ratings of clinical experts in a blinded and nonblinded setting. RESULTS: During clinical application, the algorithm achieved a sensitivity of 74.1% and a specificity of 82.2%. Discrimination on prospective data (area under the receiver-operating characteristic curve = 0.86) was as good as in the test dataset, but calibration was poor. The predictions correlated strongly with delirium risk perceived by experts in the blinded (r = 0.81) and nonblinded (r = 0.62) settings. A major advantage of our setting was the timely prediction without additional data entry. DISCUSSION: The implemented machine learning algorithm achieved a stable performance predicting delirium in high agreement with expert ratings, but improvement of calibration is needed. Future research should evaluate the acceptance of implemented machine learning algorithms by health professionals. CONCLUSIONS: Our study provides new insights into the implementation process of a machine learning algorithm into a clinical workflow and demonstrates its predictive power for delirium. Stefanie Jauk, Diether Kramer, Birgit Großauer, Susanne Rienmüller, Alexander Avian, Andrea Berghold, Werner Leodolter, Stefan Schulz 0001 |
J. Am. Medical Informatics Assoc. | 8 |
| 2018 | Towards an Ontology of Religious and Spiritual BeliefabstractRealist ontologies claim to represent what exists. However, human behaviour and culture is deeply influenced by religious and spiritual belief, whose veracity is highly controversial. Such beliefs are nevertheless known to have substantial impact on well-being, social behaviour, health and disease. It is therefore desirable to be able to represent beliefs within a principled ontology using the Web Ontology Language (OWL) as a standardised representation language. This paper demonstrates how a realist ontology, expressed in description logics, can deal with such entities without requiring consensus about their existence in reality. We present several ontology design patterns which allow for taxonomically arranging elements of religious or spiritual belief systems. This provides a framework on which data on particular beliefs can be better standardised for research in humanities, social research and life sciences. Stefan Schulz 0001, Ludger Jansen |
FOIS | 1 |
| 2017 | Ontological Representation of Laboratory Test Observables: Challenges and Perspectives in the SNOMED CT Observable Entity Model Adoption
Mélissa Mary, Lina Fatima Soualmia, Xavier Gansel, Stéfan Jacques Darmoni, Daniel Karlsson, Stefan Schulz 0001 |
AIME | 6 |
| 2017 | Is the Application of SNOMED CT Concept Model sufficiently Quality Assured?
Jean Marie Rodrigues, Stefan Schulz 0001, Bassim Mizen, Alan L. Rector, sofyane Serir |
AMIA | 2 |
| 2017 | Validating EHR clinical models using ontology patterns
Catalina Martínez-Costa, Stefan Schulz 0001 |
J. Biomed. Informatics | 2 |
| 2015 | SEMCARE - Semantic Data Platform for Healthcare
Philipp Daumke, Claudia Riede, Thomas Fassbender, Angel Honrado, Markus Kreuzthaler, Pablo López-García, Stefan Schulz 0001, Erik M. van Mulligen, Herman van Haagen, Jan A. Kors, Hanney Gonna, Xinkai Wang 0002, Elijah Behr |
AMIA | 7 |
| 2015 | JUFIT: A Configurable Rule Engine for Filtering and Generating New Multilingual UMLS Terms
Johannes Hellrich, Stefan Schulz 0001, Sven Buechel, Udo Hahn |
AMIA | 2 |
| 2015 | Knowledge Extraction from MEDLINE by Combining Clustering with Natural Language Processing
José Antonio Miñarro-Giménez, Markus Kreuzthaler, Stefan Schulz 0001 |
AMIA | 3 |
| 2015 | Feasibility of an ontology driven tumor-node-metastasis classifier application: A study on colorectal cancerabstractThe objectives of this work are (1) to develop a classifier application for tumor staging based on a formal representation of the Tumor-Node-Metastasis classification system (TNM), and (2) to show the feasibility of this approach on real data. This paper presents a classifier application for colorectal tumors based on the TNM-O ontology. It was developed in the JAVA using the OWL-API. The TNM-O uses the Foundational Model of Anatomy for representing anatomical entities and BioTopLite2 as a domain-top-level ontology. The classifier application processes input data via a user interface or tabular data. The classification starts with the creation of RDF Individuals for each pathological information item formally described in the ontology. These Individuals are then classified by the HermiT Description Logics reasoner by A-Box classification. A dataset with 382 entries was provided by the pathology department of a university hospital. It was automatically classified with regard to metastatic regional lymph nodes. Results or expert classification by pathologists and automatic classification were compared. The automatic process helped to detect and explain inconsistencies between expert and automatic classifications. This work, we demonstrate the use of semantic technologies in a TNM classifier application separating underlying medical knowledge represented in OWL from process logics. The presented prototypical TNM classifier application shows the potential to be integrated in larger software systems. Fábio França, Stefan Schulz 0001, Peter Bronsert, Paulo Novais, Martin Boeker |
INISTA | 2 |
| 2015 | Semantic enrichment of clinical models towards semantic interoperability. The heart failure summary use caseabstractOBJECTIVE: To improve semantic interoperability of electronic health records (EHRs) by ontology-based mediation across syntactically heterogeneous representations of the same or similar clinical information. MATERIALS AND METHODS: Our approach is based on a semantic layer that consists of: (1) a set of ontologies supported by (2) a set of semantic patterns. The first aspect of the semantic layer helps standardize the clinical information modeling task and the second shields modelers from the complexity of ontology modeling. We applied this approach to heterogeneous representations of an excerpt of a heart failure summary. RESULTS: Using a set of finite top-level patterns to derive semantic patterns, we demonstrate that those patterns, or compositions thereof, can be used to represent information from clinical models. Homogeneous querying of the same or similar information, when represented according to heterogeneous clinical models, is feasible. DISCUSSION: Our approach focuses on the meaning embedded in EHRs, regardless of their structure. This complex task requires a clear ontological commitment (ie, agreement to consistently use the shared vocabulary within some context), together with formalization rules. These requirements are supported by semantic patterns. Other potential uses of this approach, such as clinical models validation, require further investigation. CONCLUSION: We show how an ontology-based representation of a clinical summary, guided by semantic patterns, allows homogeneous querying of heterogeneous information structures. Whether there are a finite number of top-level patterns is an open question. Catalina Martínez-Costa, Ronald Cornet, Daniel Karlsson, Stefan Schulz 0001, Dipak Kalra |
J. Am. Medical Informatics Assoc. | 4 |
| 2015 | Secondary use of electronic health records for building cohort studies through top-down information extraction
Markus Kreuzthaler, Stefan Schulz 0001, Andrea Berghold |
J. Biomed. Informatics | 2 |
| 2014 | An Ontological Analysis of Reference in Health Record StatementsabstractThe relation between an information entity and its referent can be described as a second-order statement, as long as the referent is a type. This is typical for medical discourse such as diagnostic statements in electronic health records (EHRs), which often express hypotheses or probability assertions about the existence of an instance of, e.g. a disease type. This paper presents several approximations using description logics and a query language, the entailments of which are checked against a reference standard. Their pros and cons are discussed in the light of formal ontology and logic. Stefan Schulz 0001, Catalina Martínez-Costa, Daniel Karlsson, Ronald Cornet, Mathias Brochhausen, Alan L. Rector |
FOIS | 1 |
| 2013 | Ontology-Based Reengineering of the SNOMED CT Context Model
Catalina Martínez-Costa, Stefan Schulz 0001 |
AIME | 2 |
| 2013 | Evaluation of the OQuaRE framework for ontology quality
Astrid Duque-Ramos, Jesualdo Tomás Fernández-Breis, Miguela Iniesta-Moreno, Michel Dumontier, Mikel Egaña Aranguren, Stefan Schulz 0001, Nathalie Aussenac-Gilles, Robert Stevens 0001 |
Expert Syst. Appl. | 6 |
| 2012 | Metonymies in Medical Terminologies. A SNOMED CT Case Study
Markus Kreuzthaler, Stefan Schulz 0001 |
AMIA | 2 |
| 2012 | Competing Interpretations of Disorder Codes in SNOMED CT and ICD
Stefan Schulz 0001, Alan L. Rector, Jean Marie Rodrigues, Kent A. Spackman |
AMIA | 1 |
| 2012 | Creation and use of Language Resources in a Question-Answering eHealth System
Ulrich Andersen, Anna Braasch, Lina Henriksen, Csaba Huszka, Anders Johannsen, Lars Kayser, Bente Maegaard, Ole Norgaard, Stefan Schulz 0001, Jürgen Wedekind |
LREC | 9 |
| 2012 | Process attributes in bio-ontologiesabstractBACKGROUND: Biomedical processes can provide essential information about the (mal-) functioning of an organism and are thus frequently represented in biomedical terminologies and ontologies, including the GO Biological Process branch. These processes often need to be described and categorised in terms of their attributes, such as rates or regularities. The adequate representation of such process attributes has been a contentious issue in bio-ontologies recently; and domain ontologies have correspondingly developed ad hoc workarounds that compromise interoperability and logical consistency. RESULTS: We present a design pattern for the representation of process attributes that is compatible with upper ontology frameworks such as BFO and BioTop. Our solution rests on two key tenets: firstly, that many of the sorts of process attributes which are biomedically interesting can be characterised by the ways that repeated parts of such processes constitute, in combination, an overall process; secondly, that entities for which a full logical definition can be assigned do not need to be treated as primitive within a formal ontology framework. We apply this approach to the challenge of modelling and automatically classifying examples of normal and abnormal rates and patterns of heart beating processes, and discuss the expressivity required in the underlying ontology representation language. We provide full definitions for process attributes at increasing levels of domain complexity. CONCLUSIONS: We show that a logical definition of process attributes is feasible, though limited by the expressivity of DL languages so that the creation of primitives is still necessary. This finding may endorse current formal upper-ontology frameworks as a way of ensuring consistency, interoperability and clarity. André Queiroz de Andrade, Ward Blondé, Janna Hastings, Stefan Schulz 0001 |
BMC Bioinform. | 4 |
| 2012 | Usability-driven pruning of large ontologies: the case of SNOMED CTabstractOBJECTIVES: To study ontology modularization techniques when applied to SNOMED CT in a scenario in which no previous corpus of information exists and to examine if frequency-based filtering using MEDLINE can reduce subset size without discarding relevant concepts. MATERIALS AND METHODS: Subsets were first extracted using four graph-traversal heuristics and one logic-based technique, and were subsequently filtered with frequency information from MEDLINE. Twenty manually coded discharge summaries from cardiology patients were used as signatures and test sets. The coverage, size, and precision of extracted subsets were measured. RESULTS: Graph-traversal heuristics provided high coverage (71-96% of terms in the test sets of discharge summaries) at the expense of subset size (17-51% of the size of SNOMED CT). Pre-computed subsets and logic-based techniques extracted small subsets (1%), but coverage was limited (24-55%). Filtering reduced the size of large subsets to 10% while still providing 80% coverage. DISCUSSION: Extracting subsets to annotate discharge summaries is challenging when no previous corpus exists. Ontology modularization provides valuable techniques, but the resulting modules grow as signatures spread across subhierarchies, yielding a very low precision. CONCLUSION: Graph-traversal strategies and frequency data from an authoritative source can prune large biomedical ontologies and produce useful subsets that still exhibit acceptable coverage. However, a clinical corpus closer to the specific use case is preferred when available. Pablo López-García, Martin Boeker, Arantza Illarramendi, Stefan Schulz 0001 |
J. Am. Medical Informatics Assoc. | 4 |
| 2011 | Ontology patterns for tabular representations of biomedical knowledge on neglected tropical diseasesabstractMOTIVATION: Ontology-like domain knowledge is frequently published in a tabular format embedded in scientific publications. We explore the re-use of such tabular content in the process of building NTDO, an ontology of neglected tropical diseases (NTDs), where the representation of the interdependencies between hosts, pathogens and vectors plays a crucial role. RESULTS: As a proof of concept we analyzed a tabular compilation of knowledge about pathogens, vectors and geographic locations involved in the transmission of NTDs. After a thorough ontological analysis of the domain of interest, we formulated a comprehensive design pattern, rooted in the biomedical domain upper level ontology BioTop. This pattern was implemented in a VBA script which takes cell contents of an Excel spreadsheet and transforms them into OWL-DL. After minor manual post-processing, the correctness and completeness of the ontology was tested using pre-formulated competence questions as description logics (DL) queries. The expected results could be reproduced by the ontology. The proposed approach is recommended for optimizing the acquisition of ontological domain knowledge from tabular representations. AVAILABILITY AND IMPLEMENTATION: Domain examples, source code and ontology are freely available on the web at http://www.cin.ufpe.br/~ntdo. CONTACT: [email protected]. Filipe Santana da Silva, Daniel Schober, Zulma Medeiros, Fred Freitas, Stefan Schulz 0001 |
Bioinform. | 5 |
| 2011 | Unintended consequences of existential quantifications in biomedical ontologiesabstractBACKGROUND: The Open Biomedical Ontologies (OBO) Foundry is a collection of freely available ontologically structured controlled vocabularies in the biomedical domain. Most of them are disseminated via both the OBO Flatfile Format and the semantic web format Web Ontology Language (OWL), which draws upon formal logic. Based on the interpretations underlying OWL description logics (OWL-DL) semantics, we scrutinize the OWL-DL releases of OBO ontologies to assess whether their logical axioms correspond to the meaning intended by their authors. RESULTS: We analyzed ontologies and ontology cross products available via the OBO Foundry site http://www.obofoundry.org for existential restrictions (someValuesFrom), from which we examined a random sample of 2,836 clauses.According to a rating done by four experts, 23% of all existential restrictions in OBO Foundry candidate ontologies are suspicious (Cohens' κ = 0.78). We found a smaller proportion of existential restrictions in OBO Foundry cross products are suspicious, but in this case an accurate quantitative judgment is not possible due to a low inter-rater agreement (κ = 0.07). We identified several typical modeling problems, for which satisfactory ontology design patterns based on OWL-DL were proposed. We further describe several usability issues with OBO ontologies, including the lack of ontological commitment for several common terms, and the proliferation of domain-specific relations. CONCLUSIONS: The current OWL releases of OBO Foundry (and Foundry candidate) ontologies contain numerous assertions which do not properly describe the underlying biological reality, or are ambiguous and difficult to interpret. The solution is a better anchoring in upper ontologies and a restriction to relatively few, well defined relation types with given domain and range constraints. Martin Boeker, Ilinca Tudose, Janna Hastings, Daniel Schober, Stefan Schulz 0001 |
BMC Bioinform. | 5 |
| 2010 | What are chemical structures and their relations?abstractIn chemistry, advances in computational technologies have allowed research into molecules that have not been synthesized yet, and may never be, to become widespread. These are described in terms of their structures, which are expressed as chemical graphs. Chemical graphs are a representational artifact, and as such are of a different ontological nature than the molecular entities which they describe. Janna Hastings, Colin R. Batchelor, Christoph Steinbeck, Stefan Schulz 0001 |
FOIS | 4 |
| 2009 | Detecting Underspecification in SNOMED CT Concept Definitions Through Natural Language Processing
Edson José Pacheco, Holger Stenzhorn, Percy Nohama, Jan Paetzold, Stefan Schulz 0001 |
AMIA | 5 |
| 2009 | Alignment of the UMLS semantic network with BioTop: methodology and assessmentabstractMOTIVATION: For many years, the Unified Medical Language System (UMLS) semantic network (SN) has been used as an upper-level semantic framework for the categorization of terms from terminological resources in biomedicine. BioTop has recently been developed as an upper-level ontology for the biomedical domain. In contrast to the SN, it is founded upon strict ontological principles, using OWL DL as a formal representation language, which has become standard in the semantic Web. In order to make logic-based reasoning available for the resources annotated or categorized with the SN, a mapping ontology was developed aligning the SN with BioTop. METHODS: The theoretical foundations and the practical realization of the alignment are being described, with a focus on the design decisions taken, the problems encountered and the adaptations of BioTop that became necessary. For evaluation purposes, UMLS concept pairs obtained from MEDLINE abstracts by a named entity recognition system were tested for possible semantic relationships. Furthermore, all semantic-type combinations that occur in the UMLS Metathesaurus were checked for satisfiability. RESULTS: The effort-intensive alignment process required major design changes and enhancements of BioTop and brought up several design errors that could be fixed. A comparison between a human curator and the ontology yielded only a low agreement. Ontology reasoning was also used to successfully identify 133 inconsistent semantic-type combinations. AVAILABILITY: BioTop, the OWL DL representation of the UMLS SN, and the mapping ontology are available at http://www.purl.org/biotop/. Stefan Schulz 0001, Elena Beisswanger, László van den Hoek, Olivier Bodenreider, Erik M. van Mulligen |
Bioinform. | 1 |
| 2008 | Evaluation of a Document Search Engine in a Clinical Department System
Stefan Schulz 0001, Philipp Daumke, Pascal Fischer, Marcel Müller |
AMIA | 1 |
| 2008 | The ontology of biological taxaabstractMOTIVATION: The classification of biological entities in terms of species and taxa is an important endeavor in biology. Although a large amount of statements encoded in current biomedical ontologies is taxon-dependent there is no obvious or standard way for introducing taxon information into an integrative ontology architecture, supposedly because of ongoing controversies about the ontological nature of species and taxa. RESULTS: In this article, we discuss different approaches on how to represent biological taxa using existing standards for biomedical ontologies such as the description logic OWL DL and the Open Biomedical Ontologies Relation Ontology. We demonstrate how hidden ambiguities of the species concept can be dealt with and existing controversies can be overcome. A novel approach is to envisage taxon information as qualities that inhere in biological organisms, organism parts and populations. AVAILABILITY: The presented methodology has been implemented in the domain top-level ontology BioTop, openly accessible at http://purl.org/biotop. BioTop may help to improve the logical and ontological rigor of biomedical ontologies and further provides a clear architectural principle to deal with biological taxa information. Stefan Schulz 0001, Holger Stenzhorn, Martin Boeker |
ISMB | 1 |
| 2008 | Model Formulation: Modeling Functional Neuroanatomy for an Anatomy Information SystemabstractOBJECTIVE: Existing neuroanatomical ontologies, databases and information systems, such as the Foundational Model of Anatomy (FMA), represent outgoing connections from brain structures, but cannot represent the "internal wiring" of structures and as such, cannot distinguish between different independent connections from the same structure. Thus, a fundamental aspect of Neuroanatomy, the functional pathways and functional systems of the brain such as the pupillary light reflex system, is not adequately represented. This article identifies underlying anatomical objects which are the source of independent connections (collections of neurons) and uses these as basic building blocks to construct a model of functional neuroanatomy and its functional pathways. DESIGN: The basic representational elements of the model are unnamed groups of neurons or groups of neuron segments. These groups, their relations to each other, and the relations to the objects of macroscopic anatomy are defined. The resulting model can be incorporated into the FMA. MEASUREMENTS: The capabilities of the presented model are compared to the FMA and the Brain Architecture Management System (BAMS). RESULTS: Internal wiring as well as functional pathways can correctly be represented and tracked. CONCLUSION: This model bridges the gap between representations of single neurons and their parts on the one hand and representations of spatial brain structures and areas on the other hand. It is capable of drawing correct inferences on pathways in a nervous system. The object and relation definitions are related to the Open Biomedical Ontology effort and its relation ontology, so that this model can be further developed into an ontology of neuronal functional systems. Jörg Niggemann, Andreas Gebert, Stefan Schulz 0001 |
J. Am. Medical Informatics Assoc. | 3 |
| 2007 | Replacing SEP-Triplets in SNOMED CT Using Tractable Description Logic Operators
Boontawee Suntisrivaraporn, Franz Baader, Stefan Schulz 0001, Kent A. Spackman |
AIME | 3 |
| 2007 | The @neurIST Ontology of Intracranial Aneurysms: Providing Terminological Services for an Integrated IT Infrastructure
Martin Boeker, Holger Stenzhorn, Kai Kumpf, Philippe Bijlenga, Stefan Schulz 0001, Susanne Hanser |
AMIA | 5 |
| 2007 | Ontological foundations for biomedical sciences
Udo Hahn, Stefan Schulz 0001 |
Artif. Intell. Medicine | 2 |
| 2007 | Towards the ontological foundations of symbolic biological theories
Stefan Schulz 0001, Udo Hahn |
Artif. Intell. Medicine | 1 |
| 2007 | Spatial location and its relevance for terminological inferences in bio-ontologiesabstractBACKGROUND: An adequate and expressive ontological representation of biological organisms and their parts requires formal reasoning mechanisms for their relations of physical aggregation and containment. RESULTS: We demonstrate that the proposed formalism allows to deal consistently with "role propagation along non-taxonomic hierarchies", a problem which had repeatedly been identified as an intricate reasoning problem in biomedical ontologies. CONCLUSION: The proposed approach seems to be suitable for the redesign of compositional hierarchies in (bio)medical terminology systems which are embedded into the framework of the OBO (Open Biological Ontologies) Relation Ontology and are using knowledge representation languages developed by the Semantic Web community. Stefan Schulz 0001, Kornél G. Markó, Udo Hahn |
BMC Bioinform. | 1 |
| 2006 | Towards a Multilingual Medical Lexicon
Kornél G. Markó, Robert H. Baud, Pierre Zweigenbaum, Lars Borin, Magnus Merkel, Stefan Schulz 0001 |
AMIA | 6 |
| 2006 | Towards an Upper-Level Ontology for Molecular Biology
Stefan Schulz 0001, Elena Beisswanger, Joachim Wermter, Udo Hahn |
AMIA | 1 |
| 2006 | From GENIA to BIOTOP - Towards a Top-Level Ontology for Biology
Stefan Schulz 0001, Elena Beisswanger, Udo Hahn, Joachim Wermter, Anand Kumar 0005, Holger Stenzhorn |
FOIS | 1 |
| 2006 | Language Specific and Topic Focused Web Crawling
Olena Medelyan, Stefan Schulz 0001, Jan Paetzold, Michael Poprat, Kornél G. Markó |
LREC | 2 |
| 2006 | Semantic Atomicity and Multilinguality in the Medical Domain: Design Considerations for the MorphoSaurus Subword Lexicon
Stefan Schulz 0001, Kornél G. Markó, Philipp Daumke, Udo Hahn, Susanne Hanser, Percy Nohama, Roosewelt L. Andrade, Edson José Pacheco, Martin Romacker |
LREC | 1 |
| 2006 | Biomedical ontologies: What part-of is and isn't
Stefan Schulz 0001, Anand Kumar 0005, Thomas Bittner |
J. Biomed. Informatics | 1 |
| 2005 | Unsupervised Multilingual Word Sense Disambiguation via an Interlingua
Kornél G. Markó, Stefan Schulz 0001, Udo Hahn |
AAAI | 2 |
| 2005 | Interchanging Lexical Information for a Multilingual Dictionary
Robert H. Baud, Mikael Nyström, Lars Borin, Roger Evans, Stefan Schulz 0001, Pierre Zweigenbaum |
AMIA | 5 |
| 2005 | Multilingual Biomedical Dictionary
Philipp Daumke, Kornél G. Markó, Michael Poprat, Stefan Schulz 0001 |
AMIA | 4 |
| 2005 | A CLIR Interface to a Web Search Engine
Philipp Daumke, Stefan Schulz 0001, Kornél G. Markó |
AMIA | 2 |
| 2005 | How to Distinguish Parthood from Location in Bio-Ontologies
Stefan Schulz 0001, Philipp Daumke, Barry Smith 0001, Udo Hahn |
AMIA | 1 |
| 2005 | Anatomical Information Science
Barry Smith 0001, José L. V. Mejino Jr., Stefan Schulz 0001, Anand Kumar 0005, Cornelius Rosse |
COSIT | 3 |
| 2005 | Cross-Language Mining for Acronyms and Their Completions from the Web
Udo Hahn, Philipp Daumke, Stefan Schulz 0001, Kornél G. Markó |
Discovery Science | 3 |
| 2005 | Subword Clusters as Light-Weight Interlingua for Multilingual Document RetrievalabstractWe introduce a light-weight interlingua for a cross-language document retrieval system in the medical domain. It is composed of equivalence classes of semantically primitive, language-specific subwords which are clustered by interlingual and intralingual synonymy. Each subword cluster represents a basic conceptual entity of the language-independent interlingua. Documents, as well as queries, are mapped to this interlingua level on which retrieval operations are performed. Evaluation experiments reveal that this interlingua-based retrieval model outperforms a direct translation approach. Udo Hahn, Kornél G. Markó, Stefan Schulz 0001 |
MTSummit | 3 |
| 2005 | A CLIR interface to a web search engineabstractNo abstract available. Philipp Daumke, Stefan Schulz 0001, Kornél G. Markó |
SIGIR | 2 |
| 2005 | Bootstrapping dictionaries for cross-language information retrievalabstractThe bottleneck for dictionary-based cross-language information retrieval is the lack of comprehensive dictionaries, in particular for many different languages. We here introduce a methodology by which multilingual dictionaries (for Spanish and Swedish) emerge automatically from simple seed lexicons. These seed lexicons are automatically generated, by cognate mapping, from (previously manually constructed) Portuguese and German as well as English sources. Lexical and semantic hypotheses are then validated and new ones iteratively generated by making use of co-occurrence patterns of hypothesized translation synonyms in parallel corpora. We evaluate these newly derived dictionaries on a large medical document collection within a cross-language retrieval setting. Kornél G. Markó, Stefan Schulz 0001, Olena Medelyan, Udo Hahn |
SIGIR | 2 |
| 2005 | Part-whole representation and reasoning in formal biomedical ontologies
Stefan Schulz 0001, Udo Hahn |
Artif. Intell. Medicine | 1 |
| 2004 | Learning Indexing Patterns from One Language for the Benefit of Others
Udo Hahn, Kornél G. Markó, Stefan Schulz 0001 |
AAAI | 3 |
| 2004 | Mereological Semantics for Bio-Ontologies
Udo Hahn, Stefan Schulz 0001, Kornél G. Markó |
AAAI | 2 |
| 2004 | Cognate Mapping - A Heuristic Strategy for the Semi-Supervised Acquisition of a Spanish Lexicon from a Portuguese Seed Lexicon
Stefan Schulz 0001, Kornél G. Markó, Eduardo Sbrissia, Percy Nohama, Udo Hahn |
COLING | 1 |
| 2004 | Representing Natural Kinds by Spatial Inclusion and Containment
Stefan Schulz 0001, Udo Hahn |
ECAI | 1 |
| 2004 | Parthood as Spatial Inclusion - Evidence from biomedical Conceptualizations
Stefan Schulz 0001, Udo Hahn |
KR | 1 |
| 2003 | Logic-based Remodeling of the Digital Anatomist Foundational Model
Rainer Beck, Stefan Schulz 0001 |
AMIA | 2 |
| 2003 | Cross-language MeSH Indexing using Morpho-Semantic Normalization
Kornél G. Markó, Philipp Daumke, Stefan Schulz 0001, Udo Hahn |
AMIA | 3 |
| 2002 | A knowledge representation view on biomedical structure and function
Stefan Schulz 0001, Udo Hahn |
AMIA | 1 |
| 2002 | MORPHOSAURUS: Crosslingual Medical Text Retrieval by Subword Indexing
Stefan Schulz 0001, Rüdiger Klar, Udo Hahn, Martin Romacker, Percy Nohama, Lúcio J. Dias Matias |
AMIA | 1 |
| 2002 | The German specialist lexicon
Gesa Weske-Heck, Albrecht Zaiß, Matthias Zabel, Stefan Schulz 0001, Wolfgang Giere, Michael Schopen, Rüdiger Klar |
AMIA | 4 |
| 2002 | Turning Lead into Gold? Feeding a Formal Knowledge Base with Informal Conceptual Knowledge
Udo Hahn, Stefan Schulz 0001 |
EKAW | 2 |
| 2002 | Necessary Parts and Wholes in Bio-Ontologies
Stefan Schulz 0001 |
KR | 1 |
| 2002 | Towards Very Large Ontologies for Medical Language Processing
Udo Hahn, Stefan Schulz 0001 |
LREC | 2 |
| 2002 | 240, 000 concepts and relations-towards mega knowledge bases for real-world applicationsabstractWe describe an ontology engineering methodology by which conceptual knowledge is extracted from an informal medical thesaurus (UMLS) and automatically converted into a formal description logics system (LOOM). Our approach consists of four steps: concept definitions are automatically generated from the UMLS, Integrity checking of taxonomic and partonomic hierarchies is performed by LOOM's terminological classifier, cycles and inconsistencies are eliminated, as well as incremental refinement of the evolving knowledge base is performed by a domain expert. We report an experiments with a very large knowledge base composed of 164,000 concepts and 76,000 relations. Udo Hahn, Stefan Schulz 0001 |
SMC (2) | 2 |
| 2001 | Parts, Locations, and Holes - Formal Reasoning about Anatomical Structures
Stefan Schulz 0001, Udo Hahn |
AIME | 1 |
| 2001 | Subword segmentation-leveling out morphological variations for medical document retrieval
Udo Hahn, Martin Honeck, Michael Piotrowski, Stefan Schulz 0001 |
AMIA | 4 |
| 2001 | Bidirectional mereological reasoning in anatomical knowledge bases
Stefan Schulz 0001 |
AMIA | 1 |
| 2001 | PILLS: A Multilingual Authoring System for Patient Information
Donia Scott, Nadjet Bouayad-Agha, Richard Power, Stefan Schulz 0001, Rainer Beck, Dawn Murphy, Rose Lockwood |
AMIA | 4 |
| 2001 | Mereotopological reasoning about parts and (w)holes in bio-ontologiesabstractWe here deal with mereotopological properties of parts and associated wholes, locations and empty spaces (holes), with particular reference to biological structures. Our considerations lead to a basic ontology which contains 'solid object', 'hole' and 'boundary' as mutually disjoint primitives. Formally, we embed the relations 'part-of' and 'location-of' into a parsimonious description logic (ALC) and emulate partonomic and spatial reasoning involving these relations by terminological subsumption. In contrast to common conceptualizations, we do not distinguish between solids and the regions they occupy, as well as we allow solids to have holes as proper parts. In order to support these modeling decisions, we discuss various concrete examples from human anatomy. Stefan Schulz 0001, Udo Hahn |
FOIS | 1 |
| 2001 | A Search Engine for Morphologically Complex Languages
Udo Hahn, Martin Honeck, Stefan Schulz 0001 |
IDA | 3 |
| 2000 | Automated coding of diagnoses-three methods compared
Pius Franz, Albrecht Zaiß, Stefan Schulz 0001, Udo Hahn, Rüdiger Klar |
AMIA | 3 |
| 2000 | MedSynDiKATe-design considerations for an ontology-based medical text understanding system
Udo Hahn, Martin Romacker, Stefan Schulz 0001 |
AMIA | 3 |
| 2000 | Modeling anatomical spatial relations with description logics
Stefan Schulz 0001, Udo Hahn, Martin Romacker |
AMIA | 1 |
| 2000 | Knowledge Engineering by Large-Scale Knowledge Reuse - Experience from the Medical Domain
Stefan Schulz 0001, Udo Hahn |
KR | 1 |
| 1999 | Streamlining semantic interpretation for medical narratives
Martin Romacker, Stefan Schulz 0001, Udo Hahn |
AMIA | 2 |
| 1999 | Automatic Import and Manual Refinement of Medical Knowledge
Stefan Schulz 0001, Giovanni Faggioli, Martin Romacker, Udo Hahn |
AMIA | 1 |
| 1999 | How knowledge drives understandingmatching medical ontologies with the needs of medical language processing
Udo Hahn, Martin Romacker, Stefan Schulz 0001 |
Artif. Intell. Medicine | 3 |
| 1998 | Controlled Evaluation of a Computer-Based Atlas of Histopathology
Stefan Schulz 0001, Thomas Auhuber, Ulrich Schrader, Rüdiger Klar |
AMIA | 1 |
| 1998 | Part-whole reasoning in medical ontologies revisited-introducing SEP triplets into classification-based description logics
Stefan Schulz 0001, Martin Romacker, Udo Hahn |
AMIA | 1 |