Udo Hahn

dblp:h/UdoHahn · DBLP profile ↗
← Back
144ranked-venue papers
49as first author
6since 2021 · last 2026
0000-0002-5052-0245ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 103 · 36 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 8 first-authorDatabases, data management, data science and information retrieval · 23 · 13 first-authorGraphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-authorHuman-computer interaction and ubiquitous computing · 4 · 3 first-authorTheory of computation · 4Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
21 papers
Information extraction and text analysis · 77% Efficient and distributed learning · 8% Knowledge representation and reasoning · 6%
Databases, data mining, and information retrieval
6 papers
Information retrieval · 97% Machine learning and data management · 3% Database system architecture and tuning · 0%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 94% Medical and health informatics · 6%

Topics — the 30 heaviest of 64, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
emotion recognition
0.512021
Towards Label-Agnostic Emotion Embeddings · EMNLP (1) 2021
Natural language and speech › Information extraction and text analysis › lexical resources › lexical resource construction
emotion lexicon
0.412020
Learning and Evaluating Emotion Lexicons for 91 Languages · ACL 2020
Natural language and speech › Information extraction and text analysis
sentiment and emotion analysis
0.412020
Learning and Evaluating Emotion Lexicons for 91 Languages · ACL 2020
Information retrieval › document retrieval › domain-specific retrieval › biomedical information retrieval
precision medicine search
0.412020
What Makes a Top-Performing Precision Medicine Search Engine?: Tracing Main System Features in a Systematic Way · SIGIR 2020
Information retrieval
retrieval evaluation
0.412020
What Makes a Top-Performing Precision Medicine Search Engine?: Tracing Main System Features in a Systematic Way · SIGIR 2020
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.112021
Towards Label-Agnostic Emotion Embeddings · EMNLP (1) 2021
Bioinformatics and computational biology › knowledge representation in biology
biomedical ontology
0.122009
MaHCO: an ontology of the major histocompatibility complex for immunoinformatic applications and text mining · Bioinform. 2009
Parthood as Spatial Inclusion - Evidence from biomedical Conceptualizations · KR 2004
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology
0.132004
Parthood as Spatial Inclusion - Evidence from biomedical Conceptualizations · KR 2004
Mereological Semantics for Bio-Ontologies · AAAI 2004
Knowledge Engineering by Large-Scale Knowledge Reuse - Experience from the Medical Domain · KR 2000
Natural language and speech › Information extraction and text analysis
data annotation
0.112010
A Cognitive Cost Model of Annotations Based on Eye-Tracking Data · ACL 2010
Natural language and speech › Information extraction and text analysis
dependency graph encoding
0.112010
Evaluating the Impact of Alternative Dependency Graph Encodings on Solving Event Extraction Tasks · EMNLP 2010
Natural language and speech › Information extraction and text analysis
event extraction
0.112010
Evaluating the Impact of Alternative Dependency Graph Encodings on Solving Event Extraction Tasks · EMNLP 2010
Machine learning › Efficient and distributed learning › active learning
semi-supervised active learning
0.112009
Semi-Supervised Active Learning for Sequence Labeling · ACL/IJCNLP 2009
Natural language and speech › Information extraction and text analysis
sequence labeling
0.112009
Semi-Supervised Active Learning for Sequence Labeling · ACL/IJCNLP 2009
Bioinformatics and computational biology › biomedical text mining
gene name disambiguation
0.112009
High-performance gene name normalization with GENO · Bioinform. 2009
Bioinformatics and computational biology
immunoinformatics
0.112009
MaHCO: an ontology of the major histocompatibility complex for immunoinformatic applications and text mining · Bioinform. 2009
Machine learning › Efficient and distributed learning
active learning
0.112008
Multi-Task Active Learning for Linguistic Annotations · ACL 2008
Natural language and speech › Information extraction and text analysis › data annotation
linguistic annotation
0.112008
Multi-Task Active Learning for Linguistic Annotations · ACL 2008
Machine learning › Learning paradigms
multi-task learning
0.112008
Multi-Task Active Learning for Linguistic Annotations · ACL 2008
Machine learning › Efficient and distributed learning › active learning
annotation cost reduction
0.112007
An Approach to Text Corpus Construction which Cuts Annotation Costs and Maintains Reusability of Annotated Data · EMNLP-CoNLL 2007
Natural language and speech › Information extraction and text analysis
dataset construction
0.112007
An Approach to Text Corpus Construction which Cuts Annotation Costs and Maintains Reusability of Annotated Data · EMNLP-CoNLL 2007
Natural language and speech › Information extraction and text analysis › phrase extraction
collocation extraction
0.112006
You Can't Beat Frequency (Unless You Use Linguistic Knowledge) - A Qualitative Evaluation of Association Measures for Collocation and Term Extraction · ACL 2006
Natural language and speech › Information extraction and text analysis › phrase extraction
terminology extraction
0.112006
You Can't Beat Frequency (Unless You Use Linguistic Knowledge) - A Qualitative Evaluation of Association Measures for Collocation and Term Extraction · ACL 2006
Natural language and speech › Information extraction and text analysis › word sense disambiguation
multilingual word sense disambiguation
0.112005
Unsupervised Multilingual Word Sense Disambiguation via an Interlingua · AAAI 2005
Natural language and speech › Information extraction and text analysis
word sense disambiguation
0.112005
Unsupervised Multilingual Word Sense Disambiguation via an Interlingua · AAAI 2005
Information retrieval
cross-language information retrieval
0.112005
Bootstrapping dictionaries for cross-language information retrieval · SIGIR 2005
Information retrieval
multilingual dictionary construction
0.112005
Bootstrapping dictionaries for cross-language information retrieval · SIGIR 2005
Information retrieval
indexing
0.012004
Learning Indexing Patterns from One Language for the Benefit of Others · AAAI 2004
Natural language and speech › Information extraction and text analysis › word sense disambiguation
metonymy resolution
0.012002
Understanding metonymies in discourse · Artif. Intell. 2002
Natural language and speech › Information extraction and text analysis
natural language semantics
0.012002
Understanding metonymies in discourse · Artif. Intell. 2002
Natural language and speech › Information extraction and text analysis › coreference resolution
anaphora resolution
0.021997
On the Interaction of Metonymies and Anaphora · IJCAI (2) 1997
Functional Centering · ACL 1996

Methods — techniques the papers use, named apart from their topics

contrastive learning · 0.5stop word filtering · 0.4query expansion · 0.4keyword boosting · 0.4embedding model · 0.4bilingual word translation model · 0.4SMAC · 0.4BM25 · 0.4eye-tracking data analysis · 0.2active learning · 0.2symbolic methods · 0.1statistical methods · 0.1semi-supervised learning · 0.1semantic profiling · 0.1semantic annotation · 0.1ontology engineering · 0.1multi-task learning · 0.1statistical hypothesis testing · 0.1
YearPublicationVenuePosition
2026 Developing the German Medical Text Corpus (GeMTeX): Legal Compliance and Semantic Enrichment
Justin Hofenbitzer, Christina Lohr, Andrea Riedel, Rebekka Kiser, Aliaksandra Shutsko, Abanoub Abdelmalak, Peter Klügl, Jutta Romberg, Sarah Riepenhausen, Miriam Schechner, Jakob Faller, Frank A. Meineke, Luise Modersohn, Markus Löffler, Juliane Fluck, Udo Hahn, Stefan Schulz 0001, Martin Boeker
LREC16
2026 Emotion Embeddings - Learning Stable and Homogeneous Abstractions From Heterogeneous Affective Datasets
abstract
Human emotion is expressed in many communication modalities and media formats and so their computational study is equally diversified into natural language processing, speech processing, computer vision, etc. Similarly, a large variety of representation formats has been used in previous research to describe emotions (polarity scales, basic emotion categories, dimensional approaches, appraisal theory, etc.) and has led to an ever proliferating diversity of datasets, predictive models, and software tools for emotion analysis. Because of these two distinct types of heterogeneity, at the expressional and representational level, there is a dire need to unify previous work on increasingly diverging data and label types. This article presents such a unifying computational model. We propose a training procedure that learns a shared latent representation for emotions, so-called emotion embeddings, independent of different natural languages, communication modalities, media or representation label formats, and even disparate model architectures. Experiments on a wide range of heterogeneous affective datasets indicate that this approach yields the desired interoperability for the sake of reusability, interpretability and flexibility, without penalizing prediction quality. Code and data are archived under DOI:10.5281/zenodo.7405327.
Sven Buechel, Udo Hahn
IEEE Trans. Affect. Comput.2
2022 GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers
abstract
Despite remarkable advances in the development of language resources over the recent years, there is still a shortage of annotated, publicly available corpora covering (German) medical language. With the initial release of the German Guideline Program in Oncology NLP Corpus (GGPONC), we have demonstrated how such corpora can be built upon clinical guidelines, a widely available resource in many natural languages with a reasonable coverage of medical terminology. In this work, we describe a major new release for GGPONC. The corpus has been substantially extended in size and re-annotated with a new annotation scheme based on SNOMED CT top level hierarchies, reaching high inter-annotator agreement (γ=.94). Moreover, we annotated elliptical coordinated noun phrases and their resolutions, a common language phenomenon in (not only German) scientific documents. We also trained BERT-based named entity recognition models on this new data set, which achieve high performance on short, coarse-grained entity spans (F1=.89), while the rate of boundary errors increases for long entity spans. GGPONC is freely available through a data use agreement. The trained named entity recognition models, as well as the detailed annotation guide, are also made publicly available.
Florian Borchert, Christina Lohr, Luise Modersohn, Jonas Witt, Thomas Langer, Markus Follmann, Matthias Gietzelt, Bert Arnrich, Udo Hahn, Matthieu-P. Schapranow
LREC9
2022 "Beste Grüße, Maria Meyer" - Pseudonymization of Privacy-Sensitive Information in Emails
abstract
The exploding amount of user-generated content has spurred NLP research to deal with documents from various digital communication formats (tweets, chats, emails, etc.). Using these texts as language resources implies complying with legal data privacy regulations. To protect the personal data of individuals and preclude their identification, we employ pseudonymization. More precisely, we identify those text spans that carry information revealing an individual’s identity (e.g., names of persons, locations, phone numbers, or dates) and subsequently substitute them with synthetically generated surrogates. Based on CodE Alltag, a German-language email corpus, we address two tasks. The first task is to evaluate various architectures for the automatic recognition of privacy-sensitive entities in raw data. The second task examines the applicability of pseudonymized data as training data for such systems since models learned on original data cannot be published for reasons of privacy protection. As outputs of both tasks, we, first, generate a new pseudonymized version of CodE Alltag compliant with the legal requirements of the General Data Protection Regulation (GDPR). Second, we make accessible a tagger for recognizing privacy-sensitive information in German emails and similar text genres, which is trained on already pseudonymized data.
Elisabeth Eder, Michael Wiegand, Ulrike Krieg-Holz, Udo Hahn
LREC4
2021 Acquiring a Formality-Informed Lexical Resource for Style Analysis
abstract
To track different levels of formality in written discourse, we introduce a novel type of lexicon for the German language, with entries ordered by their degree of (in)formality.We start with a set of words extracted from traditional lexicographic resources, extend it by sentence-based similarity computations, and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation.We submit this lexicon to an intrinsic evaluation related to the best regression models and their effect on predicting formality scores and complement our investigation by an extrinsic evaluation of formality on a German-language email corpus.
Elisabeth Eder, Ulrike Krieg-Holz, Udo Hahn
EACL3
2021 Towards Label-Agnostic Emotion Embeddings
abstract
Research in emotion analysis is scattered across different label formats (e.g., polarity types, basic emotion categories, and affective dimensions), linguistic levels (word vs. sentence vs. discourse), and, of course, (few wellresourced but much more under-resourced) natural languages and text genres (e.g., product reviews, tweets, news).The resulting heterogeneity makes data and software developed under these conflicting constraints hard to compare and challenging to integrate.To resolve this unsatisfactory state of affairs we here propose a training scheme that learns a shared latent representation of emotion independent from different label formats, natural languages, and even disparate model architectures.Experiments on a wide range of datasets indicate that this approach yields the desired interoperability without penalizing prediction quality.Code and data are archived under
Sven Buechel, Luise Modersohn, Udo Hahn
EMNLP (1)3
2020 Learning and Evaluating Emotion Lexicons for 91 Languages
abstract
Emotion lexicons describe the affective meaning of words and thus constitute a centerpiece for advanced sentiment and emotion analysis.Yet, manually curated lexicons are only available for a handful of languages, leaving most languages of the world without such a precious resource for downstream applications.Even worse, their coverage is often limited both in terms of the lexical units they contain and the emotional variables they feature.In order to break this bottleneck, we here introduce a methodology for creating almost arbitrarily large emotion lexicons for any target language.Our approach requires nothing but a source language emotion lexicon, a bilingual word translation model, and a target language embedding model.Fulfilling these requirements for 91 languages, we are able to generate representationally rich high-coverage lexicons comprising eight emotional variables with more than 100k lexical entries each.We evaluated the automatically generated lexicons against human judgment from 26 datasets, spanning 12 typologically diverse languages, and found that our approach produces results in line with state-of-the-art monolingual approaches to lexicon creation and even surpasses human reliability for some languages and variables.Code and data are available at github.
Sven Buechel, Susanna Rücker, Udo Hahn
ACL3
2020 CodE Alltag 2.0 - A Pseudonymized German-Language Email Corpus
abstract
The vast amount of social communication distributed over various electronic media channels (tweets, blogs, emails, etc.), so-called user-generated content (UGC), creates entirely new opportunities for today’s NLP research. Yet, data privacy concerns implied by the unauthorized use of these text streams as a data resource are often neglected. In an attempt to reconciliate the diverging needs of unconstrained raw data use and preservation of data privacy in digital communication, we here investigate the automatic recognition of privacy-sensitive stretches of text in UGC and provide an algorithmic solution for the protection of personal data via pseudonymization. Our focus is directed at the de-identification of emails where personally identifying information does not only refer to the sender but also to those people, locations, dates, and other identifiers mentioned in greetings, boilerplates and the content-carrying body of emails. We evaluate several de-identification procedures and systems on two hitherto non-anonymized German-language email corpora (CodE AlltagS+d and CodE AlltagXL), and generate fully pseudonymized versions for both (CodE Alltag 2.0) in which personally identifying information of all social actors addressed in these mails has been camouflaged (to the greatest extent possible).
Elisabeth Eder, Ulrike Krieg-Holz, Udo Hahn
LREC3
2020 ProGene - A Large-scale, High-Quality Protein-Gene Annotated Benchmark Corpus
abstract
Genes and proteins constitute the fundamental entities of molecular genetics. We here introduce ProGene (formerly called FSU-PRGE), a corpus that reflects our efforts to cope with this important class of named entities within the framework of a long-lasting large-scale annotation campaign at the Jena University Language & Information Engineering (JULIE) Lab. We assembled the entire corpus from 11 subcorpora covering various biological domains to achieve an overall subdomain-independent corpus. It consists of 3,308 MEDLINE abstracts with over 36k sentences and more than 960k tokens annotated with nearly 60k named entity mentions. Two annotators strove for carefully assigning entity mentions to classes of genes/proteins as well as families/groups, complexes, variants and enumerations of those where genes and proteins are represented by a single class. The main purpose of the corpus is to provide a large body of consistent and reliable annotations for supervised training and evaluation of machine learning algorithms in this relevant domain. Furthermore, we provide an evaluation of two state-of-the-art baseline systems — BioBert and flair — on the ProGene corpus. We make the evaluation datasets and the trained models available to encourage comparable evaluations of new methods in the future.
Erik Faessler, Luise Modersohn, Christina Lohr, Udo Hahn
LREC4
2020 Allgemeine Musikalische Zeitung as a Searchable Online Corpus
abstract
The massive digitization efforts related to historical newspapers over the past decades have focused on mass media sources and ordinary people as their primary recipients. Much less attention has been paid to newspapers published for a more specialized audience, e.g., those aiming at scholarly or cultural exchange within intellectual communities much narrower in scope, such as newspapers devoted to music criticism, arts or philosophy. Only some few of these specialized newspapers have been digitized up until now, but they are usually not well curated in terms of digitization quality, data formatting, completeness, redundancy (de-duplication), supply of metadata, and, hence, searchability. This paper describes our approach to eliminate these drawbacks for a major German-language newspaper resource of the Romantic Age, the Allgemeine Musikalische Zeitung (General Music Gazette). We here focus on a workflow that copes with a posteriori digitization problems, inconsistent OCRing and index building for searchability. In addition, we provide a user-friendly graphic interface to empower content-centric access to this (and other) digital resource(s) adopting open-source software for the purpose of Web presentation.
Bernd Kampe, Tinghui Duan, Udo Hahn
LREC3
2020 What Makes a Top-Performing Precision Medicine Search Engine?: Tracing Main System Features in a Systematic Way
abstract
From 2017 to 2019 the Text REtrieval Conference (TREC) held a challenge task on precision medicine using documents from medical publications (PubMed) and clinical trials. Despite lots of performance measurements carried out in these evaluation campaigns, the scientific community is still pretty unsure about the impact individual system features and their weights have on the overall system performance. In order to overcome this explanatory gap, we first determined optimal feature configurations using the Sequential Model-based Algorithm Configuration (SMAC) program and applied its output to a BM25-based search engine. We then ran an ablation study to systematically assess the individual contributions of relevant system features: BM25 parameters, query type and weighting schema, query expansion, stop word filtering, and keyword boosting. For evaluation, we employed the gold standard data from the three TREC Precision Medicine (TREC-PM) installments to evaluate the effectiveness of different features using the commonly shared infNDCG metric.
Erik Faessler, Michel Oleynik, Udo Hahn
SIGIR3
2018 CDA-compliant section annotation of German-language discharge summaries: Guideline development, annotation campaign, section classification
Christina Lohr, Stephanie Luther, Franz Matthies, Udo Hahn
AMIA4
2018 Emotion Representation Mapping for Automatic Lexicon Construction (Mostly) Performs on Human Level
abstract
Emotion Representation Mapping (ERM) has the goal to convert existing emotion ratings from one representation format into another one, e.g., mapping Valence-Arousal-Dominance annotations for words or sentences into Ekman’s Basic Emotions and vice versa. ERM can thus not only be considered as an alternative to Word Emotion Induction (WEI) techniques for automatic emotion lexicon construction but may also help mitigate problems that come from the proliferation of emotion representation formats in recent years. We propose a new neural network approach to ERM that not only outperforms the previous state-of-the-art. Equally important, we present a refined evaluation methodology and gather strong evidence that our model yields results which are (almost) as reliable as human annotations, even in cross-lingual settings. Based on these results we generate new emotion ratings for 13 typologically diverse languages and claim that they have near-gold quality, at least.
Sven Buechel, Udo Hahn
COLING2
2018 Annotation Data Management with JeDIS
abstract
This paper introduces the Jena Document Information System (JeDIS). The focus lies on its capability to partition annotation graphs into modules. Annotation modules are defined in terms of types from the annotation schema. Modules allow easy manipulation of their annotations (deletion or update) and the creation of alternative annotations of individual documents even for annotation formalisms that by design do not support this feature.
Erik Faessler, Udo Hahn
DocEng2
2018 Representation Mapping: A Novel Approach to Generate High-Quality Multi-Lingual Emotion Lexicons
Sven Buechel, Udo Hahn
LREC2
2018 Sharing Copies of Synthetic Clinical Corpora without Physical Distribution - A Case Study to Get Around IPRs and Privacy Constraints Featuring the German JSYNCC Corpus
Christina Lohr, Sven Buechel, Udo Hahn
LREC3
2018 Word Emotion Induction for Multiple Languages as a Deep Multi-Task Learning Problem
abstract
Predicting the emotional value of lexical items is a well-known problem in sentiment analysis.While research has focused on polarity for quite a long time, meanwhile this early focus has been shifted to more expressive emotion representation models (such as Basic Emotions or Valence-Arousal-Dominance).This change resulted in a proliferation of heterogeneous formats and, in parallel, often smallsized, non-interoperable resources (lexicons and corpus annotations).In particular, the limitations in size hampered the application of deep learning methods in this area because they typically require large amounts of input data.We here present a solution to get around this language data bottleneck by rephrasing word emotion induction as a multi-task learning problem.In this approach, the prediction of each independent emotion dimension is considered as an individual task and hidden layers are shared between these dimensions.We investigate whether multi-task learning is more advantageous than single-task learning for emotion prediction by comparing our model against a wide range of alternative emotion and polarity induction methods featuring 9 typologically diverse languages and a total of 15 conditions.Our model turns out to outperform each one of them.Against all odds, the proposed deep learning approach yields the largest gain on the smallest data sets, merely composed of one thousand samples.
Sven Buechel, Udo Hahn
NAACL-HLT2
2017 A Flexible Mapping Scheme for Discrete and Dimensional Emotion Representations
Sven Buechel, Udo Hahn
CogSci2
2016 Bad Company - Neighborhoods in Neural Embedding Spaces Considered Harmful
abstract
We assess the reliability and accuracy of (neural) word embeddings for both modern and historical English and German. Our research provides deeper insights into the empirically justified choice of optimal training methods and parameters. The overall low reliability we observe, nevertheless, casts doubt on the suitability of word neighborhoods in embedding spaces as a basis for qualitative conclusions on synchronic and diachronic lexico-semantic matters, an issue currently high up in the agenda of Digital Humanities.
Johannes Hellrich, Udo Hahn
COLING2
2016 Emotion Analysis as a Regression Problem - Dimensional Models and Their Implications on Emotion Representation and Metrical Evaluation
abstract
Emotion analysis (EA) and sentiment analysis are closely related tasks differing in the psychological phenomenon they aim to catch. We address fine-grained models for EA which treat the computation of the emotional status of narrative documents as a regression rather than a classification problem, as performed by coarse-grained approaches. We introduce Ekman's Basic Emotions (BE) and Russell and Mehrabian's Valence-Arousal-Dominance (VAD) model—two major schemes of emotion representation following opposing lines of psychological research, i.e., categorical and dimensional models—and discuss problems when BEs are used in a regression approach. We present the first natural language system thoroughly evaluated for fine-grained emotion analysis using the VAD scheme. Although we only employ simple BOW features, we reach correlation values up until r = .65 with human annotations. Furthermore, we show that the prevailing evaluation methodology relying solely on Pearson's correlation coefficient r is deficient which leads us to the introduction of a complementary error-based metric. Due to the lack of comparable (VAD-based) systems, we, finally, introduce a novel method of mapping between VAD and BE emotion representations to create a reasonable basis for comparison. This enables us to evaluate VAD output against human BE judgments and, thus, allows for a more direct comparison with existing BE-based emotion analysis systems. Even with this, admittedly, error-prone transformation step our VAD-based system achieves state-of-the-art performance in three out of six emotion categories, out-performing all existing BE-based systems but one.
Sven Buechel, Udo Hahn
ECAI2
2016 UIMA-Based JCoRe 2.0 Goes GitHub and Maven Central ― State-of-the-Art Software Resource Engineering and Distribution of NLP Pipelines
Udo Hahn, Franz Matthies, Erik Faessler, Johannes Hellrich
LREC1
2016 CodE Alltag: A German-Language E-Mail Corpus
Ulrike Krieg-Holz, Christian Schuschnig, Franz Matthies, Benjamin Redling, Udo Hahn
LREC5
2015 JUFIT: A Configurable Rule Engine for Filtering and Generating New Multilingual UMLS Terms
Johannes Hellrich, Stefan Schulz 0001, Sven Buechel, Udo Hahn
AMIA4
2014 Fostering Multilinguality in the UMLS: A Computational Approach to Terminology Expansion for Multiple Languages
Johannes Hellrich, Udo Hahn
AMIA2
2014 An Empirically Grounded Approach to Extend the Linguistic Coverage and Lexical Diversity of Verbal Probabilities
Christine Engelmann, Udo Hahn
CogSci2
2014 Disclose Models, Hide the Data - How to Make Use of Confidential Corpora without Seeing Sensitive Raw Data
Erik Faessler, Johannes Hellrich, Udo Hahn
LREC3
2014 Collaboratively Annotating Multilingual Parallel Corpora in the Biomedical Domain―some MANTRAs
Johannes Hellrich, Simon Clematide, Udo Hahn, Dietrich Rebholz-Schuhmann
LREC3
2014 Enhancing Multilingual Biomedical Terminologies via Machine Translation from Parallel Corpora
Johannes Hellrich, Udo Hahn
NLDB2
2014 Grounding Epistemic Modality in Speakers' Judgments
Udo Hahn, Christine Engelmann
PRICAI1
2012 Active Learning-Based Corpus Annotation - The PathoJen Experience
Udo Hahn, Elena Beisswanger, Ekaterina Buyko, Erik Faessler
AMIA1
2012 Iterative Refinement and Quality Checking of Annotation Guidelines - How to Deal Effectively with Semantically Sloppy Named Entity Types, such as Pathological Phenomena
Udo Hahn, Elena Beisswanger, Ekaterina Buyko, Erik Faessler, Jenny Traumüller, Susann Schröder, Kerstin Hornbostel
LREC1
2012 CALBC: Releasing the Final Corpora
Senay Kafkas, Ian Lewin, David Milward, Erik M. van Mulligen, Jan A. Kors, Udo Hahn, Dietrich Rebholz-Schuhmann
LREC6
2012 Mining the pharmacogenomics literature - a survey of the state of the art
abstract
This article surveys efforts on text mining of the pharmacogenomics literature, mainly from the period 2008 to 2011. Pharmacogenomics (or pharmacogenetics) is the field that studies how human genetic variation impacts drug response. Therefore, publications span the intersection of research in genotypes, phenotypes and pharmacology, a topic that has increasingly become a focus of active research in recent years. This survey covers efforts dealing with the automatic recognition of relevant named entities (e.g. genes, gene variants and proteins, diseases and other pathological phenomena, drugs and other chemicals relevant for medical treatment), as well as various forms of relations between them. A wide range of text genres is considered, such as scientific publications (abstracts, as well as full texts), patent texts and clinical narratives. We also discuss infrastructure and resources needed for advanced text analytics, e.g. document corpora annotated with corresponding semantic metadata (gold standards and training data), biomedical terminologies and ontologies providing domain-specific background knowledge at different levels of formality and specificity, software architectures for building complex and scalable text analytics pipelines and Web services grounded to them, as well as comprehensive ways to disseminate and interact with the typically huge amounts of semiformal knowledge structures extracted by text mining tools. Finally, we consider some of the novel applications that have already been developed in the field of pharmacogenomic text mining and point out perspectives for future research.
Udo Hahn, Kevin Cohen 0001, Yael Garten, Nigam H. Shah
Briefings Bioinform.1
2011 Towards Automatic Pathway Generation from Biological Full-Text Publications
Ekaterina Buyko, Jörg Linde, Steffen Priebe, Udo Hahn
IDA4
2011 U-Compare bio-event meta-service: compatible BioNLP event extraction services
abstract
BACKGROUND: Bio-molecular event extraction from literature is recognized as an important task of bio text mining and, as such, many relevant systems have been developed and made available during the last decade. While such systems provide useful services individually, there is a need for a meta-service to enable comparison and ensemble of such services, offering optimal solutions for various purposes. RESULTS: We have integrated nine event extraction systems in the U-Compare framework, making them intercompatible and interoperable with other U-Compare components. The U-Compare event meta-service provides various meta-level features for comparison and ensemble of multiple event extraction systems. Experimental results show that the performance improvements achieved by the ensemble are significant. CONCLUSIONS: While individual event extraction systems themselves provide useful features for bio text mining, the U-Compare meta-service is expected to improve the accessibility to the individual systems, and to enable meta-level uses over multiple event extraction systems such as comparison and ensemble.
Yoshinobu Kano, Jari Björne, Filip Ginter, Tapio Salakoski, Ekaterina Buyko, Udo Hahn, Kevin Cohen 0001, Karin Verspoor, Christophe Roeder, Lawrence Hunter, Halil Kilicoglu, Sabine Bergler, Sofie Van Landeghem, Thomas Van Parys, Yves Van de Peer, Makoto Miwa, Sophia Ananiadou, Mariana L. Neves, Alberto D. Pascual-Montano, Arzucan Özgür, Dragomir R. Radev, Sebastian Riedel 0001, Rune Sætre, Hong-Woo Chun, Jin-Dong Kim, Sampo Pyysalo, Tomoko Ohta, Jun'ichi Tsujii
BMC Bioinform.6
2011 Syntactic Simplification and Semantic Enrichment - Trimming Dependency Graphs for Event Extraction
abstract
In our approach to event extraction, dependency graphs constitute the fundamental data structure for knowledge capture. Two types of trimming operations pave the way to more effective relation extraction. First, we simplify the syntactic representation structures resulting from parsing by pruning informationally irrelevant lexical material from dependency graphs. Second, we enrich informationally relevant lexical material in the simplified dependency graphs with additional semantic meta data at several layers of conceptual granularity. These two aggregation operations on linguistic representation structures are intended to avoid overfitting of machine learning-based classifiers which we use for event extraction (besides manually curated dictionaries). Given this methodological framework, the corresponding JReX system developed by the JulieLab Team from Friedrich-Schiller-Universität Jena (Germany) scored on 2nd rank among 24 competing teams for Task 1 in the “BioNLP’09 Shared Task on Event Extraction,” with 45.8% recall, 47.5% precision and 46.7% F1-score on all 3,182 events. In more recent experiments, based on slight modifications of JReX and using the same data sets, we were able to achieve 45.9% recall, 57.7% precision, and 51.1% F1-score.
Ekaterina Buyko, Erik Faessler, Joachim Wermter, Udo Hahn
Comput. Intell.4
2010 A Cognitive Cost Model of Annotations Based on Eye-Tracking Data
Katrin Tomanek, Udo Hahn, Steffen Lohmann, Jürgen Ziegler 0001
ACL2
2010 Evaluating the Impact of Alternative Dependency Graph Encodings on Solving Event Extraction Tasks
Ekaterina Buyko, Udo Hahn
EMNLP2
2010 The GeneReg Corpus for Gene Expression Regulation Events - An Overview of the Corpus and its In-Domain and Out-of-Domain Interoperability
Ekaterina Buyko, Elena Beisswanger, Udo Hahn
LREC3
2010 The CALBC Silver Standard Corpus for Biomedical Named Entities - A Study in Harmonizing the Contributions from Four Independent Named Entity Taggers
Dietrich Rebholz-Schuhmann, Antonio Jimeno-Yepes, Erik M. van Mulligen, Ning Kang 0002, Jan A. Kors, David Milward, Peter T. Corbett, Ekaterina Buyko, Katrin Tomanek, Elena Beisswanger, Udo Hahn
LREC11
2010 Annotation Time Stamps - Temporal Metadata from the Linguistic Annotation Process
Katrin Tomanek, Udo Hahn
LREC2
2010 Introduction to Linguistic Annotation and Text Analytics Graham Wilcock (University of Helsinki) Princeton, NJ: Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 2, No. 1), 2009, x+149 pp; paperbound, ISBN 978-1-59829-738-6, $40.00; ebook, ISBN 978-1-59829-739-3, $30.00 or by subscription
Udo Hahn
Comput. Linguistics1
2009 Semi-Supervised Active Learning for Sequence Labeling
Katrin Tomanek, Udo Hahn
ACL/IJCNLP2
2009 Reducing class imbalance during active learning for named entity annotation
abstract
In lots of natural language processing tasks, the classes to be dealt with often occur heavily imbalanced in the underlying data set and classifiers trained on such skewed data tend to exhibit poor performance for low-frequency classes. We introduce and compare different approaches to reduce class imbalance by design within the context of active learning (AL). Our goal is to compile more balanced data sets up front during annotation time when AL is used as a strategy to acquire training material. We situate our approach in the context of named entity recognition. Our experiments reveal that we can indeed reduce class imbalance and increase the performance of classifiers on minority classes while preserving a good overall performance in terms of macro F-score.
Katrin Tomanek, Udo Hahn
K-CAP2
2009 Do User Appreciate Novel Interface Features for Literature Search? - A User Study in the Life Sciences Domain
abstract
Faced with the challenges to design an easy-to-use, immediately comprehensible and powerful expert-user interface to search very large document collections in the life sciences, we developed several system prototypes. Their main features were faceting of the domain vocabulary for browsing and searching, flexible search-state-dependent drilling of the terminological hierarchy, dynamic query term auto-completions, and highlighting of matched terms (including synonyms and spelling variants). Under lab conditions we then evaluated these features in several task-based scenarios using camera recordings, thinking-aloud protocols and questionnaires. The results reveal that faceting and highlighting were very well received, while auto-completions seemed less important or were misconceptualized as spelling aids.
Joachim Wermter, Udo Hahn, Rico Landefeld, Anne Schneider
SMC2
2009 MaHCO: an ontology of the major histocompatibility complex for immunoinformatic applications and text mining
abstract
MOTIVATION: The high level of polymorphism associated with the major histocompatibility complex (MHC) poses a challenge to organizing associated bioinformatic data, particularly in the area of hematopoietic stem cell transplantation. Thus, this area of research has great potential to profit from the ongoing development of biomedical ontologies, which offer structure and definition to MHC-data related communication and portability issues. RESULTS: We introduce the design considerations, methodological foundations and implementational issues underlying MaHCO, an ontology which represents the alleles and encoded molecules of the major histocompatibility complex. Importantly for human immunogenetics, it includes a detailed level of human leukocyte antigen (HLA) classification. We then present an ontology browser, search interfaces for immunogenetic fact and document retrieval, and the specification of an annotation language for semantic metadata, based on MaHCO. These use cases are intended to demonstrate the utility of ontology-driven bioinformatics in the field of immunogenetics. AVAILABILITY AND IMPLEMENTATION: The MaHCO Ontology is available via the BioPortal: http://www.bioontology.org/tools/portal/bioportal.html, and at: http://purl.org/stemnet/.
David S. DeLuca, Elena Beisswanger, Joachim Wermter, Peter A. Horn, Udo Hahn, Rainer Blasczyk
Bioinform.5
2009 High-performance gene name normalization with GENO
abstract
MOTIVATION: The recognition and normalization of textual mentions of gene and protein names is both particularly important and challenging. Its importance lies in the fact that they constitute the crucial conceptual entities in biomedicine. Their recognition and normalization remains a challenging task because of widespread gene name ambiguities within species, across species, with common English words and with medical sublanguage terms. RESULTS: We present GeNo, a highly competitive system for gene name normalization, which obtains an F-measure performance of 86.4% (precision: 87.8%, recall: 85.0%) on the BioCreAtIvE-II test set, thus being on a par with the best system on that task. Our system tackles the complex gene normalization problem by employing a carefully crafted suite of symbolic and statistical methods, and by fully relying on publicly available software and data resources, including extensive background knowledge based on semantic profiling. A major goal of our work is to present GeNo's architecture in a lucid and perspicuous way to pave the way to full reproducibility of our results. AVAILABILITY: GeNo, including its underlying resources, will be available from www.julielab.de. It is also currently deployed in the Semedico search engine at www.semedico.org.
Joachim Wermter, Katrin Tomanek, Udo Hahn
Bioinform.3
2008 Multi-Task Active Learning for Linguistic Annotations
Roi Reichart, Katrin Tomanek, Udo Hahn, Ari Rappoport
ACL3
2008 Are Morpho-Syntactic Features More Predictive for the Resolution of Noun Phrase Coordination Ambiguity than Lexico-Semantic Similarity Scores?
Ekaterina Buyko, Udo Hahn
COLING2
2008 Semantic Annotations for Biology: a Corpus Development Initiative at the Jena University Language & Information Engineering (JULIE) Lab
Udo Hahn, Elena Beisswanger, Ekaterina Buyko, Michael Poprat, Katrin Tomanek, Joachim Wermter
LREC1
2008 Approximating Learning Curves for Active-Learning-Driven Annotation
Katrin Tomanek, Udo Hahn
LREC2
2007 An Ontology for Major Histocompatibility Complex (MHC) Alleles and Molecules
Elena Beisswanger, David S. DeLuca, Rainer Blasczyk, Udo Hahn
AMIA4
2007 An Approach to Text Corpus Construction which Cuts Annotation Costs and Maintains Reusability of Annotated Data
Katrin Tomanek, Joachim Wermter, Udo Hahn
EMNLP-CoNLL3
2007 Ontological foundations for biomedical sciences
Udo Hahn, Stefan Schulz 0001
Artif. Intell. Medicine1
2007 Towards the ontological foundations of symbolic biological theories
Stefan Schulz 0001, Udo Hahn
Artif. Intell. Medicine2
2007 Spatial location and its relevance for terminological inferences in bio-ontologies
abstract
BACKGROUND: An adequate and expressive ontological representation of biological organisms and their parts requires formal reasoning mechanisms for their relations of physical aggregation and containment. RESULTS: We demonstrate that the proposed formalism allows to deal consistently with "role propagation along non-taxonomic hierarchies", a problem which had repeatedly been identified as an intricate reasoning problem in biomedical ontologies. CONCLUSION: The proposed approach seems to be suitable for the redesign of compositional hierarchies in (bio)medical terminology systems which are embedded into the framework of the OBO (Open Biological Ontologies) Relation Ontology and are using knowledge representation languages developed by the Semantic Web community.
Stefan Schulz 0001, Kornél G. Markó, Udo Hahn
BMC Bioinform.3
2006 You Can't Beat Frequency (Unless You Use Linguistic Knowledge) - A Qualitative Evaluation of Association Measures for Collocation and Term Extraction
abstract
In the past years, a number of lexical association measures have been studied to help extract new scientific terminology or general-language collocations. The implicit assumption of this research was that newly designed term measures involving more sophisticated statistical criteria would outperform simple counts of co-occurrence frequencies. We here explicitly test this assumption. By way of four qualitative criteria, we show that purely statistics-based measures reveal virtually no difference compared with frequency of occurrence counts, while linguistically more informed metrics do reveal such a marked difference.
Joachim Wermter, Udo Hahn
ACL2
2006 Towards an Upper-Level Ontology for Molecular Biology
Stefan Schulz 0001, Elena Beisswanger, Joachim Wermter, Udo Hahn
AMIA4
2006 From GENIA to BIOTOP - Towards a Top-Level Ontology for Biology
Stefan Schulz 0001, Elena Beisswanger, Udo Hahn, Joachim Wermter, Anand Kumar 0005, Holger Stenzhorn
FOIS3
2006 Semantic Atomicity and Multilinguality in the Medical Domain: Design Considerations for the MorphoSaurus Subword Lexicon
Stefan Schulz 0001, Kornél G. Markó, Philipp Daumke, Udo Hahn, Susanne Hanser, Percy Nohama, Roosewelt L. Andrade, Edson José Pacheco, Martin Romacker
LREC4
2006 Semantic Mining in Biomedicine (Introduction to the papers selected from the SMBM 2005 Symposium, Hinxton, U.K., April 2005)
abstract
Researchers working in the life sciences domain in the past years have witnessed an enormous growth of literature—for the whole field as well as for their highly specialized areas of expertise. Only small portions of the biomedical knowledge are accessible in a structured way, i.e. through formatted databases. These few pieces of textually encoded knowledge that have gone into databases are, by default, manually extracted from documents and manually inserted into databases after careful curation efforts by highly skilled domain experts. Still, the vast majority of biomedical knowledge captured in texts is not at disposal when biomedical databases are queried. Life scientists have realized this loss of possibly highly relevant information and devised various forms of support. The weakest one is provided by information retrieval (IR) systems [for a life-science-centred survey; cf. Hersh (2002)]. Given a user-formulated query the terms from this query are appropriately matched with the terms occurring in documents from a large collection (e.g. the, currently, 14 million abstracts from Medline). Documents matching the query (up to a specified degree) are returned to the user for closer inspection and, possibly, ranked by some relevance-based sorting criterion (e.g. closeness of match). Information extraction (IE) provides a more powerful alternative that has mainly been developed in other areas different from molecular biology by the natural language processing community. IE aims at directly extracting relevant information from natural language documents [usually original text snippets, sentences, relevant phrases or even quasi-logical propositions, such as predicate–argument structures—for a general survey, cf. Gaizauskas and Wilks (1998) and for a life-science-centred view, cf. Blaschke et al. (2002) and Hoffmann et al. (2005)]. Unlike the output of IR systems, which only list relevant documents, IE systems provide immediate access to relevant information pieces via pre-specified information templates. This is achieved, however, at the price of supplying rather sophisticated language processing methodologies [e.g. taggers, chunkers, light semantic interpreters and information extraction rules; cf. for a survey, Hahn and Wermter (2006)], domain-specific developments and resources (e.g. databases and ontologies) and machine learning methodologies usually lack in IR systems. The evaluation of the degree of achievements from a biomedical perspective is an issue of active research. The IR stream is currently mainly investigated in the TREC (Text Retrieval Conference) Genomics track (Author Webpage), whereas there are several challenge evaluation platforms for IE that deal with often complementary problems from a biological perspective, the most important, currently, being the BioCreAtIvE (Critical Assessment of Information Extraction systems in Biology) contest (Author Webpage) [surveyed in Hirschman et al. (2005); see also Blaschke et al. (2005)]. Both forms of activities, IR as well as IE, are often labelled as text mining but miss a major extra requirement, namely the knowledge discovery perspective usually attributed to text mining procedures as well (Hearst, 1999). In particular, this relates to the identification and elimination of redundant knowledge as well as the recognition of (user-new?, expert-new? and community-new?) novel information. This value-adding, summarizing and selective aspect of text mining could be particularly helpful in taming the flood of literature for biomedical researchers, and will certainly be the focus of new developments in the years to come. The challenge evaluations, however, have already revealed some of the most pressing research problems for text analysis in the biomedical domain. In particular, biomedical terminology is extremely hard to deal with, in part because of the poor introduction of standards. It starts from identifying biological terms in a document (terms have a complex internal structure and are often composed of multiple, up to four or five, words), and leads to determining their conceptual type (e.g. genes, proteins and cell lines) and the way they are relationally linked (e.g. in terms of taxonomies or partonomies that in biology are often related with the organization of protein families). Further on, concrete factual biomedical knowledge (sometimes called relation mining) is also hard to extract from documents (e.g. ‘protein X inhibits protein Y’). This step is crucial for any sort of automated functional annotation in biological databases. Once this kind of knowledge has been successfully captured on a large scale making thousands of these propositions available, another severe follow-up problem arises, namely how to communicate this mass of information in a concise, comprehensible and, finally, useful way to the researcher in the laboratory. For this purpose, text mining systems have simply borrowed visualization techniques that were originally developed for numerical data mining. However, symbolic abstraction mechanisms leading, e.g. to the automatic generation of pathway diagrams from this huge dataset are still an area that requires further developments. The above-mentioned research problems have motivated the creation of a Network of Excellence—‘Semantic Interoperability and Data Mining in Biomedicine’ (Semantic Mining, Author Webpage)—which has been funded by the European Community since 2004 under the FP6 Programme ‘Integrating and Strengthening the European Research Area’. The NoE has initiated a series of conferences dedicated to these particular challenges of data mining and text mining in the life sciences [for a life-science-centred survey, cf. Ananiadou and McNaught (2006)]. The first of these symposia was held under the title ‘Semantic Mining in Biomedicine’ (SMBM) in Hinxton (Cambridgeshire, UK) from April 10–13, 2005 organized by Stefan Schulz, Freiburg University Hospital, and Dietrich Rebholz-Schuhmann, EBI-EMBL, Hinxton (see Author Webpage). A specific feature of SMBM meetings is their focus on content-oriented methodologies and semantic resources—either controlled vocabularies, terminologies and formal domain ontologies, or conceptually as well as propositionally annotated corpora—in order to improve text-based biomedical knowledge management, e.g. through document classification, text or fact retrieval, information extraction, or (real) text mining. Also methodologies being discussed should look at applications to real-world problems in molecular biology and biomedicine [for a review of systems currently operational in this domain, see Krallinger and Valencia (2005)]. We had the honour of chairing the programme committee that comprised 21 scientists who evaluated the 28 submissions and selected 12 papers for their presentation in the conference. Four outstanding papers were selected for publication in Bioinformatics, after additional extensive reviews and revisions. Seven full papers plus the abstracts of these selected presentations appeared in the proceedings of the conference [Hahn and Valencia (2005)]. The selected papers cover research performed under the following headings: (1) entity identification—identification of gene names, (2) text classification classification—assignment of sentences to known Gene Ontology (GO) and Medical Subject Headings (MeSH) classes, (3) identification of relations in text—extracting phosphorylation and gene control networks; and (4) identification of new concepts—proposing new GO categories and their corresponding associated genes. Automatic Term List Generation for Entity Tagging by Ted Sandler, Andrew I. Schein and Lyle H. Ungar from the University of Pennsylvania. The basic problem of term characterization is tackled here with an unsupervised approach based on clustering terms (gene names) using additional context information. The clustering approach is related to the distributional clustering technique published previously and the context information provided include neighbouring and syntactic relations. The basic sources of information were sentences from the Biocreative gene tagging challenge and a set of two million Medline abstracts. The results are significantly better than those obtained with standard taggers based on dictionaries of genes. Interestingly enough, the results are still far from matching those obtained in other domains such as newswire information, most probably owing to the additional complexity of biological nomenclature. Automatic Assignment of Biomedical Categories: Toward a Generic Approach by Patrick Ruch from the University Hospitals of Geneva. Describes new results on the automatical assignment of biomedical categories with a system that is designed to be largely data-independent. The system includes a pattern-based identification and vector space retrieval engine, and uses both stems and linguistically motivated information, and it is applied to the classification of sentences in MeSH and GO classes. The results are compared with those obtained in the related BioCreative task. Extraction of Regulatory Gene/Protein Networks from Medline by Jasmin Saric, Lars Juhl Jensen, Rossitza Ouzounova, Isabel Rojas and Peer Bork, from EML Research and EMBL both in Heidelberg. The authors address the problem of extracting two key types of biological relations, which are the regulators of protein function by phosphorylation and the control of gene expression. Their rule-based String-IE system uses organism-specific lexicons that are incorporated in the training of a part-of-speech tagger that uses the GENIA corpus as background information. In practice, the system is able to extract 3319 phosphorylations or gene expression relations, with a sustained level of accuracy across different organisms. Automatic Extension of GO with Flexible Identification of Candidate Terms by Jin-Bok Lee, Jung-jae Kim and Jong C. Park from KAIST in Daejeon, Korea. The authors tackle the problem of identifying new GO concepts using existing GO concepts and their relations in text. The proposed new terms are compared with those created by human experts in subsequent releases of GO. This type of approaches can be useful for speeding up the process of annotation, and for increasing the number of categories in which GO concepts can be divided when they have a large number of genes assigned.
Udo Hahn, Alfonso Valencia
Bioinform.1
2006 Towards new information resources for public health - From WordNet to MedicalWordNet
Christiane Fellbaum, Udo Hahn, Barry Smith 0001
J. Biomed. Informatics2
2005 Unsupervised Multilingual Word Sense Disambiguation via an Interlingua
Kornél G. Markó, Stefan Schulz 0001, Udo Hahn
AAAI3
2005 How to Distinguish Parthood from Location in Bio-Ontologies
Stefan Schulz 0001, Philipp Daumke, Barry Smith 0001, Udo Hahn
AMIA4
2005 Effective Grading of Termhood in Biomedical Literature
Joachim Wermter, Udo Hahn
AMIA2
2005 Cross-Language Mining for Acronyms and Their Completions from the Web
Udo Hahn, Philipp Daumke, Stefan Schulz 0001, Kornél G. Markó
Discovery Science1
2005 Massive Biomedical Term Discovery
Joachim Wermter, Udo Hahn
Discovery Science2
2005 Finding new terminology in very large corpora
abstract
Most technical and scientific terms are comprised of complex, multi-word noun phrases but certainly not all noun phrases are technical or scientific terms. The distinction of specific terminology from common non-specific noun phrases can be based on the observation that terms reveal a much lesser degree of distributional variation than non-specific noun phrases. We formalize the limited paradigmatic modifiability of terms and, subsequently, test the corresponding algorithm on bigram, trigram and quadgram noun phrases extracted from a 104-million-word biomedical text corpus. Using an already existing and community-wide curated biomedical terminology as an evaluation gold standard, we show that our algorithm significantly outperforms standard term identification measures and, therefore, qualifies as a high-performant building block for any terminology identification system. We also provide empirical evidence that the superiority of our approach, beyond a 10-million-word threshold, is essentially domain- and corpus-size-independent.
Joachim Wermter, Udo Hahn
K-CAP2
2005 Subword Clusters as Light-Weight Interlingua for Multilingual Document Retrieval
abstract
We introduce a light-weight interlingua for a cross-language document retrieval system in the medical domain. It is composed of equivalence classes of semantically primitive, language-specific subwords which are clustered by interlingual and intralingual synonymy. Each subword cluster represents a basic conceptual entity of the language-independent interlingua. Documents, as well as queries, are mapped to this interlingua level on which retrieval operations are performed. Evaluation experiments reveal that this interlingua-based retrieval model outperforms a direct translation approach.
Udo Hahn, Kornél G. Markó, Stefan Schulz 0001
MTSummit1
2005 Bootstrapping dictionaries for cross-language information retrieval
abstract
The bottleneck for dictionary-based cross-language information retrieval is the lack of comprehensive dictionaries, in particular for many different languages. We here introduce a methodology by which multilingual dictionaries (for Spanish and Swedish) emerge automatically from simple seed lexicons. These seed lexicons are automatically generated, by cognate mapping, from (previously manually constructed) Portuguese and German as well as English sources. Lexical and semantic hypotheses are then validated and new ones iteratively generated by making use of co-occurrence patterns of hypothesized translation synonyms in parallel corpora. We evaluate these newly derived dictionaries on a large medical document collection within a cross-language retrieval setting.
Kornél G. Markó, Stefan Schulz 0001, Olena Medelyan, Udo Hahn
SIGIR4
2005 Part-whole representation and reasoning in formal biomedical ontologies
Stefan Schulz 0001, Udo Hahn
Artif. Intell. Medicine2
2004 Learning Indexing Patterns from One Language for the Benefit of Others
Udo Hahn, Kornél G. Markó, Stefan Schulz 0001
AAAI1
2004 Mereological Semantics for Bio-Ontologies
Udo Hahn, Stefan Schulz 0001, Kornél G. Markó
AAAI1
2004 High-Performance Tagging on Medical Texts
Udo Hahn, Joachim Wermter
COLING1
2004 Cognate Mapping - A Heuristic Strategy for the Semi-Supervised Acquisition of a Spanish Lexicon from a Portuguese Seed Lexicon
Stefan Schulz 0001, Kornél G. Markó, Eduardo Sbrissia, Percy Nohama, Udo Hahn
COLING5
2004 Collocation Extraction Based on Modifiability Statistics
Joachim Wermter, Udo Hahn
COLING2
2004 Representing Natural Kinds by Spatial Inclusion and Containment
Stefan Schulz 0001, Udo Hahn
ECAI2
2004 Parthood as Spatial Inclusion - Evidence from biomedical Conceptualizations
Stefan Schulz 0001, Udo Hahn
KR2
2004 Pumping Documents Through a Domain and Genre Classification Pipeline
Udo Hahn, Joachim Wermter
LREC1
2004 An Annotated German-Language Medical Text Corpus as Language Resource
Joachim Wermter, Udo Hahn
LREC2
2004 Tagging Medical Documents with High Accuracy
Udo Hahn, Joachim Wermter
PRICAI1
2003 Cross-language MeSH Indexing using Morpho-Semantic Normalization
Kornél G. Markó, Philipp Daumke, Stefan Schulz 0001, Udo Hahn
AMIA4
2002 A knowledge representation view on biomedical structure and function
Stefan Schulz 0001, Udo Hahn
AMIA2
2002 MORPHOSAURUS: Crosslingual Medical Text Retrieval by Subword Indexing
Stefan Schulz 0001, Rüdiger Klar, Udo Hahn, Martin Romacker, Percy Nohama, Lúcio J. Dias Matias
AMIA3
2002 Turning Lead into Gold? Feeding a Formal Knowledge Base with Informal Conceptual Knowledge
Udo Hahn, Stefan Schulz 0001
EKAW1
2002 Towards Very Large Ontologies for Medical Language Processing
Udo Hahn, Stefan Schulz 0001
LREC1
2002 240, 000 concepts and relations-towards mega knowledge bases for real-world applications
abstract
We describe an ontology engineering methodology by which conceptual knowledge is extracted from an informal medical thesaurus (UMLS) and automatically converted into a formal description logics system (LOOM). Our approach consists of four steps: concept definitions are automatically generated from the UMLS, Integrity checking of taxonomic and partonomic hierarchies is performed by LOOM's terminological classifier, cycles and inconsistencies are eliminated, as well as incremental refinement of the evolving knowledge base is performed by a domain expert. We report an experiments with a very large knowledge base composed of 164,000 concepts and 76,000 relations.
Udo Hahn, Stefan Schulz 0001
SMC (2)1
2002 Understanding metonymies in discourse
Katja Markert, Udo Hahn
Artif. Intell.2
2002 The Theory and Practice of Discourse Parsing and Summarization by Daniel Marcu
Udo Hahn
Comput. Linguistics1
2002 An integrated, dual learner for grammars and ontologies
Udo Hahn, Kornél G. Markó
Data Knowl. Eng.1
2001 Semantic Interpretation of Medical Language - Quantitative Analysis and Qualitative Yield
Martin Romacker, Udo Hahn
AIME2
2001 Parts, Locations, and Holes - Formal Reasoning about Anatomical Structures
Stefan Schulz 0001, Udo Hahn
AIME2
2001 Subword segmentation-leveling out morphological variations for medical document retrieval
Udo Hahn, Martin Honeck, Michael Piotrowski, Stefan Schulz 0001
AMIA1
2001 Empirical data for the semantic interpretation of prepositional phrases in medical documents
Martin Romacker, Udo Hahn
AMIA2
2001 Mereotopological reasoning about parts and (w)holes in bio-ontologies
abstract
We here deal with mereotopological properties of parts and associated wholes, locations and empty spaces (holes), with particular reference to biological structures. Our considerations lead to a basic ontology which contains 'solid object', 'hole' and 'boundary' as mutually disjoint primitives. Formally, we embed the relations 'part-of' and 'location-of' into a parsimonious description logic (ALC) and emulate partonomic and spatial reasoning involving these relations by terminological subsumption. In contrast to common conceptualizations, we do not distinguish between solids and the regions they occupy, as well as we allow solids to have holes as proper parts. In order to support these modeling decisions, we discuss various concrete examples from human anatomy.
Stefan Schulz 0001, Udo Hahn
FOIS2
2001 A Search Engine for Morphologically Complex Languages
Udo Hahn, Martin Honeck, Stefan Schulz 0001
IDA1
2001 Joint knowledge capture for grammars and ontologies
abstract
We introduce a methodology for automating the maintenance and growth of domain-specific concept taxonomies and grammatical class hierarchies simultaneously, based on knowledge capture from natural language texts. The assimilation process is centered around the linguistic and conceptual `quality' of various forms of evidence underlying the generation, assessment and on-going refinement of lexical and concept hypotheses. On the basis of the strength of evidence, hypotheses are ranked according to plausibility, and the most reasonable ones are selected for assimilation into the given lexical class hierarchy and domain ontology.
Udo Hahn, Kornél G. Markó
K-CAP1
2000 Automated coding of diagnoses-three methods compared
Pius Franz, Albrecht Zaiß, Stefan Schulz 0001, Udo Hahn, Rüdiger Klar
AMIA4
2000 MedSynDiKATe-design considerations for an ontology-based medical text understanding system
Udo Hahn, Martin Romacker, Stefan Schulz 0001
AMIA1
2000 Modeling anatomical spatial relations with description logics
Stefan Schulz 0001, Udo Hahn, Martin Romacker
AMIA2
2000 An Integrated Model of Semantic and Conceptual Interpretation from Dependency Structures
Udo Hahn, Martin Romacker
COLING1
2000 Knowledge Engineering by Large-Scale Knowledge Reuse - Experience from the Medical Domain
Stefan Schulz 0001, Udo Hahn
KR2
2000 Coping with Different Types of Ambiguity Using a Uniform Context Handling Mechanism
Martin Romacker, Udo Hahn
NLDB2
2000 Content management in the SYNDIKATE system - How technical documents are automatically transformed to text knowledge bases
Udo Hahn, Martin Romacker
Data Knowl. Eng.1
1999 Streamlining semantic interpretation for medical narratives
Martin Romacker, Stefan Schulz 0001, Udo Hahn
AMIA3
1999 Automatic Import and Manual Refinement of Medical Knowledge
Stefan Schulz 0001, Giovanni Faggioli, Martin Romacker, Udo Hahn
AMIA4
1999 Text Understanding for Knowledge Base Generation in the SYNDIKATE System
Udo Hahn, Martin Romacker
DEXA1
1999 Lean Semantic Interpretation
Martin Romacker, Udo Hahn, Katja Markert
IJCAI2
1999 Scalable Temporal Reasoning
Steffen Staab, Udo Hahn
IJCAI2
1999 How knowledge drives understandingmatching medical ontologies with the needs of medical language processing
Udo Hahn, Martin Romacker, Stefan Schulz 0001
Artif. Intell. Medicine1
1999 Functional Centering - Grounding Referential Coherence in Information Structure
Michael Strube 0001, Udo Hahn
Comput. Linguistics2
1998 Part-whole reasoning in medical ontologies revisited-introducing SEP triplets into classification-based description logics
Stefan Schulz 0001, Martin Romacker, Udo Hahn
AMIA3
1998 Knowledge Generation from Texts
Udo Hahn
ECAI1
1998 Text Summarization Based on Terminological Logics
Udo Hahn, Ulrich Reimer
ECAI1
1997 Centering in-the-Large: Computing Referential Discourse Segments
abstract
We specify an algorithm that builds up a hierarchy of referential discourse segments from local centering data. The spatial extension and nesting of these discourse segments constrain the reachability of potential antecedents of an anaphoric expression beyond the local level of adjacent center pairs. Thus, the centering model is scaled up to the level of the global referential structure of discourse. An empirical evaluation of the algorithm is supplied.
Udo Hahn, Michael Strube 0001
ACL1
1997 Text structures in medical text processing: empirical evidence and a text understanding prototype
Udo Hahn, Martin Romacker
AMIA1
1997 Knowledge Mining from Textual Sources
abstract
Article Knowledge mining from textual sources Share on Authors: Udo Hahn Computational Linguistics Lab - Text Knowledge Engineering Group, Freiburg University, Werthmannplatz 1, D-79085 Freiburg, Germany Computational Linguistics Lab - Text Knowledge Engineering Group, Freiburg University, Werthmannplatz 1, D-79085 Freiburg, GermanyView Profile , Klemens Schnattinger Computational Linguistics Lab - Text Knowledge Engineering Group, Freiburg University, Werthmannplatz 1, D-79085 Freiburg, Germany Computational Linguistics Lab - Text Knowledge Engineering Group, Freiburg University, Werthmannplatz 1, D-79085 Freiburg, GermanyView Profile Authors Info & Claims CIKM '97: Proceedings of the sixth international conference on Information and knowledge managementJanuary 1997 Pages 83–90https://doi.org/10.1145/266714.266865Online:01 January 1997Publication History 5citation572DownloadsMetricsTotal Citations5Total Downloads572Last 12 Months10Last 6 weeks2 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Udo Hahn, Klemens Schnattinger
CIKM1
1997 Text Knowledge Engineering by Qualitative Terminological Learning
Udo Hahn, Klemens Schnattinger
DEXA1
1997 Intelligent Text Analysis for Dynamically Maintaining and Updating Domain Knowledge Bases
Klemens Schnattinger, Udo Hahn
IDA2
1997 On the Interaction of Metonymies and Anaphora
Katja Markert, Udo Hahn
IJCAI (2)2
1997 "Tall", "Good", "High" - Compared to What?
Steffen Staab, Udo Hahn
IJCAI (2)2
1997 Deep Knowledge Discovery from Natural Language Texts
Udo Hahn, Klemens Schnattinger
KDD1
1996 Functional Centering
abstract
Based on empirical evidence from a free word order language (German) we propose a fundamental revision of the principles guiding the ordering of discourse entities in the forward-looking centers within the centering model. We claim that grammatical role criteria should be replaced by indicators of the functional information structure of the utterances, i.e., the distinction between context-bound and unbound discourse elements. This claim is backed up by an empirical evaluation of functional centering.
Michael Strube 0001, Udo Hahn
ACL2
1996 Bridging Textual Ellipses
Udo Hahn, Michael Strube 0001, Katja Markert
COLING1
1996 Restricted Parallelism in Object-Oriented Lexical Parsing
Peter Neuhaus, Udo Hahn
COLING2
1996 A Conceptual Reasoning Approach to Textual Ellipsis
Udo Hahn, Katja Markert, Michael Strube 0001
ECAI1
1995 ParseTalk about Sentence- and Text-Level Anaphora
Michael Strube 0001, Udo Hahn
EACL2
1994 Concurrent Lexicalized Dependency Parsing: The ParseTalk Model
Norbert Bröker, Udo Hahn, Susanne Schacht
COLING2
1994 Concurrent Lexicalized Dependency Parsing: A Behavioral View On ParseTalk Events
Susanne Schacht, Udo Hahn, Norbert Bröker
COLING2
1994 Tracking the Evolution of Concepts in Dynamic Worlds
Udo Hahn, Manfred Klenner
DEXA1
1994 Concept Versioning: A Methodology for Tracking Evolutionary Concept Drift in Dynamic Concept Systems
Manfred Klenner, Udo Hahn
ECAI2
1994 Concurrent, object-oriented natural language parsing: the ParseTalk model
Udo Hahn, Susanne Schacht, Norbert Bröker
Int. J. Hum. Comput. Stud.1
1992 On Text Coherence Parsing
Udo Hahn
COLING1
1991 Teamwork Support in a Knowledge-Based Information Systems Environment
abstract
Development assistance for interactive database applications (DAIDA) is an experimental environment for the knowledge-assisted development and maintenance of database-intensive information systems from object-oriented requirements and specifications. Within the DAIDA framework, an approach to integrate different tasks encountered in software projects via a conceptual modeling strategy has been developed. Emphasis is put on integrating the semantics of the software development domain with aspects of group work, on social strategies to negotiate problems by argumentation, and on assigning responsibilities for task fulfillment by way of contracting. The implementation of a prototype is demonstrated with a sample session.>
Udo Hahn, Matthias Jarke, Thomas Rose 0001
IEEE Trans. Software Eng.1
1990 Panel: Hypertext: "Growing Up?"
abstract
This panel will employ two different interpretations of the phrase “growing up” to address areas of common interest between hypertext and information retrieval researchers. First, the panelists will question whether or not hypertext is “growing up” as a scientific discipline; They will discuss characteristics that separate hypertext research from other related disciplines. Second, the panelists will discuss the problems encountered when a hypertext system “grows up” in size and complexity; They will discuss the very real problems expected when representing and integrating large knowledge bases, accommodating multiple users, and distributing single logical hypertexts across multiple physical sites.
Mark E. Frisse, Maristella Agosti, Marie-France Bruandet, Udo Hahn, Stephen F. Weiss
SIGIR4
1990 Topic parsing: Accounting for text macro structures in full-text analysis
Udo Hahn
Inf. Process. Manag.1
1989 CoNeX: Coordination and Negotiation Support for Expert Teams in Project Management
Udo Hahn, Matthias Jarke
ECSCW1
1989 Making understanders out of parsers: Semantically driven parsing as a key concept for realistic text understanding applications
abstract
Semantically driven natural language parsers have found wide-spread application as a text processing methodology for knowledge-based information retrieval systems. It is argued that this parsing technique particularly corresponds to the requirements inherent to large-scale text analysis. Unfortunately, this approach suffers from several shortcomings which demand a thorough reformulation of its paradigm. Incorporating principles from conceptual analysis and word expert parsing in a model of lexically distributed text parsing, the focus of the modifications proposed in this article, is on a clean declarative separation of linguistio and other knowledge representation levels, abstraction mechanisms leading to a small collection of specification primitives for the parser, and an attempt to incorporate linguistic generalizations and modularization principles into the design of a semantic text grammar. A sample parse illustrates the operation and linguistic coverage of a lexically distributed text parser based on these theoretical considerations with respect to the semantic analysis of noun groups, simple assertional sentences, nominal anaphora, and textual ellipsis.
Udo Hahn
Int. J. Intell. Syst.1
1989 Book review
Udo Hahn
Mach. Transl.2
1986 Topic Essentials
Udo Hahn, Ulrich Reimer
COLING1
1986 A Generalized Word Expert Model of Lexically Distributed Text Parsing
Udo Hahn
ECAI1
1984 Textual Expertise In Word Experts: An Approach To Text Parsing Based On Topic/Comment Monitoring
abstract
In this paper prototype versions of two word experts for text analysis are dealt with which demonstrate that word experts are a feasible tool for parsing texts on the level of text cohesion as well as text coherence. The analysis is based on two major knowledge sources: context information is modelled in terms of a frame knowledge base, while the co-text keeps record of the linear sequencing of text analysis. The result of text parsing consists of a text graph reflecting the thematic organization of topics in a text.
Udo Hahn
COLING1
1984 Computing text Constituency: An Algorithmic Approach to the Generation of Text Graphs
Udo Hahn, Ulrich Reimer
SIGIR1
1983 A Formal Approach to the Semantics of a Frame Data Model
Ulrich Reimer, Udo Hahn
IJCAI2