VLDB 2026 Research / reviewers in the wild / expert
Jannik Strötgen
dblp:28/8510
· DBLP profile ↗
33ranked-venue papers
10as first author
12since 2021 · last 2025
0000-0002-7961-4798ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 7 first-author · 9 since 2021Databases, data management, data science and information retrieval · 12 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language ModelsabstractMingyang Wang, Heike Adel, Lukas Lange, Yihong Liu, Ercong Nie, Jannik Strötgen, Hinrich Schuetze. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Mingyang Wang 0003, Heike Adel, Lukas Lange, Yihong Liu 0001, Ercong Nie, Jannik Strötgen, Hinrich Schütze |
ACL (1) | 6 |
| 2025 | Language Mixing in Reasoning Language Models: Patterns, Impact, and Internal CausesabstractReasoning language models (RLMs) excel at complex tasks by leveraging a chain-of-thought process to generate structured intermediate steps.However, language mixing, i.e., reasoning steps containing tokens from languages other than the prompt, has been observed in their outputs and shown to affect performance, though its impact remains debated.We present the first systematic study of language mixing in RLMs, examining its patterns, impact, and internal causes across 15 languages, 7 task difficulty levels, and 18 subject areas, and show how all three factors influence language mixing.Moreover, we demonstrate that the choice of reasoning language significantly affects performance: forcing models to reason in Latin or Han scripts via constrained decoding notably improves accuracy.Finally, we show that the script composition of reasoning traces closely aligns with that of the model's internal representations, indicating that language mixing reflects latent processing preferences in RLMs.Our findings provide actionable insights for optimizing multilingual reasoning and open new directions for controlling reasoning languages to build more interpretable and adaptable RLMs. 1 4 This overthinking behavior is also observed in prior work such as Cuadron et al. (2025). Mingyang Wang 0003, Lukas Lange, Heike Adel, Yunpu Ma, Jannik Strötgen, Hinrich Schütze |
EMNLP | 5 |
| 2025 | Explainable Zero-Shot Visual Question Answering via Logic-Based ReasoningabstractVisual Question Answering (VQA) is the task of answering natural language questions about images, which is a challenge for AI systems. To enhance adaptability and reduce training overhead, we address VQA in a zero-shot setting by leveraging pre-trained neural modules without additional fine-tuning. Our proposed hybrid neurosymbolic framework, whose capabilities are demonstrated on the challenging GQA dataset, integrates neural and symbolic components through logic-based reasoning via Answer-Set Programming. Specifically, our pipeline employs large language models for semantic parsing of input questions, followed by the generation of a scene graph that captures relevant visual content. Interpretable rules then operate on the symbolic representations of both the question and the scene graph to derive an answer. Our framework provides a key advantage: it enables full transparency into the reasoning process. Using an existing explanation tool, we illustrate how our method fosters trust by making decisions interpretable and facilitates error analysis when predictions are incorrect. Beyond explaining its own reasoning, our framework can also explain answers from more opaque models by integrating their answers into our system, enabling broader interpretability in VQA. Thomas Eiter, Jan Hadl, Nelson Higuera, Lukas Lange, Johannes Oetsch, Bileam Scheuvens, Jannik Strötgen |
NeSy | 7 |
| 2023 | Multilingual Normalization of Temporal Expressions with Masked Language ModelsabstractThe detection and normalization of temporal expressions is an important task and preprocessing step for many applications.However, prior work on normalization is rule-based, which severely limits the applicability in realworld multilingual settings, due to the costly creation of new rules.We propose a novel neural method for normalizing temporal expressions based on masked language modeling.Our multilingual method outperforms prior rule-based systems in many languages, and in particular, for low-resource languages with performance improvements of up to 33 F 1 on average compared to the state of the art. Lukas Lange, Jannik Strötgen, Heike Adel, Dietrich Klakow |
EACL | 2 |
| 2023 | GradSim: Gradient-Based Language Grouping for Effective Multilingual TrainingabstractMost languages of the world pose low-resource challenges to natural language processing models.With multilingual training, knowledge can be shared among languages.However, not all languages positively influence each other and it is an open research question how to select the most suitable set of languages for multilingual training and avoid negative interference among languages whose characteristics or data distributions are not compatible.In this paper, we propose GradSim, a language grouping method based on gradient similarity.Our experiments on three diverse multilingual benchmark datasets show that it leads to the largest performance gains compared to other similarity measures and it is better correlated with cross-lingual model performance.As a result, we set the new state of the art on AfriSenti, a benchmark dataset for sentiment analysis on low-resource African languages.In our extensive analysis, we further reveal that besides linguistic features, the topics of the datasets play an important role for language grouping and that lower layers of transformer models encode language-specific features while higher layers capture task-specific information. Mingyang Wang 0003, Heike Adel, Lukas Lange, Jannik Strötgen, Hinrich Schütze |
EMNLP | 4 |
| 2022 | Three Real-World Datasets and Neural Computational Models for Classification Tasks in Patent LandscapingabstractPatent Landscaping, one of the central tasks of intellectual property management, includes selecting and grouping patents according to userdefined technical or application-oriented criteria.While recent transformer-based models have been shown to be effective for classifying patents into taxonomies such as CPC or IPC, there is yet little research on how to support real-world Patent Landscape Studies (PLSs) using natural language processing methods.With this paper, we release three labeled datasets for PLS-oriented classification tasks covering two diverse domains.We provide a qualitative analysis and report detailed corpus statistics.Most research on neural models for patents has been restricted to leveraging titles and abstracts.We compare strong neural and non-neural baselines, proposing a novel model that takes into account textual information from the patents' full texts as well as embeddings created based on the patents' CPC labels.We find that for PLS-oriented classification tasks, going beyond title and abstract is crucial, CPC labels are an effective source of information, and combining all features yields the best results. Subhash Chandra Pujari, Jannik Strötgen, Mark Giereth, Michael Gertz 0001, Annemarie Friedrich |
EMNLP | 2 |
| 2022 | Enhancing Knowledge Bases with Quantity FactsabstractMachine knowledge about the world’s entities should include quantity properties, such as heights of buildings, running times of athletes, energy efficiency of car models, energy production of power plants, and more. State-of-the-art knowledge bases (KBs), such as Wikidata, cover many relevant entities but often miss the corresponding quantities. Prior work on extracting quantity facts from web contents focused on high precision for top-ranked outputs, but did not tackle the KB coverage issue. This paper presents a recall-oriented approach which aims to close this gap in knowledge-base coverage. Our method is based on iterative learning for extracting quantity facts, with two novel contributions to boost recall for KB augmentation without sacrificing the quality standards of the knowledge base. The first contribution is a query expansion technique to capture a larger pool of fact candidates. The second contribution is a novel technique for harnessing observations on value distributions for self-consistency. Experiments with extractions from more than 13 million web documents demonstrate the benefits of our method. Vinh Thinh Ho, Daria Stepanova 0001, Dragan Milchevski, Jannik Strötgen, Gerhard Weikum |
WWW | 4 |
| 2022 | CLIN-X: pre-trained language models and a study on cross-task transfer for concept extraction in the clinical domainabstractMOTIVATION: The field of natural language processing (NLP) has recently seen a large change toward using pre-trained language models for solving almost any task. Despite showing great improvements in benchmark datasets for various tasks, these models often perform sub-optimal in non-standard domains like the clinical domain where a large gap between pre-training documents and target documents is observed. In this article, we aim at closing this gap with domain-specific training of the language model and we investigate its effect on a diverse set of downstream tasks and settings. RESULTS: We introduce the pre-trained CLIN-X (Clinical XLM-R) language models and show how CLIN-X outperforms other pre-trained transformer models by a large margin for 10 clinical concept extraction tasks from two languages. In addition, we demonstrate how the transformer model can be further improved with our proposed task- and language-agnostic model architecture based on ensembles over random splits and cross-sentence context. Our studies in low-resource and transfer settings reveal stable model performance despite a lack of annotated data with improvements of up to 47 F1 points when only 250 labeled sentences are available. Our results highlight the importance of specialized language models, such as CLIN-X, for concept extraction in non-standard domains, but also show that our task-agnostic model architecture is robust across the tested tasks and languages so that domain- or task-specific adaptations are not required. AVAILABILITY AND IMPLEMENTATION: The CLIN-X language models and source code for fine-tuning and transferring the model are publicly available at https://github.com/boschresearch/clin_x/ and the huggingface model hub. Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow |
Bioinform. | 3 |
| 2021 | A Multi-task Approach to Neural Multi-label Hierarchical Patent Classification Using Transformers
Subhash Chandra Pujari, Annemarie Friedrich, Jannik Strötgen |
ECIR (1) | 3 |
| 2021 | FAME: Feature-Based Adversarial Meta-Embeddings for Robust Input RepresentationsabstractCombining several embeddings typically improves performance in downstream tasks as different embeddings encode different information.It has been shown that even models using embeddings from transformers still benefit from the inclusion of standard word embeddings.However, the combination of embeddings of different types and dimensions is challenging.As an alternative to attention-based meta-embeddings, we propose feature-based adversarial meta-embeddings (FAME) with an attention function that is guided by features reflecting word-specific properties, such as shape and frequency, and show that this is beneficial to handle subword-based embeddings.In addition, FAME uses adversarial training to optimize the mappings of differently-sized embeddings to the same space.We demonstrate that FAME works effectively across languages and domains for sequence labeling and sentence classification, in particular in lowresource settings.FAME sets the new state of the art for POS tagging in 27 languages, various NER settings and question classification in different domains. Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow |
EMNLP (1) | 3 |
| 2021 | To Share or not to Share: Predicting Sets of Sources for Model Transfer LearningabstractIn low-resource settings, model transfer can help to overcome a lack of labeled data for many tasks and domains.However, predicting useful transfer sources is a challenging problem, as even the most similar sources might lead to unexpected negative transfer results.Thus, ranking methods based on task and text similarity -as suggested in prior workmay not be sufficient to identify promising sources.To tackle this problem, we propose a new approach to automatically determine which and how many sources should be exploited.For this, we study the effects of model transfer on sequence labeling across various domains and tasks and show that our methods based on model similarity and support vector machines are able to predict promising sources, resulting in performance increases of up to 24 F 1 points. Lukas Lange, Jannik Strötgen, Heike Adel, Dietrich Klakow |
EMNLP (1) | 2 |
| 2021 | A Survey on Recent Approaches for Natural Language Processing in Low-Resource ScenariosabstractMichael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, Dietrich Klakow |
NAACL-HLT | 4 |
| 2020 | Closing the Gap: Joint De-Identification and Concept Extraction in the Clinical DomainabstractExploiting natural language processing in the clinical domain requires de-identification, i.e., anonymization of personal information in texts. However, current research considers de-identification and downstream tasks, such as concept extraction, only in isolation and does not study the effects of de-identification on other tasks. In this paper, we close this gap by reporting concept extraction performance on automatically anonymized data and investigating joint models for de-identification and concept extraction. In particular, we propose a stacked model with restricted access to privacy-sensitive information and a multitask model. We set the new state of the art on benchmark datasets in English (96.1% F1 for de-identification and 88.9% F1 for concept extraction) and Spanish (91.4% F1 for concept extraction). Lukas Lange, Heike Adel, Jannik Strötgen |
ACL | 3 |
| 2020 | Fast Computation of Explanations for Inconsistency in Large-Scale Knowledge GraphsabstractKnowledge graphs (KGs) are essential resources for many applications including Web search and question answering. As KGs are often automatically constructed, they may contain incorrect facts. Detecting them is a crucial, yet extremely expensive task. Prominent solutions detect and explain inconsistency in KGs with respect to accompanying ontologies that describe the KG domain of interest. Compared to machine learning methods they are more reliable and human-interpretable but scale poorly on large KGs. In this paper, we present a novel approach to dramatically speed up the process of detecting and explaining inconsistency in large KGs by exploiting KG abstractions that capture prominent data patterns. Though much smaller, KG abstractions preserve inconsistency and their explanations. Our experiments with large KGs (e.g., DBpedia and Yago) demonstrate the feasibility of our approach and show that it significantly outperforms the popular baseline. Trung Kien Tran, Mohamed H. Gad-Elrab, Daria Stepanova 0001, Evgeny Kharlamov, Jannik Strötgen |
WWW | 5 |
| 2019 | "A Buster Keaton of Linguistics": First Automated Approaches for the Extraction of Vossian AntonomasiaabstractMichel Schwab, Robert Jäschke, Frank Fischer, Jannik Strötgen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Michel Schwab, Robert Jäschke, Frank Fischer 0005, Jannik Strötgen |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Generating Semantic Aspects for QueriesabstractLarge document collections can be hard to explore if the user presents her information need in a limited set of keywords. Ambiguous intents arising out of these short queries often result in long-winded query sessions and many query reformulations. To alleviate this problem, in this work, we propose the novel concept of semantic aspects (e.g., $${\langle }\{\textsf {michael\text {-}phelps}\}, \{\textsf {athens, beijing, london}\}, [2004,2016] \rangle $$ for the ambiguous query ) and present the xFactor algorithm that generates them from annotations in documents. Semantic aspects uplift document contents into a meaningful structured representation, thereby allowing the user to sift through many documents without the need to read their contents. The semantic aspects are created by the analysis of semantic annotations in the form of temporal, geographic, and named entity annotations. We evaluate our approach on a novel testbed of over 5,000 aspects on Web-scale document collections amounting to more than 450 million documents. Our results show the xFactor algorithm finds relevant aspects for highly ambiguous queries. Dhruv Gupta 0002, Klaus Berberich, Jannik Strötgen, Demetris Zeinalipour |
ESWC | 3 |
| 2018 | TEQUILA: Temporal Question Answering over Knowledge BasesabstractQuestion answering over knowledge bases (KB-QA) poses challenges in handling complex questions that need to be decomposed into sub-questions. An important case, addressed here, is that of temporal questions, where cues for temporal relations need to be discovered and handled. We present TEQUILA, an enabler method for temporal QA that can run on top of any KB-QA engine. TEQUILA has four stages. It detects if a question has temporal intent. It decomposes and rewrites the question into non-temporal sub-questions and temporal constraints. Answers to sub-questions are then retrieved from the underlying KB-QA engine. Finally, TEQUILA uses constraint reasoning on temporal intervals to compute final answers to the full question. Comparisons against state-of-the-art baselines show the viability of our method. Zhen Jia 0002, Abdalghani Abujabal, Rishiraj Saha Roy, Jannik Strötgen, Gerhard Weikum |
CIKM | 4 |
| 2018 | KRAUTS: A German Temporally Annotated News Corpus
Jannik Strötgen, Anne-Lyse Minard, Lukas Lange, Manuela Speranza, Bernardo Magnini |
LREC | 1 |
| 2018 | Domain-Sensitive Temporal Tagging By Jannik Strötgen, Michael Gertz (Max Planck Institute for Informatics, Heidelberg University) Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 36), 2016, xvii+133 pp; paperback, ISBN 978-1-62705-495-1; ebook, ISBN 978-1-62705-499-7; doi: 10.2200/S00721ED1V01Y201606HLT036, $55.00abstractUnderstanding time as expressed in text is an important goal of natural language understanding and extremely important for many applications, including information extraction, information retrieval, and question answering. This book provides a comprehensive overview, the challenges, available data resources, and existing systems for the task, with a special emphasis on sensitivity of temporal tagging to domains. This is a well-written book with contents well structured and organized. The discussions on temporal tagging for different domains are valuable and inspirational not only to readers interested in the specific subject of temporal tagging, but also to a broader range of readers interested in Natural Language Processing (NLP) in general. Although it is well known that domain changes often cause significant performance reductions of various NLP systems, little work has been conducted to understand and further explain specific ways such influences have been applied. To a great extent, this book fills the gap in the context of temporal tagging.This book consists of six chapters, which I will group into three parts. The first three chapters give a clear introduction to temporal tagging, including subtasks, characteristics of time, realizations of temporal expressions, data annotation standards, data sets, and evaluation metrics. This initial part provides sufficent background knowledge for further discussion of domain influences on temporal tagging. The fourth chapter is the core of the book, and it identifies four major domains—news-style, narrative-style, colloquial-style, and autonomic-style documents—and elaborates on unique characteristics of each domain and their implications on temporal tagging. The last two chapters describe a list of temporal taggers (full-fledged or focusing on one stage of temporal tagging), compare their designs (rule-based vs. learning-based) and their capabilities of addressing multiple domains and even multiple languages, and conclude with future directions of temporal tagging.Chapter 1 specifies the two subtasks of temporal tagging, temporal expression extraction, and normalization, and explains that temporal tagging can be viewed as a specific type of named entity recognition and normalization. Chapter 1 also briefly describes several temporal tagging applications, including information extraction, information retrieval, and question answering.Chapter 2 clearly describes key characteristics of time. Specifically, time can be normalized and temporal information can be organized in a hierarchical structure based on their granularities. The chapter goes on to describe four categories of temporal expressions in real text (i.e., date, time, duration, and set). I found the discussion on differences between a point in time that may have a duration and a duration of time very interesting. The subsection on realizations of temporal expressions is at the core of this chapter, and clearly defines four types of temporal expressions—explicit, implicit, relative, and underspecified expressions. It is important to distinguish between these types before we examine differences of temporal tagging across domains. Note that recognizing and normalizing each type of temporal expression requires different strategies and is at a different difficulty level. Meanwhile, the authors discuss uncertainty or fuzziness of some temporal expressions. For instance, in “He visited Germany in 2010,” it is rather unlikely that the visit took place the whole year. The exact point or period in 2010 is not known.Chapter 3 surveys annotation standards, describes several evaluation metrics, and provides a comprehensive list of research competitions and annotated news-style corpora. Although I found the description of annotation standards to be generally well thought out and organized, I occasionally felt it was difficult to understand some of the description. For instance, it is difficult to immediately understand the major differences between TIMEX2 and TIMEX3, the description of TIMEX3, and its various tags and abstract tags with no extent in TIMEX3. More examples would be helpful. Instead of sequentially reading each section in this chapter, it may help to first read Section 3.4 on annotated corpora. This chapter also includes an extensive description of different metrics used for measuring temporal tagging performance. The list of research competitions covers all the recent major efforts that were indicated by their adopted corpora, including news-style corpora (MUC, ACE, and TempEval), biomedical texts, QA TempEval (news, wiki, and blogs), and multi-language annotated corpora.Chapter 4 is the core of the book, and defines four major types of domains, examines their unique characteristics, and discusses strategies of temporal tagging for each domain. The chapter starts by describing specific characteristics of news-style documents, which is the dominant type in most studies on temporal tagging. This is followed by a general discussion of genres or domains, covering news, Wikipedia, dialogs, short messages, and clinical reports. Specifically, this book defines a domain as a group of documents that have the same characteristics relevant for the task of temporal tagging. After providing a list of annotated corpora that includes non-news texts, the chapter identifies four broad types of domains—news-style, narrative-style, colloquial-style, and autonomic-style documents—which I found fascinating, especially the fourth domain that features unresolvable time expressions due to local time frames. The discussions of unique features for each domain with respect to temporal tagging are clear and well organized, and directly lead to strategies as suggested by the authors for addressing the task in each domain. Note that each domain is defined in a broad sense and covers texts created in several scenarios. For instance, the news-style documents include “not only news articles but also many other types of documents (e.g., letters and formal blog posts), which are written similarly and thus belong to the same domain from a temporal tagging point of view.”Chapter 5 describes several widely used temporal taggers, both rule-based and learning-based taggers, and compares their temporal tagging performance on documents in different domains. Clearly, taggers prepared or trained for a particular domain do not perform well when applied to a different domain. Consistent with the authors’ vision, this chapter emphasizes that temporal taggers developed with all the domains in mind are preferable. This chapter also argues for developing highly multilingual temporal taggers.Chapter 6 concludes the book by pointing out more directions for the future work of temporal tagging, in order to achieve accurate and complete temporal understanding.To summarize, this book provides timely references to most recent advances of temporal tagging, including both annotated corpora from different domains and systems, as well as insightful discussions and vision in terms of categorizing effects of domains and designing generalized temporal taggers. This book is recommended not only to students, researchers, and developers who work in the field of temporal tagging, but also to a wide range of readers interested in NLP in general. Jannik Strötgen, Michael Gertz 0001, Graeme Hirst, Ruihong Huang |
Comput. Linguistics | 1 |
| 2016 | Credibility Assessment of Textual Claims on the WebabstractThere is an increasing amount of false claims in news, social media, and other web sources. While prior work on truth discovery has focused on the case of checking factual statements, this paper addresses the novel task of assessing the credibility of arbitrary claims made in natural-language text - in an open-domain setting without any assumptions about the structure of the claim, or the community where it is made. Our solution is based on automatically finding sources in news and social media, and feeding these into a distantly supervised classifier for assessing the credibility of a claim (i.e., true or fake). For inference, our method leverages the joint interaction between the language of articles about the claim and the reliability of the underlying web sources. Experiments with claims from the popular website snopes.com and from reported cases of Wikipedia hoaxes demonstrate the viability of our methods and their superior accuracy over various baselines. Kashyap Popat, Subhabrata Mukherjee, Jannik Strötgen, Gerhard Weikum |
CIKM | 3 |
| 2016 | GATE-Time: Extraction of Temporal Expressions and Events
Leon Derczynski, Jannik Strötgen, Diana Maynard, Mark A. Greenwood, Manuel Jung |
LREC | 2 |
| 2016 | As Time Goes By: Comprehensive Tagging of Textual Phrases with Temporal ScopesabstractTemporal expressions (TempEx's for short) are increasingly important in search, question answering, information extraction, and more. Techniques for identifying and normalizing explicit temporal expressions work well, but are not designed for and cannot cope with textual phrases that denote named events, such as "Clinton's term as secretary of state". This paper addresses the problem of detecting such temponyms, inferring their temporal scopes, and mapping them to events in a knowledge base if present there. Erdal Kuzey, Vinay Setty, Jannik Strötgen, Gerhard Weikum |
WWW | 3 |
| 2015 | A Baseline Temporal Tagger for all LanguagesabstractTemporal taggers are usually developed for a certain language. Besides English, only few languages have been addressed, and only the temporal tagger HeidelTime covers several languages. While this tool was manually extended to these lan-guages, there have been earlier approaches for automatic extensions to a single tar-get language. In this paper, we present an approach to extend HeidelTime to all lan-guages in the world. Our evaluation shows promising results, in particular consider-ing that our approach neither requires lan-guage skills nor training data, but results Jannik Strötgen, Michael Gertz 0001 |
EMNLP | 1 |
| 2015 | Time and information retrieval: Introduction to the special issue
Leon Derczynski, Jannik Strötgen, Ricardo Campos 0001, Omar Alonso |
Inf. Process. Manag. | 2 |
| 2014 | Chinese Temporal Tagging with HeidelTimeabstractTemporal information is important for many NLP tasks, and there has been extensive research on temporal tagging with a particular focus on English texts.Recently, other languages have also been addressed, e.g., HeidelTime was extended to process eight languages.Chinese temporal tagging has achieved less attention, and no Chinese temporal tagger is publicly available.In this paper, we address the full task of Chinese temporal tagging (extraction and normalization) by developing Chinese HeidelTime resources.Our evaluation on a publicly available corpus -which we also partially re-annotated due to its rather low quality -demonstrates the effectiveness of our approach, and we outperform a recent approach to normalize temporal expressions.The Chinese HeidelTime resource as well as the corrected corpus are made publicly available. Hui Li 0017, Jannik Strötgen, Julian Zell, Michael Gertz 0001 |
EACL | 2 |
| 2014 | Computational Narratology: Extracting Tense Clusters from Narrative Texts
Thomas Bögel, Jannik Strötgen, Michael Gertz 0001 |
LREC | 2 |
| 2014 | Extending HeidelTime for Temporal Expressions Referring to Historic Dates
Jannik Strötgen, Thomas Bögel, Julian Zell, Ayser Armiti, Tran Van Canh, Michael Gertz 0001 |
LREC | 1 |
| 2014 | Time for More Languages: Temporal Tagging of Arabic, Italian, Spanish, and VietnameseabstractMost of the research on temporal tagging so far is done for processing English text documents. There are hardly any multilingual temporal taggers supporting more than two languages. Recently, the temporal tagger HeidelTime has been made publicly available, supporting the integration of new languages by developing language-dependent resources without modifying the source code. In this article, we describe our work on developing such resources for two Asian and two Romance languages: Arabic, Vietnamese, Spanish, and Italian. While temporal tagging of the two Romance languages has been addressed before, there has been almost no research on Arabic and Vietnamese temporal tagging so far. Furthermore, we analyze language-dependent challenges for temporal tagging and explain the strategies we followed to address them. Our evaluation results on publicly available and newly annotated corpora demonstrate the high quality of our new resources for the four languages, which we make publicly available to the research community. Jannik Strötgen, Ayser Armiti, Tran Van Canh, Julian Zell, Michael Gertz 0001 |
ACM Trans. Asian Lang. Inf. Process. | 1 |
| 2013 | Proximity2-aware ranking for textual, temporal, and geographic queriesabstractTemporal and geographic information needs are frequent and important but not well served by standard IR systems. Recent approaches address such needs by extracting and normalizing temporal and geographic expressions from documents. They calculate specific scores for the temporal and/ or geographic parts of a query. However, all approaches assume independence between the different query parts. In this paper, we present a new model to rank documents according to combined textual, temporal, and geographic queries. The independence assumption between the query parts is eliminated by calculating proximity scores. Thus, documents are regarded to be more relevant if terms and expressions satisfying the different query parts occur close to each other in a document. As our evaluations based on the NTCIR-GeoTime data show, our proposed model outperforms baseline models that do not use proximity information. Jannik Strötgen, Michael Gertz 0001 |
CIKM | 1 |
| 2012 | Retro: Time-Based Exploration of Product Reviews
Jannik Strötgen, Omar Alonso, Michael Gertz 0001 |
ECIR | 1 |
| 2012 | Temporal Tagging on Different Domains: Challenges, Strategies, and Gold Standards
Jannik Strötgen, Michael Gertz 0001 |
LREC | 1 |
| 2011 | An event-centric model for multilingual document similarityabstractDocument similarity measures play an important role in many document retrieval and exploration tasks. Over the past decades, several models and techniques have been developed to determine a ranked list of documents similar to a given query document. Interestingly, the proposed approaches typically rely on extensions to the vector space model and are rarely suited for multilingual corpora. Jannik Strötgen, Michael Gertz 0001, Conny Junghans |
SIGIR | 1 |
| 2010 | TimeTrails: A System for Exploring Spatio-Temporal Information in DocumentsabstractSpatial and temporal data have become ubiquitous in many application domains such as the Geosciences or life sciences. Sophisticated database management systems are employed to manage such structured data. However, an important source of spatio-temporal information that has not been fully utilized are unstructured text documents. In documents, combinations of temporal and spatial expressions form events, which can be mapped to a database structure and organized into trajectories that can be explored. In this context, the coupling of information retrieval techniques with spatio-temporal database concepts leads to new ways for managing and exploring document collections. In this demonstration, we present TimeTrails, a system for the extraction, querying, storage, and exploration of spatio-temporal information embedded in text documents. The user can query a document collection, and TimeTrails visualizes the spatio-temporal information extracted from relevant documents as document trajectories, resulting in a map-based view of documents. This view helps the user to explore the temporal and spatial content of documents in a meaningful way and to further restrict search results using spatial and temporal predicates. Jannik Strötgen, Michael Gertz 0001 |
Proc. VLDB Endow. | 1 |