VLDB 2026 Research / reviewers in the wild / expert
Friederike Klan
dblp:k/FriederikeKlan
· DBLP profile ↗
7ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0002-1856-7334ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toponym resolution leveraging lightweight and open-source large language models and geo-knowledgeabstractToponym resolution is crucial for extracting geographic information from natural language texts, such as social media posts and news articles. Despite the advancements in current methods, including state-of-the-art deep learning solutions like GENRE and a sophisticated voting system that integrates seven individual methods, further enhancing their accuracy is essential. To achieve this goal, we propose a novel method that combines lightweight and open-source large language models and geo-knowledge. Specifically, we first fine-tune Mistral (7B), Baichuan2 (7B), Llama2 (7B & 13B), and Falcon (7B) to estimate toponyms’ unambiguous reference (e.g., city, state, country) given their contexts. Subsequently, we correct inaccuracies in generated references and determine their geo-coordinates via sequentially querying GeoNames, Nominatim, and ArcGIS geocoders until a successful geocoding result is achieved. Our methods demonstrate enhanced performance compared to 20 existing methods, as evidenced across seven challenging datasets including 83,365 toponyms worldwide, with the Mistral-based method leading, followed by Baichuan2, Llama2, and Falcon-based methods. Specifically, the Mistral-based method achieves an Accuracy@161km of 0.91, surpassing GENRE, the best individual method, by 17% and the seven-methods composite voting system by 7%. Moreover, our methods are computationally efficient, operable on one general GPU, have modest memory requirements (14 GB for 7B models and 27 GB for 13B models), and exceed both GENRE and the voting system in inferring speed. Xuke Hu, Jens Kersten, Friederike Klan, Sheikh Mastura Farzana |
Int. J. Geogr. Inf. Sci. | 3 |
| 2024 | DLRGeoTweet: A comprehensive social media geocoding corpus featuring fine-grained placesabstractEvery day, many short text messages on social media are generated in response to real-world events, providing a valuable resource for various domains such as emergency response and traffic management. Since exact coordinates of social media posts are rarely attached by users, accurately recognizing and resolving fine-grained place names, such as home addresses and Points of Interest, from these posts is crucial for understanding the precise locations of critical events, such as rescue requests. This task, known as geoparsing, involves toponym recognition and toponym resolution or geocoding. However, existing social media datasets for evaluating geoparsing approaches often lack sufficient fine-grained place names with associated geo-coordinates or linked to gazetteers, making evaluating, comparing, and training geocoding methods for such locations challenging. Moreover, the absence of supportive annotation tools compounds this challenge. To address these gaps, we implemented a lightweight Python tool leveraging Nominatim. Using this tool, we annotated a comprehensive X (formerly Twitter) geocoding corpus called DLRGeoTweet. The corpus underwent a rigorous cross-validation process to guarantee its quality. This corpus includes a total of 7,364 tweets and 12,510 places, of which 6,012 are fine-grained. It comprises two global datasets encompassing worldwide events and three local datasets related to local events such as the 2017 Hurricane Harvey. The annotation process spanned over ten months and required approximately 1000 person-hours to complete. We then evaluate 15 latest and representative geocoding approaches, including many deep learning-based, on DLRGeoTweet. The results highlight the inherent challenges in resolving fine-grained places accurately. Despite increasing access constraints to Twitter data, our corpus’s focus on short, informal text makes it a valuable resource for geocoding across multiple social media platforms. Xuke Hu, Tobias Elßner, Shiyu Zheng, Helen Ngonidzashe Serere, Jens Kersten, Friederike Klan, Qinjun Qiu |
Inf. Process. Manag. | 6 |
| 2022 | GazPNE: annotation-free deep learning for place name extraction from microblogs leveraging gazetteer and synthetic data by rulesabstractExtracting precise location information from microblogs is a crucial task in many applications, particularly in disaster response, revealing where damages are, where people need assistance, and where help can be found. A crucial prerequisite to location extraction is place name extraction. In this paper, we present GazPNE: a hybrid approach to place name extraction which fuses rules, gazetteers, and deep learning techniques without requiring any manually annotated data. The core of the approach is to learn the intrinsic characteristics of multi-word place names with deep learning from gazetteers. Specifically, GazPNE consists of a rule-based system to select n-grams from the microblogs that potentially contain place names, and a C-LSTM model that decides if the selected n-gram is a place name or not. The C-LSTM is trained on 388.1 million examples containing 6.8 million positive examples with US and Indian place names extracted from OpenStreetMap and 381.3 million negative examples synthesized by rules. We evaluate GazPNE against the SoTA on a manually annotated 4,500 tweet dataset which contains 9,026 place names from three foods: 2016 in Louisiana (US), 2016 in Houston (US), and 2015 in Chennai (India). GazPNE achieves SotA performance on the test data with an F1 of 0.84. Xuke Hu, Hussein Al-Olimat, Jens Kersten, Matti Wiegmann, Friederike Klan, Yeran Sun, Hongchao Fan |
Int. J. Geogr. Inf. Sci. | 5 |
| 2022 | GazPNE2: A General Place Name Extractor for Microblogs Fusing Gazetteers and Pretrained Transformer ModelsabstractThe concept of “human as sensors” defines a new sensing model, in which humans act as sensors by contributing their observations, perceptions, and sensations. This is crucial for the development of Social Internet of Things, which is an integral part of cyber-physical–social systems. Online social media platforms, as the most active places where users act as social sensors, are responsive to real-world events and are useful for gathering situational information in real time. Unfortunately, posts rarely contain structured geographic information, thus hindering their usage for contributing to various challenges, such as emergency response. We address this limitation by introducing a general approach for extracting place names from tweets, named GazPNE2. It combines global gazetteers (i.e., OpenStreetMap and GeoNames), deep learning, and pretrained transformer models (i.e., BERT and BERTweet), which requires no manually annotated data. It can extract place names at both coarse (e.g., city) and fine-grained (e.g., street and POI) levels and place names with abbreviations. To fully evaluate GazPNE2 and compare it with 11 competing approaches, we use 19 public tweet data sets, containing 38 802 tweets and 22 197 places across the world. The results show GazPNE2 achieves a much higher F1 (0.8) than the other approaches. Furthermore, we apply GazPNE2 to three large unannotated tweet data sets related to over 20 crisis events (e.g., coronavirus disease 2019), containing 560 040 tweets. An F1 of 0.84 is achieved on 3000 tweets, which are randomly selected from the three data sets and then manually annotated. Code and data are available on GitHub page:https://github.com/uhuohuy/GazPNE2. Xuke Hu, Zhiyong Zhou 0005, Yeran Sun, Jens Kersten, Friederike Klan, Hongchao Fan, Matti Wiegmann |
IEEE Internet Things J. | 5 |
| 2016 | OAPT: A Tool for Ontology Analysis and Partitioning
Alsayed Algergawy, Samira Babalou, Friederike Klan, Birgitta König-Ries |
EDBT | 3 |
| 2015 | A sequence-based tree similarity searchabstractTree-structured data are pervasively growing and exploiting them based on similarity is essential for a broad number of applications. Therefore, there has been a growing need to develop high-performance techniques to efficiently look for similar trees across a large number of trees. To this end, in this paper, we present a new sequence-based approach for tree similarity search that exploits both the structural and the content characteristics of tree-structured data. In particular, we transform tree data into sequence representations using a modified Prüfer sequence that constructs a one-to-one mapping between tree data and their sequence representations. We introduce a new tree sequence distance based on the structural information of the data tree, which filters out a set of false positive candidates. We then introduce a refinement step exploiting the content information of data trees. The preliminarily experimental results show that our algorithm achieves high performance. Our method is especially suitable for accelerating similarity computation in clustering and/or classification of large numbers of trees in massive datasets. Alsayed Algergawy, Friederike Klan |
RCIS | 2 |
| 2010 | Enabling trust-aware semantic web service selection a flexible and personalized approachabstractIn today's online markets, consumers need support in finding providers that offer the products or services they need and that are trustworthy. While Semantic Web Services (SWS) research addresses the first problem (discovering functionally suitable service providers), it neglects the second. Hence, several attempts have been made to complement service retrieval techniques based on semantic matchmaking with trust-establishing techniques that leverage collaborative consumer feedback. However, the diversity and multi-faceted nature of SWS impose special requirements on the underlying feedback mechanism, in particular w.r.t. their flexibility and expressiveness. Existing approaches only partially meet those requirements. In this paper, we will therefore propose a trust-establishing mechanism for Semantic Web Services that allows to assess a service provider's trustworthiness with respect to various service aspects and is flexible enough to adjust to various kinds of services and consumer requirements. Friederike Klan, Birgitta König-Ries |
iiWAS | 1 |