VLDB 2026 Research / reviewers in the wild / expert
Jens Kersten
dblp:142/1638
· DBLP profile ↗
9ranked-venue papers in the field
0as first author
8since 2021 · last 2026
0000-0002-4735-7360ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5Database Systems & Data Management · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 4th International Workshop on Geographic Information Extraction from Texts (GeoExT 2026)
Xuke Hu, Ludovic Moncla, Jens Kersten, Anna M. Kruspe, Inhye Kong |
ECIR (3) | 3 |
| 2026 | Toponym resolution leveraging lightweight and open-source large language models and geo-knowledgeabstractToponym resolution is crucial for extracting geographic information from natural language texts, such as social media posts and news articles. Despite the advancements in current methods, including state-of-the-art deep learning solutions like GENRE and a sophisticated voting system that integrates seven individual methods, further enhancing their accuracy is essential. To achieve this goal, we propose a novel method that combines lightweight and open-source large language models and geo-knowledge. Specifically, we first fine-tune Mistral (7B), Baichuan2 (7B), Llama2 (7B & 13B), and Falcon (7B) to estimate toponyms’ unambiguous reference (e.g., city, state, country) given their contexts. Subsequently, we correct inaccuracies in generated references and determine their geo-coordinates via sequentially querying GeoNames, Nominatim, and ArcGIS geocoders until a successful geocoding result is achieved. Our methods demonstrate enhanced performance compared to 20 existing methods, as evidenced across seven challenging datasets including 83,365 toponyms worldwide, with the Mistral-based method leading, followed by Baichuan2, Llama2, and Falcon-based methods. Specifically, the Mistral-based method achieves an Accuracy@161km of 0.91, surpassing GENRE, the best individual method, by 17% and the seven-methods composite voting system by 7%. Moreover, our methods are computationally efficient, operable on one general GPU, have modest memory requirements (14 GB for 7B models and 27 GB for 13B models), and exceed both GENRE and the voting system in inferring speed. Xuke Hu, Jens Kersten, Friederike Klan, Sheikh Mastura Farzana |
Int. J. Geogr. Inf. Sci. | 2 |
| 2026 | Extracting and analysing geographic information from natural language textsabstract1. Traditionally, geographic information used in spatial analysis and research was almost exclusively produced by a relatively small set of governmental and commercial actors, and took the form of ... Xuke Hu, Ross Purves, Ludovic Moncla, Jens Kersten, Kristin Stock |
Int. J. Geogr. Inf. Sci. | 4 |
| 2025 | 3rd International Workshop on Geographic Information Extraction from Texts (GeoExT 2025)
Xuke Hu, Ross Purves, Ludovic Moncla, Jens Kersten, Anna M. Kruspe |
ECIR (5) | 4 |
| 2024 | 2nd International Workshop on Geographic Information Extraction from Texts (GeoExT 2024)
Xuke Hu, Ross Purves, Ludovic Moncla, Jens Kersten, Kristin Stock |
ECIR (5) | 4 |
| 2024 | DLRGeoTweet: A comprehensive social media geocoding corpus featuring fine-grained placesabstractEvery day, many short text messages on social media are generated in response to real-world events, providing a valuable resource for various domains such as emergency response and traffic management. Since exact coordinates of social media posts are rarely attached by users, accurately recognizing and resolving fine-grained place names, such as home addresses and Points of Interest, from these posts is crucial for understanding the precise locations of critical events, such as rescue requests. This task, known as geoparsing, involves toponym recognition and toponym resolution or geocoding. However, existing social media datasets for evaluating geoparsing approaches often lack sufficient fine-grained place names with associated geo-coordinates or linked to gazetteers, making evaluating, comparing, and training geocoding methods for such locations challenging. Moreover, the absence of supportive annotation tools compounds this challenge. To address these gaps, we implemented a lightweight Python tool leveraging Nominatim. Using this tool, we annotated a comprehensive X (formerly Twitter) geocoding corpus called DLRGeoTweet. The corpus underwent a rigorous cross-validation process to guarantee its quality. This corpus includes a total of 7,364 tweets and 12,510 places, of which 6,012 are fine-grained. It comprises two global datasets encompassing worldwide events and three local datasets related to local events such as the 2017 Hurricane Harvey. The annotation process spanned over ten months and required approximately 1000 person-hours to complete. We then evaluate 15 latest and representative geocoding approaches, including many deep learning-based, on DLRGeoTweet. The results highlight the inherent challenges in resolving fine-grained places accurately. Despite increasing access constraints to Twitter data, our corpus’s focus on short, informal text makes it a valuable resource for geocoding across multiple social media platforms. Xuke Hu, Tobias Elßner, Shiyu Zheng, Helen Ngonidzashe Serere, Jens Kersten, Friederike Klan, Qinjun Qiu |
Inf. Process. Manag. | 5 |
| 2023 | Geographic Information Extraction from Texts (GeoExT)
Xuke Hu, Yingjie Hu 0001, Bernd Resch, Jens Kersten |
ECIR (3) | 4 |
| 2022 | GazPNE: annotation-free deep learning for place name extraction from microblogs leveraging gazetteer and synthetic data by rulesabstractExtracting precise location information from microblogs is a crucial task in many applications, particularly in disaster response, revealing where damages are, where people need assistance, and where help can be found. A crucial prerequisite to location extraction is place name extraction. In this paper, we present GazPNE: a hybrid approach to place name extraction which fuses rules, gazetteers, and deep learning techniques without requiring any manually annotated data. The core of the approach is to learn the intrinsic characteristics of multi-word place names with deep learning from gazetteers. Specifically, GazPNE consists of a rule-based system to select n-grams from the microblogs that potentially contain place names, and a C-LSTM model that decides if the selected n-gram is a place name or not. The C-LSTM is trained on 388.1 million examples containing 6.8 million positive examples with US and Indian place names extracted from OpenStreetMap and 381.3 million negative examples synthesized by rules. We evaluate GazPNE against the SoTA on a manually annotated 4,500 tweet dataset which contains 9,026 place names from three foods: 2016 in Louisiana (US), 2016 in Houston (US), and 2015 in Chennai (India). GazPNE achieves SotA performance on the test data with an F1 of 0.84. Xuke Hu, Hussein Al-Olimat, Jens Kersten, Matti Wiegmann, Friederike Klan, Yeran Sun, Hongchao Fan |
Int. J. Geogr. Inf. Sci. | 3 |
| 2014 | Airborne near-real-time monitoring of assembly and parking areas in case of large-scale public events and natural disastersabstractA critical requirement for an effective and coordinated response by public entities tasked with management, security, and relief during large-scale public events or natural disasters is the availability of current situational information. However, today there is a lack of comprehensive operational systems allowing a near-real-time (NRT) collection, visualization, and provision of situational information for larger areas. In this study a methodological framework is proposed, which allows an NRT extraction and visualization of situational information based on aerial image acquisition. The framework combines digital image analysis using a generic supervised information extraction approach based on statistical modeling with a downstream web-based visualization component realized through an automatic update of web services. Even though being applicable for different scenarios, the workflow will be demonstrated for the specific use-case of a NRT monitoring of open spaces including assembly and parking areas. Compared to other approaches, image analysis results indicate a high robustness and a low demand for computational power sources (7 seconds per image). Due to a high degree of automation, the proposed workflow contributes to a NRT ‘end-to-end’ monitoring system, which was developed within the VABENE (German acronym for ‘traffic management under large-scale public events and disaster conditions’) project covering all parts from the acquisition of raw aerial imagery to the dissemination of information products to end-users. Hannes Römer, Jens Kersten, Ralph Kiefl, Stefan Plattner, Alexander Mager, Stefan Voigt |
Int. J. Geogr. Inf. Sci. | 2 |