EDBT 2026 Demo / reviewers in the wild / expert
Xuke Hu
dblp:183/5437
· DBLP profile ↗
10ranked-venue papers in the field
9as first author
10since 2021 · last 2026
0000-0002-5649-0243ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 5 (4 first)Information Retrieval & Web Search · 5 (5 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 4th International Workshop on Geographic Information Extraction from Texts (GeoExT 2026)
Xuke Hu, Ludovic Moncla, Jens Kersten, Anna M. Kruspe, Inhye Kong |
ECIR (3) | 1 |
| 2026 | Toponym resolution leveraging lightweight and open-source large language models and geo-knowledgeabstractToponym resolution is crucial for extracting geographic information from natural language texts, such as social media posts and news articles. Despite the advancements in current methods, including state-of-the-art deep learning solutions like GENRE and a sophisticated voting system that integrates seven individual methods, further enhancing their accuracy is essential. To achieve this goal, we propose a novel method that combines lightweight and open-source large language models and geo-knowledge. Specifically, we first fine-tune Mistral (7B), Baichuan2 (7B), Llama2 (7B & 13B), and Falcon (7B) to estimate toponyms’ unambiguous reference (e.g., city, state, country) given their contexts. Subsequently, we correct inaccuracies in generated references and determine their geo-coordinates via sequentially querying GeoNames, Nominatim, and ArcGIS geocoders until a successful geocoding result is achieved. Our methods demonstrate enhanced performance compared to 20 existing methods, as evidenced across seven challenging datasets including 83,365 toponyms worldwide, with the Mistral-based method leading, followed by Baichuan2, Llama2, and Falcon-based methods. Specifically, the Mistral-based method achieves an Accuracy@161km of 0.91, surpassing GENRE, the best individual method, by 17% and the seven-methods composite voting system by 7%. Moreover, our methods are computationally efficient, operable on one general GPU, have modest memory requirements (14 GB for 7B models and 27 GB for 13B models), and exceed both GENRE and the voting system in inferring speed. Xuke Hu, Jens Kersten, Friederike Klan, Sheikh Mastura Farzana |
Int. J. Geogr. Inf. Sci. | 1 |
| 2026 | Extracting and analysing geographic information from natural language textsabstract1. Traditionally, geographic information used in spatial analysis and research was almost exclusively produced by a relatively small set of governmental and commercial actors, and took the form of ... Xuke Hu, Ross Purves, Ludovic Moncla, Jens Kersten, Kristin Stock |
Int. J. Geogr. Inf. Sci. | 1 |
| 2025 | 3rd International Workshop on Geographic Information Extraction from Texts (GeoExT 2025)
Xuke Hu, Ross Purves, Ludovic Moncla, Jens Kersten, Anna M. Kruspe |
ECIR (5) | 1 |
| 2024 | 2nd International Workshop on Geographic Information Extraction from Texts (GeoExT 2024)
Xuke Hu, Ross Purves, Ludovic Moncla, Jens Kersten, Kristin Stock |
ECIR (5) | 1 |
| 2024 | DLRGeoTweet: A comprehensive social media geocoding corpus featuring fine-grained placesabstractEvery day, many short text messages on social media are generated in response to real-world events, providing a valuable resource for various domains such as emergency response and traffic management. Since exact coordinates of social media posts are rarely attached by users, accurately recognizing and resolving fine-grained place names, such as home addresses and Points of Interest, from these posts is crucial for understanding the precise locations of critical events, such as rescue requests. This task, known as geoparsing, involves toponym recognition and toponym resolution or geocoding. However, existing social media datasets for evaluating geoparsing approaches often lack sufficient fine-grained place names with associated geo-coordinates or linked to gazetteers, making evaluating, comparing, and training geocoding methods for such locations challenging. Moreover, the absence of supportive annotation tools compounds this challenge. To address these gaps, we implemented a lightweight Python tool leveraging Nominatim. Using this tool, we annotated a comprehensive X (formerly Twitter) geocoding corpus called DLRGeoTweet. The corpus underwent a rigorous cross-validation process to guarantee its quality. This corpus includes a total of 7,364 tweets and 12,510 places, of which 6,012 are fine-grained. It comprises two global datasets encompassing worldwide events and three local datasets related to local events such as the 2017 Hurricane Harvey. The annotation process spanned over ten months and required approximately 1000 person-hours to complete. We then evaluate 15 latest and representative geocoding approaches, including many deep learning-based, on DLRGeoTweet. The results highlight the inherent challenges in resolving fine-grained places accurately. Despite increasing access constraints to Twitter data, our corpus’s focus on short, informal text makes it a valuable resource for geocoding across multiple social media platforms. Xuke Hu, Tobias Elßner, Shiyu Zheng, Helen Ngonidzashe Serere, Jens Kersten, Friederike Klan, Qinjun Qiu |
Inf. Process. Manag. | 1 |
| 2023 | Geographic Information Extraction from Texts (GeoExT)
Xuke Hu, Yingjie Hu 0001, Bernd Resch, Jens Kersten |
ECIR (3) | 1 |
| 2022 | GazPNE: annotation-free deep learning for place name extraction from microblogs leveraging gazetteer and synthetic data by rulesabstractExtracting precise location information from microblogs is a crucial task in many applications, particularly in disaster response, revealing where damages are, where people need assistance, and where help can be found. A crucial prerequisite to location extraction is place name extraction. In this paper, we present GazPNE: a hybrid approach to place name extraction which fuses rules, gazetteers, and deep learning techniques without requiring any manually annotated data. The core of the approach is to learn the intrinsic characteristics of multi-word place names with deep learning from gazetteers. Specifically, GazPNE consists of a rule-based system to select n-grams from the microblogs that potentially contain place names, and a C-LSTM model that decides if the selected n-gram is a place name or not. The C-LSTM is trained on 388.1 million examples containing 6.8 million positive examples with US and Indian place names extracted from OpenStreetMap and 381.3 million negative examples synthesized by rules. We evaluate GazPNE against the SoTA on a manually annotated 4,500 tweet dataset which contains 9,026 place names from three foods: 2016 in Louisiana (US), 2016 in Houston (US), and 2015 in Chennai (India). GazPNE achieves SotA performance on the test data with an F1 of 0.84. Xuke Hu, Hussein Al-Olimat, Jens Kersten, Matti Wiegmann, Friederike Klan, Yeran Sun, Hongchao Fan |
Int. J. Geogr. Inf. Sci. | 1 |
| 2021 | Tagging the main entrances of public buildings based on OpenStreetMap and binary imbalanced learningabstractDetermining the location of a building’s entrance is crucial to location-based services, such as wayfinding for pedestrians. Unfortunately, entrance information is often missing from current mainstream map providers such as Google Maps. Frequently, automatic approaches for detecting building entrances are based on street-level images that are not widely available. To address this issue, we propose a more general approach for inferring the main entrances of public buildings based on the association between spatial elements extracted from OpenStreetMap. In particular, we adopt three binary classification approaches, weighted random forest, balanced random forest, and smooth-boost to model the association relationship. There are two types of features considered in the classification: intrinsic features derived from building footprints and extrinsic features derived from spatial contexts, such as roads, green spaces, bicycle parking areas, and neighboring buildings. We conducted extensive experiments on 320 public buildings with an average perimeter of 350 m. The experimental results showed that the locations of building entrances estimated by the weighted random forest and balanced random forest models have a mean linear distance error of 21 m and a mean path distance error of 22 m, ruling out 90% of the incorrect locations of the main entrance of buildings. Xuke Hu, Alexey Noskov, Hongchao Fan, Tessio Novack, Hao Li 0019, Fuqiang Gu, Jianga Shang, Alexander Zipf |
Int. J. Geogr. Inf. Sci. | 1 |
| 2021 | Indoor mapping and modeling by parsing floor plan imagesabstractA large proportion of indoor spatial data is generated by parsing floor plans. However, a mature and automatic solution for generating high-quality building elements (e.g., walls and doors) and space partitions (e.g., rooms) is still lacking. In this study, we present a two-stage approach to indoor mapping and modeling (IMM) from floor plan images. The first stage vectorizes the building elements on the floor plan images and the second stage repairs the topological inconsistencies between the building elements, separates indoor spaces, and generates indoor maps and models. To reduce the shape complexity of indoor boundary elements, i.e., walls and openings, we harness the regularity of the boundary elements and extract them as rectangles in the first stage. Furthermore, to resolve the overlaps and gaps of the vectorized results, we propose an optimization model that adjusts the rectangle vertex coordinates to conform to the topological constraints. Experiments demonstrate that our approach achieves a considerable improvement in room detection without conforming to Manhattan World Assumption. Our approach also outputs instance-separate walls with consistent topology, which enables direct modeling into Industry Foundation Classes (IFC) or City Geography Markup Language (CityGML). Jianga Shang, Pan Chen 0004, Sisi Zlatanova, Xuke Hu, Zhiyong Zhou 0005 |
Int. J. Geogr. Inf. Sci. | 5 |