VLDB 2026 Research / reviewers in the wild / expert
Martin Werner 0001
dblp:87/2697-1
· DBLP profile ↗
18ranked-venue papers in the field
5as first author
14since 2021 · last 2026
0000-0002-6951-8022ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 16 (5 first)Data Mining & Knowledge Discovery · 1Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TrajGen: Demonstrating a Tool for Interactive Trajectory Generation in the Browser
Paul M. Walther, Xuanshu Luo, Balthasar Teuscher, Martin Werner 0001 |
MDM | 4 |
| 2026 | TrajGen: Approaches to the Artificial Generation of Trajectory Datasets
Paul M. Walther, Balthasar Teuscher, Xuanshu Luo, Martin Werner 0001 |
MDM | 4 |
| 2026 | Learning Bloom Filters: A Review
Paul M. Walther, Martin Werner 0001 |
PAKDD (4) | 2 |
| 2026 | Triple-objective cross-view geolocalization of disaster-related VGI: the case of Hurricane IanabstractVolunteered geographic information (VGI) often contains rich geolocations that are crucial for disaster response and post-disaster assessment. However, existing studies on VGI geolocalization have not fully used the potential of multi-source and multimodal data. In this paper, we constructed a multimodal disaster dataset (MultiIan) and developed two novel methods (i.e. StaGeo and TriGeo) to enhance the cross-view geolocalization accuracy of disaster-related VGI. MultiIan comprised VGI texts and images, street view imagery (SVI) and remote sensing imagery (RSI). Large language models (LLMs) were used to extract the implicit geoinformation from VGI texts for geotagging. StaGeo was developed using staged training with ConvNeXt and vision transformer (ViT), while TriGeo used VGI ↔ SVI ↔ RSI triple-objective joint training of the ViT based on DINOv2. Using SVI to link VGI and RSI, our methods significantly improved the geolocalization accuracy of VGI across various train–test splits in MultiIan. With a typical 8:2 data split, StaGeo achieved Recall@1, Recall@5, Recall@10 and Recall@1% of 54.93%, 71.27%, 77.93% and 80.33%, respectively. TriGeo further improved these metrics, achieving 62.87%, 85.55%, 90.54% and 90.89%, respectively. These findings demonstrate significant advancements in our cross-view geolocalization methods, enabling timely geolocation to support rapid decision-making in emergency response and promoting the broader application of GeoAI in geospatial analysis. Wenping Yin, Fabian Deuser, Xuanshu Luo, Martin Werner 0001, Hao Li 0019, Yong Xue |
Int. J. Geogr. Inf. Sci. | 6 |
| 2025 | Entropy-Driven Curriculum for Multi-Task Training in Human Mobility Prediction
Tianye Fang, Xuanshu Luo, Martin Werner 0001 |
IEEE Big Data | 3 |
| 2025 | Enhancing Contrastive Learning for Geolocalization by Discovering Hard Negatives on SemivariogramsabstractAccurate and robust image-based geo-localization at a global scale is challenging due to diverse environments, visually ambiguous scenes, and the lack of distinctive landmarks in many regions. While contrastive learning methods show promising performance by aligning features between street-view images and corresponding locations, they neglect the underlying spatial dependency in the geographic space. As a result, they fail to address the issue of false negatives - image pairs that are both visually and geographically similar but labeled as negatives, and struggle to effectively distinguish hard negatives, which are visually similar but geographically distant. To address this issue, we propose a novel spatially regularized contrastive learning strategy that integrates a semivariogram, which is a geostatistical tool for modeling how spatial correlation changes with distance. We fit the semivariogram by relating the distance of images in feature space to their geographical distance, capturing the expected visual content in a spatial correlation. With the fitted semivariogram, we define the expected visual dissimilarity at a given spatial distance as reference to identify hard negatives and false negatives. We integrate this strategy into GeoCLIP and evaluate it on the OSV5M dataset, demonstrating that explicitly modeling spatial priors improves image-based geo-localization performance, particularly at finer granularity. Boyi Chen, Fabian Deuser, Johann Maximilian Zollner, Martin Werner 0001 |
SIGSPATIAL/GIS | 5 |
| 2025 | Human Mobility Prediction with Multi-Task Curriculum TrainingabstractEffective human mobility modeling and prediction constitute the core prerequisites for various location-based applications. To encourage research in this direction, the ACM SIGSPATIAL Cup 2025 posed the challenge of predicting human mobility trajectories from a sparse multi-city dataset. This paper presents our solution, MoBERT, a BERT-like model that adapts and leverages mobility semantics with additional direction and distance prediction, providing supplementary supervision signals for robust feature learning. MoBERT models are trained in stages through curriculum learning, where augmented trajectories are ordered by increasing mobility entropy for training with progressively increasing difficulty. The final score of our method is 0.14609, as measured by average GEO-BLEU distances across four cities. Finally, we analyze the results and discuss insights from our approach. Tianye Fang, Xuanshu Luo, Paul M. Walther, Martin Werner 0001 |
SIGSPATIAL/GIS | 4 |
| 2025 | SHM-DB - A Shared Memory Database for FAIR Geospatial Algorithm DevelopmentabstractThis paper introduces the shared memory database SHM-DB and demonstrates how this can be used for developing point cloud processing algorithms on comparably large datasets with improved software quality. The SHM-DB acts as a component that completely disentangles all aspects of the computational process such that individual aspects can be implemented in a very reusable form. As a side-effect, the approach introduces a novel dimension of algorithmic transparency to the development system as the system visualizes intermediate results while an algorithm is running. Despite demonstrating the system on point cloud data, the system is generic and can be used beyond point clouds for other types of challenging data. Martin Werner 0001, Katharina Anders |
SIGSPATIAL/GIS | 1 |
| 2024 | Calculating Upstream Relation in Spatial Networks Under Path ConstraintsabstractAssessing the resilience of utility networks and ensuring human safety in indoor environments during emergencies relies on the identification of the upstream relation in a spatial network. This concept is essential for detecting vulnerabilities and maintaining operational continuity in the event of unforeseen disruptions by finding alternative paths under various constraints and accessibility restrictions. As spatial networks grow increasingly complex, it is no longer practical to comprehensively analyze all paths within a spatial semantic graph, e.g., in applications related to geographic information systems and navigation services. This paper presents a novel computational method-different from the conventional biconnected components-for analyzing spatial networks independent of any number of constraints. The defined approach employs a fast graph search algorithm based on shortest path queries with increasing graph weights. It integrates the constrained shortest path search with weighted edges and a parallelized linear search scalable to large networks. Unlike Block-Cut trees, our algorithm efficiently determines the upstream relation with directly embedded path constraints by identifying all simple paths that fulfill the specified constraints without entirely reconstructing the graph. We demonstrate our solution's efficiency across two infrastructure networks: a critical utility network and an indoor navigation map with different path constraints. The source code is available at https://github.com/tum-bgd/2024-sigspatial-upstream Wejdene Mansour, Martin Werner 0001 |
SIGSPATIAL/GIS | 2 |
| 2023 | Bavaria Buildings - A Novel Dataset for Building Footprint Extraction, Instance Segmentation, and Data Quality EstimationabstractBavaria Buildings is a large, analysis-ready dataset providing openly available co-registered 40cm aerial imagery of Upper Bavaria paired with building footprint information. The Bavaria Buildings dataset (BBD) contains 18205 orthophotos of 2500 × 2500 pixels, where each pixel covers 40cm × 40cm in space (Digitales Orthophoto 40cm - DOP40). The dataset has been pre-processed and co-registered and also provides a set of 5.5 million image tiles of 250 × 250 pixels ready for deep learning and image analysis tasks. For each image tile, we provide two segmentation masks; one based on the official building footprints (Hausumringe) data as published by the Free State of Bavaria and one based on a historic OpenStreetMap (OSM) extract dating to 2021. The dataset is ready for essential analysis tasks, such as detection, segmentation, instance extraction, footprint geometry extraction, multimodal localization, and multimodal data quality assessment of buildings in Bavaria. We plan to update the dataset with each major re-publication of the upstream data sources to foster change detection research in the future. The BBD is available at https://doi.org/10.14459/2023mp1709451. Martin Werner 0001, Hao Li 0019, Johann Maximilian Zollner, Balthasar Teuscher, Fabian Deuser |
SIGSPATIAL/GIS | 1 |
| 2023 | Rethink Geographical Generalizability with Unsupervised Self-Attention Model Ensemble: A Case Study of OpenStreetMap Missing Building Detection in AfricaabstractThe recent advance of adapting pre-trained task-agnostic artificial intelligence (AI) models leads to great successes in downstream tasks via fine-tuning, or low-resource (i.e., few-shot and zero-shot) learning. However, when adapting such pre-trained AI models to geographical applications, it is still challenging to find the "sweet spot" of the model's generalizability and specializability (e.g., geographic generalizability v.s. spatial heterogeneity). For instance, a building detection task may require vision models with different parameters across different geographic areas of the world. In this paper, we rethink this interesting topic, namely Geographical Generalizability of GeoAI models, with a case study of detecting OpenStreetMap (OSM) missing buildings across different countries in sub-Saharan Africa. We consider a real-world scenario, in which we first train a Single-Shot Multibox Detection (SSD) base model for OSM missing building detection in Kakola, Tanzania, where a previous humanitarian mapping project of OSM was organized to map all possible buildings. Then we extrapolate this base model using Few-Shot Transfer Learning (FSTL) to a set of areas in the proximity of the test area in Cameroon. Here, we develop a Geographical Weighted Model Ensemble (GWME) method to improve Geographical Generalizability of GeoAI models. Moreover, we compare four unsupervised model ensemble weighting strategies: 1) Average weighting, 2) Image similarity weighting, 3) Geographical distance weighting, and 4) Self-attention-based weighting. Experiments show promising results of the proposed GWME method, which implicitly generates model weights from their location embedding and image feature embedding in an unsupervised manner. More specifically, the self-attention-based model ensemble achieves the highest performance. The results shed inspiring light on improving the generalizability and replicability of GeoAI models across geographic areas. Data and code are available at https://github.com/tum-bgd/GWME. Hao Li 0019, Jiapan Wang, Johann Maximilian Zollner, Gengchen Mai, Ni Lao, Martin Werner 0001 |
SIGSPATIAL/GIS | 6 |
| 2023 | Exploring GeoAI Methods for Supraglacial Lake Mapping on Greenland Ice SheetabstractThe ACM SIGSPATIAL Cup 2023 proposed the challenge to identify and map supraglacial lakes in Greenland in satellite imagery. The peculiarities of supraglacial lakes pose a hard problem for semantic segmentation and object detection tasks because the definition of a lake is ill-fitted to the inner workings of such approaches. For example, lakes are often covered by ice and snow and narrow streams can connect distinct lakes, which is not directly translatable to the semantic segmentation of water. It is also not well-posed for object detection, especially the identity relation - what is a lake, what is not (yet) a lake, and what are two lakes is challenging. In this context, we worked on adapting semantic segmentation using the Segment Anything Model and instance segmentation using Mask R-CNN to the setting. The latter ended up superior in our own evaluation and even got ranked second among all participants. We are proud that our approach has led to competitive performance. The source code is available from https://github.com/tum-bgd/GISCup23. Xuanshu Luo, Paul M. Walther, Wejdene Mansour, Balthasar Teuscher, Johann Maximilian Zollner, Hao Li 0019, Martin Werner 0001 |
SIGSPATIAL/GIS | 7 |
| 2021 | Trajectory Similarity using CompressionabstractIn this paper, we present a novel approach for trajectory similarity based on Kolmogorov complexity approximated by a lossy compression of the original trajectory data using selected features compressed into a concise memory representation by means of a Bloom filter. Given the importance of trajectory data, a linear-time distance measure with all theoretical guarantees implied by a proper metric is very powerful if it is capturing enough detail for important trajectory mining tasks. This stack of feature extraction combined with feature embedding in a Bloom filter constitutes a lossy compression for trajectory data which can easily be extended with other discrete data like travel mode. In addition, this compression has the needed properties for efficient calculation of a normalized compression distance (NCD) which approximates Kolmogorov complexity. We evaluate this novel trajectory distance measurement using very simple features and k-nearest-neighbor classification on selected realworld datasets with remarkable classification accuracies. Furthermore, we argue that the distance measure is very suited to geospatial big data applications as each trajectory is first transformed into few bits using the lossy compression stack. At time of comparison, the original trajectory geometry is not needed, instead, the sketches suffice. Despite very compressible parameters (equal or less than 1024 bit per trajectory) and very simple features, we already reach classification accuracies for real-world trajectory classification tasks of more than 80% across various datasets. Gabriel Dax, Martin Werner 0001 |
MDM | 2 |
| 2021 | Improving persistence based trajectory simplificationabstractIn this paper, we propose a novel linear time online algorithm for simplification of spatial trajectories. Trajectory simplification plays a major role in movement data analytics, in contexts such as reducing the communication overhead of tracking applications, keeping big data collections manageable, or harmonizing the number of points per trajectory. We follow the framework of topological persistence in order to detect a set of important points for the shape of the trajectory from local geometry information. Topological is meant in the mathematical sense in this paper and should not be confused with geographic topology. Our approach is able to prune pairs of non-persistent features in angle-representation of the trajectory. We show that our approach outperforms previous work, including multiresolution simplification (MRS) by a significant margin over a wide range of datasets without increasing computational complexity. In addition, we compare our novel algorithm with Douglas Peucker which is widely respected for its high-quality simplifications. We conclude that some datasets are better simplified using persistence-based methods and others are more difficult, but that the variations between the three considered variants of persistence-based simplification are small. In summary, this concludes that our novel pruning rule Segment-Distance Simplification (SDS) leads to more compact simplification results compared to β-pruning persistence and multiresolution simplification at similar quality levels in comparison to Douglas Peucker over a wide range of datasets. Moritz Laass, Marie Kiermeier, Martin Werner 0001 |
MDM | 3 |
| 2019 | GloBiMaps - A Probabilistic Data Structure for In-Memory Processing of Global Raster DatasetsabstractIn the last decade, more and more spatial data has been acquired on a global scale due to satellite missions, social media, and coordinated governmental activities. This observational data suffers from huge storage footprints and makes global analysis challenging. Therefore, many information products have been designed in which observations are turned into global maps showing features such as land cover or land use, often with only a few discrete values and sparse spatial coverage like only within cities. Martin Werner 0001 |
SIGSPATIAL/GIS | 1 |
| 2018 | DCount - A Probabilistic Algorithm for Accurately Disaggregating Building Occupant Counts into Room CountsabstractSensing accurately the number of occupants in the rooms of a building enables many important applications for smart building operation and energy management. A range of sensor technologies has been studied and applied to the problem. However, it is costly to achieve high accuracy by instrumenting all rooms in a building with dedicated occupant sensors. In this paper, we propose a new concept for estimating accurate room-level counts of occupants. The idea is to disaggregate accurate building-level counts via existing common sensors available at the room level. This solution is cost-effective as it scales to large buildings without requiring dedicated sensors in each room. We propose an algorithm named DCount that implements this concept. Our results document that DCount can provide room-level counts with a low normalized root mean squared error of 0.93. This is a major improvement compared to a state-of-the-art algorithm using common sensors and ventilation rate measurements resulting in a normalized root mean squared error of 1.54 on the same data set. Further more, we demonstrate how the results enable occupant-driven analysis of plug-load consumption which is one out of many applications using accurate room-level counts of occupants we hope to enable by proposing DCount. Mikkel Baun Kjærgaard, Martin Werner 0001, Fisayo Caleb Sangogboye, Krzysztof Arendt |
MDM | 2 |
| 2015 | The Gaussian Bloom Filter
Martin Werner 0001, Mirco Schönfeld |
DASFAA (1) | 1 |
| 2015 | BACR: set similarities with lower bounds and application to spatial trajectoriesabstractThis paper proposes a length-independent feature representation of sets of strings based on Bloom filters called BACR for similarity search in databases. Further, we show how a Z-curve-based discretization of geospatial trajectories can be used in order to search for similar trajectories in large databases. Additionally to the already-known estimation of the size of the union and the intersection of sets from Bloom filters, we propose a way to calculate an upper bound for the intersection and a lower bound for the union of sets. Consequently, we show that the Jaccard distance and many other similarity measures allow for a lower bound. This makes exact similarity search on large databases of this type feasible. Finally, we show that the Jaccard distance is incompatible with the union of sets and replace the Jaccard distance appropriately in a way such that even collections of sets of strings can be represented with a single BACR feature vector at least for similarity search applications. The algorithms are thoroughly evaluated and motivated by real-world examples. Martin Werner 0001 |
SIGSPATIAL/GIS | 1 |