EDBT 2026 Demo / reviewers in the wild / expert
Stefan Leyk
dblp:48/5378
· DBLP profile ↗
7ranked-venue papers in the field
1as first author
2since 2021 · last 2021
0000-0001-9180-4853ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 2Big Data, Cloud & Distributed Data Systems · 2Data Mining & Knowledge Discovery · 1Knowledge Engineering, Semantic Web & Information Systems · 1Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Guided Generative Models using Weak Supervision for Detecting Object Spatial Arrangement in Overhead ImagesabstractThe increasing availability and accessibility of numerous overhead images allows us to estimate and assess the spatial arrangement of groups of geospatial target objects, which can benefit many applications, such as traffic monitoring and agricultural monitoring. Spatial arrangement estimation is the process of identifying the areas which contain the desired objects in overhead images. Traditional supervised object detection approaches can estimate accurate spatial arrangement but require large amounts of bounding box annotations. Recent semi-supervised clustering approaches can reduce manual labeling but still require annotations for all object categories in the image. This paper presents the target-guided generative model (TGGM), under the Variational Auto-encoder (VAE) framework, which uses Gaussian Mixture Models (GMM) to estimate the distributions of both hidden and decoder variables in VAE. Modeling both hidden and decoder variables by GMM reduces the required manual annotations significantly for spatial arrangement estimation. Unlike existing approaches that the training process can only update the GMM as a whole in the optimization iterations (e.g., a "minibatch"), TGGM allows the update of individual GMM components separately in the same optimization iteration. Optimizing GMM components separately allows TGGM to exploit the semantic relationships in spatial data and requires only a few labels to initiate and guide the generative process. Our experiments shows that TGGM achieves results comparable to the state-of-the-art semi-supervised methods and outperformes unsupervised methods by 10% based on the F1scores, while requiring significantly fewer labeled data. Yao-Yi Chiang, Stefan Leyk, Johannes H. Uhl, Craig A. Knoblock |
IEEE BigData | 3 |
| 2021 | A Label Correction Algorithm Using Prior Information for Automatic and Accurate Geospatial Object RecognitionabstractThousands of scanned historical topographic maps contain valuable information covering long periods of time, such as how the hydrography of a region has changed over time. Efficiently unlocking the information in these maps requires training a geospatial objects recognition system, which needs a large amount of annotated data. Overlapping geo-referenced external vector data with topographic maps according to their coordinates can annotate the desired objects’ locations in the maps automatically. However, directly overlapping the two datasets causes misaligned and false annotations because the publication years and coordinate projection systems of topographic maps are different from the external vector data. We propose a label correction algorithm, which leverages the color information of maps and the prior shape information of the external vector data to reduce misaligned and false annotations. The experiments show that the precision of annotations from the proposed algorithm is 10% higher than the annotations from a state-of-the-art algorithm. Consequently, recognition results using the proposed algorithm’s annotations achieve 9% higher correctness than using the annotations from the state-of-the-art algorithm. Yao-Yi Chiang, Stefan Leyk, Johannes H. Uhl, Craig A. Knoblock |
IEEE BigData | 3 |
| 2020 | Building Linked Spatio-Temporal Data from Vectorized Historical Maps
Basel Shbita, Craig A. Knoblock, Yao-Yi Chiang, Johannes H. Uhl, Stefan Leyk |
ESWC | 6 |
| 2020 | An Automatic Approach for Generating Rich, Linked Geo-Metadata from Historical Map ImagesabstractHistorical maps contain detailed geographic information difficult to find elsewhere covering long-periods of time (e.g., 125 years for the historical topographic maps in the US). However, these maps typically exist as scanned images without searchable metadata. Existing approaches making historical maps searchable rely on tedious manual work (including crowd-sourcing) to generate the metadata (e.g., geolocations and keywords). Optical character recognition (OCR) software could alleviate the required manual work, but the recognition results are individual words instead of location phrases (e.g., "Black'' and "Mountain'' vs. "Black Mountain''). This paper presents an end-to-end approach to address the real-world problem of finding and indexing historical map images. This approach automatically processes historical map images to extract their text content and generates a set of metadata that is linked to large external geospatial knowledge bases. The linked metadata in the RDF (Resource Description Framework) format support complex queries for finding and indexing historical maps, such as retrieving all historical maps covering mountain peaks higher than 1,000 meters in California. We have implemented the approach in a system called mapKurator. We have evaluated mapKurator using historical maps from several sources with various map styles, scales, and coverage. Our results show significant improvement over the state-of-the-art methods. The code has been made publicly available as modules of the Kartta Labs project at https://github.com/kartta-labs/Project. Zekun Li 0007, Yao-Yi Chiang, Sasan Tavakkol, Basel Shbita, Johannes H. Uhl, Stefan Leyk, Craig A. Knoblock |
KDD | 6 |
| 2020 | Automatic alignment of contemporary vector data and georeferenced historical maps using reinforcement learningabstractWith large amounts of digital map archives becoming available, automatically extracting information from scanned historical maps is needed for many domains that require long-term historical geographic data. Convolutional Neural Networks (CNN) are powerful techniques that can be used for extracting locations of geographic features from scanned maps if sufficient representative training data are available. Existing spatial data can provide the approximate locations of corresponding geographic features in historical maps and thus be useful to annotate training data automatically. However, the feature representations, publication date, production scales, and spatial reference systems of contemporary vector data are typically very different from those of historical maps. Hence, such auxiliary data cannot be directly used for annotation of the precise locations of the features of interest in the scanned historical maps. This research introduces an automatic vector-to-raster alignment algorithm based on reinforcement learning to annotate precise locations of geographic features on scanned maps. This paper models the alignment problem using the reinforcement learning framework, which enables informed, efficient searches for matching features without pre-processing steps, such as extracting specific feature signatures (e.g. road intersections). The experimental results show that our algorithm can be applied to various features (roads, water lines, and railroads) and achieve high accuracy. Yao-Yi Chiang, Stefan Leyk, Johannes H. Uhl, Craig A. Knoblock |
Int. J. Geogr. Inf. Sci. | 3 |
| 2018 | Enhancing areal interpolation frameworks through dasymetric refinement to create consistent population estimates across censusesabstractTo assess micro-scale population dynamics effectively, demographic variables should be available over temporally consistent small area units. However, fine-resolution census boundaries often change between survey years. This research advances areal interpolation methods with dasymetric refinement to create accurate consistent population estimates in 1990 and 2000 (source zones) within tract boundaries of the 2010 census (target zones) for five demographically distinct counties in the U.S. Three levels of dasymetric refinement of source and target zones are evaluated. First, residential parcels are used as a binary ancillary variable prior to regular areal interpolation methods. Second, Expectation Maximization (EM) and its data-extended version leverage housing types of residential parcels as a related ancillary variable. Finally, a third refinement strategy to mitigate the overestimation effect of large residential parcels in rural areas uses road buffers and developed land cover classes. Results suggest the effectiveness of all three levels of dasymetric refinement in reducing estimation errors. They provide a first insight into the potential accuracy improvement achievable in varying geographic and demographic settings but also through the combination of different refinement strategies in parts of a study area. Such improved consistent population estimates are the basis for advanced spatio-temporal demographic research. Hamidreza Zoraghein, Stefan Leyk |
Int. J. Geogr. Inf. Sci. | 2 |
| 2010 | Colors of the past: color image segmentation in historical topographic maps based on homogeneity
Stefan Leyk, Ruedi Boesch |
GeoInformatica | 1 |