Daniel Glake

dblp:210/8175 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2022
0000-0003-4575-9267ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2022 Operator Placement for Spatio-temporal Tasks
abstract
The amount of publicly available Spatio-temporal (ST) data is growing daily and possesses an increasing degree of complexity in more and more use cases. Besides spatial queries such as intersection, the requirements of current applications like Digital Twins (DT) go beyond the limits of a single data processing platform and need to combine a variety of queries with filtering ( e.g., k -NN), aggregation (e.g., counting), ranking (e.g., page-rank), clustering (e.g., k-means, ST-DBSCAN) and more, on ST-models. Since existing ST-platforms are highly specialized for a subset of these operations, it seems logical to distribute the data and queries across several of these systems. However, efficient p rocessing a cross d ifferent s ystems i s still a major challenge in polyglot data management and often demands manual query planning. To solve the automatic planning of those complex queries, we present an approach for cross-platform processing of ST-tasks that uses a symmetric join to handle platform heterogeneity and includes a novel algorithm for operator placement based on a latency model. Although the underlying problem is NP-hard and additional network transfers slow down the overall processing time, experiments on real-world tasks for DTs have shown that cross-platform processing can speed up well-known ST-tasks compared to the expensive query reformulations performed by state-of-the-art ST single-platform solutions.
Daniel Glake, Mareike Schmidt, Felix Kiehn, Fabian Panse, Ulfia Clemen, Thomas Clemen, Norbert Ritter
IEEE Big Data1
2022 Spatio-temporal Trajectory Learning using Simulation Systems
abstract
Spatio-temporal trajectories are essential factors for systems used in public transport, social ecology, and many other disciplines where movement is a relevant dynamic process. Each trajectory describes multiple state changes over time, induced by individual decision-making, based on psychological and social factors with physical constraints. Since a crucial factor of such systems is to reason about the potential trajectories in a closed environment, the primary problem is the realistic replication of individual decision making. Mental factors are often uncertain, not available or cannot be observed in reality. Thus, models for data generation must be derived from abstract studies using probabilities. To solve these problems, we present Multi-Agent-Trajectory-Learning (MATL), a state transition model to learn and generate human-like Spatio-temporal trajectory data. MATL combines Generative Adversarial Imitation Learning (GAIL) with a simulation system that uses constraints given by an agent-based model (Aℬℳ). We use GAIL to learn policies in conjunction with the Aℬℳ, resulting in a novel concept of individual decision making. Experiments with standard trajectory predictions show that our approach produces similar results to real-world observations.
Daniel Glake, Fabian Panse, Ulfia Clemen, Thomas Clemen, Norbert Ritter
CIKM1
2022 Adaptive Partitioning for Distributed Multi-Agent Simulations
abstract
Agent-based modeling and simulation is an essential paradigm for complex, data-intensive research questions to absorb and process emergent insights from often large-scale scenarios. That demands its execution within a distributed simulation system. One critical factor of those runtime systems is distributing and partitioning involved agents. Unsuitable partitioning schemes lead to a computing load that needs to be synchronized continuously, which bears the risk of drastic performance reductions. Work from load balancing has produced a series of distribution classes that rely on geometric distribution. The partitioning decomposes the agent environment so that when the agent leaves one partition, it instantly switches to another partition. However, many simulations cannot be partitioned in this spatial way, for example, because no spatial reference exists or agents interact on different temporal-granularity levels.
Daniel Glake, Florian Ocker, Ulfia Clemen, Thomas Clemen
SIGSIM-PADS1
2022 Polyglot Data Management: State of the Art & Open Challenges
abstract
Due to the increasing variety of the current database landscape, polyglot data management has become a hot research topic in recent years. The underlying idea is to combine the benefits of different data stores behind a predefined set of common interfaces and thus address use cases that individual stores cannot meet. This can be accomplished using different approaches which vary greatly in terms of capabilities, functionality, and architectural concepts. This tutorial provides a detailed overview of the current state of research in polyglot data management. We motivate its use by showing the high diversity of existing data stores and discussing three use cases in which individual stores are insufficient. Thereafter, we present different taxonomies for classifying polyglot data systems and give a detailed review of a number of selected systems. Finally, we compare these systems based on their features and discuss open challenges that still need to be addressed in future research.
Felix Kiehn, Mareike Schmidt, Daniel Glake, Fabian Panse, Wolfram Wingerath, Benjamin Wollmer, Martin Poppinga, Norbert Ritter
Proc. VLDB Endow.3
2021 Hierarchical Semantics Matching For Heterogeneous Spatio-temporal Sources
abstract
Spatio-temporal data are semantically valuable information used for various analytical tasks to identify spatially relevant and temporally limited correlations within a domain. The increasing availability and data acquisition from multiple sources with their typically high heterogeneity are getting more and more attention. However, these sources often lack interconnecting shared keys, making their integration a challenging problem. For example, publicly available parking data that consist of point data on parking facilities with fluctuating occupancy and static location data on parking spaces cannot be directly correlated. Both data sets describe two different aspects from distinct sources in which parking spaces and fluctuating occupancy are part of the same semantic model object. Especially for ad hoc analytical tasks on integrated models, these missing relationships cannot be handled using join operations as usual in relational databases. The reason lies in the lack of equijoin relationships, comparing for equality of strings and additional overhead in loading data up before processing. This paper addresses the optimization problem of finding suitable partners in the absence of equijoin relations for heterogeneous spatio-temporal data, applicable to ad hoc analytics. We propose a graph-based approach that achieves good recall and performance scaling via hierarchically separating the semantics along spatial, temporal, and domain-specific dimensions. We evaluate our approach using public data, showing that it is suitable for many standard join scenarios and highlighting its limitations.
Daniel Glake, Norbert Ritter, Florian Ocker, Nima Ahmady-Moghaddam, Daniel Osterholz, Ulfia Clemen, Thomas Clemen
CIKM1
2021 Multi-Agent Systems and Digital Twins for Smarter Cities
abstract
An intelligent combination of the Internet of Things (IoT) and approaches to modeling and simulation is one of the most challenging endeavors for future cities, manufacturing industries, and predictive maintenance. Digital Twins take on a unique role here. However, the question of what a Digital Twin is and what differentiates it from a regular model is still open. We present an experimental setup for integrating an existing simulation model of Hamburg's traffic system with the city's real-time sensor network. The Digital Twin is implemented using the large-scale multi-agent framework MARS. The entire process from the model description to retrieving real-time data from the IoT sensors and incorporating it in the simulation is presented. As a first prototypical example, a multi-modal mobility model was connected to real-world bike-sharing locations in Hamburg. We find that the combination of multi-agent systems and IoT sensors as a Digital Twin shows enormous potential for city planners, policy stakeholders, and other decision-makers. By correcting the course of a simulation via real-time data, the corridor-of-uncertainty that is intrinsic to some simulation models' use can be reduced significantly. Furthermore, any divergence of simulated and sampled data can lead to a deeper understanding of complex adaptive systems like big cities.
Thomas Clemen, Nima Ahmady-Moghaddam, Ulfia Clemen, Florian Ocker, Daniel Osterholz, Jonathan Ströbele, Daniel Glake
SIGSIM-PADS7