VLDB 2026 Research / reviewers in the wild / expert
Michael Gertz 0001
dblp:g/MichaelGertz
· DBLP profile ↗
68ranked-venue papers in the field
4as first author
6since 2021 · last 2025
0000-0003-4530-6110ORCID · verified
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 38 (3 first)Information Retrieval & Web Search · 19Data Mining & Knowledge Discovery · 6Other / Interdisciplinary · 4 (1 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ClusterChat: Multi-Feature Search for Corpus Exploration
Ashish Chouhan, Saifeldin Mandour, Michael Gertz 0001 |
SIGIR | 3 |
| 2024 | QuantPlorer: Exploration of Quantities in Text
Satya Almasian, Alexander Kosnac, Michael Gertz 0001 |
ECIR (5) | 3 |
| 2023 | TrendTracker: Temporal, network-based exploration of long-term Twitter trendsabstractTrendTracker is a web application for the network-based and temporal exploration of long-term social media trends. Topical trends, represented as a series of hashtag co-occurrence networks, can interactively be explored while the user is provided with detailed trend analysis insights. This approach has several benefits compared to alternative trend visualization and exploration methods, such as ranked lists of trending keywords, as it provides the user with additional context-sensitive information. To showcase the TrendTracker application, we leverage a Twitter dataset of German political actors and demonstrate the system's capabilities in various ways. For example, the user is able to investigate a single trend from multiple perspectives, such as the trend's temporal development over time, including its topical shifts and changes in popularity. Also, given the network-based trend visualization, the user can intuitively understand the different facets of a trend and how these are interrelated. Thereby, individual hashtags and relationships can be tracked over time as well. Furthermore, the TrendTracker application allows the user to compare trends. This way, differences in the trends' temporal evolution or topical alignment can be uncovered. The demo is publicly available via the following URL: https://trend-tracker.ifi.uni-heidelberg.de. John Ziegler, Johannes Sindlinger, Marina Walther, Michael Gertz 0001 |
ASONAM | 4 |
| 2023 | Who Is behind a Trend? Temporal Analysis of Interactions among Trend Participants on TwitterabstractTrends are a fundamental component of today's fast-evolving media landscape. Still, a lot of questions about who participates in such trends remain unanswered. Are trends driven by individual actors, or do interactions between actors reveal community structures? If so, do those structures change during the life cycle of a trend or between topically similar trends? In short: Who is behind a trend? This paper contributes to a better understanding of these questions and, in general, actor networks underlying trends on social media. As a case study, we leverage a large Twitter dataset from the EURO2020 soccer competition to detect and analyze topical trends. Our novel Gaussian fitting method allows separating trend life cycles into up- and down-trend components, as well as determining the duration of trends. An event-based evaluation proves good performance results. Given separate trend stages and topically similar trends at different points in time, we conduct a temporal analysis of the actor networks during trends. Our findings not only reveal a large overlap of participants between successive trends but also indicate large variations within a trend life cycle. Furthermore, actor networks seem to be centred around a small number of dominant users and communities. Those users also show large stability across similar trends over time. In contrast, temporally stable community structures are neither found within nor across topically similar trends. John Ziegler, Michael Gertz 0001 |
ICWSM | 2 |
| 2022 | QFinder: A Framework for Quantity-centric RankingabstractQuantities shape our understanding of measures and values, and they are an important means to communicate the properties of objects. Often, search queries contain numbers as retrieval units, e.g., "iPhone that costs less than 800 Euros''. Yet, modern search engines lack a proper understanding of numbers and units. In queries and documents, search engines handle them as normal keywords and therefore are ignorant of relative conditions between numbers, such as greater than or less than, or, more generally, the numerical proximity of quantities. In this work, we demonstrate QFinder, our quantity-centric framework for ranking search results for queries with quantity constraints. We also open-source our new ranking method as an Elasticsearch plug-in for future use. Our demo is available at: https://qfinder.ifi.uni-heidelberg.de/ Satya Almasian, Milena Bruseva, Michael Gertz 0001 |
SIGIR | 3 |
| 2022 | Online DATEing: A Web Interface for Temporal AnnotationsabstractDespite more than two decades of research on temporal tagging and temporal relation extraction, usable tools for annotating text remain very basic and hard to set up from an average end-user perspective, limiting the applicability of developments to a selected group of invested researchers. In this work, we aim to increase the accessibility of temporal tagging systems by presenting an intuitive web interface, called "Online DATEing", which simplifies the interaction with existing temporal annotation frameworks. Our system integrates several approaches in a single interface and streamlines the process of importing (and tagging) groups of documents, as well as making it accessible through a programmatic API. It further enables users to interactively investigate and visualize tagged texts, and is designed with an extensible API for the inclusion of new models or data formats. A web demonstration of our tool is available at https://onlinedating.ifi.uni-heidelberg.de and public code accessible at https://github.com/satya77/Temporal_Tagger_Service. Dennis Aumiller, Satya Almasian, David Pohl, Michael Gertz 0001 |
SIGIR | 4 |
| 2020 | TiCCo: Time-Centric Content ExplorationabstractTime is a natural way to order information and can be utilized to summarize events and to construct a chronology of contents within a document collection in many application domains. Structuring the sequence of events along a timeline allows users to grasp information at-a-glance, which enables them to get familiar with a topic in only a short amount of time and can hence support the analysis of more complicated and heterogeneous textual data. The manual construction of timelines, however, is a tedious and error-prone task, leading to static timeline representations that limit users to a passive role. In this paper, TiCCo, an automated extraction pipeline from arbitrary English and German text collections, is provided and presented to the user in an interactive manner. This puts the user in an active role in which she not only absorbs knowledge, but also influences in which ways the information is presented to her. In-depth investigations of a specific point in time are augmented by utilizing time-centric co-occurrence graphs that further summarize information extracted from a document collection, and enable users to explore the chronology of events by allowing them to interact with the constructed graphs as well as the underlying documents. Philip Hausner, Dennis Aumiller, Michael Gertz 0001 |
CIKM | 3 |
| 2020 | A Versatile Hypergraph Model for Document CollectionsabstractEfficiently and effectively representing large collections of text is of central importance to information retrieval tasks such as summarization and search. Since models for these tasks frequently rely on an implicit graph structure of the documents or their contents, graph-based document representations are naturally appealing. For tasks that consider the joint occurrence of words or entities, however, existing document representations often fall short in capturing cooccurrences of higher order, higher multiplicity, or at varying proximity levels. Furthermore, while numerous applications benefit from structured knowledge sources, external data sources are rarely considered as integral parts of existing document models. Andreas Spitz, Dennis Aumiller, Bálint Soproni, Michael Gertz 0001 |
SSDBM | 4 |
| 2019 | Word Embeddings for Entity-Annotated Texts
Satya Almasian, Andreas Spitz, Michael Gertz 0001 |
ECIR (1) | 3 |
| 2019 | Retrieving Multi-Entity Associations: An Evaluation of Combination Modes for Word EmbeddingsabstractWord embeddings have gained significant attention as learnable representations of semantic relations between words, and have been shown to improve upon the results of traditional word representations. However, little effort has been devoted to using embeddings for the retrieval of entity associations beyond pairwise relations. In this paper, we use popular embedding methods to train vector representations of an entity-annotated news corpus, and evaluate their performance for the task of predicting entity participation in news events versus a traditional word cooccurrence network as a baseline. To support queries for events with multiple participating entities, we test a number of combination modes for the embedding vectors. While we find that even the best combination modes for word embeddings do not quite reach the performance of the full cooccurrence network, especially for rare entities, we observe that different embedding methods model different types of relations, thereby indicating the potential for ensemble methods. Gloria Feher, Andreas Spitz, Michael Gertz 0001 |
SIGIR | 3 |
| 2019 | TopExNet: Entity-Centric Network Topic Exploration in News StreamsabstractThe recent introduction of entity-centric implicit network representations of unstructured text offers novel ways for exploring entity relations in document collections and streams efficiently and interactively. Here, we present TopExNet as a tool for exploring entity-centric network topics in streams of news articles. The application is available as a web service at https://topexnet.ifi.uni-heidelberg.de. Andreas Spitz, Satya Almasian, Michael Gertz 0001 |
WSDM | 3 |
| 2018 | Entity-Centric Topic Extraction and Exploration: A Network-Based Approach
Andreas Spitz, Michael Gertz 0001 |
ECIR | 2 |
| 2018 | Efficient anti-community detection in complex networksabstractModeling the relations between the components of complex systems as networks of vertices and edges is a commonly used method in many scientific disciplines that serves to obtain a deeper understanding of the systems themselves. In particular, the detection of densely connected communities in these networks is frequently used to identify functionally related components, such as social circles in networks of personal relations or interactions between agents in biological networks. Traditionally, communities are considered to have a high density of internal connections, combined with a low density of external edges between different communities. However, not all naturally occurring communities in complex networks are characterized by this notion of structural equivalence, such as groups of energy states with shared quantum numbers in networks of spectral line transitions. In this paper, we focus on this inverse task of detecting anti-communities that are characterized by an exceptionally low density of internal connections and a high density of external connections. While anti-communities have been discussed in the literature for anecdotal applications or as a modification of traditional community detection, no rigorous investigation of algorithms for the problem has been presented. To this end, we introduce and discuss a broad range of possible approaches and evaluate them with regard to efficiency and effectiveness on a range of real-world and synthetic networks. Furthermore, we show that the presence of a community and anti-community structure are not mutually exclusive, and that even networks with a strong traditional community structure may also contain anti-communities. Sebastian Lackner, Andreas Spitz, Matthias Weidemüller, Michael Gertz 0001 |
SSDBM | 4 |
| 2018 | Numerically stable parallel computation of (co-)varianceabstractWith the advent of big data, we see an increasing interest in computing correlations in huge data sets with both many instances and many variables. Essential descriptive statistics such as the variance, standard deviation, covariance, and correlation can suffer from a numerical instability known as "catastrophic cancellation" that can lead to problems when naively computing these statistics with a popular textbook equation. While this instability has been discussed in the literature already 50 years ago, we found that even today, some high-profile tools still employ the instable version. Erich Schubert, Michael Gertz 0001 |
SSDBM | 2 |
| 2018 | MetaStore: an adaptive metadata management framework for heterogeneous metadata models
Ajinkya Prabhune, Rainer Stotzka, Vaibhav Sakharkar, Jürgen Hesser, Michael Gertz 0001 |
Distributed Parallel Databases | 5 |
| 2018 | P-PIF: a ProvONE provenance interoperability framework for analyzing heterogeneous workflow specifications and provenance traces
Ajinkya Prabhune, Aaron Zweig, Rainer Stotzka, Jürgen Hesser, Michael Gertz 0001 |
Distributed Parallel Databases | 5 |
| 2018 | Editorial: Advances in spatial and temporal databases
Michael Gertz 0001, Matthias Renz, Xiaofang Zhou 0001 |
GeoInformatica | 1 |
| 2017 | Intrinsic t-Stochastic Neighbor Embedding for Visualization and Outlier Detection - A Remedy Against the Curse of Dimensionality?
Erich Schubert, Michael Gertz 0001 |
SISAP | 2 |
| 2017 | Efficient online extraction of keywords for localized events in twitter
Hamed Abdelhaq, Michael Gertz 0001, Ayser Armiti |
GeoInformatica | 2 |
| 2016 | MetaStore: A metadata framework for scientific data repositoriesabstractIn this paper, we present MetaStore, a metadata management framework for scientific data repositories. Scientific experiments are generating a deluge of data and metadata. Metadata is critical for scientific research, as it enables discovering, analysing, reusing, and sharing of scientific data. Moreover, metadata produced by scientific experiments is heterogeneous and subject to frequent changes, demanding a flexible data model. Currently, there does not exist an adaptive and a generic solution that is capable of handling heterogeneous metadata models. To address this challenge, we present MetaStore, an adaptive metadata management framework based on a NoSQL database. To handle heterogeneous metadata models and standards, the MetaStore automatically generates the necessary software code (services) and extends the functionality of the framework. To leverage the functionality of NoSQL databases, the MetaStore framework allows full-text search over metadata through automated creation of indexes. Finally, a dedicated REST service is provided for efficient harvesting (sharing) of metadata using the METS metadata standard over the OAI-PMH protocol. Ajinkya Prabhune, Hasebullah Ansari, Anil Keshav, Rainer Stotzka, Michael Gertz 0001, Jürgen Hesser |
IEEE BigData | 5 |
| 2016 | Terms over LOAD: Leveraging Named Entities for Cross-Document Extraction and Summarization of EventsabstractReal world events, such as historic incidents, typically contain both spatial and temporal aspects and involve a specific group of persons. This is reflected in the descriptions of events in textual sources, which contain mentions of named entities and dates. Given a large collection of documents, however, such descriptions may be incomplete in a single document, or spread across multiple documents. In these cases, it is beneficial to leverage partial information about the entities that are involved in an event to extract missing information. In this paper, we introduce the LOAD model for cross-document event extraction in large-scale document collections. The graph-based model relies on co-occurrences of named entities belonging to the classes locations, organizations, actors, and dates and puts them in the context of surrounding terms. As such, the model allows for efficient queries and can be updated incrementally in negligible time to reflect changes to the underlying document collection. We discuss the versatility of this approach for event summarization, the completion of partial event information, and the extraction of descriptions for named entities and dates. We create and provide a LOAD graph for the documents in the English Wikipedia from named entities extracted by state-of-the-art NER tools. Based on an evaluation set of historic data that include summaries of diverse events, we evaluate the resulting graph. We find that the model not only allows for near real-time retrieval of information from the underlying document collection, but also provides a comprehensive framework for browsing and summarizing event data. Andreas Spitz, Michael Gertz 0001 |
SIGIR | 2 |
| 2016 | Geometric Graph Indexing for Similarity Search in Scientific DatabasesabstractSearching a database for similar graphs is a critical task in many scientific applications, such as in drug discovery, geoinformatics, or pattern recognition. Typically, graph edit distance is used to estimate the similarity of non-identical graphs, which is a very hard task. Several indexing structures and lower bound distances have been proposed to prune the search space. Most of them utilize the number of edit operations and assume graphs with a discrete label alphabet that has a certain canonical order. Unfortunately, such assumptions cannot be guaranteed for geometric graphs where vertices have coordinates in some two dimensional space. Ayser Armiti, Michael Gertz 0001 |
SSDBM | 2 |
| 2015 | Beyond Friendships and Followers: The Wikipedia Social NetworkabstractMost traditional social networks rely on explicitly given relations between users, their friends and followers. In this paper, we go beyond well structured data repositories and create a person-centric network from unstructured text -- the Wikipedia Social Network. To identify persons in Wikipedia, we make use of interwiki links, Wikipedia categories and person related information available in Wikidata. From the co-occurrences of persons on a Wikipedia page we construct a large-scale person-centric network and provide a weighting scheme for the relationship of two persons based on the distances of their mentions within the text. We extract key characteristics of the network such as centrality, clustering coefficient and component sizes for which we find values that are typical for social networks. Using state-of-the-art algorithms for community detection in massive networks, we identify interesting communities and evaluate them against Wikipedia categories. The Wikipedia social network developed this way provides an important source for future social analysis tasks. Johanna Geiß, Andreas Spitz, Michael Gertz 0001 |
ASONAM | 3 |
| 2015 | Breaking the News: Extracting the Sparse Citation Network Backbone of Online News ArticlesabstractNetworks of online news articles and blog posts are some of the most commonly used data sets in network science. As a result, they have become a vital piece of network analysis and are used for the evaluation of algorithms that work on large networks, or serve as examples in the analysis of information diffusion and propagation. Similarly, scientific citation networks are part of the bedrock upon which much of modern network analysis is built and have been studied for decades. In this paper, we show that the backbone inherent to networks of online news articles shares significant structural similarities to scientific citation networks once the noise of spurious links is stripped away. We present a data set of news articles that, while it is extremely sparse and lightweight, still contains information relevant to the propagation of information in mass media and is remarkably similar to scientific citation networks, thus opening the door to the use of established methodologies from scientometrics and bibliometrics in the analysis of online news propagation. Andreas Spitz, Michael Gertz 0001 |
ASONAM | 2 |
| 2015 | Character retrieval of vectorized cuneiform scriptabstractMotivated by the increasing demand for computerized analysis of documents within the Digital Humanities we present an approach to automating handwritten cuneiform character recognition on vectorized cuneiform tablets. Cuneiform is one of the oldest handwritten scripts used for more than three millennia. In previous work we have shown how to extract vector drawings from 3D-models of cuneiform tablets similar to those manually drawn over digital photographs. We approach the problem of recognizing these characters by applying pattern matching against the basic structural features of cuneiform, the wedge-shaped impressions. Then, we find an optimal assignment between the wedge configuration of two characters w.r.t. wedge shape and position. The similarity of two characters is measured by the quality of the assignment. We compare our method against well known methods for handwritten character recognition with favorable results for our method. Bartosz Bogacz, Michael Gertz 0001, Hubert Mara |
ICDAR | 2 |
| 2014 | An Event-Based Framework for the Semantic Annotation of Locations
Michael Gertz 0001, Christian Sengstock |
ADBIS | 2 |
| 2014 | rLinkTopic: A probabilistic model for discovering regional LinkTopic communitiesabstractAlthough geographic and regional aspects of communities find many practical applications, e.g., in social studies and marketing, to date, existing approaches to community detection have paid little attention to these features when analyzing social network data. To address these shortcomings, we introduce the concept of regional LinkTopic communities and propose a novel probabilistic model for extracting such communities. Our model jointly considers the spatio-temporal proximity of users in terms of the messages they post over time, together with contextual links and message topics to determine communities. The model allows users to have a membership in more than just one community, an important feature when discovering communities based on topics. Each community derived by our approach is not only described by a mixture of topics but also by its regional properties. Using data from Twitter, we demonstrate the effectiveness of our model in extracting regional LinkTopic communities, which are described in terms of both geographic locations and coherent topics. The experimental results show that our model outperforms related models that only use links and topics to extract communities by the measure of perplexity. Tran Van Canh, Michael Gertz 0001 |
ASONAM | 2 |
| 2014 | Geometric graph matching and similarity: a probabilistic approachabstractFinding common structures is vital for many graph-based applications, such as road network analysis, pattern recognition, or drug discovery. Such a task is formalized as the inexact graph matching problem, which is known to be NP-hard. Several graph matching algorithms have been proposed to find approximate solutions. However, such algorithms still face many problems in terms of memory consumption, runtime, and tolerance to changes in graph structure or labels. Ayser Armiti, Michael Gertz 0001 |
SSDBM | 2 |
| 2013 | Mining Periodic Event Patterns from RDF Datasets
Van Quoc Anh Le, Michael Gertz 0001 |
ADBIS | 2 |
| 2013 | Spatial Itemset Mining: A Framework to Explore Itemsets in Geographic Space
Christian Sengstock, Michael Gertz 0001 |
ADBIS | 2 |
| 2013 | A spatial LDA model for discovering regional communitiesabstractModels and techniques for the extraction and analysis of communities from social network data have become a major area of research. Most of the prominent approaches exploit the link structure among users based on, e.g., information about followers or the exchange of messages among users. However, there are also other types of information that are useful for extracting communities from social network data, such as geographic information associated with postings and users. In this paper, we present a novel approach to discover so-called regional communities. Motivated by the fact that more and more postings to social networks include the geo-location of users, we claim that communities also form even if their users do not necessarily interact but are posting (similar) messages in both spatial and temporal proximity. To discover such regional communities we propose a generative probabilistic model based on spatial latent Dirichlet allocation (SLDA) that unveils not only regional communities but also topics associated with these communities. We demonstrate the effectiveness of our approach using Twitter data and compare the properties of communities detected that way with communities discovered by approaches using link graphs. Tran Van Canh, Michael Gertz 0001 |
ASONAM | 2 |
| 2013 | Proximity2-aware ranking for textual, temporal, and geographic queriesabstractTemporal and geographic information needs are frequent and important but not well served by standard IR systems. Recent approaches address such needs by extracting and normalizing temporal and geographic expressions from documents. They calculate specific scores for the temporal and/ or geographic parts of a query. However, all approaches assume independence between the different query parts. In this paper, we present a new model to rank documents according to combined textual, temporal, and geographic queries. The independence assumption between the query parts is eliminated by calculating proximity scores. Thus, documents are regarded to be more relevant if terms and expressions satisfying the different query parts occur close to each other in a document. As our evaluations based on the NTCIR-GeoTime data show, our proposed model outperforms baseline models that do not use proximity information. Jannik Strötgen, Michael Gertz 0001 |
CIKM | 2 |
| 2013 | Spatio-temporal characteristics of bursty words in Twitter streamsabstractSocial networking and microblogging services such as Twitter provide a continuous source of data from which useful information can be extracted. The detection and characterization of bursty words play an important role in processing such data, as bursty words might hint to events or trending topics of social importance upon which actions can be triggered. While there are several approaches to extract bursty words from the content of messages, there is only little work that deals with the dynamics of continuous streams of messages, in particular messages that are geo-tagged. Hamed Abdelhaq, Michael Gertz 0001, Christian Sengstock |
SIGSPATIAL/GIS | 2 |
| 2013 | Efficient geometric graph matching using vertex embeddingabstractFor many applications such as road network analysis and image processing, it is critical to study spatial properties of objects in addition to object relationships. Geometric graphs provide a suitable modeling framework for such applications, where vertices are located in some 2D space. For applications where the similarity between the structures of different graphs plays an important role, typically, inexact graph matching algorithms are employed. However, graph matching algorithms face many problems such as scalability with respect to graph size and less tolerance to changes in graph structure or labels. Ayser Armiti, Michael Gertz 0001 |
SIGSPATIAL/GIS | 2 |
| 2013 | A probablistic model for spatio-temporal signal extraction from social mediaabstractIt is nowadays possible to access a huge and increasing stream of social media records. Recently, such data has been used to infer about spatio-temporal phenomena by treating the records as proxy observations of the real world. However, since such observations are heavily uncertain and their spatio-temporal distribution is highly heterogeneous, extracting meaningful signals from such data is a challenging task. In this paper, we present a probabilistic model to extract spatio-temporal distributions of phenomena (called spatio-temporal signals) from social media. Our approach models spatio-temporal and semantic knowledge about real-world phenomena embedded in records on the basis of conditional probability distributions in a Bayesian network. Through this, we realize a generic and comprehensive model where knowledge and uncertainties about spatio-temporal phenomena can be described in a modular and extensible fashion. We show that existing models for the extraction of spatio-temporal phenomena distributions from social media are particular instances of our model. We quantitatively evaluate instances of our model by comparing the spatio-temporal distributions of extracted phenomena from a large Twitter data set to their real-world distributions. The results clearly show that our model allows to extract better spatio-temporal signals in terms of quality and robustness. Christian Sengstock, Michael Gertz 0001, Florian Flatow, Hamed Abdelhaq |
SIGSPATIAL/GIS | 2 |
| 2013 | Reliable Spatio-temporal Signal Extraction and Exploration from Human Activity Records
Christian Sengstock, Michael Gertz 0001, Hamed Abdelhaq, Florian Flatow |
SSTD | 2 |
| 2013 | EvenTweet: Online Localized Event Detection from TwitterabstractMicroblogging services such as Twitter, Facebook, and Foursquare have become major sources for information about real-world events. Most approaches that aim at extracting event information from such sources typically use the temporal context of messages. However, exploiting the location information of georeferenced messages, too, is important to detect localized events, such as public events or emergency situations. Users posting messages that are close to the location of an event serve as human sensors to describe an event. In this demonstration, we present a novel framework to detect localized events in real-time from a Twitter stream and to track the evolution of such events over time. For this, spatio-temporal characteristics of keywords are continuously extracted to identify meaningful candidates for event descriptions. Then, localized event information is extracted by clustering keywords according to their spatial similarity. To determine the most important events in a (recent) time frame, we introduce a scoring scheme for events. We demonstrate the functionality of our system, called Even-Tweet, using a stream of tweets from Europe during the 2012 UEFA European Football Championship. Hamed Abdelhaq, Christian Sengstock, Michael Gertz 0001 |
Proc. VLDB Endow. | 3 |
| 2012 | Retro: Time-Based Exploration of Product Reviews
Jannik Strötgen, Omar Alonso, Michael Gertz 0001 |
ECIR | 3 |
| 2012 | Latent geographic feature extraction from social mediaabstractIn this work we present a framework for the unsupervised extraction of latent geographic features from georeferenced social media. A geographic feature represents a semantic dimension of a location and can be seen as a sensor that measures a signal of geographic semantics. Our goal is to extract a small number of informative geographic features from social media, to describe and explore geographic space, and for subsequent spatial analysis, e.g., in market research. We propose a framework that, first, transforms the unstructured and noisy geographic information in social media into a high-dimensional multivariate signal of geographic semantics. Then, we use dimensionality reduction to extract latent geographic features. We conduct experiments using two large-scale Flickr data sets covering the LA area and the US. We show that dimensionality reduction techniques extracting sparse latent features find dimensions with higher informational value. In addition, we show that prior normalization can be used as a parameter in the exploration process to extract features representing different geographic characteristics, that is, landmarks, regional phenomena, or global phenomena. Christian Sengstock, Michael Gertz 0001 |
SIGSPATIAL/GIS | 2 |
| 2011 | Exploration and comparison of geographic information sources using distance statisticsabstractGiven the steadily increasing amount of geographic information on the Web, there is a strong need for suitable methods in exploratory data analysis that can be used to efficiently describe the characteristics of such large-scale, often noisy datasets. Existing methods in spatial data mining focus primarily on mining patterns describing spatial proximity relationships such as co-location patterns or spatial associations rules. Christian Sengstock, Michael Gertz 0001 |
GIS | 2 |
| 2011 | An event-centric model for multilingual document similarityabstractDocument similarity measures play an important role in many document retrieval and exploration tasks. Over the past decades, several models and techniques have been developed to determine a ranked list of documents similar to a given query document. Interestingly, the proposed approaches typically rely on extensions to the vector space model and are rarely suited for multilingual corpora. Jannik Strötgen, Michael Gertz 0001, Conny Junghans |
SIGIR | 2 |
| 2011 | Enhancing Document Snippets Using Temporal Information
Omar Alonso, Michael Gertz 0001, Ricardo Baeza-Yates |
SPIRE | 2 |
| 2010 | Temporal Analysis of Document Collections: Framework and Applications
Omar Alonso, Michael Gertz 0001, Ricardo Baeza-Yates |
SPIRE | 2 |
| 2010 | TimeTrails: A System for Exploring Spatio-Temporal Information in DocumentsabstractSpatial and temporal data have become ubiquitous in many application domains such as the Geosciences or life sciences. Sophisticated database management systems are employed to manage such structured data. However, an important source of spatio-temporal information that has not been fully utilized are unstructured text documents. In documents, combinations of temporal and spatial expressions form events, which can be mapped to a database structure and organized into trajectories that can be explored. In this context, the coupling of information retrieval techniques with spatio-temporal database concepts leads to new ways for managing and exploring document collections. In this demonstration, we present TimeTrails, a system for the extraction, querying, storage, and exploration of spatio-temporal information embedded in text documents. The user can query a document collection, and TimeTrails visualizes the spatio-temporal information extracted from relevant documents as document trajectories, resulting in a map-based view of documents. This view helps the user to explore the temporal and spatial content of documents in a meaningful way and to further restrict search results using spatial and temporal predicates. Jannik Strötgen, Michael Gertz 0001 |
Proc. VLDB Endow. | 2 |
| 2009 | Clustering and exploring search results using timeline constructionsabstractTime is an important dimension of any information space and can be very useful in information retrieval and in particular clustering and exploration of search results. Search result clustering is a feature integrated in some of today's search engines, allowing users to further explore search results. However, only little work has been done on exploiting temporal information embedded in documents for the presentation, clustering, and exploration of search results along well-defined timelines. In this paper, we present an add-on to traditional information retrieval applications in which we exploit various temporal information associated with documents to present and cluster documents along timelines. Temporal information expressed in the form of, e.g., date and time tokens or temporal references, appear in documents as part of the textual context or metadata. Using temporal entity extraction techniques, we show how temporal expressions are made explicit and used in the construction of multiple-granularity timelines. We discuss how hit-list based search results can be clustered according to temporal aspects, anchored in the constructed timelines, and how time-based document clusters can be used to explore search results that include temporal snippets. We also outline a prototypical implementation and evaluation that demonstrates the feasibility and functionality of our framework. Omar Alonso, Michael Gertz 0001, Ricardo Baeza-Yates |
CIKM | 2 |
| 2009 | Efficiently managing large-scale raster species distribution data in PostgreSQLabstractSpecies distribution data play an important role in biodiversity related research, especially in exploring relationships with the environment. In the recent years, both the number of species being explored and the spatial resolution of species distribution data are increasing fast. It is thus imperative to develop database systems that allow users to efficiently query such large-scale data based on spatial and non-spatial (e.g., taxonomic and phylogenetics) criteria.In this paper, we present our approach to building such a system by integrating several components, including a quadtree representation of binary raster data, tree path indexing and query processing in PostgreSQL, and window decomposition techniques for spatial queries. Our unique contribution is in associating species identifiers with intermediate quadtree nodes and query optimization for multiple independent queries after window query decomposition. Our system enables PostgreSQL to support binary raster data without requiring any changes to the database backend and is suitable for managing large-scale species distribution data.Our experiments using 4000+ bird species distribution data related to the Western hemisphere show that the proposed approach in associating species identifiers with quadtree nodes reduces the number of database tuples by more than 1/3 and the average identifiers to be associated with each tuple from 110.6 to 4.8, a significant improvement compared to classic quadtree-based approaches. With respect to query optimization, optimized queries are 6--9.5 times faster than the baseline queries for average query response times and 5.5--8.3 times faster than the baseline queries for maximum query response times for four query window sizes ranging from 0.1 to 5.0 degrees. Our query optimization techniques thus make the system suitable for many interactive applications for querying and exploring species distribution data. Michael Gertz 0001, Le Gruenwald |
GIS | 2 |
| 2009 | ORDEN: outlier region detection and exploration in sensor networksabstractSensor networks play a central role in applications that monitor variables in geographic areas such as the traffic volume on roads or the temperature in the environment. A key feature users are often interested in when employing such systems is the detection of unusual phenomena, that is, anomalous values measured by the sensors. In this demonstration, we present a system, called ORDEN, that allows for the detection and (visual) exploration of outliers and anomalous events in sensor networks in real-time. In particular, the system constructs outlier regions from anomalous sensor measurements to provide for a comprehensive description of the spatial extent of phenomena of interest. With our system, users can interactively explore displayed outlier regions and investigate the heterogeneity within individual regions using different parameter and threshold settings. Using real-world sensor data streams from different application domains, we demonstrate the effectiveness and utility of our system. Conny Junghans, Michael Gertz 0001 |
SIGMOD Conference | 2 |
| 2009 | Constraint-Based Learning of Distance Functions for Object Trajectories
Michael Gertz 0001 |
SSDBM | 2 |
| 2008 | Expertise identification and visualization from CVSabstractAs software evolves over time, the identification of expertise becomes an important problem. Component ownership and team awareness of such ownership are signals of solid project. Ownership and ownership awareness are also issues in open-source software (OSS) projects. Indeed, the membership in OSS projects is dynamic with team members arriving and leaving. In large open source projects, specialists who know the system very well are considered experts. How can one identify the experts in a project by mining a particular repository like the source code? Have they gotten help from other people? Omar Alonso, Premkumar T. Devanbu, Michael Gertz 0001 |
MSR | 3 |
| 2008 | Real-Time Integration of Geospatial Raster and Point Data Streams
Carlos Rueda, Michael Gertz 0001 |
SSDBM | 2 |
| 2007 | Modeling satellite image streams for change analysisabstractFast detection of changes in environmental remotely sensed data is a major requirement in the Earth sciences, especially in natural disaster related scenarios. As satellite, transmission, and network technologies continue to improve, the real-time stream processing and delivery of geospatial data from remote sensors requires a systematic approach for change analysis and visualization in a streaming fashion. Although various approaches have been formulated to model the inherent spatial-temporal-spectral complexity of remotely sensed satellite data, there are still challenging peculiarities that demand a precise characterization in the context of environmental change detection. Carlos Rueda, Michael Gertz 0001 |
GIS | 2 |
| 2007 | Search results using timeline visualizationsabstractNo abstract available. Omar Alonso, Michael Gertz 0001, Ricardo Baeza-Yates |
SIGIR | 2 |
| 2007 | Modeling and Querying Vague Spatial Objects Using Shapelets
Daniel Zinn, Jim Bosch, Michael Gertz 0001 |
VLDB | 3 |
| 2006 | Optimization of multiple continuous queries over streaming satellite dataabstractRemotely sensed data, in particular satellite imagery, play many important roles in environmental applications. In particular applications that study rapid changes in the environment require frequent access to these data. For continuous data products, users are often interested in formulating continuous queries that deliver results for each incoming image. In the presence of multiple continuous queries, there is clearly an opportunity to share common intermediate data and thus, increase the overall processing speed of the system.Based on the widely used GRASS, this paper describes a system that realizes multiple query processing using two major components. A query optimizer maintains the current set of active continuous queries. Queries are organized into a single processing plan designed to share intermediate results. For each new image from the stream, the optimizer generates an execution plan specific to the active queries. The query executor then rewrites this plan into a set of geospatial processing steps and executes the plan.We detail experiments using data from NOAA's GOES. Continuous queries are defined in a way similar to the OGC WMS query specification. Using predicted query patterns over the visible hemisphere of GOES, experimental results indicate that multiple-query optimized plans can improve performance significantly when compared to queries that are executed separately. Quinn J. Hart, Michael Gertz 0001 |
GIS | 2 |
| 2006 | Clustering of search results using temporal attributesabstractClustering of search results is an important feature in many of today's information retrieval applications. The notion of hit list clustering appears in Web search engines and enterprise search engines as a mechanism that allows users to further explore the coverage of a query. However, there has been little work on exposing temporal attributes for constructing and presentation of clusters. These attributes appear in documents as part of the textual content, e.g., as a date and time token or as a temporal reference in a sentence. In this paper, we outline a model and describe a prototype that shows the main ideas. Omar Alonso, Michael Gertz 0001 |
SIGIR | 2 |
| 2006 | An Extensible Infrastructure for Processing Distributed Geospatial Data StreamsabstractAlthough the processing of data streams has been the focus of many research efforts in several areas, the case of remotely sensed streams in scientific contexts has received little attention. We present an extensible architecture to compose streaming image processing pipelines spanning multiple nodes on a network using a scientific workflow approach. This architecture includes (i) a mechanism for stream query dispatching so new streams can be dynamically generated from within individual processing nodes as a result of local or remote requests, and (ii) a mechanism for making the resulting streams externally available. As complete processing image pipelines can be cascaded across multiple interconnected nodes in a dynamic, scientist-driven way, the approach facilitates the reuse of data and the scalability of computations. We demonstrate the advantages of our infrastructure with a toolset of stream operators acting on remotely sensed data streams for realtime change detection Carlos Rueda, Michael Gertz 0001, Bertram Ludäscher, Bernd Hamann |
SSDBM | 2 |
| 2006 | Integrating document and data retrieval based on XML
Jan-Marco Bremer, Michael Gertz 0001 |
VLDB J. | 2 |
| 2005 | Evaluation of a Dynamic Tree Structure for Indexing Query Regions on Streaming Geospatial Data
Quinn J. Hart, Michael Gertz 0001 |
SSTD | 2 |
| 2005 | Querying Streaming Geospatial Image Data: The GeoStreams Project
Quinn J. Hart, Michael Gertz 0001 |
SSDBM | 2 |
| 2003 | On Distributing XML Repositories
Jan-Marco Bremer, Michael Gertz 0001 |
WebDB | 2 |
| 2002 | Reverse Engineering for Web Data: From Visual to Semantic StructureabstractDespite the advancement of XML, the majority of documents on the Web is still marked up with HTML for visual rendering purposes only, thus building a huge amount of legacy data. In order to facilitate querying Web based data in a way more efficient and effective than just keyword based retrieval, enriching such Web documents with both structure and semantics is necessary. We describe a novel approach to the integration of topic specific HTML documents into a repository of XML documents. In particular, we describe how topic specific HTML documents are transformed into XML documents. The proposed document transformation and semantic element tagging process utilizes document restructuring rules and minimum information about the topic in the form of concepts. For the resulting XML documents, a majority schema is derived that describes common structures among the documents in the form of a DTD. We explore and discuss different techniques, and rules for document conversion and majority schema discovery. We finally demonstrate the feasibility and effectiveness of our approach by applying it to a set of resume HTML documents gathered by a Web crawler. Christina Yip Chung, Michael Gertz 0001, Neel Sundaresan |
ICDE | 2 |
| 2002 | Annotating Scientific Images: A Concept-Based ApproachabstractData annotations are an important kind of metadata that occur in the form of externally assigned descriptions of particular features in Web accessible documents. Such metadata are eventually used in data retrieval tasks on heterogeneous, possible distributed Web-accessible documents. In this paper, we present the model and realization of an annotation framework that scientists can employ to semantically enrich different types of documents, primarily scientific images made available through an image repository. Although we employ ontology like structures, called concepts, for metadata schemes used in annotations, our primary focus is on how concepts are actually used to annotate images and regions of interest, respectively, that exhibit features of interest to a researcher. It turns out that the combined consideration of domain specific concepts and annotated regions in images provides interesting means to analyze the usage of metadata regarding certain correctness and plausibility criteria. We detail our annotation management framework in the context of the Human Brain Project in which Neuroscientists record their observations on specific brain structures, and share and exchange information through concept-based annotations associated with images. Michael Gertz 0001, Kai-Uwe Sattler, Fredric Gorin, Michael A. Hogarth, Jim Stone |
SSDBM | 1 |
| 2002 | XQuery/IR: Integrating XML Document and Data Retrieval
Jan-Marco Bremer, Michael Gertz 0001 |
WebDB | 2 |
| 2001 | Quixote: Building XML Repositories from Topic Specific Web Documents
Christina Yip Chung, Michael Gertz 0001, Neel Sundaresan |
WebDB | 2 |
| 2001 | Authentic Publication of XML Document DataabstractWith XML becoming a major standard for the exchange and sharing of data, many applications now rely on XML to publish and utilize Web data. Data sources (data owners), however often do not have the means to provide high performance query mechanisms, serving hundreds of thousands of client queries per day. Furthermore, it is too costly for them to maintain a secure Web information system infrastructure preventing intruders from tampering with the integrity of the data. As a solution to this problem, data owners can distribute XML document data to data publishers who then handle queries from clients on behalf of the data owners. In this paper, we present a novel authentic publication scheme for XML document data that allows clients to efficiently verify query results from data publishers for correctness and completeness. For this, data owners simply provide summary signatures of their data to clients who then use the signatures in addition to verification objects provided by a publisher to authenticate query results. Our approach supports two of the most common forms of XML queries, path and selection queries. The approach is reliable where no trust from publishers is required, and is practical with a reasonably small overhead, both in terms of time and size of the storage. April Kwong, Michael Gertz 0001 |
WISE (1) | 2 |
| 2001 | Semantic integrity support in SQL: 1999 and commercial (object-)relational database management systems
Can Türker, Michael Gertz 0001 |
VLDB J. | 2 |
| 1996 | Deriving Optimized Integrity Monitoring Triggers from Dynamic Integrity Constraints
Michael Gertz 0001, Udo W. Lipeck |
Data Knowl. Eng. | 1 |
| 1993 | Deriving Integrity Maintaining Triggers from Transition GraphsabstractMethods for deriving constraint maintaining triggers from dynamic integrity constraints represented by transition graphs are presented. The methods reduce integrity monitoring to checking changing static conditions according to life cycle situations. Thus, triggers have to be generated from these graphs, which depend not only on the operations that have occurred in a transaction, but also on the situations that have been reached by the objects mentioned in the constraints. The techniques presented work for dynamic constraints and their corresponding transition graphs as well as for simple static constraints. Only passive reactions (rollbacks) to constraint violations are provided by the trigger patterns, but the systematic generation of such patterns should help the database designer in identifying possible active reactions for repairing constraint violations.> Michael Gertz 0001, Udo W. Lipeck |
ICDE | 1 |