Alexander Markowetz

dblp:m/AlexanderMarkowetz · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
1since 2021 · last 2024
0000-0002-2621-722XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 4 first-authorArtificial intelligence and machine learning · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Graph data management · 24% Query processing and optimization · 19% Data stream processing · 19%
Network and information security
1 paper
Privacy and data protection · 100%
Human-computer interaction and pervasive computing
1 paper
Ubiquitous computing and smart environments · 100%

Topics — the 11 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data stream processing
continuous query processing
0.112009
Keyword search over relational tables and streams · ACM Trans. Database Syst. 2009
Graph data management
graph indexing
0.112009
Reachability Indexes for Relational Keyword Search · ICDE 2009
Information retrieval › keyword search
keyword search over relational databases
0.112009
Keyword search over relational tables and streams · ACM Trans. Database Syst. 2009
Graph data management › graph indexing
reachability indexing
0.112009
Reachability Indexes for Relational Keyword Search · ICDE 2009
Query processing and optimization › keyword query processing
relational keyword search
0.112009
Reachability Indexes for Relational Keyword Search · ICDE 2009
Ubiquitous computing and smart environments › mobile computing
mobile app usage analysis
0.112016
Differentiating smartphone users by app usage · UbiComp 2016
Ubiquitous computing and smart environments
mobile computing
0.112016
Differentiating smartphone users by app usage · UbiComp 2016
Data stream processing › continuous query processing
keyword search on data streams
0.112007
Keyword search on relational data streams · SIGMOD Conference 2007
Spatial and temporal data management › temporal databases
temporal aggregation
0.012001
Efficient Computation of Temporal Aggregates with Range Predicates · PODS 2001
Indexing and storage engines
multiversion index
0.012008
On computing temporal aggregates with range predicates · ACM Trans. Database Syst. 2008
Spatial and temporal data management
spatial query processing
0.012006
Efficient query processing in geographic web search engines · SIGMOD Conference 2006

Methods — techniques the papers use, named apart from their topics

hamming distance analysis · 0.5behavioral fingerprinting · 0.5operator-based query processing · 0.1join reachability indexing · 0.1graph-based query processing · 0.1multiversion b+-tree · 0.1SB-tree · 0.1schema adaptation · 0.1operator mesh · 0.1text and spatial indexing · 0.1
YearPublicationVenuePosition
2024 DeFaktS: A German Dataset for Fine-Grained Disinformation Detection through Social Media Framing
abstract
In today’s rapidly evolving digital age, disinformation poses a significant threat to public sentiment and socio-political dynamics. To address this, we introduce a new dataset “DeFaktS”, designed to understand and counter disinformation within German media. Distinctively curated across various news topics, DeFaktS offers an unparalleled insight into the diverse facets of disinformation. Our dataset, containing 105,855 posts with 20,008 meticulously labeled tweets, serves as a rich platform for in-depth exploration of disinformation’s diverse characteristics. A key attribute that sets DeFaktS apart is, its fine-grain annotations based on polarized categories. Our annotation framework, grounded in the textual characteristics of news content, eliminates the need for external knowledge sources. Unlike most existing corpora that typically assign a singular global veracity value to news, our methodology seeks to annotate every structural component and semantic element of a news piece, ensuring a comprehensive and detailed understanding. In our experiments, we employed a mix of classical machine learning and advanced transformer-based models. The results underscored the potential of DeFaktS, with transformer models, especially the German variant of BERT, exhibiting pronounced effectiveness in both binary and fine-grained classifications.
Shaina Ashraf, Isabel Bezzaoui, Ionut Andone, Alexander Markowetz, Jonas Fegert, Lucie Flek
LREC/COLING4
2017 Impact of location-based games on phone usage and movement: a case study on Pokémon GO
abstract
Pokémon GO was a short lived mobile location-based gaming phenomenon. After its launch in July 2016, it quickly reached 500 million installs, but afterwards interest faded. As part of a large scale "in the wild" mobile phone study we have recorded phone usage and location measurements between June and September 2016. We investigate who were the people who installed and played Pokémon GO and what effects it had on their behaviour. We chose as a middle point the start of playing the game, and selected users that had activity for at least two weeks before and two weeks after it. In this work we present our findings on a sample of 2, 861 users. We compare demographic characteristics and Big Five personality traits of these users with 7, 904 non-playing users from the same time period. The general daily phone usage of players increased on average by 27 minutes, which represents 16% per day. In terms of large scale movement patterns, these did not change, with regard to diameter and total path length per day.
Ionut Andone, Konrad Blaszkiewicz, Matthias Böhmer 0001, Alexander Markowetz
MobileHCI4
2016 Three-hop distance estimation in social graphs
abstract
In this paper, we study a 3-hop approach to distance estimation that uses two intermediate landmarks, where each landmark only stores distances to vertices in its local neighborhood and to the other landmarks. We show how to suitably represent and compress the distance data stored for each landmark, for the 2-hop and 3-hop case. Overall, we find that 3-hop methods achieve modest but promising improvement in some cases, while being comparable or slightly worse than 2-hop methods in others. Furthermore, our light compression schemes improve the practical applicability of both the 2-hop and 3-hop methods.
Pascal Welke, Alexander Markowetz, Torsten Suel, Maria Christoforaki
IEEE BigData2
2016 Differentiating smartphone users by app usage
abstract
Tracking users across websites and apps is as desirable to the marketing industry as it is unalluring to users. The central challenge lies in identifying users from the perspective of different apps/sites. While there are methods to identify users via technical settings of their phones, these are prone to countermeasures. Yet, in this paper, we show that it is possible to differentiate users via their set of used apps, their app signature. To this end, we investigate the app usage of 46726 participants from the Menthal project. Even limiting our observation to the 500 globally most frequent apps results in unique signatures for 99.67% of users. Furthermore, even under this restriction, the average minimum Hamming distance to the closest other user is 25.93. Avoiding identification would thus require a massive change in the behavior of a user. Indeed, 99.4% of all users have unique usage patterns among the top 60 globally used apps. In contrast to previous work, this paper differentiates between users based on behavior instead of technical parameters. It thus opens an entirely new discussion regarding privacy.
Pascal Welke, Ionut Andone, Konrad Blaszkiewicz, Alexander Markowetz
UbiComp4
2011 Text vs. space: efficient geo-search query processing
abstract
Many web search services allow users to constrain text queries to a geographic location (e.g., yoga classes near Santa Monica). Important examples include local search engines such as Google Local and location-based search services for smart phones. Several research groups have studied the efficient execution of queries mixing text and geography; their approaches usually combine inverted lists with a spatial access method such as an R-tree or space-filling curve. In this paper, we take a fresh look at this problem. We feel that previous work has often focused on the spatial aspect at the expense of performance considerations in text processing, such as inverted index access, compression, and caching. We describe new and existing approaches and discuss their different perspectives. We then compare their performance in extensive experiments on large document collections. Our results indicate that a query processor that combines state-of-the-art text processing techniques with a simple coarse-grained spatial structure can outperform existing approaches by up to two orders of magnitude. In fact, even a naive approach that first uses a simple inverted index and then filters out any documents outside the query range outperforms many previous methods.
Maria Christoforaki, Jinru He, Constantinos Dimopoulos, Alexander Markowetz, Torsten Suel
CIKM4
2009 Reachability Indexes for Relational Keyword Search
abstract
Due to its considerable ease of use, relational keyword search (R-KWS) has become increasingly popular. Its simplicity, however, comes at the cost of intensive query processing. Specifically, R-KWS explores a vast search space, comprised of all possible combinations of keyword occurrences in any attribute of every table. Existing systems follow two general methodologies for query processing: (i) graph based, which traverses a materialized data graph, and (ii) operator based, which executes relational operator trees on an underlying DBMS. In both cases, computations are largely wasted on graph traversals or operator tree executions that fail to return results. Motivated by this observation, we introduce a comprehensive framework for reachability indexing that eliminates such fruitless operations. We describe a range of indexes that capture various types of join reachability. Extensive experiments demonstrate that the proposed techniques significantly improve performance, often by several orders of magnitude.
Alexander Markowetz, Yin Yang 0001, Dimitris Papadias
ICDE1
2009 Keyword search over relational tables and streams
abstract
Relational Keyword Search (R-KWS) provides an intuitive way to query relational data without requiring SQL, or knowledge of the underlying schema. In this article we describe a comprehensive framework for R-KWS covering snapshot queries on conventional tables and continuous queries on relational streams. Our contributions are summarized as follows: (i) We provide formal semantics, addressing the temporal validity and order of results, spanning uniformly over tables and streams; (ii) we investigate two general methodologies for query processing, graph based and operator based , that resolve several problems of previous approaches; and (iii) we develop a range of algorithms and optimizations covering both methodologies. We demonstrate the effectiveness of R-KWS, as well as the significant performance benefits of the proposed techniques, through extensive experiments with static and streaming datasets.
Alexander Markowetz, Yin Yang 0001, Dimitris Papadias
ACM Trans. Database Syst.1
2008 On computing temporal aggregates with range predicates
abstract
Computing temporal aggregates is an important but costly operation for applications that maintain time-evolving data (data warehouses, temporal databases, etc.) Due to the large volume of such data, performance improvements for temporal aggregate queries are critical. Previous approaches have aggregate predicates that involve only the time dimension. In this article we examine techniques to compute temporal aggregates that include key-range predicates as well ( range-temporal aggregates ). In particular we concentrate on the SUM aggregate, while COUNT is a special case. To handle arbitrary key ranges, previous methods would need to keep a separate index for every possible key range. We propose an approach based on a new index structure called the Multiversion SB-Tree , which incorporates features from both the SB-Tree and the Multiversion B+--tree, to handle arbitrary key-range temporal aggregate queries. We analyze the performance of our approach and present experimental results that show its efficiency. Furthermore, we address a novel and practical variation called functional range-temporal aggregates. Here, the value of any record is a function over time. The meaning of aggregates is altered such that the contribution of a record to the aggregate result is proportional to the size of the intersection between the record's time interval and the query time interval. Both analytical and experimental results show the efficiency of our result.
Alexander Markowetz, Vassilis J. Tsotras, Dimitrios Gunopulos, Bernhard Seeger
ACM Trans. Database Syst.2
2007 Keyword search on relational data streams
abstract
Increasing monitoring of transactions, environmental parameters, homeland security, RFID chips and interactions of online users rapidly establishes new data sources and application scenarios. In this paper, we propose keyword search on relational data streams (S-KWS) as an effective way for querying in such intricate and dynamic environments. Compared to conventional query methods, S-KWS has several benefits. First, it allows search for combinations of interesting terms without a-priori knowledge of the data streams in which they appear. Second, it hides the schema from the user and allows it to change, without the need for query re-writing. Finally, keyword queries are easy to express. Our contributions are summarized as follows. (i) We provide formal semantics for S-KWS, addressing the temporal validity and order of results. (ii) We propose an efficient algorithm for generating operator trees, applicable to arbitrary schemas. (iii) We integrate these trees into an operator mesh that shares common expressions. (iv) We develop techniques that utilize the operator mesh for efficient query processing. The techniques adapt dynamically to changes in the schema and input characteristics. Finally, (v) we present methods for purging expired tuples, minimizing either CPU, or memory requirements.
Alexander Markowetz, Yin Yang 0001, Dimitris Papadias
SIGMOD Conference1
2006 Efficient query processing in geographic web search engines
abstract
Geographic web search engines allow users to constrain and order search results in an intuitive manner by focusing a query on a particular geographic region. Geographic search technology, also called local search, has recently received significant interest from major search engine companies. Academic research in this area has focused primarily on techniques for extracting geographic knowledge from the web. In this paper, we study the problem of efficient query processing in scalable geographic search engines. Query processing is a major bottleneck in standard web search engines, and the main reason for the thousands of machines used by the major engines. Geographic search engine query processing is different in that it requires a combination of text and spatial data processing techniques. We propose several algorithms for efficient query processing in geographic search engines, integrate them into an existing web search query processor, and evaluate them on large sets of real data and query traces.
Torsten Suel, Alexander Markowetz
SIGMOD Conference3
2005 Design and Implementation of a Geographic Search Engine
Alexander Markowetz, Torsten Suel, Xiaohui Long, Bernhard Seeger
WebDB1
2001 Efficient Computation of Temporal Aggregates with Range Predicates
abstract
A temporal aggregation query is an important but costly operation for applications that maintain time-evolving data (data warehouses, temporal databases, etc.). Due to the large volume of such data, performance improvements for temporal aggregation queries are critical. In this paper we examine techniques to compute temporal aggregates that include key-range predicates (range temporal aggregates). In particular we concentrate on SUM, COUNT and AVG aggregates. This problem is novel; to handle arbitrary key ranges, previous methods would need to keep a separate index for every possible key range. We propose an approach based on a new index structure called the Multiversion SB-Tree, which incorporates features from both the SB-Tree and the Multiversion B-Tree, to handle arbitrary key-range temporal SUM, COUNT and AVG queries. We analyze the performance of our approach and present experimental results that show its efficiency.
Alexander Markowetz, Vassilis J. Tsotras, Dimitrios Gunopulos, Bernhard Seeger
PODS2