Karl Aberer

dblp:a/KarlAberer · DBLP profile ↗
← Back
145ranked-venue papers in the field
17as first author
17since 2021 · last 2025
0000-0003-3005-7342ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 72 (11 first)Information Retrieval & Web Search · 41 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 10 (2 first)Data Mining & Knowledge Discovery · 8Big Data, Cloud & Distributed Data Systems · 7Other / Interdisciplinary · 4Business Process & Enterprise Data · 3 (1 first)
YearPublicationVenuePosition
2025 SAGESSE: A System for Argument Generation, Extraction and Structuring of Social Exchanges
abstract
Online debates provide critical insights into public opinion and societal trends, yet the unstructured nature of these discussions presents significant challenges for analysis. In this paper, we present SAGESSE, a novel argumentation parsing pipeline tailored to Reddit debates, leveraging the capabilities of large language models to structure and interpret complex online arguments. SAGESSE generates detailed argument maps that organize debates systematically, offering a clearer understanding of discourse dynamics. We have developed a web application where users can select controversial Reddit topics and visualize the corresponding argument maps generated from user comments. This tool has the potential to aid analysts, policymakers, and researchers in tracking debate progress, gauging public sentiment, and identifying influential arguments. The web application is available at https://modemos.epfl.ch/sagesse/ .
Nicolas Almerge, Matteo Santelmo, Ilker Gül, Amin Asadi Sarijalou, Rémi Lebret, Léo Laugier, Karl Aberer
WSDM7
2024 Fast-FedUL: A Training-Free Federated Unlearning with Provable Skew Resilience
Trong Bang Nguyen, Phi-Le Nguyen, Thanh Tam Nguyen, Matthias Weidlich 0001, Nguyen Quoc Viet Hung, Karl Aberer
ECML/PKDD (5)7
2023 Efficient and Effective Multi-Modal Queries through Heterogeneous Network Embedding (Extended Abstract)
abstract
Recent information retrieval (IR) systems answer a multi-modal query by considering it as a set of separate uni-modal queries. However, depending on the chosen operationalisation, such an approach is inefficient or ineffective. It either requires multiple passes over the data or leads to inaccuracies since the relations between data modalities are neglected in the relevance assessment. To mitigate these challenges, we present an IR system that has been designed to answer genuine multi-modal queries. It relies on a heterogeneous network embedding, so that features from diverse modalities can be incorporated when representing both, a query and the data over which it shall be evaluated. An experimental evaluation using diverse real-world and synthetic datasets illustrates that our approach returns twice the amount of relevant information compared to baseline techniques, while scaling to large multi-modal databases.
Thanh Tam Nguyen, Chi Thang Duong, Hongzhi Yin, Matthias Weidlich 0001, Son T. Mai, Karl Aberer, Nguyen Quoc Viet Hung
ICDE6
2023 Misleading Repurposing on Twitter
abstract
We present the first in-depth and large-scale study of misleading repurposing, in which a malicious user changes the identity of their social media account via, among other things, changes to the profile attributes in order to use the account for a new purpose while retaining their followers. We propose a definition for the behavior and a methodology that uses supervised learning on data mined from the Internet Archive's Twitter Stream Grab to flag repurposed accounts. We found over 100,000 accounts that may have been repurposed. Of those, 28% were removed from the platform after 2 years, thereby confirming their inauthenticity. We also characterize repurposed accounts and found that they are more likely to be repurposed after a period of inactivity and deleting old tweets. We also provide evidence that adversaries target accounts with high follower counts to repurpose, and some make them have high follower counts by participating in follow-back schemes. The results we present have implications for the security and integrity of social media platforms, for data science studies in how historical data is considered, and for society at large in how users can be deceived about the popularity of an opinion. The data and the code is available at https://github.com/tugrulz/MisleadingRepurposing.
Tugrulcan Elmas, Rebekah Overdorf, Karl Aberer
ICWSM3
2023 SciLander: Mapping the Scientific News Landscape
abstract
The COVID-19 pandemic has fueled the spread of misinformation on social media and the Web as a whole. The phenomenon dubbed `infodemic' has taken the challenges of information veracity and trust to new heights by massively introducing seemingly scientific and technical elements into misleading content. Despite the existing body of work on modeling and predicting misinformation, the coverage of very complex scientific topics with inherent uncertainty and an evolving set of findings, such as COVID-19, provides many new challenges that are not easily solved by existing tools. To address these issues, we introduce SciLander, a method for learning representations of news sources reporting on science-based topics. We extract four heterogeneous indicators for the sources; two generic indicators that capture (1) the copying of news stories between sources, and (2) the use of the same terms to mean different things (semantic shift), and two scientific indicators that capture (1) the usage of jargon and (2) the stance towards specific citations. We use these indicators as signals of source agreement, sampling pairs of positive (similar) and negative (dissimilar) samples, and combine them in a unified framework to train unsupervised news source embeddings with a triplet margin loss objective. We evaluate our method on a novel COVID-19 dataset containing nearly 1M news articles from 500 sources spanning a period of 18 months since the beginning of the pandemic in 2020. Our results show that the features learned by our model outperform state-of-the-art baseline methods on the task of news veracity classification. Furthermore, a clustering analysis suggests that the learned representations encode information about the reliability, political leaning, and partisanship bias of these sources.
Maurício Gruppi, Panayiotis Smeros, Sibel Adali, Carlos Castillo 0001, Karl Aberer
ICWSM5
2023 Firearms on Twitter: A Novel Object Detection Pipeline
abstract
Social media is an important source of real-time imagery concerning world events. One subset of social media posts which may be of particular interest are those featuring firearms. These posts can give insight into weapon movements, troop activity and civilian safety. Object detection tools offer important opportunities for insight into these images. Unfortunately, these images can be visually complex, poorly lit and generally challenging for object detection models. We present an analysis of existing gun detection datasets, and find that these datasets to not effectively address the challenge of gun detection on real-life images. Following this, we present a novel object detection pipeline. We train our pipeline on a number of datasets including one created for this investigation made up of Twitter images of the Russo-Ukrainian War. We compare the performance of our model as trained on the different datasets to baseline numbers provided by original authors as well as a YOLO v5 benchmark. We find that our model outperforms the state-of-the-art benchmarks on contextually rich, real-life-derived imagery of firearms.
Ryan Harvey, Rémi Lebret, Stéphane Massonnet, Karl Aberer, Gianluca Demartini
ICWSM4
2023 Efficient Integration of Multi-Order Dynamics and Internal Dynamics in Stock Movement Prediction
abstract
Advances in deep neural network (DNN) architectures have enabled new prediction techniques for stock market data. Unlike other multivariate time-series data, stock markets show two unique characteristics: (i) multi-order dynamics, as stock prices are affected by strong non-pairwise correlations (e.g., within the same industry); and (ii) internal dynamics, as each individual stock shows some particular behaviour. Recent DNN-based methods capture multi-order dynamics using hypergraphs, but rely on the Fourier basis in the convolution, which is both inefficient and ineffective. In addition, they largely ignore internal dynamics by adopting the same model for each stock, which implies a severe information loss.
Minh Hieu Nguyen 0003, Thanh Tam Nguyen, Phi-Le Nguyen, Matthias Weidlich 0001, Nguyen Quoc Viet Hung, Karl Aberer
WSDM7
2023 Scalable maximal subgraph mining with backbone-preserving graph convolutions
abstract
Maximal subgraph mining is increasingly important in various domains, including bioinformatics, genomics, and chemistry, as it helps identify common characteristics among a set of graphs and enables their classification into different categories. Existing approaches for identifying maximal subgraphs typically rely on traversing a graph lattice. However, in practice, these approaches are limited to relatively small subgraphs due to the exponential growth of the search space and the NP-completeness of the underlying subgraph isomorphism test. In this work, we propose SCAMA , an approach that addresses these limitations by adopting a divide-and-conquer strategy for efficient mining of maximal subgraphs. Our approach involves initially partitioning a graph database into equivalence classes using bootstrapped backbones, which are tree-shaped frequent subgraphs. We then introduce a learning process based on a novel graph convolutional network (GCN) to extract maximal backbones for each equivalence class. A critical insight of our approach is that by estimating each maximal backbone directly in the embedding space, we can avoid the exponential traversal of the graph lattice. From the extracted maximal backbones, we construct the maximal frequent subgraphs. Furthermore, we outline how SCAMA can be extended to perform top- k largest frequent subgraph mining and how the discovered patterns facilitate graph classification. Our experimental results demonstrate the effectiveness of SCAMA in identifying almost perfectly maximal frequent subgraphs, while exhibiting approximately 10 times faster performance compared to the best baseline technique.
Matthias Weidlich 0001, Thanh Tho Quan, Hongzhi Yin, Karl Aberer, Nguyen Quoc Viet Hung
Inf. Sci.6
2022 WayPop Machine: A Wayback Machine to Investigate Popularity and Root Out Trolls
abstract
Contrary to celebrities who owe their popularity online to their activity offline, malicious users such as trolls have to gain fame on social media through the social media itself. The exact reasons that a certain user has become popular are often obscure especially when the popularity was gained illicitly through means such as fake amplification of content. In this paper, we develop a methodology for uncovering why an account has become popular and present an open source tool that encapsulates this methodology. This tool aims to aid others in uncovering malicious accounts which have artificially gained many followers and to distinguish such accounts from those which gained followers and popularity honestly.
Tugrulcan Elmas, Thomas Romain Ibanez, Alexandre Hutter, Rebekah Overdorf, Karl Aberer
ASONAM5
2022 Characterizing Retweet Bots: The Case of Black Market Accounts
Tugrulcan Elmas, Rebekah Overdorf, Karl Aberer
ICWSM3
2022 Preface - Special Issue on Misinformation on the Web
Karl Aberer, Ioannis Katakis 0001, Nguyen Quoc Viet Hung, Hongzhi Yin
Inf. Syst.1
2022 Efficient and Effective Multi-Modal Queries Through Heterogeneous Network Embedding
abstract
The heterogeneity of today’s Web sources requires information retrieval (IR) systems to handle multi-modal queries. Such queries define a user’s information needs by different data modalities, such as keywords, hashtags, user profiles, and other media. Recent IR systems answer such a multi-modal query by considering it as a set of separate uni-modal queries. However, depending on the chosen operationalisation, such an approach is inefficient or ineffective. It either requires multiple passes over the data or leads to inaccuracies since the relations between data modalities are neglected in the relevance assessment. To mitigate these challenges, we present an IR system that has been designed to answer genuine multi-modal queries. It relies on a heterogeneous network embedding, so that features from diverse modalities can be incorporated when representing both, a query and the data over which it shall be evaluated. By embedding a query and the data in the same vector space, the relations across modalities are made explicit and exploited for more accurate query evaluation. At the same time, multi-modal queries are answered with a single pass over the data. An experimental evaluation using diverse real-world and synthetic datasets illustrates that our approach returns twice the amount of relevant information compared to baseline techniques, while scaling to large multi-modal databases.
Chi Thang Duong, Thanh Tam Nguyen, Hongzhi Yin, Matthias Weidlich 0001, Son T. Mai, Karl Aberer, Nguyen Quoc Viet Hung
IEEE Trans. Knowl. Data Eng.6
2021 SciClops: Detecting and Contextualizing Scientific Claims for Assisting Manual Fact-Checking
abstract
This paper describes SciClops, a method to help combat online scientific misinformation. Although automated fact-checking methods have gained significant attention recently, they require pre-existing ground-truth evidence, which, in the scientific context, is sparse and scattered across a constantly-evolving scientific literature. Existing methods do not exploit this literature, which can effectively contextualize and combat science-related fallacies. Furthermore, these methods rarely require human intervention, which is essential for the convoluted and critical domain of scientific misinformation.
Panayiotis Smeros, Carlos Castillo 0001, Karl Aberer
CIKM3
2021 A Dataset of State-Censored Tweets
Tugrulcan Elmas, Rebekah Overdorf, Karl Aberer
ICWSM3
2021 Recommendation on Live-Streaming Platforms: Dynamic Availability and Repeat Consumption
abstract
Live-streaming platforms broadcast user-generated video in real-time. Recommendation on these platforms shares similarities with traditional settings, such as a large volume of heterogeneous content and highly skewed interaction distributions. However, several challenges must be overcome to adapt recommendation algorithms to live-streaming platforms: first, content availability is dynamic which restricts users to choose from only a subset of items at any given time; during training and inference we must carefully handle this factor in order to properly account for such signals, where ‘non-interactions’ reflect availability as much as implicit preference. Streamers are also fundamentally different from ‘items’ in traditional settings: repeat consumption of specific channels plays a significant role, though the content itself is fundamentally ephemeral.
Jérémie Rappaz, Julian J. McAuley, Karl Aberer
RecSys3
2021 Efficient Streaming Subgraph Isomorphism with Graph Neural Networks
abstract
Queries to detect isomorphic subgraphs are important in graph-based data management. While the problem of subgraph isomorphism search has received considerable attention for the static setting of a single query, or a batch thereof, existing approaches do not scale to a dynamic setting of a continuous stream of queries. In this paper, we address the scalability challenges induced by a stream of subgraph isomorphism queries by caching and re-use of previous results. We first present a novel subgraph index based on graph embeddings that serves as the foundation for efficient stream processing. It enables not only effective caching and re-use of results, but also speeds-up traditional algorithms for subgraph isomorphism in case of cache misses. Moreover, we propose cache management policies that incorporate notions of reusability of query results. Experiments using real-world datasets demonstrate the effectiveness of our approach in handling isomorphic subgraph search for streams of queries.
Chi Thang Duong, Dung Hoang, Hongzhi Yin, Matthias Weidlich 0001, Nguyen Quoc Viet Hung, Karl Aberer
Proc. VLDB Endow.6
2021 Scalable Robust Graph Embedding with Spark
abstract
Graph embedding aims at learning a vector-based representation of vertices that incorporates the structure of the graph. This representation then enables inference of graph properties. Existing graph embedding techniques, however, do not scale well to large graphs. While several techniques to scale graph embedding using compute clusters have been proposed, they require continuous communication between the compute nodes and cannot handle node failure. We therefore propose a framework for scalable and robust graph embedding based on the MapReduce model, which can distribute any existing embedding technique. Our method splits a graph into subgraphs to learn their embeddings in isolation and subsequently reconciles the embedding spaces derived for the subgraphs. We realize this idea through a novel distributed graph decomposition algorithm. In addition, we show how to implement our framework in Spark to enable efficient learning of effective embeddings. Experimental results illustrate that our approach scales well, while largely maintaining the embedding quality.
Chi Thang Duong, Dung Hoang, Hongzhi Yin, Matthias Weidlich 0001, Nguyen Quoc Viet Hung, Karl Aberer
Proc. VLDB Endow.6
2020 Graph Embeddings for One-pass Processing of Heterogeneous Queries
abstract
Effective information retrieval (IR) relies on the ability to comprehensively capture a user's information needs. Traditional IR systems are limited to homogeneous queries that define the information to retrieve by a single modality. Support for heterogeneous queries that combine different modalities has been proposed recently. Yet, existing approaches for heterogeneous querying are computationally expensive, as they require several passes over the data to construct a query answer.In this paper, we propose an IR system that overcomes the computational challenges imposed by heterogeneous queries by adopting graph embeddings. Specifically, we propose graph-based models in which both, data and queries, incorporate information of different modalities. Then, we show how either representation is transformed into a graph embedding in the same space, capturing relations between information of different modalities. By grounding query processing in graph embeddings, we enable processing of heterogeneous queries with a single pass over the data representation. Our experiments on several real-world and synthetic datasets illustrate that our technique is able to return twice the amount of relevant information in comparison with several baselines, while being scalable to large-scale data.
Chi Thang Duong, Hongzhi Yin, Dung Hoang, Minn Hung Nguyen, Matthias Weidlich 0001, Nguyen Quoc Viet Hung, Karl Aberer
ICDE7
2020 SciLens News Platform: A System for Real-Time Evaluation of News Articles
abstract
We demonstrate the SciLens News Platform, a novel system for evaluating the quality of news articles. The SciLens News Platform automatically collects contextual information about news articles in real-time and provides quality indicators about their validity and trustworthiness. These quality indicators derive from i) social media discussions regarding news articles, showcasing the reach and stance towards these articles, and ii) their content and their referenced sources, showcasing the journalistic foundations of these articles. Furthermore, the platform enables domain-experts to review articles and rate the quality of news sources. This augmented view of news articles, which combines automatically extracted indicators and domain-expert reviews, has provably helped the platform users to have a better consensus about the quality of the underlying articles. The platform is built in a distributed and robust fashion and runs operationally handling daily thousands of news articles. We evaluate the SciLens News Platform on the emerging topic of COVID-19 where we highlight the discrepancies between low and high-quality news outlets based on three axes, namely their newsroom activity, evidence seeking and social engagement. A live demonstration of the platform can be found here: http://scilens.epfl.ch.
Angelika Romanou, Panayiotis Smeros, Carlos Castillo 0001, Karl Aberer
Proc. VLDB Endow.4
2019 A Dynamic Embedding Model of the Media Landscape
abstract
Information about world events is disseminated through a wide variety of news channels, each with specific considerations in the choice of their reporting. Although the multiplicity of these outlets should ensure a variety of viewpoints, recent reports suggest that the rising concentration of media ownership may void this assumption. This observation motivates the study of the impact of ownership on the global media landscape and its influence on the coverage the actual viewer receives. To this end, the selection of reported events has been shown to be informative about the high-level structure of the news ecosystem. However, existing methods only provide a static view into an inherently dynamic system, providing underperforming statistical models and hindering our understanding of the media landscape as a whole.
Jérémie Rappaz, Dylan Bourgeois, Karl Aberer
WWW3
2019 SciLens: Evaluating the Quality of Scientific News Articles Using Social Media and Scientific Literature Indicators
abstract
This paper describes, develops, and validates SciLens, a method to evaluate the quality of scientific news articles. The starting point for our work are structured methodologies that define a series of quality aspects for manually evaluating news. Based on these aspects, we describe a series of indicators of news quality. According to our experiments, these indicators help non-experts evaluate more accurately the quality of a scientific news article, compared to non-experts that do not have access to these indicators. Furthermore, SciLens can also be used to produce a completely automated quality score for an article, which agrees more with expert evaluators than manual evaluations done by non-experts. One of the main elements of SciLens is the focus on both content and context of articles, where context is provided by (1) explicit and implicit references on the article to scientific literature, and (2) reactions in social media referencing the article. We show that both contextual elements can be valuable sources of information for determining article quality. The validation of SciLens, done through a combination of expert and non-expert annotation, demonstrates its effectiveness for both semi-automatic and automatic quality evaluation of scientific news.
Panayiotis Smeros, Carlos Castillo 0001, Karl Aberer
WWW3
2019 Handling probabilistic integrity constraints in pay-as-you-go reconciliation of data models
abstract
Data models capture the structure and characteristic properties of data entities, e.g., in terms of a database schema or an ontology. They are the backbone of diverse applications, reaching from information integration , through peer-to-peer systems and electronic commerce to social networking . Many of these applications involve models of diverse data sources. Effective utilisation and evolution of data models, therefore, calls for matching techniques that generate correspondences between their elements. Various such matching tools have been developed in the past. Yet, their results are often incomplete or erroneous, and thus need to be reconciled, i.e., validated by an expert. This paper analyses the reconciliation process in the presence of large collections of data models, where the network induced by generated correspondences shall meet consistency expectations in terms of integrity constraints. We specifically focus on how to handle data models that show some internal structure and potentially differ in terms of their assumed level of abstraction. We argue that such a setting calls for a probabilistic model of integrity constraints, for which satisfaction is preferred, but not required. In this work, we present a model for probabilistic constraints that enables reasoning on the correctness of individual correspondences within a network of data models, in order to guide an expert in the validation process. To support pay-as-you-go reconciliation, we also show how to construct a set of high-quality correspondences, even if an expert validates only a subset of all generated correspondences. We demonstrate the efficiency of our techniques for real-world datasets comprising database schemas and ontologies from various application domains.
Nguyen Quoc Viet Hung, Matthias Weidlich 0001, Thanh Tam Nguyen, Zoltán Miklós 0001, Karl Aberer, Avigdor Gal, Bela Stantic
Inf. Syst.5
2018 Latent Structure in Collaboration: The Case of Reddit r/place
Jérémie Rappaz, Michele Catasta, Robert West 0001, Karl Aberer
ICWSM4
2017 Taxonomy Induction Using Hypernym Subsequences
abstract
We propose a novel, semi-supervised approach towards domain taxonomy induction from an input vocabulary of seed terms. Unlike all previous approaches, which typically extract direct hypernym edges for terms, our approach utilizes a novel probabilistic framework to extract hypernym subsequences. Taxonomy induction from extracted subsequences is cast as an instance of the minimum-cost flow problem on a carefully designed directed graph. Through experiments, we demonstrate that our approach outperforms state-of-the-art taxonomy induction approaches across four languages. Importantly, we also show that our approach is robust to the presence of noise in the input vocabulary. To the best of our knowledge, this robustness has not been empirically proven in any previous approach.
Rémi Lebret, Hamza Harkous, Karl Aberer
CIKM4
2017 Efficient Document Filtering Using Vector Space Topic Expansion and Pattern-Mining: The Case of Event Detection in Microposts
abstract
Automatically extracting information from social media is challenging given that social content is often noisy, ambiguous, and inconsistent. However, as many stories break on social channels first before being picked up by mainstream media, developing methods to better handle social content is of utmost importance. In this paper, we propose a robust and effective approach to automatically identify microposts related to a specific topic defined by a small sample of reference documents. Our framework extracts clusters of semantically similar microposts that overlap with the reference documents, by extracting combinations of key features that define those clusters through frequent pattern mining. This allows us to construct compact and interpretable representations of the topic, dramatically decreasing the computational burden compared to classical clustering and k-NN-based machine learning techniques and producing highly-competitive results even with small training sets (less than 1'000 training objects). Our method is efficient and scales gracefully with large sets of incoming microposts. We experimentally validate our approach on a large corpus of over 60M microposts, showing that it significantly outperforms state-of-the-art techniques.
Julia Proskurnia, Ruslan Mavlyutov, Carlos Castillo 0001, Karl Aberer, Philippe Cudré-Mauroux
CIKM4
2017 Predicting the Success of Online Petitions Leveraging Multidimensional Time-Series
abstract
Applying classical time-series analysis techniques to online content is challenging, as web data tends to have data quality issues and is often incomplete, noisy, or poorly aligned. In this paper, we tackle the problem of predicting the evolution of a time series of user activity on the web in a manner that is both accurate and interpretable, using related time series to produce a more accurate prediction. We test our methods in the context of predicting signatures for online petitions using data from thousands of petitions posted on The Petition Site - one of the largest platforms of its kind. We observe that the success of these petitions is driven by a number of factors, including promotion through social media channels and on the front page of the petitions platform. We propose an interpretable model that incorporates seasonality, aging effects, self-excitation, and external effects. The interpretability of the model is important for understanding the elements that drives the activity of an online content. We show through an extensive empirical evaluation that our model is significantly better at predicting the outcome of a petition than state-of-the-art techniques.
Julia Proskurnia, Przemyslaw A. Grabowicz, Ryota Kobayashi, Carlos Castillo 0001, Philippe Cudré-Mauroux, Karl Aberer
WWW6
2017 Argument discovery via crowdsourcing
Nguyen Quoc Viet Hung, Chi Thang Duong, Thanh Tam Nguyen, Matthias Weidlich 0001, Karl Aberer, Hongzhi Yin, Xiaofang Zhou 0001
VLDB J.5
2017 Answer validation for generic crowdsourcing tasks with minimal efforts
Nguyen Quoc Viet Hung, Chi Thang Duong, Thanh Tam Nguyen, Matthias Weidlich 0001, Karl Aberer, Hongzhi Yin, Xiaofang Zhou 0001
VLDB J.5
2016 Leveraging user expertise in collaborative systems for annotating energy datasets
abstract
While tasks such as segmenting images or determining the sentiment expressed in a sentence can be assigned to regular users, some others require background knowledge and thus, the selection of expert users. In the case of energy datasets, acquiring data represents an obstacle to develop data-driven methods, due to prohibitive monetary and time costs linked to the instrumentation of households in order to monitor the energy consumption. More so, most datasets only contain pure power time series, despite labels being required to determine when a device is in use from when it is idle (incurring stand-by consumption or being off), and by extension to separate human activities triggering the consumption from the baseline consumption. We build upon our Collaborative Annotation Framework for Energy Datasets (CAFED) to evaluate and distinguish the performance of expert users against that of regular users. Through a user study with curated benchmark annotation tasks, we provide data-driven and efficient techniques to detect weak and adversarial workers and promote users when the contributors' user-base is limited. Additionally, we show that if carefully selected, the seed gold standard tasks can be reduced to a small number of tasks that are representative enough to determine the user's expertise and predict crowd-combined annotations with high precision.
Hông-Ân Sandlin, Felix Rauchenstein, Tri Kurniawan Wijaya, Karl Aberer, Nuno Nunes 0001
IEEE BigData4
2016 Estimating human interactions with electrical appliances for activity-based energy savings recommendations
abstract
Since the power consumption of different electrical appliances in a household can be recorded by individual smart meters, it becomes possible to start considering in more detail the interactions of the residents with those devices throughout the day. Appliances' usages should not be considered as independent events, but rather as enablers for activities. Leveraging activity knowledge over time will allow us to design personalized energy efficient measures. We envision the design of future ambient intelligence systems, where the smart home can optimize the energy consumption in regards to the lifestyles of its residents and the smart grid's needs. In this work, we propose an automated method for determining when an electrical device is triggered by households' residents solely from its power trace. Knowing when an appliance is in use is required for identifying recurrent patterns that could later be understood as activities.
Hông-Ân Sandlin, Tri Kurniawan Wijaya, Karl Aberer, Nuno Nunes 0001
IEEE BigData3
2016 Temporal association rules for electrical activity detection in residential homes
abstract
Attaining energy efficiency requires understanding human behaviors triggering energy consumption within households. In conjunction to providing appliance-level feedback, targeting human activities that involve the usage of electrical appliances can provide a higher abstraction level to bring awareness to the electricity wastage. In this paper, we make use of a large dataset with appliance- and circuit-level power data and provide a framework for determining temporal sequential association rules. Sequences of time intervals where the appliances are in usage can vary in their order, duration and the time elapsed between these events. Our contribution consists in providing a full pipeline for mining frequent sequential itemsets and a novel way to discover the time windows during which these sequences of events occur and to capture their variance in terms of duration and order. Our method is data-driven and relies on the data's statistical properties and allows us to avoid an exhaustive search for the time windows' sizes, by relying instead on machine learning techniques to identify and predict those time windows.
Hông-Ân Sandlin, Tri Kurniawan Wijaya, Karl Aberer, Nuno Nunes 0001
IEEE BigData3
2016 Data Summarization with Social Contexts
abstract
While social data is being widely used in various applications such as sentiment analysis and trend prediction, its sheer size also presents great challenges for storing, sharing and processing such data. These challenges can be addressed by data summarization which transforms the original dataset into a smaller, yet still useful, subset. Existing methods find such subsets with objective functions based on data properties such as representativeness or informativeness but do not exploit social contexts, which are distinct characteristics of social data. Further, till date very little work has focused on topic preserving data summarization, despite the abundant work on topic modeling. This is a challenging task for two reasons. First, since topic model is based on latent variables, existing methods are not well-suited to capture latent topics. Second, it is difficult to find such social contexts that provide valuable information for building effective topic-preserving summarization model. To tackle these challenges, in this paper, we focus on exploiting social contexts to summarize social data while preserving topics in the original dataset. We take Twitter data as a case study. Through analyzing Twitter data, we discover two social contexts which are important for topic generation and dissemination, namely (i) CrowdExp topic score that captures the influence of both the crowd and the expert users in Twitter and (ii) Retweet topic score that captures the influence of Twitter users' actions. We conduct extensive experiments on two real-world Twitter datasets using two applications. The experimental results show that, by leveraging social contexts, our proposed solution can enhance topic-preserving data summarization and improve application performance by up to 18%.
Hao Zhuang 0002, Rameez Rahman, Xia Ben Hu, Tian Guo 0002, Pan Hui 0001, Karl Aberer
CIKM6
2016 Robust Online Time Series Prediction with Recurrent Neural Networks
abstract
Time series forecasting for streaming data plays an important role in many real applications, ranging from IoT systems, cyber-networks, to industrial systems and healthcare. However the real data is often complicated with anomalies and change points, which can lead the learned models deviating from the underlying patterns of the time series, especially in the context of online learning mode. In this paper we present an adaptive gradient learning method for recurrent neural networks (RNN) to forecast streaming time series in the presence of anomalies and change points. We explore the local features of time series to automatically weight the gradients of the loss of the newly available observations with distributional properties of the data in real time. We perform extensive experimental analysis on both synthetic and real datasets to evaluate the performance of the proposed method.
Tian Guo 0002, Zhao Xu 0001, Xin Yao 0001, Karl Aberer, Koichi Funaya
DSAA5
2016 A Distributed Mining Framework for Influence in Evolving Entities
Tian Guo 0002, Karl Aberer
EDBT2
2016 Efficient Distributed Decision Trees for Robust Regression
Tian Guo 0002, Konstantin Kutzkov, Mohamed Ahmed 0001, Jean-Paul Calbimonte, Karl Aberer
ECML/PKDD (2)5
2016 TripleWave: Spreading RDF Streams on the Web
Andrea Mauri 0001, Jean-Paul Calbimonte, Daniele Dell'Aglio, Marco Balduini, Marco Brambilla 0001, Emanuele Della Valle, Karl Aberer
ISWC (2)7
2016 Contextualized ranking of entity types based on knowledge graphs
Alberto Tonon, Michele Catasta, Roman Prokofyev, Gianluca Demartini, Karl Aberer, Philippe Cudré-Mauroux
J. Web Semant.5
2015 A collaborative framework for annotating energy datasets
abstract
Targeting human activities responsible for the energy consumption instead of focusing solely on single appliance feedback for achieving energy efficiency in residential homes would link human behaviors to the resulting energy consumption. To this end, learning when appliances are in an active or idle state and the related user activity is crucial. Until smart appliances become widespread and can communicate their internal state, identifying when the residents interact with the appliances has to be determined from the available information that can be recorded from these devices. Developing and validating learning models require ground truth in the form of annotations to indicate when an appliance is active or idle. Launching data collection campaigns to incorporate these missing ground truth data involves careful planning before the roll-out of the experiment. Prohibitive costs for the hardware and time investment to monitor the deployed equipment are necessary for quality data. As such, publicly released datasets containing appliance-level data offer a basis for most researchers. This paper addresses these challenges by providing a collaborative web-based framework to retrofit labeling on existing datasets. The platform is publicly available, applies the wisdom of the crowd in the realm of energy research and leverages gamification techniques to encourage users' active contribution. The access to the platform and furthermore to the expert manually labeled dataset intends to enable future research and foster more collaboration in this area.
Hông-Ân Sandlin, Tri Kurniawan Wijaya, Karl Aberer, Nuno Nunes 0001
IEEE BigData3
2015 SigCO: Mining significant correlations via a distributed real-time computation engine
abstract
The dramatic rise of time-series data produced in a variety of contexts, such as stock markets, mobile sensing, sensor networks, data centre monitoring, etc., has fuelled the development of large-scale distributed real-time computation systems (e.g., Apache Storm, Samza, Spark Streaming, S4, etc.). However, it is still unclear how certain time series mining tasks could be performed using such new emerging systems. In this paper, we focus on the task of efficiently discovering statistically significant correlations among a large number of time series via a distributed realtime computation engine. We propose a framework referred to as SigCO. In SigCO, we put forward a novel partition-aware data shuffling, which is able to adaptively shuffle time series data only to the relevant nodes of the distributed real-time computation engine. On the other hand, in SigCO we design a δ-hypercube structure based correlation computation approach which is capable of pruning unnecessary correlation computations. Finally, our extensive experimental evaluations on real and synthetic datasets establish that SigCO outperforms the baseline approaches in terms of diverse performance metrics.
Tian Guo 0002, Jean-Paul Calbimonte, Hao Zhuang 0002, Karl Aberer
IEEE BigData4
2015 Cluster-based aggregate forecasting for residential electricity demand using smart meter data
abstract
While electricity demand forecasting literature has focused on large, industrial, and national demand, this paper focuses on short-term (1 and 24 hour ahead) electricity demand forecasting for residential customers at the individual and aggregate level. Since electricity consumption behavior may vary between households, we first build a feature universe, and then apply Correlation-based Feature Selection to select features relevant to each household. Additionally, smart meter data can be used to obtain aggregate forecasts with higher accuracy using the so-called Cluster-based Aggregate Forecasting (CBAF) strategy, i.e., by first clustering the households, forecasting the clusters' energy consumption separately, and finally aggregating the forecasts. We found that the improvement provided by CBAF depends not only on the number of clusters, but also more importantly on the size of the customer base.
Tri Kurniawan Wijaya, Matteo Vasirani, Samuel Humeau 0003, Karl Aberer
IEEE BigData4
2015 Fast Distributed Correlation Discovery Over Streaming Time-Series Data
abstract
The dramatic rise of time-series data in a variety of contexts, such as social networks, mobile sensing, data centre monitoring, etc., has fuelled interest in obtaining real-time insights from such data using distributed stream processing systems. One such extremely valuable insight is the discovery of correlations in real-time from large-scale time-series data. A key challenge in discovering correlations is that the number of time-series pairs that have to be analyzed grows quadratically in the number of time-series, giving rise to a quadratic increase in both computation cost and communication cost between the cluster nodes in a distributed environment. To tackle the challenge, we propose a framework called AEGIS. AEGIS exploits well-established statistical properties to dramatically prune the number of time-series pairs that have to be evaluated for detecting interesting correlations. Our extensive experimental evaluations on real and synthetic datasets establish the efficacy of AEGIS over baselines.
Tian Guo 0002, Saket Sathe 0001, Karl Aberer
CIKM3
2015 Tag-Based Paper Retrieval: Minimizing User Effort with Diversity Awareness
Nguyen Quoc Viet Hung, Do Son Thanh, Thanh Tam Nguyen, Karl Aberer
DASFAA (1)4
2015 An Evaluation of Diversification Techniques
Chi Thang Duong, Thanh Tam Nguyen, Nguyen Quoc Viet Hung, Karl Aberer
DEXA (2)4
2015 SMART: A tool for analyzing and reconciling schema matching networks
abstract
Schema matching supports data integration by establishing correspondences between the attributes of independently designed database schemas. In recent years, various tools for automatic pair-wise matching of schemas have been developed. Since the matching process is inherently uncertain, the correspondences generated by such tools are often validated by a human expert. In this work, we consider scenarios in which attribute correspondences are identified in a network of schemas and not only in a pairwise setting. Here, correspondences between different schemas are interrelated, so that incomplete and erroneous matching results propagate in the network and the validation of a correspondence by an expert has ripple effects. To analyse and reconcile such matchings in schema networks, we present the Schema Matching Analyzer and Reconciliation Tool (SMART). It allows for the definition of network-level integrity constraints for the matching and, based thereon, detects and visualizes inconsistencies of the matching. The tool also supports the reconciliation of a matching by guiding an expert in the validation process and by offering semi-automatic conflict-resolution techniques.
Nguyen Quoc Viet Hung, Thanh Tam Nguyen, Vinh Tuan Chau, Tri Kurniawan Wijaya, Zoltán Miklós 0001, Karl Aberer, Avigdor Gal, Matthias Weidlich 0001
ICDE6
2015 Result selection and summarization for Web Table search
abstract
The amount of information available on the Web has been growing dramatically, raising the importance of techniques for searching the Web. Recently, Web Tables emerged as a model, which enables users to search for information in a structured way. However, effective presentation of results for Web Table search requires (1) selecting a ranking of tables that acknowledges the diversity within the search result; and (2) summarizing the information content of the selected tables concisely but meaningful. In this paper, we formalize these requirements as the diversified table selection problem and the structured table summarization problem. We show that both problems are computationally intractable and, thus, present heuristic algorithms to solve them. For these algorithms, we prove salient performance guarantees, such as near-optimality, stability, and fairness. Our experiments with real-world collections of thousands of Web Tables highlight the scalability of our techniques. We achieve improvements up to 50% in diversity and 10% in relevance over baselines for Web Table selection, and reduce the information loss induced by table summarization by up to 50%. In a user study, we observed that our techniques are preferred over alternative solutions.
Thanh Tam Nguyen, Nguyen Quoc Viet Hung, Matthias Weidlich 0001, Karl Aberer
ICDE4
2015 Comparing Events Coverage in Online News and Social Media: The Case of Climate Change
Alexandra Olteanu, Carlos Castillo 0001, Nicholas Diakopoulos, Karl Aberer
ICWSM4
2015 ERICA: Expert Guidance in Validating Crowd Answers
abstract
Crowdsourcing became an essential tool for a broad range of Web applications. Yet, the wide-ranging levels of expertise of crowd workers as well as the presence of faulty workers call for quality control of the crowdsourcing result. To this end, many crowdsourcing platforms feature a post-processing phase, in which crowd answers are validated by experts. This approach incurs high costs though, since expert input is a scarce resource. To support the expert in the validation process, we present a tool for \emph{ExpeRt guidance In validating Crowd Answers (ERICA)}. It allows us to guide the expert's work by collecting input on the most problematic cases, thereby achieving a set of high quality answers even if the expert does not validate the complete answer set. The tool also supports the task requester in selecting the most cost-efficient allocation of the budget between the expert and the crowd.
Nguyen Quoc Viet Hung, Chi Thang Duong, Matthias Weidlich 0001, Karl Aberer
SIGIR4
2015 Minimizing Efforts in Validating Crowd Answers
abstract
In recent years, crowdsourcing has become essential in a wide range of Web applications. One of the biggest challenges of crowdsourcing is the quality of crowd answers as workers have wide-ranging levels of expertise and the worker community may contain faulty workers. Although various techniques for quality control have been proposed, a post-processing phase in which crowd answers are validated is still required. Validation is typically conducted by experts, whose availability is limited and who incur high costs. Therefore, we develop a probabilistic model that helps to identify the most beneficial validation questions in terms of both, improvement of result correctness and detection of faulty workers. Our approach allows us to guide the expert's work by collecting input on the most problematic cases, thereby achieving a set of high quality answers even if the expert does not validate the complete answer set. Our comprehensive evaluation using both real-world and synthetic datasets demonstrates that our techniques save up to 50% of expert efforts compared to baseline methods when striving for perfect result correctness. In absolute terms, for most cases, we achieve close to perfect correctness after expert input has been sought for only 20\% of the questions.
Nguyen Quoc Viet Hung, Chi Thang Duong, Matthias Weidlich 0001, Karl Aberer
SIGMOD Conference4
2015 Time- and Space-Efficient Sliding Window Top-k Query Processing
abstract
A sliding window top-k ( top-k/w ) query monitors incoming data stream objects within a sliding window of size w to identify the k highest-ranked objects with respect to a given scoring function over time. Processing of such queries is challenging because, even when an object is not a top-k/w object at the time when it enters the processing system, it might become one in the future. Thus a set of potential top-k/w objects has to be stored in memory while its size should be minimized to efficiently cope with high data streaming rates. Existing approaches typically store top-k/w and candidate sliding window objects in a k-skyband over a two-dimensional score-time space. However, due to continuous changes of the k-skyband, its maintenance is quite costly. Probabilistic k-skyband is a novel data structure storing data stream objects from a sliding window with significant probability to become top-k/w objects in future. Continuous probabilistic k-skyband maintenance offers considerably improved runtime performance compared to k-skyband maintenance, especially for large values of k , at the expense of a small and controllable error rate. We propose two possible probabilistic k-skyband usages: ( i ) When it is used to process all sliding window objects, the resulting top-k/w algorithm is approximate and adequate for processing random-order data streams. ( ii ) When probabilistic k-skyband is used to process only a subset of most recent sliding window objects, it can improve the runtime performance of continuous k-skyband maintenance, resulting in a novel exact top-k/w algorithm. Our experimental evaluation systematically compares different top-k/w processing algorithms and shows that while competing algorithms offer either time efficiency at the expanse of space efficiency or vice-versa, our algorithms based on the probabilistic k-skyband are both time and space efficient.
Kresimir Pripuzic, Ivana Podnar Zarko, Karl Aberer
ACM Trans. Database Syst.3
2014 Online Indexing and Distributed Querying Model-View Sensor Data in the Cloud
Tian Guo 0002, Thanasis G. Papaioannou, Hao Zhuang 0002, Karl Aberer
DASFAA (1)4
2014 Privacy-Preserving Schema Reuse
Nguyen Quoc Viet Hung, Do Son Thanh, Thanh Tam Nguyen, Karl Aberer
DASFAA (2)4
2014 Pay-as-you-go reconciliation in schema matching networks
abstract
Schema matching is the process of establishing correspondences between the attributes of database schemas for data integration purposes. Although several automatic schema matching tools have been developed, their results are often incomplete or erroneous. To obtain a correct set of correspondences, a human expert is usually required to validate the generated correspondences. We analyze this reconciliation process in a setting where a number of schemas needs to be matched, in the presence of consistency expectations about the network of attribute correspondences. We develop a probabilistic model that helps to identify the most uncertain correspondences, thus allowing us to guide the expert's work and collect his input about the most problematic cases. As the availability of such experts is often limited, we develop techniques that can construct a set of good quality correspondences with a high probability, even if the expert does not validate all the necessary correspondences. We demonstrate the efficiency of our techniques through extensive experimentation using real-world datasets.
Nguyen Quoc Viet Hung, Thanh Tam Nguyen, Zoltán Miklós 0001, Karl Aberer, Avigdor Gal, Matthias Weidlich 0001
ICDE4
2014 Pattern-Wise Trust Assessment of Sensor Data
abstract
One of the most important tasks of a sensor network (SN) is to detect occurrences of interesting events in the monitored environment. However, data measured by SN is often affected by errors. We investigate the problem of assessing trustworthiness (trust) of a sensor value (tested value) in the presence of events and errors. A usual approach is to express the trust as a deviation of the tested value from a reference value (a normal value). State of the art approaches aim at defining the reference value in terms of a context consisting of values of spatially proximate sensors that are correlated with the tested value. However, they trade accuracy for simplicity and use a fixed context consisting of values of a fixed neighborhood (e.g., All values within a circular neighborhood of radius r). Therefore, such a fixed context fails in most practical cases by under or overestimating the reference values. We present the first pattern-wise method (PW) for trust assessment of sensor data that addresses the limitations of the state of the art approaches by departing from the idea of the fixed neighborhood. We consider a variable neighborhood that consists of an arbitrary subset of the spatially proximate sensors. We define the context as a frequent spatial pattern consisting of values of the variable neighborhood that frequently co-occurs with the tested value in the stream of sensor values. We define the trust as a belief (probability) that the tested value is correct given selected features of a frequent pattern consisting of the context and the tested value. We compute trust as the output of the logistic regression, where the input variables consist of the following features of the pattern: (I) the relative frequency, (II) the conditional probability of the tested value given the context and (III) the size of the variable neighborhood. Experimental results confirmed superiority of the proposed method over the state of the art method.
Robert Gwadera, Mehdi Riahi, Karl Aberer
MDM (1)3
2014 Towards a dynamic top-N recommendation framework
abstract
Real world large-scale recommender systems are always dynamic: new users and items continuously enter the system, and the status of old ones (e.g., users' preference and items' popularity) evolve over time. In order to handle such dynamics, we propose a recommendation framework consisting of an online component and an offline component, where the newly arrived items are processed by the online component such that users are able to get suggestions for fresh information, and the influence of longstanding items is captured by the offline component. Based on individual users' rating behavior, recommendations from the two components are combined to provide top-N recommendation. We formulate recommendation problem as a ranking problem where learning to rank is applied to extend upon matrix factorization to optimize item rankings by minimizing a pairwise loss function. Furthermore, to better model interactions between users and items, Latent Dirichlet Allocation is incorporated to fuse rating information and textual information. Real data based experiments demonstrate that our approach outperforms the state-of-the-art models by at least 61.21% and 50.27% in terms of mean average precision (MAP) and normalized discounted cumulative gain (NDCG) respectively.
Xin Liu 0027, Karl Aberer
RecSys2
2014 Consumer Segmentation and Knowledge Extraction from Smart Meter and Survey Data
abstract
Many electricity suppliers around the world are deploying smart meters to gather fine-grained spatiotemporal consumption data and to effectively manage the collective demand of their consumer base. In this paper, we introduce a structured framework and a discriminative index that can be used to segment the consumption data along multiple contextual dimensions such as locations, communities, seasons, weather patterns, holidays, etc. The generated segments can enable various higher-level applications such as usage-specific tariff structures, theft detection, consumer-specific demand response programs, etc. Our framework is also able to track consumers’ behavioral changes, evaluate different temporal aggregations, and identify main characteristics which define a cluster.
Tri Kurniawan Wijaya, Tanuja Ganu, Dipanjan Chakraborty 0001, Karl Aberer, Deva P. Seetharam
SDM4
2014 Comparing the Predictive Capability of Social and Interest Affinity for Recommendations
Alexandra Olteanu, Anne-Marie Kermarrec, Karl Aberer
WISE (1)3
2014 User-side adaptive protection of location privacy in participatory sensing
Berker Agir, Thanasis G. Papaioannou, Rammohan Narendula, Karl Aberer, Jean-Pierre Hubaux
GeoInformatica4
2014 Top-k/w publish/subscribe: A publish/subscribe model for continuous top-k processing over data streams
Kresimir Pripuzic, Ivana Podnar Zarko, Karl Aberer
Inf. Syst.3
2014 TransactiveDB: Tapping into Collective Human Memories
abstract
Database Management Systems (DBMSs) have been rapidly evolving in the recent years, exploring ways to store multi-structured data or to involve human processes during query execution. In this paper, we outline a future avenue for DBMSs supporting transactive memory queries that can only be answered by a collection of individuals connected through a given interaction graph. We present TransactiveDB and its ecosystem, which allow users to pose queries in order to reconstruct collective human memories. We describe a set of new transactive operators including TUnion, TFill, TJoin, and TProjection. We also describe how TransactiveDB leverages transactive operators---by mixing query execution, social network analysis and human computation---in order to effectively and efficiently tap into the memories of all targeted users.
Michele Catasta, Alberto Tonon, Djellel Eddine Difallah, Gianluca Demartini, Karl Aberer, Philippe Cudré-Mauroux
Proc. VLDB Endow.5
2014 B-hist: Entity-centric search over personal web browsing history
Michele Catasta, Alberto Tonon, Gianluca Demartini, Jean-Eudes Ranvier, Karl Aberer, Philippe Cudré-Mauroux
J. Web Semant.5
2013 Model-view sensor data management in the cloud
abstract
Infinite nature of sensor data poses a serious challenge for query processing even in a cloud infrastructure. Model-based sensor data approximation reduces the amount of data for query processing, but all modeled segments need to be scanned, in the worst case. In this paper, we propose an innovative index for modeled segments in key-value stores, namely KVI-index. KVI-index has an in-memory tree component and a secondary structure materialized in the key-value store that maps the tree nodes to the modeled data segments. Then, we introduce a KVI-index-Scan-MapReduce hybrid approach to perform efficient query processing. As proved by a series of experiments in a real private cloud infrastructure, our approach outperforms in query response time and index updating efficiency both Hadoop-based parallel processing of the raw sensor data and multiple alternative indexing approaches of model-view data.
Tian Guo 0002, Thanasis G. Papaioannou, Karl Aberer
IEEE BigData3
2013 Personalized point-of-interest recommendation by mining users' preference transition
abstract
Location-based social networks (LBSNs) offer researchers rich data to study people's online activities and mobility patterns. One important application of such studies is to provide personalized point-of-interest (POI) recommendations to enhance user experience in LBSNs. Previous solutions directly predict users' preference on locations but fail to provide insights about users' preference transitions among locations. In this work, we propose a novel category-aware POI recommendation model, which exploits the transition patterns of users' preference over location categories to improve location recommendation accuracy. Our approach consists of two stages: (1) preference transition (over location categories) prediction, and (2) category-aware POI recommendation. Matrix factorization is employed to predict a user's preference transitions over categories and then her preference on locations in the corresponding categories. Real data based experiments demonstrate that our approach outperforms the state-of-the-art POI recommendation models by at least 39.75% in terms of recall.
Xin Liu 0027, Yong Liu 0020, Karl Aberer, Chunyan Miao
CIKM3
2013 On Leveraging Crowdsourcing Techniques for Schema Matching Networks
Nguyen Quoc Viet Hung, Thanh Tam Nguyen, Zoltán Miklós 0001, Karl Aberer
DASFAA (2)4
2013 Web Credibility: Features Exploration and Credibility Prediction
Alexandra Olteanu, Stanislav Peshterliev, Xin Liu 0027, Karl Aberer
ECIR4
2013 Utility-driven data acquisition in participatory sensing
abstract
Participatory sensing (PS) is becoming a popular data acquisition means for interesting emerging applications. However, as data queries from these applications increase, the sustainability of this platform for multiple concurrent applications is at stake. In this paper, we consider the problem of efficient data acquisition in PS when queries of different types come from different applications. We effectively deal with the issues related to resource constraints, user privacy, data reliability, and uncontrolled mobility. We formulate the problem as multi-query optimization and propose efficient heuristics for its effective solution for the various query types and mixes that enable sustainable sensing. Based on simulations with real and artificial data traces, we found that our heuristic algorithms outperform baseline approaches in a multitude of settings considered.
Mehdi Riahi, Thanasis G. Papaioannou, Immanuel Trummer, Karl Aberer
EDBT4
2013 Minimizing Human Effort in Reconciling Match Networks
Nguyen Quoc Viet Hung, Tri Kurniawan Wijaya, Zoltán Miklós 0001, Karl Aberer, Eliezer Levy, Victor Shafran, Avigdor Gal, Matthias Weidlich 0001
ER4
2013 AFFINITY: Efficiently querying statistical measures on time-series data
abstract
Computing statistical measures for large databases of time series is a fundamental primitive for querying and mining time-series data [1]-[6]. This primitive is gaining importance with the increasing number and rapid growth of time series databases. In this paper, we introduce a framework for efficient computation of statistical measures by exploiting the concept of affine relationships. Affine relationships can be used to infer statistical measures for time series, from other related time series, instead of computing them directly; thus, reducing the overall computational cost significantly. The resulting methods exhibit at least one order of magnitude improvement over the best known methods. To the best of our knowledge, this is the first work that presents an unified approach for computing and querying several statistical measures at once. Our approach exploits affine relationships using three key components. First, the AFCLST algorithm clusters the time-series data, such that high-quality affine relationships could be easily found. Second, the SYMEX algorithm uses the clustered time series and efficiently computes the desired affine relationships. Third, the SCAPE index structure produces a many-fold improvement in the performance of processing several statistical queries by seamlessly indexing the affine relationships. Finally, we establish the effectiveness of our approaches by performing comprehensive experimental evaluation on real datasets.
Saket Sathe 0001, Karl Aberer
ICDE2
2013 TripEneer: User-Based Travel Plan Recommendation Application
Surender Reddy Yerva, Flavia Grosan, Alexandru Tandrau, Karl Aberer
ICWSM4
2013 TRank: Ranking Entity Types Using the Web of Data
Alberto Tonon, Michele Catasta, Gianluca Demartini, Philippe Cudré-Mauroux, Karl Aberer
ISWC (1)5
2013 BATC: a benchmark for aggregation techniques in crowdsourcing
abstract
As the volumes of AI problems involving human knowledge are likely to soar, crowdsourcing has become essential in a wide range of world-wide-web applications. One of the biggest challenges of crowdsourcing is aggregating the answers collected from crowd workers; and thus, many aggregate techniques have been proposed. However, given a new application, it is difficult for users to choose the best-suited technique as well as appropriate parameter values since each of these techniques has distinct performance characteristics depending on various factors (e.g. worker expertise, question difficulty). In this paper, we develop a benchmarking tool that allows to (i) simulate the crowd and (ii) evaluate aggregate techniques in different aspects (accuracy, sensitivity to spammers, etc.). We believe that this tool will be able to serve as a practical guideline for both researchers and software developers. While researchers can use our tool to assess existing or new techniques, developers can reuse its components to reduce the development complexity.
Nguyen Quoc Viet Hung, Thanh Tam Nguyen, Ngoc Tran Lam, Karl Aberer
SIGIR4
2013 An Evaluation of Aggregation Techniques in Crowdsourcing
Nguyen Quoc Viet Hung, Thanh Tam Nguyen, Ngoc Tran Lam, Karl Aberer
WISE (2)4
2013 SoCo: a social network aided context-aware recommender system
abstract
Contexts and social network information have been proven to be valuable information for building accurate recommender system. However, to the best of our knowledge, no existing works systematically combine diverse types of such information to further improve recommendation quality. In this paper, we propose SoCo, a novel context-aware recommender system incorporating elaborately processed social network information. We handle contextual information by applying random decision trees to partition the original user-item-rating matrix such that the ratings with similar contexts are grouped. Matrix factorization is then employed to predict missing preference of a user for an item using the partitioned matrix. In order to incorporate social network information, we introduce an additional social regularization term to the matrix factorization objective function to infer a user's preference for an item by learning opinions from his/her friends who are expected to share similar tastes. A context-aware version of Pearson Correlation Coefficient is proposed to measure user similarity. Real datasets based experiments show that SoCo improves the performance (in terms of root mean square error) of the state-of-the-art context-aware recommender system and social recommendation model by 15.7% and 12.2% respectively.
Xin Liu 0027, Karl Aberer
WWW2
2013 EnviroMeter: A Platform for Querying Community-Sensed Data
abstract
Efficiently querying data collected from Large-area Community driven Sensor Networks (LCSNs) is a new and challenging problem. In our previous works, we proposed adaptive techniques for learning models (e.g., statistical, nonparametric, etc.) from such data, considering the fact that LCSN data is typically geo-temporally skewed. In this paper, we present a demonstration of EnviroMeter. EnviroMeter uses our adaptive model creation techniques for processing continuous queries on community-sensed environmental pollution data. Subsequently, it efficiently pushes current pollution updates to GPS-enabled smartphones (through its Android application) or displays it via a web-interface. We experimentally demonstrate that our model-based query processing approach is orders of magnitude efficient than processing the queries over indexed raw data.
Saket Sathe 0001, Arthur Oviedo, Dipanjan Chakraborty 0001, Karl Aberer
Proc. VLDB Endow.4
2013 Semantic trajectories: Mobility data computation and annotation
abstract
With the large-scale adoption of GPS equipped mobile sensing devices, positional data generated by moving objects (e.g., vehicles, people, animals) are being easily collected. Such data are typically modeled as streams of spatio-temporal (x,y,t) points, called trajectories . In recent years trajectory management research has progressed significantly towards efficient storage and indexing techniques, as well as suitable knowledge discovery. These works focused on the geometric aspect of the raw mobility data. We are now witnessing a growing demand in several application sectors (e.g., from shipment tracking to geo-social networks) on understanding the semantic behavior of moving objects. Semantic behavior refers to the use of semantic abstractions of the raw mobility data, including not only geometric patterns but also knowledge extracted jointly from the mobility data and the underlying geographic and application domains information. The core contribution of this article lies in a semantic model and a computation and annotation platform for developing a semantic approach that progressively transforms the raw mobility data into semantic trajectories enriched with segmentations and annotations. We also analyze a number of experiments we did with semantic trajectories in different domains.
Zhixian Yan, Dipanjan Chakraborty 0001, Christine Parent, Stefano Spaccapietra, Karl Aberer
ACM Trans. Intell. Syst. Technol.5
2013 An Evaluation of Model-Based Approaches to Sensor Data Compression
abstract
As the volumes of sensor data being accumulated are likely to soar, data compression has become essential in a wide range of sensor-data applications. This has led to a plethora of data compression techniques for sensor data, in particular model-based approaches have been spotlighted due to their significant compression performance. These methods, however, have never been compared and analyzed under the same setting, rendering a "right" choice of compression technique for a particular application very difficult. Addressing this problem, this paper presents a benchmark that offers a comprehensive empirical study on the performance comparison of the model-based compression techniques. Specifically, we reimplemented several state-of-the-art methods in a comparable manner, and measured various performance factors with our benchmark, including compression ratio, computation time, model maintenance cost, approximation quality, and robustness to noisy data. We then provide in-depth analysis of the benchmark results, obtained by using 11 different real data sets consisting of 346 heterogeneous sensor data signals. We believe that the findings from the benchmark will be able to serve as a practical guideline for applications that need to compress sensor data.
Nguyen Quoc Viet Hung, Hoyoung Jeung, Karl Aberer
IEEE Trans. Knowl. Data Eng.3
2012 Model-Based Similarity Measure in TimeCloud
Thanh-Nguyen Ngo, Hoyoung Jeung, Karl Aberer
APWeb3
2012 A decentralized recommender system for effective web credibility assessment
abstract
An overwhelming and growing amount of data is available online. The problem of untrustworthy online information is augmented by its high economic potential and its dynamic nature, e.g. transient domain names, dynamic content, etc. In this paper, we address the problem of assessing the credibility of web pages by a decentralized social recommender system. Specifically, we concurrently employ i) item-based collaborative filtering (CF) based on specific web page features, ii) user-based CF based on friend ratings and iii) the ranking of the page in search results. These factors are appropriately combined into a single assessment based on adaptive weights that depend on their effectiveness for different topics and different fractions of malicious ratings. Simulation experiments with real traces of web page credibility evaluations suggest that our hybrid approach outperforms both its constituent components and classical content-based classification approaches.
Thanasis G. Papaioannou, Jean-Eudes Ranvier, Alexandra Olteanu, Karl Aberer
CIKM4
2012 Cloud based social and sensor data fusion
Surender Reddy Yerva, Hoyoung Jeung, Karl Aberer
FUSION3
2012 OptiMoS: Optimal Sensing for Mobile Sensors
abstract
Both sensor coverage maximization and energy cost minimization are the fundamental requirements in the design of real-life mobile sensing applications, e.g., (1) deploying environmental sensors (like CO2, fine particle measurement) on public transports to monitor air pollution, (2) analyzing smart phone embedded sensors (like GPS, accelerometer) to recognize people daily activities. However sensor coverage and energy cost contradict each other: the higher frequency mobile sensing takes, the more energy is used, and vise versa. In this paper, we design a novel two-step mobile sensing process ("OptiMoS") to achieve optimal mobile sensing that can effectively balance sensor coverage and energy cost. In the first step, OptiMoS divides the continuous mobile sensor readings into several segments, where the readings in one segment are highly-correlated rather than readings amongst different segments. In the second step, OptiMoS identifies optimal sampling for the sensor readings in each segment, where the selected readings can guarantee reasonably high sensor coverage with limited sampling rate. Various greedy & near-optimal segmentation and sampling methods are designed in OptiMoS, and are evaluated using real-life environmental data from mobile sensors.
Zhixian Yan, Julien Eberle, Karl Aberer
MDM3
2012 Social and Sensor Data Fusion in the Cloud
abstract
This paper explores the potential of fusing social and sensor data in the cloud, presenting a practice - a travel recommendation system that offers the predicted mood information of people on where and when users wish to travel. The system is built upon a conceptual framework that allows to blend the heterogeneous social and sensor data for integrated analysis, extracting weather-dependant people's mood information from Twitter and meteorological sensor data streams. In order to handle massively streaming data, the system employs various cloud-serving systems, such as Hadoop, HBase, and GSN. Using this scalable system, we performed heavy ETL as well as filtering jobs, resulting in 12 million tweets over four months. We then derived a rich set of interesting findings through the data fusion, proving that our approach is effective and scalable, which can serve as an important basis in fusing social and sensor data in the cloud.
Surender Reddy Yerva, Jonnahtan Saltarin, Hoyoung Jeung, Karl Aberer
MDM4
2012 Tag Recommendation for Large-Scale Ontology-Based Information Systems
Roman Prokofyev, Alexey Boyarsky, Oleg Ruchayskiy, Karl Aberer, Gianluca Demartini, Philippe Cudré-Mauroux
ISWC (2)4
2012 TweetSpector: entity-based retrieval of tweets
abstract
TweetSpector is a tool for demonstrating entity-based of retrieval of tweets. The various features of this tool include: entity profile creation, real-time tweet classification, active improvement of the created profiles through user feedback, and the dashboard displaying different metrics.
Surender Reddy Yerva, Zoltán Miklós 0001, Flavia Grosan, Alexandru Tandrau, Karl Aberer
SIGIR5
2012 Enabling Query Technologies for the Semantic Sensor Web
abstract
Sensor networks are increasingly being deployed in the environment for many different purposes. The observations that they produce are made available with heterogeneous schemas, vocabularies and data formats, making it difficult to share and reuse this data, for other purposes than those for which they were originally set up. The authors propose an ontology-based approach for providing data access and query capabilities to streaming data sources, allowing users to express their needs at a conceptual level, independent of implementation and language-specific details. In this article, the authors describe the theoretical foundations and technologies that enable exposing semantically enriched sensor metadata, and querying sensor observations through SPARQL extensions, using query rewriting and data translation techniques according to mapping languages, and managing both pull and push delivery modes.
Jean-Paul Calbimonte, Hoyoung Jeung, Óscar Corcho, Karl Aberer
Int. J. Semantic Web Inf. Syst.4
2012 Quality-aware similarity assessment for entity matching in Web data
Surender Reddy Yerva, Zoltán Miklós 0001, Karl Aberer
Inf. Syst.3
2011 SeMiTri: a framework for semantic annotation of heterogeneous trajectories
abstract
GPS devices allow recording the movement track of the moving object they are attached to. This data typically consists of a stream of spatio-temporal (x,y,t) points. For application purposes the stream is transformed into finite subsequences called trajectories. Existing knowledge extraction algorithms defined for trajectories mainly assume a specific context (e.g. vehicle movements) or analyze specific parts of a trajectory (e.g. stops), in association with data from chosen geographic sources (e.g. points-of-interest, road networks). We investigate a more comprehensive semantic annotation framework that allows enriching trajectories with any kind of semantic data provided by multiple 3rd party sources.
Zhixian Yan, Dipanjan Chakraborty 0001, Christine Parent, Stefano Spaccapietra, Karl Aberer
EDBT5
2011 Advanced search, visualization and tagging of sensor metadata
abstract
As sensors continue to proliferate, the capabilities of effectively querying not only sensor data but also its metadata becomes important in a wide range of applications. This paper demonstrates a search system that utilizes various techniques and tools for querying sensor metadata and visualizing the results. Our system provides an easy-to-use query interface, built upon semantic technologies where users can freely store and query their metadata. Going beyond basic keyword search, the system provides a variety of advanced functionalities tailored for sensor metadata search; ordering search results according to our ranking mechanism based on the PageRank algorithm, recommending pages that contain relevant metadata information to given search conditions, presenting search results using various visualization tools, and offering dynamic hypergraphs and tag clouds of metadata. The system has been running as a real application and its effectiveness has been proved by a number of users.
Ioannis K. Paparrizos, Hoyoung Jeung, Karl Aberer
ICDE3
2011 Creating probabilistic databases from imprecise time-series data
abstract
Although efficient processing of probabilistic databases is a well-established field, a wide range of applications are still unable to benefit from these techniques due to the lack of means for creating probabilistic databases. In fact, it is a challenging problem to associate concrete probability values with given time-series data for forming a probabilistic database, since the probability distributions used for deriving such probability values vary over time. In this paper, we propose a novel approach to create tuple-level probabilistic databases from (imprecise) time-series data. To the best of our knowledge, this is the first work that introduces a generic solution for creating probabilistic databases from arbitrary time series, which can work in online as well as offline fashion. Our approach consists of two key components. First, the dynamic density metrics that infer time-dependent probability distributions for time series, based on various mathematical models. Our main metric, called the GARCH metric, can robustly capture such evolving probability distributions regardless of the presence of erroneous values in a given time series. Second, the Ω-View builder that creates probabilistic databases from the probability distributions inferred by the dynamic density metrics. For efficient processing, we introduce the σ-cache that reuses the information derived from probability values generated at previous times. Extensive experiments over real datasets demonstrate the effectiveness of our approach.
Saket Sathe 0001, Hoyoung Jeung, Karl Aberer
ICDE3
2011 Towards Online Multi-model Approximation of Time Series
abstract
The increasing use of sensor technology for various monitoring applications (e.g. air-pollution, traffic, climate-change, etc.) has led to an unprecedented volume of streaming data that has to be efficiently aggregated, stored and retrieved. Real-time model-based data approximation and filtering is a common solution for reducing the storage (and communication) overhead. However, the selection of the most efficient model depends on the characteristics of the data stream, namely rate, burstiness, data range, etc., which cannot be always known a priori for (mobile) sensors and they can even dynamically change. In this paper, we investigate the innovative concept of efficiently combining multiple approximation models in real-time. Our approach dynamically adapts to the properties of the data stream and approximates each data segment with the most suitable model. As experimentally proved, our multi-model approximation approach always produces fewer or equal data segments than those of the best individual model, and thus provably achieves higher data compression ratio than individual linear models.
Thanasis G. Papaioannou, Mehdi Riahi, Karl Aberer
Mobile Data Management (1)3
2010 The gist of everything new: personalized top-k processing over web 2.0 streams
abstract
Web 2.0 portals have made content generation easier than ever with millions of users contributing news stories in form of posts in weblogs or short textual snippets as in Twitter. Efficient and effective filtering solutions are key to allow users stay tuned to this ever-growing ocean of information, releasing only relevant trickles of personal interest. In classical information filtering systems, user interests are formulated using standard IR techniques and data from all available information sources is filtered based on a predefined absolute quality-based threshold. In contrast to this restrictive approach which may still overwhelm the user with the returned stream of data, we envision a system which continuously keeps the user updated with only the top-k relevant new information. Freshness of data is guaranteed by considering it valid for a particular time interval, controlled by a sliding window. Considering relevance as relative to the existing pool of new information creates a highly dynamic setting. We present POL-filter which together with our maintenance module constitute an efficient solution to this kind of problem. We show by comprehensive performance evaluations using real world data, obtained from a weblog crawl, that our approach brings performance gains compared to state-of-the-art.
Parisa Haghani, Sebastian Michel 0001, Karl Aberer
CIKM3
2010 Automatic construction and multi-level visualization of semantic trajectories
abstract
With the prevalence of GPS-embedded mobile devices, enormous amounts of mobility data are being collected in the form of trajectory - a stream of (x,y,t) points. Such trajectories are of heterogeneous entities - vehicles, people, animals, parcels etc. Most applications primarily analyze raw trajectory data and extract geometric patterns. Real-life applications however, need a far more comprehensive, semantic representation of trajectories. This paper demonstrates the automatic construction and visualization capabilities of SeMiTri - a system we built that exploits 3rd party information sources containing geographic information, to semantically enrich trajectories. The construction stack encapsulates several spatio-temporal data integration and mining techniques to automatically compute and annotate all meaningful parts of heterogeneous trajectories. The visualization interface exhibits different levels of data abstraction, from low-level raw trajectories (i.e. the initial GPS trace) to high-level semantic trajectories (i.e. the sequence of interesting places where moving objects have passed and/or stayed).
Zhixian Yan, Lazar Spremic, Dipanjan Chakraborty 0001, Christine Parent, Stefano Spaccapietra, Karl Aberer
GIS6
2010 Cost-efficient and differentiated data availability guarantees in data clouds
abstract
Failures of any type are common in current datacenters. As data scales up, its availability becomes more complex, while different availability levels per application or per data item may be required. In this paper, we propose a self-managed key-value store that dynamically allocates the resources of a data cloud to several applications in a cost-efficient and fair way. Our approach offers and dynamically maintains multiple differentiated availability guarantees to each different application despite failures. We employ a virtual economy, where each data partition acts as an individual optimizer and chooses whether to migrate, replicate or remove itself based on net benefit maximization regarding the utility offered by the partition and its storage and maintenance cost. Comprehensive experimental evaluations suggest that our solution is highly scalable and adaptive to query rate variations and to resource upgrades/failures.
Nicolas Bonvin, Thanasis G. Papaioannou, Karl Aberer
ICDE3
2010 Continuous query evaluation over distributed sensor networks
abstract
In this paper we address the problem of processing continuous multi-join queries, over distributed data streams. Our approach makes use of existing work in the field of publish/subscribe systems. We show how these principles can be ported to our envisioned architectural model by enriching the common query model with location dependent attributes. We allow users to subscribe to a set of sensor attributes, a service that requires processing multi-join correlation queries. The goal is to decrease the overall network traffic consumption by removing redundant subscriptions and eliminating unrequested events close to the publishing sensors. This is non-trivial, especially in the presence of multi-join queries without any central control mechanism. Our approach is based on the concept of filter-split-forward phases for efficient subscription filtering and placement inside the network. We report on a performance evaluation using a real-world dataset, showing the improvements over the state-of-the-art, as we reduce the overall data traffic by half.
Oana Jurca, Sebastian Michel 0001, Alexandre Herrmann, Karl Aberer
ICDE4
2009 Evaluating top-k queries over incomplete data streams
abstract
We study the problem of continuous monitoring of top-k queries over multiple non-synchronized streams. Assuming a sliding window model, this general problem has been a well addressed research topic in recent years. Most approaches, however, assume synchronized streams where all attributes of an object are known simultaneously to the query processing engine. In many streaming scenarios though, different attributes of an item are reported in separate non-synchronized streams which do not allow for exact score calculations. We present how the traditional notion of object dominance changes in this case such that the k dominance set still includes all and only those objects which have a chance of being among the top-k results in their life time. Based on this, we propose an exact algorithm which builds on generating multiple instances of the same object in a way that enables efficient object pruning. We show that even with object pruning the necessary storage for exact evaluation of top-k queries is linear in the size of the sliding window. As data should reside in main memory to provide fast answers in an online fashion and cope with high stream rates, storing all this data may not be possible with limited resources. We present an approximate algorithm which leverages correlation statistics of pairs of streams to evict more objects while maintaining accuracy. We evaluate the efficiency of our proposed algorithms with extensive experiments.
Parisa Haghani, Sebastian Michel 0001, Karl Aberer
CIKM3
2009 Distributed similarity search in high dimensions using locality sensitive hashing
abstract
In this paper we consider distributed K-Nearest Neighbor (KNN) search and range query processing in high dimensional data. Our approach is based on Locality Sensitive Hashing (LSH) which has proven very efficient in answering KNN queries in centralized settings. We consider mappings from the multi-dimensional LSH bucket space to the linearly ordered set of peers that jointly maintain the indexed data and derive requirements to achieve high quality search results and limit the number of network accesses. We put forward two such mappings that come with these salient properties: being locality preserving so that buckets likely to hold similar data are stored on the same or neighboring peers and having a predictable output distribution to ensure fair load balancing. We show how to leverage the linearly aligned data for efficient KNN search and how to efficiently process range queries which is, to the best of our knowledge, not possible in existing LSH schemes. We show by comprehensive performance evaluations using real world data that our approach brings major performance and accuracy gains compared to state-of-the-art.
Parisa Haghani, Sebastian Michel 0001, Karl Aberer
EDBT3
2009 Towards integrated and efficient scientific sensor data processing: a database approach
abstract
In this work, we focus on managing scientific environmental data, which are measurement readings collected from wireless sensors. In environmental science applications, raw sensor data often need to be validated, interpolated, aligned and aggregated before being used to construct meaningful result sets. Due to the lack of a system that integrates all the necessary processing steps, scientists often resort to multiple tools to manage and process the data, which can severely affect the efficiency of their work. In this paper, we propose a new data processing framework, HyperGrid, to address the problem. HyperGrid adopts a generic data model and a generic query processing and optimization framework. It offers an integrated environment to store, query, analyze and visualize scientific datasets. The experiments on real query set and data set show that the framework not only introduces little processing overhead, but also provides abundant opportunities to optimize the processing cost and thus significantly enhances the processing efficiency.
Ji Wu 0011, Yongluan Zhou, Karl Aberer, Kian-Lee Tan
EDBT3
2009 Neighborhood-Based Tag Prediction
Adriana Budura, Sebastian Michel 0001, Philippe Cudré-Mauroux, Karl Aberer
ESWC4
2009 Environmental Monitoring 2.0
abstract
A sensor network data gathering and visualization infrastructure is demonstrated, comprising of global sensor networks (GSN) middleware and Microsoft SensorMap. Users are invited to actively participate in the process of monitoring real-world deployments and can inspect measured data in the form of contour plots overlayed onto a high resolution map and a digital topographic model. Users can go back in time virtually to search for interesting events or simply to visualize the temporal dependencies of the data. The system presented is not only interesting and visually enticing for non-expert users but brings substantial benefits to environmental scientists. The easily installed data acquisition component as well as the powerful data sharing and visualization platform opens up new ground in collaborative data gathering and interpretation in the spirit of Web 2.0 applications.
Sebastian Michel 0001, Ali Salehi, Liqian Luo, Nicholas Dawes, Karl Aberer, Guillermo Barrenetxea, Mathias Bavay, Aman Kansal, K. Ashwin Kumar, Suman Nath, Marc Parlange, Stewart Tansley, Catharine van Ingen, Feng Zhao 0001, Yongluan Zhou
ICDE5
2009 Knowing When to Slide - Efficient Scheduling for Sliding Window Processing
abstract
We consider sliding window query execution scheduling in stream processing engines. Sliding windows are an essential building block to limit the query focus at a particular part of the stream, based either on value count or time ranges. These so called sliding window predicates specify the execution condition for the query. Due to the often massive amount of registered queries, efficient algorithms to check these predicates are essential. While there exists a comprehensive set of works on the stream processing techniques, the actual algorithms to intelligently decide on the sliding behaviors is not extensively addressed in the existing works. In this paper we propose a set of algorithms for managing and sharing sliding decisions. This work introduces the concept of the batch sliding and sliding graphs to improve the sliding decision of the stream processing engines. Our algorithms can be efficiently used in large-scale stream processing systems where data arrives at high rates and a large number of user queries are registered to these data streams. Our evaluation results show the suitability of this approach in the real world applications.
Ali Salehi, Mehdi Riahi, Sebastian Michel 0001, Karl Aberer
Mobile Data Management4
2009 idMesh: graph-based disambiguation of linked data
abstract
We tackle the problem of disambiguating entities on the Web. We propose a user-driven scheme where graphs of entities -- represented by globally identifiable declarative artifacts -- self-organize in a dynamic and probabilistic manner. Our solution has the following two desirable properties: i) it lets end-users freely define associations between arbitrary entities and ii) it probabilistically infers entity relationships based on uncertain links using constraint-satisfaction mechanisms. We outline the interface between our scheme and the current data Web, and show how higher-layer applications can take advantage of our approach to enhance search and update of information relating to online entities. We describe a decentralized infrastructure supporting efficient and scalable entity disambiguation and demonstrate the practicability of our approach in a deployment over several hundreds of machines.
Philippe Cudré-Mauroux, Parisa Haghani, Michael Jost 0003, Karl Aberer, Hermann de Meer
WWW4
2009 Scalable Delivery of Stream Query Results
abstract
Continuous queries over data streams typically produce large volumes of result streams. To scale up the system, one should carefully study the problem of delivering the result streams to the end users, which, unfortunately, is often overlooked in existing systems. In this paper, we leverage Distributed Publish/Subscribe System (DPSS), a scalable data dissemination infrastructure, for efficient stream query result delivery. To take advantage of DPSS's multicast-like data dissemination architecture, one has to exploit the common contents among different result streams and maximize the sharing of their delivery. Hence, we propose to merge the user queries into a few representative queries whose results subsume those of the original ones, and disseminate the result streams of these representative queries through the DPSS. To realize this approach, we study the stream query containment theories and propose efficient query grouping and merging algorithms. The proposed approach is non-intrusive and hence can be easily implemented as a middleware to be incorporated into existing stream processing systems. A prototype is developed on top of an open-source stream processing system and results of an extensive performance study on real datasets verify the efficacy of the proposed techniques.
Yongluan Zhou, Ali Salehi, Karl Aberer
Proc. VLDB Endow.3
2008 To tag or not to tag -: harvesting adjacent metadata in large-scale tagging systems
abstract
We present HAMLET, a suite of principles, scoring models and algorithms to automatically propagate metadata along edges in a document neighborhood. As a showcase scenario we consider tag prediction in community-based Web 2.0 tagging applications. Experiments using real-world data demonstrate the viability of our approach in large-scale environments where tags are scarce. To the best of our knowledge, HAMLET is the first system to promote an efficient and precise reuse of shared metadata in highly dynamic, large-scale Web 2.0 tagging systems.
Adriana Budura, Sebastian Michel 0001, Philippe Cudré-Mauroux, Karl Aberer
SIGIR4
2008 LSH At Large - Distributed KNN Search in High Dimensions
Parisa Haghani, Sebastian Michel 0001, Philippe Cudré-Mauroux, Karl Aberer
WebDB4
2008 Effective Usage of Computational Trust Models in Rational Environments
abstract
Reputation-based trust models using statistical learning have been intensively studied for distributed systems where peers behave maliciously. However practical applications of such models in environments with both malicious and rational behaviors are still very little understood. This paper studies the relation between accuracy of a computational trust model and its ability to effectively enforce cooperation among rational agents. We provide theoretical results showing under which conditions cooperation emerges when using a trust learning algorithms with given accuracy and how cooperation can be still sustained while reducing cost and accuracy of those algorithms. We then verify and extend these theoretical results to a variety of settings involving honest, malicious and strategic players through extensive simulation. These results will enable a much more targeted, cost-effective and realistic design for decentralized trust management systems, such as needed for peer-to-peer systems and electronic commerce.
Le-Hung Vu, Karl Aberer
Web Intelligence2
2008 AlvisP2P: scalable peer-to-peer text retrieval in a structured P2P network
abstract
In this paper we present the AlvisP2P IR engine, which enables efficient retrieval with multi-keyword queries from a global document collection available in a P2P network. In such a network, each peer publishes its local index and invests a part of its local computing resources (storage, CPU, bandwidth) to maintain a fraction of a global P2P index. This investment is rewarded by the network-wide accessibility of the local documents via the global search facility. The AlvisP2P engine uses an optimized overlay network and relies on novel indexing/retrieval mechanisms that ensure low bandwidth consumption, thus enabling unlimited network growth. Our demonstration shows how an easy-to-install AlvisP2P client can be used to join an existing P2P network, index local (text or even multimedia) documents with collection-specific indexing mechanisms, and control access rights to them.
Toan Luu, Gleb Skobeltsyn, Fabius Klemm, Maroje Puh, Ivana Podnar Zarko, Martin Rajman, Karl Aberer
Proc. VLDB Endow.7
2008 PicShark: mitigating metadata scarcity through large-scale P2P collaboration
Philippe Cudré-Mauroux, Adriana Budura, Manfred Hauswirth, Karl Aberer
VLDB J.4
2007 Oscar: A Data-Oriented Overlay For Heterogeneous Environments
abstract
Quite a few data-oriented overlay networks have been designed in recent years. These designs often (implicitly) assume various homogeneity which seriously limit their usability in real world. In this paper we present some performance results of the Oscar overlay, which simultaneously deals with heterogeneity as observed in the Internet (capacity of computers, bandwidth) as well as non-uniformity observed in data-oriented applications.
Sarunas Girdzijauskas, Anwitaman Datta, Karl Aberer
ICDE3
2007 Scalable Peer-to-Peer Web Retrieval with Highly Discriminative Keys
abstract
The suitability of peer-to-peer (P2P) approaches for full-text Web retrieval has recently been questioned because of the claimed unacceptable bandwidth consumption induced by retrieval from very large document collections. In this contribution we formalize a novel indexing/retrieval model that achieves high performance, cost-efficient retrieval by indexing with highly discriminative keys (HDKs) stored in a distributed global index maintained in a structured P2P network. HDKs correspond to carefully selected terms and term sets appearing in a small number of collection documents. We provide a theoretical analysis of the scalability of our retrieval model and report experimental results obtained with our HDK-based P2P retrieval engine. These results show that, despite increased indexing costs, the total traffic generated with the HDK approach is significantly smaller than the one obtained with distributed single-term indexing strategies. Furthermore, our experiments show that the retrieval performance obtained with a random set of real queries is comparable to the one of centralized, single-term solution using the best state-of-the-art BM25 relevance computation scheme. Finally, our scalability analysis demonstrates that the HDK approach can scale to large networks of peers indexing Web-size document collections, thus opening the way towards viable, truly-decentralized Web retrieval.
Ivana Podnar Zarko, Martin Rajman, Toan Luu, Fabius Klemm, Karl Aberer
ICDE5
2007 An Extensible and Personalized Approach to QoS-enabled Service Discovery
abstract
We present an extensible and customizable framework for the autonomous discovery of Semantic Web services based on their QoS properties. Using semantic technologies, users can specify the QoS matching model and customize the ranking of services flexibly according to their preferences. The formal modeling of the discovery process as a query execution plan facilitates the introduction of different discovery algorithms and the automatic generation of parallelized matchmaking evaluations. This enables adapting our approach to unpredictable arrival rates of user queries and scales up to high numbers of published service descriptions.
Le-Hung Vu, Fábio Porto 0001, Karl Aberer, Manfred Hauswirth
IDEAS3
2007 Smart Earth: From Pervasive Observation to Trusted Information
abstract
We introduce a model for information management in a world where information from the physical environment is gathered through a ubiquitous wireless infrastructure. Thus the Internet information space and the physical environment become increasingly entangled in what we call a smart Earth. We illustrate technical challenges and applications from the work of the Swiss national competence centre of research in mobile information and communication systems (NCCR-MICS). We identify the three layers of data access networks, semantic overlay networks and social networks as essential building blocks and observe that each of these network layers will be largely self-organized. Finally we argue that understanding the interplay among these self-organizing network layers will be an important research challenge for the future.
Karl Aberer
MDM1
2007 Infrastructure for Data Processing in Large-Scale Interconnected Sensor Networks
abstract
With the price of wireless sensor technologies diminishing rapidly we can expect large numbers of autonomous sensor networks being deployed in the near future. These sensor networks will typically not remain isolated but the need of interconnecting them on the network level to enable integrated data processing will arise, thus realizing the vision of a global "sensor Internet." This requires a flexible middleware layer which abstracts from the underlying, heterogeneous sensor network technologies and supports fast and simple deployment and addition of new platforms, facilitates efficient distributed query processing and combination of sensor data, provides support for sensor mobility, and enables the dynamic adaption of the system configuration during runtime with minimal (zero-programming) effort. This paper describes the global sensor networks (GSN) middleware which addresses these goals. We present GSN's conceptual model, abstractions, and architecture, and demonstrate the efficiency of the implementation through experiments with typical high-load application profiles. The GSN implementation is available from http://gsn.sourceforge.net/.
Karl Aberer, Manfred Hauswirth, Ali Salehi
MDM1
2007 Web text retrieval with a P2P query-driven index
abstract
In this paper, we present a query-driven indexing/retrieval strategy for efficient full text retrieval from large document collections distributed within a structured P2P network. Our indexing strategy is based on two important properties: (1) the generated distributed index stores posting lists for carefully chosen indexing term combinations, and (2) the posting lists containing too many document references are truncated to a bounded number of their top-ranked elements. These two properties guarantee acceptable storage and bandwidth requirements, essentially because the number of indexing term combinations remains scalable and the transmitted posting lists never exceed a constant size. However, as the number of generated term combinations can still become quite large, we also use term statistics extracted from available query logs to index only such combinations that are frequently present in user queries. Thus, by avoiding the generation of superfluous indexing term combinations, we achieve an additional substantial reduction in bandwidth and storage consumption. As a result, the generated distributed index corresponds to a constantly evolving query-driven indexing structure that efficiently follows current information needs of the users. More precisely, our theoretical analysis and experimental results indicate that, at the price of a marginal loss in retrieval quality for rare queries, the generated index size and network traffic remain manageable even for web-size document collections. Furthermore, our experiments show that at the same time the achieved retrieval quality is fully comparable to the one obtained with a state-of-the-art centralized query engine.
Gleb Skobeltsyn, Toan Luu, Ivana Podnar Zarko, Martin Rajman, Karl Aberer
SIGIR5
2007 Self-Organizing Schema Mappings in the GridVine Peer Data Management System
Philippe Cudré-Mauroux, Suchit Agarwal, Adriana Budura, Parisa Haghani, Karl Aberer
VLDB5
2007 Query-driven indexing for peer-to-peer text retrieval
abstract
We describe a query-driven indexing framework for scalable text retrieval over structured P2P networks. To cope with the bandwidth consumption problem that has been identified as the major obstacle for full-text retrieval in P2P networks, we truncate posting lists associated with indexing features to a constant size storing only top-k ranked document references. To compensate for the loss of information caused by the truncation, we extend the set of indexing features with carefully chosen term sets. Indexing term sets are selected based on the query statistics extracted from query logs, thus we index only such combinations that are a) frequently present in user queries and b) non-redundant w.r.t the rest of the index. The distributed index is compact and efficient as it constantly evolves adapting to the current query popularity distribution. Moreover, it is possible to control the tradeoff between the storage/bandwidth requirements and the quality of query answering by tuning the indexing parameters. Our theoretical analysis and experimental results indicate that we can indeed achieve scalable P2P text retrieval for very large document collections and deliver good retrieval performance.
Gleb Skobeltsyn, Toan Luu, Karl Aberer, Martin Rajman, Ivana Podnar Zarko
WWW3
2006 Data Management in the Social Web
Karl Aberer
EDBT1
2006 Probabilistic Message Passing in Peer Data Management Systems
abstract
Until recently, most data integration techniques involved central components, e.g., global schemas, to enable transparent access to heterogeneous databases. Today, however, with the democratization of tools facilitating knowledge elicitation in machine-processable formats, one cannot rely on global, centralized schemas anymore as knowledge creation and consumption are getting more and more dynamic and decentralized. Peer Data Management Systems (PDMS) provide an answer to this problem by eliminating the central semantic component and considering instead compositions of local, pair-wise mappings to propagate queries from one database to the others. PDMS approaches proposed so far make the implicit assumption that all mappings used in this way are correct. This obviously cannot be taken as granted in typical PDMS settings where mappings can be created (semi) automatically by independent parties. In this work, we propose a totally decentralized, efficient message passing scheme to automatically detect erroneous mappings in PDMS. Our scheme is based on a probabilistic model where we take advantage of transitive closures of mapping operations to confront local belief on the correctness of a mapping against evidences gathered around the network. We show that our scheme can be efficiently embedded in any PDMS and provide a preliminary evaluation of our techniques on sets of both automatically-generated and real-world schemas.
Philippe Cudré-Mauroux, Karl Aberer, Andras Feher
ICDE2
2006 A Middleware for Fast and Flexible Sensor Network Deployment
Karl Aberer, Manfred Hauswirth, Ali Salehi
VLDB1
2006 Mapping Moving Landscapes by Mining Mountains of Logs: Novel Techniques for Dependency Model Generation
Mirko Steinle, Karl Aberer, Sarunas Girdzijauskas, Christian Lovis
VLDB2
2006 Query optimization in XML structured-document databases
Dunren Che, Karl Aberer, M. Tamer Özsu
VLDB J.2
2005 Semantic Overlay Networks
Karl Aberer, Philippe Cudré-Mauroux
VLDB1
2005 Indexing Data-oriented Overlay Networks
Karl Aberer, Anwitaman Datta, Manfred Hauswirth, Roman Schmidt
VLDB1
2005 Opportunities from Open Source Search
abstract
Internet search has a strong business model that permits a free service to users, so it is difficult to see why, if at all, there should be open source offerings as well. This paper first discusses open source search and a rationale for the computer science community at large to get involved. Because there is no shortage of core open source components for at least some of the tasks involved, the Alvis Consortium is building infrastructure for open source search engines using peer-to-peer and subject specific technology as its core, based on this rationale. We view open source search as a rich future playground in which information extraction and retrieval components can be used and intelligent agents can operate.
Wray L. Buntine, Karl Aberer, Ivana Podnar Zarko, Martin Rajman
Web Intelligence2
2004 Emergent Semantics Principles and Issues
Karl Aberer, Philippe Cudré-Mauroux, Aris M. Ouksel, Tiziana Catarci, Mohand-Said Hacid, Arantza Illarramendi, Vipul Kashyap, Massimo Mecella, Eduardo Mena, Erich J. Neuhold, Olga De Troyer, Thomas Risse 0001, Monica Scannapieco, Fèlix Saltor, Luca De Santis, Stefano Spaccapietra, Steffen Staab, Rudi Studer
DASFAA1
2004 GridVine: Building Internet-Scale Semantic Overlay Networks
Karl Aberer, Philippe Cudré-Mauroux, Manfred Hauswirth, Tim Van Pelt
ISWC1
2004 Efficient, Self-Contained Handling of Identity in Peer-to-Peer Systems
abstract
Identification is an essential building block for many services in distributed information systems. The quality and purpose of identification may differ, but the basic underlying problem is always to bind a set of attributes to an identifier in a unique and deterministic way. Name/directory services, such as DNS, X.500, or UDDI, are a well-established concept to address this problem in distributed information systems. However, none of these services addresses the specific requirements of peer-to-peer systems with respect to dynamism, decentralization, and maintenance. We propose the implementation of directories using a structured peer-to-peer overlay network and apply this approach to support self-contained maintenance of routing tables with dynamic IP addresses in structured P2P systems. Thus, we keep routing tables intact without affecting the organization of the overlay networks, making it logically independent of the underlying network infrastructure. Even though the directory is self-referential, since it uses its own service to maintain itself, we show that it is robust due to a self-healing capability. For security, we apply a combination of PGP-like public key distribution and a quorum-based query scheme. We describe the algorithm as implemented in the P-Grid P2P lookup system (http:// www.p-grid.org/) and give a detailed analysis and simulation results demonstrating the efficiency and robustness of our approach.
Karl Aberer, Anwitaman Datta, Manfred Hauswirth
IEEE Trans. Knowl. Data Eng.1
2003 A Framework for Decentralized Ranking in Web Information Retrieval
Karl Aberer, Jie Wu 0012
APWeb1
2003 Swarm Intelligent Surfing in the Web
Jie Wu 0012, Karl Aberer
ICWE2
2003 The chatty web: emergent semantics through gossiping
abstract
This paper describes a novel approach for obtaining semantic interoperability among data sources in a bottom-up, semi-automatic manner without relying on pre-existing, global semantic models. We assume that large amounts of data exist that have been organized and annotated according to local schemas. Seeing semantics as a form of agreement, our approach enables the participating data sources to incrementally develop global agreement in an evolutionary and completely decentralized process that solely relies on pair-wise, local interactions: Participants provide translations between schemas they are interested in and can learn about other translations by routing queries (gossiping). To support the participants in assessing the semantic quality of the achieved agreements we develop a formal framework that takes into account both syntactic and semantic criteria. The assessment process is incremental and the quality ratings are adjusted along with the operation of the system. Ultimately, this process results in global agreement, i.e., the semantics that all participants understand. We discuss strategies to efficiently find translations and provide results from a case study to justify our claims. Our approach applies to any system which provides a communication infrastructure (existing websites or databases, decentralized systems, P2P systems) and offers the opportunity to study semantic interoperability as a global phenomenon in a network of information sharing parties.
Karl Aberer, Philippe Cudré-Mauroux, Manfred Hauswirth
WWW1
2003 Start making sense: The Chatty Web approach for global semantic agreements
Karl Aberer, Philippe Cudré-Mauroux, Manfred Hauswirth
J. Web Semant.1
2002 On the efficient evaluation of relaxed queries in biological databases
abstract
In this paper, a new technique is developed to support the query relaxation in biological databases. Query relaxation is required due to the fact that queries tend not to be expressed exactly by the users, especially in scientific databases such as biological databases, in which complex domain knowledge is heavily involved. To treat this problem, we propose the concept of the so-called fuzzy equivalence classes to capture important kinds of domain knowledge that is used to relax queries. This concept is further integrated with the canonical techniques for pattern searching such as the position tree and automaton theory. As a result, fuzzy queries produced through relaxation can be efficiently evaluated. This method has been successfully utilized in a practical biological database - the GPCRDB.
Yangjun Chen, Dunren Che, Karl Aberer
CIKM3
2002 P2P Information Systems
Karl Aberer, Manfred Hauswirth
ICDE1
2001 Managing Trust in a Peer-2-Peer Information System
abstract
Managing trust is a problem of particular importance in peer-to-peer environments where one frequently encounters unknown agents. Existing methods for trust management, that are based on reputation, focus on the semantic properties of the trust model. They do not scale as they either rely on a central database or require to maintain global knowledge at each agent to provide data on earlier interactions. In this paper we present an approach that addresses the problem of reputation-based trust management at both the data management and the semantic level. We employ at both levels scalable data structures and algorithms that require no central control and allow to assess trust by computing an agents reputation from its former interactions with other agents. Thus the meethod can be implemented in a peer-to-peer environment and scales well for very large numbers of participants. We expect that scalable methods for trust management are an important factor, if fully decentralized peer-to-peer systems should become the platform for more serious applications than simple file exchange.
Karl Aberer, Zoran Despotovic
CIKM1
1999 Adaptive Outsourcing in Cross-Organizational Workflows
Justus Klingemann, Jürgen Wäsch, Karl Aberer
CAiSE3
1999 Combining Pat-Trees and Signature Files for Query Evaluation in Document Databases
Yangjun Chen, Karl Aberer
DEXA2
1999 A Heuristics-Based Approach to Query Optimization in Structured Document Databases
abstract
The number of documents published via the World Wide Web in the form of SGML/HTML has been rapidly growing for years. Efficient, declarative access mechanisms for this type of document-structured documents in general-are becoming of great importance. This paper reports our most recent advance in pursuit of the effective processing and optimization of structured document queries, which are important for large repositories of structured documents. Our methodology emphasizes applying exclusively deterministic transformations on query expressions to achieve the best possible optimization efficiency. A new approach is thus proposed that facilitates the exploitation of the DTD (document type definition) knowledge, structural properties and structure indices of structured documents for the purpose of fast query optimization.
Dunren Che, Karl Aberer
IDEAS2
1999 A Query System in a Biological Database
abstract
We present a query system that has been implemented in a practical biological database-GPCRDB. Distinguishing features of this system include: smart query relaxation and smooth integration of navigation with conventional language based query functions. Query relaxation is required due to the fact that queries are not always effective (in other words, expected results are frequently not achieved), particularly in scientific databases like biological databases, in which complex domain knowledge is heavily used. On the other hand, navigation capability is desired as complex data sets are involved, especially in a WWW based environment where multiple hyperlinks are often employed. For efficient implementation, the "fuzzy equivalence class" concept has been applied that captures an important type of domain knowledge.
Dunren Che, Yangjun Chen, Karl Aberer
SSDBM3
1999 The Advanced Web Query System of GPCRDB
abstract
GPCRDB (G-Protein Coupled Receptors DataBase) is an advanced data management system for G-protein coupled receptors (GPCRs). It collects various data related to all aspects of GPCRs in a variety of forms: HTML pages, formatted source data files, 3D structures, ordinary database tables, etc. The GPCRDB query system is designed to facilitate academic and industrial users in the areas of pharmacology and/or biology to search for information in an integrated, easy-to-operate environment. Currently, the query system is operable on the World Wide Web at.
Dunren Che, Yangjun Chen, Karl Aberer, Hannelore Eisner
SSDBM3
1998 Layered Index Structures in Document Database Systems
abstract
LSIR
Yangjun Chen, Karl Aberer
CIKM2
1997 Admissible Record-Oriented Evaluation Plans for Declarative Updates
Gisela Fischer, Karl Aberer
ADBIS2
1997 Databases and the Web: What's in it for Databases? (Panel)
abstract
Database technology and the Web, as they exist today, have obviously many relationships. However, the opinions about these relationships vary widely and accordingly database technology plays many roles in the World Wide Web (WWW). The spectrum ranges from viewing the Web as one huge database for which data management problems need to be tackled to considering the WWW just as a platform for database system interoperability and access. In this panel we try to explore highly probable scenarios for the role of database technology in the WWW.
Erich J. Neuhold, Karl Aberer
ICDE2
1997 Structured Document Storage and Refined Declarative and Navigational Access Mechanisms in HyperStorM
Klemens Böhm, Karl Aberer, Erich J. Neuhold, Xiaoya Yang
VLDB J.2
1996 Applying a Flexible OODBMS-IRS-Coupling for Structured Document Handling
abstract
In document management systems, it is desirable to provide content-based access to documents going beyond regular expression search in addition to access based on structural characteristics or associated attributes. We present a new approach for coupling OODBMSs (object-oriented database management systems) and IRSs (information retrieval systems) that provides enhanced flexibility and functionality as compared to coupling approaches reported from the literature. Our approach allows one to decide freely to which document collections, that are used as retrieval context, document objects belong, which text contents they provide for retrieval, and how they derive their associated retrieval values, either directly from the retrieval machine or from the values of related objects. Especially, we show how, in this approach, different strategies can be applied to hierarchically structured documents, possibly avoiding redundancy and IRS or OODBMS peculiarities. Content-based and structural queries can be freely combined within the OODBMS query language.
Marc Volz, Karl Aberer, Klemens Böhm
ICDE2
1996 HyperStorM - Administering Structured Documents Using Object-Oriented Database Technology
abstract
No abstract available.
Klemens Böhm, Karl Aberer
SIGMOD Conference2
1995 Semantic Query Optimization for Methods in Object-Oriented Database Systems
abstract
Although the main difference between the relational and the object-oriented data model is the possibility to define object behavior, query optimization techniques in object-oriented database systems are mainly based on the structural part of objects. We claim that the optimization potential emerging from methods has been strongly underestimated so far. In this paper we concentrate on the question of how semantic knowledge about methods can be considered in query optimization. We rely on the algebraic and rule-based approach for query optimization and present a framework that allows to integrate schema-specific knowledge by tailoring the query optimizer according to the particular application's needs. We sketch an implementation of our concepts within the OODBMS VODAK using the Volcano optimizer generator.>
Karl Aberer, Gisela Fischer
ICDE1
1994 Designing a User-Oriented Query Modification Facility in Object-Oriented Database Systems
Karl Aberer, Wolfgang Klas, António L. Furtado 0001
CAiSE1
1994 An Object-Oriented Database Application for HyTime Document Storage
abstract
LSIR
Klemens Böhm, Karl Aberer
CIKM2