EDBT 2026 Demo / reviewers in the wild / expert
Abraham Bernstein
dblp:b/AbrahamBernstein
· DBLP profile ↗
64ranked-venue papers in the field
7as first author
17since 2021 · last 2025
0000-0002-0128-4602ORCID · verified
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 29 (5 first)Information Retrieval & Web Search · 16Database Systems & Data Management · 8 (1 first)Data Mining & Knowledge Discovery · 6Other / Interdisciplinary · 3 (1 first)Big Data, Cloud & Distributed Data Systems · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Informfully Recommenders - Reproducibility Framework for Diversity-aware Intra-session RecommendationsabstractNorm-aware recommender systems have gained increased attention, especially for diversity optimization. The recommender systems community has well-established experimentation pipelines that support reproducible evaluations by facilitating models' benchmarking and comparisons against state-of-the-art methods. However, to the best of our knowledge, there is currently no reproducibility framework to support thorough norm-driven experimentation at the pre-processing, in-processing, post-processing, and evaluation stages of the recommender pipeline. To address this gap, we present Informfully Recommenders, a first step towards a normative reproducibility framework that focuses on diversity-aware design built on Cornac. Our extension provides an end-to-end solution for implementing and experimenting with normative and general-purpose diverse recommender systems that cover 1) dataset pre-processing, 2) diversity-optimized models, 3) dedicated intrasession item re-ranking, and 4) an extensive set of diversity metrics. We demonstrate the capabilities of our extension through an extensive offline experiment in the news domain. Lucien Heitz, Oana Inel, Abraham Bernstein |
RecSys | 4 |
| 2025 | D-RDW: Diversity-Driven Random Walks for News Recommender SystemsabstractThis paper introduces Diversity-Driven Random Walks (D-RDW), a lightweight algorithm and re-ranking technique that generates diverse news recommendations. D-RDW is a societal recommender, which combines the diversification capabilities of the traditional random walk algorithms with customizable target distributions of news article properties. In doing so, our model provides a transparent approach for editors to incorporate norms and values into the recommendation process. D-RDW shows enhanced performance across key diversity metrics that consider the articles’ sentiment and political party mentions when compared to state-of-the-art neural models. Furthermore, D-RDW proves to be more computationally efficient than existing approaches. Lucien Heitz, Oana Inel, Abraham Bernstein |
RecSys | 4 |
| 2025 | NLQxform-UI: An Interactive and Intuitive Scholarly Question Answering SystemabstractMost scholarly search services only provide basic text-matching or similarity-based searches, with limited operations that require manual configuration, such as sorting and filtering by specific metadata attributes. These capabilities are insufficient for researchers who often have queries that involve complex constraints and operations, such as ''enumerating the authors of a given paper along with the venues where they have published other papers.'' In this work, we develop an interactive and intuitive scholarly question answering system called NLQxform-UI, which allows users to pose complex queries in the form of natural language questions. It is capable of automatically translating these questions into SPARQL queries that can be executed over the DBLP knowledge graph to retrieve expected answers. Furthermore, the users can interact with each step of the answering process and browse the final results in a web-based interface. A video recording of our system is available at https://youtu.be/elq8CPykiyk Additionally, the system has been completely open-sourced: https://github.com/ruijie-wang-uzh/NLQxform-UI Ruijie Wang 0003, Zhiruo Zhang, Luca Rossetto, Florian Ruosch, Abraham Bernstein |
SIGIR | 5 |
| 2025 | Whom do Explanations Serve? A Systematic Literature Survey of User Characteristics in Explainable Recommender Systems EvaluationabstractAdding explanations to recommender systems is said to have multiple benefits, such as increasing user trust or system transparency. Previous work from other application areas suggests that specific user characteristics impact the users’ perception of the explanation. However, we rarely find this type of evaluation for recommender systems explanations. This paper addresses this gap by surveying 124 papers in which recommender systems explanations were evaluated in user studies. We analyzed their participant descriptions and study results where the impact of user characteristics on the explanation effects was measured. Our findings suggest that the results from the surveyed studies predominantly cover specific users who do not necessarily represent the users of recommender systems in the evaluation domain. This may seriously hamper the generalizability of any insights we may gain from current studies on explanations in recommender systems. We further find inconsistencies in the data reporting, which impacts the reproducibility of the reported results. Hence, we recommend actions to move toward a more inclusive and reproducible evaluation. Kathrin Wardatzky, Oana Inel, Luca Rossetto, Abraham Bernstein |
Trans. Recomm. Syst. | 4 |
| 2024 | QAGCN: Answering Multi-relation Questions via Single-Step Implicit Reasoning over Knowledge Graphs
Ruijie Wang 0003, Luca Rossetto, Michael Cochez, Abraham Bernstein |
ESWC (1) | 4 |
| 2024 | Fast and Adaptive Questionnaires for Voting Advice Applications
Fynn Bachmann, Cristina Sarasua, Abraham Bernstein |
ECML/PKDD (10) | 3 |
| 2024 | Informfully - Research Platform for Reproducible User StudiesabstractThis paper presents Informfully, a research platform for content distribution and user studies. Informfully allows to push algorithmically curated text, image, audio, and video content to users and automatically generates a detailed log of their consumption history. As such, it serves as an open-source platform for conducting user experiments to investigate the impact of item recommendations on users’ consumption behavior. The platform was designed to accommodate different experiment types through versatility, ease of use, and scalability. It features three core components: 1) a front end for displaying and interacting with recommended items, 2) a back end for researchers to create and maintain user experiments, and 3) a simple JSON-based exchange format for ranked item recommendations to interface with third-party frameworks. We provide a system overview and outline the three core components of the platform. A sample workflow is shown for conducting field studies incorporating multiple user groups, personalizing recommendations, and measuring the effect of algorithms on user engagement. We present evidence for the versatility, ease of use, and scalability of Informfully by showcasing previous studies that used our platform. Lucien Heitz, Julian A. Croci, Madhav Sachdeva, Abraham Bernstein |
RecSys | 4 |
| 2024 | SciHyp: A Fine-Grained Dataset Describing Hypotheses and Their Components from Scientific Articles
Rosni Vasu, Cristina Sarasua, Abraham Bernstein |
ISWC (3) | 3 |
| 2023 | Deliberative Diversity for News Recommendations: Operationalization and Experimental User StudyabstractNews recommender systems are an increasingly popular field of study that attracts a growing interdisciplinary research community. As these systems play an essential role in our daily lives, the mechanisms behind their curation processes are under scrutiny. In the area of personalized news, many platforms make design choices driven by economic incentives. In contrast to such systems that optimize for financial gain, there can be norm-driven diversity systems that prioritize normative and democratic goals. However, their impact on users in terms of inducing behavioral change or influencing knowledge is still understudied. In this paper, we contribute to the field of news recommender system design by conducting a user study that examines the impact of these normative approaches. We a.) operationalize the notion of a deliberative public sphere for news recommendations, show b.) the impact on news usage, and c.) the influence on political knowledge, attitudes and voting behavior. We find that exposure to small parties is associated with an increase in knowledge about their candidates and that intensive news consumption about a party can change the direction of attitudes of readers towards the issues of the party. Lucien Heitz, Juliane A. Lischka, Rana Abdullah, Laura Laugwitz, Hendrik Meyer, Abraham Bernstein |
RecSys | 6 |
| 2023 | Towards the Web of Embeddings: Integrating multiple knowledge graph embedding spaces with FedCoderabstractThe Semantic Web is distributed yet interoperable: Distributed since resources are created and published by a variety of producers, tailored to their specific needs and knowledge; Interoperable as entities are linked across resources, allowing to use resources from different providers in concord. Complementary to the explicit usage of Semantic Web resources, embedding methods made them applicable to machine learning tasks. Subsequently, embedding models for numerous tasks and structures have been developed, and embedding spaces for various resources have been published. The ecosystem of embedding spaces is distributed but not interoperable: Entity embeddings are not readily comparable across different spaces. To parallel the Web of Data with a Web of Embeddings, we must thus integrate available embedding spaces into a uniform space. Current integration approaches are limited to two spaces and presume that both of them were embedded with the same method — both assumptions are unlikely to hold in the context of a Web of Embeddings. In this paper, we present FedCoder— an approach that integrates multiple embedding spaces via a latent space. We assert that linked entities have a similar representation in the latent space so that entities become comparable across embedding spaces. FedCoder employs an autoencoder to learn this latent space from linked as well as non-linked entities. Our experiments show that FedCoder substantially outperforms state-of-the-art approaches when faced with different embedding models, that it scales better than previous methods in the number of embedding spaces, and that it improves with more graphs being integrated whilst performing comparably with current approaches that assumed joint learning of the embeddings and were, usually, limited to two sources. Our results demonstrate that FedCoder is well adapted to integrate the distributed, diverse, and large ecosystem of embeddings spaces into an interoperable Web of Embeddings. Matthias Baumgartner, Daniele Dell'Aglio, Heiko Paulheim, Abraham Bernstein |
J. Web Semant. | 4 |
| 2022 | Evaluation of Algorithms for Interaction-Sparse Recommendations: Neural Networks don't Always Win
Yasamin Klingler, Claude Lehmann, João Pedro Monteiro, Carlo Saladin, Abraham Bernstein, Kurt Stockinger |
EDBT | 5 |
| 2022 | A framework for differentially-private knowledge graph embeddings
Xiaolin Han 0002, Daniele Dell'Aglio, Tobias Grubenmann, Reynold Cheng, Abraham Bernstein |
J. Web Semant. | 5 |
| 2022 | Visualising the effects of ontology changes and studying their understanding with ChImpabstractDue to the Semantic Web’s decentralised nature, ontology engineers rarely know all applications that leverage their ontology. Consequently, they are unaware of the full extent of possible consequences that changes might cause to the ontology. Our goal is to lessen the gap between ontology engineers and users by investigating ontology engineers’ understanding of ontology changes’ impact at editing time. Hence, this paper introduces the Protégé plugin ChImp which we use to reach our goal. We elicited requirements for ChImp through a questionnaire with ontology engineers. We then developed ChImp according to these requirements and it displays all changes of a given session and provides selected information on said changes and their effects. For each change, it computes a number of metrics on both the ontology and its materialisation. It displays those metrics on both the originally loaded ontology at the beginning of the editing session and the current state to help ontology engineers understand the impact of their changes. We investigated the informativeness of materialisation impact measures, the meaning of severe impact, and also the usefulness of ChImp in an online user study with 36 ontology engineers. We asked the participants to solve two ontology engineering tasks – with and without ChImp (assigned in random order) – and answer in-depth questions about the applied changes as well as the materialisation impact measures. We found that ChImp increased the participants’ understanding of change effects and that they felt better informed. Answers also suggest that the proposed measures were useful and informative. We also learned that the participants consider different outcomes of changes severe, but most would define severity based on the amount of changes to the materialisation compared to its size. The participants also acknowledged the importance of quantifying the impact of changes and that the study will affect their approach of editing ontologies. Romana Pernisch, Daniele Dell'Aglio, Mirko Serbak, Rafael S. Gonçalves 0001, Abraham Bernstein |
J. Web Semant. | 5 |
| 2021 | Single Point Incremental Fourier Transform on 2D Data StreamsabstractIn radio astronomy, antennas monitor portions of the sky to collect radio signals. The antennas produce data streams that are of high volume and velocity (~2.5 GB/s) and the inverse Fourier transform is used to convert the collected signals into sky images that astrophysicists use to conduct their research. Applying the inverse Fourier transform in a streaming setting, however, is not ideal since its computational complexity is quadratic in the size of the image.In this article, we propose the Single Point Incremental Fourier Transform (SPIFT), a novel incremental algorithm to produce sequences of sky images. SPIFT computes the Fourier transform for a new signal in a linear number of complex multiplications by exploiting twiddle factors, multiplicative constant coefficients. We prove that twiddle factors are periodic and show how circular shifts can be exploited to reuse multiplication results. The cost of the additive operations can be curbed by exploiting the embarrassingly parallel nature of the additions, which modern big data streaming frameworks can leverage to compute slices of the image in parallel. Our experiments suggest that SPIFT can efficiently generate sequences of sky images: it computes the complex multiplications 4 to 12x faster than the Discrete Fourier Transform, and its parallelisation of the additive operations shows linear speedup. Muhammad Saad 0006, Abraham Bernstein, Michael H. Böhlen, Daniele Dell'Aglio |
ICDE | 2 |
| 2021 | Toward Measuring the Resemblance of Embedding Models for Evolving OntologiesabstractUpdates on ontologies affect the operations built on top of them. But not all changes are equal: some updates drastically change the result of operations; others lead to minor variations, if any. Hence, estimating the impact of a change ex-ante is highly important, as it might make ontology engineers aware of the consequences of their action during editing. However, in order to estimate the impact of changes, we need to understand how to measure them. Romana Pernisch, Daniele Dell'Aglio, Abraham Bernstein |
K-CAP | 3 |
| 2021 | Random Walks with Erasure: Diversifying Personalized Recommendations on Social and Information NetworksabstractMost existing personalization systems promote items that match a user’s previous choices or those that are popular among similar users. This results in recommendations that are highly similar to the ones users are already exposed to, resulting in their isolation inside familiar but insulated information silos. In this context, we develop a novel recommendation framework with a goal of improving information diversity using a modified random walk exploration of the user-item graph. We focus on the problem of political content recommendation, while addressing a general problem applicable to personalization tasks in other social and information networks. Bibek Paudel, Abraham Bernstein |
WWW | 2 |
| 2021 | Beware of the hierarchy - An analysis of ontology evolution and the materialisation impact for biomedical ontologiesabstractOntologies are becoming a key component of numerous applications and research fields. But knowledge captured within ontologies is not static. Some ontology updates potentially have a wide ranging impact; others only affect very localised parts of the ontology and their applications. Investigating the impact of the evolution gives us insight into the editing behaviour but also signals ontology engineers and users how the ontology evolution is affecting other applications. However, such research is in its infancy. Hence, we need to investigate the evolution itself and its impact on the simplest of applications: the materialisation. In this work, we define impact measures that capture the effect of changes on the materialisation. In the future, the impact measures introduced in this work can be used to investigate how aware the ontology editors are about consequences of changes. By introducing five different measures, which focus either on the change in the materialisation with respect to the size or on the number of changes applied, we are able to quantify the consequences of ontology changes. To see these measures in action, we investigate the evolution and its impact on materialisation for nine open biomedical ontologies, most of which adhere to the EL++ description logic. Our results show that these ontologies evolve at varying paces but no statistically significant difference between the ontologies with respect to their evolution could be identified. We identify three types of ontologies based on the types of complex changes which are applied to them throughout their evolution. The impact on the materialisation is the same for the investigated ontologies, bringing us to the conclusion that the effect of changes on the materialisation can be generalised to other similar ontologies. Further, we found that the materialised concept inclusion axioms experience most of the impact induced by changes to the class inheritance of the ontology and other changes only marginally touch the materialisation. Romana Pernisch, Daniele Dell'Aglio, Abraham Bernstein |
J. Web Semant. | 3 |
| 2020 | Differentially Private Stream Processing for the Semantic WebabstractData often contains sensitive information, which poses a major obstacle to publishing it. Some suggest to obfuscate the data or only releasing some data statistics. These approaches have, however, been shown to provide insufficient safeguards against de-anonymisation. Recently, differential privacy (DP), an approach that injects noise into the query answers to provide statistical privacy guarantees, has emerged as a solution to release sensitive data. This study investigates how to continuously release privacy-preserving histograms (or distributions) from online streams of sensitive data by combining DP and semantic web technologies. We focus on distributions, as they are the basis for many analytic applications. Specifically, we propose SihlQL, a query language that processes RDF streams in a privacy-preserving fashion. SihlQL builds on top of SPARQL and the w-event DP framework. We show how some peculiarities of w-event privacy constrain the expressiveness of SihlQL queries. Addressing these constraints, we propose an extension of w-event privacy that provides answers to a larger class of queries while preserving their privacy. To evaluate SihlQL, we implemented a prototype engine that compiles queries to Apache Flink topologies and studied its privacy properties using real-world data from an IPTV provider and an online e-commerce web site. Daniele Dell'Aglio, Abraham Bernstein |
WWW | 2 |
| 2019 | Collaborative Streaming: Trust Requirements for Price SharingabstractStream Processing (SP) is an important Big Data technology enabling continuous querying of data streams. The stream setting offers the opportunity to exploit synergies and, theoretically, share the access and processing costs between multiple different collaborators. But what should be the monetary contribution of each consumer when they do not trust each other and have varying valuations of the differing outcomes? In this article, we present Collaborative Stream Processing (CSP), a model where the costs, which are set exogenously by providers, are shared between multiple consumers, the collaborators. For this, we identify three important requirements for CSP to establish trust between the collaborators and propose a CSP algorithm, ENCSPA, adhering to these requirements. Based on the collaborators' outcome valuations and the costs of the raw data streams, ENCSPA computes the payment for each collaborator. At the same time, ENCSPA ensures that no collaborator has an incentive to manipulate the system by providing misinformation about her/his value, budget, or time limit. We show that ENCSPA can calculate payments in a reasonable amount of time for up to one thousand collaborators. Tobias Grubenmann, Daniele Dell'Aglio, Abraham Bernstein |
IEEE BigData | 3 |
| 2019 | Interaction Embeddings for Prediction and Explanation in Knowledge GraphsabstractKnowledge graph embedding aims to learn distributed representations for entities and relations, and is proven to be effective in many applications. Crossover interactions -- bi-directional effects between entities and relations --- help select related information when predicting a new triple, but haven't been formally discussed before. In this paper, we propose CrossE, a novel knowledge graph embedding which explicitly simulates crossover interactions. It not only learns one general embedding for each entity and relation as most previous methods do, but also generates multiple triple specific embeddings for both of them, named interaction embeddings. We evaluate embeddings on typical link prediction tasks and find that CrossE achieves state-of-the-art results on complex and more challenging datasets. Furthermore, we evaluate embeddings from a new perspective -- giving explanations for predicted triples, which is important for real applications. In this work, an explanation for a triple is regarded as a reliable closed-path between the head and the tail entity. Compared to other baselines, we show experimentally that CrossE, benefiting from interaction embeddings, is more capable of generating reliable explanations to support its predictions. Wen Zhang 0015, Bibek Paudel, Wei Zhang 0127, Abraham Bernstein, Huajun Chen |
WSDM | 4 |
| 2019 | Iteratively Learning Embeddings and Rules for Knowledge Graph ReasoningabstractReasoning is essential for the development of large knowledge graphs, especially for completion, which aims to infer new triples based on existing ones. Both rules and embeddings can be used for knowledge graph reasoning and they have their own advantages and difficulties. Rule-based reasoning is accurate and explainable but rule learning with searching over the graph always suffers from efficiency due to huge search space. Embedding-based reasoning is more scalable and efficient as the reasoning is conducted via computation between embeddings, but it has difficulty learning good representations for sparse entities because a good embedding relies heavily on data richness. Based on this observation, in this paper we explore how embedding and rule learning can be combined together and complement each other's difficulties with their advantages. We propose a novel framework IterE iteratively learning embeddings and rules, in which rules are learned from embeddings with proper pruning strategy and embeddings are learned from existing triples and new triples inferred by rules. Evaluations on embedding qualities of IterE show that rules help improve the quality of sparse entity embeddings and their link prediction results. We also evaluate the efficiency of rule learning and quality of rules from IterE compared with AMIE+, showing that IterE is capable of generating high quality rules more efficiently. Experiments show that iteratively learning embeddings and rules benefit each other during learning and prediction. Wen Zhang 0015, Bibek Paudel, Jiaoyan Chen 0001, Wei Zhang 0127, Abraham Bernstein, Huajun Chen |
WWW | 7 |
| 2019 | A comparative survey of recent natural language interfaces for databasesabstractOver the last few years, natural language interfaces (NLI) for databases have gained significant traction both in academia and industry. These systems use very different approaches as described in recent survey papers. However, these systems have not been systematically compared against a set of benchmark questions in order to rigorously evaluate their functionalities and expressive power. In this paper, we give an overview over 24 recently developed NLIs for databases. Each of the systems is evaluated using a curated list of ten sample questions to show their strengths and weaknesses. We categorize the NLIs into four groups based on the methodology they are using: keyword-, pattern-, parsing- and grammar-based NLI. Overall, we learned that keyword-based systems are enough to answer simple questions. To solve more complex questions involving subqueries, the system needs to apply some sort of parsing to identify structural dependencies. Grammar-based systems are overall the most powerful ones, but are highly dependent on their manually designed rules. In addition to providing a systematic analysis of the major systems, we derive lessons learned that are vital for designing NLIs that can answer a wide range of user questions. Katrin Affolter, Kurt Stockinger, Abraham Bernstein |
VLDB J. | 3 |
| 2018 | Distributed Stream Consistency Checking
Shen Gao, Daniele Dell'Aglio, Jeff Z. Pan, Abraham Bernstein |
ICWE | 4 |
| 2018 | Aligning Knowledge Base and Document Embedding Models Using Regularized Multi-Task Learning
Matthias Baumgartner, Wen Zhang 0015, Bibek Paudel, Daniele Dell'Aglio, Huajun Chen, Abraham Bernstein |
ISWC (1) | 6 |
| 2018 | Financing the Web of Data with Delayed-Answer AuctionsabstractThe World Wide Web is a massive network of interlinked documents. One of the reasons the World Wide Web is so successful is the fact that most content is available free of any charge. Inspired by the success of the World Wide Web, the Web of Data applies the same strategy of interlinking to data. To this point, most of data in the Web of Data is also free of charge. The fact that the data is freely available raises the question of financing these services, however. As we will discuss in this paper, advertisement and donations cannot easily be applied to this new setting. To create incentives to subsidize data providers, we propose that sponsors should pay the providers to promote sponsored data. In return, sponsored data will be privileged over non-sponsored data. Since it is not possible to enforce a certain ordering on the data the user will receive, we propose to split up the data into different batches and deliver these batches with different delays. In this way, we can privilege sponsored data without withholding any non-sponsored data from the user. In this paper, we introduce a new concept of a delayed-answer auction, where sponsors can pay to prioritize their data. We introduce a new model which captures the particular situation when a user access data in the Web of Data. We show how the weighted Vickrey-Clarke-Groves auction mechanism can be applied to our scenario and we discuss how certain parameters can influence the nature of our auction. With our new concept, we build a first step to a free yet financial sustainable Web of Data. Tobias Grubenmann, Abraham Bernstein, Dmitry Moor, Sven Seuken |
WWW | 2 |
| 2017 | Expert Estimates for Feature Relevance are ImperfectabstractAn early step in the knowledge discovery process is deciding on what data to look at when trying to predict a given target variable. Most of KDD so far is focused on the workflow after data has been obtained, or settings where data is readily available and easily integrable for model induction. However, in practice, this is rarely the case, and many times data requires cleaning and transformation before it can be used for feature selection and knowledge discovery. In such environments, it would be costly to obtain and integrate data that is not relevant to the predicted target variable. To reduce the risk of such scenarios in practice, we often rely on experts to estimate the value of potential data based on its meta information (e.g. its description). However, as we will find in this paper, experts perform abysmally at this task. We therefore developed a methodology, KrowDD, to help humans estimate how relevant a dataset might be based on such meta data. We evaluate KrowDD on 3 real-world problems and compare its relevancy estimates with data scientists' and domain experts'. Our findings indicate large possible cost savings when using our tool in bias-free environments, which may pave the way for lowering the cost of classifier design in practice. Patrick M. De Boer, Marcel C. Bühler, Abraham Bernstein |
DSAA | 3 |
| 2017 | Break the Windows: Explicit State Management for Stream Processing Systems
Alessandro Margara, Daniele Dell'Aglio, Abraham Bernstein |
EDBT | 3 |
| 2017 | Fewer Flops at the Top: Accuracy, Diversity, and Regularization in Two-Class Collaborative FilteringabstractIn most existing recommender systems, implicit or explicit interactions are treated as positive links and all unknown interactions are treated as negative links. The goal is to suggest new links that will be perceived as positive by users. However, as signed social networks and newer content services become common, it is important to distinguish between positive and negative preferences. Even in existing applications, the cost of a negative recommendation could be high when people are looking for new jobs, friends, or places to live. Bibek Paudel, Thilo Haas, Abraham Bernstein |
RecSys | 3 |
| 2017 | Challenges of Source Selection in the WoD
Tobias Grubenmann, Abraham Bernstein, Dmitry Moor, Sven Seuken |
ISWC (1) | 2 |
| 2016 | Planning Ahead: Stream-Driven Linked-Data Access Under Update-Budget Constraints
Shen Gao, Daniele Dell'Aglio, Soheila Dehghanzadeh, Abraham Bernstein, Emanuele Della Valle, Alessandra Mileo |
ISWC (1) | 4 |
| 2016 | PPLib: Toward the Automated Generation of Crowd Computing Programs Using Process Recombination and Auto-ExperimentationabstractCrowdsourcing is increasingly being adopted to solve simple tasks such as image labeling and object tagging, as well as more complex tasks, where crowd workers collaborate in processes with interdependent steps. For the whole range of complexity, research has yielded numerous patterns for coordinating crowd workers in order to optimize crowd accuracy, efficiency, and cost. Process designers, however, often don't know which pattern to apply to a problem at hand when designing new applications for crowdsourcing. In this article, we propose to solve this problem by systematically exploring the design space of complex crowdsourced tasks via automated recombination and auto-experimentation for an issue at hand. Specifically, we propose an approach to finding the optimal process for a given problem by defining the deep structure of the problem in terms of its abstract operators, generating all possible alternatives via the (re)combination of the abstract deep structure with concrete implementations from a Process Repository, and then establishing the best alternative via auto-experimentation. To evaluate our approach, we implemented PPLib (pronounced “People Lib”), a program library that allows for the automated recombination of known processes stored in an easily extensible Process Repository. We evaluated our work by generating and running a plethora of process candidates in two scenarios on Amazon's Mechanical Turk followed by a meta-evaluation, where we looked at the differences between the two evaluations. Our first scenario addressed the problem of text translation, where our automatic recombination produced multiple processes whose performance almost matched the benchmark established by an expert translation. In our second evaluation, we focused on text shortening; we automatically generated 41 crowd process candidates, among them variations of the well-established Find-Fix-Verify process. While Find-Fix-Verify performed well in this setting, our recombination engine produced five processes that repeatedly yielded better results. We close the article by comparing the two settings where the Recombinator was used, and empirically show that the individual processes performed differently in the two settings, which led us to contend that there is no unifying formula, hence emphasizing the necessity for recombination. Patrick M. De Boer, Abraham Bernstein |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2015 | Workload scheduling in distributed stream processors using graph partitioningabstractWith ever increasing data volumes, large compute clusters that process data in a distributed manner have become prevalent in industry. For distributed stream processing platforms (such as Storm) the question of how to distribute workload to available machines, has important implications for the overall performance of the system. We present a workload scheduling strategy that is based on a graph partitioning algorithm. The scheduler is application agnostic: it collects the communication behavior of running applications and creates the schedules by partitioning the resulting communication graph using the METIS graph partitioning software. As we build upon graph partitioning algorithms that have been shown to scale to very large graphs, our approach can cope with topologies with millions of tasks. While the experiments in this paper assume static data loads, our approach could also be used in a dynamic setting. We implemented our proposed algorithm for the Storm stream processing system and evaluated it on a commodity cluster with up to 80 machines. The evaluation was conducted on four different use cases - three using synthetic data loads and one application that processes real data. We compared our algorithm against two state-of-the-art scheduler implementations and show that our approach offers significant improvements in terms of resource utilization, enabling higher throughput at reduced network loads. We show that these improvements can be achieved while maintaining a balanced workload in terms of CPU usage and bandwidth consumption across the cluster. We also found that the performance advantage increases with message size, providing an important insight for stream-processing approaches based on micro-batching. Lorenz Fischer, Abraham Bernstein |
IEEE BigData | 2 |
| 2015 | Approximate Continuous Query Answering over Streams and Dynamic Linked Data Sets
Soheila Dehghanzadeh, Daniele Dell'Aglio, Shen Gao, Emanuele Della Valle, Alessandra Mileo, Abraham Bernstein |
ICWE | 6 |
| 2015 | Blockbusters and Wallflowers: Accurate, Diverse, and Scalable Recommendations with Random WalksabstractUser satisfaction is often dependent on providing accurate and diverse recommendations. In this paper, we explore scalable algorithms that exploit random walks as a sampling technique to obtain diverse recommendations without compromising on accuracy. Specifically, we present a novel graph vertex ranking recommendation algorithm called RP^3_beta that re-ranks items based on 3-hop random walk transition probabilities. We show empirically, that RP^3_beta provides accurate recommendations with high long-tail item frequency at the top of the recommendation list. We also present scalable approximate versions of RP^3_beta and the two most accurate previously published vertex ranking algorithms based on random walk transition probabilities and show that these approximations converge with increasing number of samples. Fabian Christoffel, Bibek Paudel, Chris Newell, Abraham Bernstein |
RecSys | 4 |
| 2015 | Timely Semantics: A Study of a Stream-Based Ranking System for Entity Relationships
Lorenz Fischer, Roi Blanco, Peter Mika, Abraham Bernstein |
ISWC (2) | 4 |
| 2015 | Random Walk TripleRush: Asynchronous Graph Querying and SamplingabstractMost Semantic Web applications rely on querying graphs, typically by using SPARQL with a triple store. Increasingly, applications also analyze properties of the graph structure to compute statistical inferences. The current Semantic Web infrastructure, however, does not efficiently support such operations. This forces developers to extract the relevant data for external statistical post-processing. In this paper we propose to rethink query execution in a triple store as a highly parallelized asynchronous graph exploration on an active index data structure. This approach also allows to integrate SPARQL-querying with the sampling of graph properties. Philip Stutz, Bibek Paudel, Mihaela Verman, Abraham Bernstein |
WWW | 4 |
| 2014 | The CLOCK Data-Aware Eviction Approach: Towards Processing Linked Data Streams with Limited Resources
Shen Gao, Thomas Scharrenbach, Abraham Bernstein |
ESWC | 3 |
| 2014 | "Semantics Inside!" But Let's Not Tell the Data Miners: Intelligent Support for Data Mining
Jörg-Uwe Kietz, Floarea Serban, Simon Fischer 0001, Abraham Bernstein |
ESWC | 4 |
| 2014 | Behavior-Based Quality Assurance in Crowdsourcing MarketsabstractQuality assurance in crowdsourcing markets has appeared to be an acute problem over the last years. We propose a quality control method inspired by Statistical Process Control (SPC), commonly used to control output quality in production processes and characterized by relying on time-series data. Behavioral traces of users may play a key role in evaluating the performance of work done on crowdsourcing platforms. Therefore, in our experiment we explore fifteen behavioral traces for their ability to recognize the drop in work quality. Preliminary results indicate that our method has a high potential for real-time detection and signaling a drop in work quality. Michael Feldman 0001, Abraham Bernstein |
HCOMP | 2 |
| 2014 | Querying a messy web of data with Avalanche
Cosmin Basca, Abraham Bernstein |
J. Web Semant. | 2 |
| 2013 | Seven Commandments for Benchmarking Semantic Flow Processing Systems
Thomas Scharrenbach, Jacopo Urbani, Alessandro Margara, Emanuele Della Valle, Abraham Bernstein |
ESWC | 5 |
| 2012 | Semantic Web/LD at a Crossroads: Into the Garbage Can or To Theory?
Abraham Bernstein |
ESWC | 1 |
| 2011 | SAWSDL-iMatcher: A customizable and effective Semantic Web Service matchmaker
Dengping Wei, Ting Wang 0009, Ji Wang 0001, Abraham Bernstein |
J. Web Semant. | 4 |
| 2010 | When process data quality affects the number of bugs: Correlations in software engineering datasetsabstractSoftware engineering process information extracted from version control systems and bug tracking databases are widely used in empirical software engineering. In prior work, we showed that these data are plagued by quality deficiencies, which vary in its characteristics across projects. In addition, we showed that those deficiencies in the form of bias do impact the results of studies in empirical software engineering. While these findings affect software engineering researchers the impact on practitioners has not yet been substantiated. In this paper we, therefore, explore (i) if the process data quality and characteristics have an influence on the bug fixing process and (ii) if the process quality as measured by the process data has an influence on the product (i.e., software) quality. Specifically, we analyze six Open Source as well as two Closed Source projects and show that process data quality and characteristics have an impact on the bug fixing process: the high rate of empty commit messages in Eclipse, for example, correlates with the bug report quality. We also show that the product quality - measured by number of bugs reported - is affected by process data quality measures. These findings have the potential to prompt practitioners to increase the quality of their software process and its associated data quality. Adrian Bachmann, Abraham Bernstein |
MSR | 2 |
| 2010 | Signal/Collect: Graph Algorithms for the (Semantic) Web
Philip Stutz, Abraham Bernstein, William W. Cohen |
ISWC (1) | 2 |
| 2010 | Evaluating the usability of natural language query languages and interfaces to Semantic Web knowledge bases
Esther Kaufmann, Abraham Bernstein |
J. Web Semant. | 2 |
| 2010 | Semantic web enabled software analysis
Jonas Tappolet, Christoph Kiefer, Abraham Bernstein |
J. Web Semant. | 3 |
| 2009 | Applied Temporal RDF: Efficient Temporal Querying of RDF Data with SPARQL
Jonas Tappolet, Abraham Bernstein |
ESWC | 2 |
| 2009 | Tracking concept drift of software projects using defect prediction qualityabstractDefect prediction is an important task in the mining of software repositories, but the quality of predictions varies strongly within and across software projects. In this paper we investigate the reasons why the prediction quality is so fluctuating due to the altering nature of the bug (or defect) fixing process. Therefore, we adopt the notion of a concept drift, which denotes that the defect prediction model has become unsuitable as set of influencing features has changed - usually due to a change in the underlying bug generation process (i.e., the concept). We explore four open source projects (Eclipse, OpenOffice, Netbeans and Mozilla) and construct file-level and project-level features for each of them from their respective CVS and Bugzilla repositories. We then use this data to build defect prediction models and visualize the prediction quality along the time axis. These visualizations allow us to identify concept drifts and - as a consequence - phases of stability and instability expressed in the level of defect prediction quality. Further, we identify those project features, which are influencing the defect prediction quality using both a tree induction-algorithm and a linear regression model. Our experiments uncover that software systems are subject to considerable concept drifts in their evolution history. Specifically, we observe that the change in number of authors editing a file and the number of defects fixed by them contribute to a project's concept drift and therefore influence the defect prediction quality. Our findings suggest that project managers using defect prediction models for decision making should be aware of the actual phase of stability or instability due to a potential concept drift. Jayalath B. Ekanayake, Jonas Tappolet, Harald C. Gall, Abraham Bernstein |
MSR | 4 |
| 2008 | The Creation and Evaluation of iSPARQL Strategies for Matchmaking
Christoph Kiefer, Abraham Bernstein |
ESWC | 2 |
| 2008 | Adding Data Mining Support to SPARQL Via Statistical Relational Learning Methods
Christoph Kiefer, Abraham Bernstein, André Locher |
ESWC | 2 |
| 2008 | How to learn enough data mining to be dangerous in 60 minutesabstractThe field of data mining provides some methods highly relevant to researchers when mining software repositories. Whether one predicts bug locations, discovers hidden architectural structures and software patterns, or identifies experts of modules, data mining algorithms are usually the working horses for these studies. The goal of this tutorial is to convey some of the most relevant theoretical foundations and practical issues when using data mining algorithms.The tutorial will first discuss the usual data mining tasks (prediction, filtering, smoothing, and elucidation of the most likely explanation or structure). Then, it will introduce a general framework for data mining paving the way to explain the functionality of some of the most used data mining algorithms. The tutorial will close with an overview over the typical evaluation methods for induced results and a number of pointers for further study. Where possible, it will use examples from software engineering. Abraham Bernstein |
MSR | 1 |
| 2008 | Enhancing Semantic Web Services with Inheritance
Simon Ferndriger, Abraham Bernstein, Jin Song Dong 0001, Yuzhang Feng, Yuan-Fang Li, Jane Hunter 0001 |
ISWC | 2 |
| 2008 | SPARQL basic graph pattern optimization using selectivity estimationabstractIn this paper, we formalize the problem of Basic Graph Pattern (BGP) optimization for SPARQL queries and main memory graph implementations of RDF data. We define and analyze the characteristics of heuristics for selectivity-based static BGP optimization. The heuristics range from simple triple pattern variable counting to more sophisticated selectivity estimation techniques. Customized summary statistics for RDF data enable the selectivity estimation of joined triple patterns and the development of efficient heuristics. Using the Lehigh University Benchmark (LUBM), we evaluate the performance of the heuristics for the queries provided by the LUBM and discuss some of them in more details. Markus Stocker, Andy Seaborne, Abraham Bernstein, Christoph Kiefer, Dave Reynolds |
WWW | 3 |
| 2008 | Hexastore: sextuple indexing for semantic web data managementabstractDespite the intense interest towards realizing the Semantic Web vision, most existing RDF data management schemes are constrained in terms of efficiency and scalability. Still, the growing popularity of the RDF format arguably calls for an effort to offset these drawbacks. Viewed from a relational-database perspective, these constraints are derived from the very nature of the RDF data model, which is based on a triple format. Recent research has attempted to address these constraints using a vertical-partitioning approach, in which separate two-column tables are constructed for each property. However, as we show, this approach suffers from similar scalability drawbacks on queries that are not bound by RDF property value. In this paper, we propose an RDF storage scheme that uses the triple nature of RDF as an asset. This scheme enhances the vertical partitioning idea and takes it to its logical conclusion. RDF data is indexed in six possible ways, one for each possible ordering of the three RDF elements. Each instance of an RDF element is associated with two vectors; each such vector gathers elements of one of the other types, along with lists of the third-type resources attached to each vector element. Hence, a sextuple-indexing scheme emerges. This format allows for quick and scalable general-purpose query processing; it confers significant advantages (up to five orders of magnitude) compared to previous approaches for RDF data management, at the price of a worst-case five-fold increase in index space. We experimentally document the advantages of our approach on real-world and synthetic data sets with practical queries. Cathrin Weiss, Panagiotis Karras, Abraham Bernstein |
Proc. VLDB Endow. | 3 |
| 2007 | The NExT System: Towards True Dynamic Adaptations of Semantic Web Service Compositions
Abraham Bernstein, Michael Dänzer |
ESWC | 1 |
| 2007 | Semantic Process Retrieval with iSPARQL
Christoph Kiefer, Abraham Bernstein, Mark Klein 0001, Markus Stocker |
ESWC | 2 |
| 2006 | Detecting Similarities in Ontologies with the SOQA-SimPack Toolkit
Patrick Ziegler, Christoph Kiefer, Christoph Sturm, Klaus R. Dittrich, Abraham Bernstein |
EDBT | 5 |
| 2006 | Entropy-based Concept Shift DetectionabstractWhen monitoring sensory data (e.g., from a wearable device) the context oftentimes changes abruptly: people move from one situation (e.g., working quietly in their office) to another (e.g., being interrupted by one's manager). These context changes can be treated like concept shifts, since the underlying data generator (the concept) changes while moving from one context situation to another. We present an entropy based measure for data streams that is suitable to detect concept shifts in a reliable, noise-resistant, fast, and computationally efficient way. We assess the entropy measure under different concept shift conditions. To support our claims we illustrate the concept shift behavior of the stream entropy. We also present a simple algorithm control approach to show how useful and reliable the information obtained by the entropy measure is compared to a ensemble learner as well as an experimentally inferred upper limit. Our analysis is based on three large synthetic data sets representing real, virtual, and a combination of both concept drifts under different noise conditions (up to 50%). Last but not least, we demonstrate the usefulness of the entropy based measure context switch indication in a real world application in the context-awareness/wearable computing domain. Peter Vorburger, Abraham Bernstein |
ICDM | 2 |
| 2006 | GINO - A Guided Input Natural Language Ontology Editor
Abraham Bernstein, Esther Kaufmann |
ISWC | 1 |
| 2006 | Generic similarity detection in ontologies with the SOQA-SimPack toolkitabstractOntologies are increasingly used to represent the intended real-world semantics of data and services in information systems. Unfortunately, different data sources often do not relate to the same ontologies when describing their semantics. Consequently, it is desirable to have information about the similarity between ontology concepts for ontology alignment and integration. In this demo, we present the SOQA-SimPack Toolkit (SST), an ontology language independent Java API that enables generic similarity detection and visualization in ontologies. We demonstrate SST's usefulness with the SOQA-SimPack Toolkit Browser that allows users to graphically perform similarity calculations in ontologies. Patrick Ziegler, Christoph Kiefer, Christoph Sturm, Klaus R. Dittrich, Abraham Bernstein |
SIGMOD Conference | 5 |
| 2005 | Querying Ontologies: A Controlled English Interface for End-Users
Abraham Bernstein, Esther Kaufmann, Anne Göhring, Christoph Kiefer |
ISWC | 1 |
| 2005 | Toward Intelligent Assistance for a Data Mining Process: An Ontology-Based Approach for Cost-Sensitive ClassificationabstractA data mining (DM) process involves multiple stages. A simple, but typical, process might include preprocessing data, applying a data mining algorithm, and postprocessing the mining results. There are many possible choices for each stage, and only some combinations are valid. Because of the large space and nontrivial interactions, both novices and data mining specialists need assistance in composing and selecting DM processes. Extending notions developed for statistical expert systems we present a prototype intelligent discovery assistant (IDA), which provides users with 1) systematic enumerations of valid DM processes, in order that important, potentially fruitful options are not overlooked, and 2) effective rankings of these valid processes by different criteria, to facilitate the choice of DM processes to execute. We use the prototype to show that an IDA can indeed provide useful enumerations and effective rankings in the context of simple classification processes. We discuss how an IDA could be an important tool for knowledge sharing among a team of data miners. Finally, we illustrate the claims with a demonstration of cost-sensitive classification using a more complicated process and data from the 1998 KDDCUP competition. Abraham Bernstein, Foster J. Provost, Shawndra Hill |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2002 | Towards High-Precision Service Retrieval
Abraham Bernstein, Mark Klein 0001 |
ISWC | 1 |