Siarhei Bykau

dblp:62/7525 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 4 first-authorArtificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Data integration and cleaning · 72% Web and social media mining · 14% Spatial and temporal data management · 5%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning › data fusion
conflict resolution
0.622018
A Framework to Integrate User Feedback for Rapid Conflict Resolution · ICDE 2018
Staging User Feedback toward Rapid Conflict Resolution in Data Fusion · SIGMOD Conference 2017
Data integration and cleaning
data fusion
0.622018
A Framework to Integrate User Feedback for Rapid Conflict Resolution · ICDE 2018
Staging User Feedback toward Rapid Conflict Resolution in Data Fusion · SIGMOD Conference 2017
Data integration and cleaning › data preprocessing
data cleaning
0.412020
Detecting and Preventing Confused Labels in Crowdsourced Data · Proc. VLDB Endow. 2020
Data integration and cleaning › data quality
label noise
0.412020
Detecting and Preventing Confused Labels in Crowdsourced Data · Proc. VLDB Endow. 2020
Web and social media mining › event detection
controversy detection
0.212015
Fine-grained controversy detection in Wikipedia · ICDE 2015
Web and social media mining › user-generated content
user-generated content analysis
0.212015
Fine-grained controversy detection in Wikipedia · ICDE 2015
Database theory
query answering
0.212013
A query answering system for data with evolution relationships · SIGMOD Conference 2013
Spatial and temporal data management
temporal databases
0.212013
A query answering system for data with evolution relationships · SIGMOD Conference 2013
Data mining › crowdsourcing
crowdsourced data
0.112020
Detecting and Preventing Confused Labels in Crowdsourced Data · Proc. VLDB Endow. 2020
Data integration and cleaning
truth discovery
0.112017
Staging User Feedback toward Rapid Conflict Resolution in Data Fusion · SIGMOD Conference 2017
Data integration and cleaning
data quality
0.112015
Fine-grained controversy detection in Wikipedia · ICDE 2015

Methods — techniques the papers use, named apart from their topics

ranking algorithm · 0.3information theory · 0.3decision theory · 0.3value of perfect information · 0.3decision-theoretic ordering · 0.3edit history analysis · 0.2steiner forest · 0.2
YearPublicationVenuePosition
2020 Detecting and Preventing Confused Labels in Crowdsourced Data
Evgeny Krivosheev, Siarhei Bykau, Fabio Casati, Sunil Prabhakar 0001
Proc. VLDB Endow.2
2018 A Framework to Integrate User Feedback for Rapid Conflict Resolution
abstract
Data fusion addresses the problem of consolidating data from disparate information providers into a single unified interface. The different data sources often provide conflicting information for the same data item. Recently, several automated data fusion models have been proposed to resolve conflicts and identify correct data. Although quite effective, these data fusion models do not achieve a close-to-perfect accuracy. We present the demonstration of a system that leverages users as first-class citizens to confirm data conflicts and rapidly improve the effectiveness of fusion. This demonstration is built on solutions proposed in our previous work [1]. To utilize the user judiciously, our system presents claims in an order that is the most beneficial to effectiveness of fusion across data items. We describe ranking algorithms that are built on concepts from information theory and decision theory, and do not need access to ground truth. We describe the user input framework and demonstrate how conflict resolution can be expedited with minimal feedback from the user. We show that: (a) the framework can be easily adopted to existing data fusion models without any internal changes to the models, and (b) the framework can integrate both perfect and imperfect feedback from users.
Romila Pradhan, Siarhei Bykau, Sunil Prabhakar 0001
ICDE2
2017 Staging User Feedback toward Rapid Conflict Resolution in Data Fusion
abstract
In domains such as the Web, sensor networks and social media, sources often provide conflicting information for the same data item. Several data fusion techniques have been proposed recently to resolve conflicts and identify correct data. The performance of these fusion systems, while quite accurate, is far from perfect. In this paper, we propose to leverage user feedback for validating data conflicts and rapidly improving the performance of fusion. To present the most beneficial data items for the user to validate, we take advantage of the level of consensus among sources, and the output of fusion to generate an effective ordering of items. We first evaluate data items individually, and then define a novel decision-theoretic framework based on the concept of value of perfect information (VPI) to order items by their ability to boost the performance of fusion. We further derive approximate formulae to scale up the decision-theoretic framework to large-scale data. We empirically evaluate our algorithms on three real-world datasets with different characteristics, and show that the accuracy of fusion can be significantly improved even while requesting feedback on a few data items. We also show that the performance of the proposed methods depends on the characteristics of data, and assess the trade-off between the amount of feedback acquired, and the effectiveness and efficiency of the methods.
Romila Pradhan, Siarhei Bykau, Sunil Prabhakar 0001
SIGMOD Conference2
2017 "Tell me more" using Ladders in Wikipedia
abstract
We focus on the problem of "tell me more" information related to a given fact in Wikipedia. We use the novel notion of role to link information in an infobox with different places in the text of the same Wikipedia page (space) as well as information across different revisions of the page (time). In this way, it is possible to link together pieces of information that may not represent the same real world entity, yet have served in the same role. To achieve this, we introduce a novel structure called ladder that allows such spatial and temporal linking and we show how to effectively and efficiently construct such structures from Wikipedia data.
Siarhei Bykau, Divesh Srivastava, Yannis Velegrakis
WebDB1
2015 Fine-grained controversy detection in Wikipedia
abstract
The advent of Web 2.0 gave birth to a new kind of application where content is generated through the collaborative contribution of many different users. This form of content generation is believed to generate data of higher quality since the “wisdom of the crowds” makes its way into the data. However, a number of specific data quality issues appear within such collaboratively generated data. Apart from normal updates, there are cases of intentional harmful changes known as vandalism as well as naturally occurring disagreements on topics which don't have an agreed upon viewpoint, known as controversies. While much work has focused on identifying vandalism, there has been little prior work on detecting controversies, especially at a fine granularity. Knowing about controversies when processing user-generated content is essential to understand the quality of the data and the trust that should be given to them. Controversy detection is a challenging task, since in the highly dynamic context of user updates, one needs to differentiate among normal updates, vandalisms and actual controversies. We describe a novel technique that finds these controversial issues by analyzing the edits that have been performed on the data over time. We apply the developed technique on Wikipedia, the world's largest known collaboratively generated database and we show that our approach has higher precision and recall than baseline approaches as well as is capable of finding previously unknown controversies.
Siarhei Bykau, Flip Korn, Divesh Srivastava, Yannis Velegrakis
ICDE1
2013 A query answering system for data with evolution relationships
abstract
Evolving data has attracted considerable research attention. Researchers have focused on modeling and querying of schema/instance-level structural changes, such as, insertion, deletion and modification of attributes. Databases with such a functionality are known as temporal databases. A limitation of the temporal databases is that they treat changes as independent events, while often the appearance (or elimination) of some structure in the database is the result of an evolution of some existing structure. We claim that maintaining the causal relationship between the two structures is of major importance since it allows additional reasoning to be performed and answers to be generated for queries that previously had no answers. We present the TrenDS, a system for exploiting the evolution relationships between the structures in the database. In particular, our system combines different structures that are associated through evolution relationships into virtual structures to be used during query answering. The virtual structures define ``possible'' database instances, in a fashion similar to the possible worlds in the probabilistic databases. TrenDS uses a query answering mechanism that allows queries to be answered over these possible databases without materializing them. Evaluation of such queries raises many technical challenges, since it requires the discovery of Steiner forests on the evolution graphs.
Siarhei Bykau, Flavio Rizzolo, Yannis Velegrakis
SIGMOD Conference1
2011 Supporting queries spanning across phases of evolving artifacts using Steiner forests
abstract
The problem of managing evolving data has attracted considerable research attention. Researchers have focused on the modeling and querying of schema/instance-level structural changes, such as, addition, deletion and modification of attributes. Databases with such a functionality are known as temporal databases. A limitation of the temporal databases is that they treat changes as independent events, while often the appearance (or elimination) of some structure in the database is the result of an evolution of some existing structure. We claim that maintaining the causal relationship between the two structures is of major importance since it allows additional reasoning to be performed and answers to be generated for queries that previously had no answers. We present here a novel framework for exploiting the evolution relationships between the structures in the database. In particular, our system combines different structures that are associated through evolution relationships into virtual structures to be used during query answering. The virtual structures define "possible" database instances, in a fashion similar to the possible worlds in the probabilistic databases. The framework includes a query answering mechanism that allows queries to be answered over these possible databases without materializing them. Evaluation of such queries raises many interesting technical challenges, since it requires the discovery of Steiner forests on the evolution graphs. On this problem we have designed and implemented a new dynamic programming algorithm with exponential complexity in the size of the input query and polynomial complexity in terms of both the attribute and the evolution data sizes.
Siarhei Bykau, John Mylopoulos, Flavio Rizzolo, Yannis Velegrakis
CIKM1
2011 The Papyrus Digital Library: Discovering History in the News
Akrivi Katifori, Charalampos Nikolaou, Manolis Platakis, Yannis E. Ioannidis, A. Tympas, Manolis Koubarakis, Nikos Sarris, V. Tountopoulos, Efstratios Tzoannos, Siarhei Bykau, Nadzeya Kiyavitskaya, Chrisa Tsinaraki, Yannis Velegrakis
TPDL10
2009 Modeling Associations through Intensional Attributes
Andrea Presa, Yannis Velegrakis, Flavio Rizzolo, Siarhei Bykau
ER4
2009 Modeling Concept Evolution: A Historical Perspective
Flavio Rizzolo, Yannis Velegrakis, John Mylopoulos, Siarhei Bykau
ER4