Mohamed Yakout

dblp:48/725 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 6 first-authorArtificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Data integration and cleaning · 61% Knowledge graphs · 23% Information retrieval · 16%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Network and information security
1 paper
Cryptographic protocols and secure computation · 77% Privacy and data protection · 23%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning › data preprocessing › data cleaning
data repair
0.432013
Don't be SCAREd: use SCalable Automatic REpairing with maximal likelihood and bounded changes · SIGMOD Conference 2013
Guided data repair · Proc. VLDB Endow. 2011
GDR: a system for guided data repair · SIGMOD Conference 2010
Knowledge graphs
entity linking
0.222010
Behavior Based Record Linkage · Proc. VLDB Endow. 2010
Efficient Private Record Linkage · ICDE 2009
Information retrieval › user interaction
user feedback
0.222011
Guided data repair · Proc. VLDB Endow. 2011
GDR: a system for guided data repair · SIGMOD Conference 2010
Data integration and cleaning › table understanding › table annotation
attribute discovery
0.112012
InfoGather: entity augmentation and attribute discovery by holistic matching with web tables · SIGMOD Conference 2012
Data integration and cleaning › data curation
entity augmentation
0.112012
InfoGather: entity augmentation and attribute discovery by holistic matching with web tables · SIGMOD Conference 2012
Data integration and cleaning
entity resolution
0.112010
Behavior Based Record Linkage · Proc. VLDB Endow. 2010
Visualization and visual analytics › visual analytics
spatiotemporal visual analytics
0.112010
A Visual Analytics Approach to Understanding Spatiotemporal Hotspots · IEEE Trans. Vis. Comput. Graph. 2010
Knowledge graphs › entity linking
privacy-preserving record linkage
0.112009
Efficient Private Record Linkage · ICDE 2009
Cryptographic protocols and secure computation
secure multiparty computation
0.112009
Efficient Private Record Linkage · ICDE 2009
Privacy and data protection
privacy-preserving data analysis
0.012009
Efficient Private Record Linkage · ICDE 2009

Methods — techniques the papers use, named apart from their topics

decision theory · 0.2active learning · 0.2cryptographic protocols · 0.2statistical machine learning · 0.2maximum likelihood · 0.2value of information · 0.1demographic filtering · 0.1candidate generation · 0.1alert detection algorithms · 0.1
YearPublicationVenuePosition
2015 Holistic entity matching across knowledge graphs
abstract
Entity matching is the problem of determining if two entities in a data set refer to the same real-world object. In the last decade a growing number of large-scale knowledge bases have been created online. Tools for automatically aligning these sources would make it possible to unify them in a structured knowledge and to answer complex queries. Here we present Holistic Entity Matching (HolisticEM), an algorithm based on Personalized Page Rank for aligning instances in large knowledge bases. It consists of two steps. First, a graph of potential matching pairs is constructed; second, local and global information from the relationship graph is propagated via Personalized Page Rank. We demonstrate that HolisticEM performs competitively and can efficiently handle databases with 110M and 203M entities accurately resolving 1.6M of matching entity pairs.
Maria Pershina, Mohamed Yakout, Kaushik Chakrabarti
IEEE BigData2
2013 Don't be SCAREd: use SCalable Automatic REpairing with maximal likelihood and bounded changes
abstract
Various computational procedures or constraint-based methods for data repairing have been proposed over the last decades to identify errors and, when possible, correct them. However, these approaches have several limitations including the scalability and quality of the values to be used in replacement of the errors. In this paper, we propose a new data repairing approach that is based on maximizing the likelihood of replacement data given the data distribution, which can be modeled using statistical machine learning techniques. This is a novel approach combining machine learning and likelihood methods for cleaning dirty databases by value modification. We develop a quality measure of the repairing updates based on the likelihood benefit and the amount of changes applied to the database. We propose SCARE (SCalable Automatic REpairing), a systematic scalable framework that follows our approach. SCARE relies on a robust mechanism for horizontal data partitioning and a combination of machine learning techniques to predict the set of possible updates. Due to data partitioning, several updates can be predicted for a single record based on local views on each data partition. Therefore, we propose a mechanism to combine the local predictions and obtain accurate final predictions. Finally, we experimentally demonstrate the effectiveness, efficiency, and scalability of our approach on real-world datasets in comparison to recent data cleaning approaches.
Mohamed Yakout, Laure Berti-Équille, Ahmed K. Elmagarmid
SIGMOD Conference1
2012 InfoGather: entity augmentation and attribute discovery by holistic matching with web tables
abstract
The Web contains a vast corpus of HTML tables, specifically entity attribute tables. We present three core operations, namely entity augmentation by attribute name, entity augmentation by example and attribute discovery, that are useful for "information gathering" tasks (e.g., researching for products or stocks). We propose to use web table corpus to perform them automatically. We require the operations to have high precision and coverage, have fast (ideally interactive) response times and be applicable to any arbitrary domain of entities. The naive approach that attempts to directly match the user input with the web tables suffers from poor precision and coverage.
Mohamed Yakout, Kris Ganjam, Kaushik Chakrabarti, Surajit Chaudhuri
SIGMOD Conference1
2011 Guided data repair
abstract
In this paper we present GDR, a Guided Data Repair framework that incorporates user feedback in the cleaning process to enhance and accelerate existing automatic repair techniques while minimizing user involvement. GDR consults the user on the updates that are most likely to be beneficial in improving data quality. GDR also uses machine learning methods to identify and apply the correct updates directly to the database without the actual involvement of the user on these specific updates. To rank potential updates for consultation by the user, we first group these repairs and quantify the utility of each group using the decision-theory concept of value of information (VOI). We then apply active learning to order updates within a group based on their ability to improve the learned model. User feedback is used to repair the database and to adaptively refine the training set for the model. We empirically evaluate GDR on a real-world dataset and show significant improvement in data quality using our user guided repairing process. We also, assess the trade-off between the user efforts and the resulting data quality.
Mohamed Yakout, Ahmed K. Elmagarmid, Jennifer Neville, Mourad Ouzzani, Ihab F. Ilyas
Proc. VLDB Endow.1
2010 GDR: a system for guided data repair
abstract
Improving data quality is a time-consuming, labor-intensive and often domain specific operation. Existing data repair approaches are either fully automated or not efficient in interactively involving the users. We present a demo of GDR, a Guided Data Repair system that uses a novel approach to efficiently involve the user alongside automatic data repair techniques to reach better data quality as quickly as possible. Specifically, GDR generates data repairs and acquire feedback on them that would be most beneficial in improving the data quality. GDR quantifies the data quality benefit of generated repairs by combining mechanisms from decision theory and active learning. Based on these benefit scores, groups of repairs are ranked and displayed to the user. User feedback is used to train a machine learning component to eventually replace the user in deciding on the validity of a suggested repair. We describe how the generated repairs are ranked and displayed to the user in a "useful-looking" way and demonstrate how data quality can be effectively improved with minimal feedback from the user.
Mohamed Yakout, Ahmed K. Elmagarmid, Jennifer Neville, Mourad Ouzzani
SIGMOD Conference1
2010 Behavior Based Record Linkage
abstract
In this paper, we present a new record linkage approach that uses entity behavior to decide if potentially different entities are in fact the same. An entity's behavior is extracted from a transaction log that records the actions of this entity with respect to a given data source. The core of our approach is a technique that merges the behavior of two possible matched entities and computes the gain in recognizing behavior patterns as their matching score. The idea is that if we obtain a well recognized behavior after merge, then most likely, the original two behaviors belong to the same entity as the behavior becomes more complete after the merge. We present the necessary algorithms to model entities' behavior and compute a matching score for them. To improve the computational efficiency of our approach, we precede the actual matching phase with a fast candidate generation that uses a "quick and dirty" matching method. Extensive experiments on real data show that our approach can significantly enhance record linkage quality while being practical for large transaction logs.
Mohamed Yakout, Ahmed K. Elmagarmid, Hazem Elmeleegy, Mourad Ouzzani, Alan Qi
Proc. VLDB Endow.1
2010 A Visual Analytics Approach to Understanding Spatiotemporal Hotspots
abstract
As data sources become larger and more complex, the ability to effectively explore and analyze patterns among varying sources becomes a critical bottleneck in analytic reasoning. Incoming data contain multiple variables, high signal-to-noise ratio, and a degree of uncertainty, all of which hinder exploration, hypothesis generation/exploration, and decision making. To facilitate the exploration of such data, advanced tool sets are needed that allow the user to interact with their data in a visual environment that provides direct analytic capability for finding data aberrations or hotspots. In this paper, we present a suite of tools designed to facilitate the exploration of spatiotemporal data sets. Our system allows users to search for hotspots in both space and time, combining linked views and interactive filtering to provide users with contextual information about their data and allow the user to develop and explore their hypotheses. Statistical data models and alert detection algorithms are provided to help draw user attention to critical areas. Demographic filtering can then be further applied as hypotheses generated become fine tuned. This paper demonstrates the use of such tools on multiple geospatiotemporal data sets.
Ross Maciejewski, Stephen Rudolph, Ryan Hafen, Ahmad M. Abusalah, Mohamed Yakout, Mourad Ouzzani, William S. Cleveland, Shaun J. Grannis, David S. Ebert
IEEE Trans. Vis. Comput. Graph.5
2009 Efficient Private Record Linkage
abstract
Record linkage is the computation of the associations among records of multiple databases. It arises in contexts like the integration of such databases, online interactions and negotiations, and many others. The autonomous entities who wish to carry out the record matching computation are often reluctant to fully share their data. In such a framework where the entities are unwilling to share data with each other, the problem of carrying out the linkage computation without full data exchange has been called private record linkage. Previous private record linkage techniques have made use of a third party. We provide efficient techniques for private record linkage that improve on previous work in that (i) they make no use of a third party; (ii) they achieve much better performance than that of previous schemes in terms of execution time and quality of output (i.e., practically without false negatives and minimal false positives). Our software implementation provides experimental validation of our approach and the above claims.
Mohamed Yakout, Mikhail J. Atallah, Ahmed K. Elmagarmid
ICDE1