EDBT 2026 Demo / reviewers in the wild / expert
Paul Suganthan G. C.
dblp:137/2708 · also Paul Suganthan
· DBLP profile ↗
10ranked-venue papers
2as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Data integration and cleaning · 81% Information retrieval · 9% Database system architecture and tuning · 6% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 100% |
Topics — the 8 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning
entity matching |
1.8 | 6 | 2019 | Entity Matching Meets Data Science: A Progress Report from the Magellan Project · SIGMOD Conference 2019 CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity Matching · Proc. VLDB Endow. 2018 Smurf: Self-Service String Matching Using Random Forests · Proc. VLDB Endow. 2018 |
Data integration and cleaning › entity resolution
blocking and matching pipeline |
0.5 | 2 | 2016 | Magellan: Toward Building Entity Matching Management Systems over Data Science Stacks · Proc. VLDB Endow. 2016 Magellan: Toward Building Entity Matching Management Systems · Proc. VLDB Endow. 2016 |
Data integration and cleaning › data quality
data validation |
0.4 | 1 | 2020 | TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines · SIGMOD Conference 2020 |
Information retrieval
string matching |
0.3 | 1 | 2018 | Smurf: Self-Service String Matching Using Random Forests · Proc. VLDB Endow. 2018 |
Database system architecture and tuning › active database
rule management |
0.2 | 1 | 2015 | Why Big Data Industrial Systems Need Rules and What We Can Do About It · SIGMOD Conference 2015 |
Data integration and cleaning › data preprocessing
data cleaning |
0.1 | 2 | 2016 | Magellan: Toward Building Entity Matching Management Systems over Data Science Stacks · Proc. VLDB Endow. 2016 Magellan: Toward Building Entity Matching Management Systems · Proc. VLDB Endow. 2016 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.1 | 1 | 2017 | Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud Services · SIGMOD Conference 2017 |
Data mining › predictive modeling
classification |
0.1 | 1 | 2015 | Why Big Data Industrial Systems Need Rules and What We Can Do About It · SIGMOD Conference 2015 |
Methods — techniques the papers use, named apart from their topics
crowdsourcing · 1.7machine learning · 0.9interactive labeling · 0.7query optimization · 0.6operator implementation · 0.6data visualization · 0.5distinct sampling · 0.4data profiling · 0.4random forest · 0.3active learning · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Entity Image and Mixed-Modal Image Retrieval DatasetsabstractDespite advances in multimodal learning, challenging benchmarks for mixed-modal image retrieval that combines visual and textual information are lacking. This paper introduces a novel benchmark to rigorously evaluate image retrieval that demands deep cross-modal contextual understanding. We present two new datasets: the Entity Image Dataset (EI), providing canonical images for Wikipedia entities, and the Mixed-Modal Image Retrieval Dataset (MMIR), derived from the WIT dataset. The MMIR benchmark features two challenging query types requiring models to ground textual descriptions in the context of provided visual entities: single entity-image queries (one entity image with descriptive text) and multi-entity-image queries (multiple entity images with relational text). We empirically validate the benchmark's utility as both a training corpus and an evaluation set for mixed-modal retrieval. The quality of both datasets is further affirmed through crowd-sourced human annotations. The datasets are accessible through the GitHub page: https://github.com/google-research-datasets/wit-retrieval. Cristian-Ioan Blaga, Paul Suganthan G. C., Sahil Dua, Krishna Srinivasan, Enrique Alfonseca, Péter Dornbach, Tom Duerig, Imed Zitouni |
LREC | 2 |
| 2020 | TensorFlow Data Validation: Data Analysis and Validation in Continuous ML PipelinesabstractMachine Learning (ML) research has primarily focused on improving the accuracy and efficiency of the training algorithms while paying much less attention to the equally important problem of understanding, validating, and monitoring the data fed to ML. Irrespective of the ML algorithms used, data errors can adversely affect the quality of the generated model. This indicates that we need to adopt a data-centric approach to ML that treats data as a first-class citizen, on par with algorithms and infrastructure which are the typical building blocks of ML pipelines. In this demonstration we showcase TensorFlow Data Validation (TFDV), a scalable data analysis and validation system for ML that we have developed at Google and recently open-sourced. This system is deployed in production as an integral part of TFX - an end-to-end machine learning platform at Google. It is used by hundreds of product teams at Google and has received significant attention from the open-source community as well. Emily Caveness, Paul Suganthan G. C., Zhuo Peng, Neoklis Polyzotis, Sudip Roy 0002, Martin Zinkevich |
SIGMOD Conference | 2 |
| 2019 | Entity Matching Meets Data Science: A Progress Report from the Magellan ProjectabstractEntity matching (EM) finds data instances that refer to the same real-world entity. In 2015, we started the Magellan project at UW-Madison, joint with industrial partners, to build EM systems. Most current EM systems are stand-alone monoliths. In contrast, Magellan borrows ideas from the field of data science (DS), to build a new kind of EM systems, which is an ecosystem of interoperable tools. \em This paper provides a progress report on the past 3.5 years of Magellan, focusing on the system aspects and on how ideas from the field of data science have been adapted to the EM context. We argue why EM can be viewed as a special class of DS problems, and thus can benefit from system building ideas in DS. We discuss how these ideas have been adapted to build \pymatcher\ and \cloudmatcher, EM tools for power users and lay users. These tools have been successfully used in 21 EM tasks at 12 companies and domain science groups, and have been pushed into production for many customers. We report on the lessons learned, and outline a new envisioned Magellan ecosystem, which consists of not just on-premise Python tools, but also interoperable microservices deployed, executed, and scaled out on the cloud, using tools such as Dockers and Kubernetes. Yash Govind, Pradap Konda, Paul Suganthan G. C., Philip Martinkus, Palaniappan Nagarajan, Aravind Soundararajan, Sidharth Mudgal, Jeffrey R. Ballard, Haojun Zhang, Adel Ardalan, Sanjib Das, Derek Paulsen, Amanpreet Singh Saini, Erik Paulson 0001, Youngchoon Park, Marshall Carter, Mingju Sun, Glenn Fung, AnHai Doan |
SIGMOD Conference | 3 |
| 2018 | MatchCatcher: A Debugger for Blocking in Entity Matching
Pradap Konda, Paul Suganthan G. C., AnHai Doan, Benjamin Snyder, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Vijay Raghavendra |
EDBT | 3 |
| 2018 | Smurf: Self-Service String Matching Using Random ForestsabstractWe argue that more attention should be devoted to developing self-service string matching (SM) solutions, which lay users can easily use. We show that Falcon, a self-service entity matching (EM) solution, can be applied to SM and is more accurate than current self-service SM solutions. However, Falcon often asks lay users to label many string pairs (e.g., 770-1050 in our experiments). This is expensive, can significantly compound labeling mistakes, and takes a long time. We developed Smurf, a self-service SM solution that reduces the labeling effort by 43-76%, yet achieves comparable F 1 accuracy. The key to make Smurf possible is a novel solution to efficiently execute a random forest (that Smurf learns via active learning with the lay user) over two sets of strings. This solution uses RDBMS-style plan optimization to reuse computations across the trees in the forest. As such, Smurf significantly advances self-service SM and raises interesting future directions for self-service EM and scalable random forest execution over structured data. Paul Suganthan G. C., Adel Ardalan, AnHai Doan, Aditya Akella |
Proc. VLDB Endow. | 1 |
| 2018 | CloudMatcher: A Hands-Off Cloud/Crowd Service for Entity MatchingabstractAs data science applications proliferate, more and more lay users must perform data integration (DI) tasks, which used to be done by sophisticated CS developers. Thus, it is increasingly critical that we develop hands-off DI services, which lay users can use to perform such tasks without asking for help from developers. We propose to demonstrate such a service. Specifically, we will demonstrate CloudMatcher, a hands-off cloud/crowd service for entity matching (EM). To use CloudMatcher to match two tables, a lay user only needs to upload them to the CloudMatcher's Web page then iteratively label a set of tuple pairs as match/no-match. Alternatively, the user can enlist a crowd of workers to label the pairs. In either case, the lay user can easily perform EM end-to-end without having to involve any developers. Cloud-Matcher has been used in several domain science projects at UW-Madison and at several organizations, and is scheduled to be deployed in a large company in Summer 2018. In the demonstration we will show how easy it is for lay users to perform EM (either via interactive labeling or crowdsourcing), how users can easily create and experiment with a range of EM workflows, and how CloudMatcher can scale to many concurrent users and large datasets. Yash Govind, Erik Paulson 0001, Palaniappan Nagarajan, Paul Suganthan G. C., AnHai Doan, Youngchoon Park, Glenn Fung, Devin Conathan, Marshall Carter, Mingju Sun |
Proc. VLDB Endow. | 4 |
| 2017 | Falcon: Scaling Up Hands-Off Crowdsourced Entity Matching to Build Cloud ServicesabstractMany works have applied crowdsourcing to entity matching (EM). While promising, these approaches are limited in that they often require a developer to be in the loop. As such, it is difficult for an organization to deploy multiple crowdsourced EM solutions, because there are simply not enough developers. To address this problem, a recent work has proposed Corleone, a solution that crowdsources the entire EM workflow, requiring no developers. While promising, Corleone is severely limited in that it does not scale to large tables. We propose Falcon, a solution that scales up the hands-off crowdsourced EM approach of Corleone, using RDBMS-style query execution and optimization over a Hadoop cluster. Specifically, we define a set of operators and develop efficient implementations. We translate a hands-off crowdsourced EM workflow into a plan consisting of these operators, optimize, then execute the plan. These plans involve both machine and crowd activities, giving rise to novel optimization techniques such as using crowd time to mask machine time. Extensive experiments show that Falcon can scale up to tables of millions of tuples, thus providing a practical solution for hands-off crowdsourced EM, to build cloud-based EM services. Sanjib Das, Paul Suganthan G. C., AnHai Doan, Jeffrey F. Naughton, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, Vijay Raghavendra, Youngchoon Park |
SIGMOD Conference | 2 |
| 2016 | Magellan: Toward Building Entity Matching Management SystemsabstractEntity matching (EM) has been a long-standing challenge in data management. Most current EM works focus only on developing matching algorithms. We argue that far more efforts should be devoted to building EM systems. We discuss the limitations of current EM systems, then present as a solution Magellan, a new kind of EM systems. Magellan is novel in four important aspects. (1) It provides how-to guides that tell users what to do in each EM scenario, step by step. (2) It provides tools to help users do these steps; the tools seek to cover the entire EM pipeline, not just matching and blocking as current EM systems do. (3) Tools are built on top of the data analysis and Big Data stacks in Python, allowing Magellan to borrow a rich set of capabilities in data cleaning, IE, visualization, learning, etc. (4) Magellan provides a powerful scripting environment to facilitate interactive experimentation and quick "patching" of the system. We describe research challenges raised by Magellan, then present extensive experiments with 44 students and users at several organizations that show the promise of the Magellan approach. Pradap Konda, Sanjib Das, Paul Suganthan G. C., AnHai Doan, Adel Ardalan, Jeffrey R. Ballard, Fatemah Panahi, Haojun Zhang, Jeffrey F. Naughton, Shishir Prasad, Ganesh Krishnan, Rohit Deep, Vijay Raghavendra |
Proc. VLDB Endow. | 3 |
| 2016 | Magellan: Toward Building Entity Matching Management Systems over Data Science StacksabstractEntity matching (EM) has been a long-standing challenge in data management. Most current EM works, however, focus only on developing matching algorithms. We argue that far more efforts should be devoted to building EM systems. We discuss the limitations of current EM systems, then present Magellan, a new kind of EM systems that addresses these limitations. Magellan is novel in four important aspects. (1) It provides a how-to guide that tells users what to do in each EM scenario, step by step. (2) It provides tools to help users do these steps; the tools seek to cover the entire EM pipeline, not just matching and blocking as current EM systems do. (3) Tools are built on top of the data science stacks in Python, allowing Magellan to borrow a rich set of capabilities in data cleaning, IE, visualization, learning, etc. (4) Magellan provide a powerful scripting environment to facilitate interactive experimentation and allow users to quickly write code to "patch" the system. We have extensively evaluated Magellan with 44 students and users at various organizations. In this paper we propose demonstration scenarios that show the promise of the Magellan approach. Pradap Konda, Sanjib Das, Paul Suganthan G. C., AnHai Doan, Adel Ardalan, Jeffrey R. Ballard, Fatemah Panahi, Haojun Zhang, Jeffrey F. Naughton, Shishir Prasad, Ganesh Krishnan, Rohit Deep, Vijay Raghavendra |
Proc. VLDB Endow. | 3 |
| 2015 | Why Big Data Industrial Systems Need Rules and What We Can Do About ItabstractBig Data industrial systems that address problems such as classification, information extraction, and entity matching very commonly use hand-crafted rules. Today, however, little is understood about the usage of such rules. In this paper we explore this issue. We discuss how these systems differ from those considered in academia. We describe default solutions, their limitations, and reasons for using rules. We show examples of extensive rule usage in industry. Contrary to popular perceptions, we show that there is a rich set of research challenges in rule generation, evaluation, execution, optimization, and maintenance. We discuss ongoing work at WalmartLabs and UW-Madison that illustrate these challenges. Our main conclusions are (1) using rules (together with techniques such as learning and crowdsourcing) is fundamental to building semantics-intensive Big Data systems, and (2) it is increasingly critical to address rule management, given the tens of thousands of rules industrial systems often manage today in an ad-hoc fashion. Paul Suganthan G. C., Krishna Gayatri K., Haojun Zhang, Frank Yang, Narasimhan Rampalli, Shishir Prasad, Esteban Arcaute, Ganesh Krishnan, Rohit Deep, Vijay Raghavendra, AnHai Doan |
SIGMOD Conference | 1 |