EDBT 2026 Demo / reviewers in the wild / expert
Sudip Roy 0002
dblp:44/4775-2
· DBLP profile ↗
15ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0002-0535-0531ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 3 first-authorArtificial intelligence and machine learning · 2Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
8 papers |
Machine learning and data management · 43% Data integration and cleaning · 28% Data models and query languages · 16% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 67% Program synthesis and code generation · 30% Program analysis · 3% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Distributed systems · 45% Hardware accelerators and domain-specific architectures · 30% Cloud and datacenter computing · 14% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 25 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
neural architecture search |
0.6 | 1 | 2022 | Neural architecture search using property guided synthesis · Proc. ACM Program. Lang. 2022 |
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
device placement |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Machine learning › Efficient and distributed learning
distributed training |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Data integration and cleaning › data quality
data validation |
0.4 | 1 | 2020 | TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines · SIGMOD Conference 2020 |
Compilers and program optimization › compiler optimization
computation graph optimization |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Compilers and program optimization › compiler optimization
machine learning for compiler optimization |
0.4 | 1 | 2020 | Transferable Graph Optimizers for ML Compilers · NeurIPS 2020 |
Machine learning and data management
machine learning lifecycle management |
0.3 | 1 | 2017 | TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017 |
Machine learning and data management
machine learning pipeline |
0.3 | 1 | 2017 | Data Management Challenges in Production Machine Learning · SIGMOD Conference 2017 |
Machine learning and data management › machine learning systems
machine learning platform |
0.3 | 1 | 2017 | TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017 |
Machine learning and data management
training data management |
0.3 | 1 | 2017 | Data Management Challenges in Production Machine Learning · SIGMOD Conference 2017 |
Machine learning and data management
training data quality |
0.3 | 1 | 2017 | Data Management Challenges in Production Machine Learning · SIGMOD Conference 2017 |
Data integration and cleaning › data discovery
dataset discovery |
0.2 | 1 | 2016 | Goods: Organizing Google's Datasets · SIGMOD Conference 2016 |
Data integration and cleaning › metadata management
metadata extraction |
0.2 | 1 | 2016 | Goods: Organizing Google's Datasets · SIGMOD Conference 2016 |
Cloud and datacenter computing › datacenter operations
cloud system operations |
0.2 | 1 | 2015 | PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015 |
Distributed systems › distributed database
distributed transactions |
0.2 | 1 | 2015 | The Homeostasis Protocol: Avoiding Transaction Coordination Through Program Analysis · SIGMOD Conference 2015 |
Distributed systems › anomaly detection
log-based anomaly detection |
0.2 | 1 | 2015 | PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015 |
Distributed systems › anomaly detection
performance anomaly diagnosis |
0.2 | 1 | 2015 | PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015 |
Distributed systems › consistency models
strong consistency |
0.2 | 1 | 2015 | The Homeostasis Protocol: Avoiding Transaction Coordination Through Program Analysis · SIGMOD Conference 2015 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2015 | PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015 |
Information retrieval › query processing
query matching |
0.1 | 1 | 2012 | Entangled queries: Enabling declarative data-driven coordination · ACM Trans. Database Syst. 2012 |
Data models and query languages › SQL
SQL extension |
0.1 | 1 | 2011 | Entangled queries: enabling declarative data-driven coordination · SIGMOD Conference 2011 |
Transaction processing and concurrency control
transaction models |
0.1 | 1 | 2011 | Entangled Transactions · Proc. VLDB Endow. 2011 |
Distributed and cloud data management
large-scale data management |
0.1 | 1 | 2016 | Goods: Organizing Google's Datasets · SIGMOD Conference 2016 |
Program analysis
static analysis |
0.1 | 1 | 2015 | The Homeostasis Protocol: Avoiding Transaction Coordination Through Program Analysis · SIGMOD Conference 2015 |
Cloud and datacenter computing
cloud service management |
0.1 | 1 | 2015 | PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015 |
Methods — techniques the papers use, named apart from their topics
program property abstraction · 1.1evolutionary algorithm · 1.1sequential attention · 0.9graph neural network · 0.9deep reinforcement learning · 0.9program analysis · 0.4homeostasis protocol · 0.4distinct sampling · 0.4data profiling · 0.4SQL extension · 0.4data validation · 0.3data cleaning · 0.3relationship inference · 0.2metadata crawling · 0.2unsupervised anomaly detection · 0.2log mining · 0.2hypothesis generation · 0.2static analysis · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Neural architecture search using property guided synthesisabstractNeural architecture search (NAS) has become an increasingly important tool within the deep learning community in recent years, yielding many practical advancements in the design of deep neural network architectures. However, most existing approaches operate within highly structured design spaces, and hence (1) explore only a small fraction of the full search space of neural architectures while also (2) requiring significant manual effort from domain experts. In this work, we develop techniques that enable efficient NAS in a significantly larger design space. In particular, we propose to perform NAS in an abstract search space of program properties. Our key insights are as follows: (1) an abstract search space can be significantly smaller than the original search space, and (2) architectures with similar program properties should also have similar performance; thus, we can search more efficiently in the abstract search space. To enable this approach, we also introduce a novel efficient synthesis procedure, which performs the role of concretizing a set of promising program properties into a satisfying neural architecture. We implement our approach, αNAS, within an evolutionary framework, where the mutations are guided by the program properties. Starting with a ResNet-34 model, αNAS produces a model with slightly improved accuracy on CIFAR-10 but 96% fewer parameters. On ImageNet, αNAS is able to improve over Vision Transformer (30% fewer FLOPS and parameters), ResNet-50 (23% fewer FLOPS, 14% fewer parameters), and EfficientNet (7% fewer FLOPS and parameters) without any degradation in accuracy. Charles Jin, Phitchaya Mangpo Phothilimthana, Sudip Roy 0002 |
Proc. ACM Program. Lang. | 3 |
| 2021 | A Flexible Approach to Autotuning Multi-Pass Machine Learning CompilersabstractSearch-based techniques have been demonstrated effective in solving complex optimization problems that arise in domain-specific compilers for machine learning (ML). Unfortunately, deploying such techniques in production compilers is impeded by two limitations. First, prior works require factorization of a computation graph into smaller subgraphs over which search is applied. This decomposition is not only non-trivial but also significantly limits the scope of optimization. Second, prior works require search to be applied in a single stage in the compilation flow, which does not fit with the multi-stage layered architecture of most production ML compilers. This paper presents Xtat, an autotuner for production ML compilers that can tune both graph-level and subgraph-level optimizations across multiple compilation stages. Xtat applies Xtat-M, a flexible search methodology that defines a search formulation for joint optimizations by accurately modeling the interactions between different compiler passes. Xtat tunes tensor layouts, operator fusion decisions, tile sizes, and code generation parameters in XLA, a production ML compiler, using various search strategies. In an evaluation across 150 ML training and inference models on Tensor Processing Units (TPUs) at Google, Xtat offers up to 2.4x and an average 5% execution time speedup over the heavily-optimized XLA compiler. Phitchaya Mangpo Phothilimthana, Amit Sabne, Nikhil Sarda, Karthik Srinivasa Murthy, Yanqi Zhou, Christof Angermueller, Michael Burrows, Sudip Roy 0002, Ketan Mandke, Rezsa Farahani, Yu Emma Wang, Berkin Ilbeyi, Blake A. Hechtman, Bjarke Roune, Yuanzhong Xu, Samuel J. Kaufman |
PACT | 8 |
| 2020 | Transferable Graph Optimizers for ML CompilersabstractMost compilers for machine learning (ML) frameworks need to solve many correlated optimization problems to generate efficient machine code. Current ML compilers rely on heuristics based algorithms to solve these optimization problems one at a time. However, this approach is not only hard to maintain but often leads to sub-optimal solutions especially for newer model architectures. Existing learning based approaches in the literature are sample inefficient, tackle a single optimization problem, and do not generalize to unseen graphs making them infeasible to be deployed in practice. To address these limitations, we propose an end-to-end, transferable deep reinforcement learning method for computational graph optimization (GO), based on a scalable sequential attention mechanism over an inductive graph neural network. GO generates decisions on the entire graph rather than on each individual node autoregressively, drastically speeding up the search compared to prior methods. Moreover, we propose recurrent attention layers to jointly optimize dependent graph optimization tasks and demonstrate 33%-60% speedup on three graph optimization tasks compared to TensorFlow default optimization. On a diverse set of representative graphs consisting of up to 80,000 nodes, including Inception-v3, Transformer-XL, and WaveNet, GO achieves on average 21% improvement over human experts and 18% improvement over the prior state of the art with 15x faster convergence, on a device placement task evaluated in real systems. Yanqi Zhou, Sudip Roy 0002, AmirAli Abdolrashidi, Daniel Wong 0001, Peter C. Ma, Qiumin Xu, Hanxiao Liu, Mangpo Phitchaya Phothilimtha, Anna Goldie, Azalia Mirhoseini, James Laudon |
NeurIPS | 2 |
| 2020 | TensorFlow Data Validation: Data Analysis and Validation in Continuous ML PipelinesabstractMachine Learning (ML) research has primarily focused on improving the accuracy and efficiency of the training algorithms while paying much less attention to the equally important problem of understanding, validating, and monitoring the data fed to ML. Irrespective of the ML algorithms used, data errors can adversely affect the quality of the generated model. This indicates that we need to adopt a data-centric approach to ML that treats data as a first-class citizen, on par with algorithms and infrastructure which are the typical building blocks of ML pipelines. In this demonstration we showcase TensorFlow Data Validation (TFDV), a scalable data analysis and validation system for ML that we have developed at Google and recently open-sourced. This system is deployed in production as an integral part of TFX - an end-to-end machine learning platform at Google. It is used by hundreds of product teams at Google and has received significant attention from the open-source community as well. Emily Caveness, Paul Suganthan G. C., Zhuo Peng, Neoklis Polyzotis, Sudip Roy 0002, Martin Zinkevich |
SIGMOD Conference | 5 |
| 2017 | TFX: A TensorFlow-Based Production-Scale Machine Learning PlatformabstractCreating and maintaining a platform for reliably producing and deploying machine learning models requires careful orchestration of many components---a learner for generating models based on training data, modules for analyzing and validating both data as well as models, and finally infrastructure for serving models in production. This becomes particularly challenging when data changes over time and fresh models need to be produced continuously. Unfortunately, such orchestration is often done ad hoc using glue code and custom scripts developed by individual teams for specific use cases, leading to duplicated effort and fragile systems with high technical debt. Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc 0001, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy 0002, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Martin Zinkevich |
KDD | 17 |
| 2017 | Data Management Challenges in Production Machine LearningabstractThe tutorial discusses data-management issues that arise in the context of machine learning pipelines deployed in production. Informed by our own experience with such largescale pipelines, we focus on issues related to understanding, validating, cleaning, and enriching training data. The goal of the tutorial is to bring forth these issues, draw connections to prior work in the database literature, and outline the open research questions that are not addressed by prior art. Neoklis Polyzotis, Sudip Roy 0002, Steven Euijong Whang, Martin Zinkevich |
SIGMOD Conference | 2 |
| 2016 | Goods: Organizing Google's DatasetsabstractEnterprises increasingly rely on structured datasets to run their businesses. These datasets take a variety of forms, such as structured files, databases, spreadsheets, or even services that provide access to the data. The datasets often reside in different storage systems, may vary in their formats, may change every day. In this paper, we present GOODS, a project to rethink how we organize structured datasets at scale, in a setting where teams use diverse and often idiosyncratic ways to produce the datasets and where there is no centralized system for storing and querying them. GOODS extracts metadata ranging from salient information about each dataset (owners, timestamps, schema) to relationships among datasets, such as similarity and provenance. It then exposes this metadata through services that allow engineers to find datasets within the company, to monitor datasets, to annotate them in order to enable others to use their datasets, and to analyze relationships between them. We discuss the technical challenges that we had to overcome in order to crawl and infer the metadata for billions of datasets, to maintain the consistency of our metadata catalog at scale, and to expose the metadata to users. We believe that many of the lessons that we learned are applicable to building large-scale enterprise-level data-management systems in general. Alon Y. Halevy, Flip Korn, Natasha F. Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy 0002, Steven Euijong Whang |
SIGMOD Conference | 6 |
| 2015 | PerfAugur: Robust diagnostics for performance anomalies in cloud servicesabstractCloud platforms involve multiple independently developed components, often executing on diverse hardware configurations and across multiple data centers. This complexity makes tracking various key performance indicators (KPIs) and manual diagnosing of anomalies in system behavior both difficult and expensive. In this paper, we describe PerfAugur, an automated system for mining service logs to identify anomalies and help formulate data-driven hypotheses. PerfAugur includes a suite of efficient mining algorithms for detecting significant anomalies in system behavior, along with potential explanations for such anomalies, without the need for an explicit supervision signal. We perform extensive experimental evaluation using both synthetic and real-life data sets, and present detailed case studies showing the impact of this technology on operations of the Windows Azure Service. Sudip Roy 0002, Arnd Christian König, Igor Dvorkin |
ICDE | 1 |
| 2015 | The Homeostasis Protocol: Avoiding Transaction Coordination Through Program AnalysisabstractDatastores today rely on distribution and replication to achieve improved performance and fault-tolerance. But correctness of many applications depends on strong consistency properties--something that can impose substantial overheads, since it requires coordinating the behavior of multiple nodes. This paper describes a new approach to achieving strong consistency in distributed systems while minimizing communication between nodes. The key insight is to allow the state of the system to be inconsistent during execution, as long as this inconsistency is bounded and does not affect transaction correctness. In contrast to previous work, our approach uses program analysis to extract semantic information about permissible levels of inconsistency and is fully automated. We then employ a novel homeostasis protocol to allow sites to operate independently, without communicating, as long as any inconsistency is governed by appropriate treaties between the nodes. We discuss mechanisms for optimizing treaties based on workload characteristics to minimize communication, as well as a prototype implementation and experiments that demonstrate the benefits of our approach on common transactional benchmarks. Sudip Roy 0002, Lucja Kot, Gabriel Bender, Bailu Ding, Hossein Hojjat, Christoph Koch 0001, Nate Foster, Johannes Gehrke |
SIGMOD Conference | 1 |
| 2013 | Quantum Databases
Sudip Roy 0002, Lucja Kot, Christoph Koch 0001 |
CIDR | 1 |
| 2012 | Entangled queries: Enabling declarative data-driven coordinationabstractMany data-driven social and Web applications involve collaboration and coordination. The vision of Declarative Data-Driven Coordination (D3C), proposed in Kot et al. [2010], is to support coordination in the spirit of data management: to make it data-centric and to specify it using convenient declarative languages. This article introduces entangled queries , a language that extends SQL by constraints that allow for the coordinated choice of result tuples across queries originating from different users or applications. It is nontrivial to define a declarative coordination formalism without arriving at the general (NP-complete) Constraint Satisfaction Problem from AI. In this article, we propose an efficiently enforceable syntactic safety condition that we argue is at the sweet spot where interesting declarative power meets applicability in large-scale data management systems and applications. The key computational problem of D3C is to match entangled queries to achieve coordination. We present an efficient matching algorithm which statically analyzes query workloads and merges coordinating entangled queries into compound SQL queries. These can be sent to a standard database system and return only coordinated results. We present the overall architecture of an implemented system that contains our evaluation algorithm. We also describe a proof-of-concept Facebook application we have built on top of this system to allow friends to coordinate flight plans. Finally, we evaluate the performance of the matching algorithm experimentally on realistic coordination workloads. Nitin Gupta 0003, Lucja Kot, Sudip Roy 0002, Gabriel Bender, Johannes Gehrke, Christoph Koch 0001 |
ACM Trans. Database Syst. | 3 |
| 2011 | Coordination through querying in the youtopia systemabstractIn a previous paper, we laid out the vision of declarative data-driven coordination (D3C) where users are provided with novel abstractions that enable them to communicate and coordinate through declarative specifications [3]. Nitin Gupta 0003, Lucja Kot, Gabriel Bender, Sudip Roy 0002, Johannes Gehrke, Christoph Koch 0001 |
SIGMOD Conference | 4 |
| 2011 | Entangled queries: enabling declarative data-driven coordinationabstractMany data-driven social and Web applications involve collaboration and coordination. The vision of declarative data-driven coordination (D3C), proposed in [9], is to support coordination in the spirit of data management: to make it data-centric and to specify it using convenient declarative languages. This paper introduces entangled queries, a language that extends SQL by constraints that allow for the coordinated choice of result tuples across queries originating from different users or applications. Nitin Gupta 0003, Lucja Kot, Sudip Roy 0002, Gabriel Bender, Johannes Gehrke, Christoph Koch 0001 |
SIGMOD Conference | 3 |
| 2011 | Entangled Transactions
Nitin Gupta 0003, Milos Nikolic 0001, Sudip Roy 0002, Gabriel Bender, Lucja Kot, Johannes Gehrke, Christoph Koch 0001 |
Proc. VLDB Endow. | 3 |
| 2008 | A layout-aware physical design method for constructing feasible QCA circuitsabstractQuantum-dot Cellular Automata (QCA) is an emerging computing paradigm, in which logical operations as well as signal transmission occurs due to Coulombic charge interaction between neighbouring QCA cells, moderated by a 4-phase QCA clock potential. Thermodynamic constraints like the number of QCA cells in a clocking zone must be obeyed to obtain a logically correct and feasible QCA circuit. These constraints depend on various design factors like total wirelength in a circuit, height of a clocking zone etc. which are not available until actual circuit layout is obtained. In this paper, the various design automation problems assosciated with obtaining a feasible QCA layout are addressed. The layout generation problem is formulated as embedding the netlist digraph in an orthogonal grid, which provides an abstraction of the actual physical layout to be obtained. Novel graph theoretic algorithms are proposed to perform placement and global routing and various design parameters like clock rate, wasted area and total wirelength are used to estimate the quality of the layout obtained. Also, planarization methods are used to remove all wire crossings, which are expensive to fabricate. The methods applied on a large number of MCNC'93 and ISCAS'89 benchmarks show good results. Mayur Bubna, Sudip Roy 0002, Naresh Shenoy, Subhra Mazumdar 0002 |
ACM Great Lakes Symposium on VLSI | 2 |