Sudip Roy 0002

dblp:44/4775-2 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0002-0535-0531ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 3 first-authorArtificial intelligence and machine learning · 2Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
8 papers
Machine learning and data management · 43% Data integration and cleaning · 28% Data models and query languages · 16%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 67% Program synthesis and code generation · 30% Program analysis · 3%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 45% Hardware accelerators and domain-specific architectures · 30% Cloud and datacenter computing · 14%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 25 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
neural architecture search
0.612022
Neural architecture search using property guided synthesis · Proc. ACM Program. Lang. 2022
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
device placement
0.412020
Transferable Graph Optimizers for ML Compilers · NeurIPS 2020
Machine learning › Efficient and distributed learning
distributed training
0.412020
Transferable Graph Optimizers for ML Compilers · NeurIPS 2020
Data integration and cleaning › data quality
data validation
0.412020
TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines · SIGMOD Conference 2020
Compilers and program optimization › compiler optimization
computation graph optimization
0.412020
Transferable Graph Optimizers for ML Compilers · NeurIPS 2020
Compilers and program optimization › compiler optimization
machine learning for compiler optimization
0.412020
Transferable Graph Optimizers for ML Compilers · NeurIPS 2020
Machine learning and data management
machine learning lifecycle management
0.312017
TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017
Machine learning and data management
machine learning pipeline
0.312017
Data Management Challenges in Production Machine Learning · SIGMOD Conference 2017
Machine learning and data management › machine learning systems
machine learning platform
0.312017
TFX: A TensorFlow-Based Production-Scale Machine Learning Platform · KDD 2017
Machine learning and data management
training data management
0.312017
Data Management Challenges in Production Machine Learning · SIGMOD Conference 2017
Machine learning and data management
training data quality
0.312017
Data Management Challenges in Production Machine Learning · SIGMOD Conference 2017
Data integration and cleaning › data discovery
dataset discovery
0.212016
Goods: Organizing Google's Datasets · SIGMOD Conference 2016
Data integration and cleaning › metadata management
metadata extraction
0.212016
Goods: Organizing Google's Datasets · SIGMOD Conference 2016
Cloud and datacenter computing › datacenter operations
cloud system operations
0.212015
PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015
Distributed systems › distributed database
distributed transactions
0.212015
The Homeostasis Protocol: Avoiding Transaction Coordination Through Program Analysis · SIGMOD Conference 2015
Distributed systems › anomaly detection
log-based anomaly detection
0.212015
PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015
Distributed systems › anomaly detection
performance anomaly diagnosis
0.212015
PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015
Distributed systems › consistency models
strong consistency
0.212015
The Homeostasis Protocol: Avoiding Transaction Coordination Through Program Analysis · SIGMOD Conference 2015
Performance modeling and evaluation
workload characterization
0.212015
PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015
Information retrieval › query processing
query matching
0.112012
Entangled queries: Enabling declarative data-driven coordination · ACM Trans. Database Syst. 2012
Data models and query languages › SQL
SQL extension
0.112011
Entangled queries: enabling declarative data-driven coordination · SIGMOD Conference 2011
Transaction processing and concurrency control
transaction models
0.112011
Entangled Transactions · Proc. VLDB Endow. 2011
Distributed and cloud data management
large-scale data management
0.112016
Goods: Organizing Google's Datasets · SIGMOD Conference 2016
Program analysis
static analysis
0.112015
The Homeostasis Protocol: Avoiding Transaction Coordination Through Program Analysis · SIGMOD Conference 2015
Cloud and datacenter computing
cloud service management
0.112015
PerfAugur: Robust diagnostics for performance anomalies in cloud services · ICDE 2015

Methods — techniques the papers use, named apart from their topics

program property abstraction · 1.1evolutionary algorithm · 1.1sequential attention · 0.9graph neural network · 0.9deep reinforcement learning · 0.9program analysis · 0.4homeostasis protocol · 0.4distinct sampling · 0.4data profiling · 0.4SQL extension · 0.4data validation · 0.3data cleaning · 0.3relationship inference · 0.2metadata crawling · 0.2unsupervised anomaly detection · 0.2log mining · 0.2hypothesis generation · 0.2static analysis · 0.1
YearPublicationVenuePosition
2022 Neural architecture search using property guided synthesis
abstract
Neural architecture search (NAS) has become an increasingly important tool within the deep learning community in recent years, yielding many practical advancements in the design of deep neural network architectures. However, most existing approaches operate within highly structured design spaces, and hence (1) explore only a small fraction of the full search space of neural architectures while also (2) requiring significant manual effort from domain experts. In this work, we develop techniques that enable efficient NAS in a significantly larger design space. In particular, we propose to perform NAS in an abstract search space of program properties. Our key insights are as follows: (1) an abstract search space can be significantly smaller than the original search space, and (2) architectures with similar program properties should also have similar performance; thus, we can search more efficiently in the abstract search space. To enable this approach, we also introduce a novel efficient synthesis procedure, which performs the role of concretizing a set of promising program properties into a satisfying neural architecture. We implement our approach, αNAS, within an evolutionary framework, where the mutations are guided by the program properties. Starting with a ResNet-34 model, αNAS produces a model with slightly improved accuracy on CIFAR-10 but 96% fewer parameters. On ImageNet, αNAS is able to improve over Vision Transformer (30% fewer FLOPS and parameters), ResNet-50 (23% fewer FLOPS, 14% fewer parameters), and EfficientNet (7% fewer FLOPS and parameters) without any degradation in accuracy.
Charles Jin, Phitchaya Mangpo Phothilimthana, Sudip Roy 0002
Proc. ACM Program. Lang.3
2021 A Flexible Approach to Autotuning Multi-Pass Machine Learning Compilers
abstract
Search-based techniques have been demonstrated effective in solving complex optimization problems that arise in domain-specific compilers for machine learning (ML). Unfortunately, deploying such techniques in production compilers is impeded by two limitations. First, prior works require factorization of a computation graph into smaller subgraphs over which search is applied. This decomposition is not only non-trivial but also significantly limits the scope of optimization. Second, prior works require search to be applied in a single stage in the compilation flow, which does not fit with the multi-stage layered architecture of most production ML compilers. This paper presents Xtat, an autotuner for production ML compilers that can tune both graph-level and subgraph-level optimizations across multiple compilation stages. Xtat applies Xtat-M, a flexible search methodology that defines a search formulation for joint optimizations by accurately modeling the interactions between different compiler passes. Xtat tunes tensor layouts, operator fusion decisions, tile sizes, and code generation parameters in XLA, a production ML compiler, using various search strategies. In an evaluation across 150 ML training and inference models on Tensor Processing Units (TPUs) at Google, Xtat offers up to 2.4x and an average 5% execution time speedup over the heavily-optimized XLA compiler.
Phitchaya Mangpo Phothilimthana, Amit Sabne, Nikhil Sarda, Karthik Srinivasa Murthy, Yanqi Zhou, Christof Angermueller, Michael Burrows, Sudip Roy 0002, Ketan Mandke, Rezsa Farahani, Yu Emma Wang, Berkin Ilbeyi, Blake A. Hechtman, Bjarke Roune, Yuanzhong Xu, Samuel J. Kaufman
PACT8
2020 Transferable Graph Optimizers for ML Compilers
abstract
Most compilers for machine learning (ML) frameworks need to solve many correlated optimization problems to generate efficient machine code. Current ML compilers rely on heuristics based algorithms to solve these optimization problems one at a time. However, this approach is not only hard to maintain but often leads to sub-optimal solutions especially for newer model architectures. Existing learning based approaches in the literature are sample inefficient, tackle a single optimization problem, and do not generalize to unseen graphs making them infeasible to be deployed in practice. To address these limitations, we propose an end-to-end, transferable deep reinforcement learning method for computational graph optimization (GO), based on a scalable sequential attention mechanism over an inductive graph neural network. GO generates decisions on the entire graph rather than on each individual node autoregressively, drastically speeding up the search compared to prior methods. Moreover, we propose recurrent attention layers to jointly optimize dependent graph optimization tasks and demonstrate 33%-60% speedup on three graph optimization tasks compared to TensorFlow default optimization. On a diverse set of representative graphs consisting of up to 80,000 nodes, including Inception-v3, Transformer-XL, and WaveNet, GO achieves on average 21% improvement over human experts and 18% improvement over the prior state of the art with 15x faster convergence, on a device placement task evaluated in real systems.
Yanqi Zhou, Sudip Roy 0002, AmirAli Abdolrashidi, Daniel Wong 0001, Peter C. Ma, Qiumin Xu, Hanxiao Liu, Mangpo Phitchaya Phothilimtha, Anna Goldie, Azalia Mirhoseini, James Laudon
NeurIPS2
2020 TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines
abstract
Machine Learning (ML) research has primarily focused on improving the accuracy and efficiency of the training algorithms while paying much less attention to the equally important problem of understanding, validating, and monitoring the data fed to ML. Irrespective of the ML algorithms used, data errors can adversely affect the quality of the generated model. This indicates that we need to adopt a data-centric approach to ML that treats data as a first-class citizen, on par with algorithms and infrastructure which are the typical building blocks of ML pipelines. In this demonstration we showcase TensorFlow Data Validation (TFDV), a scalable data analysis and validation system for ML that we have developed at Google and recently open-sourced. This system is deployed in production as an integral part of TFX - an end-to-end machine learning platform at Google. It is used by hundreds of product teams at Google and has received significant attention from the open-source community as well.
Emily Caveness, Paul Suganthan G. C., Zhuo Peng, Neoklis Polyzotis, Sudip Roy 0002, Martin Zinkevich
SIGMOD Conference5
2017 TFX: A TensorFlow-Based Production-Scale Machine Learning Platform
abstract
Creating and maintaining a platform for reliably producing and deploying machine learning models requires careful orchestration of many components---a learner for generating models based on training data, modules for analyzing and validating both data as well as models, and finally infrastructure for serving models in production. This becomes particularly challenging when data changes over time and fresh models need to be produced continuously. Unfortunately, such orchestration is often done ad hoc using glue code and custom scripts developed by individual teams for specific use cases, leading to duplicated effort and fragile systems with high technical debt.
Denis Baylor, Eric Breck, Heng-Tze Cheng, Noah Fiedel, Chuan Yu Foo, Zakaria Haque, Salem Haykal, Mustafa Ispir, Vihan Jain, Levent Koc 0001, Chiu Yuen Koo, Lukasz Lew, Clemens Mewald, Akshay Naresh Modi, Neoklis Polyzotis, Sukriti Ramesh, Sudip Roy 0002, Steven Euijong Whang, Martin Wicke, Jarek Wilkiewicz, Martin Zinkevich
KDD17
2017 Data Management Challenges in Production Machine Learning
abstract
The tutorial discusses data-management issues that arise in the context of machine learning pipelines deployed in production. Informed by our own experience with such largescale pipelines, we focus on issues related to understanding, validating, cleaning, and enriching training data. The goal of the tutorial is to bring forth these issues, draw connections to prior work in the database literature, and outline the open research questions that are not addressed by prior art.
Neoklis Polyzotis, Sudip Roy 0002, Steven Euijong Whang, Martin Zinkevich
SIGMOD Conference2
2016 Goods: Organizing Google's Datasets
abstract
Enterprises increasingly rely on structured datasets to run their businesses. These datasets take a variety of forms, such as structured files, databases, spreadsheets, or even services that provide access to the data. The datasets often reside in different storage systems, may vary in their formats, may change every day. In this paper, we present GOODS, a project to rethink how we organize structured datasets at scale, in a setting where teams use diverse and often idiosyncratic ways to produce the datasets and where there is no centralized system for storing and querying them. GOODS extracts metadata ranging from salient information about each dataset (owners, timestamps, schema) to relationships among datasets, such as similarity and provenance. It then exposes this metadata through services that allow engineers to find datasets within the company, to monitor datasets, to annotate them in order to enable others to use their datasets, and to analyze relationships between them. We discuss the technical challenges that we had to overcome in order to crawl and infer the metadata for billions of datasets, to maintain the consistency of our metadata catalog at scale, and to expose the metadata to users. We believe that many of the lessons that we learned are applicable to building large-scale enterprise-level data-management systems in general.
Alon Y. Halevy, Flip Korn, Natasha F. Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy 0002, Steven Euijong Whang
SIGMOD Conference6
2015 PerfAugur: Robust diagnostics for performance anomalies in cloud services
abstract
Cloud platforms involve multiple independently developed components, often executing on diverse hardware configurations and across multiple data centers. This complexity makes tracking various key performance indicators (KPIs) and manual diagnosing of anomalies in system behavior both difficult and expensive. In this paper, we describe PerfAugur, an automated system for mining service logs to identify anomalies and help formulate data-driven hypotheses. PerfAugur includes a suite of efficient mining algorithms for detecting significant anomalies in system behavior, along with potential explanations for such anomalies, without the need for an explicit supervision signal. We perform extensive experimental evaluation using both synthetic and real-life data sets, and present detailed case studies showing the impact of this technology on operations of the Windows Azure Service.
Sudip Roy 0002, Arnd Christian König, Igor Dvorkin
ICDE1
2015 The Homeostasis Protocol: Avoiding Transaction Coordination Through Program Analysis
abstract
Datastores today rely on distribution and replication to achieve improved performance and fault-tolerance. But correctness of many applications depends on strong consistency properties--something that can impose substantial overheads, since it requires coordinating the behavior of multiple nodes. This paper describes a new approach to achieving strong consistency in distributed systems while minimizing communication between nodes. The key insight is to allow the state of the system to be inconsistent during execution, as long as this inconsistency is bounded and does not affect transaction correctness. In contrast to previous work, our approach uses program analysis to extract semantic information about permissible levels of inconsistency and is fully automated. We then employ a novel homeostasis protocol to allow sites to operate independently, without communicating, as long as any inconsistency is governed by appropriate treaties between the nodes. We discuss mechanisms for optimizing treaties based on workload characteristics to minimize communication, as well as a prototype implementation and experiments that demonstrate the benefits of our approach on common transactional benchmarks.
Sudip Roy 0002, Lucja Kot, Gabriel Bender, Bailu Ding, Hossein Hojjat, Christoph Koch 0001, Nate Foster, Johannes Gehrke
SIGMOD Conference1
2013 Quantum Databases
Sudip Roy 0002, Lucja Kot, Christoph Koch 0001
CIDR1
2012 Entangled queries: Enabling declarative data-driven coordination
abstract
Many data-driven social and Web applications involve collaboration and coordination. The vision of Declarative Data-Driven Coordination (D3C), proposed in Kot et al. [2010], is to support coordination in the spirit of data management: to make it data-centric and to specify it using convenient declarative languages. This article introduces entangled queries , a language that extends SQL by constraints that allow for the coordinated choice of result tuples across queries originating from different users or applications. It is nontrivial to define a declarative coordination formalism without arriving at the general (NP-complete) Constraint Satisfaction Problem from AI. In this article, we propose an efficiently enforceable syntactic safety condition that we argue is at the sweet spot where interesting declarative power meets applicability in large-scale data management systems and applications. The key computational problem of D3C is to match entangled queries to achieve coordination. We present an efficient matching algorithm which statically analyzes query workloads and merges coordinating entangled queries into compound SQL queries. These can be sent to a standard database system and return only coordinated results. We present the overall architecture of an implemented system that contains our evaluation algorithm. We also describe a proof-of-concept Facebook application we have built on top of this system to allow friends to coordinate flight plans. Finally, we evaluate the performance of the matching algorithm experimentally on realistic coordination workloads.
Nitin Gupta 0003, Lucja Kot, Sudip Roy 0002, Gabriel Bender, Johannes Gehrke, Christoph Koch 0001
ACM Trans. Database Syst.3
2011 Coordination through querying in the youtopia system
abstract
In a previous paper, we laid out the vision of declarative data-driven coordination (D3C) where users are provided with novel abstractions that enable them to communicate and coordinate through declarative specifications [3].
Nitin Gupta 0003, Lucja Kot, Gabriel Bender, Sudip Roy 0002, Johannes Gehrke, Christoph Koch 0001
SIGMOD Conference4
2011 Entangled queries: enabling declarative data-driven coordination
abstract
Many data-driven social and Web applications involve collaboration and coordination. The vision of declarative data-driven coordination (D3C), proposed in [9], is to support coordination in the spirit of data management: to make it data-centric and to specify it using convenient declarative languages. This paper introduces entangled queries, a language that extends SQL by constraints that allow for the coordinated choice of result tuples across queries originating from different users or applications.
Nitin Gupta 0003, Lucja Kot, Sudip Roy 0002, Gabriel Bender, Johannes Gehrke, Christoph Koch 0001
SIGMOD Conference3
2011 Entangled Transactions
Nitin Gupta 0003, Milos Nikolic 0001, Sudip Roy 0002, Gabriel Bender, Lucja Kot, Johannes Gehrke, Christoph Koch 0001
Proc. VLDB Endow.3
2008 A layout-aware physical design method for constructing feasible QCA circuits
abstract
Quantum-dot Cellular Automata (QCA) is an emerging computing paradigm, in which logical operations as well as signal transmission occurs due to Coulombic charge interaction between neighbouring QCA cells, moderated by a 4-phase QCA clock potential. Thermodynamic constraints like the number of QCA cells in a clocking zone must be obeyed to obtain a logically correct and feasible QCA circuit. These constraints depend on various design factors like total wirelength in a circuit, height of a clocking zone etc. which are not available until actual circuit layout is obtained. In this paper, the various design automation problems assosciated with obtaining a feasible QCA layout are addressed. The layout generation problem is formulated as embedding the netlist digraph in an orthogonal grid, which provides an abstraction of the actual physical layout to be obtained. Novel graph theoretic algorithms are proposed to perform placement and global routing and various design parameters like clock rate, wasted area and total wirelength are used to estimate the quality of the layout obtained. Also, planarization methods are used to remove all wire crossings, which are expensive to fabricate. The methods applied on a large number of MCNC'93 and ISCAS'89 benchmarks show good results.
Mayur Bubna, Sudip Roy 0002, Naresh Shenoy, Subhra Mazumdar 0002
ACM Great Lakes Symposium on VLSI2