Srinath Shankar

dblp:47/3052 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Distributed and cloud data management · 42% Query processing and optimization · 31% Database system architecture and tuning · 20%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 81% Distributed systems · 13% Parallel and multicore computing · 6%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
0.322014
Towards Multi-Tenant Performance SLOs · IEEE Trans. Knowl. Data Eng. 2014
Data driven workflow planning in cluster management systems · HPDC 2007
Distributed and cloud data management
distributed query processing
0.212013
Split query processing in polybase · SIGMOD Conference 2013
Distributed and cloud data management
federated database
0.212013
Split query processing in polybase · SIGMOD Conference 2013
Query processing and optimization › query optimization
distributed query optimization
0.112012
Query optimization in microsoft SQL server PDW · SIGMOD Conference 2012
Database system architecture and tuning
workload management
0.112012
Towards Multi-tenant Performance SLOs · ICDE 2012
Cloud and datacenter computing
database-as-a-service
0.112012
Towards Multi-tenant Performance SLOs · ICDE 2012
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112008
Clustera: an integrated computation and data management system · Proc. VLDB Endow. 2008
Cloud and datacenter computing
job scheduling
0.112008
Clustera: an integrated computation and data management system · Proc. VLDB Endow. 2008
Distributed systems
distributed coordination and fault tolerance
0.112007
Data driven workflow planning in cluster management systems · HPDC 2007
Recommender systems › large-scale recommendation › multi-stage recommender systems
candidate generation
0.112006
Database support for matching: limitations and opportunities · SIGMOD Conference 2006
Query processing and optimization
join processing
0.112006
Database support for matching: limitations and opportunities · SIGMOD Conference 2006
Distributed and cloud data management › cloud database
database-as-a-service
0.112014
Towards Multi-Tenant Performance SLOs · IEEE Trans. Knowl. Data Eng. 2014
Query processing and optimization › query optimization
cost-based optimization
0.012013
Split query processing in polybase · SIGMOD Conference 2013
Database system architecture and tuning
parallel database system
0.012012
Query optimization in microsoft SQL server PDW · SIGMOD Conference 2012
Parallel and multicore computing › parallel architecture
massively parallel processing
0.012012
Query optimization in microsoft SQL server PDW · SIGMOD Conference 2012
Distributed systems
distributed caching
0.012007
Data driven workflow planning in cluster management systems · HPDC 2007

Methods — techniques the papers use, named apart from their topics

workload modeling · 0.4optimization · 0.4scheduling · 0.3cost optimization · 0.3mapreduce · 0.2cost-based optimization · 0.2data-aware matchmaking · 0.1adaptive scheduling · 0.1sorting · 0.1grouping · 0.1graph matching · 0.1
YearPublicationVenuePosition
2014 Towards Multi-Tenant Performance SLOs
abstract
As traditional and mission-critical relational database workloads migrate to the cloud in the form of Database-as-a-Service (DaaS), there is an increasing motivation to provide performance goals in Service Level Objectives (SLOs). Providing such performance goals is challenging for DaaS providers as they must balance the performance that they can deliver to tenants and the data center's operating costs. In general, aggressively aggregating tenants on each server reduces the operating costs but degrades performance for the tenants, and vice versa. In this paper, we present a framework that takes as input the tenant workloads, their performance SLOs, and the server hardware that is available to the DaaS provider, and outputs a cost-effective recipe that specifies how much hardware to provision and how to schedule the tenants on each hardware resource. We evaluate our method and show that it produces effective solutions that can reduce the costs for the DaaS provider while meeting performance goals.
Willis Lang, Srinath Shankar, Jignesh M. Patel, Ajay Kalhan
IEEE Trans. Knowl. Data Eng.2
2013 Split query processing in polybase
abstract
This paper presents Polybase, a feature of SQL Server PDW V2 that allows users to manage and query data stored in a Hadoop cluster using the standard SQL query language. Unlike other database systems that provide only a relational view over HDFS-resident data through the use of an external table mechanism, Polybase employs a split query processing paradigm in which SQL operators on HDFS-resident data are translated into MapReduce jobs by the PDW query optimizer and then executed on the Hadoop cluster. The paper describes the design and implementation of Polybase along with a thorough performance evaluation that explores the benefits of employing a split query processing paradigm for executing queries that involve both structured data in a relational DBMS and unstructured data in Hadoop. Our results demonstrate that while the use of a split-based query execution paradigm can improve the performance of some queries by as much as 10X, one must employ a cost-based query optimizer that considers a broad set of factors when deciding whether or not it is advantageous to push a SQL operator to Hadoop. These factors include the selectivity factor of the predicate, the relative sizes of the two clusters, and whether or not their nodes are co-located. In addition, differences in the semantics of the Java and SQL languages must be carefully considered in order to avoid altering the expected results of a query.
David J. DeWitt, Alan Halverson, Rimma V. Nehme, Srinath Shankar, Josep Aguilar-Saborit, Artin Avanes, Miro Flasza, Jim Gramling
SIGMOD Conference4
2012 Towards Multi-tenant Performance SLOs
abstract
As traditional and mission-critical relational database workloads migrate to the cloud in the form of Database-as-a-Service (DaaS), there is an increasing motivation to provide performance goals in Service Level Objectives (SLOs). Providing such performance goals is challenging for DaaS providers as they must balance the performance that they can deliver to tenants and the data center's operating costs. In general, aggressively aggregating tenants on each server reduces the operating costs but degrades performance for the tenants, and vice versa. In this paper, we present a framework that takes as input the tenant workloads, their performance SLOs, and the server hardware that is available to the DaaS provider, and outputs a cost-effective recipe that specifies how much hardware to provision and how to schedule the tenants on each hardware resource. We evaluate our method and show that it produces effective solutions that can reduce the costs for the DaaS provider while meeting performance goals.
Willis Lang, Srinath Shankar, Jignesh M. Patel, Ajay Kalhan
ICDE2
2012 Query optimization in microsoft SQL server PDW
abstract
In recent years, Massively Parallel Processors have increasingly been used to manage and query vast amounts of data. Dramatic performance improvements are achieved through distributed execution of queries across many nodes. Query optimization for such system is a challenging and important problem.
Srinath Shankar, Rimma V. Nehme, Josep Aguilar-Saborit, Mostafa Elhemali, Alan Halverson, Eric Robinson, Mahadevan Sankara Subramanian, David J. DeWitt, César A. Galindo-Legaria
SIGMOD Conference1
2010 Wimpy node clusters: what about non-wimpy workloads?
abstract
The high cost associated with powering servers has introduced new challenges in improving the energy efficiency of clusters running data processing jobs. Traditional high-performance servers are largely energy inefficient due to various factors such as the over-provisioning of resources. The increasing trend to replace traditional high-performance server nodes with low-power low-end nodes in clusters has recently been touted as a solution to the cluster energy problem. However, the key tacit assumption that drives such a solution is that the proportional scale-out of such low-power cluster nodes results in constant scaleup in performance. This paper studies the validity of such an assumption using measured price and performance results from a low-power Atom-based node and a traditional Xeon-based server and a number of published parallel scaleup results. Our results show that in most cases, computationally complex queries exhibit disproportionate scaleup characteristics which potentially makes scale-out with low-end nodes an expensive and lower performance solution.
Willis Lang, Jignesh M. Patel, Srinath Shankar
DaMoN3
2008 Clustera: an integrated computation and data management system
abstract
This paper introduces Clustera, an integrated computation and data management system. In contrast to traditional cluster-management systems that target specific types of workloads, Clustera is designed for extensibility, enabling the system to be easily extended to handle a wide variety of job types ranging from computationally-intensive, long-running jobs with minimal I/O requirements to complex SQL queries over massive relational tables. Another unique feature of Clustera is the way in which the system architecture exploits modern software building blocks including application servers and relational database systems in order to realize important performance, scalability, portability and usability benefits. Finally, experimental evaluation suggests that Clustera has good scale-up properties for SQL processing, that Clustera delivers performance comparable to Hadoop for MapReduce processing and that Clustera can support higher job throughput rates than previously published results for the Condor and CondorJ2 batch computing systems.
David J. DeWitt, Erik Paulson 0001, Eric Robinson, Jeffrey F. Naughton, Joshua Royalty, Srinath Shankar, Andrew Krioukov
Proc. VLDB Endow.6
2007 Data driven workflow planning in cluster management systems
abstract
Traditional scientific computing has been associated with harnessing computation cycles within and across clusters of machines. In recent years, scientific applications have become increasingly data-intensive. This is especially true in the fields of astronomy and high energy physics. Furthermore, the lowered cost of disks and commodity machines has led to a dramatic increase in the amount of free disk space spread across machines in a cluster. This space is not being exploited by traditional distributed computing tools. In this paper we have evaluated ways to improve the data management capabilities of Condor, a popular distributed computing system. We have augmented the Condor system by providing the capability to store data used and produced by workflows on the disks of machines in the cluster. We have also replaced the Condor matchmaker with a new workflow planning framework that is cognizant of dependencies between jobs in a workflow and exploits these new data storage capabilities to produce workflow schedules. We show that our data caching and workflow planning framework can significantly reduce response times for data-intensive workflows by reducing data transfer over the network in a cluster. We also consider ways in which this planning framework can be made adaptive in a dynamic, multi-user, failure-prone environment.
Srinath Shankar, David J. DeWitt
HPDC1
2006 Database support for matching: limitations and opportunities
abstract
We define a match join of R and S with predicate θ to be a subset of the θ-join of R and S such that each tuple of R and S contributes to at most one result tuple. Match joins and their generalizations belong to a broad class of matching problems that have attracted a great deal of attention in disciplines including operations research and theoretical computer science. Instances of these problems arise in practice in resource allocation scenarios. To the best of our knowledge no one uses an RDBMS as a tool to help solve these problems; our goal in this paper is to explore whether or not this needs to be the case. We show that the simple approach of computing the full θ-join and then applying standard graph-matching algorithms to the result is ineffective for all but the smallest of problem instances. By contrast, a closer study shows that the DBMS primitives of grouping, sorting, and joining can be exploited to yield efficient match join operations. This suggests that RDBMSs can play a role in matching related problems beyond merely serving as expensive file systems exporting data sets to external user programs.
Ameet Kini, Srinath Shankar, Jeffrey F. Naughton, David J. DeWitt
SIGMOD Conference2