Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ronnie Chaiken

dblp:94/4948 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-authorSystems, architecture and hardware · 1Computer networks · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Distributed and cloud data management · 40% Query processing and optimization · 38% Database system architecture and tuning · 22%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 85% Storage systems · 15%
Computer networks
1 paper
Datacenter networks · 50% Network measurement and analytics · 50%
Network and information security
1 paper
Systems and software security · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management › mapreduce
mapreduce query processing
0.112012
SCOPE: parallel databases meet MapReduce · VLDB J. 2012
Database system architecture and tuning
parallel database system
0.112012
SCOPE: parallel databases meet MapReduce · VLDB J. 2012
Query processing and optimization › query optimization
cost-based optimization
0.112010
Incorporating partitioning and parallel plans into the SCOPE optimizer · ICDE 2010
Query processing and optimization
query optimization
0.112010
Incorporating partitioning and parallel plans into the SCOPE optimizer · ICDE 2010
Network measurement and analytics
traffic matrix estimation
0.112009
The nature of data center traffic: measurements & analysis · Internet Measurement Conference 2009
Distributed systems › service-oriented architecture › service management
service migration
0.112006
The SMART way to migrate replicated stateful services · EuroSys 2006
Distributed systems › replication
state machine replication
0.112006
The SMART way to migrate replicated stateful services · EuroSys 2006
Distributed systems
fault tolerance
0.122006
FARSITE: Federated, Available, and Reliable Storage for an Incompletely Trusted Environment · OSDI 2002
The SMART way to migrate replicated stateful services · EuroSys 2006
Systems and software security
untrusted platform
0.012002
FARSITE: Federated, Available, and Reliable Storage for an Incompletely Trusted Environment · OSDI 2002
Storage systems
distributed storage
0.012002
FARSITE: Federated, Available, and Reliable Storage for an Incompletely Trusted Environment · OSDI 2002
Distributed systems › fault tolerance
high availability
0.012006
The SMART way to migrate replicated stateful services · EuroSys 2006

Methods — techniques the papers use, named apart from their topics

mapreduce · 0.1transformation-based optimization · 0.1socket-level logging · 0.1query compilation · 0.1SQL-like declarative language · 0.1federation · 0.1pipelining · 0.1migration protocol · 0.1
YearPublicationVenuePosition
2012 SCOPE: parallel databases meet MapReduce
Jingren Zhou 0001, Nicolas Bruno, Ming-Chuan Wu, Per-Åke Larson, Ronnie Chaiken, Darren Shakib
VLDB J.5
2010 Incorporating partitioning and parallel plans into the SCOPE optimizer
abstract
Massive data analysis on large clusters presents new opportunities and challenges for query optimization. Data partitioning is crucial to performance in this environment. However, data repartitioning is a very expensive operation so minimizing the number of such operations can yield very significant performance improvements. A query optimizer for this environment must therefore be able to reason about data partitioning including its interaction with sorting and grouping. SCOPE is a SQL-like scripting language used at Microsoft for massive data analysis. A transformation-based optimizer is responsible for converting scripts into efficient execution plans for the Cosmos distributed computing platform. In this paper, we describe how reasoning about data partitioning is incorporated into the SCOPE optimizer. We show how relational operators affect partitioning, sorting and grouping properties and describe how the optimizer reasons about and exploits such properties to avoid unnecessary operations. In most optimizers, consideration of parallel plans is an afterthought done in a postprocessing step. Reasoning about partitioning enables the SCOPE optimizer to fully integrate consideration of parallel, serial and mixed plans into the cost-based optimization. The benefits are illustrated by showing the variety of plans enabled by our approach.
Jingren Zhou 0001, Per-Åke Larson, Ronnie Chaiken
ICDE3
2009 The nature of data center traffic: measurements & analysis
abstract
We explore the nature of traffic in data centers, designed to support the mining of massive data sets. We instrument the servers to collect socket-level logs, with negligible performance impact. In a 1500 server operational cluster, we thus amass roughly a petabyte of measurements over two months, from which we obtain and report detailed views of traffic and congestion conditions and patterns. We further consider whether traffic matrices in the cluster might be obtained instead via tomographic inference from coarser-grained counter data.
Srikanth Kandula, Sudipta Sengupta, Albert G. Greenberg, Parveen Patel, Ronnie Chaiken
Internet Measurement Conference5
2008 SCOPE: easy and efficient parallel processing of massive data sets
abstract
Companies providing cloud-scale services have an increasing need to store and analyze massive data sets such as search logs and click streams. For cost and performance reasons, processing is typically done on large clusters of shared-nothing commodity machines. It is imperative to develop a programming model that hides the complexity of the underlying system but provides flexibility by allowing users to extend functionality to meet a variety of requirements. In this paper, we present a new declarative and extensible scripting language, SCOPE (Structured Computations Optimized for Parallel Execution), targeted for this type of massive data analysis. The language is designed for ease of use with no explicit parallelism, while being amenable to efficient parallel execution on large clusters. SCOPE borrows several features from SQL. Data is modeled as sets of rows composed of typed columns. The select statement is retained with inner joins, outer joins, and aggregation allowed. Users can easily define their own functions and implement their own versions of operators: extractors (parsing and constructing rows from a file), processors (row-wise processing), reducers (group-wise processing), and combiners (combining rows from two inputs). SCOPE supports nesting of expressions but also allows a computation to be specified as a series of steps, in a manner often preferred by programmers. We also describe how scripts are compiled into efficient, parallel execution plans and executed on large clusters.
Ronnie Chaiken, Bob Jenkins, Per-Åke Larson, Bill Ramsey, Darren Shakib, Simon Weaver, Jingren Zhou 0001
Proc. VLDB Endow.1
2006 The SMART way to migrate replicated stateful services
abstract
Many stateful services use the replicated state machine approach for high availability. In this approach, a service runs on multiple machines to survive machine failures. This paper describes SMART, a new technique for changing the set of machines where such a service runs, i.e., migrating the service. SMART improves upon existing techniques in three important ways. First, SMART allows migrations that replace non-failed machines. Thus, SMART enables load balancing and lets an automated system replace failed machines. Such autonomic migration is an important step toward full autonomic operation, in which administrators play a minor role and need not be available twenty-four hours a day, seven days a week. Second, SMART can pipeline concurrent requests, a useful performance optimization. Third, prior published migration techniques are described in insufficient detail to admit implementation, whereas our description of SMART is complete. In addition to describing SMART, we also demonstrate its practicality by implementing it, evaluating our implementation’s performance, and using it to build a consistent, replicated, migratable file system. Our experiments demonstrate the performance advantage of pipelining concurrent requests, and show that migration has only a minor and temporary effect on performance.
Jacob R. Lorch, Atul Adya, William J. Bolosky, Ronnie Chaiken, John R. Douceur, Jon Howell
EuroSys4
2002 FARSITE: Federated, Available, and Reliable Storage for an Incompletely Trusted Environment
Atul Adya, William J. Bolosky, Miguel Castro 0001, Gerald Cermak, Ronnie Chaiken, John R. Douceur, Jon Howell, Jacob R. Lorch, Marvin Theimer, Roger Wattenhofer
OSDI5