James Cipar

dblp:34/1302 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Storage systems · 39% Cloud and datacenter computing · 33% Distributed systems · 22%
Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 39% Data models and query languages · 30% Query processing and optimization · 30%

Topics — the 19 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
distributed storage
0.422014
Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014
SpringFS: bridging agility and performance in elastic distributed storage · FAST 2014
Storage systems
file systems
0.332012
File system virtual appliances: Portable file system implementations · ACM Trans. Storage 2012
Contributing storage using the transparent file system · ACM Trans. Storage 2007
TFS: A Transparent File System for Contributory Storage · FAST 2007
Cloud and datacenter computing
big data analytics
0.212014
Exploiting Bounded Staleness to Speed Up Big Data Analytics · USENIX ATC 2014
Distributed systems › consistency models
bounded staleness
0.212014
Exploiting Bounded Staleness to Speed Up Big Data Analytics · USENIX ATC 2014
Cloud and datacenter computing
cloud storage
0.212014
SpringFS: bridging agility and performance in elastic distributed storage · FAST 2014
Storage systems
data migration
0.212014
Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014
Cloud and datacenter computing › datacenter storage
elastic storage
0.212014
Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014
Distributed systems
distributed machine learning
0.212013
More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server · NIPS 2013
Parallel and multicore computing
parallel programming models
0.212013
More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server · NIPS 2013
Distributed systems › distributed machine learning
parameter server
0.212013
More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server · NIPS 2013
Data models and query languages
NoSQL database
0.112012
LazyBase: trading freshness for performance in a scalable database · EuroSys 2012
Distributed and cloud data management › large-scale data management
scalable database systems
0.112012
LazyBase: trading freshness for performance in a scalable database · EuroSys 2012
Storage systems › file systems › file system design
transparent file system
0.122007
Contributing storage using the transparent file system · ACM Trans. Storage 2007
TFS: A Transparent File System for Contributory Storage · FAST 2007
Cloud and datacenter computing
virtualization
0.112012
File system virtual appliances: Portable file system implementations · ACM Trans. Storage 2012
Cloud and datacenter computing › virtualization
virtual machine
0.112012
File system virtual appliances: Portable file system implementations · ACM Trans. Storage 2012
Storage systems › distributed storage
peer-to-peer storage
0.122007
Contributing storage using the transparent file system · ACM Trans. Storage 2007
TFS: A Transparent File System for Contributory Storage · FAST 2007
Distributed systems
resource sharing
0.112006
Transparent Contribution of Memory · USENIX ATC, General Track 2006
Distributed systems
replication
0.012007
Contributing storage using the transparent file system · ACM Trans. Storage 2007
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.012006
Transparent Contribution of Memory · USENIX ATC, General Track 2006

Methods — techniques the papers use, named apart from their topics

virtual machine isolation · 0.3FS-agnostic proxy · 0.3trace analysis · 0.2read offloading · 0.2passive migration · 0.2bounded write offloading · 0.2stale synchronous parallel · 0.2parameter server · 0.2update scanning · 0.1pipelining · 0.1batching · 0.1background storage allocation · 0.1memory donation · 0.1
YearPublicationVenuePosition
2014 SpringFS: bridging agility and performance in elastic distributed storage
Lianghong Xu, James Cipar, Elie Krevat, Alexey Tumanov, Nitin Gupta 0001, Michael A. Kozuch, Gregory R. Ganger
FAST2
2014 Exploiting Bounded Staleness to Speed Up Big Data Analytics
Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai 0003, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, Eric P. Xing
USENIX ATC2
2014 Agility and Performance in Elastic Distributed Storage
abstract
Elastic storage systems can be expanded or contracted to meet current demand, allowing servers to be turned off or used for other tasks. However, the usefulness of an elastic distributed storage system is limited by its agility: how quickly it can increase or decrease its number of servers. Due to the large amount of data they must migrate during elastic resizing, state of the art designs usually have to make painful trade-offs among performance, elasticity, and agility. This article describes the state of the art in elastic storage and a new system, called SpringFS, that can quickly change its number of active servers, while retaining elasticity and performance goals. SpringFS uses a novel technique, termed bounded write offloading , that restricts the set of servers where writes to overloaded servers are redirected. This technique, combined with the read offloading and passive migration policies used in SpringFS, minimizes the work needed before deactivation or activation of servers. Analysis of real-world traces from Hadoop deployments at Facebook and various Cloudera customers and experiments with the SpringFS prototype confirm SpringFS’s agility, show that it reduces the amount of data migrated for elastic resizing by up to two orders of magnitude, and show that it cuts the percentage of active servers required by 67--82%, outdoing state-of-the-art designs by 6--120%.
Lianghong Xu, James Cipar, Elie Krevat, Alexey Tumanov, Nitin Gupta 0001, Michael A. Kozuch, Gregory R. Ganger
ACM Trans. Storage2
2013 Solving the Straggler Problem with Bounded Staleness
James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Gregory R. Ganger, Garth A. Gibson, Kimberly Keeton, Eric P. Xing
HotOS1
2013 More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server
abstract
We propose a parameter server system for distributed ML, which follows a Stale Synchronous Parallel (SSP) model of computation that maximizes the time computational workers spend doing useful work on ML algorithms, while still providing correctness guarantees. The parameter server provides an easy-to-use shared interface for read/write access to an ML model's values (parameters and variables), and the SSP model allows distributed workers to read older, stale versions of these values from a local cache, instead of waiting to get them from a central storage. This significantly increases the proportion of time workers spend computing, as opposed to waiting. Furthermore, the SSP model ensures ML algorithm correctness by limiting the maximum age of the stale values. We provide a proof of correctness under SSP, as well as empirical results demonstrating that the SSP model achieves faster algorithm convergence on several different ML problems, compared to fully-synchronous and asynchronous schemes.
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B. Gibbons, Garth A. Gibson, Gregory R. Ganger, Eric P. Xing
NIPS2
2012 alsched: algebraic scheduling of mixed workloads in heterogeneous clouds
abstract
As cloud resources and applications grow more heterogeneous, allocating the right resources to different tenants' activities increasingly depends upon understanding tradeoffs regarding their individual behaviors. One may require a specific amount of RAM, another may benefit from a GPU, and a third may benefit from executing on the same rack as a fourth. This paper promotes the need for and an approach for accommodating diverse tenant needs, based on having resource requests indicate any soft (i.e., when certain resource types would be better, but are not mandatory) and hard constraints in the form of composable utility functions. A scheduler that accepts such requests can then maximize overall utility, perhaps weighted by priorities, taking into account application specifics. Experiments with a prototype scheduler, called alsched, demonstrate that support for soft constraints is important for efficiency in multi-purpose clouds and that composable utility functions can provide it.
Alexey Tumanov, James Cipar, Gregory R. Ganger, Michael A. Kozuch
SoCC2
2012 LazyBase: trading freshness for performance in a scalable database
abstract
The LazyBase scalable database system is specialized for the growing class of data analysis applications that extract knowledge from large, rapidly changing data sets. It provides the scalability of popular NoSQL systems without the query-time complexity associated with their eventual consistency models, offering a clear consistency model and explicit per-query control over the trade-off between latency and result freshness. With an architecture designed around batching and pipelining of updates, LazyBase simultaneously ingests atomic batches of updates at a very high throughput and offers quick read queries to a stale-but-consistent version of the data. Although slightly stale results are sufficient for many analysis queries, fully up-to-date results can be obtained when necessary by also scanning updates still in the pipeline. Compared to the Cassandra NoSQL system, LazyBase provides 4X--5X faster update throughput and 4X faster read query throughput for range queries while remaining competitive for point queries. We demonstrate LazyBase's tradeoff between query latency and result freshness as well as the benefits of its consistency model. We also demonstrate specific cases where Cassandra's consistency model is weaker than LazyBase's.
James Cipar, Gregory R. Ganger, Kimberly Keeton, Charles B. Morrey III, Craig A. N. Soules, Alistair C. Veitch
EuroSys1
2012 File system virtual appliances: Portable file system implementations
abstract
File system virtual appliances (FSVAs) address the portability headaches that plague file system (FS) developers. By packaging their FS implementation in a virtual machine (VM), separate from the VM that runs user applications, they can avoid the need to port the file system to each operating system (OS) and OS version. A small FS-agnostic proxy, maintained by the core OS developers, connects the FSVA to whatever OS the user chooses. This article describes an FSVA design that maintains FS semantics for unmodified FS implementations and provides desired OS and virtualization features, such as a unified buffer cache and VM migration. Evaluation of prototype FSVA implementations in Linux and NetBSD, using Xen as the virtual machine manager (VMM), demonstrates that the FSVA architecture is efficient, FS-agnostic, and able to insulate file system implementations from OS differences that would otherwise require explicit porting.
Michael Abd-El-Malek, Matthew Wachs, James Cipar, Karan Sanghi, Gregory R. Ganger, Garth A. Gibson, Michael K. Reiter
ACM Trans. Storage3
2010 Robust and flexible power-proportional storage
abstract
Power-proportional cluster-based storage is an important component of an overall cloud computing infrastructure. With it, substantial subsets of nodes in the storage cluster can be turned off to save power during periods of low utilization. Rabbit is a distributed file system that arranges its data-layout to provide ideal power-proportionality down to very low minimum number of powered-up nodes (enough to store a primary replica of available datasets). Rabbit addresses the node failure rates of large-scale clusters with data layouts that minimize the number of nodes that must be powered-up if a primary fails. Rabbit also allows different datasets to use different subsets of nodes as a building block for interference avoidance when the infrastructure is shared by multiple tenants. Experiments with a Rabbit prototype demonstrate its power-proportionality, and simulation experiments demonstrate its properties at scale.
Hrishikesh Amur, James Cipar, Varun Gupta 0004, Gregory R. Ganger, Michael A. Kozuch, Karsten Schwan
SoCC2
2007 TFS: A Transparent File System for Contributory Storage
James Cipar, Mark D. Corner, Emery D. Berger
FAST1
2007 Contributing storage using the transparent file system
abstract
Contributory applications allow users to donate unused resources on their personal computers to a shared pool. Applications such as SETI@home, Folding@home, and Freenet are now in wide use and provide a variety of services, including data processing and content distribution. However, while several research projects have proposed contributory applications that support peer-to-peer storage systems, their adoption has been comparatively limited. We believe that a key barrier to the adoption of contributory storage systems is that contributing a large quantity of local storage interferes with the principal user of the machine. To overcome this barrier, we introduce the Transparent File System (TFS). TFS provides background tasks with large amounts of unreliable storage—all of the currently available space—without impacting the performance of ordinary file access operations. We show that TFS allows a peer-to-peer contributory storage system to provide 40% more storage at twice the performance when compared to a user-space storage mechanism. We analyze the impact of TFS on replication in peer-to-peer storage systems and show that TFS does not appreciably increase the resources needed for file replication.
James Cipar, Mark D. Corner, Emery D. Berger
ACM Trans. Storage1
2006 Transparent Contribution of Memory
James Cipar, Mark D. Corner, Emery D. Berger
USENIX ATC, General Track1