EDBT 2026 Demo / reviewers in the wild / expert
James Cipar
dblp:34/1302
· DBLP profile ↗
12ranked-venue papers
5as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Storage systems · 39% Cloud and datacenter computing · 33% Distributed systems · 22% | |
| Databases, data mining, and information retrieval
1 paper |
Distributed and cloud data management · 39% Data models and query languages · 30% Query processing and optimization · 30% |
Topics — the 19 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
distributed storage |
0.4 | 2 | 2014 | Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014 SpringFS: bridging agility and performance in elastic distributed storage · FAST 2014 |
Storage systems
file systems |
0.3 | 3 | 2012 | File system virtual appliances: Portable file system implementations · ACM Trans. Storage 2012 Contributing storage using the transparent file system · ACM Trans. Storage 2007 TFS: A Transparent File System for Contributory Storage · FAST 2007 |
Cloud and datacenter computing
big data analytics |
0.2 | 1 | 2014 | Exploiting Bounded Staleness to Speed Up Big Data Analytics · USENIX ATC 2014 |
Distributed systems › consistency models
bounded staleness |
0.2 | 1 | 2014 | Exploiting Bounded Staleness to Speed Up Big Data Analytics · USENIX ATC 2014 |
Cloud and datacenter computing
cloud storage |
0.2 | 1 | 2014 | SpringFS: bridging agility and performance in elastic distributed storage · FAST 2014 |
Storage systems
data migration |
0.2 | 1 | 2014 | Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014 |
Cloud and datacenter computing › datacenter storage
elastic storage |
0.2 | 1 | 2014 | Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014 |
Distributed systems
distributed machine learning |
0.2 | 1 | 2013 | More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server · NIPS 2013 |
Parallel and multicore computing
parallel programming models |
0.2 | 1 | 2013 | More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server · NIPS 2013 |
Distributed systems › distributed machine learning
parameter server |
0.2 | 1 | 2013 | More Effective Distributed ML via a Stale Synchronous Parallel Parameter Server · NIPS 2013 |
Data models and query languages
NoSQL database |
0.1 | 1 | 2012 | LazyBase: trading freshness for performance in a scalable database · EuroSys 2012 |
Distributed and cloud data management › large-scale data management
scalable database systems |
0.1 | 1 | 2012 | LazyBase: trading freshness for performance in a scalable database · EuroSys 2012 |
Storage systems › file systems › file system design
transparent file system |
0.1 | 2 | 2007 | Contributing storage using the transparent file system · ACM Trans. Storage 2007 TFS: A Transparent File System for Contributory Storage · FAST 2007 |
Cloud and datacenter computing
virtualization |
0.1 | 1 | 2012 | File system virtual appliances: Portable file system implementations · ACM Trans. Storage 2012 |
Cloud and datacenter computing › virtualization
virtual machine |
0.1 | 1 | 2012 | File system virtual appliances: Portable file system implementations · ACM Trans. Storage 2012 |
Storage systems › distributed storage
peer-to-peer storage |
0.1 | 2 | 2007 | Contributing storage using the transparent file system · ACM Trans. Storage 2007 TFS: A Transparent File System for Contributory Storage · FAST 2007 |
Distributed systems
resource sharing |
0.1 | 1 | 2006 | Transparent Contribution of Memory · USENIX ATC, General Track 2006 |
Distributed systems
replication |
0.0 | 1 | 2007 | Contributing storage using the transparent file system · ACM Trans. Storage 2007 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 2006 | Transparent Contribution of Memory · USENIX ATC, General Track 2006 |
Methods — techniques the papers use, named apart from their topics
virtual machine isolation · 0.3FS-agnostic proxy · 0.3trace analysis · 0.2read offloading · 0.2passive migration · 0.2bounded write offloading · 0.2stale synchronous parallel · 0.2parameter server · 0.2update scanning · 0.1pipelining · 0.1batching · 0.1background storage allocation · 0.1memory donation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | SpringFS: bridging agility and performance in elastic distributed storage
Lianghong Xu, James Cipar, Elie Krevat, Alexey Tumanov, Nitin Gupta 0001, Michael A. Kozuch, Gregory R. Ganger |
FAST | 2 |
| 2014 | Exploiting Bounded Staleness to Speed Up Big Data Analytics
Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai 0003, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, Eric P. Xing |
USENIX ATC | 2 |
| 2014 | Agility and Performance in Elastic Distributed StorageabstractElastic storage systems can be expanded or contracted to meet current demand, allowing servers to be turned off or used for other tasks. However, the usefulness of an elastic distributed storage system is limited by its agility: how quickly it can increase or decrease its number of servers. Due to the large amount of data they must migrate during elastic resizing, state of the art designs usually have to make painful trade-offs among performance, elasticity, and agility. This article describes the state of the art in elastic storage and a new system, called SpringFS, that can quickly change its number of active servers, while retaining elasticity and performance goals. SpringFS uses a novel technique, termed bounded write offloading , that restricts the set of servers where writes to overloaded servers are redirected. This technique, combined with the read offloading and passive migration policies used in SpringFS, minimizes the work needed before deactivation or activation of servers. Analysis of real-world traces from Hadoop deployments at Facebook and various Cloudera customers and experiments with the SpringFS prototype confirm SpringFS’s agility, show that it reduces the amount of data migrated for elastic resizing by up to two orders of magnitude, and show that it cuts the percentage of active servers required by 67--82%, outdoing state-of-the-art designs by 6--120%. Lianghong Xu, James Cipar, Elie Krevat, Alexey Tumanov, Nitin Gupta 0001, Michael A. Kozuch, Gregory R. Ganger |
ACM Trans. Storage | 2 |
| 2013 | Solving the Straggler Problem with Bounded Staleness
James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Gregory R. Ganger, Garth A. Gibson, Kimberly Keeton, Eric P. Xing |
HotOS | 1 |
| 2013 | More Effective Distributed ML via a Stale Synchronous Parallel Parameter ServerabstractWe propose a parameter server system for distributed ML, which follows a Stale Synchronous Parallel (SSP) model of computation that maximizes the time computational workers spend doing useful work on ML algorithms, while still providing correctness guarantees. The parameter server provides an easy-to-use shared interface for read/write access to an ML model's values (parameters and variables), and the SSP model allows distributed workers to read older, stale versions of these values from a local cache, instead of waiting to get them from a central storage. This significantly increases the proportion of time workers spend computing, as opposed to waiting. Furthermore, the SSP model ensures ML algorithm correctness by limiting the maximum age of the stale values. We provide a proof of correctness under SSP, as well as empirical results demonstrating that the SSP model achieves faster algorithm convergence on several different ML problems, compared to fully-synchronous and asynchronous schemes. Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B. Gibbons, Garth A. Gibson, Gregory R. Ganger, Eric P. Xing |
NIPS | 2 |
| 2012 | alsched: algebraic scheduling of mixed workloads in heterogeneous cloudsabstractAs cloud resources and applications grow more heterogeneous, allocating the right resources to different tenants' activities increasingly depends upon understanding tradeoffs regarding their individual behaviors. One may require a specific amount of RAM, another may benefit from a GPU, and a third may benefit from executing on the same rack as a fourth. This paper promotes the need for and an approach for accommodating diverse tenant needs, based on having resource requests indicate any soft (i.e., when certain resource types would be better, but are not mandatory) and hard constraints in the form of composable utility functions. A scheduler that accepts such requests can then maximize overall utility, perhaps weighted by priorities, taking into account application specifics. Experiments with a prototype scheduler, called alsched, demonstrate that support for soft constraints is important for efficiency in multi-purpose clouds and that composable utility functions can provide it. Alexey Tumanov, James Cipar, Gregory R. Ganger, Michael A. Kozuch |
SoCC | 2 |
| 2012 | LazyBase: trading freshness for performance in a scalable databaseabstractThe LazyBase scalable database system is specialized for the growing class of data analysis applications that extract knowledge from large, rapidly changing data sets. It provides the scalability of popular NoSQL systems without the query-time complexity associated with their eventual consistency models, offering a clear consistency model and explicit per-query control over the trade-off between latency and result freshness. With an architecture designed around batching and pipelining of updates, LazyBase simultaneously ingests atomic batches of updates at a very high throughput and offers quick read queries to a stale-but-consistent version of the data. Although slightly stale results are sufficient for many analysis queries, fully up-to-date results can be obtained when necessary by also scanning updates still in the pipeline. Compared to the Cassandra NoSQL system, LazyBase provides 4X--5X faster update throughput and 4X faster read query throughput for range queries while remaining competitive for point queries. We demonstrate LazyBase's tradeoff between query latency and result freshness as well as the benefits of its consistency model. We also demonstrate specific cases where Cassandra's consistency model is weaker than LazyBase's. James Cipar, Gregory R. Ganger, Kimberly Keeton, Charles B. Morrey III, Craig A. N. Soules, Alistair C. Veitch |
EuroSys | 1 |
| 2012 | File system virtual appliances: Portable file system implementationsabstractFile system virtual appliances (FSVAs) address the portability headaches that plague file system (FS) developers. By packaging their FS implementation in a virtual machine (VM), separate from the VM that runs user applications, they can avoid the need to port the file system to each operating system (OS) and OS version. A small FS-agnostic proxy, maintained by the core OS developers, connects the FSVA to whatever OS the user chooses. This article describes an FSVA design that maintains FS semantics for unmodified FS implementations and provides desired OS and virtualization features, such as a unified buffer cache and VM migration. Evaluation of prototype FSVA implementations in Linux and NetBSD, using Xen as the virtual machine manager (VMM), demonstrates that the FSVA architecture is efficient, FS-agnostic, and able to insulate file system implementations from OS differences that would otherwise require explicit porting. Michael Abd-El-Malek, Matthew Wachs, James Cipar, Karan Sanghi, Gregory R. Ganger, Garth A. Gibson, Michael K. Reiter |
ACM Trans. Storage | 3 |
| 2010 | Robust and flexible power-proportional storageabstractPower-proportional cluster-based storage is an important component of an overall cloud computing infrastructure. With it, substantial subsets of nodes in the storage cluster can be turned off to save power during periods of low utilization. Rabbit is a distributed file system that arranges its data-layout to provide ideal power-proportionality down to very low minimum number of powered-up nodes (enough to store a primary replica of available datasets). Rabbit addresses the node failure rates of large-scale clusters with data layouts that minimize the number of nodes that must be powered-up if a primary fails. Rabbit also allows different datasets to use different subsets of nodes as a building block for interference avoidance when the infrastructure is shared by multiple tenants. Experiments with a Rabbit prototype demonstrate its power-proportionality, and simulation experiments demonstrate its properties at scale. Hrishikesh Amur, James Cipar, Varun Gupta 0004, Gregory R. Ganger, Michael A. Kozuch, Karsten Schwan |
SoCC | 2 |
| 2007 | TFS: A Transparent File System for Contributory Storage
James Cipar, Mark D. Corner, Emery D. Berger |
FAST | 1 |
| 2007 | Contributing storage using the transparent file systemabstractContributory applications allow users to donate unused resources on their personal computers to a shared pool. Applications such as SETI@home, Folding@home, and Freenet are now in wide use and provide a variety of services, including data processing and content distribution. However, while several research projects have proposed contributory applications that support peer-to-peer storage systems, their adoption has been comparatively limited. We believe that a key barrier to the adoption of contributory storage systems is that contributing a large quantity of local storage interferes with the principal user of the machine. To overcome this barrier, we introduce the Transparent File System (TFS). TFS provides background tasks with large amounts of unreliable storage—all of the currently available space—without impacting the performance of ordinary file access operations. We show that TFS allows a peer-to-peer contributory storage system to provide 40% more storage at twice the performance when compared to a user-space storage mechanism. We analyze the impact of TFS on replication in peer-to-peer storage systems and show that TFS does not appreciably increase the resources needed for file replication. James Cipar, Mark D. Corner, Emery D. Berger |
ACM Trans. Storage | 1 |
| 2006 | Transparent Contribution of Memory
James Cipar, Mark D. Corner, Emery D. Berger |
USENIX ATC, General Track | 1 |