Avery Ching

dblp:97/6329 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
2since 2021 · last 2021
0000-0002-1165-7871ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 6 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Storage systems · 26% High-performance computing · 23% Distributed systems · 20%
Databases, data mining, and information retrieval
1 paper
Graph data management · 67% Distributed and cloud data management · 33%

Topics — the 14 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Interconnection networks and networks-on-chip › interprocessor communication
all-to-all communication
0.312018
Riffle: optimized shuffle service for large-scale data analytics · EuroSys 2018
Distributed systems › distributed data processing
data shuffling
0.312018
Riffle: optimized shuffle service for large-scale data analytics · EuroSys 2018
Storage systems › i/o optimization
disk i/o optimization
0.312018
Riffle: optimized shuffle service for large-scale data analytics · EuroSys 2018
High-performance computing › data-intensive computing
large-scale data analytics
0.312018
Riffle: optimized shuffle service for large-scale data analytics · EuroSys 2018
Cloud and datacenter computing › big data platform
shuffle service
0.312018
Riffle: optimized shuffle service for large-scale data analytics · EuroSys 2018
Graph data management › graph processing
graph processing systems
0.212015
One Trillion Edges: Graph Processing at Facebook-Scale · Proc. VLDB Endow. 2015
Graph data management › graph processing
large-scale graph processing
0.212015
One Trillion Edges: Graph Processing at Facebook-Scale · Proc. VLDB Endow. 2015
Storage systems › file systems › distributed file system
parallel file system
0.232007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Noncontiguous locking techniques for parallel file systems · SC 2007
Exploring I/O Strategies for Parallel Sequence-Search Tools with S3aSim · HPDC 2006
High-performance computing
parallel i/o
0.232007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Exploring I/O Strategies for Parallel Sequence-Search Tools with S3aSim · HPDC 2006
Noncontiguous locking techniques for parallel file systems · SC 2007
Memory systems › cache management › storage caching
client-side caching
0.112007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Storage systems
file systems
0.112007
Noncontiguous locking techniques for parallel file systems · SC 2007
Storage systems › storage performance
write performance
0.112007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Memory systems
cache coherence
0.012007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
High-performance computing
scientific computing systems
0.012007
Noncontiguous locking techniques for parallel file systems · SC 2007

Methods — techniques the papers use, named apart from their topics

pregel model extensions · 0.4shuffle optimization · 0.3write-behind · 0.1thread-based caching · 0.1high-level access pattern information · 0.1simulation · 0.1
YearPublicationVenuePosition
2021 Twins: BFT Systems Made Robust
Shehar Bano, Alberto Sonnino, Andrey Chursin, Dmitri Perelman, Zekun Li 0009, Avery Ching, Dahlia Malkhi
OPODIS6
2021 Brief Announcement: Twins - BFT Systems Made Robust
abstract
Twins is an effective strategy for generating test scenarios with Byzantine [Lamport et al., 1982] nodes in order to find flaws in Byzantine Fault Tolerant (BFT) systems. Twins finds flaws in the design or implementation of BFT protocols that may cause correctness issues. The main idea of Twins is the following: running twin instances of a node that use correct, unmodified code and share the same network identity and credentials allows to emulate most interesting Byzantine behaviors. Because a twin executes normal, unmodified node code, building Twins only requires a thin wrapper over an existing distributed system designed for Byzantine tolerance. To emulate material, interesting scenarios with Byzantine nodes, it instantiates one or more twin copies of the node, giving the twins the same identities and network credentials as the original node. To the rest of the system, the node and all its twins appear indistinguishable from a single node behaving in a "questionable" manner. This approach generates many interesting Byzantine behaviors, including equivocation, double voting, and losing internal state, while forgoing uninteresting behavior scenarios that can be filtered at the transport layer, such as producing semantically invalid messages. Building on configurations with twin nodes, Twins systematically generates scenarios with Byzantine nodes via enumeration over protocol rounds and communication patterns among nodes. Despite this being inherently exponential, one new flaw and several known flaws were materialized by Twins in the arena of BFT consensus protocols. In all cases, protocols break within fewer than a dozen protocol rounds, hence it is realistic for the Twins approach to expose the problems. In two of these cases, it took the community more than a decade to discover protocol flaws that Twins would have surfaced within minutes. Additionally, Twins has been incorporated into the continuous release testing process of a production setting (DiemBFT) in which it can execute 44M Twins-generated scenarios daily.
Shehar Bano, Alberto Sonnino, Andrey Chursin, Dmitri Perelman, Zekun Li 0009, Avery Ching, Dahlia Malkhi
DISC6
2018 Riffle: optimized shuffle service for large-scale data analytics
abstract
The rapidly growing size of data and complexity of analytics present new challenges for large-scale data processing systems. Modern systems keep data partitions in memory for pipelined operators, and persist data across stages with wide dependencies on disks for fault tolerance. While processing can often scale well by splitting jobs into smaller tasks for better parallelism, all-to-all data transfer---called shuffle operations---become the scaling bottleneck when running many small tasks in multi-stage data analytics jobs. Our key observation is that this bottleneck is due to the superlinear increase in disk I/O operations as data volume increases.
Ergin Seyfe, Avery Ching, Michael J. Freedman
EuroSys4
2018 Generating Synthetic Social Graphs with Darwini
abstract
Synthetic graph generators facilitate research in graph algorithms and graph processing systems by providing access to graphs that resemble real social networks while addressing privacy and security concerns. Nevertheless, their practical value lies in their ability to capture important metrics of real graphs, such as degree distribution and clustering properties. Graph generators must also be able to produce such graphs at the scale of real-world industry graphs, that is, hundreds of billions or trillions of edges. In this paper, we propose Darwini, a graph generator that captures a number of core characteristics of real graphs. Importantly, given a source graph, it can reproduce the degree distribution and, unlike existing approaches, the local clustering coefficient distribution. Furthermore, Darwini maintains a number of metrics, such as graph assortativity, eigenvalues, and others. Comparing Darwini with state-of-the-art generative models, we show that it can reproduce these characteristics more accurately. Finally, we provide an open source implementation of Darwini on the vertex-centric Apache Giraph model that can generate synthetic graphs with up to 3 trillion edges.
Sergey Edunov, Dionysios Logothetis, Cheng Wang 0001, Avery Ching, Maja Kabiljo
ICDCS4
2015 One Trillion Edges: Graph Processing at Facebook-Scale
abstract
Analyzing large graphs provides valuable insights for social networking and web companies in content ranking and recommendations. While numerous graph processing systems have been developed and evaluated on available benchmark graphs of up to 6.6B edges, they often face significant difficulties in scaling to much larger graphs. Industry graphs can be two orders of magnitude larger - hundreds of billions or up to one trillion edges. In addition to scalability challenges, real world applications often require much more complex graph processing workflows than previously evaluated. In this paper, we describe the usability, performance, and scalability improvements we made to Apache Giraph, an open-source graph processing system, in order to use it on Facebook-scale graphs of up to one trillion edges. We also describe several key extensions to the original Pregel model that make it possible to develop a broader range of production graph applications and workflows as well as improve code reuse. Finally, we report on real-world operations as well as performance characteristics of several large-scale production applications.
Avery Ching, Sergey Edunov, Maja Kabiljo, Dionysios Logothetis, Sambavi Muthukrishnan
Proc. VLDB Endow.1
2007 Green Supercomputing in a Desktop Box
abstract
The advent of the Beowulf cluster in 1994 provided dedicated compute cycles, i.e., supercomputing for the masses, as a cost-effective alternative to large supercomputers, i.e., supercomputing for the few. However as the cluster movement matured, these clusters became like their large-scale supercomputing brethren - a shared (and power-hungry) datacenter resource that must reside in a actively-cooled machine room in order to operate properly. The above observation, coupled with the increasing performance gap between the PC and supercomputer, provides the motivation for a "green supercomputer" in a desktop box. Thus, this paper presents and evaluates such an architectural solution: a 12-node personal desktop supercomputer that offers an interactive environment for developing parallel codes and achieves 14 Gflops on Linpack but sips only 185 watts of power at load - all this in the approximate form factor of a Sun SPARCstation 1 pizza box.
Wu-chun Feng, Avery Ching, Chung-Hsing Hsu
IPDPS2
2007 Improving MPI Independent Write Performance Using A Two-Stage Write-Behind Buffering Method
abstract
Many large-scale production applications often have very long executions times and require periodic data checkpoints in order to save the state of the computation for program restart and/or tracing application progress. These write-only operations often dominate the overall application runtime, which makes them a good optimization target. Existing approaches for write-behind data buffering at the MPI I/O level have been proposed, but challenges still exist for addressing system-level I/O issues. We propose a two-stage write-behind buffering scheme for handing checkpoint operations. The first-stage of buffering accumulates write data for better network utilization and the second-stage of buffering enables the alignment for the write requests to the file stripe boundaries. Aligned I/O requests avoid file lock contention that can seriously degrade I/O performance. We present our performance evaluation using BTIO benchmarks on both GPFS and Lustre file systems. With the two-stage buffering, the performance of BTIO through MPI independent I/O is significantly improved and even surpasses that of collective I/O.
Wei-keng Liao, Avery Ching, Kenin Coloma, Alok N. Choudhary, Mahmut T. Kandemir
IPDPS2
2007 An Implementation and Evaluation of Client-Side File Caching for MPI-IO
abstract
Client-side file caching has long been recognized as a file system enhancement to reduce the amount of data transfer between application processes and I/O servers. However, caching also introduces cache coherence problems when a file is simultaneously accessed by multiple processes. Existing coherence controls tend to treat the client processes independently and ignore the aggregate I/O access pattern. This causes a serious performance degradation for parallel I/O applications. In this paper we discuss our new implementation and present an extended performance evaluation on GPFS and Lustre parallel file systems. In addition to comparing our methods to traditional approaches, we examine the performance of MPI-IO caching under direct I/O mode to bypass the underlying file system cache. We also investigate the performance impact of two file domain partitioning methods to MPI collective I/O operations: one which creates a balanced workload and the other which aligns accesses to the file system stripe size. In our experiments, alignment results in better performance by reducing file lock contention. When the cache page size is set to a multiple of the stripe size, MPI-IO caching inherits the same advantage and produces significantly improved I/O bandwidth.
Wei-keng Liao, Avery Ching, Kenin Coloma, Alok N. Choudhary, Lee Ward
IPDPS2
2007 Noncontiguous locking techniques for parallel file systems
abstract
Many parallel scientific applications use high-level I/O APIs that offer atomic I/O capabilities. Atomic I/O in current parallel file systems is often slow when multiple processes simultaneously access interleaved, shared files. Current atomic I/O solutions are not optimized for handling noncontiguous access patterns because current locking systems have a fixed file system block-based granularity and do not leverage high-level access pattern information.
Avery Ching, Wei-keng Liao, Alok N. Choudhary, Robert B. Ross, Lee Ward
SC1
2007 Using MPI file caching to improve parallel write performance for large-scale scientific applications
abstract
Typical large-scale scientific applications periodically write checkpoint files to save the computational state throughout execution. Existing parallel file systems improve such write-only I/O patterns through the use of client-side file caching and write-behind strategies. In distributed environments where files are rarely accessed by more than one client concurrently, file caching has achieved significant success; however, in parallel applications where multiple clients manipulate a shared file, cache coherence control can serialize I/O. We have designed a thread based caching layer for the MPI I/O library, which adds a portable caching system closer to user applications so more information about the application’s I/O patterns is available for better coherence control. We demonstrate the impact of our caching solution on parallel write performance with a comprehensive evaluation that includes a set of widely used I/O benchmarks and production application I/O kernels. 1.
Wei-keng Liao, Avery Ching, Kenin Coloma, Arifa Nisar, Alok N. Choudhary, Jacqueline Chen, Ramanan Sankaran, Scott Klasky
SC2
2006 Scalable Approaches for Supporting MPI-IO Atomicity
abstract
Scalable atomic and parallel access to noncontiguous regions of a file is essential to exploit high performance I/O as required by large-scale applications. Parallel I/O frameworks such as MPI I/O conceptually allow I/O to be defined on regions of a file using derived datatypes. Access to regions of a file can be automatically computed on a perprocessor basis using the datatype, resulting in a list of (offset, length) pairs. We describe three approaches for implementing lock serving (whole file, region locking, and byterange locking) and compare the various approaches using three noncontiguous I/O benchmarks. We present the details of the lock server architecture and describe the implementation of a fully-functional prototype that makes use of a lightweight message passing library and red/black trees.
Peter M. Aarestad, Avery Ching, George K. Thiruvathukal, Alok N. Choudhary
CCGRID2
2006 A New Flexible MPI Collective I/O Implementation
abstract
The MPI-IO standard creates a huge opportunity to break out of the traditional file system I/O methods. As a software layer between the user and the file system, an MPI-IO library can potentially optimize I/O on behalf of the user with little to no user intervention. This is all possible because of the rich data description and communication infrastructure MPI-2 offers. Powerful data descriptions and some of the other desirable features of MPI-2, however, make MPI-IO challenging to implement. By creating a new collective I/O implementation that allows developers to easily tinker and play with new optimizations or combinations of different techniques, research can proceed faster and be quickly and reliably deployed
Kenin Coloma, Avery Ching, Alok N. Choudhary, Wei-keng Liao, Robert B. Ross, Rajeev Thakur, Lee Ward
CLUSTER2
2006 Exploring I/O Strategies for Parallel Sequence-Search Tools with S3aSim
abstract
Parallel sequence-search tools are rising in popularity among computational biologists. With the rapid growth of sequence databases, database segmentation is the trend of the future for such search tools. While I/O currently is not a significant bottleneck for parallel sequence-search tools, future technologies including faster processors, customized computational hardware such as FPGAs, improved search algorithms, and exponentially growing databases emphasize an increasing need for efficient parallel I/O in future parallel sequence-search tools. Our paper focuses on examining different I/O strategies for these future tools in a modern parallel file system (PVFS2). Because implementing and comparing various I/O algorithms in every search tool is labor-intensive and time-consuming, we introduce S3aSim, a general simulation framework for sequence-search which allows us to quickly implement, test, and profile various I/O strategies. We examine a variety of I/O strategies (e.g., master-writing and various worker-writing strategies using individual and collective I/O methods) for storing result data in sequence-search tools such as mpiBLAST, pioBLAST, and parallel HMMer. Our experiments fully detail the interaction of computing and I/O within a full application simulation as opposed to typical I/O-only benchmarks
Avery Ching, Wu-chun Feng, Heshan Lin, Xiaosong Ma, Alok N. Choudhary
HPDC1
2006 Evaluating I/O characteristics and methods for storing structured scientific data
abstract
Many large-scale scientific simulations generate large, structured multi-dimensional datasets. Data is stored at various intervals on high performance I/O storage systems for checkpointing, post-processing, and visualization. Data storage is very I/O intensive and can dominate the overall running time of an application, depending on the characteristics of the I/O access pattern. Our NCIO benchmark determines how I/O characteristics greatly affect performance (up to 2 orders of magnitude) and provides scientific application developers with guidelines for improvement. In this paper, we examine the impact of various I/O parameters and methods when using the MPI-IO interface to store structured scientific data in an optimized parallel file system.
Avery Ching, Alok N. Choudhary, Wei-keng Liao, Lee Ward, Neil Pundit
IPDPS1
2003 Noncontiguous I/O Accesses Through MPI-IO
abstract
I/O performance remains a weakness of parallel computing systems today. While this weakness is partly attributed to rapid advances in other system components, I/O interfaces available to programmers and the I/O methods supported by file systems have traditionally not matched efficiently with the types of I/O operations that scientific applications perform, particularly noncontiguous accesses. The MPI-IO interface allows for rich descriptions of the I/O patterns desired for scientific applications and implementations such as ROMIO have taken advantage of this ability while remaining limited by underlying file system methods. A method of noncontiguous data access, list I/O, was recently implemented in the Parallel Virtual File System (PVFS). We implement support for this interface in the ROMIO MPI-IO implementation. Through a suite of noncontiguous I/O tests we compared ROMIO list I/O to current methods of ROMIO noncontiguous access and found that the list I/O interface provides performance benefits in many noncontiguous cases.
Avery Ching, Alok N. Choudhary, Kenin Coloma, Wei-keng Liao, Robert B. Ross, William Gropp
CCGRID1
2003 Efficient Structured Data Access in Parallel File Systems
abstract
Parallel scientific applications store and retrieve very large, structured datasets. Directly supporting these structured accesses is an important step in providing high-performance I/O solutions for these applications. High-level interfaces such as HDF5 and Parallel netCDF provide convenient APIs for accessing structured datasets, and the MPI-IO interface also supports efficient access to structured data. However, parallel file systems do not traditionally support such access. In this work we present an implementation of structured data access support in the context of the parallel virtual file system (PVFS). We call this support "datatype I/O" because of its similarity to MPI datatypes. This support is built by using a reusable datatype-processing component from the MPICH2 MPI implementation. We describe how this component is leveraged to efficiently process structured data representations resulting from MPI-IO operations. We quantitatively assess the solution using three test applications. We also point to further optimizations in the processing path that could be leveraged for even more efficient operation.
Avery Ching, Alok N. Choudhary, Wei-keng Liao, Robert B. Ross, William Gropp
CLUSTER1
2002 Noncontiguous I/O through PVFS
abstract
With the tremendous advances in processor and memory technology, I/O has risen to become the bottleneck in high-performance computing for many applications. The development of parallel file systems has helped to ease the performance gap, but I/O still remains an area needing significant performance improvement. Research has found that noncontiguous I/O access patterns in scientific applications combined with current file system methods, to perform these accesses lead to unacceptable performance for large data sets. To enhance performance of noncontiguous I/O, we have created list I/O, a native version of noncontiguous I/O. We have used the Parallel Virtual File System (PVFS) to implement our ideas. Our research and experimentation shows that list I/O outperforms current noncontiguous I/O access methods in most I/O situations and can substantially enhance the performance of real-world scientific applications.
Avery Ching, Alok N. Choudhary, Wei-keng Liao, Robert B. Ross, William Gropp
CLUSTER1