Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Kenin Coloma

dblp:17/3689 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
0since 2021 · last 2007
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 35% High-performance computing · 33% Storage systems · 32%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › file systems › distributed file system
parallel file system
0.232007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Scalable Design and Implementations for MPI Parallel Overlapping I/O · IEEE Trans. Parallel Distributed Syst. 2006
Collective caching: application-aware client-side file caching · HPDC 2005
High-performance computing
parallel i/o
0.232007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Scalable Design and Implementations for MPI Parallel Overlapping I/O · IEEE Trans. Parallel Distributed Syst. 2006
Collective caching: application-aware client-side file caching · HPDC 2005
Memory systems › cache management › storage caching
client-side caching
0.122007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Collective caching: application-aware client-side file caching · HPDC 2005
Memory systems
cache coherence
0.132007
Scalable Design and Implementations for MPI Parallel Overlapping I/O · IEEE Trans. Parallel Distributed Syst. 2006
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Collective caching: application-aware client-side file caching · HPDC 2005
Storage systems › storage performance
write performance
0.112007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
High-performance computing › parallel i/o
MPI-IO
0.112006
Scalable Design and Implementations for MPI Parallel Overlapping I/O · IEEE Trans. Parallel Distributed Syst. 2006
Memory systems › cache management › storage caching
cooperative caching
0.112005
Collective caching: application-aware client-side file caching · HPDC 2005
High-performance computing
performance optimization at scale
0.012006
Scalable Design and Implementations for MPI Parallel Overlapping I/O · IEEE Trans. Parallel Distributed Syst. 2006

Methods — techniques the papers use, named apart from their topics

write-behind · 0.1thread-based caching · 0.1process-rank ordering · 0.1persistent file domains · 0.1graph coloring · 0.1collective caching · 0.1
YearPublicationVenuePosition
2007 Improving MPI Independent Write Performance Using A Two-Stage Write-Behind Buffering Method
abstract
Many large-scale production applications often have very long executions times and require periodic data checkpoints in order to save the state of the computation for program restart and/or tracing application progress. These write-only operations often dominate the overall application runtime, which makes them a good optimization target. Existing approaches for write-behind data buffering at the MPI I/O level have been proposed, but challenges still exist for addressing system-level I/O issues. We propose a two-stage write-behind buffering scheme for handing checkpoint operations. The first-stage of buffering accumulates write data for better network utilization and the second-stage of buffering enables the alignment for the write requests to the file stripe boundaries. Aligned I/O requests avoid file lock contention that can seriously degrade I/O performance. We present our performance evaluation using BTIO benchmarks on both GPFS and Lustre file systems. With the two-stage buffering, the performance of BTIO through MPI independent I/O is significantly improved and even surpasses that of collective I/O.
Wei-keng Liao, Avery Ching, Kenin Coloma, Alok N. Choudhary, Mahmut T. Kandemir
IPDPS3
2007 An Implementation and Evaluation of Client-Side File Caching for MPI-IO
abstract
Client-side file caching has long been recognized as a file system enhancement to reduce the amount of data transfer between application processes and I/O servers. However, caching also introduces cache coherence problems when a file is simultaneously accessed by multiple processes. Existing coherence controls tend to treat the client processes independently and ignore the aggregate I/O access pattern. This causes a serious performance degradation for parallel I/O applications. In this paper we discuss our new implementation and present an extended performance evaluation on GPFS and Lustre parallel file systems. In addition to comparing our methods to traditional approaches, we examine the performance of MPI-IO caching under direct I/O mode to bypass the underlying file system cache. We also investigate the performance impact of two file domain partitioning methods to MPI collective I/O operations: one which creates a balanced workload and the other which aligns accesses to the file system stripe size. In our experiments, alignment results in better performance by reducing file lock contention. When the cache page size is set to a multiple of the stripe size, MPI-IO caching inherits the same advantage and produces significantly improved I/O bandwidth.
Wei-keng Liao, Avery Ching, Kenin Coloma, Alok N. Choudhary, Lee Ward
IPDPS3
2007 Using MPI file caching to improve parallel write performance for large-scale scientific applications
abstract
Typical large-scale scientific applications periodically write checkpoint files to save the computational state throughout execution. Existing parallel file systems improve such write-only I/O patterns through the use of client-side file caching and write-behind strategies. In distributed environments where files are rarely accessed by more than one client concurrently, file caching has achieved significant success; however, in parallel applications where multiple clients manipulate a shared file, cache coherence control can serialize I/O. We have designed a thread based caching layer for the MPI I/O library, which adds a portable caching system closer to user applications so more information about the application’s I/O patterns is available for better coherence control. We demonstrate the impact of our caching solution on parallel write performance with a comprehensive evaluation that includes a set of widely used I/O benchmarks and production application I/O kernels. 1.
Wei-keng Liao, Avery Ching, Kenin Coloma, Arifa Nisar, Alok N. Choudhary, Jacqueline Chen, Ramanan Sankaran, Scott Klasky
SC3
2006 A New Flexible MPI Collective I/O Implementation
abstract
The MPI-IO standard creates a huge opportunity to break out of the traditional file system I/O methods. As a software layer between the user and the file system, an MPI-IO library can potentially optimize I/O on behalf of the user with little to no user intervention. This is all possible because of the rich data description and communication infrastructure MPI-2 offers. Powerful data descriptions and some of the other desirable features of MPI-2, however, make MPI-IO challenging to implement. By creating a new collective I/O implementation that allows developers to easily tinker and play with new optimizations or combinations of different techniques, research can proceed faster and be quickly and reliably deployed
Kenin Coloma, Avery Ching, Alok N. Choudhary, Wei-keng Liao, Robert B. Ross, Rajeev Thakur, Lee Ward
CLUSTER1
2006 Scalable Design and Implementations for MPI Parallel Overlapping I/O
abstract
We investigate the Message Passing Interface Input/Output (MPI I/O) implementation issues for two overlapping access patterns: the overlaps among processes within a single I/O operation and the overlaps across a sequence of I/O operations. The former case considers whether I/O atomicity can be obtained in the overlapping regions. The latter focuses on the file consistency problem on parallel machines with client-side file caching enabled. Traditional solutions for both overlapping I/O problems use whole file or byte-range file locking to ensure exclusive access to the overlapping regions and bypass the file system cache. Unfortunately, not only can file locking serialize I/O, but it can also increase the aggregate communication overhead between clients and I/O servers. For atomicity, we first differentiate MPI's requirements from the Portable Operating System Interface (POSIX) standard and propose two scalable approaches, graph coloring and process-rank ordering, which can resolve access conflicts and maintain I/O parallelism. For solving the file consistency problem across multiple I/O operations, we propose a method called Persistent File Domains, which tackles cache coherency with additional information and coordination to guarantee safe cache access without using file locks.
Wei-keng Liao, Kenin Coloma, Alok N. Choudhary, Lee Ward, Eric Russell, Neil Pundit
IEEE Trans. Parallel Distributed Syst.2
2005 Collective caching: application-aware client-side file caching
abstract
Parallel file subsystems in today's high-performance computers adopt many I/O optimization strategies that were designed for distributed systems. These strategies, for instance client-side file caching, treat each I/O request process independently, due to the consideration that clients are unlikely related with each other in a distributed environment. However, it is inadequate to apply such strategies directly in the high-performance computers where most of the I/O requests come from the processes that work on the same parallel applications. We believe that client-side could perform more effectively if the subsystem is aware of the process scope of an application and regards all the application processes as a single client. In this paper, we propose the idea of caching which coordinates the application processes to manage cache data and achieve cache coherence without involving the I/O servers. To demonstrate this idea, we implemented a collective subsystem at user space as a library, which can be incorporated into any message passing interface implementation to increase its portability. The performance evaluation is presented with three I/O benchmarks on an IBM SP using its native parallel file system, GPFS. Our results show significant performance enhancement obtained by collective over the traditional approaches.
Wei-keng Liao, Kenin Coloma, Alok N. Choudhary, Lee Ward, Eric Russell, Sonja Tideman
HPDC2
2004 Scalable High-level Caching for Parallel I/O
abstract
Summary form only given. In order for I/O systems to achieve high performance in a parallel environment, they must either sacrifice client-side file caching, or keep caching and deal with complex coherency issues. The most common technique for dealing with cache coherency in multiclient file caching environments uses file locks to bypass the client-side cache. Aside from effectively disabling cache usage, file locking is sometimes unavailable on larger systems. The high-level abstraction layer of MPI allows us to tackle cache coherency with additional information and coordination without using file locks. By approaching the cache coherency issue further up, the underlying I/O accesses can be modified in such a way as to ensure access to coherent data while satisfying the user's I/O request. We can effectively exploit the benefits of a file system's client-side cache while minimizing its management costs.
Kenin Coloma, Alok N. Choudhary, Wei-keng Liao, Lee Ward, Eric Russell, Neil Pundit
IPDPS1
2003 Noncontiguous I/O Accesses Through MPI-IO
abstract
I/O performance remains a weakness of parallel computing systems today. While this weakness is partly attributed to rapid advances in other system components, I/O interfaces available to programmers and the I/O methods supported by file systems have traditionally not matched efficiently with the types of I/O operations that scientific applications perform, particularly noncontiguous accesses. The MPI-IO interface allows for rich descriptions of the I/O patterns desired for scientific applications and implementations such as ROMIO have taken advantage of this ability while remaining limited by underlying file system methods. A method of noncontiguous data access, list I/O, was recently implemented in the Parallel Virtual File System (PVFS). We implement support for this interface in the ROMIO MPI-IO implementation. Through a suite of noncontiguous I/O tests we compared ROMIO list I/O to current methods of ROMIO noncontiguous access and found that the list I/O interface provides performance benefits in many noncontiguous cases.
Avery Ching, Alok N. Choudhary, Kenin Coloma, Wei-keng Liao, Robert B. Ross, William Gropp
CCGRID3
2003 Scalable Implementations of MPI Atomicity for Concurrent Overlapping I/O
abstract
For concurrent I/O operations, atomicity defines the results in the overlapping file regions simultaneously read/written by requesting processes. Atomicity has been well studied at the file system level, such as POSIX standard. We investigate the problems arising from the implementation of MPI atomicity for concurrent overlapping write access and provide two programming solutions. Since the MPI definition of atomicity differs from the POSIX one, an implementation that simply relies on the POSIX file systems does not guarantee correct MPI semantics. To have a correct implementation of atomic I/O in MPI, we examine the efficiency of three approaches: I) file locking, 2) graph-coloring, and 3) process-rank ordering. Performance complexity for these methods are analyzed and their experimental results are presented for file systems including NFS, SGI's XFS, and IBM's GPFS.
Wei-keng Liao, Alok N. Choudhary, Kenin Coloma, George K. Thiruvathukal, Lee Ward, Eric Russell, Neil Pundit
ICPP3