Arifa Nisar

dblp:33/3710 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Storage systems · 48% High-performance computing · 30% Memory systems · 13%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
parallel i/o
0.222012
Delegation-Based I/O Mechanism for High Performance Computing Systems · IEEE Trans. Parallel Distributed Syst. 2012
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Storage systems › file systems › distributed file system
parallel file system
0.232012
Scaling parallel I/O performance through I/O delegate and caching system · SC 2008
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Delegation-Based I/O Mechanism for High Performance Computing Systems · IEEE Trans. Parallel Distributed Syst. 2012
Storage systems › i/o optimization
parallel i/o optimization
0.112008
Scaling parallel I/O performance through I/O delegate and caching system · SC 2008
Memory systems › cache management › storage caching
client-side caching
0.112007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Storage systems › storage performance
write performance
0.112007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007
Parallel and multicore computing › synchronization
lock contention
0.012012
Delegation-Based I/O Mechanism for High Performance Computing Systems · IEEE Trans. Parallel Distributed Syst. 2012
Distributed systems › fault tolerance
checkpointing
0.012008
Scaling parallel I/O performance through I/O delegate and caching system · SC 2008
Memory systems
cache coherence
0.012007
Using MPI file caching to improve parallel write performance for large-scale scientific applications · SC 2007

Methods — techniques the papers use, named apart from their topics

static file domain partitioning · 0.1performance evaluation · 0.1collective i/o · 0.1caching · 0.1MPI-IO · 0.1write-behind · 0.1thread-based caching · 0.1
YearPublicationVenuePosition
2012 Delegation-Based I/O Mechanism for High Performance Computing Systems
abstract
Massively parallel applications often require periodic data checkpointing for program restart and post-run data analysis. Although high performance computing systems provide massive parallelism and computing power to fulfill the crucial requirements of the scientific applications, the I/O tasks of high-end applications do not scale. Strict data consistency semantics adopted from traditional file systems are inadequate for homogeneous parallel computing platforms. For high performance parallel applications independent I/O is critical, particularly if checkpointing data are dynamically created or irregularly partitioned. In particular, parallel programs generating a large number of unrelated I/O accesses on large-scale systems often face serious I/O serializations introduced by lock contention and conflicts at file system layer. As these applications may not be able to utilize the I/O optimizations requiring process synchronization, they pose a great challenge for parallel I/O architecture and software designs. We propose an I/O mechanism to bridge the gap between scientific applications and parallel storage systems. A static file domain partitioning method is developed to align the I/O requests and produce a client-server mapping that minimizes the file lock acquisition costs and eliminates the lock contention. Our performance evaluations of production application I/O kernels demonstrate scalable performance and achieve high I/O bandwidths.
Arifa Nisar, Wei-keng Liao, Alok N. Choudhary
IEEE Trans. Parallel Distributed Syst.1
2009 Using Subfiling to Improve Programming Flexibility and Performance of Parallel Shared-file I/O
abstract
There are two popular parallel I/O programming styles used by modern scientific computational applications: unique-file and shared-file. Unique-file I/O usually gives satisfactory performance, but its major drawback is that managing a large number of files can overwhelm the task of post-simulation data processing. Shared-file I/O produces fewer files and allows arrays partitioned among processes to be saved in the canonical order. As the number of processors on modern parallel machines increases into thousands and more, the problem size and in turn the global array size also increase proportionally. It is not practical to manage files of size each larger than a few hundreds of GB. Hence, to seek a middle ground between these two I/O styles, we propose a subfiling scheme that divides a large multi-dimensional global array into smaller subarrays, each saved in a smaller file, named subfile. Subfiling is implemented on top of MPI-IO. We also incorporate it into the parallel netCDF library in order to preserve the partitioning information in the netCDF file header, so that the global array can later be reconstructed. In addition, since the subfiling scheme decreases the number of processes sharing a file, it can reduce the overhead of file system's data consistency control. Our experimental results with several I/O benchmarks show that subfiling can provide improved I/O performance.
Kui Gao, Wei-keng Liao, Arifa Nisar, Alok N. Choudhary, Robert B. Ross, Robert Latham
ICPP3
2009 High Performance Parallel/Distributed Biclustering Using Barycenter Heuristic
abstract
Biclustering refers to simultaneous clustering of objects and their features. Use of biclustering is gaining momentum in areas such as text mining, gene expression analysis and collaborative filtering. Due to requirements for high performance in large scale data processing applications such as Collaborative filtering in E-commerce systems and large scale genome-wide gene expression analysis in microarray experiments, a high performance prallel/distributed solution for biclustering problem is highly desirable. Recently, Ahmad et al [1] showed that Bipartite Spectral Partitioning, which is a popular technique for biclustering, can be reformulated as a graph drawing problem where objective is to minimize Hall's energy of the bipartite graph representation of the input data. They showed that optimal solution to this problem is achieved when nodes are placed at the barycenter of their neighbors. In this paper, we provide a parallel algorithm for biclustering based on this formulation. We show that parallel energy minimization using barycenter heuristic is embarrassingly parallel. The challenge is to design a bi-cluster identification algorithm which is scalable as well as accurate. We show that our parallel implementation is not just extremely scalable, it is comparable in accuracy as well with serial implementation. We have evaluated proposed parallel biclustering algorithm with large synthetic data sets on upto 256 processors. Experimental evaluation shows large superlinear speedups, scalability and high level of accuracy.
Arifa Nisar, Waseem Ahmad, Wei-keng Liao, Alok N. Choudhary
SDM1
2008 Scaling parallel I/O performance through I/O delegate and caching system
abstract
Increasingly complex scientific applications require massive parallelism to achieve the goals of fidelity and high computational performance. Such applications periodically offload checkpointing data to file system for post-processing and program resumption. As a side effect of high degree of parallelism, I/O contention at servers doesn't allow overall performance to scale with increasing number of processors. To bridge the gap between parallel computational and I/O performance, we propose a portable MPI-IO layer where certain tasks, such as file caching, consistency control, and collective I/O optimization are delegated to a small set of compute nodes, collectively termed as I/O Delegate nodes. A collective cache design is incorporated to resolve cache coherence and hence alleviates the lock contention at I/O servers. By using popular parallel I/O benchmark and application I/O kernels, our experimental evaluation indicates considerable performance improvement with a small percentage of compute resources reserved for I/O.
Arifa Nisar, Wei-keng Liao, Alok N. Choudhary
SC1
2007 Using MPI file caching to improve parallel write performance for large-scale scientific applications
abstract
Typical large-scale scientific applications periodically write checkpoint files to save the computational state throughout execution. Existing parallel file systems improve such write-only I/O patterns through the use of client-side file caching and write-behind strategies. In distributed environments where files are rarely accessed by more than one client concurrently, file caching has achieved significant success; however, in parallel applications where multiple clients manipulate a shared file, cache coherence control can serialize I/O. We have designed a thread based caching layer for the MPI I/O library, which adds a portable caching system closer to user applications so more information about the application’s I/O patterns is available for better coherence control. We demonstrate the impact of our caching solution on parallel write performance with a comprehensive evaluation that includes a set of widely used I/O benchmarks and production application I/O kernels. 1.
Wei-keng Liao, Avery Ching, Kenin Coloma, Arifa Nisar, Alok N. Choudhary, Jacqueline Chen, Ramanan Sankaran, Scott Klasky
SC4