EDBT 2026 Demo / reviewers in the wild / expert
Sarp Oral
dblp:61/5178 · also Sarp H. Oral
· DBLP profile ↗
47ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0001-8745-7078ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 3 first-author · 10 since 2021Computer networks · 7 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5Software engineering, systems software and programming languages · 4Databases, data management, data science and information retrieval · 3 · 1 first-authorArtificial intelligence and machine learning · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling the memory wall using mixed-precision - HPG-MxP on an exascale machineabstractMixed-precision algorithms have been proposed as a way for scientific computing to benefit from some of the gains seen for artificial intelligence (AI) on recent high performance computing (HPC) platforms. A few applications dominated by dense matrix operations have seen substantial speedups by utilizing low precision formats such as FP16. However, a majority of scientific simulation applications are memory bandwidth limited. Beyond preliminary studies, the practical gain from using mixed-precision algorithms on a given HPC system is largely unclear. Aditya Kashi, Nicholson Koukpaizan, Hao Lu 0001, Michael A. Matheson, Sarp Oral, Feiyi Wang |
SC | 5 |
| 2025 | Lustre Unveiled: Evolution, Design, Advancements, and Current TrendsabstractThe Lustre filesystem serves as a vital element in high-performance parallel storage, meeting the rising demands of scientific, research, and enterprise environments. Widely deployed across HPC environments, ranging from small-scale applications in AI/ML, to domains like oil and gas, drug discovery, and meteorology, and manufacturing, Lustre addresses the universal challenge of efficiently accessing vast and ever-increasing volumes of data. Lustre is the filesystem of choice on six out of the top 10 fastest supercomputers in the world today, over 65% of the top 100, and also for over 60% of the top 500. Despite its widespread popularity, there is a lack of a complete and up-to-date reference, covering Lustre’s evolution, design, and various advancements made over the years. In this journal, we aim to fill this gap by providing a comprehensive journey of Lustre, including its history with significant contributions to HPC, detailed architecture and design elements, exploration of advancements added through its evolution, and future directions. Additionally, we present a comparison of Lustre with other prominent storage technologies of the era. To illustrate the current state of Lustre, we analyze several filesystem trends, including utilization, performance, and usage patterns on Orion, the Lustre filesystem on the first exascale supercomputer Frontier. We hope that this journal serves as a comprehensive educational reference for the current and future generations interested in HPC filesystem storage aspects. Anjus George, Andreas Dilger, Michael J. Brim, Rick Mohr, Amir Shehata, Jong Choi 0001, Ahmad Maroof Karimi, Jesse Hanley, James Simmons, Dominic Manno, Verónica G. Vergara Larrea, Sarp Oral, Christopher Zimmer 0001 |
ACM Trans. Storage | 12 |
| 2024 | Integrating quantum computing resources into scientific HPC ecosystems
Thomas L. Beck, Alessandro Baroni 0003, Ryan S. Bennink, Gilles Buchs, Eduardo Antonio Coello Pérez, Markus Eisenbach 0002, Rafael Ferreira da Silva, Muralikrishnan Gopalakrishnan Meena, Kalyana C. Gottiparthi, Peter Groszkowski, Travis S. Humble, Ryan Landfield, Ketan Maheshwari, Sarp Oral, Michael A. Sandoval, Amir Shehata, In-Saeng Suh, Christopher Zimmer 0001 |
Future Gener. Comput. Syst. | 14 |
| 2024 | Tarazu: An Adaptive End-to-end I/O Load-balancing Framework for Large-scale Parallel File SystemsabstractThe imbalanced I/O load on large parallel file systems affects the parallel I/O performance of high-performance computing (HPC) applications. One of the main reasons for I/O imbalances is the lack of a global view of system-wide resource consumption. While approaches to address the problem already exist, the diversity of HPC workloads combined with different file striping patterns prevents widespread adoption of these approaches. In addition, load-balancing techniques should be transparent to client applications. To address these issues, we propose Tarazu , an end-to-end control plane where clients transparently and adaptively write to a set of selected I/O servers to achieve balanced data placement. Our control plane leverages real-time load statistics for global data placement on distributed storage servers, while our design model employs trace-based optimization techniques to minimize latency for I/O load requests between clients and servers and to handle multiple striping patterns in files. We evaluate our proposed system on an experimental cluster for two common use cases: the synthetic I/O benchmark IOR and the scientific application I/O kernel HACC-I/O. We also use a discrete-time simulator with real HPC application traces from emerging workloads running on the Summit supercomputer to validate the effectiveness and scalability of Tarazu in large-scale storage environments. The results show improvements in load balancing and read performance of up to 33% and 43%, respectively, compared to the state-of-the-art. Arnab Kumar Paul, Sarah Neuwirth, Bharti Wadhwa, Feiyi Wang, Sarp Oral, Ali Raza Butt |
ACM Trans. Storage | 5 |
| 2023 | UnifyFS: A User-level Shared File System for Unified Access to Distributed Local StorageabstractWe introduce UnifyFS, a user-level file system that aggregates node-local storage tiers available on high performance computing (HPC) systems and makes them available to HPC applications under a unified namespace. UnifyFS employs transparent I/O interception, so it does not require changes to application code and is compatible with commonly used HPC I/O libraries. The design of UnifyFS supports the predominant HPC I/O workloads and is optimized for bulk-synchronous I/O patterns. Furthermore, UnifyFS provides customizable file system semantics to flexibly adapt its behavior for diverse I/O workloads and storage devices. In this paper, we discuss the unique design goals and architecture of UnifyFS and evaluate its performance on a leadership-class HPC system. In our experimental results, we demonstrate that UnifyFS exhibits excellent scaling performance for write operations and can improve the performance of application checkpoint operations by as much as 3× versus a tuned configuration. Michael J. Brim, Adam Moody, Seung-Hwan Lim, Ross G. Miller, Swen Böhm, Cameron Stanavige, Kathryn Mohror, Sarp Oral |
IPDPS | 8 |
| 2023 | Frontier: Exploring ExascaleabstractAs the US Department of Energy (DOE) computing facilities began deploying petascale systems in 2008, DOE was already setting its sights on exascale. In that year, DARPA published a report on the feasibility of reaching exascale. The report authors identified several key challenges in the pursuit of exascale including power, memory, concurrency, and resiliency. That report informed the DOE's computing strategy for reaching exascale. With the deployment of Oak Ridge National Laboratory's Frontier supercomputer, we have officially entered the exascale era. In this paper, we discuss Frontier's architecture, how it addresses those challenges, and describe some early application results from Oak Ridge Leadership Computing Facility's Center of Excellence and the Exascale Computing Project. Scott Atchley, Christopher Zimmer 0001, Jack Lange, David E. Bernholdt, Verónica G. Vergara Larrea, Michael J. Brim, Reuben D. Budiardja, Sunita Chandrasekaran, Markus Eisenbach 0002, Thomas M. Evans 0001, Matthew Ezell, Nicholas Frontiere, Antigoni Georgiadou, Joseph Glenski, Philipp Grete, Steven P. Hamilton, John K. Holmen, Axel Huebl, Daniel A. Jacobson, Wayne Joubert, Kim H. McMahon, Elia Merzari, Stan G. Moore, Andrew Myers 0001, Stephen Nichols, Sarp Oral, Thomas Papatheodore, Danny Perez, David M. Rogers 0001, Evan Schneider, Jean-Luc Vay, P. K. Yeung |
SC | 27 |
| 2022 | Hvac: Removing I/O Bottleneck for Large-Scale Deep Learning ApplicationsabstractScientific communities are increasingly adopting deep learning (DL) models in their applications to accelerate scientific discovery processes. However, with rapid growth in the computing capabilities of HPC supercomputers, large-scale DL applications have to spend a significant portion of training time performing I/O to a parallel storage system. Previous research works have investigated optimization techniques such as prefetching and caching. Unfortunately, there exist non-trivial challenges to adopting the existing solutions on HPC supercomputers for large-scale DL training applications, which include non-performance and/or failures at extreme scale, lack of portability and generality in design, complex deployment methodology, and being limited to a specific application or dataset. To address these challenges, we propose High-Velocity AI Cache (HVAC), a distributed read-cache layer that targets and fully exploits the node-local storage or near node-local storage technology. HVAC seamlessly accelerates read I/O by aggregating node-local or near node-local storage, avoiding metadata lookups and file locking while preserving portability in the application code. We deploy and evaluate HVAC on 1,024 nodes (with over 6000 NVIDIA V100 GPUS) of the Summit supercomputer. In particular, we evaluate the scalability, efficiency, accuracy, and load distribution of HVAC compared to GPFS and XFS-on-NVMe. With four different DL applications, we observe an average 25 % performance improvement atop GPFS and 9% drop against XFS-on-NVMe, which scale linearly and are considered the performance upper bound. We envision HVAC as an important caching library for upcoming HPC supercomputers such as Frontier. Awais Khan 0002, Arnab Kumar Paul, Christopher Zimmer 0001, Sarp Oral, Sajal Dash, Scott Atchley, Feiyi Wang |
CLUSTER | 4 |
| 2022 | Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production LoadabstractScientific computing workloads at HPC facilities have been shifting from traditional numerical simulations to AI/ML applications for training and inference while processing and producing ever-increasing amounts of scientific data. To address the growing need for increased storage capacity, lower access latency, and higher bandwidth, emerging technologies such as non-volatile memory are integrated into supercomputer I/O subsystems. With these emerging trends, we need a better understanding of the multilayer supercomputer I/O systems and ways to use these subsystems efficiently. In this work, we study the I/O access patterns and performance characteristics of two representative supercomputer I/O subsystems. Through an extensive analysis of year-long I/O logs on each system, we report new observations in I/O reads and writes, unbalanced use of storage system layers, and new trends in user behaviors at the HPC I/O middleware stack. Jean Luca Bez, Ahmad Maroof Karimi, Arnab Kumar Paul, Surendra Byna, Philip H. Carns, Sarp Oral, Feiyi Wang, Jesse Hanley |
HPDC | 7 |
| 2021 | Battle of the Defaults: Extracting Performance Characteristics of HDF5 under Production LoadabstractPopular parallel I/O libraries, such as HDF5, provide tuning parameters to obtain superior performance. However, the selection of effective parameters on production systems is complex due to the interdependence of I/O software and file system layers. Hence, application developers typically use the default parameters and often experience poor I/O performance. This work conducts a benchmarking-based analysis on the HDF5 behaviors with a wide variety of I/O patterns to extract performance characteristics under the production workload. To make the analysis well controlled, we exercise I/O benchmarks on POSIX-IO, MPI-IO, and HDF5 using the same I/O patterns and in the same jobs. To address high performance variability in production environments, we repeat the benchmarks across I/O patterns, storage devices, and time intervals. Based on the results, we identified consistent HDF5 behaviors that appropriate configurations and operations on dataset layout and file-metadata placement can improve performance significantly. We apply our findings and evaluate the tuned I/O library on two supercomputers: Summit and Cori. The results show that our tuned parameters can achieve more than 10× I/O performance speedup than that with default parameters on both systems, suggesting the effectiveness, stability, and generality of our solution. Houjun Tang, Surendra Byna, Jesse Hanley, Quincey Koziol, Tonglin Li, Sarp Oral |
CCGRID | 7 |
| 2021 | Interpreting Write Performance of Supercomputer I/O Systems with Regression ModelsabstractThis work seeks to advance the state of the art in HPC I/O performance analysis and interpretation. In particular, we demonstrate effective techniques to: (1) model output performance in the presence of I/O interference from production loads; (2) build features from write patterns and key parameters of the system architecture and configurations; (3) employ suitable machine learning algorithms to improve model accuracy. We train models with five popular regression algorithms and conduct experiments on two distinct production HPC platforms. We find that the lasso and random forest models predict output performance with high accuracy on both of the target systems. We also explore use of the models to guide adaptation in I/O middleware systems, and show potential for improvements of at least 15% from model-guided adaptation on 70% of samples, and improvements up to 10 x on some samples for both of the target systems. Zilong Tan, Philip H. Carns, Jeffrey S. Chase, Kevin Harms, Jay F. Lofstead, Sarp Oral, Sudharshan S. Vazhkudai, Feiyi Wang |
IPDPS | 7 |
| 2020 | Characterizing Output Bottlenecks of a Production Supercomputer: Analysis and ImplicationsabstractThis article studies the I/O write behaviors of the Titan supercomputer and its Lustre parallel file stores under production load. The results can inform the design, deployment, and configuration of file systems along with the design of I/O software in the application, operating system, and adaptive I/O libraries. We propose a statistical benchmarking methodology to measure write performance across I/O configurations, hardware settings, and system conditions. Moreover, we introduce two relative measures to quantify the write-performance behaviors of hardware components under production load. In addition to designing experiments and benchmarking on Titan, we verify the experimental results on one real application and one real application I/O kernel, XGC and HACC IO, respectively. These two are representative and widely used to address the typical I/O behaviors of applications. In summary, we find that Titan’s I/O system is variable across the machine at fine time scales. This variability has two major implications. First, stragglers lessen the benefit of coupled I/O parallelism (striping). Peak median output bandwidths are obtained with parallel writes to many independent files, with no striping or write sharing of files across clients (compute nodes). I/O parallelism is most effective when the application—or its I/O libraries—distributes the I/O load so that each target stores files for multiple clients and each client writes files on multiple targets in a balanced way with minimal contention. Second, our results suggest that the potential benefit of dynamic adaptation is limited. In particular, it is not fruitful to attempt to identify “good locations” in the machine or in the file system: component performance is driven by transient load conditions and past performance is not a useful predictor of future performance. For example, we do not observe diurnal load patterns that are predictable. Sarp Oral, Christopher Zimmer 0001, Jong Choi 0001, David Dillow, Scott Klasky, Jay F. Lofstead, Norbert Podhorszki, Jeffrey S. Chase |
ACM Trans. Storage | 2 |
| 2019 | Learning from Five-year Resource-Utilization Data of Titan SystemabstractTitan was the flagship supercomputer at the Oak Ridge Leadership Computing Facility (OLCF). It was deployed in late 2012, became the fastest supercomputer in the world and was retired on August 2, 2019. With Titan's mission complete, this paper provides a first-order examination of the usage of its critical resources (CPU, Memory, GPU, and I/O) over a five-year production period (2015-2019). In particular, we show quantitatively that the majority of CPU time was spent on the large-scale jobs, which is consistent with the policy of driving ground-breaking science through leadership computing. We also corroborate the general observation of the low CPU-memory usage with 95% jobs utilizing only 15% or less available memory. Additionally, we correlate the increase of total job submissions and the decrease of GPU-enabled jobs during 2016 with the GPU reliability issue which impacted the large-scale runs. We further show the surprising read/write ratio over the five-year period, which contradicts the general mindset of the large-scale simulation machines being “write-heavy”. This understanding will have potential impact on how we design our next-generation large-scale storage systems. We believe that our analyses and findings are going to be of great interest to the high-performance computing (HPC) community at large. Feiyi Wang, Sarp Oral, Satyabrata Sen, Neena Imam |
CLUSTER | 2 |
| 2019 | Data Jockey: Automatic Data Management for HPC Multi-tiered Storage SystemsabstractWe present the design and implementation of Data Jockey, a data management system for HPC multi-tiered storage systems. As a centralized data management control plane, Data Jockey automates bulk data movement and placement for scientific workflows and integrates into existing HPC storage infrastructures. Data Jockey simplifies data management by eliminating human effort in programming complex data movements, laying datasets across multiple storage tiers when supporting complex workflows, which in turn increases the usability of multitiered storage systems emerging in modern HPC data centers. Specifically, Data Jockey presents a new data management scheme called “goal driven data management” that can automatically infer low-level bulk data movement plans from declarative high-level goal statements that come from the lifetime of iterative runs of scientific workflows. While doing so, Data Jockey aims to minimize data wait times by taking responsibility for datasets that are unused or to be used, and aggressively utilizing the capacity of the upper, higher performant storage tiers. We evaluated a prototype implementation of Data Jockey under a synthetic workload based on a year's worth of Oak Ridge Leadership Computing Facility's (OLCF) operational logs. Our evaluations suggest that Data Jockey leads to higher utilization of the upper storage tiers while minimizing the programming effort of data movement compared to human involved, per-domain adhoc data management scripts. Woong Shin, Christopher Brumgard, Sudharshan S. Vazhkudai, Devarshi Ghoshal, Sarp Oral, Lavanya Ramakrishnan |
IPDPS | 6 |
| 2019 | iez: Resource Contention Aware Load Balancing for Large-Scale Parallel File SystemsabstractParallel I/O performance is crucial to sustaining scientific applications on large-scale High-Performance Computing (HPC) systems. However, I/O load imbalance in the underlying distributed and shared storage systems can significantly reduce overall application performance. There are two conflicting challenges to mitigate this load imbalance: (i) optimizing systemwide data placement to maximize the bandwidth advantages of distributed storage servers, i.e., allocating I/O resources efficiently across applications and job runs; and (ii) optimizing client-centric data movement to minimize I/O load request latency between clients and servers, i.e., allocating I/O resources efficiently in service to a single application and job run. Moreover, existing approaches that require application changes limit wide-spread adoption in commercial or proprietary deployments. We propose iez, an “end-to-end control plane” where clients transparently and adaptively write to a set of selected I/O servers to achieve balanced data placement. Our control plane leverages realtime load information for distributed storage server global data placement while our design model leverages trace-based optimization techniques to minimize I/O load request latency between clients and servers. We evaluate our proposed system on an experimental cluster for two common use cases: synthetic I/O benchmark IOR for large sequential writes and a scientific application I/O kernel, HACC-I/O. Results show read and write performance improvements of up to 34% and 32%, respectively, compared to the state of the art. Bharti Wadhwa, Arnab Kumar Paul, Sarah Neuwirth, Feiyi Wang, Sarp Oral, Ali Raza Butt, Jon Bernard, Kirk W. Cameron |
IPDPS | 5 |
| 2019 | End-to-end I/O portfolio for the summit supercomputing ecosystemabstractThe I/O subsystem for the Summit supercomputer, No. 1 on the Top500 list, and its ecosystem of analysis platforms is composed of two distinct layers, namely the in-system layer and the center-wide parallel file system layer (PFS), Spider 3. The in-system layer uses node-local SSDs and provides 26.7 TB/s for reads, 9.7 TB/s for writes, and 4.6 billion IOPS to Summit. The Spider 3 PFS layer uses IBM's Spectrum Scale™ and provides 2.5 TB/s and 2.6 million IOPS to Summit and other systems. While deploying them as two distinct layers was operationally efficient, it also presented usability challenges in terms of multiple mount points and lack of transparency in data movement. To address these challenges, we have developed novel end-to-end I/O solutions for the concerted use of the two storage layers. We present the I/O subsystem architecture, the end-to-end I/O solution space, their design considerations and our deployment experience. Sarp Oral, Sudharshan S. Vazhkudai, Feiyi Wang, Christopher Zimmer 0001, Christopher Brumgard, Jesse Hanley, George Markomanolis, Ross G. Miller, Dustin Leverman, Scott Atchley, Verónica G. Vergara Larrea |
SC | 1 |
| 2019 | Are we witnessing the spectre of an HPC meltdown?abstractSummary We measure and analyze the performance observed when running applications and benchmarks before and after the Meltdown and Spectre fixes have been applied to the Cray supercomputers and supporting systems at the Oak Ridge Leadership Computing Facility (OLCF). Of particular interest is the effect of these fixes on applications selected from the OLCF portfolio when running at scale. This comprehensive study presents results from experiments run on Titan, Eos, Cumulus, and Percival supercomputers at the OLCF. The results from this study are useful for HPC users running on Cray supercomputers and serve to better understand the impact that these two vulnerabilities have on diverse HPC workloads at scale. Verónica G. Vergara Larrea, Michael J. Brim, Wayne Joubert, Swen Böhm, Matthew B. Baker, Oscar R. Hernandez, Sarp Oral, James Simmons, Don E. Maxwell |
Concurr. Comput. Pract. Exp. | 7 |
| 2018 | Towards Exascale Computing for High Energy Physics: The ATLAS Experience at ORNLabstractTraditionally, the ATLAS experiment at Large Hadron Collider (LHC) has utilized distributed resources as provided by the Worldwide LHC Computing Grid (WLCG) to support data distribution, data analysis and simulations. For example, the ATLAS experiment uses a geographically distributed grid of approximately 200,000 cores continuously (250 000 cores at peak), (over 1,000 million core-hours per year) to process, simulate, and analyze its data (todays total data volume of ATLAS is more than 300 PB). After the early success in discovering a new particle consistent with the long-awaited Higgs boson, ATLAS is continuing the precision measurements necessary for further discoveries. Planned high-luminosity LHC upgrade and related ATLAS detector upgrades, that are necessary for physics searches beyond Standard Model, pose serious challenge for ATLAS computing. Data volumes are expected to increase at higher energy and luminosity, causing the storage and computing needs to grow at a much higher pace than the flat budget technology evolution (see Fig. 1). The need for simulation and analysis will overwhelm the expected capacity of WLCG computing facilities unless the range and precision of physics studies will be curtailed. V. Ananthraj, Kaushik De, Shantenu Jha, Alexei Klimentov, Danila Oleynik, Sarp Oral, André Merzky, Ruslan Mashinistov, Sergey Panitkin, P. Svirin, Matteo Turilli, Jack C. Wells, Sean R. Wilkinson |
eScience | 6 |
| 2018 | Modeling Impact of Execution Strategies on Resource UtilizationabstractThe analysis of the hundreds of petabytes of raw and derived HEP (High Energy Physics) data will necessitate exascale computing. In addition to unprecedented volume, these data are distributed over hundreds of computing centers. In response to these application requirement, as well as performance requirement by using parallel processing (i.e., parallelism), and as a consequence of technology trends, there has been an increase in the uptake of supercomputers by HEP projects. Alexey A. Poyda, Mikhail Titov, Alexei Klimentov, Jack C. Wells, Sarp Oral, Kaushik De, Danila Oleynik, Shantenu Jha |
eScience | 5 |
| 2018 | The design, deployment, and evaluation of the CORAL pre-exascale systems
Sudharshan S. Vazhkudai, Bronis R. de Supinski, Arthur S. Bland, Al Geist, James C. Sexton, James A. Kahle, Christopher Zimmer 0001, Scott Atchley, Sarp Oral, Don E. Maxwell, Verónica G. Vergara Larrea, Adam Bertsch, Robin Goldstone, Wayne Joubert, Christopher M. Chambreau, David Appelhans, Robert Blackmore, Ben Casses, George Chochia, Gene Davison, Matthew Ezell, Thomas Gooding, Elsa Gonsiorowski, Leopold Grinberg, Bill Hanson, Bill Hartner, Ian Karlin, Matthew L. Leininger, Dustin Leverman, Chris Marroquin, Adam Moody, Martin Ohmacht, Ramesh Pankajakshan, Fernando Pizzano, James H. Rogers, Bryan S. Rosenburg, Drew Schmidt, Mallikarjun Shankar, Feiyi Wang, Py Watson, Bob Walkup, Lance D. Weems, Junqi Yin |
SC | 9 |
| 2017 | I/O load balancing for big data HPC applicationsabstractHigh Performance Computing (HPC) big data problems require efficient distributed storage systems. However, at scale, such storage systems often experience load imbalance and resource contention due to two factors: the bursty nature of scientific application I/O; and the complex I/O path that is without centralized arbitration and control. For example, the extant Lustre parallel file system-that supports many HPC centers-comprises numerous components connected via custom network topologies, and serves varying demands of a large number of users and applications. Consequently, some storage servers can be more loaded than others, which creates bottlenecks and reduces overall application I/O performance. Existing solutions typically focus on per application load balancing, and thus are not as effective given their lack of a global view of the system. In this paper, we propose a data-driven approach to load balance the I/O servers at scale, targeted at Lustre deployments. To this end, we design a global mapper on Lustre Metadata Server, which gathers runtime statistics from key storage components on the I/O path, and applies Markov chain modeling and a minimum-cost maximum-flow algorithm to decide where data should be placed. Evaluation using a realistic system simulator and a real setup shows that our approach yields better load balancing, which in turn can improve end-to-end performance. Arnab Kumar Paul, Arpit Goyal, Feiyi Wang, Sarp Oral, Ali Raza Butt, Michael J. Brim, Sangeetha B. Srinivasa |
IEEE BigData | 4 |
| 2017 | High-Throughput Computing on High-Performance Platforms: A Case StudyabstractThe computing systems used by LHC experiments has historically consisted of the federation of hundreds to thousands of distributed resources, ranging from small to mid-size re-source. In spite of the impressive scale of the existing distributed computing solutions, the federation of small to mid-size resources will be insufficient to meet projected future demands. This paper is a case study of how the ATLAS experiment has embraced Titan - a DOE leadership facility in conjunction with traditional distributed high-throughput computing to reach sustained production scales of approximately 52M core-hours a years. The three main contributions of this paper are: (i) a critical evaluation of design and operational considerations to support the sustained, scalable and production usage of Titan; (ii) a preliminary characterization of a next generation executor for PanDA to support new workloads and advanced execution modes; and (iii) early lessons for how current and future experimental and observational systems can be integrated with production supercomputers and other platforms in a general and extensible manner. Danila Oleynik, Sergey Panitkin, Matteo Turilli, Alessio Angius, Sarp Oral, Kaushik De, Alexei Klimentov, Jack C. Wells, Shantenu Jha |
eScience | 5 |
| 2017 | Predicting Output Performance of a Petascale SupercomputerabstractIn this paper, we develop a predictive model useful for output performance prediction of supercomputer file systems under production load. Our target environment is Titan---the 3rd fastest supercomputer in the world---and its Lustre-based multi-stage write path. We observe from Titan that although output performance is highly variable at small time scales, the mean performance is stable and consistent over typical application run times. Moreover, we find that output performance is non-linearly related to its correlated parameters due to interference and saturation on individual stages on the path. These observations enable us to build a predictive model of expected write times of output patterns and I/O configurations, using feature transformations to capture non-linear relationships. We identify the candidate features based on the structure of the Lustre/Titan write path, and use feature transformation functions to produce a model space with 135,000 candidate models. By searching for the minimal mean square error in this space we identify a good model and show that it is effective. Yezhou Huang, Jeffrey S. Chase, Jong Choi 0001, Scott Klasky, Jay F. Lofstead, Sarp Oral |
HPDC | 7 |
| 2017 | Automatic and Transparent Resource Contention Mitigation for Improving Large-Scale Parallel File System PerformanceabstractProportional to the scale increases in HPC systems, many scientific applications are becoming increasingly data intensive, and parallel I/O has become one of the dominant factors impacting the large-scale HPC application performance. On a typical large-scale HPC system, we have observed that the lack of a global workload coordination coupled with the shared nature of storage systems cause load imbalance and resource contention over the end-to-end I/O paths resulting in severe performance degradation. I/O load imbalance on HPC systems is generally a self-inflicted wound and mostly occurs between the I/O paths and resources consumed by each individual job. In this paper, we introduce TAPP-IO, a dynamic, shared load balancing framework for mitigating resource contention. TAPP-IO extends our previous work and solves two major limitations: First, it transparently intercepts file creation calls during runtime to balance the workload over all available storage targets. The usage of TAPP-IO requires no application source code modifications and is independent from any I/O middleware. The framework can be applied to almost any HPC platform and is suitable for systems that lack a centralized file system resource manager. Second, the framework proposes a new placement strategy to support not only file-per-process I/O, but also single shared file I/O. This opens the door to a new class of scientific applications that can leverage the placement library for improved performance. We demonstrate the effectiveness of our integration on the Titan system at the Oak Ridge National Laboratory. Our experiments with a synthetic benchmark and real-world HPC workload show that, even in a noisy production environment, TAPP-IO can improve large-scale application performance significantly. Sarah Neuwirth, Feiyi Wang, Sarp Oral, Ulrich Brüning 0001 |
ICPADS | 3 |
| 2017 | GUIDE: a scalable information directory service to collect, federate, and analyze logs for operational insights into a leadership HPC facilityabstractIn this paper, we describe the GUIDE framework used to collect, federate, and analyze log data from the Oak Ridge Leadership Computing Facility (OLCF), and how we use that data to derive insights into facility operations. We collect system logs and extract monitoring data at every level of the various OLCF subsystems, and have developed a suite of pre-processing tools to make the raw data consumable. The cleansed logs are then ingested and federated into a central, scalable data warehouse, Splunk, that offers storage, indexing, querying, and visualization capabilities. We have further developed and deployed a set of tools to analyze these multiple disparate log streams in concert and derive operational insights. We describe our experience from developing and deploying the GUIDE infrastructure, and deriving valuable insights on the various subsystems, based on two years of operations in the production OLCF environment. Sudharshan S. Vazhkudai, Ross G. Miller, Devesh Tiwari, Christopher Zimmer 0001, Feiyi Wang, Sarp Oral, Raghul Gunasekaran, Deryl Steinert |
SC | 6 |
| 2017 | Optimizing checkpoint data placement with guaranteed burst buffer endurance in large-scale hierarchical storage systems
Lipeng Wan 0001, Qing Cao 0001, Feiyi Wang, Sarp Oral |
J. Parallel Distributed Comput. | 4 |
| 2016 | Using Balanced Data Placement to Address I/O Contention in Production EnvironmentsabstractDesigned for capacity and capability, HPC I/O systems are inherently complex and shared among multiple, concurrent jobs competing for resources. Lack of centralized coordination and control often render the end-to-end I/O paths vulnerable to load imbalance and contention. With the emergence of data-intensive HPC applications, storage systems are further contended for performance and scalability. This paper proposes to unify two key approaches to tackle the imbalanced use of I/O resources and to achieve an end-to-end I/O performance improvement in the most transparent way. First, it utilizes a topology-aware, Balanced Placement I/O method (BPIO) for mitigating resource contention. Second, it takes advantage of the platform-neutral ADIOS middleware, which provides a flexible I/O mechanism for scientific applications. By integrating BPIO with ADIOS, referred to as Aequilibro, we obtain an end-to-end and per job I/O performance improvement for ADIOS-enabled HPC applications without requiring any code changes. Aequilibro can be applied to almost any HPC platform and is mostly suitable for systems that lack a centralized file system resource manager. We demonstrate the effectiveness of our integration on the Titan system at the Oak Ridge National Laboratory. Our experiments with a synthetic benchmark and real-world HPC workload show that, even in a noisy production environment, Aequilibro can improve large-scale application performance significantly. Sarah Neuwirth, Feiyi Wang, Sarp Oral, Sudharshan S. Vazhkudai, James H. Rogers, Ulrich Brüning 0001 |
SBAC-PAD | 3 |
| 2015 | TRIO: Burst Buffer Based I/O OrchestrationabstractThe growing computing power on leadership HPC systems is often accompanied by ever-escalating failure rates. Checkpointing is a common defensive mechanism used by scientific applications for failure recovery. However, directly writing the large and bursty checkpointing dataset to parallel file systems can incur significant I/O contention on storage servers. Such contention in turn degrades bandwidth utilization of storage servers and prolongs the average job I/O time of concurrent applications. Recently burst buffers have been proposed as an intermediate layer to absorb the bursty I/O traffic from compute nodes to storage backend. But an I/O orchestration mechanism is still desirable to efficiently move checkpointing data from burst buffers to storage backend. In this paper, we propose a burst buffer based I/O orchestration framework, named TRIO, to intercept and reshape the bursty writes for better sequential write traffic to storage servers. Meanwhile, TRIO coordinates the flushing orders among concurrent burst buffers to alleviate the contention on storage server. Our experimental results demonstrated that TRIO could efficiently utilize storage bandwidth and reduce the average job I/O time by 37% on average for data-intensive applications in typical checkpointing scenarios. Teng Wang 0001, Sarp Oral, Michael Pritchard, Bin Wang 0019, Weikuan Yu |
CLUSTER | 2 |
| 2015 | A practical approach to reconciling availability, performance, and capacity in provisioning extreme-scale storage systemsabstractThe increasing data demands from high-performance computing applications significantly accelerate the capacity, capability and reliability requirements of storage systems. As systems scale, component failures and repair times increase, significantly impacting data availability. A wide array of decision points must be balanced in designing such systems. Lipeng Wan 0001, Feiyi Wang, Sarp Oral, Devesh Tiwari, Sudharshan S. Vazhkudai, Qing Cao 0001 |
SC | 3 |
| 2014 | BurstMem: A high-performance burst buffer system for scientific applicationsabstractThe growth of computing power on large-scale systems requires commensurate high-bandwidth I/O systems. Many parallel file systems are designed to provide fast sustainable I/O in response to applications' soaring requirements. To meet this need, a novel system is imperative to temporarily buffer the bursty I/O and gradually flush datasets to long-term parallel file systems. In this paper, we introduce the design of BurstMem, a high-performance burst buffer system. BurstMem provides a storage framework with efficient storage and communication management strategies. Our experiments demonstrate that BurstMem is able to speed up the I/O performance of scientific applications by up to 8.5× on leadership computer systems. Teng Wang 0001, Sarp Oral, Yandong Wang 0001, Bradley W. Settlemyer, Scott Atchley, Weikuan Yu |
IEEE BigData | 2 |
| 2014 | Improving large-scale storage system performance via topology-aware and balanced data placementabstractWith the advent of big data, the I/O subsystems of large-scale compute clusters are becoming a center of focus. More applications are putting greater demands on end-to-end I/O performance. These subsystems are often complex in design. They comprise of multiple hardware and software layers to cope with the increasing capacity, capability, and scalability requirements of data intensive applications. However, the sharing nature of storage resources and the intrinsic interactions across these layers make it a great challenge to realize end-to-end performance gains. This paper proposes a topology-aware strategy to balance the load across resources, to improve the per-application I/O performance. We demonstrate the effectiveness of our algorithm on an extreme-scale compute cluster, Titan, at the Oak Ridge Leadership Computing Facility (OLCF). Our experiments with both synthetic benchmarks and a real-world application show that, even under congestion, our proposed algorithm can improve large-scale application I/O performance significantly, resulting in both a reduction in application run time as well as a higher resolution of simulation run. Feiyi Wang, Sarp Oral, Saurabh Gupta 0002, Devesh Tiwari, Sudharshan S. Vazhkudai |
ICPADS | 2 |
| 2014 | SSD-optimized workload placement with adaptive learning and classification in HPC environmentsabstractIn recent years, non-volatile memory devices such as SSD drives have emerged as a viable storage solution due to their increasing capacity and decreasing cost. Due to the unique capability and capacity requirements in large scale HPC (High Performance Computing) storage environment, a hybrid configuration (SSD and HDD) may represent one of the most available and balanced solutions considering the cost and performance. Under this setting, effective data placement as well as movement with controlled overhead become a pressing challenge. In this paper, we propose an integrated object placement and movement framework and adaptive learning algorithms to address these issues. Specifically, we present a method that shuffle data objects across storage tiers to optimize the data access performance. The method also integrates an adaptive learning algorithm where realtime classification is employed to predict the popularity of data object accesses, so that they can be placed on, or migrate between SSD or HDD drives in the most efficient manner. We discuss preliminary results based on this approach using a simulator we developed to show that the proposed methods can dynamically adapt storage placements and access pattern as workloads evolve to achieve the best system level performance such as throughput. Lipeng Wan 0001, Zheng Lu 0005, Qing Cao 0001, Feiyi Wang, Sarp Oral, Bradley W. Settlemyer |
MSST | 5 |
| 2014 | Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File SystemsabstractThe Oak Ridge Leadership Computing Facility (OLCF) has deployed multiple large-scale parallel file systems (PFS) to support its operations. During this process, OLCF acquired significant expertise in large-scale storage system design, file system software development, technology evaluation, benchmarking, procurement, deployment, and operational practices. Based on the lessons learned from each new PFS deployment, OLCF improved its operating procedures, and strategies. This paper provides an account of our experience and lessons learned in acquiring, deploying, and operating large-scale parallel file systems. We believe that these lessons will be useful to the wider HPC community. Sarp Oral, James Simmons, Jason Hill, Dustin Leverman, Feiyi Wang, Matthew Ezell, Ross G. Miller, Douglas Fuller, Raghul Gunasekaran, Youngjae Kim 0001, Saurabh Gupta 0002, Devesh Tiwari, Sudharshan S. Vazhkudai, James H. Rogers, David Dillow, Galen M. Shipman, Arthur S. Bland |
SC | 1 |
| 2014 | Coordinating Garbage Collectionfor Arrays of Solid-State DrivesabstractAlthough solid-state drives (SSDs) offer significant performance improvements over hard disk drives (HDDs) for a number of workloads, they can exhibit substantial variance in request latency and throughput as a result of garbage collection (GC). When GC conflicts with an I/O stream, the stream can make no forward progress until the GC cycle completes. GC cycles are scheduled by logic internal to the SSD based on several factors such as the pattern, frequency, and volume of write requests. When SSDs are used in a RAID with currently available technology, the lack of coordination of the SSD-local GC cycles amplifies this performance variance. We propose a global garbage collection (GGC) mechanism to improve response times and reduce performance variability for a RAID of SSDs. We include a high-level design of SSD-aware RAID controller and GGC-capable SSD devices and algorithms to coordinate the GGC cycles. We develop reactive and proactive GC coordination algorithms and evaluate their I/O performance and block erase counts for various workloads. Our simulations show that GC coordination by a reactive scheme improves average response time and reduces performance variability for a wide variety of enterprise workloads. For bursty, write-dominated workloads, response time was improved by 69 percent and performance variability was reduced by 71 percent. We show that a proactive GC coordination algorithm can further improve the I/O response times by up to 9 percent and the performance variability by up to 15 percent. We also observe that it could increase the lifetimes of SSDs with some workloads (e.g., Financial) by reducing the number of block erase counts by up to 79 percent relative to a reactive algorithm for write-dominant enterprise workloads. Youngjae Kim 0001, Junghee Lee 0004, Sarp Oral, David Dillow, Feiyi Wang, Galen M. Shipman |
IEEE Trans. Computers | 3 |
| 2013 | Preemptible I/O Scheduling of Garbage Collection for Solid State DrivesabstractUnlike hard disks, flash devices use out-of-place updates operations and require a garbage collection (GC) process to reclaim invalid pages to create free blocks. This GC process is a major cause of performance degradation when running concurrently with other I/O operations as internal bandwidth is consumed to reclaim these invalid pages. The invocation of the GC process is generally governed by a low watermark on free blocks and other internal device metrics that different workloads meet at different intervals. This results in an I/O performance that is highly dependent on workload characteristics. In this paper, we examine the GC process and propose a semipreemptible GC (PGC) scheme that allows GC processing to be preempted while pending I/O requests in the queue are serviced. Moreover, we further enhance flash performance by pipelining internal GC operations and merge them with pending I/O requests whenever possible. Our experimental evaluation of this semi-PGC scheme with realistic workloads demonstrates both improved performance and reduced performance variability. Write-dominant workloads show up to a 66.56% improvement in average response time with a 83.30% reduced variance in response time compared to the non-PGC scheme. In addition, we explore opportunities of a new NAND flash device that supports suspend/resume commands for read, write, and erase operations for fully PGC (F-PGC). Our experiments with an F-PGC enabled flash device show that request response time can be improved by up to 14.57% compared to semi-PGC. Junghee Lee 0004, Youngjae Kim 0001, Galen M. Shipman, Sarp Oral, Jongman Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2012 | Characterizing output bottlenecks in a supercomputerabstractSupercomputer I/O loads are often dominated by writes. HPC (High Performance Computing) file systems are designed to absorb these bursty outputs at high bandwidth through massive parallelism. However, the delivered write bandwidth often falls well below the peak. This paper characterizes the data absorption behavior of a center-wide shared Lustre parallel file system on the Jaguar supercomputer. We use a statistical methodology to address the challenges of accurately measuring a shared machine under production load and to obtain the distribution of bandwidth across samples of compute nodes, storage targets, and time intervals. We observe and quantify limitations from competing traffic, contention on storage servers and I/O routers, concurrency limitations in the client compute node operating systems, and the impact of variance (stragglers) on coupled output such as striping. We then examine the implications of our results for application performance and the design of I/O middleware systems on shared supercomputers. Jeffrey S. Chase, David Dillow, Oleg Drokin, Scott Klasky, Sarp Oral, Norbert Podhorszki |
SC | 6 |
| 2011 | Enhancing I/O throughput via efficient routing and placement for large-scale parallel file systemsabstractAs storage systems get larger to meet the demands of petascale systems, careful planning must be applied to avoid congestion points and extract the maximum performance. In addition, the large data sets generated by such systems makes it desirable for all compute resources to have common access to this data without needing to copy it to each machine. This paper describes a method of placing I/O close to the storage nodes to minimize contention on Cray's SeaStar2+ network, and extends it to a routed Lustre configuration to gain the same benefits when running against a center-wide file system. Our experiments using half of the resources of Spider - the center-wide file system at the Oak Ridge Leadership Computing Facility - show that I/O write bandwidth can be improved by up to 45% (from 71.9 to 104 GB/s) for a direct-attached configuration and by 137% (47.6 GB/s to 115 GB/s) for a routed configuration. We demonstrated up to 20.7% reduction in run-time for production scientific applications. With the full Spider system, we demonstrated over 240 GB/s of aggregate bandwidth using our techniques. David Dillow, Galen M. Shipman, Sarp Oral, Youngjae Kim 0001 |
IPCCC | 3 |
| 2011 | A semi-preemptive garbage collector for solid state drivesabstractNAND flash memory is a preferred storage media for various platforms ranging from embedded systems to enterprise-scale systems. Flash devices do not have any mechanical moving parts and provide low-latency access. They also require less power compared to rotating media. Unlike hard disks, flash devices use out-of-update operations and they require a garbage collection (GC) process to reclaim invalid pages to create free blocks. This GC process is a major cause of performance degradation when running concurrently with other I/O operations as internal bandwidth is consumed to reclaim these invalid pages. The invocation of the GC process is generally governed by a low watermark on free blocks and other internal device metrics that different workloads meet at different intervals. This results in I/O performance that is highly dependent on workload characteristics. In this paper, we examine the GC process and propose a semi-preemptive GC scheme that can preempt on-going GC processing and service pending I/O requests in the queue. Moreover, we further enhance flash performance by pipelining internal GC operations and merge them with pending I/O requests whenever possible. Our experimental evaluation of this semi-preemptive GC sheme with realistic workloads demonstrate both improved performance and reduced performance variability. Write-dominant workloads show up to a 66.56% improvement in average response time with a 83.30% reduced variance in response time compared to the non-preemptive GC scheme. Junghee Lee 0004, Youngjae Kim 0001, Galen M. Shipman, Sarp Oral, Feiyi Wang, Jongman Kim |
ISPASS | 4 |
| 2011 | Harmonia: A globally coordinated garbage collector for arrays of Solid-State DrivesabstractSolid-State Drives (SSDs) offer significant performance improvements over hard disk drives (HDD) on a number of workloads. The frequency of garbage collection (GC) activity is directly correlated with the pattern, frequency, and volume of write requests, and scheduling of GC is controlled by logic internal to the SSD. SSDs can exhibit significant performance degradations when garbage collection (GC) conflicts with an ongoing I/O request stream. When using SSDs in a RAID array, the lack of coordination of the local GC processes amplifies these performance degradations. No RAID controller or SSD available today has the technology to overcome this limitation. This paper presents Harmonia, a Global Garbage Collection (GGC) mechanism to improve response times and reduce performance variability for a RAID array of SSDs. Our proposal includes a high-level design of SSD-aware RAID controller and GGC-capable SSD devices, as well as algorithms to coordinate the global GC cycles. Our simulations show that this design improves response time and reduces performance variability for a wide variety of enterprise workloads. For bursty, write dominant workloads response time was improved by 69% while performance variability was reduced by 71%. Youngjae Kim 0001, Sarp Oral, Galen M. Shipman, Junghee Lee 0004, David Dillow, Feiyi Wang |
MSST | 2 |
| 2010 | Efficient Object Storage Journaling in a Distributed Parallel File System
Sarp Oral, Feiyi Wang, David Dillow, Galen M. Shipman, Ross G. Miller, Oleg Drokin |
FAST | 1 |
| 2008 | Empirical Analysis of a Large-Scale Hierarchical Storage System
Weikuan Yu, Sarp Oral, Shane Canon, Jeffrey S. Vetter, Ramanan Sankaran |
Euro-Par | 2 |
| 2008 | Performance characterization and optimization of parallel I/O on the Cray XTabstractThis paper presents an extensive characterization, tuning, and optimization of parallel I/O on the Cray XT supercomputer, named Jaguar, at Oak Ridge National Laboratory. We have characterized the performance and scalability for different levels of storage hierarchy including a single Lustre object storage target, a single S2A storage couplet, and the entire system. Our analysis covers both data- and metadata-intensive I/O patterns. In particular, for small, non-contiguous data- intensive I/O on Jaguar, we have evaluated several parallel I/O techniques, such as data sieving and two- phase collective I/O, and shed light on their effectiveness. Based on our characterization, we have demonstrated that it is possible, and often prudent, to improve the I/O performance of scientific benchmarks and applications by tuning and optimizing I/O. For example, we demonstrate that the I/O performance of the S3D combustion application can be improved at large scale by tuning the I/O system to avoid a bandwidth degradation of 49% with 8192 processes when compared to 4096 processes. We have also shown that the performance of Flash I/O can be improved by 34% by tuning the collective I/O parameters carefully. Weikuan Yu, Jeffrey S. Vetter, Sarp Oral |
IPDPS | 3 |
| 2003 | Performance analysis of HP AlphaServer ES80 vs. SAN-based clustersabstractThe last decade has introduced various affordable computing platforms to the parallel computing community. Distributed shared-memory systems and clusters built with commercial-off-the-shelf (COTS) parts and interconnected with high-performance networks have proven to be serious alternatives to expensive supercomputers in terms of both performance and cost. HP's new AlphaServer ES80 is an example of distributed shared-memory systems, while SCI and Myrinet are the two most widely used high-performance interconnects in building parallel-computing clusters. In this study, we experimentally compare the performance of these parallel computer systems. The emphasis is pointing out the strengths and weakness of the HPs AlphaServer ES80 in comparison with high-performance SCI and Myrinet clusters. We evaluated the systems in terms of sustainable memory bandwidth, interprocess communication and overall parallel computation performance using various widely-accepted benchmarks such as STREAM, PALLAS PMB-MP1, and NAS2.3 parallel suite. It was observed that the HPs AlphaServer ES80, executing Linux, provides remarkable computing power while its communication subsystem cannot handle heavy loads of small messages as effectively. Burt Gordon, Sarp Oral, Hung-Hsun Su, Alan D. George |
IPCCC | 2 |
| 2003 | A User-level Multicast Performance Comparison of Scalable Coherent Interface and Myrinet InterconnectsabstractThis paper compares and evaluates the multicast performance of two of the most widely deployed system-area networks (SANs), Dolphin's scalable coherent interface (SCI) and Myricom's Myrinet. Both networks deliver low latency and high bandwidth to applications, but do not support multicast in hardware. We compared SCI and Myrinet in terms of their user-level performance using various software-based multicast algorithms under various networking and multicasting scenarios. The strengths and weaknesses of each network are comparatively presented in terms of numerous metrics, such as multicast completion latency, CPU utilization, link concentration and concurrency. Sarp Oral, Alan D. George |
LCN | 1 |
| 2002 | Gigabit COTS Ethernet Switch Evaluation for AvionicsabstractThe evolving network needs of both commercial and military aircraft are expanding the requirements from slow speeds (1 Mbps) to over 1 Gbps for many new aircraft designs, driven primarily by advanced video technology. Low-cost Ethernet switch technology is a key enabler for advanced aircraft system interconnection. Our research focuses on the evaluation of several gigabit Ethernet (GE) COTS switches to meet avionic requirements. The results indicate that GE COTS switches are beginning to approach the required performance for future avionic applications. Jack L. Meier, Alan D. George, Sarp Oral |
LCN | 4 |
| 2002 | A Comparative Throughput Analysis of Scalable Coherent Interface and MyrinetabstractIt has become increasingly popular to construct large parallel computers by connecting many inexpensive nodes built with commercial-off-the-shelf (COTS) parts. These clusters can be built at a much lower cost than traditional supercomputers of comparable performance. A key decision that will greatly affect the overall performance of the cluster is the method used to connect the nodes together. Choosing the best interconnect and topology is not at all trivial since performance and cost will change as the system size is scaled. This paper presents throughput models used for the analysis and comparison of performance in two leading system area networks (SAN), Myrinet and Scalable Coherent Interface (SCI). First, analytical models for throughput are developed by determining the theoretical bandwidth of all internal buses and links that are part of the interconnect architecture. Then, experiments are conducted to measure the actual bandwidth available at each of these components, and the models are calibrated so they accurately represent the experimental results. Finally, the models are used to compare the maximum throughput of Myrinet and SCI systems with respect to system size and overall dollar cost. Sarp Millich, Alan D. George, Sarp Oral |
LCN | 3 |
| 2002 | Multicast Performance Analysis for High-Speed Torus NetworksabstractOverall efficiency of high-performance computing clusters not only relies on the computing power of the individual nodes, but also on the performance that the underlying network can provide to the computational application. Although modern high-performance networks, especially system area networks (SAN), have high unicast performance, they do not support multicast communication in hardware. This research experimentally evaluates the performance of various protocols for unicast-based and path-based multicast communication on high-speed torus networks. Software-based multicast performance results of selected algorithms on a 16-node Scalable Coherent Interface (SCI) torus are given. The strengths and weaknesses of the various protocols are illustrated in terms of startup and completion latency, CPU utilization, and link utilization and concurrency. Sarp Oral, Alan D. George |
LCN | 1 |
| 2002 | Design and Analysis of a Dynamically Reconfigurable Network ProcessorabstractThe combination of high-performance processing power and flexibility found in network processors (NPs) has made them a good solution for today's packet processing needs. Similarly, the emerging technology of reconfigurable computing (RC) has made advances in packet processing as well as other point-solution markets. Current NP designs offer configurable elements but generally do not use dynamic RC techniques for run-time reconfiguration. Incorporating RC into NP designs to enhance packet processing is a natural progression for both of these emerging technologies. This paper presents the simulation results of a novel design for a RC-enhanced NP based on the Intel IXP1200 NIC design philosophy. The enhanced NP's performance is compared to that of the baseline NP in terms of three normalized traffic patterns and a case-study traffic pattern based on a military application. The results demonstrate that the enhanced NP significantly outperforms the baseline NP design in terms of latency for prioritized traffic that is non-uniform. Ian A. Troxel, Alan D. George, Sarp Oral |
LCN | 3 |