Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

George Ostrouchov

dblp:23/4445 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
3since 2021 · last 2026
0000-0002-2043-6026ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Hardware reliability and fault tolerance · 35% High-performance computing · 32% GPUs and heterogeneous computing · 15%
Theoretical computer science
1 paper
Algorithms and data structures · 100%

Topics — the 14 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware reliability and fault tolerance
failure analysis
0.412020
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
GPUs and heterogeneous computing
GPU reliability
0.412020
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
Hardware reliability and fault tolerance › reliability analysis
lifetime analysis
0.412020
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
High-performance computing
HPC storage systems
0.412019
Can I/O Variability Be Reduced on QoS-Less HPC Storage Systems? · IEEE Trans. Computers 2019
High-performance computing › parallel i/o
i/o variability
0.412019
Can I/O Variability Be Reduced on QoS-Less HPC Storage Systems? · IEEE Trans. Computers 2019
Parallel and multicore computing
parallel programming models
0.112012
OpenMP-style parallelism in data-centered multicore computing with R · PPoPP 2012
High-performance computing
supercomputing
0.112020
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
Hardware reliability and fault tolerance
system reliability
0.112020
GPU lifetimes on titan supercomputer: survival analysis and reliability · SC 2020
Storage systems
i/o scheduling
0.112019
Can I/O Variability Be Reduced on QoS-Less HPC Storage Systems? · IEEE Trans. Computers 2019
Distributed systems
fault tolerance
0.112009
A tunable holistic resiliency approach for high-performance computing systems · PPoPP 2009
Distributed systems › fault tolerance
resilience
0.112009
A tunable holistic resiliency approach for high-performance computing systems · PPoPP 2009
Algorithms and data structures › numerical linear algebra
dimensionality reduction
0.112005
On FastMap and the Convex Hull of Multivariate Data: Toward Fast and Robust Dimension Reduction · IEEE Trans. Pattern Anal. Mach. Intell. 2005
Parallel and multicore computing
parallel programming runtimes
0.012012
OpenMP-style parallelism in data-centered multicore computing with R · PPoPP 2012
Bioinformatics and computational biology
protein-protein interaction prediction
0.012003
Inference of Protein-Protein Interactions by Unlikely Profile Pair · ICDM 2003

Methods — techniques the papers use, named apart from their topics

time between failures analysis · 0.4statistical survival analysis · 0.4throttling · 0.4messaging-based re-routing · 0.4analytical modeling · 0.4mapreduce · 0.1code transformation · 0.1preemptive migration · 0.1failure prediction · 0.1checkpoint/restart · 0.1robust statistics · 0.1convex hull · 0.1statistical simulation · 0.0bootstrapping · 0.0
YearPublicationVenuePosition
2026 Modeling Spatially Correlated Failure-Time Data Under Euclidean and Toroidal Distance Functions With an Application to Titan GPU Data
Jared M. Clark, Jie Min, Yueyao Wang, Yili Hong 0001, George Ostrouchov
IEEE Trans. Reliab.5
2025 ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis
Ziming Gan, Doudou Zhou, Everett Neil Rush, Vidul Ayakulangara Panickan, Yuk-Lam Ho, George Ostrouchov, Shuting Shen, Xin Xiong 0006, Kimberly F. Greco, Chuan Hong, Clara-Lea Bonzel, Jun Wen 0001, Lauren Costa, Tianrun A. Cai, Edmon Begoli, Zongqi Xia, John Michael Gaziano, Katherine P. Liao, Kelly Cho, Tianxi Cai
J. Biomed. Informatics6
2024 Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study
abstract
GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge.
Vladyslav Oles, Anna Schmedding, George Ostrouchov, Woong Shin, Evgenia Smirni, Christian Engelmann
ICS3
2020 GPU lifetimes on titan supercomputer: survival analysis and reliability
abstract
The Cray XK7 Titan was the top supercomputer system in the world for a long time and remained critically important throughout its nearly seven year life. It was an interesting machine from a reliability viewpoint as most of its power came from 18,688 GPUs whose operation was forced to execute three rework cycles, two on the GPU mechanical assembly and one on the GPU circuitboards. We write about the last rework cycle and a reliability analysis of over 100,000 years of GPU lifetimes during Titan’s 6-year-long productive period. Using time between failures analysis and statistical survival analysis techniques, we find that GPU reliability is dependent on heat dissipation to an extent that strongly correlates with detailed nuances of the cooling architecture and job scheduling. We describe the history, data collection, cleaning, and analysis and give recommendations for future supercomputing systems. We make the data and our analysis codes publicly available.
George Ostrouchov, Don E. Maxwell, Rizwan A. Ashraf, Christian Engelmann, Mallikarjun Shankar, James H. Rogers
SC1
2019 Can I/O Variability Be Reduced on QoS-Less HPC Storage Systems?
abstract
For a production high-performance computing (HPC) system, where storage devices are shared between multiple applications and managed in a best effort manner, I/O contention is often a major problem. In this paper, we propose a balanced messaging-based re-routing in conjunction with throttling at the middleware level. This work tackles two key challenges that have not been fully resolved in the past: whether I/O variability can be reduced on a QoS-less HPC storage system, and how to design a runtime scheduling system that can scale up to a large amount of cores. The proposed scheme uses a two-level messaging system to re-route I/O requests to a less congested storage location so that write performance is improved, while limiting the impact on read by throttling re-routing. An analytical model is derived to guide the setup of optimal throttling factor. We thoroughly analyze the virtual messaging layer overhead and explore whether the in-transit buffering is effective in managing I/O variability. Contrary to the intuition, in-transit buffer cannot completely solve the problem. It can reduce the absolute variability but not the relative variability. The proposed scheme is verified against a synthetic benchmark as well as being used by production applications.
Dan Huang 0001, Qing Liu 0002, Jong Choi 0001, Norbert Podhorszki, Scott Klasky, Jeremy Logan, George Ostrouchov, Xubin He, Matthew Wolf
IEEE Trans. Computers7
2018 A View from ORNL: Scientific Data Research Opportunities in the Big Data Age
abstract
One of the core issues across computer and computational science today is adapting to, managing, and learning from the influx of "Big Data". In the commercial space, this problem has led to a huge investment in new technologies and capabilities that are well adapted to dealing with the sorts of human-generated logs, videos, texts, and other large-data artifacts that are processed and resulted in an explosion of useful platforms and languages (Hadoop, Spark, Pandas, etc.). However, translating this work from the enterprise space to the computational science and HPC community has proven somewhat difficult, in part because of some of the fundamental differences in type and scale of data and timescales surrounding its generation and use. We describe a forward-looking research and development plan which centers around the concept of making Input/Output (I/O) intelligent for users in the scientific community, whether they are accessing scalable storage or performing in situ workflow tasks. Much of our work is based on our experience with the Adaptable I/O System (ADIOS 1.X), and our next generation version of the software ADIOS 2.X [1].
Scott Klasky, Matthew Wolf, Mark Ainsworth, Chuck Atkins, Jong Choi 0001, Greg Eisenhauer, Berk Geveci, William F. Godoy, Mark Kim, James Kress, Tahsin M. Kurç, Qing Liu 0002, Jeremy Logan, Arthur B. Maccabe, Kshitij Mehta, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Eric Suchyta, Lipeng Wan 0001
ICDCS16
2018 An evaluation of the state of time synchronization on leadership class supercomputers
abstract
Summary We present a detailed examination of time agreement characteristics for nodes within extreme‐scale parallel computers. Using a software tool we introduce in this paper, we quantify attributes of clock skew among nodes in three representative high‐performance computers sited at three national laboratories. Our measurements detail the statistical properties of time agreement among nodes and how time agreement drifts over typical application execution durations. We discuss the implications of our measurements, why the current state of the field is inadequate, and propose strategies to address observed shortcomings.
Terry R. Jones, George Ostrouchov, Gregory A. Koenig, Oscar H. Mondragon, Patrick G. Bridges
Concurr. Comput. Pract. Exp.2
2017 TGE: Machine Learning Based Task Graph Embedding for Large-Scale Topology Mapping
abstract
Task mapping is an important problem in parallel and distributed computing. The goal in task mapping is to find an optimal layout of the processes of an application (or a task) onto a given network topology. We target this problem in the context of staging applications. A staging application consists of two or more parallel applications (also referred to as staging tasks) which run concurrently and exchange data over the course of computation. Task mapping becomes a more challenging problem in staging applications, because not only data is exchanged between the staging tasks, but also the processes of a staging task may exchange data with each other. We propose a novel method, called Task Graph Embedding (TGE), that harnesses the observable graph structures of parallel applications and network topologies. TGE employs a machine learning based algorithm to find the best representation of a graph, called an embedding, onto a space in which the task-to-processor mapping problem can be solved. We evaluate and demonstrate the effectiveness of TGE experimentally with the communication patterns extracted from runs of XGC, a large-scale fusion simulation code, on Titan.
Jong Choi 0001, Jeremy Logan, Matthew Wolf, George Ostrouchov, Tahsin M. Kurç, Qing Liu 0002, Norbert Podhorszki, Scott Klasky, Melissa Romanus, Manish Parashar, Michael Churchill, Choong-Seock Chang
CLUSTER4
2017 Extending Skel to Support the Development and Optimization of Next Generation I/O Systems
abstract
As the memory and storage hierarchy get deeper and more complex, it is important to have new benchmarks and evaluation tools that allow us to explore the emerging middleware solutions to use this hierarchy. Skel is a tool aimed at automating and refining this process of studying HPC I/O performance. It works by generating application I/O kernel/benchmarks as determined by a domain-specific model. This paper provides some techniques for extending Skel to address new situations and to answer new research questions. For example, we document use cases as diverse as using Skel to troubleshoot I/O performance issues for remote users, refining an I/O system model, and facilitating the development and testing of a mechanism for runtime monitoring and performance analytics. We also discuss data oriented extensions to Skel to support the study of compression techniques for Exascale scientific data management.
Jeremy Logan, Jong Choi 0001, Matthew Wolf, George Ostrouchov, Lipeng Wan 0001, Norbert Podhorszki, William F. Godoy, Scott Klasky, Erich Lohrmann, Greg Eisenhauer, Chad Wood, Kevin A. Huck
CLUSTER4
2017 Computing Just What You Need: Online Data Analysis and Reduction at Extreme Scales
Ian T. Foster, Mark Ainsworth, Bryce Allen, Julie Bessac, Franck Cappello, Jong Choi 0001, Emil M. Constantinescu, Philip E. Davis, Sheng Di, Zichao Wendy Di, Hanqi Guo 0001, Scott Klasky, Kerstin Kleese van Dam, Tahsin M. Kurç, Qing Liu 0002, Abid Malik, Kshitij Mehta, Klaus Mueller 0001, Todd S. Munson, George Ostrouchov, Manish Parashar, Tom Peterka, Line C. Pouchard, Dingwen Tao, Ozan Tugluk, Stefan M. Wild, Matthew Wolf, Justin M. Wozniak, Wei Xu 0020, Shinjae Yoo
Euro-Par20
2017 Exacution: Enhancing Scientific Data Management for Exascale
abstract
As we continue toward exascale, scientific data volume is continuing to scale and becoming more burdensome to manage. In this paper, we lay out opportunities to enhance state of the art data management techniques. We emphasize well-principled data compression, and using it to achieve progressive refinement. This can both accelerate I/O and afford the user increased flexibility when she interacts with the data. The formulation naturally maps onto enabling partitioning of the progressively improving-quality representations of a data quantity into different media-type destinations, to keep the highest priority information as close as possible to the computation, and take advantage of deepening memory/storage hierarchies in ways not previously possible. Careful monitoring is requisite to our vision, not only to verify that compression has not eliminated salient features in the data, but also to better understand the performance of massively parallel scientific applications. Increased mathematical rigor would be ideal,to help bring compression on a better-understood theoretical footing, closer to the relevant scientific theory, more aware of constraints imposed by the science, and more tightly error-controlled. Throughout, we highlight pathfinding research we have begun exploring related these topics, and comment toward future work that will be needed.
Scott Klasky, Eric Suchyta, Mark Ainsworth, Qing Liu 0002, Ben Whitney, Matthew Wolf, Jong Choi 0001, Ian T. Foster, Mark Kim, Jeremy Logan, Kshitij Mehta, Todd S. Munson, George Ostrouchov, Manish Parashar, Norbert Podhorszki, David Pugmire, Lipeng Wan 0001
ICDCS13
2017 Comprehensive Measurement and Analysis of the User-Perceived I/O Performance in a Production Leadership-Class Storage System
abstract
With the increase of the scale and intensity of the parallel I/O workloads generated by those scientific applications running on high performance computing facilities, understanding the I/O dynamics, especially the root cause of the I/O performance variability and degradation in HPC environment, have become extremely critical to the HPC community. In this paper, we run extensive I/O measuring tests on a production leadership-class storage system to capture the performance variabilities of large-scale parallel I/O. Analyzing these results and its statistic correlation revealed some valuable insights into the characteristics of the storage system and the root cause of I/O performance variability. Further, we leverage these findings and propose an I/O middleware design refactoring which can improve the performance of the parallel I/O by optimizing the data striping and placement. Our preliminary evaluation results demonstrate the proposed approach can reduce the average per-process write latency by at least 80% and the maximum per-process write latency by at least 20%.
Lipeng Wan 0001, Matthew Wolf, Feiyi Wang, Jong Choi 0001, George Ostrouchov, Scott Klasky
ICDCS5
2012 Understanding I/O Performance Using I/O Skeletal Applications
Jeremy Logan, Scott Klasky, Hasan Abbasi, Qing Liu 0002, George Ostrouchov, Manish Parashar, Norbert Podhorszki, Yuan Tian 0004, Matthew Wolf
Euro-Par5
2012 OpenMP-style parallelism in data-centered multicore computing with R
abstract
R1 is a domain specific language widely used for data analysis by the statistics community as well as by researchers in finance, biology, social sciences, and many other disciplines. As R programs are linked to input data, the exponential growth of available data makes high-performance computing with R imperative. To ease the process of writing parallel programs in R, code transformation from a sequential program to a parallel version would bring much convenience to R users. In this paper, we present our work in semi-automatic parallelization of R codes with user-added OpenMP-style pragmas. While such pragmas are used at the frontend, we take advantage of multiple parallel backends with different R packages. We provide flexibility for importing parallelism with plug-in components, impose built-in MapReduce for data processing, and also maintain code reusability. We illustrate the advantage of the on-the-fly mechanisms which can lead to significant applications in data-centered parallel computing.
Lei Jiang 0010, Pragneshkumar B. Patel, George Ostrouchov, Ferdinand Jamitzky
PPoPP3
2009 Blue Gene/L Log Analysis and Time to Interrupt Estimation
abstract
System- and application-level failures could be characterized by analyzing relevant log files. The resulting data might then be used in numerous studies on and future developments for the mission-critical and large scale computational architecture, including fields such as failure prediction, reliability modeling, performance modeling and power awareness. In this paper, system logs covering a six month period of the Blue Gene/L supercomputer were obtained and subsequently analyzed. Temporal filtering was applied to remove duplicated log messages. Optimistic and pessimistic perspectives were exerted on filtered log information to observe failure behavior within the system. Further, various time to repair factors were applied to obtain application time to interrupt, which will be exploited in further resilience modeling research.
Narate Taerat, Nichamon Naksinehaboon, Clayton Chandler, James Elliott, Chokchai Leangsuksun, George Ostrouchov, Stephen L. Scott, Christian Engelmann
ARES6
2009 A tunable holistic resiliency approach for high-performance computing systems
abstract
In order to address anticipated high failure rates, resiliency characteristics have become an urgent priority for next-generation extreme-scale high-performance computing (HPC) systems. This poster describes our past and ongoing efforts in novel fault resilience technologies for HPC. Presented work includes proactive fault resilience techniques, system and application reliability models and analyses, failure prediction, transparent process- and virtual-machine-level migration, and trade-off models for combining preemptive migration with checkpoint/restart. This poster summarizes our work and puts all individual technologies into context with a proposed holistic fault resilience framework.
Stephen L. Scott, Christian Engelmann, Geoffroy Vallée, Thomas J. Naughton, Anand Tikotekar, George Ostrouchov, Chokchai Leangsuksun, Nichamon Naksinehaboon, Raja Nassar, Mihaela Paun, Frank Mueller 0001, Chao Wang 0056, Arun Babu Nagarajan, Jyothish Varma
PPoPP6
2005 On FastMap and the Convex Hull of Multivariate Data: Toward Fast and Robust Dimension Reduction
abstract
FastMap is a dimension reduction technique that operates on distances between objects. Although only distances are used, implicitly the technique assumes that the objects are points in a p-dimensional Euclidean space. It selects a sequence of k < or = p orthogonal axes defined by distant pairs of points (called pivots) and computes the projection of the points onto the orthogonal axes. We show that FastMap uses only the outer envelope of a data set. Pivots are taken from the faces, usually vertices, of the convex hull of the data points in the original implicit Euclidean space. This provides a bridge to results in robust statistics, where the convex hull is used as a tool in multivariate outlier detection and in robust estimation methods. The connection sheds new light on the properties of FastMap, particularly its sensitivity to outliers, and provides an opportunity for a new class of dimension reduction algorithms, RobustMaps, that retain the speed of FastMap and exploit ideas in robust statistics.
George Ostrouchov, Nagiza F. Samatova
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 Reservoir-Based Random Sampling with Replacement from Data Stream
abstract
Random sampling is a widely accepted basis for estimation from large data sets that outstrip available computer memory. When the data comes as a stream, its total size is potentially infinite and usually only one pass through the data is possible. Reservoir sampling is a method of maintaining a fixed size random sample from streaming data. All reservoir schemes that have been introduced in the past are random sampling without replacement; no duplicates are allowed in a sample. This paper introduces a new method for reservoir sampling with replacement. We first prove that the proposed method indeed maintains a random sample with replacement at any given time. Then we introduce a refined version that significantly speeds up the overall sampling procedure.
George Ostrouchov, Nagiza F. Samatova, Al Geist
SDM2
2003 Inference of Protein-Protein Interactions by Unlikely Profile Pair
abstract
We note that a set of statistically "unusual" protein-profile pairs in experimentally determined database of protein-protein interactions can typify protein-protein interactions, and propose a novel method called PICUPP that sifts such protein-profile pairs using a statistical simulation. It is demonstrated that unusual Pfam and InterPro profile pairs can be extracted from the DIP database using a bootstrapping approach. We particularly illustrate that such protein-profile pairs can be used for predicting putative pairs of interacting proteins. Their prediction accuracies are around 86% and 90% when InterPro and Pfam profiles are used, respectively at 75% confidence level.
George Ostrouchov, Gong-Xin Yu, Al Geist, Andrey Gorin, Nagiza F. Samatova
ICDM2
2002 RACHET: An Efficient Cover-Based Merging of Clustering Hierarchies from Distributed Datasets
Nagiza F. Samatova, George Ostrouchov, Al Geist, Anatoli V. Melechko
Distributed Parallel Databases2
1999 Mining consumer product data via latent semantic indexing
abstract
One important focus of data mining research is in the development of algorithms for extracting valuable information from large databases in order to facilitate business decisions. This study explores a new technique for data mining – latent semantic indexing (LSI). LSI is an efficient information retrieval method for textual documents. By determining the singular value decomposition (SVD) of a large sparse term-by-document matrix, LSI constructs an approximate vector space model which represents important associative relationships between terms and documents that are not evident in individual documents. This paper explores the applicability of the LSI model to numerical databases, namely consumer product data. By properly choosing attributes of data records as terms or documents, a term-by-document frequency matrix is built from which a distribution-based indexing scheme is employed to construct a correlated distribution matrix (CDM). An LSI-like vector space model is then used to detect useful or hidden patterns in the numerical data. The extracted information can then be validated using statistical hypotheses testing or resampling. LSI is an automatic yet intelligent indexing method. Its application to numerical data introduces a promising way to discover knowledge in important commercial application areas such as retail and consumer banking.
Jingqian Jiang, Michael W. Berry, June M. Donato, George Ostrouchov, Nancy W. Grady
Intell. Data Anal.4