Vitali A. Morozov

dblp:15/10441 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
1since 2021 · last 2025
0009-0007-5620-5266ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
High-performance computing · 63% Performance modeling and evaluation · 10% Storage systems · 8%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 100%

Topics — the 28 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
scientific computing systems
1.242025
Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability · SC 2025
HACC: extreme scaling and performance across diverse architectures · SC 2013
The universe at extreme scale: multi-petaflop sky simulation on the BG/Q · SC 2012
High-performance computing
performance optimization at scale
1.232025
Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability · SC 2025
HACC: extreme scaling and performance across diverse architectures · SC 2013
The universe at extreme scale: multi-petaflop sky simulation on the BG/Q · SC 2012
High-performance computing › large-scale simulation
exascale simulation
0.912025
Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability · SC 2025
High-performance computing › performance optimization at scale
extreme-scale scalability
0.322013
HACC: extreme scaling and performance across diverse architectures · SC 2013
The universe at extreme scale: multi-petaflop sky simulation on the BG/Q · SC 2012
High-performance computing › performance engineering
performance reproducibility
0.312017
Run-to-run variability on Xeon Phi based cray XC systems · SC 2017
Performance modeling and evaluation
performance variability
0.312017
Run-to-run variability on Xeon Phi based cray XC systems · SC 2017
Electronic design automation › yield analysis
process variation modeling
0.312017
Run-to-run variability on Xeon Phi based cray XC systems · SC 2017
High-performance computing › scientific data analysis
in-situ analysis
0.312025
Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability · SC 2025
Cloud and datacenter computing › job scheduling
batch scheduling
0.212016
Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints · IEEE Trans. Parallel Distributed Syst. 2016
High-performance computing
supercomputing
0.212016
Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints · IEEE Trans. Parallel Distributed Syst. 2016
Storage systems › i/o optimization
i/o forwarding
0.222011
Topology-aware data movement and staging for I/O acceleration on Blue Gene/P supercomputing systems · SC 2011
Accelerating I/O Forwarding in IBM Blue Gene/P Systems · SC 2010
High-performance computing › supercomputing
leadership-class systems
0.222011
Topology-aware data movement and staging for I/O acceleration on Blue Gene/P supercomputing systems · SC 2011
Accelerating I/O Forwarding in IBM Blue Gene/P Systems · SC 2010
GPUs and heterogeneous computing
GPU kernel optimization
0.112012
Dataflow-driven GPU performance projection for multi-kernel transformations · SC 2012
Computational science and engineering › computational mechanics
biomechanical simulation
0.112011
A new computational paradigm in multiscale simulations: application to brain blood flow · SC 2011
Computational science and engineering › numerical simulation
multiscale simulation
0.112011
A new computational paradigm in multiscale simulations: application to brain blood flow · SC 2011
Performance modeling and evaluation
analytical modeling
0.112011
GROPHECY: GPU performance projection from CPU code skeletons · SC 2011
Storage systems › i/o optimization
collective i/o
0.112011
Topology-aware data movement and staging for I/O acceleration on Blue Gene/P supercomputing systems · SC 2011
GPUs and heterogeneous computing
GPU computing
0.112011
GROPHECY: GPU performance projection from CPU code skeletons · SC 2011
Storage systems
i/o optimization
0.112011
Topology-aware data movement and staging for I/O acceleration on Blue Gene/P supercomputing systems · SC 2011
Performance modeling and evaluation
performance prediction
0.112011
GROPHECY: GPU performance projection from CPU code skeletons · SC 2011
Storage systems
i/o scheduling
0.112010
Accelerating I/O Forwarding in IBM Blue Gene/P Systems · SC 2010
Interconnection networks and networks-on-chip › network topology › low-diameter topology
dragonfly network
0.112017
Run-to-run variability on Xeon Phi based cray XC systems · SC 2017
Performance modeling and evaluation
benchmarking
0.112016
Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints · IEEE Trans. Parallel Distributed Syst. 2016
Compilers and program optimization
loop transformation
0.012012
Dataflow-driven GPU performance projection for multi-kernel transformations · SC 2012
Compilers and program optimization › program transformation
code skeletonization
0.012011
GROPHECY: GPU performance projection from CPU code skeletons · SC 2011
Parallel and multicore computing › data distribution
domain partitioning
0.012011
A new computational paradigm in multiscale simulations: application to brain blood flow · SC 2011
Interconnection networks and networks-on-chip
network topology
0.012011
Topology-aware data movement and staging for I/O acceleration on Blue Gene/P supercomputing systems · SC 2011
High-performance computing
parallel numerical algorithms
0.012011
A new computational paradigm in multiscale simulations: application to brain blood flow · SC 2011

Methods — techniques the papers use, named apart from their topics

tree solver · 0.9separation-of-scale · 0.9multi-tiered i/o · 0.9particle-grid methods · 0.5analytical framework · 0.3performance characterization · 0.3scheduling scheme · 0.2comparative study · 0.2benchmarking · 0.2asynchronous data staging · 0.2data flow analysis · 0.1spectral element method · 0.1molecular dynamics · 0.1dissipative particle dynamics · 0.1code skeleton transformation · 0.1analytical model · 0.1
YearPublicationVenuePosition
2025 Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability
abstract
Resolving the most fundamental questions in cosmology requires simulations that match the scale, fidelity, and physical complexity demanded by next-generation sky surveys. To achieve the realism needed for this critical scientific partnership, detailed gas dynamics must be treated self-consistently with gravity for end-to-end modeling of structure formation. Exascale computing enables simulations that span survey-scale volumes while incorporating key astrophysical processes that shape complex cosmic structures. We present results from CRK-HACC, a cosmological hydrodynamics code built for extreme scalability. Using separation-of-scale techniques, GPU-resident tree solvers, in situ analysis pipelines, and multi-tiered I/O, CRK-HACCexecuted Frontier-E: a four trillion particle full-sky simulation, over an order of magnitude larger than previous efforts. The run achieved 513.1 PFLOPs peak performance, processing 46.6 billion particles per second and writing more than 100 PB of data in just over one week of runtime. Frontier-E marks a significant advance in predictive modeling for next-generation cosmological science.
Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban Rangel, Salman Habib 0002, Katrin Heitmann, Patricia Larsen, Vitali A. Morozov, Adrian Pope, Claude-André Faucher-Giguère, Antigoni Georgiadou, Damien Lebrun-Grandié, Andrey Prokopenko
SC8
2017 Run-to-run variability on Xeon Phi based cray XC systems
abstract
The increasing complexity of HPC systems has introduced new sources of variability, which can contribute to significant differences in run-to-run performance of applications. With components at various levels of the system contributing variability, application developers and system users are now faced with the difficult task of running and tuning their applications in an environment where run-to-run performance measurements can vary by as much as a factor of two to three. In this study, we classify, quantify, and present ways to mitigate the sources of run-to-run variability on Cray XC systems with Intel Xeon Phi processors and a dragonfly interconnect. We further demonstrate that the code-tuning performance observed in a variability-mitigating environment correlates with the performance observed in production running conditions.
Sudheer Chunduri, Kevin Harms, Scott Parker, Vitali A. Morozov, Samuel Oshin, Naveen Cherukuri, Kalyan Kumaran
SC4
2016 Improving Data Transfer Throughput with Direct Search Optimization
abstract
Improving data transfer throughput over high-speed long-distance networks has become increasingly difficult. Numerous factors such as nondeterministic congestion, dynamics of the transfer protocol, and multiuser and multitask source and destination endpoints, as well as interactions among these factors, contribute to this difficulty. A promising approach to improving throughput consists in using parallel streams at the application layer. We formulate and solve the problem of choosing the number of such streams from a mathematical optimization perspective. We propose the use of direct search methods, a class of easy-to-implement and light-weight mathematical optimization algorithms, to improve the performance of data transfers by dynamically adapting the number of parallel streams in a manner that does not require domain expertise, instrumentation, analytical models, or historic data. We apply our method to transfers performed with the GridFTP protocol, and illustrate the effectiveness of the proposed algorithm when used within Globus, a state-of-the-art data transfer tool, on production WAN links and servers. We show that when compared to user default settings our direct search methods can achieve up to 10x performance improvement under certain conditions. We also show that our method can overcome performance degradation due to external compute and network load on source end points, a common scenario at high performance computing facilities.
Prasanna Balaprakash, Vitali A. Morozov, Rajkumar Kettimuthu, Kalyan Kumaran, Ian T. Foster
ICPP2
2016 Workflow performance improvement using model-based scheduling over multiple clusters and clouds
Ketan Maheshwari, Eun-Sung Jung, Jiayuan Meng, Vitali A. Morozov, Venkatram Vishwanath, Rajkumar Kettimuthu
Future Gener. Comput. Syst.4
2016 Improving Batch Scheduling on Blue Gene/Q by Relaxing Network Allocation Constraints
abstract
As systems scale toward exascale, many resources will become increasingly constrained. While some of these resources have historically been explicitly allocated, many-such as network bandwidth, I/O bandwidth, or power-have not. As systems continue to evolve, we expect many such resources to become explicitly managed. This change will pose critical challenges to resource management and job scheduling. In this paper, we explore the potential of relaxing network allocation constraints for Blue Gene systems. Our objective is to improve the batch scheduling performance, where the partition-based interconnect architecture provides a unique opportunity to explicitly allocate network resources to jobs. This paper makes three major contributions. The first is substantial benchmarking of parallel applications, focusing on assessing application sensitivity to communication bandwidth at large scale. The second is three new scheduling schemes using relaxed network allocation and targeted at balancing individual job performance with overall system performance. The third is a comparative study of our scheduling schemes versus the existing scheduler on Mira, a 48-rack Blue Gene/Q system at Argonne National Laboratory. Specifically, we use job traces collected from this production system.
Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan, Paul M. Rich, Wei Tang 0001, Vitali A. Morozov, Narayan Desai
IEEE Trans. Parallel Distributed Syst.6
2015 Improving Batch Scheduling on Blue Gene/Q by Relaxing 5D Torus Network Allocation Constraints
abstract
As systems scale toward exactable, many resources will become increasingly constrained. While some of these resources have historically been explicitly allocated, many -- such as network bandwidth, I/O bandwidth, or power -- have not. As systems continue to evolve, we expect many such resources to become explicitly managed. This change will pose critical challenges to resource management and job scheduling. In this paper, we explore the potentiality of relaxing network allocation constraints for Blue Gene systems. Our objectives to improve the batch scheduling performance, where the partition-based interconnect architecture provides a unique opportunity to explicitly allocate network resources to jobs. This paper makes three major contributions. The first is substantial benchmarking of parallel applications, focusing on assessing application sensitivity to communication bandwidth at large scale. The second is two new scheduling schemes using relaxed network allocation and targeted at balancing individual job performance with overall system performance. The third is a comparative study of our scheduling schemes versus the existing one under different workloads, using job traces collected from the 48-rack Mira, an IBM Blue Gene/Q system at Argonne National Laboratory.
Zhou Zhou 0006, Xu Yang 0009, Zhiling Lan, Paul M. Rich, Wei Tang 0001, Vitali A. Morozov, Narayan Desai
IPDPS6
2015 Design and performance characterization of electronic structure calculations on massively parallel supercomputers: a case study of GPAW on the Blue Gene/P architecture
abstract
SUMMARY Density function theory (DFT) is the most widely employed electronic structure method because of its favorable scaling with system size and accuracy for a broad range of molecular and condensed‐phase systems. The advent of massively parallel supercomputers has enhanced the scientific community's ability to study larger system sizes. Ground‐state DFT calculations on ∼ 103 valence electrons using traditional algorithms can be routinely performed on present‐day supercomputers. The performance characteristics of these massively parallel DFT codes on > 104 computer cores are not well understood. The GPAW code was ported an optimized for the Blue Gene/P architecture. We present our algorithmic parallelization strategy and interpret the results for a number of benchmark test cases.Copyright © 2013 John Wiley & Sons, Ltd.
Nichols A. Romero, Christian Glinsvad, Ask Hjorth Larsen, Jussi Enkovaara, Sameer Shende, Vitali A. Morozov, Jens J. Mortensen
Concurr. Comput. Pract. Exp.6
2014 Analytically Modeling Application Execution for Software-Hardware Co-design
abstract
Software-hardware co-design has become increasingly important as the scale and complexity of both are reaching an unprecedented level. To predict and understand application behavior on emerging or conceptual systems, existing research has mostly relied on cycle-accurate micro-architecture simulators, which are known to be time-consuming and are oblivious to workloads' control flow structure. As a result, simulations are often limited to small kernels, and the first step in the co-design process is often to extract important kernels, construct mini-applications, and identify potential hardware limitations. This requires a high level understanding about the full applications' potential behavior on a future system, e.g. the most time-consuming regions, the performance bottlenecks for these regions, etc. Unfortunately, such application knowledge gained from one system may not hold true on a future system. One solution is to instrument the full application with timers and simulate it with a reasonable input size, which can be a daunting task in itself. We propose an alternative approach to gain first-order insights into hardware-dependent application behavior by trading off the accuracy of analysis for improved efficiency. By modeling the execution flows of user applications and analyzing it using target hardware's performance models, our technique requires no cycle-accurate simulation on a prospective system. In fact, our technique's analysis time does not increase with the input data size.
Jichi Guo, Jiayuan Meng, Qing Yi, Vitali A. Morozov, Kalyan Kumaran
IPDPS4
2013 Early Experience on the Blue Gene/Q Supercomputing System
abstract
The Argonne Leadership Computing Facility (ALCF) is home to Mira, a 10 PF Blue Gene/Q (BG/Q) system. The BG/Q system is the third generation in Blue Gene architecture from IBM and like its predecessors combines system-onchip technology with a proprietary interconnect (5-D torus). Each compute node has 16 augmented PowerPC A2 processor cores with support for simultaneous multithreading, 4-wide double precision SIMD, and different data prefetching mechanisms. Mira offers several new opportunities for tuning and scaling scientific applications. This paper discusses our early experience with a subset of micro-benchmarks, MPI benchmarks, and a variety of science and engineering applications running at ALCF. Both performance and power are studied and results on BG/Q is compared with its predecessor BG/P. Several lessons gleaned from tuning applications on the BG/Q architecture for better performance and scalability are shared.
Vitali A. Morozov, Kalyan Kumaran, Venkatram Vishwanath, Jiayuan Meng, Michael E. Papka
IPDPS1
2013 HACC: extreme scaling and performance across diverse architectures
abstract
Supercomputing is evolving towards hybrid and accelerator-based architectures with millions of cores. The HACC (Hardware/Hybrid Accelerated Cosmology Code) framework exploits this diverse landscape at the largest scales of problem size, obtaining high scalability and sustained performance. Developed to satisfy the science requirements of cosmological surveys, HACC melds particle and grid methods using a novel algorithmic structure that flexibly maps across architectures, including CPU/GPU, multi/many-core, and Blue Gene systems. We demonstrate the success of HACC on two very different machines, the CPU/GPU system Titan and the BG/Q systems Sequoia and Mira, attaining unprecedented levels of scalable performance. We demonstrate strong and weak scaling on Titan, obtaining up to 99.2% parallel efficiency, evolving 1.1 trillion particles. On Sequoia, we reach 13.94 PFlops (69.2% of peak) and 90% parallel efficiency on 1,572,864 cores, with 3.6 trillion particles, the largest cosmological benchmark yet performed. HACC design concepts are applicable to several other supercomputer applications.
Salman Habib 0002, Vitali A. Morozov, Nicholas Frontiere, Hal Finkel, Adrian Pope, Katrin Heitmann
SC2
2012 The universe at extreme scale: multi-petaflop sky simulation on the BG/Q
abstract
Remarkable observational advances have established a compelling cross-validated model of the Universe. Yet, two key pillars of this model -- dark matter and dark energy -- remain mysterious. Next-generation sky surveys will map billions of galaxies to explore the physics of the 'Dark Universe'. Science requirements for these surveys demand simulations at extreme scales; these will be delivered by the HACC (Hybrid/Hardware Accelerated Cosmology Code) framework. HACC's novel algorithmic structure allows tuning across diverse architectures, including accelerated and multi-core systems. On the IBM BG/Q, HACC attains unprecedented scalable performance - currently 6.23 PFlops at 62% of peak and 92% parallel efficiency on 786,432 cores (48 racks) - at extreme problem sizes with up to almost two trillion particles, larger than any cosmological simulation yet performed. HACC simulations at these scales will for the first time enable tracking individual galaxies over the entire volume of a cosmological survey.
Salman Habib 0002, Vitali A. Morozov, Hal Finkel, Adrian Pope, Katrin Heitmann, Kalyan Kumaran, Tom Peterka, Joseph A. Insley, David Daniel, Patricia K. Fasel, Nicholas Frontiere, Zarija Lukic
SC2
2012 Dataflow-driven GPU performance projection for multi-kernel transformations
abstract
Applications often have a sequence of parallel operations to be offloaded to graphics processors; each operation can become an individual GPU kernel. Developers typically explore a variety of transformations for each kernel. Furthermore, it is well known that efficient data management is critical in achieving high GPU performance and that "fusing" multiple kernels into one may greatly improve data locality. Doing so, however, requires transformations across multiple, potentially nested, parallel loops; at the same time, the original code semantics and data dependency must be preserved. Since each kernel may have distinct data access patterns, their combined dataflow can be nontrivial. As a result, the complexity of multi-kernel transformations often leads to significant effort with no guarantee of performance benefits. This paper proposes a dataflow-driven analytical framework to project GPU performance for a sequence of parallel operations. Users need only provide CPU code skeletons for a sequence of parallel loops. The framework can then automatically identify opportunities for multi-kernel transformations and data management. It is also able to project the overall performance without implementing GPU code or using physical hardware.
Jiayuan Meng, Vitali A. Morozov, Venkatram Vishwanath, Kalyan Kumaran
SC2
2011 A new computational paradigm in multiscale simulations: application to brain blood flow
abstract
Interfacing atomistic-based with continuum-based simulation codes is now required in many multiscale physical and biological systems. We present the computational advances that have enabled the first multiscale simulation on 190,740 processors by coupling a high-order (spectral element) Navier-Stokes solver with a stochastic (coarse-grained) Molecular Dynamics solver based on Dissipative Particle Dynamics (DPD). The key contributions are proper interface conditions for overlapped domains, topology-aware communication, SIMDization, multiscale visualization and a new domain partitioning for atomistic solvers. We study blood flow in a patient-specific cerebrovasculature with a brain aneurysm, and analyze the interaction of blood cells with the arterial walls endowed with a glycocalyx causing thrombus formation and eventual aneurysm rupture. The macro-scale dynamics (about 3 billion unknowns) are resolved by NεκTαr - a spectral element solver; the micro-scale flow and cell dynamics within the aneurysm are resolved by an in-house version of DPD-LAMMPS (for an equivalent of about 100 billions molecules).
Leopold Grinberg, Joseph A. Insley, Vitali A. Morozov, Michael E. Papka, George Em Karniadakis, Dmitry A. Fedosov, Kalyan Kumaran
SC3
2011 GROPHECY: GPU performance projection from CPU code skeletons
abstract
We propose GROPHECY, a GPU performance projection framework that can estimate the performance benefit of GPU acceleration without actual GPU programming or hardware. Users need only to skeletonize pieces of CPU code that are targets for GPU acceleration. Code skeletons are automatically transformed in various ways to mimic tuned GPU codes with characteristics resembling real implementations. The synthesized characteristics are used by an existing analytical model to project GPU performance. The cost and benefit of GPU development can then be estimated according to the transformed code skeleton that yields the best projected performance. With GROPHECY, users can leap toward GPU acceleration only when the cost-benefit makes sense. The framework is validated using kernel benchmarks and data-parallel codes in legacy scientific applications. The measured performance of manually tuned codes deviates from the projected performance by 17% in geometric mean.
Jiayuan Meng, Vitali A. Morozov, Kalyan Kumaran, Venkatram Vishwanath, Thomas D. Uram
SC2
2011 Topology-aware data movement and staging for I/O acceleration on Blue Gene/P supercomputing systems
abstract
There is growing concern that I/O systems will be hard pressed to satisfy the requirements of future leadership-class machines. Even current machines are found to be I/O bound for some applications. In this paper, we identify existing performance bottlenecks in data movement for I/O on the IBM Blue Gene/P (BG/P) supercomputer currently deployed at several leadership computing facilities. We improve the I/O performance by exploiting the network topology of BG/P for collective I/O, leveraging data semantics of applications and incorporating asynchronous data staging. We demonstrate the efficacy of our approaches for synthetic benchmark experiments and for application-level benchmarks at scale on leadership computing systems.
Venkatram Vishwanath, Mark Hereld, Vitali A. Morozov, Michael E. Papka
SC3
2010 Accelerating I/O Forwarding in IBM Blue Gene/P Systems
abstract
Current leadership-class machines suffer from a significant imbalance between their computational power and their I/O bandwidth. I/O forwarding is a paradigm that attempts to bridge the increasing performance and scalability gap between the compute and I/O components of leadership-class machines to meet the requirements of data-intensive applications by shipping I/O calls from compute nodes to dedicated I/O nodes. I/O forwarding is a critical component of the I/O subsystem of the IBM Blue Gene/P supercomputer currently deployed at several leadership computing facilities. In this paper, we evaluate the performance of the existing I/O forwarding mechanisms for BG/P and identify the performance bottlenecks in the current design. We augment the I/O forwarding with two approaches: I/O scheduling using a work-queue model and asynchronous data staging. We evaluate the efficacy of our approaches using microbenchmarks and application-level benchmarks on leadership class systems.
Venkatram Vishwanath, Mark Hereld, Kamil Iskra, Dries Kimpe, Vitali A. Morozov, Michael E. Papka, Robert B. Ross, Kazutomo Yoshii
SC5