EDBT 2026 Demo / reviewers in the wild / expert
Salman Habib 0002
dblp:23/6216-2
· DBLP profile ↗
10ranked-venue papers
2as first author
1since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
High-performance computing · 93% GPUs and heterogeneous computing · 7% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Computational science and engineering · 100% |
Topics — the 5 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
performance optimization at scale |
1.2 | 3 | 2025 | Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability · SC 2025 HACC: extreme scaling and performance across diverse architectures · SC 2013 The universe at extreme scale: multi-petaflop sky simulation on the BG/Q · SC 2012 |
High-performance computing
scientific computing systems |
1.2 | 3 | 2025 | Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability · SC 2025 HACC: extreme scaling and performance across diverse architectures · SC 2013 The universe at extreme scale: multi-petaflop sky simulation on the BG/Q · SC 2012 |
High-performance computing › large-scale simulation
exascale simulation |
0.9 | 1 | 2025 | Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability · SC 2025 |
High-performance computing › scientific data analysis
in-situ analysis |
0.5 | 2 | 2025 | Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in Capability · SC 2025 Large-scale compute-intensive analysis via a combined in-situ and co-scheduling workflow approach · SC 2015 |
High-performance computing › performance optimization at scale
extreme-scale scalability |
0.3 | 2 | 2013 | HACC: extreme scaling and performance across diverse architectures · SC 2013 The universe at extreme scale: multi-petaflop sky simulation on the BG/Q · SC 2012 |
Methods — techniques the papers use, named apart from their topics
tree solver · 0.9separation-of-scale · 0.9multi-tiered i/o · 0.9particle-grid methods · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cosmological Hydrodynamics at Exascale: A Trillion-Particle Leap in CapabilityabstractResolving the most fundamental questions in cosmology requires simulations that match the scale, fidelity, and physical complexity demanded by next-generation sky surveys. To achieve the realism needed for this critical scientific partnership, detailed gas dynamics must be treated self-consistently with gravity for end-to-end modeling of structure formation. Exascale computing enables simulations that span survey-scale volumes while incorporating key astrophysical processes that shape complex cosmic structures. We present results from CRK-HACC, a cosmological hydrodynamics code built for extreme scalability. Using separation-of-scale techniques, GPU-resident tree solvers, in situ analysis pipelines, and multi-tiered I/O, CRK-HACCexecuted Frontier-E: a four trillion particle full-sky simulation, over an order of magnitude larger than previous efforts. The run achieved 513.1 PFLOPs peak performance, processing 46.6 billion particles per second and writing more than 100 PB of data in just over one week of runtime. Frontier-E marks a significant advance in predictive modeling for next-generation cosmological science. Nicholas Frontiere, J. D. Emberson, Michael Buehlmann, Esteban Rangel, Salman Habib 0002, Katrin Heitmann, Patricia Larsen, Vitali A. Morozov, Adrian Pope, Claude-André Faucher-Giguère, Antigoni Georgiadou, Damien Lebrun-Grandié, Andrey Prokopenko |
SC | 5 |
| 2017 | Building Halo Merger Trees from the Q Continuum SimulationabstractCosmological N-body simulations rank among the most computationally intensive efforts today. A key challenge is the analysis of structure, substructure, and the merger history for many billions of compact particle clusters, called halos. Effectively representing the merging history of halos is essential for many galaxy formation models used to generate synthetic sky catalogs, an important application of modern cosmological simulations. Generating realistic mock catalogs requires computing the halo formation history from simulations with large volumes and billions of halos over many time steps, taking hundreds of terabytes of analysis data. We present fast parallel algorithms for producing halo merger trees and tracking halo substructure from a single-level, density-based clustering algorithm. Merger trees are created from analyzing the halo-particle membership function in adjacent snapshots, and substructure is identified by tracking the "cores" of merging halos – sets of particles near the halo center. Core tracking is performed after creating merger trees and uses the relationships found during tree construction to associate substructures with hosts. The algorithms are implemented with MPI and evaluated on a Cray XK7 supercomputer using up to 16,384 processes on data from HACC, a modern cosmological simulation framework. We present results for creating merger trees from 101 analysis snapshots taken from the Q Continuum, a large volume, high mass resolution, cosmological simulation evolving half a trillion particles. Esteban Rangel, Nicholas Frontiere, Salman Habib 0002, Katrin Heitmann, Wei-keng Liao, Ankit Agrawal 0001, Alok N. Choudhary |
HiPC | 3 |
| 2016 | An integrated visualization system for interactive analysis of large, heterogeneous cosmology dataabstractCosmological simulations produce a multitude of data types whose large scale makes them difficult to thoroughly explore in an interactive setting. One aspect of particular interest to scientists is the evolution of groups of dark matter particles, or "halos," described by merger trees. However, in order to fully understand subtleties in the merger trees, other data types derived from the simulation must be incorporated as well. In this work, we develop a novel interactive linked-view visualization system that focuses on simultaneously exploring dark matter halos, their hierarchical evolution, corresponding particle data, and other quantitative information. We employ a parallel remote renderer and a local merger tree selection tool so that users can analyze large data sets interactively. This allows scientists to assess their simulation code, understand inconsistencies in extracted data, and intuitively understand simulation behavior on all scales. We demonstrate the effectiveness of our system through a set of case studies on large-scale cosmological data from the HACC (Hardware/Hybrid Accelerated Cosmology Code) simulation framework. Annie Preston, Ramyar Ghods, Franz Sauer, Nick Leaf, Kwan-Liu Ma, Esteban Rangel, Eve Kovacs, Katrin Heitmann, Salman Habib 0002 |
PacificVis | 10 |
| 2016 | Parallel DTFE Surface Density Field ReconstructionabstractWe improve the interpolation accuracy and efficiency of the Delaunay tessellation field estimator (DTFE) for surface density field reconstruction by proposing an algorithm that takes advantage of the adaptive triangular mesh for line-of-sight integration. The costly computation of an intermediate 3D grid is completely avoided by our method and only optimally chosen interpolation points are computed, thus, the overall computational cost is significantly reduced. The algorithm is implemented as a parallel shared-memory kernel for large-scale grid rendered field reconstructions in our distributed-memory framework designed for N-body gravitational lensing simulations in large volumes. We also introduce a load balancing scheme to optimize the efficiency of processing a large number of field reconstructions. Our results show our kernel outperforms existing software packages for volume weighted density field reconstruction, achieving~10x speedup, and our load balancing algorithm gains an additional~3.6x speedup at scales with~16k processes. Esteban Rangel, Nan Li 0022, Salman Habib 0002, Tom Peterka, Ankit Agrawal 0001, Wei-keng Liao, Alok N. Choudhary |
CLUSTER | 3 |
| 2015 | Large-scale compute-intensive analysis via a combined in-situ and co-scheduling workflow approachabstractLarge-scale simulations can produce hundreds of terabytes to petabytes of data, complicating and limiting the efficiency of workflows. Traditionally, outputs are stored on the file system and analyzed in post-processing. With the rapidly increasing size and complexity of simulations, this approach faces an uncertain future. Trending techniques consist of performing the analysis in-situ, utilizing the same resources as the simulation, and/or off-loading subsets of the data to a compute-intensive analysis system. We introduce an analysis framework developed for HACC, a cosmological N-body code, that uses both in-situ and co-scheduling approaches for handling petabyte-scale outputs. We compare different analysis set-ups ranging from purely off-line, to purely in-situ to in-situ/co-scheduling. The analysis routines are implemented using the PISTON/VTK-m framework, allowing a single implementation of an algorithm that simultaneously targets a variety of GPU, multi-core, and many-core architectures. Christopher M. Sewell, Katrin Heitmann, Hal Finkel, George Zagaris, Suzanne Parete-Koon, Patricia K. Fasel, Adrian Pope, Nicholas Frontiere, Li-Ta Lo, O. E. Bronson Messer, Salman Habib 0002, James P. Ahrens |
SC | 11 |
| 2014 | Scalable Parallel I/O on a Blue Gene/Q Supercomputer Using Compression, Topology-Aware Data Aggregation, and SubfilingabstractIn this paper, we propose an approach to improving the I/O performance of an IBM Blue Gene/Q supercomputing system using a novel framework that can be integrated into high performance applications. We take advantage of the system's tremendous computing resources and high interconnection bandwidth among compute nodes to efficiently exploit I/O bandwidth. This approach focuses on lossless data compression, topology-aware data movement, and subfiling. The efficacy of this solution is demonstrated using microbenchmarks and an application-level benchmark. Huy Bui, Hal Finkel, Venkatram Vishwanath, Salman Habib 0002, Katrin Heitmann, Jason Leigh, Michael E. Papka, Kevin Harms |
PDP | 4 |
| 2013 | HACC: extreme scaling and performance across diverse architecturesabstractSupercomputing is evolving towards hybrid and accelerator-based architectures with millions of cores. The HACC (Hardware/Hybrid Accelerated Cosmology Code) framework exploits this diverse landscape at the largest scales of problem size, obtaining high scalability and sustained performance. Developed to satisfy the science requirements of cosmological surveys, HACC melds particle and grid methods using a novel algorithmic structure that flexibly maps across architectures, including CPU/GPU, multi/many-core, and Blue Gene systems. We demonstrate the success of HACC on two very different machines, the CPU/GPU system Titan and the BG/Q systems Sequoia and Mira, attaining unprecedented levels of scalable performance. We demonstrate strong and weak scaling on Titan, obtaining up to 99.2% parallel efficiency, evolving 1.1 trillion particles. On Sequoia, we reach 13.94 PFlops (69.2% of peak) and 90% parallel efficiency on 1,572,864 cores, with 3.6 trillion particles, the largest cosmological benchmark yet performed. HACC design concepts are applicable to several other supercomputer applications. Salman Habib 0002, Vitali A. Morozov, Nicholas Frontiere, Hal Finkel, Adrian Pope, Katrin Heitmann |
SC | 1 |
| 2012 | Analyzing the evolution of large scale structures in the universe with velocity based methodsabstractThe formation of cosmic structure results from the action of gravity on matter in an expanding Universe. As the evolution proceeds, the velocity field changes from being single-valued almost everywhere in space to being multi-valued over a complex web of `multistreaming' regions associated with the formation of large-scale structure (LSS) such as halos (or clumps), filaments, and sheets. Until recently, these structures have been investigated primarily via the (scalar) mass density field. In this application paper we apply data analysis and visualization techniques to cosmological simulations with the aim of studying multistreaming regions using velocity-based probes. Compared to the current practice of using density information (e.g., morphology estimators, locating overdense regions with halo finders), we show that velocity-based methods can provide useful supporting, as well as complementary, information. Because the density field and multistreaming are correlated but do not contain the same information, new and interesting information about the properties of the large-scale structure may be extracted, e.g., capturing dynamical behavior not possible with density-based estimators. Incorporating a novel method for setting thresholds for the velocity-based estimators, we study the relationships between the density field as represented by compact overdense halos and the different properties of multistreaming regions as represented by different velocity-based estimators. Uliana Popov, Eddy Chandra, Katrin Heitmann, Salman Habib 0002, James P. Ahrens, Alex T. Pang |
PacificVis | 4 |
| 2012 | The universe at extreme scale: multi-petaflop sky simulation on the BG/QabstractRemarkable observational advances have established a compelling cross-validated model of the Universe. Yet, two key pillars of this model -- dark matter and dark energy -- remain mysterious. Next-generation sky surveys will map billions of galaxies to explore the physics of the 'Dark Universe'. Science requirements for these surveys demand simulations at extreme scales; these will be delivered by the HACC (Hybrid/Hardware Accelerated Cosmology Code) framework. HACC's novel algorithmic structure allows tuning across diverse architectures, including accelerated and multi-core systems. On the IBM BG/Q, HACC attains unprecedented scalable performance - currently 6.23 PFlops at 62% of peak and 92% parallel efficiency on 786,432 cores (48 racks) - at extreme problem sizes with up to almost two trillion particles, larger than any cosmological simulation yet performed. HACC simulations at these scales will for the first time enable tracking individual galaxies over the entire volume of a cosmological survey. Salman Habib 0002, Vitali A. Morozov, Hal Finkel, Adrian Pope, Katrin Heitmann, Kalyan Kumaran, Tom Peterka, Joseph A. Insley, David Daniel, Patricia K. Fasel, Nicholas Frontiere, Zarija Lukic |
SC | 1 |
| 2011 | In-situ Sampling of a Large-Scale Particle Simulation for Interactive Visualization and AnalysisabstractAbstract We describe a simulation‐time random sampling of a large‐scale particle simulation, the RoadRunner Universe MC3cosmological simulation, for interactive post‐analysis and visualization. Simulation data generation rates will continue to be far greater than storage bandwidth rates by many orders of magnitude. This implies that only a very small fraction of data generated by a simulation can ever be stored and subsequently post‐analyzed. The limiting factors in this situation are similar to the problem in many population surveys: there aren't enough human resources to query a large population. To cope with the lack of resources, statistical sampling techniques are used to create a representative data set of a large population. Following this analogy, we propose to store a simulation‐time random sampling of the particle data for post‐analysis, with level‐of‐detail organization, to cope with the bottlenecks. A sample is stored directly from the simulation in a level‐of‐detail format for post‐visualization and analysis, which amortizes the cost of post‐processing and reduces workflow time. Additionally by sampling during the simulation, we are able to analyze the entire particle population to record full population statistics and quantify sample error. Jonathan Woodring, James P. Ahrens, J. Figg, Joanne Wendelberger, Salman Habib 0002, Katrin Heitmann |
Comput. Graph. Forum | 5 |