EDBT 2026 Demo / reviewers in the wild / expert
Sam Foreman
dblp:292/3878
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-9981-0876ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 37% High-performance computing · 29% Cloud and datacenter computing · 28% | |
| Artificial intelligence
1 paper |
Generative modeling · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Cloud and datacenter computing › configuration tuning
configuration auto-tuning |
1.0 | 1 | 2026 | GLANCED-IO: Taming I/O Optimization for Deep Learning at Scale · HPDC 2026 |
Storage systems
i/o optimization |
1.0 | 1 | 2026 | GLANCED-IO: Taming I/O Optimization for Deep Learning at Scale · HPDC 2026 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions · SC 2025 |
High-performance computing
scientific computing systems |
0.8 | 1 | 2024 | MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024 |
Storage systems › file systems › distributed file system
parallel file system |
0.3 | 1 | 2026 | GLANCED-IO: Taming I/O Optimization for Deep Learning at Scale · HPDC 2026 |
High-performance computing › large-scale training
large-scale distributed training |
0.3 | 1 | 2025 | AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions · SC 2025 |
GPUs and heterogeneous computing
GPU and heterogeneous computing |
0.2 | 1 | 2024 | MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024 |
Methods — techniques the papers use, named apart from their topics
window parallelism · 1.7swin diffusion transformer · 1.7sequence parallelism · 1.7pipeline parallelism · 1.7one-factor-at-a-time greedy exploration · 1.0auto-tuning · 1.0multimodal generative models · 0.8mixed precision · 0.8direct preference optimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GLANCED-IO: Taming I/O Optimization for Deep Learning at ScaleabstractScientific deep learning (DL) at scale typically trains on terabyte-scale datasets across thousands of accelerators, placing immense pressure on storage systems to keep pace with computation. Existing solutions respond to this demand by tuning individual I/O parameters to accelerate training performance. However, these techniques are limited by costly experiments, configuration space explosion, and inability to generalize application-specific optimizations. This leads to applications running with suboptimal configurations that reduce training efficiency, system utilization, or both. To address the challenge of finding the optimal configuration efficiently, we developed GLANCED-IO, a cross-layer I/O optimization framework that optimizes DL pipelines with high-fidelity approximation and efficient configuration space exploration. Through this work, we identified the following three key findings. First, independently optimizing either the application or system configurations leaves up to 2.4 × performance on the table for scientists to efficiently run DL pipelines on HPC systems. Second, GLANCED-IO’s one-factor-at-a-time (OFAT)-guided greedy exploration strategy achieved results comparable to more-expensive autotuning techniques while removing the pre-training required by ML-based approaches. Third, GLANCED-IO avoids executing the full application during optimization by operating on representative data subsets without GPUs, yet preserves 93% performance fidelity on average when deployed in DL pipelines. We demonstrate the efficacy of GLANCED-IO by optimizing large-scale global weather forecasting DL workloads, achieving up to 1.57 × better performance than state-of-the-art with 2.3 × fewer configuration evaluations than AIIO and 3.3 × faster optimization than DeepHyper. Ray A. O. Sinurat, William Nixon, Philip H. Carns, Huihuo Zheng, Sandeep Madireddy, Sam Foreman, Troy Arcomano, Robert B. Ross, Haryadi S. Gunawi, Hariharan Devarajan |
HPDC | 6 |
| 2025 | AERIS: Argonne Earth Systems Model for Reliable and Skillful PredictionsabstractGenerative machine learning offers new opportunities to better understand complex Earth system dynamics. Recent diffusion-based methods address spectral biases and improve ensemble calibration in weather forecasting compared to deterministic methods, yet have so far proven difficult to scale stably at high resolutions. We introduce AERIS, a 1.3 to 80B parameter pixel-level Swin diffusion transformer to address this gap, and SWiPe, a generalizable technique that composes window parallelism with sequence and pipeline parallelism to shard window-based transformers without added communication cost or increased global batch size. On Aurora (10,080 nodes), AERIS sustains 10.21 ExaFLOPS (mixed precision) and a peak performance of 11.21 ExaFLOPS with 1 × 1 patch size on the 0.25° ERA5 dataset, achieving 95.5% weak scaling efficiency, and 81.6% strong scaling efficiency. AERIS outperforms the IFS ENS and remains stable on seasonal scales to 90 days, highlighting the potential of billion-parameter diffusion models for weather and climate prediction. Väinö Hatanpää, Eugene Ku, Jason Stock, Murali Emani, Sam Foreman, Chunyong Jung, Sandeep Madireddy, Varuni Sastry 0001, Ray A. O. Sinurat, Huihuo Zheng, Sam Wheeler, Troy Arcomano, Venkatram Vishwanath, Rao Kotamarthi |
SC | 5 |
| 2024 | MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference OptimizationabstractWe present a scalable, end-to-end workflow for protein design. By augmenting protein sequences with natural language descriptions of their biochemical properties, we train generative models that can be preferentially aligned with protein fitness landscapes. Through complex experimental-and simulation-based observations, we integrate these measures as preferred parameters for generating new protein variants and demonstrate our workflow on five diverse supercomputers. We achieve >1 ExaFLOPS sustained performance in mixed precision on each supercomputer and a maximum sustained performance of 4.11 Ex-aFLOPS and peak performance of 5.57 ExaFLOPS. We establish the scientific performance of our model on two tasks: (1) across a predetermined benchmark dataset of deep mutational scanning experiments to optimize the fitness-determining mutations in the yeast protein HIS7, and (2) in optimizing the design of the enzyme malate dehydrogenase to achieve lower activation barriers (and therefore increased catalytic rates) using simulation data. Our implementation thus sets high watermarks for multimodal protein design workflows. Gautham Dharuman, Kyle Hippe, Alex Brace, Sam Foreman, Väinö Hatanpää, Varuni Sastry 0001, Huihuo Zheng, Logan T. Ward, Servesh Muralidharan, Archit Vasan, Bharat Kale, Carla M. Mann, Yun-Hsuan Cheng, Yuliana Zamora, Shengchao Liu, Chaowei Xiao, Murali Emani, Tom Gibbs, Mahidhar Tatineni, Deepak Canchi, Jerome Mitchell, Koichi Yamada, María Jesús Garzarán, Michael E. Papka, Ian T. Foster, Rick L. Stevens, Anima Anandkumar, Venkatram Vishwanath, Arvind Ramanathan |
SC | 4 |