EDBT 2026 Demo / reviewers in the wild / expert
Troy Arcomano
dblp:258/5097
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-9359-6020ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Generative modeling · 74% Deep learning architectures and training · 20% Time series and sequential data · 6% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Environmental and earth informatics · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 51% Cloud and datacenter computing · 39% High-performance computing · 10% |
Topics — the 11 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.7 | 2 | 2025 | AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions · SC 2025 OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales · NeurIPS 2025 |
Environmental and earth informatics › weather forecasting
data-driven weather forecasting |
1.6 | 2 | 2025 | OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales · NeurIPS 2025 Scaling transformer neural networks for skillful and reliable medium-range weather forecasting · NeurIPS 2024 |
Environmental and earth informatics
weather forecasting |
1.6 | 2 | 2025 | OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales · NeurIPS 2025 Scaling transformer neural networks for skillful and reliable medium-range weather forecasting · NeurIPS 2024 |
Cloud and datacenter computing › configuration tuning
configuration auto-tuning |
1.0 | 1 | 2026 | GLANCED-IO: Taming I/O Optimization for Deep Learning at Scale · HPDC 2026 |
Storage systems
i/o optimization |
1.0 | 1 | 2026 | GLANCED-IO: Taming I/O Optimization for Deep Learning at Scale · HPDC 2026 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.9 | 1 | 2025 | OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Scaling transformer neural networks for skillful and reliable medium-range weather forecasting · NeurIPS 2024 |
Storage systems › file systems › distributed file system
parallel file system |
0.3 | 1 | 2026 | GLANCED-IO: Taming I/O Optimization for Deep Learning at Scale · HPDC 2026 |
Machine learning › Generative modeling
variational autoencoder |
0.3 | 1 | 2025 | OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time Scales · NeurIPS 2025 |
High-performance computing › large-scale training
large-scale distributed training |
0.3 | 1 | 2025 | AERIS: Argonne Earth Systems Model for Reliable and Skillful Predictions · SC 2025 |
Machine learning › Time series and sequential data › time series modeling
probabilistic forecasting |
0.2 | 1 | 2024 | Scaling transformer neural networks for skillful and reliable medium-range weather forecasting · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
window parallelism · 1.7swin diffusion transformer · 1.7sequence parallelism · 1.7pipeline parallelism · 1.7per-token diffusion head · 1.7masked latent diffusion · 1.7iterative unmasking · 1.7weather-specific embedding · 1.5randomized dynamics forecast · 1.5pressure-weighted loss · 1.5one-factor-at-a-time greedy exploration · 1.0auto-tuning · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GLANCED-IO: Taming I/O Optimization for Deep Learning at ScaleabstractScientific deep learning (DL) at scale typically trains on terabyte-scale datasets across thousands of accelerators, placing immense pressure on storage systems to keep pace with computation. Existing solutions respond to this demand by tuning individual I/O parameters to accelerate training performance. However, these techniques are limited by costly experiments, configuration space explosion, and inability to generalize application-specific optimizations. This leads to applications running with suboptimal configurations that reduce training efficiency, system utilization, or both. To address the challenge of finding the optimal configuration efficiently, we developed GLANCED-IO, a cross-layer I/O optimization framework that optimizes DL pipelines with high-fidelity approximation and efficient configuration space exploration. Through this work, we identified the following three key findings. First, independently optimizing either the application or system configurations leaves up to 2.4 × performance on the table for scientists to efficiently run DL pipelines on HPC systems. Second, GLANCED-IO’s one-factor-at-a-time (OFAT)-guided greedy exploration strategy achieved results comparable to more-expensive autotuning techniques while removing the pre-training required by ML-based approaches. Third, GLANCED-IO avoids executing the full application during optimization by operating on representative data subsets without GPUs, yet preserves 93% performance fidelity on average when deployed in DL pipelines. We demonstrate the efficacy of GLANCED-IO by optimizing large-scale global weather forecasting DL workloads, achieving up to 1.57 × better performance than state-of-the-art with 2.3 × fewer configuration evaluations than AIIO and 3.3 × faster optimization than DeepHyper. Ray A. O. Sinurat, William Nixon, Philip H. Carns, Huihuo Zheng, Sandeep Madireddy, Sam Foreman, Troy Arcomano, Robert B. Ross, Haryadi S. Gunawi, Hariharan Devarajan |
HPDC | 7 |
| 2025 | OmniCast: A Masked Latent Diffusion Model for Weather Forecasting Across Time ScalesabstractAccurate weather forecasting across time scales is critical for anticipating and mitigating the impacts of climate change. Recent data-driven methods based on deep learning have achieved significant success in the medium range, but struggle at longer subseasonal-to-seasonal (S2S) horizons due to error accumulation in their autoregressive approach. In this work, we propose OmniCast, a scalable and skillful probabilistic model that unifies weather forecasting across timescales. OmniCast consists of two components: a VAE model that encodes raw weather data into a continuous, lower-dimensional latent space, and a diffusion-based transformer model that generates a sequence of future latent tokens given the initial conditioning tokens. During training, we mask random future tokens and train the transformer to estimate their distribution given conditioning and visible tokens using a per-token diffusion head. During inference, the transformer generates the full sequence of future tokens by iteratively unmasking random subsets of tokens. This joint sampling across space and time mitigates compounding errors from autoregressive approaches. The low-dimensional latent space enables modeling long sequences of future latent states, allowing the transformer to learn weather dynamics beyond initial conditions. OmniCast performs competitively with leading probabilistic methods at the medium-range timescale while being 10× to 20× faster, and achieves state-of-the-art performance at the subseasonal-to-seasonal scale across accuracy, physics-based, and probabilistic metrics. Furthermore, we demonstrate that OmniCast can generate stable rollouts up to 100 years ahead. Code and model checkpoints are available at https://github.com/tung-nd/omnicast. Troy Arcomano, Rao Kotamarthi, Ian T. Foster, Sandeep Madireddy, Aditya Grover |
NeurIPS | 3 |
| 2025 | AERIS: Argonne Earth Systems Model for Reliable and Skillful PredictionsabstractGenerative machine learning offers new opportunities to better understand complex Earth system dynamics. Recent diffusion-based methods address spectral biases and improve ensemble calibration in weather forecasting compared to deterministic methods, yet have so far proven difficult to scale stably at high resolutions. We introduce AERIS, a 1.3 to 80B parameter pixel-level Swin diffusion transformer to address this gap, and SWiPe, a generalizable technique that composes window parallelism with sequence and pipeline parallelism to shard window-based transformers without added communication cost or increased global batch size. On Aurora (10,080 nodes), AERIS sustains 10.21 ExaFLOPS (mixed precision) and a peak performance of 11.21 ExaFLOPS with 1 × 1 patch size on the 0.25° ERA5 dataset, achieving 95.5% weak scaling efficiency, and 81.6% strong scaling efficiency. AERIS outperforms the IFS ENS and remains stable on seasonal scales to 90 days, highlighting the potential of billion-parameter diffusion models for weather and climate prediction. Väinö Hatanpää, Eugene Ku, Jason Stock, Murali Emani, Sam Foreman, Chunyong Jung, Sandeep Madireddy, Varuni Sastry 0001, Ray A. O. Sinurat, Huihuo Zheng, Sam Wheeler, Troy Arcomano, Venkatram Vishwanath, Rao Kotamarthi |
SC | 13 |
| 2024 | Scaling transformer neural networks for skillful and reliable medium-range weather forecastingabstractWeather forecasting is a fundamental problem for anticipating and mitigating the impacts of climate change. Recently, data-driven approaches for weather forecasting based on deep learning have shown great promise, achieving accuracies that are competitive with operational systems. However, those methods often employ complex, customized architectures without sufficient ablation analysis, making it difficult to understand what truly contributes to their success. Here we introduce Stormer, a simple transformer model that achieves state-of-the art performance on weather forecasting with minimal changes to the standard transformer backbone. We identify the key components of Stormer through careful empirical analyses, including weather-specific embedding, randomized dynamics forecast, and pressure-weighted loss. At the core of Stormer is a randomized forecasting objective that trains the model to forecast the weather dynamics over varying time intervals. During inference, this allows us to produce multiple forecasts for a target lead time and combine them to obtain better forecast accuracy. On WeatherBench 2, Stormer performs competitively at short to medium-range forecasts and outperforms current methods beyond 7 days, while requiring orders-of-magnitude less training data and compute. Additionally, we demonstrate Stormer’s favorable scaling properties, showing consistent improvements in forecast accuracy with increases in model size and training tokens. Code and checkpoints are available at https://github.com/tung-nd/stormer. Rohan Shah, Hritik Bansal, Troy Arcomano, Romit Maulik, Rao Kotamarthi, Ian T. Foster, Sandeep Madireddy, Aditya Grover |
NeurIPS | 4 |