EDBT 2026 Demo / reviewers in the wild / expert
Amarjit Singh
dblp:14/3064
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-7081-7329ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Eliminating Python Overhead in Predictive Neural Compression: A Native C++/LibTorch Implementation of TEZipabstractSignificant data reduction is essential for extreme-scale scientific facilities, where massive spatiotemporal datasets overwhelm traditional parallel file systems and create severe I/O bottlenecks. To mitigate this, emerging hybrid HPC workloads integrate Deep Neural Network (DNN) models into scientific workflows to perform predictive data compression. TEZip (Time Evolutionary Zip) uses these DNN-based methods to predict sequences and reduce storage footprints. However, integrating dynamic Python-based ML frameworks (e.g., PyTorch) into static HPC environments introduces severe runtime and data movement overheads, preventing these hybrid applications from keeping pace with node-local data generation rates. In this work, we present the first native C++/LibTorch implementation of TEZip, an architectural co-design explicitly targeting the HPC I/O critical path. We have updated the prediction module to use advanced deep learning architectures, including ConvLSTM and PredNet, to process spatiotemporal data. Rather than a simple language translation, we redesign the tensor lifecycle management and prediction loops to embed this inference directly into the I/O stream. We identify and quantify Python runtime overhead sources unique to these neural compression workloads, and expose non-trivial design challenges in embedding LibTorch inference into parallel I/O pipelines. This native C++ architecture achieves approximately a 4 × speedup in training, a 13 × speedup in compression, and a 4.8 × speedup in decompression in three datasets, while maintaining the identical compression ratio and reconstruction quality. These improvements allow AI-driven predictive compression to run efficiently at the edge of the compute tier, significantly reducing data volume before it affects parallel file system bandwidth. Mina Yousef, Amarjit Singh, Kento Sato |
HPDC | 2 |
| 2025 | Refactoring TEZip: Integrating Python-Based Predictive Compression into an HPC C++/LibTorch EnvironmentabstractTEZip is a framework for compressing time-evolving image data using predictive deep neural networks. Until now, TEZip primarily relied on Python libraries (TensorFlow or PyTorch). This work presents a new TEZip pipeline built with C++/LibTorch for improved speed and High preformance Computer(HPC) compatibility. Mina Yousef, Amarjit Singh, Kento Sato |
HPDC | 2 |
| 2025 | What to Support When You're Compressing: The State of Practice Gaps and Opportunities for Scientific Data CompressionabstractOver the last nearly 20 years, lossy compression has become an essential aspect of HPC applications’ data pipelines, allowing them to overcome limitations in storage capacity and bandwidth and, in some cases, increase computational throughput and capacity. However, with the adoption of lossy compression comes the requirement to assess and control the impact lossy compression has on scientific outcomes. In this work, we take a major step forward in describing the state of practice and by characterizing workloads. We examine applications’ needs and compressors’ capabilities across 9 different supercomputing application domains. We present 24 takeaways that provide best practices for applications, operational impacts for facilities achieving compressed data, and gaps in application needs not addressed by production compressors that point towards opportunities for future compression research. Franck Cappello, Robert Underwood, Yuri Alexeev, Allison H. Baker, Ebru Bozdag, Martin Burtscher, Kyle Chard, Sheng Di, Kyle Gerard Felker, Paul Christopher O'Grady, Hanqi Guo 0001, Yafan Huang, Peng Jiang 0004, Sian Jin, Petter Johansson, Shaomeng Li, Xin Liang 0001, Erik Lindahl, Peter Lindstrom 0001, Zarija Lukic, Magnus Lundborg, Danylo Lykov, Masaru Nagaso, Kento Sato, Amarjit Singh, Seung Woo Son 0001, Shihui Song, William Tang 0002, Dingwen Tao, Jiannan Tian, Kazutomo Yoshii, Kai Zhao 0008 |
SC | 25 |
| 2025 | Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing
Franck Cappello, Mario C. Acosta, Emmanuel Agullo, Hartwig Anzt, Jon Calhoun 0001, Sheng Di, Luc Giraud, Thomas Grützmacher, Sian Jin, Kentaro Sano, Kento Sato, Amarjit Singh, Dingwen Tao, Jiannan Tian, Tomohiro Ueno, Robert Underwood, Frédéric Vivien, Xavier Yepes, Kazutomo Yoshii, Boyuan Zhang 0002 |
Future Gener. Comput. Syst. | 12 |