EDBT 2026 Demo / reviewers in the wild / expert
Jan Balewski
dblp:271/0984
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-1899-6526ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ML-Enabled FPGA Framework for Fast Quantum State Discrimination in Mid-Circuit Measurement RegimesabstractAccurate and low-latency quantum state discrimination is essential for protocols involving mid-circuit measurement (MCM) and conditional feed-forward. In superconducting quantum systems, conventional readout pipelines transfer measurement data to host processors for post-processing, introducing millisecond-scale delays that far exceed qubit coherence times. Neel Vora, Akel Hashim, Neelay Fruitwala, Noah Goss, Jan Balewski, K. Birgitta Whaley, Irfan Siddiqi, VP Nguyen |
ACM Great Lakes Symposium on VLSI | 7 |
| 2025 | Q-GEAR: Improving quantum simulation frameworkabstractThe rapid execution of complex quantum circuit simulations is essential for validating theoretical algorithms, thereby facilitating their successful implementation on quantum hardware. Although mainstream CPU-based platforms for circuit simulations are well established, they tend to be slower. Conversely, the adoption of GPU platforms remains limited because of the necessity for specialized quantum simulation frameworks tailored to different hardware architectures, each requiring distinct implementation and optimization strategies. Therefore, we introduced Q-Gear, a platform-agnostic framework that transforms Qiskit quantum circuits into Cuda-Q kernels. By leveraging Cuda-Q seamless execution on GPUs, Q-Gear accelerates both CPU- and GPU-based simulations by two orders of magnitude and ten times, respectively, with minimal coding effort. Furthermore, Q-Gear leverages the Cuda-Q configuration to interconnect the memory of GPUs, allowing the execution of much larger circuits beyond the memory limit set by a single GPU or CPU node. Additionally, we created and deployed a Podman container and Shifter image at Perlmutter (NERSC/LBNL), both derived from an NVIDIA public image. These public NERSC containers were optimized for the Slurm job scheduler, allowing approximately 100% utilization of up to 1,024 GPUs. We present various benchmarks for Q-Gear to demonstrate the efficiency of our computational paradigm. Ziqing Guo, Jan Balewski, Ziwen Pan |
ICPP | 2 |
| 2025 | First-principle crosstalk dynamics and Hamiltonian learning via Rabi experiments
Jan Balewski, Adam Winick, Neel Vora, David Santiago 0001, Joseph Emerson, Irfan Siddiqi |
MobiSys | 1 |
| 2023 | Automatic Qubit Characterization and Gate Optimization with QubiCabstractAs the size and complexity of a quantum computer increases, quantum bit (qubit) characterization and gate optimization become complex and time-consuming tasks. Current calibration techniques require complicated and verbose measurements to tune up qubits and gates, which cannot easily expand to the large-scale quantum systems. We develop a concise and automatic calibration protocol to characterize qubits and optimize gates using QubiC , which is an open source FPGA (field-programmable gate array)-based control and measurement system for superconducting quantum information processors. We propose multi-dimensional loss-based optimization of single-qubit gates and full XY-plane measurement method for the two-qubit CNOT gate calibration. We demonstrate the QubiC automatic calibration protocols are capable of delivering high-fidelity gates on the state-of-the-art transmon-type processor operating at the Advanced Quantum Testbed at Lawrence Berkeley National Laboratory. The single-qubit and two-qubit Clifford gate infidelities measured by randomized benchmarking are of 4.9(1.1) × 10 -4 and 1.4(3) × 10 -2 , respectively. Jan Balewski, Alexis Morvan, Kasra Nowrouzi, David Santiago 0001, Ravi Naik 0001, Brad Mitchell, Irfan Siddiqi |
ACM Trans. Quantum Comput. | 3 |
| 2022 | Using Multi-Resolution Data to Accelerate Neural Network Training in Scientific ApplicationsabstractNeural networks are powerful solutions to many scientific applications; however, they usually require long model training time due to large training data sets or large model size. Research has been focused on developing numerical optimization algorithms and parallel processing to reduce the training time. In this work, we propose a multi-resolution strategy that can reduce the training time by training the model with the reduced-resolution data samples at the beginning and later switching to the original resolution data samples. This strategy is motivated by the observation that coarser versions of many applications can be solved faster than their denser counterparts, and the solution to a coarser problem could be used to initialize the solution to the denser problem. When applying the idea to neural network training, coarse data can have a similar effect on the learning curves at the early stage as the dense data but requires less time. Once the curves no longer improve significantly, our strategy switches to using the data in original resolution. The key in this process is the ability to generate multiple resolutions of a problem automatically, which could usually be done with scientific applications with spatial and temporal continuity. We use two real-world scientific applications, CosmoFlow and DeepCAM, to evaluate the proposed mixed-resolution training strategy. Our experiment results demonstrate that the proposed training strategy effectively reduces the end-to-end training time while achieving a comparable accuracy to that of the training only with the original data. While maintaining the same model accuracy, our multi-resolution training strategy reduces the end-to-end training time up to 30% and 23% for CosmoFlow and DeepCAM, respectively. Kewei Wang 0002, Sunwoo Lee 0001, Jan Balewski, Alex Sim, Peter Nugent, Ankit Agrawal 0001, Alok N. Choudhary, Kesheng Wu, Wei-keng Liao |
CCGRID | 3 |
| 2021 | Asynchronous I/O Strategy for Large-Scale Deep Learning ApplicationsabstractMany scientific applications have started using deep learning methods for their classification or regression problems. However, for data-intensive scientific applications, I/O performance can be the major performance bottleneck. In order to effectively solve important real-world problems using deep learning methods on High-Performance Computing (HPC) systems, it is essential to address the poor I/O performance issue in large-scale neural network training. In this paper, we propose an asynchronous I/O strategy that can be generally applied to deep learning applications. Our I/O strategy employs an I/O -dedicated thread per process, that performs I/O operations independently of the training progress. The I/O thread reads many training samples at once to reduce the total number of I/O operations per epoch. Given the fixed amount of training data, the fewer the I/O operations per epoch, the shorter the overall I/O time. The I/O operations are also overlapped with the computations using the double-buffering method. We evaluate our I/O strategy using two real-world scientific applications, CosmoFlow and Neuron-Inverter. Our experimental results demonstrate that the proposed I/O strategy significantly improves the scaling performance without affecting the regression performance. Sunwoo Lee 0001, Qiao Kang, Kewei Wang 0002, Jan Balewski, Alex Sim, Ankit Agrawal 0001, Alok N. Choudhary, Peter Nugent, Kesheng Wu, Wei-keng Liao |
HiPC | 4 |
| 2021 | The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs With Hybrid ParallelismabstractWe present scalable hybrid-parallel algorithms for training large-scale 3D convolutional neural networks. Deep learning-based emerging scientific workflows often require model training with large, high-dimensional samples, which can make training much more costly and even infeasible due to excessive memory usage. We solve these challenges by extensively applying hybrid parallelism throughout the end-to-end training pipeline, including both computations and I/O. Our hybrid-parallel algorithm extends the standard data parallelism with spatial parallelism, which partitions a single sample in the spatial domain, realizing strong scaling beyond the mini-batch dimension with a larger aggregated memory capacity. We evaluate our proposed training algorithms with two challenging 3D CNNs, CosmoFlow and 3D U-Net. Our comprehensive performance studies show that good weak and strong scaling can be achieved for both networks using up to 2K GPUs. More importantly, we enable training of CosmoFlow with much larger samples than previously possible, realizing an order-of-magnitude improvement in prediction accuracy. Yosuke Oyama, Naoya Maruyama, Nikoli Dryden, Erin McCarthy, Peter Harrington, Jan Balewski, Satoshi Matsuoka, Peter Nugent, Brian Van Essen |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2020 | Atmospheric Blocking Pattern Recognition in Global Climate Model Simulation DataabstractIn this paper, we address a problem of atmospheric blocking pattern recognition in global climate model simulation data. Understanding blocking events is a crucial problem to society and natural infrastructure, as they often lead to weather extremes, such as heat waves, heavy precipitation, and the unusually poor air condition. Moreover, it is very challenging to detect these events as there is no physics-based model of blocking dynamic development that could account for their spatiotemporal characteristics. Here, we propose a new two-stage hierarchical pattern recognition method for detection and localisation of atmospheric blocking events in different regions over the globe. For both the detection stage and localisation stage, we train five different architectures of a convolutional neural network (CNN) based classifier and regressor. The results show the general pattern of the atmospheric blocking detection performance increasing significantly for the deep CNN architectures. In contrast, we see the estimation error of event location decreasing significantly in the localisation problem for the shallow CNN architectures. We demonstrate that CNN architectures tend to achieve the highest accuracy for blocking event detection and the lowest estimation error of event localisation in regions of the Northern Hemisphere than in regions of the Southern Hemisphere. Grzegorz Muszynski, Prabhat, Jan Balewski, Karthik Kashinath, Michael F. Wehner, Vitaliy Kurlin |
ICPR | 3 |