VLDB 2026 Research / reviewers in the wild / expert
Dongho Ha
dblp:290/4137
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0009-0005-4090-4025ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-author · 12 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Carbon-Aware Continuous Learning for Sustainable Real-Time Machine Learning AnalyticsabstractReal-time machine learning (ML) analytics models deployed on edge servers often experience degraded inference accuracy due to data drift. Continuous learning mitigates this by periodically retraining models using newly collected data. However, retraining incurs significant computational overhead, increasing energy consumption and carbon footprint several-fold. Existing approaches for reducing carbon footprint predominantly focus on general workload scheduling strategies, such as shifting jobs to periods or regions with lower carbon intensity. However, they neglect continuous-learning-specific parameters like data labeling model, retraining threshold, and retraining hyperparameters. Consequently, these approaches miss opportunities to further reduce the carbon footprint and enhance inference accuracy in continuous learning systems under dynamic data drift. In this paper, we propose a novel Carbon-footprint-aware Continuous Learning (CCL) scheme that minimizes carbon emissions during model retraining without sacrificing inference accuracy. Distinct from prior workload scheduling approaches, CCL adaptively adjusts labeling model, retraining threshold, and retraining hyperparameters based on predictive models that estimate data drift severity and carbon intensity dynamics. Our adaptive real-time optimization approach consistently achieves a near-optimal balance between accuracy and carbon footprint under dynamic conditions. Experimental results demonstrate that CCL reduces operational carbon footprint by up to 68.5% compared to state-of-the-art carbon-agnostic methods, with negligible accuracy degradation. Gwanjong Park, Dongho Ha, Myeongjae Jeon, Euiseong Seo |
EuroSys | 3 |
| 2026 | Reducing Page Faults via Invalidation-Based Mapping Propagation in Multi-GPU Systems
Junsung Kim, Dongho Ha, Sungwoo Kim 0003, Wonho Cho, Sungbin Kim, Yufei Ding, Won Woo Ro |
ISCA | 2 |
| 2026 | MXFFP: Microscaling Flexible Floating Point Format for Large-Scale AI Model Acceleration
Sungwoo Kim 0003, Sungbin Kim, Dongho Ha, Hyunwuk Lee, Junsung Kim, Mingu Jung, Murali Annavaram, Won Woo Ro |
ISCA | 3 |
| 2026 | DeSpa: Heterogeneous multi-core accelerators for energy-efficient dense and sparse computation at the tile level in Deep Neural Networks
Hyungjun Jang, Dongho Ha, Hyunwuk Lee, Won Woo Ro |
J. Syst. Archit. | 2 |
| 2025 | Effective Interplay between Sparsity and Quantization: From Theory to PracticeabstractThe increasing size of deep neural networks (DNNs) necessitates effective model compression to reduce their computational and memory footprints. Sparsity and quantization are two prominent compression methods that have been shown to reduce DNNs' computational and memory footprints significantly while preserving model accuracy. However, how these two methods interact when combined together remains a key question for developers, as many tacitly assume that they are orthogonal, meaning that their combined use does not introduce additional errors beyond those introduced by each method independently. In this paper, we provide the first mathematical proof that sparsity and quantization are non-orthogonal. We corroborate these results with experiments spanning a range of large language models, including the OPT and LLaMA model families (with 125M to 8B parameters), and vision models like ViT and ResNet. We show that the order in which we apply these methods matters because applying quantization before sparsity may disrupt the relative importance of tensor elements, which may inadvertently remove significant elements from a tensor. More importantly, we show that even if applied in the correct order, the compounded errors from sparsity and quantization can significantly harm accuracy. Our findings extend to the efficient deployment of large models in resource-constrained compute platforms to reduce serving cost, offering insights into best practices for applying these compression methods to maximize hardware resource efficiency without compromising accuracy. Simla Burcu Harma, Ayan Chakraborty 0005, Elizaveta Kostenok, Danila Mishin, Dongho Ha, Babak Falsafi, Martin Jaggi, Yunho Oh, Suvinay Subramanian, Amir Yazdanbakhsh |
ICLR | 5 |
| 2025 | Avant-Garde: Empowering GPUs with Scaled Numeric FormatsabstractThe escalating computational and memory demands of deep neural networks have outpaced chip density improvements, making arithmetic density a key bottleneck for GPUs.Scaled numeric formats, such as FP8 and Microscaling (MX), improve arithmetic density by applying adaptive scaling factors across varying block sizes and multiple scaling hierarchies.Unfortunately, supporting diverse scaled numeric formats often requires GPUs to rely on softwarebased implementations, increasing instruction and register overhead and degrading performance.We propose Avant-Garde, a GPU microarchitecture that natively supports diverse scaled numeric formats by converting them into a consistent single-level internal representation.Avant-Garde integrates an Operand Transformer, a hardware module that dynamically flattens multi-level scaling formats into single-level internal representations, a novel Tensor Core, and an optimized data layout to eliminate instruction and register overhead.Our evaluations show that Avant-Garde achieves up to 74% higher throughput and 44% lower execution time, while maintaining accuracy within 0.2% compared to conventional GPUs. Minseong Gil, Dongho Ha, Simla Burcu Harma, Myung Kuk Yoon, Babak Falsafi, Won Woo Ro, Yunho Oh |
ISCA | 2 |
| 2025 | BitL: A Hybrid Bit-Serial and Parallel Deep Learning Accelerator for Critical Path ReductionabstractAs deep neural networks (DNNs) advance, their computational demands have grown immensely.In this context, previous research introduced bit-wise computation to enhance silicon efficiency, along with skipping unnecessary zero-bit calculations.However, we observe that existing bit-wise approaches miss an opportunity to optimize the critical computation path, as they process groups of values sequentially from the most significant bits (MSBs) to the least significant bits (LSBs).To address this limitation, we propose BitL, a novel bit-wise computing unit designed to minimize the critical path and improve the throughput.BitL dynamically switches between horizontal and vertical data lookups across sub-tiles during Multiply-Accumulate (MAC) operations.Additionally, it presents an innovative optimization technique to maximize the utilization of computing units while switching its lookup direction.Our evaluation demonstrates that BitL delivers up to 1.92× higher throughput compared to a baseline DNN accelerator and achieves a 1.24× improvement over recent zero-bit skipping accelerators.Furthermore, BitL improves energy efficiency by 2.06× on average, with a silicon area overhead of only 5.71%. Seunghyun Lee 0003, Dongho Ha, Sungbin Kim, Sungwoo Kim 0003, Hyunwuk Lee, Won Woo Ro |
MICRO | 2 |
| 2024 | Recompiling QAOA Circuits on Various Rotational DirectionsabstractThe quantum approximate optimization algorithm (QAOA) is introduced to efficiently solve combinatorial optimization problems. Despite the promise of QAOA, the cost of executing QAOA circuits at scale for quantum advantage may still be excessive for the near-future quantum device. We observe the increasing overhead of QAOA circuit execution in the native gate translation. To execute QAOA circuits on a real quantum computing device, Hamiltonians composed of predefined specific rotations (e.g., ZZ and X) should be decomposed into finite native gates. By adopting rotational combinations that utilize native gates more directly than the standard QAOA circuit model, the execution cost on real quantum devices can be reduced. In this study, we propose Racoon (Rotational Space Virtualization for QAOA Ansatz), an algorithm-hardware co-design approach that revisits the synthesis conditions of QAOA circuits and selects alternative candidates with different rotational combinations. Our analysis of six commercial quantum processors demonstrates that applying Racoon to QAOA circuits for the 4-node Sherrington-Kirkpatrick model reduces the number of native gates by an average of 23% and up to 79%. Consequently, using Racoon results in 43% fewer training epochs, 41% lower training energy consumption, and a 6% improvement in inference on average compared to standard QAOA. Racoon consistently reduces circuit depth as the number of qubits and layers increases, achieving 123 × more circuit depth reduction compared to the recently proposed Depth First Search (DFS)-based method. Furthermore, we confirm that Racoon’s method can be extended to State-of-The-Art QAOAs with modified ansätze and to the variational quantum eigensolver (VQE). Enhyeok Jang, Dongho Ha, Seungwoo Choi 0001, Youngmin Kim 0005, Jaewon Kwon, Yongju Lee 0003, Sungwoo Ahn, Hyungseok Kim 0003, Won Woo Ro |
PACT | 2 |
| 2024 | Generalizing Ray Tracing Accelerators for Tree Traversals on GPUsabstractTree traversal is a fundamental operation in many applications, such as database indexing and physics simulations. Although tree traversals feature high parallelism, they are inherently divergent and irregular, leading to inefficient performance on GPUs. Tree traversals are also prevalent in ray tracing, which is executed on dedicated Ray-Tracing Accelerators (RTAs) in modern GPUs to mitigate inefficiencies such as control flow divergence and underutilization of memory bandwidth by irregular memory accesses. In this paper, we propose the Tree Traversal Accelerator (TTA) to replicate the success of RTAs in ray tracing for general tree traversal applications. TTAs extend RTAs to support tree structures and operations beyond those in ray tracing, such as B- Tree search and radius search algorithms, by modifying existing computing units. Despite TTAs' effectiveness, they still rely on fixed-function computations, making it challenging to support other tree-based applications such as N-Body simulation fully. Thus, we introduce TTA + as an alternative design, which modularizes the RTA computing units and makes them programmable, trading some efficiency for flexibility. With less than 1 % increase in RTA area, our proposals can achieve up to S.4x speedup for B-Tree search, 1.7x for N-Body simulation, and 1.2x for select ray-tracing applications. Dongho Ha, Lufei Liu 0001, Yuan-Hsi Chou, Seokjin Go, Won Woo Ro, Hung-Wei Tseng 0001, Tor M. Aamodt |
MICRO | 1 |
| 2024 | M3XU: Achieving High-Precision and Complex Matrix Multiplication with Low-Precision MXUsabstractBeyond the high-profile artificial intelligence and machine learning ($\mathrm{AI} / \mathrm{ML}$) workloads, the demand for high-performance matrix operations on standard and complex floating-point numbers remains strong but underserved. However, the widely adopted low-precision matrix processing units (MXUs) can only fulfill the need for AI/ML workloads, which are underutilized or idle when running applications outside their target domains. This paper presents $\mathbf{M}^{3} \mathbf{X U}$, multi-mode matrix processing units that support IEEE 754 single-precision and complex 32bit floating-point numbers. $\mathbf{M}^{3} \mathbf{X U}$ does not rely on more precise but costly multipliers. Instead, $\mathbf{M}^{3} \mathbf{X U}$ proposes a multi-step approach that extends existing MXUs for AI/ML workloads. The resulting $\mathbf{M}^{3} \mathbf{X U}$ can seamlessly upgrade existing systems without programmers’ efforts and maintain the bandwidth demand of existing memory subsystems. This paper evaluates $\mathbf{M}^{3} \mathbf{X U}$ with full-system emulation and hardware synthesis. $\mathrm{M}^{3} \mathbf{X U}$ can achieve a $3.64 \times$ speedup for 32 -bit matrix multiplications and $3.51 \times$ speedup for complex number operations on average compared with conventional vector processing units. Dongho Ha, Chen-Chien Kao, Christopher J. Hughes, Won Woo Ro, Hung-Wei Tseng 0001 |
SC | 1 |
| 2023 | R2D2: Removing ReDunDancy Utilizing Linearity of Address Generation in GPUsabstractA generally used GPU programming methodology is that adjacent threads access data in neighbor or specific-stride memory addresses and perform computations with the fetched data. This paper demonstrates that the memory addresses often exhibit a simple linear value pattern across GPU threads, as each thread uses built-in variables and constant values to compute the memory addresses. However, since the threads compute their context data individually, GPUs incur a heavy instruction overhead to calculate the memory addresses, even though they exhibit a simple pattern. We propose a GPU architecture called Removing ReDunDancy Utilizing Linearity of Address Generation (R2D2), reducing a large amount of the dynamic instruction count by detecting such linear patterns in the memory addresses and exploiting them for kernel computations. R2D2 detects linearities of the memory addresses with software support and pre-computes them before the threads execute the instructions. With the proposed scheme, each thread is able to compute its memory addresses with fewer dynamic instructions than conventional GPUs. In our evaluation, R2D2 achieves dynamic instruction reduction by 28%, 1.25x speedup, and energy consumption reduction by 17% over baseline GPU. Dongho Ha, Yunho Oh, Won Woo Ro |
ISCA | 1 |
| 2023 | TensorCV: Accelerating Inference-Adjacent Computation Using Tensor ProcessorsabstractThe advancements in AI/ML accelerators have made the core AI/ML computation relatively insignificant in application pipelines. For example, inferencing only accounts for 3% of the latency in an image-based ML pipeline with the help of Tensor Cores. The mismatch in performance growth between ML model computation and ML-adjacent computation, the producer and consumer of ML models, will become the bottleneck leading to system inefficiency. This paper presents a set of innovative algorithms to allow the entire ML-based computer vision pipelines to leverage AI/ML accelerators. Our proposed algorithms feature matrix-based operations that AI/ML accelerators specialize in. Simply compiler optimizations cannot take full advantage of hardware acceleration without revisiting algorithms. This paper implements the proposed algorithms as an open-source library, TensorCV, in a system platform with Tensor Cores. TensorCV shows a 6.12 × speedup in optimized ML-adjacent functions and saves 81 % energy consumption on modern heterogeneous computers. The code is available at https://github.com/escalab/TensorCV. Dongho Ha, Won Woo Ro, Hung-Wei Tseng 0001 |
ISLPED | 1 |
| 2023 | MAD MAcce: Supporting Multiply-Add Operations for Democratizing Matrix-Multiplication AcceleratorsabstractModern GPUs commonly employ specialized matrix multiplication units (MXUs) to accelerate matrix multiplication, the core computation of deep learning workloads. However, it is challenging to exploit the MXUs for GPGPU applications whose fundamental algorithms do not rely on matrix multiplication. Furthermore, an additional programming effort is necessary to tailor existing code or algorithms using dedicated APIs or libraries to utilize MXUs. Therefore, MXUs are often underutilized even when GPUs hunger for higher throughput. Seunghwan Sung, Sujin Hur, Sungwoo Kim 0003, Dongho Ha, Yunho Oh, Won Woo Ro |
MICRO | 4 |