Wonho Cho

dblp:362/9007 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 51% Generative modeling · 49%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.522025
Ditto: Accelerating Diffusion Model via Temporal Value Similarity · HPCA 2025
Exploiting Inherent Properties of Complex Numbers for Accelerating Complex Valued Neural Networks · MICRO 2023
Machine learning › Efficient and distributed learning
model compression
0.922025
Exploiting Inherent Properties of Complex Numbers for Accelerating Complex Valued Neural Networks · MICRO 2023
Ditto: Accelerating Diffusion Model via Temporal Value Similarity · HPCA 2025
Machine learning › Efficient and distributed learning › model compression
quantization
0.922025
Exploiting Inherent Properties of Complex Numbers for Accelerating Complex Valued Neural Networks · MICRO 2023
Ditto: Accelerating Diffusion Model via Temporal Value Similarity · HPCA 2025
Machine learning › Generative modeling
diffusion model
0.912025
Ditto: Accelerating Diffusion Model via Temporal Value Similarity · HPCA 2025
Machine learning › Generative modeling › diffusion model › diffusion model acceleration
diffusion model quantization
0.912025
Ditto: Accelerating Diffusion Model via Temporal Value Similarity · HPCA 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
diffusion model accelerator
0.912025
Ditto: Accelerating Diffusion Model via Temporal Value Similarity · HPCA 2025
Hardware accelerators and domain-specific architectures
systolic array
0.212023
Exploiting Inherent Properties of Complex Numbers for Accelerating Complex Valued Neural Networks · MICRO 2023

Methods — techniques the papers use, named apart from their topics

temporal difference processing · 1.7quantization · 1.7polar form quantization · 1.3CVNN-aware scheduling · 1.3
YearPublicationVenuePosition
2026 Reducing Page Faults via Invalidation-Based Mapping Propagation in Multi-GPU Systems
Junsung Kim, Dongho Ha, Sungwoo Kim 0003, Wonho Cho, Sungbin Kim, Yufei Ding, Won Woo Ro
ISCA4
2025 Ditto: Accelerating Diffusion Model via Temporal Value Similarity
abstract
Diffusion models achieve superior performance in image generation tasks. However, it incurs significant computation overheads due to its iterative structure. To address these overheads, we analyze this iterative structure and observe that adjacent time steps in diffusion models exhibit high value similarity, leading to narrower differences between consecutive time steps. We adapt these characteristics to a quantized diffusion model and reveal that the majority of these differences can be represented with reduced bit-width, and even zero. Based on our observations, we propose the Ditto algorithm, a difference processing algorithm that leverages temporal similarity with quantization to enhance the efficiency of diffusion models. By exploiting the narrower differences and the distributive property of layer operations, it performs full bit-width operations for the initial time step and processes subsequent steps with temporal differences. In addition, Ditto execution flow optimization is designed to mitigate the memory overhead of temporal difference processing, further boosting the efficiency of the Ditto algorithm. We also design the Ditto hardware, a specialized hardware accelerator, fully exploiting the dynamic characteristics of the proposed algorithm. As a result, the Ditto hardware achieves up to $1.5 \times$ speedup and 17.74% energy saving compared to other accelerators.
Sungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park, Won Woo Ro
HPCA3
2023 Exploiting Inherent Properties of Complex Numbers for Accelerating Complex Valued Neural Networks
abstract
Since conventional Deep Neural Networks (DNNs) use real numbers as their data, they are unable to capture the imaginary values and the correlations between real and imaginary values in applications that use complex numbers. To address this limitation, Complex Valued Neural Networks (CVNNs) have been introduced, enabling to capture the context of complex numbers for various applications such as Magnetic Resonance Imaging (MRI), radar, and sensing. CVNNs handle their data with complex numbers and adopt complex number arithmetic to their layer operations, so they exhibit distinct design challenges with real-valued DNNs. The first challenge is the data representation of the complex number, which requires two values for a single data, doubling the total data size of the networks. Moreover, due to the unique operations of the complex-valued layers, CVNNs require a specialized scheduling policy to fully utilize the hardware resources and achieve optimal performance. To mitigate the design challenges, we propose software and hardware co-design techniques that effectively resolves the memory and compute overhead of CVNNs. First, we propose Polar Form Aware Quantization (PAQ) that utilizes the characteristics of the complex number and their unique value distribution on CVNNs. Then, we propose our hardware accelerator that supports PAQ and CVNN operations. Lastly, we design a CVNN-aware scheduling scheme that optimizes the performance and resource utilization of an accelerator by aiming at the special layer operations of CVNN. PAQ achieves 62.5% data compression over CVNNs using FP16 while retaining a similar error with INT8 quantization, and our hardware support PAQ with only 2% area overhead over conventional systolic array architecture. In our evaluation, PAQ hardware with the scheduling scheme achieves a 32% lower latency and 30% lower energy consumption than other accelerators.
Hyunwuk Lee, Hyungjun Jang, Sungbin Kim, Sungwoo Kim 0003, Wonho Cho, Won Woo Ro
MICRO5