EDBT 2026 Demo / reviewers in the wild / expert
Wooyoung Jo
dblp:284/0695
· DBLP profile ↗
15ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0003-3598-3572ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 198.7 μJ/token Block Diffusion LLM Processor with Mask Token Similarity-Based Activation Reuse
Yujin Moon, Seryeong Kim, Wonhoon Park, Wooyoung Jo, Yuseon Choi, Sunjoo Whang, Hoi-Jun Yoo |
ISCAS | 5 |
| 2025 | BROCA: A Low-power and Low-latency Conversational Agent RISC-V System-on-Chip for Voice-interactive Mobile Devices
Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Haoyang Sang, Dongseok Im, Sangyeob Kim, Chaeyun Jeong, Yujin Moon, Hoi-Jun Yoo |
HCS | 1 |
| 2025 | A 32.65µm2 Spin/Area Large Scale Ising CIM with Progressive Circular Dataflow and Bi-directional eDRAM Cell ArrayabstractThis paper presents a large-scale Ising computing-in-memory (CIM) for a real-world combinatorial optimization problem (COP). While the CIM approach shows promising performance improvement compared to digital-based Ising machines, it can be only used for simple COP due to limited CIM connectivity, bit precision, and graph size. To overcome limitations, the proposed processor supports reconfigurable high-bit Chimera graph topology achieving a small cell area and high energy efficiency through three key features : 1) Progressive circular dataflow reduce area by 58.8% and improved energy efficiency by 44.7% thanks to fully reuse spin and coefficient. 2) Bi-directional eDRAM cell array support bi-directional ising computation for circular dataflow with 3T-2C eDRAM cell. Due to the signed operation with compact coefficient storing, the area was reduced by 36.1%. 3) Reconfigurable spin exchange link and reconfigurable C-2C ladder support various graph sizes and various bit precision of coefficients. In conclusion, the proposed large-scale Ising CIM achieves 32.65µm2spin area and 2.09µW effective spin power which is 3.21× and 5.17× smaller than previous state-of-the art. Jingu Lee, Sangwoo Ha, Sunjoo Whang, Soyeon Um, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 7 |
| 2025 | A 17.1 TOPS/W FP-INT Transformer Inference Accelerator with Sparsity Boosting and Output Importance-Aware ProcessingabstractThis paper presents an energy-efficient FP-INT transformer inference accelerator for diverse applications. The proposed accelerator achieves high energy efficiency with two key features as solutions: 1) Sparsity Boosting Adder Tree (SBAT) to reduce adder tree power by 23.8% by modifying the adder tree structure and booth encoding to maximize sparsity. 2) Output Importance-aware Processing (OIAP) to reduce the Floating-Point Accumulation (FP-ACC) power by 79.6%, dynamically adjusting input block sizes based on the saliency of output channels, thereby reducing the FP-ACC operations by 78.7%. The proposed accelerator is implemented in 28 nm CMOS technology and achieves 17.1 TOPS/W, leveraging the unique FP-INT characteristics on booth encoding and exploiting redundancy based on output saliency in the transformer architecture. Jeonggyu So, Seongyon Hong, Wooyoung Jo, Hoi-Jun Yoo, Donghyeon Han |
ISCAS | 4 |
| 2025 | Generative Unsupervised Anomaly Detection with Coarse-Fine Ensemble for Workload Reduction in 3D Non-contrast Brain CT of Emergency Room
Jongjun Won, Joonseo Oh, Yereen Yoo, Jieun Yum, Joonsang Lee, Joon Hyung Park, Wooyoung Jo, Yoojin Nam, Hyunki Lee, Gil-Sun Hong, Namkug Kim |
MICCAI (3) | 8 |
| 2024 | A Low-Power Large-Language-Model Processor with Big-Little Network and Implicit-Weight-Generation for On-Device AI
Sangyeob Kim, Wooyoung Jo, Seongyon Hong, Nayeong Lee, Hoi-Jun Yoo |
HCS | 3 |
| 2024 | A 28.6 mJ/iter Stable Diffusion Processor for Text-to-Image Generation with Patch Similarity-based Sparsity Augmentation and Text-based Mixed-PrecisionabstractThis paper presents an energy-efficient stable diffusion processor for text-to-image generation. While stable diffusion attained attention for high-quality image synthesis results, its inherent characteristics hinder its deployment on mobile platforms. The proposed processor achieves high throughput and energy efficiency with three key features as solutions: 1) Patch similarity-based sparsity augmentation (PSSA) to reduce external memory access (EMA) energy of self-attention score by 60.3 %, leading to 37.8 % total EMA energy reduction. 2) Text-based important pixel spotting (TIPS) to allow 44.8 % of the FFN layer workload to be processed with low-precision activation. 3) Dual-mode bit-slice core (DBSC) architecture to enhance energy efficiency in FFN layers by 43.0 %. The proposed processor is implemented in 28 nm CMOS technology and achieves 3.84 TOPS peak throughput with 225.6 mW average power consumption. In sum, 28.6 mJ/iteration highly energy-efficient text-to-image generation processor can be achieved at MS-COCO dataset. Wooyoung Jo, Seongyon Hong, Beomseok Kwon, Wonhoon Park, Hoi-Jun Yoo |
ISCAS | 2 |
| 2023 | A 332 TOPS/W Input/Weight-Parallel Computing-in-Memory Processor with Voltage-Capacitance-Ratio Cell and Time-Based ADCabstractRecent computing-in-memory (CIM) achieves high energy efficiency with charge-domain computation and multi-bit input driving. However, the previous works still require high power consumption and trade computation signal-to-noise ratio (SNR) for energy efficiency. This work proposes an energy-efficient and accurate multi-bit input/weight-parallel CIM processor with four key features: 1) a 10T2C sign-magnitude cell with voltage-capacitance-ratio (VCR) decoding for 5-bit analog inputs with only 2-level supply voltages, 2) a computation word line (CWL) charge reuse method for input driver power reduction, 3) a signal-amplifying noise canceling voltage-to-time converter (SANC-VTC) for SNR improvement, and 4) a distribution-aware time-to-digital converter (DA-TDC) for ADC power reduction. The proposed CIM processor is simulated in 28 nm CMOS technology with 1.25 mm2area. As a result, it achieves 4.44 mW power consumption and 332 TOPS/W energy efficiency with 72.43% benchmark accuracy (@ ImageNet, ResNet50, 5-bit input/5-bit weight). Seongyon Hong, Soyeon Um, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 5 |
| 2023 | A Reconfigurable 1T1C eDRAM-based Spiking Neural Network Computing-In-Memory Processor for High System-Level EfficiencyabstractSpiking Neural Network (SNN) Computing-In-Memory (CIM) was proposed for high macro-level energy efficiency. However, system-level energy efficiency is limited by EMA due to a large intermediate activation footprint requirement. To reduce the EMA, a large capacity SNN CIM is needed to load tons of weights in the CIM. This paper proposes a high-density 1T1C eDRAM-based SNN CIM processor for supporting high system-level energy efficiency with two key features: 1) High-density and low-power Reconfigurable Neuro-Cell Array (ReNCA) for memory and SNN peripheral logic using a charge pump and reusing 1T1C cell array, achieving 41% area and 90% power reduction compared to previous work. 2) Reconfigurable CIM architecture with dual-mode ReNCA and Dynamic Adjustable Neuron Link (DAN Link) for layer fusion increases system-level efficiency including intermediate and weight EMA. It achieves$10\times$higher state-of-the-art system-level energy efficiency including EMA. Seryeong Kim, Soyeon Um, Zhiyong Li 0016, Sangyeob Kim, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 7 |
| 2023 | A 15.9 mW 96.5 fps Memory-Efficient 3D Reconstruction Processor with Dilation-based TSDF Fusion and Block-Projection Cache SystemabstractA real-time dense 3D reconstruction on lightweight AR headsets is challenging since its memory access surpasses the available memory bandwidth. To solve this problem, the proposed processor integrates two key building blocks - Dilation-based TSDF (D-TSDF) fusion and Block-Projection (BP) engine. D-TSDF projects the depth map in the reverse order of voxel-to-pixel coordinate transformation and dilates it, leading to 96.61% External Memory Access (EMA) reduction with minimum map quality degradation. Second, a specialized BP engine compresses high-resolution occupancy grid by decomposing the 3D bitmap into 2D and 1D vectors, achieving$\times \mathbf{166.09}$reduced memory bandwidth. The proposed processor is implemented in 28nm CMOS technology occupying 1.27 mm2area. As a result, 96.45 fps 3D reconstruction is possible while consuming only 15.94 mW power. Hankyul Kwon, Gwangtae Park, Junha Ryu, Wooyoung Jo, Hoi-Jun Yoo |
ISCAS | 4 |
| 2023 | A 5.99 TFLOPS/W Heterogeneous CIM-NPU Architecture for an Energy Efficient Floating-Point DNN AccelerationabstractThis work presents an energy-efficient digital-based computing-in-memory (CIM) processor to support floating-point (FP) deep neural network (DNN) acceleration. Previous FP-CIM processors have two limitations. Processors with post-alignment shows low throughput due to serial operation, and the other processor with pre-alignment incurs truncation error. To resolve these problems, we focus on the statistics that outlier exists according to shift amount in pre-alignment-based FP operation. As those outlier decreases energy efficiency due to long operation cycles, it needs to be processed separately. The proposed Hetero-FP-CIM integrates both CIM arrays and shared NPU, so they compute both dense inlier and sparse outlier respectively. It also includes efficient weight caching system to avoid entire weight copy in shared NPU. The proposed Hetero-FP-CIM is simulated in 28 nm CMOS technology and occupies 2.7 mm2. As a result, it achieves 5.99 TOPS/W at ImageNet (ResNet50) with bfloat16 representation. Wonhoon Park, Junha Ryu, Soyeon Um, Wooyoung Jo, Sangyoeb Kim, Hoi-Jun Yoo |
ISCAS | 5 |
| 2022 | A 161.6 TOPS/W Mixed-mode Computing-in-Memory Processor for Energy-Efficient Mixed-Precision Deep Neural NetworksabstractA Mixed-mode Computing-in memory (CIM) processor for the mixed-precision Deep Neural Network (DNN) processing is proposed. Due to the bit-serial processing for the multi-bit data, the previous CIM processors could not exploit the energy-efficient computation of mixed-precision DNNs. This paper proposes an energy-efficient mixed-mode CIM processor with two key features: 1) Mixed-Mode Mixed-precision CIM (M3-CIM) which achieves 55.46% energy efficiency improvement. 2) Digital-CIM for In-memory MAC for the increased throughput of M3-CIM. The proposed CIM processor was simulated in 28nm CMOS technology and occupies 1.96 mm2. It achieves a state-of-the-art energy efficiency of 161.6 TOPS/W with 72.8% accuracy at ImageNet (ResNet50). Wooyoung Jo, Juhyeong Lee, Soyeon Um, Zhiyong Li 0016, Hoi-Jun Yoo |
ISCAS | 1 |
| 2022 | TSUNAMI: Triple Sparsity-Aware Ultra Energy-Efficient Neural Network Training Accelerator With Multi-Modal Iterative PruningabstractThis article proposes the TSUNAMI, which supports an energy-efficient deep-neural-network training. The TSUNAMI supports multi-modal iterative pruning to generate zeros in activation and weight. Tile-based dynamic activation pruning unit and weight memory shared pruning unit eliminate additional memory access. Coarse-zero skipping controller skips multiple unnecessary multiply-and-accumulation (MAC) operations at once, and fine-zero skipping controller skips randomly located unnecessary MAC operations. Weight sparsity balancer solves a utilization degradation caused by weight sparsity imbalance, and the workload of each convolution core is allocated by a random channel allocator. The TSUNAMI achieves an energy efficiency of 3.42 TFLOPS/W at 0.78V and 50MHz with floating-point 8-bit activation and weight. Also, it achieves an energy efficiency of 405.96 TFLOPS/W at 90% sparsity condition. Sangyeob Kim, Juhyoung Lee, Donghyeon Han, Wooyoung Jo, Hoi-Jun Yoo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | An Energy-efficient Floating-Point DNN Processor using Heterogeneous Computing Architecture with Exponent-Computing-in-MemoryabstractAbstract of Proposed FP CIM Processor (1) Heterogeneous FP Computing Arch. : Separate optimization of FP computing: Realize 2 cycles FP MAC w/ CIM (2) Exponent Computing-in-Memory: In-memory AND/NOR + BL charge reusing: Total memory power 46.4% 2) Mantissa Free Exponent Calculation: Removing redundant normalization: Total MAC power 14.4% Juhyoung Lee, Ji-Hoon Kim 0004, Wooyoung Jo, Sangyeob Kim, Donghyeon Han, Jinsu Lee, Hoi-Jun Yoo |
HCS | 3 |
| 2021 | OmniDRL: An Energy-Efficient Mobile Deep Reinforcement Learning Accelerators with Dual-mode Weight Compression and Direct Processing of Compressed DataabstractDeep Reinforcement Learning (DRL)▪ No Pre-labelled Data ➔ Training with Trial-and-errors!– Sequential decision making problems @ Unknown environments– Applications: gaming agent, autonomous systems, agent adaptation Juhyoung Lee, Sangyeob Kim, Ji-Hoon Kim 0004, Wooyoung Jo, Donghyeon Han, Hoi-Jun Yoo |
HCS | 5 |