VLDB 2026 Research / reviewers in the wild / expert
Sungju Ryu
dblp:196/4362
· DBLP profile ↗
18ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-0254-391XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 7 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NITRO: 3D NAND Flash-Based In-Storage LLM Computing with Enhanced Activation DataflowabstractIn-storage computing (ISC) has emerged as a next-generation memory architecture to relieve the data movement bottleneck between host processors and memory systems. While recent NAND flash-based processing-in-memory works leverage the high density of 3D NAND flash for deep neural networks, they primarily focus on optimizing computation inside the NAND array. Consequently, these approaches often fail to address the critical latency overhead associated with managing intermediate activation data. To overcome such a limitation, we propose a heterogeneous NAND flash-based ISC architecture with enhanced activation buffering. By buffering intermediate values in a DRAM subsystem rather than programming them into the NAND flash array, our approach effectively mitigates the high programming latency penalties. We also introduce a distributed dataflow scheme that maximizes computational parallelism through optimized plane- and bank-level data mapping. The results show that our proposed architecture achieves performance improvements, reducing inference latency by up to 86% compared to the baseline. Sanghun Shin, Gisan Ji, Sungju Ryu |
DATE | 3 |
| 2026 | E-Flash: Energy-Efficient LLM Mapping on NAND Flash-Based In-Storage Inference ComputingabstractTransformer-based deep neural networks (DNNs) have achieved remarkable success across a wide range of applications such as image and text generation tasks. However, The continuous growth of model size and memory demands imposes significant pressure on energy consumption and memory bandwidth, especially during weight access operations. To deal with such a challenge, prior studies have investigated NAND flash-based processing-in-memory (PIM) architectures, but it still experiences large energy consumption due to the significant increase in the recent model sizes. In this work, we present E-Flash, a digital NAND flash-based architecture for energy-efficient DNN weight access. E-Flash introduces a novel state-switching algorithm that reallocates frequently occurring weight patterns to low-power cell states in triple-level cell (TLC) flash memory. In addition, a cell-first allocation scheme further amplifies energy savings by aligning bit patterns within cells. Evaluation results on quantized BERT and Llama 2 models demonstrate up to 37.73% and 16.74% reduction in read energy, respectively, with negligible hardware overhead. Gisan Ji, Sanghun Shin, Jangho Baik, Wonbo Shim, Sungju Ryu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | RADiT: Redundancy-Aware Diffusion Transformer Acceleration Leveraging Timestep SimilarityabstractDiffusion Transformers (DiTs) have demonstrated unprecedented performance across various generative tasks including image and video generation. However, a large amount of computations on the inference process and iterative sampling steps in the DiT models result in high computational costs, leading to substantial latency and energy consumption challenges. To address these issues, we propose a redundancy-aware DiT (RADiT), a novel software-hardware co-optimization accelerator for DiTs that minimizes redundant operations in the iterative sampling stages. We identify data redundancy by evaluating blockwise input features and skip redundant computations by reusing results from consecutive timesteps. Furthermore, to minimize accuracy degradation and maximize computational efficiency, the Dynamic Threshold Scaling Module (DTSM) and Compress and Compare Unit (CCU) are employed in the redundancy detection process. This approach enables DiTs to achieve up to $1.8 \times$ and $1.7 \times$ faster speeds for image and video generation, respectively, without compromising quality, along with 41% and 45.5% reductions in energy consumption. Our RADiT scheme improves throughput by $1.67 \times$ and $1.76 \times$ for image and video generation tasks, respectively, while maintaining output quality and significantly reducing energy consumption. Youngjun Park, Yeonggeon Kim, Gisan Ji, Sungju Ryu |
DAC | 5 |
| 2025 | Thanos: Energy-Efficient Keyword Spotting Processor with Hybrid Time-Feature-Frequency-Domain Zero-SkippingabstractIn recent years, the keyword spotting algorithm has gained significant attention for applications such as personalized virtual assistants. However, the keyword spotting system must be always turned on to listen to the input voice for the recognition, which worsens the battery constraint problem in the edge devices. In this paper, we first analyze the sparsities in the keyword spotting computation. Based on the characteristic, we introduce the keyword spotting processor called Thanos to enable the zero-skipping scheme in the multiple keyword spotting domains to mitigate the burdensome energy consumption. Experimental results show that our hybrid-domain zero-skipping scheme reduces the latency by 80.3-87.3% and energy consumption by 48.1-79.8% on average, respectively, over the baseline architecture. Sungju Ryu |
DATE | 3 |
| 2025 | OptiRange: An Efficient ReRAM-Based PIM Accelerator with ADC Resolution OptimizationabstractProcessing-in-memory (PIM) techniques have attracted significant attention in computer architecture research due to their ability to mitigate the memory wall bottleneck for deep neural network (DNN) applications. In-memory computing (IMC) using ReRAM crossbar arrays is promising for energy efficiency but suffers from high-cost peripheral circuits, particularly analog-to-digital converters (ADCs). There are existing methods to reduce ADC energy consumption, such as lowering resolution or sharing an ADC among multiple columns. Unfortunately, these techniques reduce throughput. In this work, we propose OptiRange for PIM accelerators that enhances both energy efficiency and throughput while maintaining accuracy and model flexibility. OptiRange proposes a split-the-burden (STB) algorithm that manipulates crossbar cell values without changing the weights. STB reduces large cell values, shifts the column-sum distribution center closer to zero, and minimizes the number of low-resistance state (LRS) cells, significantly lowering the hardware-level fixed ADC resolution and improving accuracy. Furthermore, OptiRange utilizes dynamic ADC range (DAR) adaptation, determining the optimal ADC operating range per column at the software level and enabling the skipping of redundant ADC conversion cycles at runtime. An associated ADC range-based grouping (AG) strategy leverages the resulting dynamic latencies to manage synchronization and further boost system throughput. By combining weight-preserving cell manipulation with adaptive ADC operation and synchronization, OptiRange dynamically optimizes the ADC workload. Experimental results indicate that our method significantly enhances overall performance compared to conventional ReRAM-based accelerators. Sangkyu Jeon, Gisan Ji, Yeonggeon Kim, Youngjun Park, Sungju Ryu |
ICCAD | 6 |
| 2025 | E-Flash: Energy-Efficient DNN Mapping on NAND Flash Memory with State-Switching AlgorithmabstractDeep neural network (DNN) has been widely adopted in various applications. Ranging from image classification to text generation, Transformer-based models have demonstrated unprecedented performance. However, they suffer from a significant computational complexity and a large memory footprint, leading to memory-bound issues. While previous research on NAND flash-based neural network computation has been performed, these studies often encounter accuracy problems, as analog processing-in memory (PIM) operations typically lead to inaccurate results. Moreover, studies on NAND flash using single-level cell (SLC) are unable to fully leverage the advantages of efficient storage density on the multi-level cell (MLC) memory. We propose an E-Flash hardware architecture with an energy-efficient DNN mapping method. E-Flash introduces a state-switching algorithm to perform data movements between flash memory and host device in an energy-efficient manner. By reallocating the data in triple-level cell (TLC) NAND flash memory, we reduce the energy consumption during the data read operation. Experimental results demonstrate that E-Flash achieves improved energy consumption compared to baseline under significantly small area overhead for quantized BERT and Llama 2 weights by 37.73% and 16.74%, respectively. Gisan Ji, Sanghun Shin, Jangho Baik, Wonbo Shim, Sungju Ryu |
ISLPED | 5 |
| 2024 | NexusCIM: High-Throughput Multi-CIM Array Architecture with C-Mesh NoC and Hub CoresabstractThis paper introduces a high-performance Nexus-CIM architecture based on C-mesh NoC and hub cores. Conventional multi-CIM array architectures typically adopted mesh-type NoC structure. However, the mesh NoC shows NoC traffic congestion in the layer pipeline because multiple PEs must simultaneously communicate with each other, thereby leading to the PE-to-PE data communication bottleneck. To mitigate such a problem, we use the hub core instead of the previous vanilla router designs, and the hub core reduces the communication bottleneck by performing arithmetic operations such as aggregation. As a result, our NexusCIM shows higher throughput than baseline NoC designs by 1.04-8.9×. Sungju Ryu |
ICCD | 2 |
| 2024 | Statues: Energy-Efficient Video Object Detection on Edge Security Devices with Computational SkippingabstractThis paper proposes a software and hardware co-optimization method tailored to object detection on edge security devices. Object detection on video inputs requires a massive number of MAC computations, so it is difficult to implement a real-time inference task with limited computational budget on edge security devices. To relieve such computational complexity, we propose a Statues, selective computational skipping approach by scoring pixel differences in the recent inputs. If the score is low enough, we skip the rest DNN part because we expect the almost same object detection result as the previous frame. Our Statues approach maximizes energy-efficiency by 44% compared with the conventional method with negligible accuracy drop. Yeonggeon Kim, Sungju Ryu |
ISLPED | 3 |
| 2024 | Mobileware: Distributed Architecture With Channel Stationary Dataflow for MobileNet AccelerationabstractThe depthwise separable convolution, a key feature of the MobileNet models, has a different input reuse pattern from the conventional standard convolution, and a smaller number of input/weight pairs are used for a dot product, thereby leading to extremely low MAC utilization. This paper proposes a Mobileware architecture for the high-performance acceleration of the MobileNet workloads. A new channel stationary dataflow architecture distributes the on-chip buffers, and the distributed SRAMs are placed near each PE. By doing so, PEs and SRAMs can communicate with high bandwidth. Our Mobileware architecture shows 1.4-29.5× higher throughput than conventional weight stationary-based hardware architecture, and the proposed design was verified on the Xilinx ZCU102 FPGA evaluation board. Sungju Ryu, Jaeyong Jang, Youngtaek Oh, Jae-Joon Kim |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Teleport: A High-Performance ShiftNet Hardware Accelerator with Fused Layer ComputationabstractIn this paper, we introduce a high-performance ShiftNet-optimized hardware accelerator called Teleport. ShiftNet replaces the standard convolutional layers with zero-flop-based shift convolution and pointwise convolution to reduce the number of computations. However, previous hardware acceleration approaches do not support the shift convolution, and hence they mapped the shift operation to the$3\times 3$convolution, and thereby the shift layer still shows the same number of computations as the conventional convolutional layers. To mitigate such a limitation, we first fuse the shift and convolutional layers without modifying the original configuration of the ShiftNets, and the fused computations are accelerated using a custom address translator, a systolic loader, and a systolic array. Our work improved the performance by$6.1-103\times$over the previous hardware acceleration approach on the ShiftNet benchmark. Sungju Ryu |
ISLPED | 2 |
| 2023 | Binaryware: A High-Performance Digital Hardware Accelerator for Binary Neural NetworksabstractBinary neural networks (BNNs) largely reduce the memory footprint and computational complexity, so they are gaining interests on various mobile applications. In the BNNs, the first layer often accounts for the largest part of the entire computing time because the layer usually uses multi-bit multiplications. However, traditional hardware designed for BNN computing focuses primarily on the rest layers, resulting in significant performance degradation. In this brief, we introduce Binaryware architecture which achieves the high-performance computation on both the first and rest layers. Experimental results show that our Binaryware improves the throughput per compute area by 1.5–$13.3\times $on various BNN workloads. Sungju Ryu, Youngtaek Oh, Jae-Joon Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | SPRITE: Sparsity-Aware Neural Processing Unit with Constant Probability of Index-MatchingabstractSparse neural networks are widely used for memory savings. However, irregular indices of non-zero input activations and weights tend to degrade the overall system performance. This paper presents a scheme to maintain constant probability of index-matching for weight and input over a wide range of sparsity overcoming a critical limitation in previous works. A sparsity-aware neural processing unit based on the proposed scheme improves the system performance up to 6.1× compared to previous sparse convolutional neural network hardware accelerators. Sungju Ryu, Youngtaek Oh, Taesu Kim, Daehyun Ahn, Jae-Joon Kim |
DATE | 1 |
| 2021 | Mobileware: A High-Performance MobileNet Accelerator with Channel Stationary DataflowabstractMobileNet models have been gaining popularity with lighter computational load than previous convolutional neural network models. However, mapping depthwise convolutional layers in the MobileNets to the conventional hardware accelerators experiences significantly low utilization of multipliers, which leads to large performance degradation. The utilization becomes low because the depthwise convolutional layers have 1) different input/weight reuse patterns and 2) much smaller number of multiplications for each dot product computation compared to standard convolutional operations. To overcome such limitations, we propose an architecture called Mobileware which uses a channel stationary dataflow. Experimental results show that the Mobileware can achieve up to$1.4-29.5\times$higher overall system performance than the previous neural network accelerators. Sungju Ryu, Youngtaek Oh, Jae-Joon Kim |
ICCAD | 1 |
| 2021 | High-throughput Near-Memory Processing on CNNs with 3D HBM-like MemoryabstractThis article discusses the high-performance near-memory neural network (NN) accelerator architecture utilizing the logic die in three-dimensional (3D) High Bandwidth Memory– (HBM) like memory. As most of the previously reported 3D memory-based near-memory NN accelerator designs used the Hybrid Memory Cube (HMC) memory, we first focus on identifying the key differences between HBM and HMC in terms of near-memory NN accelerator design. One of the major differences between the two 3D memories is that HBM has the centralized through- silicon-via (TSV) channels while HMC has distributed TSV channels for separate vaults. Based on the observation, we introduce the Round-Robin Data Fetching and Groupwise Broadcast schemes to exploit the centralized TSV channels for improvement of the data feeding rate for the processing elements. Using synthesized designs in a 28-nm CMOS technology, performance and energy consumption of the proposed architectures with various dataflow models are evaluated. Experimental results show that the proposed schemes reduce the runtime by 16.4–39.3% on average and the energy consumption by 2.1–5.1% on average compared to conventional data fetching schemes. Naebeom Park, Sungju Ryu, Jaeha Kung 0001, Jae-Joon Kim |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2020 | Algorithm/Hardware Co-Design for In-Memory Neural Network Computing with Minimal Peripheral Circuit OverheadabstractWe propose an in-memory neural network accelerator architecture called MOSAIC which uses minimal form of peripheral circuits; 1-bit word line driver to replace DAC and 1-bit sense amplifier to replace ADC. To map multi-bit neural networks on MOSAIC architecture which has 1-bit precision peripheral circuits, we also propose a bit-splitting method to approximate the original network by separating each bit path of the multi-bit network so that each bit path can propagate independently throughout the network. Thanks to the minimal form of peripheral circuits, MOSAIC can achieve an order of magnitude higher energy and area efficiency than previous in-memory neural network accelerators. Yulhwa Kim, Sungju Ryu, Jae-Joon Kim |
DAC | 3 |
| 2019 | BitBlade: Area and Energy-Efficient Precision-Scalable Neural Network Accelerator with Bitwise SummationabstractDeep Neural Networks (DNNs) have various performance requirements and power constraints depending on applications. To maximize the energy-efficiency of hardware accelerators for different applications, the accelerators need to support various bit-width configurations. When designing bit-reconfigurable accelerators, each PE must have variable shift-addition logic, which takes a large amount of area and power. This paper introduces an area and energy efficient precision-scalable neural network accelerator (BitBlade), which reduces the control overhead for variable shift-addition using bitwise summation method. The proposed BitBlade, when synthesized in a 28nm CMOS technology, showed reduction in area by 41% and in energy by 36-46% compared to the state-of-the-art precision-scalable architecture [14]. Sungju Ryu, Wooseok Yi, Jae-Joon Kim |
DAC | 1 |
| 2019 | Feedforward-Cutset-Free Pipelined Multiply-Accumulate Unit for the Machine Learning AcceleratorabstractMultiply-accumulate (MAC) computations account for a large part of machine learning accelerator operations. The pipelined structure is usually adopted to improve the performance by reducing the length of critical paths. An increase in the number of flip-flops due to pipelining, however, generally results in significant area and power increase. A large number of flip-flops are often required to meet the feedforward-cutset rule. Based on the observation that this rule can be relaxed in machine learning applications, we propose a pipelining method that eliminates some of the flip-flops selectively. The simulation results show that the proposed MAC unit achieved a 20% energy saving and a 20% area reduction compared with the conventional pipelined MAC. Sungju Ryu, Naebeom Park, Jae-Joon Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Low design overhead timing error correction scheme for elastic clock methodologyabstractThe elastic clock scheme is a robust design methodology to ensure timing closure under PVT variation using locally generated clocks and handshaking protocol. However, it still has a chance of timing errors due to delay mismatch between the data-path and delay replica. In this paper, we propose a low design overhead timing error correction scheme tailored to elastic clock. In the proposed scheme, a timing error can be corrected within a cycle using clock stretching. The proposed scheme shows 40.3× and 4.6× reduction in timing margin with 9.1% and 9.0% area overhead over the synchronous baseline and elastic clock design, respectively. Sungju Ryu, Jongeun Koo, Jae-Joon Kim |
ISLPED | 1 |