EDBT 2026 Demo / reviewers in the wild / expert
Dongsuk Jeon
dblp:28/9878
· DBLP profile ↗
30ranked-venue papers
2as first author
25since 2021 · last 2026
0000-0002-0395-8076ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 16 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLUX: Frequency Scaling with Layer-wise Utilization for Energy-Efficient NPU Execution (WIP)abstractWith the widespread adoption of Deep Neural Networks (DNNs), Neural Processing Units (NPUs) are emerging as energy-efficient alternatives to GPUs through parallel processing and high data reuse. However, since diverse deep learning kernels have different memory and computation resource requirements, a utilization imbalance between memory and computation resources often occurs. Inho Lee 0002, Ky Yeop Lim, Hyejun Kim, Beomseok Kim, Dongsuk Jeon, Hunjun Lee, Yongjun Park 0001 |
LCTES | 5 |
| 2026 | Design Approaches for Efficient Parallel Pseudo-Random Ternary Sequence Generation
Dongkwon Lee, Hyungil Chae, Dongsuk Jeon |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | An Area- and Energy-Efficient Point-Based 3-D Object Detection Processor With Feature-Decoupled PointNet for Autonomous Driving
Changyu Seong, Sunwoo Lee 0005, Beomseok Kim, Dongsuk Jeon |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | Weighted Coding Scheme for Noise Reduction in Silicon Interposer of HBMabstractHigh-bandwidth memory (HBM) has enabled substantial advancements in bandwidth-intensive applications, including large-scale artificial intelligence models. HBM is typically integrated with other systems-on-chip (SoCs) through a silicon interposer, and the primary bottleneck in aggressively scaling the bandwidth in next-generation HBM systems stems from significant crosstalk caused by closely spaced, high-density parallel interconnects in the interposer. While crosstalk avoidance code (CAC) has emerged as a promising solution, prior CAC schemes suffer from low bit efficiency and considerable hardware overhead. This article proposes an efficient CAC scheme, WITCH, which employs a novel weighted coding strategy. Unlike prior approaches that consider all channels identically, WITCH assigns different weights to channels based on their relative positions in the array, enabling more bit-efficient crosstalk suppression. We also present WITCH with additional shielding (WITCH-AS), an extension of WITCH that incorporates additional shielding to further reduce crosstalk levels. Our coding schemes achieve high bit efficiency, up to 17.3% higher than state-of-the-art techniques while providing an identical level of crosstalk reduction. Simulation results using an industry-proven channel model demonstrate that WITCH and WITCH-AS improve eye height by 10.1%–49.4% and 17.1%–51.1%, respectively. Furthermore, we propose an area- and energy-efficient hardware implementation that can be integrated into real-world HBM systems. The design reduces area and critical path delay by 31.0% and 28.2%, respectively, compared with conventional designs. Finally, we propose a compatible simultaneous switching output (SSO) noise mitigation technique that can be seamlessly integrated into WITCH to further enhance signal integrity under high-speed, high-density operating conditions. Sangouk Jeon, Seoyoon Jang, Kwanghyun Shin, Dongkwon Lee, Hankyu Chi, Wookjin Shin, Changhyun Pyo, Jaeha Kim, Dongsuk Jeon |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2026 | An Energy-Efficient Block-Convolution-Based Super-Resolution Processor With Tiling Artifact Reduction
Byeungseok Yoo, Sungjin Park 0003, Sunwoo Lee 0005, Dongsuk Jeon |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2026 | Processing-in-Memory Architecture for In-Storage Information Retrieval on 3-D NAND Flash Solid-State DrivesabstractRecent advances in retrieval-augmented generation (RAG) highlight retrieval latency, dominated by storage access, and large embedding-index size as the major bottlenecks. Although processing-in-memory (PIM) based on 3-Dnandflash and binary passage retriever (BPR) algorithm offer potential solutions for reducing latency and memory usage, naively integrating BPR algorithm into an SSD based on prior in-memory search (IMS) architectures results in negligible performance gain with substantial hardware overhead. This stems from the fact that prior IMS architectures are optimized only for search, leaving candidate readout misaligned with thenandpage direction and requiring excessive readcycles, while BPR's original top-$L$selection further requires high-precision ADCs and sorting logic. To address these limitations, we propose an algorithm–hardware co-designed in-storage information retrieval architecture. We introduce a hardware-efficient optimized BPR algorithm with a threshold-based selection and two-stage in-storage processing approach to minimize data transfer overheads. We further present an IMS architecture with page-aligned encoding that efficiently supports both search and candidate read operations, eliminating the read cycle penalty. We also present a lightweight in situ error monitoring scheme that bypasses error-correcting code (ECC) decoding while preserving error-aware refresh triggering. Experiments show that our design reduces the amount of data read fromnandflash by 7.94–$1330.40\times $compared with host-side processing baselines and further reduces controller-to-host data transfer, while maintaining the retrieval accuracy. It also reduces total operationcycles by up to 96.93% compared with prior IMS architectures, and lowers error-detection power by 94.71% and total read energy by 17.22% compared with a conventional ECC decoder for error monitoring. Huiwon Yun, Inho Jeong, Kyoyun Lee, Jae Yong Lee 0004, Myoungjun Chun, Jihong Kim 0001, Dongsuk Jeon |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | WITCH: WeIghTed Coding Scheme for Crosstalk Reduction in High Bandwidth MemoryabstractHigh bandwidth memory (HBM) has enabled a breakthrough in bandwidth-bound applications, including large-scale artificial intelligence models. HBM is typically connected to other SoCs through a silicon interposer. However, the increasing density of the parallel interconnect wires incurs significant amount of crosstalk, hindering bandwidth improvement in the next-generation HBMs. While Crosstalk Avoidance Code (CAC) has emerged as a solution to mitigate crosstalk, prior CAC schemes suffer from low bit efficiency and significant hardware overhead. This paper proposes an efficient CAC scheme, WITCH. It employs a new coding system, weighted coding, which gives a different emphasis to each channel according to its relative position in the channel array. This enables crosstalk reduction with higher bit efficiency than prior CACs treating all channels in the array equally. The extended version of WITCH, WITCH-AS, is also proposed with additional shielding for further crosstalk reduction. Our coding system shows high bit efficiency of 91.2--91.7% and 84.3--84.6% for WITCH and WITCH-AS, which is up to 20.8% higher than the state-of-the-art schemes while preserving the same crosstalk level reduction. We have shown through simulations using an industry-proven channel model that WITCH and WITCH-AS improve the eye heights by 10.1--49.4% and 17.1--51.1% respectively. In addition, this paper presents an efficient hardware implementation of our coding schemes which shows 28.2% lower critical path delay and 31.0% smaller area than conventional implementation, proving itself a practical solution for HBMs. Seoyoon Jang, Sangouk Jeon, Kwanghyun Shin, Dongkwon Lee, Hankyu Chi, Wookjin Shin, Changhyun Pyo, Jaeha Kim, Dongsuk Jeon |
ASP-DAC | 9 |
| 2025 | Efficient Key Switching Accelerator for Fully Homomorphic EncryptionabstractFully Homomorphic Encryption (FHE) is a cryptosystem of the newest generation that allows limitless computation on encrypted data, enabling privacy-preserving computing in cloud services. While the Key Switching (KS) operation is the key bottleneck in FHE and has its unique, complex data access patterns, little attention has been given to maximizing efficiency for this particular key operation through dedicated hardware. This paper presents an efficient KS accelerator for FHE employing various design techniques on multiple levels. The proposed techniques lead to (i) an efficient modular multiplier showing 47.8% and 46.0% reduction each in area and power compared with naïve Barrett modular multiplier, lightweight NTT units with (ii) an efficient twiddle factor generator (TFG) reducing on-chip memory space for twiddle factors in NTT units from O(N) to O(logN) with minimal overhead and (iii) conflict-free addressing scheme to replace dual-port memories into single-port ones without performance degradation, and (iv) bandwidth-efficient core design leading to 38.7% lower external memory access compared with baseline. Implemented in a 28nm CMOS process, the design occupies 5.13×3.71mm2 and consumes 5.2 mJ per KS operation, achieving 11.0× acceleration of KS operation compared with the state-of-the-art CPU implementation and 12.7× higher energy efficiency in KS operation compared with state-of-the-art FHE processors or KS accelerators on various platforms. Seoyoon Jang, Sungjin Park 0003, Dongsuk Jeon |
ASP-DAC | 3 |
| 2025 | PaCA: Partial Connection Adaptation for Efficient Fine-TuningabstractPrior parameter-efficient fine-tuning (PEFT) algorithms reduce memory usage and computational costs of fine-tuning large neural network models by training only a few additional adapter parameters, rather than the entire model. However, the reduction in computational costs due to PEFT does not necessarily translate to a reduction in training time; although the computational costs of the adapter layers are much smaller than the pretrained layers, it is well known that those two types of layers are processed sequentially on GPUs, resulting in significant latency overhead. LoRA and its variants avoid this latency overhead by merging the low-rank adapter matrices with the pretrained weights during inference. However, those layers cannot be merged during training since the pretrained weights must remain frozen while the low-rank adapter matrices are updated continuously over the course of training. Furthermore, LoRA and its variants do not reduce activation memory, as the first low-rank adapter matrix still requires the input activations to the pretrained weights to compute weight gradients. To mitigate this issue, we propose **Pa**rtial **C**onnection **A**daptation (**PaCA**), which fine-tunes randomly selected partial connections within the pretrained weights instead of introducing adapter layers in the model. PaCA not only enhances training speed by eliminating the time overhead due to the sequential processing of the adapter and pretrained layers but also reduces activation memory since only partial activations, rather than full activations, need to be stored for gradient computation. Compared to LoRA, PaCA reduces training time by 22% and total memory usage by 16%, while maintaining comparable accuracy across various fine-tuning scenarios, such as fine-tuning on the MMLU dataset and instruction tuning on the Oasst1 dataset. PaCA can also be combined with quantization, enabling the fine-tuning of large models such as LLaMA3.1-70B. In addition, PaCA enables training with 23% longer sequence and improves throughput by 16\% on both NVIDIA A100 GPU and INTEL Gaudi2 HPU compared to LoRA. The code is available at [https://github.com/WooSunghyeon/paca](https://github.com/WooSunghyeon/paca). Sunghyeon Woo, Sol Namkung, Sunwoo Lee 0005, Inho Jeong, Beomseok Kim, Dongsuk Jeon |
ICLR | 6 |
| 2025 | Input-adaptive Mixed-Precision Framework for Efficient Object DetectionabstractTo reduce computational redundancy inherent in fixed-bit-width quantization, input-adaptive quantization dynamically adjusts the bit-width of network parameters based on the difficulty of the given input. However, estimating image difficulty for object detection is a non-trivial task, as multiple detection results may occur within a single image. In this paper, we propose an input-adaptive mixed-precision framework that automatically adjusts the bit-width of each layer in the target model based on the characteristics of an input image. For searching optimal bit configurations, the framework employs a reward function that considers both the difficulty of a single image and the computational cost. Experimental results demonstrate that the proposed method outperforms prior quantization methods with fixed bit-widths. Isaac Jeong, Ji-Ye Jeon, Dongsuk Jeon |
ISCAS | 4 |
| 2025 | LogSimViT: Logarithmic Similar Pattern Skipping for Hardware Acceleration of ViTabstractVision Transformer (ViT) has emerged as a key neural network architecture in computer vision, but its substantial computation and power overheads pose significant challenges, particularly in resource-constrained environments. To address these limitations, we propose Log-similar Pattern Skipping (LPS) algorithm and its accelerator architecture. LPS exploits the redundancy in logarithmically quantized weights to regularly bypass half of FC operations, which accounts for 54% of the total latency of DeiT-Tiny, Small, and Base models on average. Our method and its hardware efficiently accelerate ViT models, achieving a speedup of 1.71× compared to an edge GPU, and 1.70× speedup and 1.33× higher energy efficiency compared to state-of-the-art ViT accelerators, with minimal accuracy loss. Isaac Jeong, Seongho Jeong, Dongsuk Jeon |
ISCAS | 4 |
| 2025 | Cost-Effective Reconfigurable MCM with Common-Value Elimination and AlignmentabstractMultiple constant multiplication (MCM) generates area/energy-efficient circuit with additions and shifts for FIR filters and DNN inference acceleration. Reconfigurable MCM (ReMCM) enables time-multiplexing execution of MCM circuits for multiple constant sets by reusing adders across different sets. Finding a shift-add-mux circuit in ReMCM, however, generally relies on a rigid order of constants, which may cause unnecessary topology conflicts and thereby incur a power/area overhead. To address the problem, this work proposes an Adaptive Grouping Compatible Graph Synthesis (AG-CGS) framework for ReMCM, which reorders constants to enhance power/area efficiency. Employing an (index, value) data structure for a constant, AG-CGS introduces two novel techniques in ReMCM: intra-set common value elimination (Intra-SCE) and inter-set common value alignment (Inter-SCA). Consequently, AG-CGS enables searching for a cost-effective circuit associated with reordered coefficients to reduce redundant computations. The experimental results demonstrate that AG-CGS reduces area by up to 58.1%, 16.3%, and 13.8%, and energy by up to 63.1%, 15.2%, and 12.1%, compared to [1], [2], and [3], respectively. Xuan Truong Nguyen, Dongsuk Jeon |
ISCAS | 3 |
| 2025 | HiFC: High-efficiency Flash-based KV Cache Swapping for Scaling LLM InferenceabstractLarge‑language‑model inference with long contexts often produces key–value (KV) caches whose footprint exceeds the capacity of high‑bandwidth memory on a GPU. Prior LLM inference frameworks such as vLLM mitigate this pressure by swapping KV cache pages to host DRAM. However, the high cost of large DRAM pools makes this solution economically unattractive. Although offloading to SSDs can be a cost-effective way to expand memory capacity relative to DRAM, conventional frameworks such as FlexGen experience a substantial throughput drop since the data path that routes SSD traffic through CPU to GPU is severely bandwidth-constrained. To overcome these limitations, we introduce HiFC, a novel DRAM‑free swapping scheme that enables direct access to SSD-resident memory with low latency and high effective bandwidth. HiFC stores KV pages in pseudo-SLC (pSLC) regions of commodity NVMe SSDs, sustaining high throughput under sequential I/O and improving write endurance by up to 8$\times$. Leveraging GPU Direct Storage, HiFC enables direct transfers between SSD and GPU, bypassing host DRAM and alleviating PCIe bottlenecks. HiFC employs fine-grained block mapping to confine writes to high-performance pSLC zones, stabilizing latency and throughput under load. HiFC achieves inference throughput comparable to DRAM-based swapping under diverse long-context workloads, such as NarrativeQA, while significantly lowering the memory expansion cost of a GPU server system by 4.5$\times$ over three years. Inho Jeong, Sunghyeon Woo, Sol Namkung, Dongsuk Jeon |
NeurIPS | 4 |
| 2025 | FDAM: Filter-Dedicated Approximate Multiplier Design for Real-Time CNN AccelerationabstractVideo super-resolution (VSR) is widely used in various high-definition applications, such as HDTVs and smartphones, requiring a dedicated upscaling technique for real-time full-HD generation. To reduce on-chip buffers for large-size output feature maps, a streaming VSR accelerator may employ an output stationary dataflow, leading to a large energy consumption caused by frequent filter switching. To mitigate this, we introduce a new filter-dedicated multiplier design for real-time VSR acceleration. We replace costly multipliers with adders, shifters, and multiplexers (MUXes), referred to as unified multiple constant multiplications (UMCM). The conventional UMCM, however, may incur considerable area/power overhead due to the unified topology constraint among different filter sets. To address this problem, we propose a new approximated MCM (AMCM) problem to relax the constraint and an approximate compatible graph synthesis (A-CGS) framework to efficiently solve AMCM by jointly searching for approximated filters and constructing a unified graph. Additionally, we suggest a lightweight fine-tuning method by freezing approximated filters and only fine-tuning biases, which can recover the original model’s accuracy within a few epochs. Experimental results with synthetic data demonstrate that AMCM reduces the area by up to 49.8%, 44.8%, and 40.3% when considering constant sets of 2, 4, and 8, respectively. Our designs with SR application achieve up to a 73.3% reduction in energy consumption. Experiments with the Set5 and Set14 datasets show that our model with bias correction achieves similar restoration performance compared to the eight-bit models. Xuan Truong Nguyen, Dongsuk Jeon |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | CLAT: A Clustering-Based Attention Transformer Accelerator for Low-Latency Text Generation in LLMsabstractTransformer-based large language models (LLMs) excel in text generation but face challenges like memory bandwidth bottlenecks and large key-value (KV) cache sizes as context lengths grow, impacting low-latency performance. Existing accelerators adopt parallelism, model compression, and sparsity exploitation but often fail to fully utilize token-specific sparsity, limiting their effectiveness for long context lengths. CLAT addresses these issues with a low-overhead clustering algorithm that identifies relevant key vector clusters for each token’s query, omitting less relevant vectors with minimal impact. It optimizes memory bandwidth using routing for single-batch inference and introduces scheduling techniques to reduce attention layer latency. Additionally, CLAT compresses model parameters to 4-bit precision and KV caches to 8-bit precision, supported by a multi-precision MAC structure that avoids extra overhead. Validated on Llama2-7B, OPT-6.7B, and Llama3-8B models, CLAT reduces attention layer latency by up to 88.6% and overall text generation latency by up to 34.9%. It improves single-batch text generation throughput by$1.66\times $to$2.42\times $over an A100 GPU, demonstrating significant performance gains. Sunwoo Lee 0005, Beomseok Kim, Jeongwoo Park 0001, Dongsuk Jeon |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | ALAM: Averaged Low-Precision Activation for Memory-Efficient Training of Transformer ModelsabstractOne of the key challenges in deep neural network training is the substantial amount of GPU memory required to store activations obtained in the forward pass. Various Activation-Compressed Training (ACT) schemes have been proposed to mitigate this issue; however, it is challenging to adopt those approaches in recent transformer-based large language models (LLMs), which experience significant performance drops when the activations are deeply compressed during training. In this paper, we introduce ALAM, a novel ACT framework that utilizes average quantization and a lightweight sensitivity calculation scheme, enabling large memory saving in LLMs while maintaining training performance. We first demonstrate that compressing activations into their group average values minimizes the gradient variance. Employing this property, we propose Average Quantization which provides high-quality deeply compressed activations with an effective precision of less than 1 bit and improved flexibility of precision allocation. In addition, we present a cost-effective yet accurate sensitivity calculation algorithm that solely relies on the L2 norm of parameter gradients, substantially reducing memory overhead due to sensitivity calculation. In experiments, the ALAM framework significantly reduces activation memory without compromising accuracy, achieving up to a 10$\times$ compression rate in LLMs. Sunghyeon Woo, Sunwoo Lee 0005, Dongsuk Jeon |
ICLR | 3 |
| 2024 | DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward PropagationabstractLarge language models (LLMs) have achieved significant success across various domains. However, training these LLMs typically involves substantial memory and computational costs during both forward and backward propagation. While parameter-efficient fine-tuning (PEFT) considerably reduces the training memory associated with parameters, it does not address the significant computational costs and activation memory. In this paper, we propose Dropping Backward Propagation (DropBP), a novel approach designed to reduce computational costs and activation memory while maintaining accuracy. DropBP randomly drops layers during backward propagation, which is essentially equivalent to training shallow submodules generated by undropped layers and residual connections. Additionally, DropBP calculates the sensitivity of each layer to assign an appropriate drop rate, thereby stabilizing the training process. DropBP is not only applicable to full fine-tuning but can also be orthogonally integrated with all types of PEFT by dropping layers during backward propagation. Specifically, DropBP can reduce training time by 44% with comparable accuracy to the baseline, accelerate convergence to the same perplexity by 1.5$\times$, and enable training with a sequence length 6.2$\times$ larger on a single NVIDIA-A100 GPU. Furthermore, our DropBP enabled a throughput increase of 79% on a NVIDIA A100 GPU and 117% on an Intel Gaudi2 HPU. The code is available at [https://github.com/WooSunghyeon/dropbp](https://github.com/WooSunghyeon/dropbp). Sunghyeon Woo, Baeseong Park, Byeongwook Kim, Minjung Jo, Se Jung Kwon, Dongsuk Jeon, Dongsoo Lee |
NeurIPS | 6 |
| 2023 | Learning with Auxiliary Activation for Memory-Efficient Training
Sunghyeon Woo, Dongsuk Jeon |
ICLR | 2 |
| 2023 | A Real-Time Object Detection Processor With xnor-Based Variable-Precision Computing UnitabstractObject detection algorithms using convolutional neural networks (CNNs) achieve high detection accuracy, but it is challenging to realize real-time object detection due to their high computational complexity, especially on resource-constrained mobile platforms. In this article, we propose an algorithm-hardware co-optimization approach to designing a real-time object detection system. We first develop a compact object detection model based on a binarized neural network (BNN), which employs a new layer structure, the DenseToRes layer, to mitigate information loss due to deep quantization. We also propose an efficient object detection processor that runs object detection with high throughput using limited hardware rescources. We develop a resource-efficient processing unit supporting variable precision with minimal hardware overheads. Implemented in field-programmable gate array (FPGA), the object detection processor achieves 64.51 frames/s throughput with 64.92 mean average precision (mAP) accuracy. Compared to prior FPGA-based designs for object detection, our design achieves high throughput with competitive accuracy and lower hardware implementation costs. Kukbyung Kim, Woohyun Ahn, Dongsuk Jeon |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | Toward Efficient Low-Precision Training: Data Format Optimization and Hysteresis Quantization
Sunwoo Lee 0005, Jeongwoo Park 0001, Dongsuk Jeon |
ICLR | 3 |
| 2022 | An Automatic Circuit Design Framework for Level Shifter CircuitsabstractAlthough design automation is a key enabler of modern large-scale digital systems, automating the transistor-level circuit design process still remains a challenge. Some recent works suggest that deep learning algorithms could be adopted to find optimal transistor dimensions in relatively small circuitry such as analog amplifiers. However, those approaches are not capable of exploring different circuit structures to meet the given design constraints. In this work, we propose an automatic circuit design framework that can generate practical circuit structures from scratch as well as optimize the size of each transistor, considering performance and reliability. We employ the framework to design level shifter circuits, and the experimental results show that the framework produces novel level shifter circuit topologies and the automatically optimized designs achieve$2.8\times $–$5.3\times $lower power-delay product (PDP) than prior arts designed by human experts. Jiwoo Hong, Dongsuk Jeon |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | A 270-mA Self-Calibrating-Clocked Output-Capacitor-Free LDO With 0.15-1.15V Output Range and 0.183-fs FoMabstractThis article proposes a fully integrated output-capacitor-free low-dropout regulator (LDO) for mobile applications. To overcome the limited output voltage range of typical analog LDOs, our design uses a rail-to-rail voltage-difference-to-time-converter (VDTC) and a charge pump (CP) to achieve a wide output range. Using a self-calibrating clock generator (SCCG) removes the need for an external clock source and adaptively tunes the clock frequency, enabling fast transient responses while minimizing quiescent current. A tunable undershoot compensator (TUC) mitigates voltage droop by detecting the drop in the output voltage due to a sharp increase in load current and compensating the output voltage immediately. The proposed LDO is fabricated in a 65-nm low power (LP) CMOS process and demonstrates a maximum load current capacity of 270 mA. The input and output voltage ranges of the LDO are 0.5–1.2 and 0.15–1.15 V, respectively, with 12.7-$\mu \text{A}$quiescent current and 99.99% peak current efficiency. The measured undershoot and settling time are 150 mV and 100 ns at a slew rate of 200 mA/3 ns, respectively, achieving a figure of merit (FoM) of 0.183 fs. Dongsuk Jeon |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | Real-Time Denoising and Dereverberation wtih Tiny Recurrent U-NetabstractModern deep learning-based models have seen outstanding performance improvement with speech enhancement tasks. The number of parameters of state-of-the-art models, however, is often too large to be deployed on devices for real-world applications. To this end, we propose Tiny Recurrent U-Net (TRU-Net), a lightweight online inference model that matches the performance of current state-of- the-art models. The size of the quantized version of TRU-Net is 362 kilobytes, which is small enough to be deployed on edge devices. In addition, we combine the small-sized model with a new masking method called phase-aware ß-sigmoid mask, which enables simultaneous denoising and dereverberation. Results of both objective and subjective evaluations have shown that our model can achieve competitive performance with the current state-of-the-art models on benchmark datasets using fewer parameters by orders of magnitude. Hyeong-Seok Choi, Sungjin Park 0003, Jie Hwan Lee, Hoon Heo, Dongsuk Jeon, Kyogu Lee |
ICASSP | 5 |
| 2021 | Activation Sharing with Asymmetric Paths Solves Weight Transport Problem without Bidirectional ConnectionabstractOne of the reasons why it is difficult for the brain to perform backpropagation (BP) is the weight transport problem, which argues forward and feedback neurons cannot share the same synaptic weights during learning in biological neural networks. Recently proposed algorithms address the weight transport problem while providing good performance similar to BP in large-scale networks. However, they require bidirectional connections between the forward and feedback neurons to train their weights, which is observed to be rare in the biological brain. In this work, we propose an Activation Sharing algorithm that removes the need for bidirectional connections between the two types of neurons. In this algorithm, hidden layer outputs (activations) are shared across multiple layers during weight updates. By applying this learning rule to both forward and feedback networks, we solve the weight transport problem without the constraint of bidirectional connections, also achieving good performance even on deep convolutional neural networks for various datasets. In addition, our algorithm could significantly reduce memory access overhead when implemented in hardware. Sunghyeon Woo, Jeongwoo Park 0001, Jiwoo Hong, Dongsuk Jeon |
NeurIPS | 4 |
| 2021 | Dynamic Block-Wise Local Learning Algorithm for Efficient Neural Network Training
Gwangho Lee, Sunwoo Lee 0005, Dongsuk Jeon |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | An Area-Efficient 128-Channel Spike Sorting Processor for Real-Time Neural Recording With 0.175µW/Channel in 65-nm CMOSabstractThis paper presents a power- and area-efficient spike sorting processor (SSP) for real-time neural recordings. The proposed SSP includes novel detection, feature extraction, and improved K-means algorithms for better clustering accuracy, online clustering performance, and lower power and smaller area per channel. Time-multiplexed registers are utilized in the detector for dynamic power reduction. Finally, an ultra-low-voltage 8T static random access memory (SRAM) is developed to reduce area and leakage consumption when compared to D flip-flop-based memory. The proposed SSP, fabricated in 65-nm CMOS process technology, consumes only 0.175 μW/channel when processing 128 input channels at 3.2 MHz and 0.54 V, which is the lowest among the compared state-of-the-art SSPs. The proposed SSP also occupies 0.003 mm2/channel, which allows 333 channels/mm2. Anh-Tuan Do, Seyed Mohammad Ali Zeinolabedin, Dongsuk Jeon, Dennis Sylvester, Tony Tae-Hyoung Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | Enhancing Reliability of Analog Neural Network ProcessorsabstractRecently, analog and mixed-signal neural network processors have been extensively studied due to their better energy efficiency and small footprint. However, analog computing is more vulnerable to circuit nonidealities such as process variation than their digital counterparts. On-chip calibration circuits can be adopted to measure and compensate for those effects, but it leads to unavoidable area and power overheads. In this brief, we propose a variation and noise-tolerant learning algorithm and postsilicon process variation compensation technique which does not require any additional monitoring circuitry. The proposed techniques reduce the accuracy degradation in the corrupted fully connected network down to 1% under large amount of variations including 10% unit capacitor mismatch, 8-mVrmscomparator noise and 20-mVrmscomparator offset. Suhong Moon, Kwanghyun Shin, Dongsuk Jeon |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | A low-power VGA full-frame feature extraction processorabstractThis paper proposes an energy-efficient VGA full-frame feature extraction processor design. It is based on the SURF algorithm and makes various algorithmic modifications to improve efficiency and reduce hardware overhead while maintaining extraction performance. Low clock frequency and deep parallelism derived from a one-sample-per-cycle matched-throughput architecture provide significantly larger room for voltage scaling and enables full-frame extraction. The proposed design consumes 4.7mW at 400mV and achieves 72% higher energy efficiency than prior work. Dongsuk Jeon, Yejoong Kim, Inhee Lee 0001, Zhengya Zhang, David T. Blaauw, Dennis Sylvester |
ICASSP | 1 |
| 2011 | Pipeline strategy for improving optimal energy efficiency in ultra-low voltage designabstractThis paper investigates pipelining methodologies for the ultra low voltage regime. Based on an analytical model and simulations, we propose a pipelining technique that provides higher energy efficiency and performance than conventional approaches to ultra low voltage design. Two-phase latch based design and sequential circuit optimizations are also proposed to further improve energy efficiency and performance. Silicon results demonstrate a 16b multiplier using the approaches in 65nm CMOS improve energy efficiency by 30% and performance by 60%. Mingoo Seok, Dongsuk Jeon, Chaitali Chakrabarti, David T. Blaauw, Dennis Sylvester |
DAC | 2 |
| 2011 | Energy-optimized high performance FFT processorabstractThis paper proposes an ultra low energy FFT processor suitable for sensor applications. The processor is based on R4MDC but achieves full utilization of computational elements. It has two parallel datapaths that increase throughput by a factor of 2 and also enable high memory utilization. The proposed design is implemented in 65nm CMOS technology and post-layout simulation including parasitic capacitances shows it achieves 9.25× higher energy efficiency than state-of-the-art FFT processors and high throughput relative to past subthreshold circuit implementations. Dongsuk Jeon, Mingoo Seok, Chaitali Chakrabarti, David T. Blaauw, Dennis Sylvester |
ICASSP | 1 |