VLDB 2026 Research / reviewers in the wild / expert
Jong Hwan Ko
dblp:168/6308
· DBLP profile ↗
83ranked-venue papers
7as first author
66since 2021 · last 2026
0000-0003-4434-4318ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 36 · 4 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-author · 25 since 2021Artificial intelligence and machine learning · 28 · 25 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 7 since 2021Computer networks · 4 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data Flow-Aware Weight Remapping for Efficient Fault Tolerance in ReRAM-Based AcceleratorsabstractResistive random-access memory (ReRAM)-based in-memory computing (IMC) systems offer significant advantages for efficient neural network inference. However, these systems are vulnerable to stuck-at faults (SAFs), which degrade inference accuracy-a challenge that becomes more pronounced in multilevel cell (MLC) configurations. A conventional fault mitigation technique, array-wise weight remapping (AWR), addresses SAFs but incurs significant hardware overhead. To overcome these limitations, we propose pseudo-array-wise weight remapping (PAWR), a novel method that integrates mux-wise weight remapping (MWR) and mux group remapping (MGR) to achieve costefficient fault tolerance. Experimental results demonstrate that even at a high 20% SAF rate, PAWR achieves accuracies within 0.7% of AWR, while significantly reducing area overhead by 86.5% and energy overhead by 72.4% compared to AWR. Hyeonsu Bang, Kang Eun Jeon, Jong Hwan Ko |
ASP-DAC | 3 |
| 2026 | Learnable Center-Based Quantization for Efficient Analog PIM with Reduced ADC PrecisionabstractProcessing-in-memory (PIM) architectures have shown significant potential for accelerating deep neural network (DNN) inference by performing matrix-vector multiplications directly within memory. However, achieving high precision often requires high-resolution analog-to-digital converters (ADCs), which can increase energy consumption and limit overall efficiency. To address this, we propose a learnable center-based quantization (LCQ) technique that minimizes the range of partial sums in PIM arrays. This reduction in the range of partial sums decreases the ADC resolution requirements, enabling accurate low-bit quantization while maintaining energy efficiency. Our framework directly models ADC precision constraints within the training process without requiring extensive retraining. Experimental results on DNN models such as ResNet20 and ResNet18 with CIFAR-10/ImageNet datasets demonstrate that LCQ significantly enhances energy efficiency while maintaining competitive accuracy compared to previous techniques for efficient analog PIM. LCQ improves both accuracy and energy efficiency, reducing ADC resolution requirements and enabling practical low-bit quantization. Sangheum Yeon, Jong Hwan Ko |
ASP-DAC | 2 |
| 2026 | Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource VarietiesabstractLow-resource language varieties used by specific groups remain neglected in the development of Multilingual Language Models.A great deal of cross-lingual research focuses on inter-lingual language transfer which strives to align allied varieties and minimize differences between them.However, for low-resource varieties, linguistic dissimilarity is also an important cue allowing generalization to unseen varieties.Unlike prior approaches, we propose a two-stage Language Generalization framework that focuses on capturing variety-specific cues while also exploiting rich overlap offered by high-resource source variety.First, we propose TOPPing, a source-selection method specifically designed for low-resource varieties.Second, we suggest a lightweight VAÇAÍ-Bowl architecture that learns variety-specific attributes with one branch while a parallel branch captures variety-invariant attributes using adversarial training.We evaluate our framework on structural prediction tasks, which are among the few tasks available, as proxy for performance on other downstream tasks.Using VAÇAÍ-Bowl with TOPPing yields an average 54.62% improvement in the dependency parsing task, which serves as a proxy for performance on other downstream tasks across 10 low-resource varieties.1 How can we transfer knowledge from source variety to target variety?Alignment Only variety-specific features are collapsed variety-specific features are preserved Variety Representation Variety-Aware Generalization Training Process Jinju Kim, Haeji Jung, Youjeong Roh, Jong Hwan Ko, David R. Mortensen |
CoNLL | 4 |
| 2026 | RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNs
Hanum Ko, Sangheum Yeon, Jong Hwan Ko, Jungrae Kim |
ISCA | 3 |
| 2026 | Dataflow-Preserving, Overhead-Free Weight Remapping for Fault-Tolerant ReRAM-Based In-Memory ComputingabstractResistive random-access memory (ReRAM)-based in-memory computing (IMC) systems provide high energy efficiency and storage density for deep neural network (DNN) acceleration, but stuck-at faults (SAFs) substantially degrade reliability. Weight remapping (WR) can mitigate SAFs; however, existing approaches either ignore dataflow consistency or require additional hardware to restore it. We propose FREEMAP, an overhead-free WR algorithm that preserves dataflow consistency without runtime operations or hardware modifications. FREEMAP combines layer-wise filter reordering (LFR), which globally reorders filters across a layer for fault resilience, with row group remapping (RGR), which realigns the next layer's weight rows to the reordered outputs. Across diverse models and datasets, FREEMAP eliminates the hardware overhead of conventional WR, reducing area and energy to 0.02×-0.06× and 0.04×-0.14×, respectively. Compared with state-of-the-art dataflow-aware WR, it further reduces area and energy to 0.45×-0.47× and 0.52×-0.63× while maintaining comparable accuracy. Hyeonsu Bang, Jong Hwan Ko |
ISLPED | 2 |
| 2026 | HyperSPACE: Sparse-Adder-Compatible Encoding for Efficient Hyperdimensional Computing on Digital CIM Arrays
Yeong Hwan Oh, Do Yeong Kang, Juhong Park, Chanwook Hwang, Kang Eun Jeon, Jong Hwan Ko |
ISLPED | 6 |
| 2026 | A Neural-Feedback-Driven Event Camera for Robust and Efficient Vision Processing
Jaehyeon So, Houk Lee, Chanwook Hwang, Woosung Chung, Jong Hwan Ko |
ISLPED | 6 |
| 2026 | Single-step Diffusion for Image Compression at Ultra-Low BitratesabstractAlthough there have been significant advancements in image compression techniques, such as standard and learned codecs, these methods still suffer from severe quality degradation at extremely low bits per pixel. While recent diffusion-based models provided enhanced generative performance at low bitrates, they often yield limited perceptual quality and prohibitive decoding latency due to multiple denoising steps. In this paper, we propose the single-step diffusion model for image compression that delivers high perceptual quality and fast decoding at ultra-low bitrates. Our approach incorporates two key innovations: (i) Vector-Quantized Residual (VQ-Residual) training, which factorizes a structural base code and a learned residual in latent space, capturing both global geometry and high-frequency details; and (ii) rate-aware noise modulation, which tunes denoising strength to match the desired bitrate. Extensive experiments show that ours achieves comparable compression performance to state-of-the-art methods while improving decoding speed by about 50× compared to prior diffusion-based methods, greatly enhancing the practicality of generative codecs. Chanung Park, Joo Chan Lee, Jong Hwan Ko |
WACV | 3 |
| 2026 | Adaptive Quantum Transformer: Adapter-Based Parameter Reuse for Dynamic Qubit Scaling in Vision TransformersabstractTransformers excel in various domains but often demand massive computational resources. Quantum Neural Networks (QNNs) combine classical and quantum computing to mitigate these costs. However, QNNs typically fix the number of qubits, forcing complete retraining whenever the qubit count changes and leading to substantial inefficiency. We propose an Adaptive Quantum Transformer (AQT) to resolve this limitation by introducing an adapter module — originally designed for large language models (LLMs) — into a quantum vision Transformer. This module reuses parameters when the qubit count expands or contracts, preserving learned states. We validate AQT in an autoscaling cloud environment, demonstrating dynamic qubit scaling without sacrificing accuracy or retraining from scratch. Results on vision tasks reveal that AQT achieves higher accuracy than a conventional quantum vision Transformer while converging faster and requiring fewer computational resources. Notably, AQT provides over triple the efficiency in training time and better stability under fluctuating qubit availability, suggesting a promising avenue for practical, cost-effective quantum–classical hybrid intelligence. Hyochan Kim, Jong Hwan Ko |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2026 | Efficient Vision Transformer Inference via UDP for Edge-Cloud Collaboration: An Adaptive Loss Detection Approach
Hyochan Kim, Jong Hwan Ko |
J. Comput. Sci. Technol. | 2 |
| 2026 | RUnQuant: High-resolution weight quantization via unanchored weight decomposition in column-wise granularity for CIM accelerators
Kang Eun Jeon, Yulhwa Kim, Jong Hwan Ko |
J. Syst. Archit. | 4 |
| 2026 | EPS: Efficient Patch Sampling for Video Overfitting in Deep Super-Resolution Model TrainingabstractLeveraging the overfitting property of deep neural networks (DNNs) is trending in video delivery systems to enhance video quality within bandwidth limits. Existing approaches transmit overfitted super-resolution (SR) model streams for low-resolution (LR) bitstreams, which are used to reconstruct high-resolution (HR) videos at the decoder. Although these approaches show promising results, the huge computational costs of training a large number of video frames limit their practical applications. To overcome this challenge, we propose an efficient patch sampling method named EPS for video SR network overfitting, which identifies the most valuable training patches from video frames. To this end, we first present two low-complexity Discrete Cosine Transform (DCT)-based spatial-temporal features to measure the complexity score of each patch directly. By analyzing the histogram distribution of these features, we then categorize all possible patches into different clusters and select training patches from the cluster with the highest spatial-temporal information. The number of sampled patches is adaptive based on the video content, addressing the trade-off between training complexity and efficiency. Our method reduces the number of training patches by 75.00% to 91.69%, depending on the resolution and number of clusters, while preserving high video quality and greatly improving training efficiency. Our method speeds up patch sampling by up to 82.1× compared to the state-of-the-art patch sampling technique (EMT). Yiying Wei, Hadi Amirpour, Jong Hwan Ko, Christian Timmerer |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Multi-Frame ISP: Enhancing Vision-Based Tasks with RAW and Infrared VideosabstractRecent vision-based downstream tasks perform well under standard conditions but degrade in low-light or high-exposure environments due to their reliance on well-lit datasets aligned with the human visual system. To address this, we propose a Multi-Frame ISP method that utilizes RAW and IR video frames for enhanced robustness. RAW images retain rich environmental details for low-light scenarios, while IR images remain unaffected by visible light. Unlike traditional ISP pipelines or paired translation models, our approach introduces Global ISP for image-wide color correction by Selective ISP module and Local ISP for region-specific enhancement using temporal information. Experiments on the RAW video datasets ImageVID and YouTubeVOS, and the IR dataset FLIR, show that our method outperforms conventional ISP and translation models by dynamically adapting to lighting variations, thereby enhancing object detection and segmentation. Ji Seok Kim, Jong Chul Ahn, Sang Min Kim, Jong Hwan Ko |
AVSS | 4 |
| 2025 | Test-Time Fine-Tuning of Image Compression Models for Multi-Task AdaptabilityabstractThe field of computer vision was initially inspired by the human visual system and has progressively expanded to include a broader range of machine vision applications. Consequently, image compressors should be designed to effectively accommodate not only human visual perception but also machine vision tasks, including closed-set scenarios that enable pre-training and open-set scenarios that involve previously unseen tasks at test time. Many recent studies effectively address both human visual perception and closed-set machine vision tasks simultaneously but struggle to handle open-set machine vision tasks. To address this issue, this paper proposes a fully instance-specific test time fine-tuning (TTFT) for adapting learned image compression (LIC) to both closed-set and open-set machine vision tasks effectively. With our method, a large-scale LIC model, originally trained for human perception, is adapted to the target task through TTFT using Singular Value Decomposition based Low Rank Adaptation (SVD-LoRA). During TTFT, the decoder adopts a modified learning scheme that focuses exclusively on training the singular values, which helps prevent excessive bitstream overhead. This enables fully instance-specific optimization for the target task, even for open-set tasks. Experimental results demonstrate that the proposed method effectively adapts the backbone compressor to diverse machine vision tasks, outperforming competing methods. The code is available at project page. Unki Park, Seongmoon Jeong, Youngchan Jang, Gyeong-Moon Park, Jong Hwan Ko |
CVPR | 5 |
| 2025 | Low-Rank Compression for IMC ArraysabstractIn this study, we address the challenge of low-rank model compression in the context of in-memory computing (IMC) architectures. Traditional pruning approaches, while effective in model size reduction, necessitate additional peripheral circuitry to manage complex dataflows and mitigate dislocation issues, leading to increased area and energy overheads. To circumvent these drawbacks, we propose leveraging low-rank compression techniques, which, unlike pruning, streamline the dataflow and seamlessly integrate with IMC architectures. However, low-rank compression presents its own set of challenges, namely i) suboptimal IMC array utilization and ii) compromised accuracy. To address these issues, we introduce a novel approach i) employing shift and duplicate kernel (SDK) mapping technique, which exploits idle IMC columns for parallel processing, and ii) group lowrank convolution, which mitigates the information imbalance in the decomposed matrices. Our experimental results demonstrate that our proposed method achieves up to 2.5× speedup or +20.9% accuracy boost over existing pruning techniques. Kang Eun Jeon, Johnny Rhe, Jong Hwan Ko |
DATE | 3 |
| 2025 | MEMHD: Memory-Efficient Multi-Centroid Hyperdimensional Computing for Fully-Utilized In-Memory Computing ArchitecturesabstractThe implementation of Hyperdimensional Computing (HDC) on In-Memory Computing (IMC) architectures faces significant challenges due to the mismatch between high-dimensional vectors and IMC array sizes, leading to inefficient memory utilization and increased computation cycles. This paper presents MEMHD, a Memory-Efficient Multi-centroid HDC framework designed to address these challenges. MEMHD introduces a clustering-based initialization method and quantization-aware iterative learning for multi-centroid associative memory. Through these approaches and its overall architecture, MEMHD achieves a significant reduction in memory requirements while maintaining or improving classification accuracy. Our approach achieves full utilization of IMC arrays and enables one-shot (or few-shot) associative search. Experimental results demonstrate that MEMHD outperforms state-of-the-art binary HDC models, achieving up to 13.69% higher accuracy with the same memory usage, or 13.25x more memory efficiency at the same accuracy level. Moreover, MEMHD reduces computation cycles by up to 80x and array usage by up to 71x compared to baseline IMC mapping methods when mapped to 128x128 IMC arrays, while significantly improving energy and computation cycle efficiency. Do Yeong Kang, Yeong Hwan Oh, Chanwook Hwang, Jinhee Kim, Kang Eun Jeon, Jong Hwan Ko |
DATE | 6 |
| 2025 | Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory AcceleratorsabstractCompute-in-memory (CIM) is an efficient method for implementing deep neural networks (DNNs) but suffers from substantial overhead from analog-to-digital converters (ADCs), especially as ADC precision increases. Low-precision ADCs can reduce this overhead but introduce partial-sum quantization errors degrading accuracy. Additionally, low-bit weight constraints, imposed by cell limitations and the need for multiple cells for higher-bit weights, present further challenges. While fine-grained partial-sum quantization has been studied to lower ADC resolution effectively, weight granularity, which limits overall partial-sum quantized accuracy, remains underexplored. This work addresses these challenges by aligning weight and partial-sum quantization granularities at the column-wise level. Our method improves accuracy while maintaining dequantization overhead, simplifies training by removing two-stage processes, and ensures robustness to memory cell variations via independent column-wise scale factors. We also propose an open-source CIM-oriented convolution framework to handle fine-grained weights and partial-sums efficiently, incorporating a novel tiling method and group convolution. Experimental results on ResNet-20 (CIFAR-10, CIFAR-100) and ResNet-18 (ImageNet) show accuracy improvements of 0.99%, 2.69%, and 1.01%, respectively, compared to the best-performing related works. Additionally, variation analysis reveals the robustness of our method against memory cell variations. These findings highlight the effectiveness of our quantization scheme in enhancing accuracy and robustness while maintaining hardware efficiency in CIM-based DNN implementations. Our code is available at https://github.com/jiyoonkm/ColumnQuant. Kang Eun Jeon, Yulhwa Kim, Jong Hwan Ko |
DATE | 4 |
| 2025 | Row-Column Hybrid Grouping for Fault-Resilient Multi-Bit Weight Representation on IMC ArraysabstractThis paper addresses two critical challenges in analog In-Memory Computing (IMC) systems that limit their scalability and deployability: the computational unreliability caused by stuck-at faults (SAFs) and the high compilation overhead of existing fault-mitigation algorithms, namely Fault-Free (FF). To overcome these limitations, we first propose a novel multi-bit weight representation technique, termed row-column hybrid grouping, which generalizes conventional column grouping by introducing redundancy across both rows and columns. This structural redundancy enhances fault tolerance and can be effectively combined with existing fault-mitigation solutions. Second, we design a compiler pipeline that reformulates the fault-aware weight decomposition problem as an Integer Linear Programming (ILP) task, enabling fast and scalable compilation through off-the-shelf solvers. Further acceleration is achieved through theoretical insights that identify fault patterns amenable to trivial solutions, significantly reducing computation. Experimental results on convolutional networks and small language models demonstrate the effectiveness of our approach, achieving up to 8%p improvement in accuracy, 150 × faster compilation, and 2 × energy efficiency gain compared to existing baselines. Kang Eun Jeon, Sangheum Yeon, Jinhee Kim, Hyeonsu Bang, Johnny Rhe, Jong Hwan Ko |
ICCAD | 6 |
| 2025 | MSQ: Memory-Efficient Bit Sparsification Quantization
Seokho Han, Seoyeon Yoon, Jinhee Kim, Dongwei Wang, Kang Eun Jeon, Huanrui Yang, Jong Hwan Ko |
ICCV | 7 |
| 2025 | Do Not Mimic My Voice : Speaker Identity Unlearning for Zero-Shot Text-to-SpeechabstractThe rapid advancement of Zero-Shot Text-to-Speech (ZS-TTS) technology has enabled high-fidelity voice synthesis from minimal audio cues, raising significant privacy and ethical concerns. Despite the threats to voice privacy, research to selectively remove the knowledge to replicate unwanted individual voices from pre-trained model parameters has not been explored. In this paper, we address the new challenge of speaker identity unlearning for ZS-TTS systems. To meet this goal, we propose the first machine unlearning frameworks for ZS-TTS, especially Teacher-Guided Unlearning (TGU), designed to ensure the model forgets designated speaker identities while retaining its ability to generate accurate speech for other speakers. Our proposed methods incorporate randomness to prevent consistent replication of forget speakers' voices, assuring unlearned identities remain untraceable. Additionally, we propose a new evaluation metric, speaker-Zero Retrain Forgetting (spk-ZRF). This assesses the model's ability to disregard prompts associated with forgotten speakers, effectively neutralizing its knowledge of these voices. The experiments conducted on the state-of-the-art model demonstrate that TGU prevents the model from replicating forget speakers' voices while maintaining high quality for other speakers.
The demo is available at https://speechunlearn.github.io/ . Taesoo Kim, Jinju Kim, Jong Hwan Ko, Gyeong-Moon Park |
ICML | 4 |
| 2025 | Event-based Neural Spike Detection Using Spiking Neural Networks for Neuromorphic iBMI SystemsabstractImplantable brain-machine interfaces (iBMIs) are evolving to record from thousands of neurons wirelessly but face challenges in data bandwidth, power consumption, and implant size. We propose a novel Spiking Neural Network Spike Detector (SNN-SPD) that processes event-based neural data generated via delta modulation and pulse count modulation, converting signals into sparse events. By leveraging the temporal dynamics and inherent sparsity of spiking neural networks, our method improves spike detection performance while maintaining low computational overhead suitable for implantable devices. Our experimental results demonstrate that the proposed SNN-SPD achieves an accuracy of 95.72% at high noise levels (standard deviation 0.2), which is about 2% higher than the existing Artificial Neural Network Spike Detector (ANN-SPD). Moreover, SNN-SPD requires only 0.41% of the computation and about 26.62% of the weight parameters compared to ANN-SPD, with zero multiplications. This approach balances efficiency and performance, enabling effective data compression and power savings for next-generation iBMIs. Chanwook Hwang, Biyan Zhou, Ye Ke, Vivek Mohan, Jong Hwan Ko, Arindam Basu |
ISCAS | 5 |
| 2025 | TruncQuant: Truncation-Ready Quantization for DNNs with Flexible Weight Bit PrecisionabstractThe deployment of deep neural networks on edge devices is a challenging task due to the increasing complexity of state-of-the-art models, requiring efforts to reduce model size and inference latency. Recent studies explore models operating at diverse quantization settings to find the optimal point that balances computational efficiency and accuracy. Truncation, an effective approach for achieving lower bit precision mapping, enables a single model to adapt to various hardware platforms with little to no cost. However, formulating a training scheme for deep neural networks to withstand the associated errors introduced by truncation remains a challenge, as the current quantization-aware training schemes are not designed for the truncation process. We propose TruncQuant, a novel truncation-ready training scheme allowing flexible bit precision through bit-shifting in runtime. We achieve this by aligning TruncQuant with the output of the truncation process, demonstrating strong robustness across bit-width settings, and offering an easily implementable training scheme within existing quantization-aware frameworks. Our code is released at https://github.com/a2jinhee/TruncQuant. Jinhee Kim, Seoyeon Yoon, Joo Chan Lee, Kang Eun Jeon, Jong Hwan Ko |
ISLPED | 6 |
| 2025 | Efficient Multi-bit Quantization Network Training via Weight Bias Correction and Bit-wise Coreset SamplingabstractMulti-bit quantization networks enable flexible deployment of deep neural networks by supporting multiple precision levels within a single model. However, existing approaches suffer from significant training overhead as full-dataset updates are repeated for each supported bit-width, resulting in a cost that scales linearly with the number of precisions. Additionally, extra fine-tuning stages are often required to support additional or intermediate precision options, further compounding the overall training burden. To address this issue, we propose two techniques that greatly reduce the training overhead without compromising model utility: (i) Weight bias correction enables shared batch normalization and eliminates the need for fine-tuning by neutralizing quantization-induced bias across bit-widths and aligning activation distributions; and (ii) Bit-wise coreset sampling strategy allows each child model to train on a compact, informative subset selected via gradient-based importance scores by exploiting the implicit knowledge transfer phenomenon. Experiments on CIFAR-10/100, TinyImageNet, and ImageNet-1K with both ResNet and ViT architectures demonstrate that our method achieves competitive or superior accuracy while reducing training time up to 7.88×. Jinhee Kim, Jae Jun An, Kang Eun Jeon, Jong Hwan Ko |
NeurIPS | 4 |
| 2025 | Optimized Minimal 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has emerged as a powerful representation for real-time, high-performance rendering, enabling a wide range of applications. However, representing 3D scenes with numerous explicit Gaussian primitives imposes significant storage and memory overhead. Recent studies have shown that high-quality rendering can be achieved with a substantially reduced number of Gaussians when represented with high-precision attributes. Nevertheless, existing 3DGS compression methods still rely on a relatively large number of Gaussians, focusing primarily on attribute compression. This is because a smaller set of Gaussians becomes increasingly sensitive to lossy attribute compression, leading to severe quality degradation. Since the number of Gaussians is directly tied to computational costs, it is essential to reduce the number of Gaussians effectively rather than only optimizing storage. In this paper, we propose Optimized Minimal Gaussians representation (OMG), which significantly reduces storage while using a minimal number of primitives. First, we determine the distinct Gaussian from the near ones, minimizing redundancy without sacrificing quality. Second, we propose a compact and precise attribute representation that efficiently captures both continuity and irregularity among primitives. Additionally, we propose a sub-vector quantization technique for improved irregularity representation, maintaining fast training with a negligible codebook size. Extensive experiments demonstrate that OMG reduces storage requirements by nearly 50% compared to the previous state-of-the-art and enables 600+ FPS rendering while maintaining high rendering quality. Our source code is available at https://maincold2.github.io/omg/. Joo Chan Lee, Jong Hwan Ko, Eunbyung Park |
NeurIPS | 2 |
| 2025 | Perturb a Model, Not an Image: Towards Robust Privacy Protection via Anti-Personalized Diffusion ModelsabstractRecent advances in diffusion models have enabled high-quality synthesis of specific subjects, such as identities or objects. This capability, while unlocking new possibilities in content creation, also introduces significant privacy risks, as personalization techniques can be misused by malicious users to generate unauthorized images. Although several studies have attempted to counter this by generating adversarially perturbed samples designed to disrupt personalization, they rely on unrealistic assumptions and become ineffective in the presence of even a few clean images or under simple image transformations. To address these challenges, we shift the protection target from the images to the diffusion model itself to hinder the personalization of specific subjects, through our novel framework called $\textbf{A}$nti-$\textbf{P}$ersonalized $\textbf{D}$iffusion $\textbf{M}$odels ($\textbf{APDM}$). We first provide a theoretical analysis demonstrating that a naive approach of existing loss functions to diffusion models is inherently incapable of ensuring convergence for robust anti-personalization. Motivated by this finding, we introduce Direct Protective Optimization (DPO), a novel loss function that effectively disrupts subject personalization in the target model without compromising generative quality. Moreover, we propose a new dual-path optimization strategy, coined Learning to Protect (L2P). By alternating between personalization and protection paths, L2P simulates future personalization trajectories and adaptively reinforces protection at each step.
Experimental results demonstrate that our framework outperforms existing methods, achieving state-of-the-art performance in preventing unauthorized personalization.
The code is available at https://github.com/KU-VGI/APDM. Tae-Young Lee, Juwon Seo, Jong Hwan Ko, Gyeong-Moon Park |
NeurIPS | 3 |
| 2025 | Continual Test-Time Fine-Tuning of Frame-Based Style Transfer Network for Video Stream DataabstractNeural network-based style transfer has long attracted research interest and continues to be an active area today. While image style transfer (IST) methods have achieved impressive advancements in arbitrary style transfer under low per-frame latency constraints, video style transfer (VST) continues to struggle due to the additional challenge of maintaining temporal consistency between frames. Despite recent rapid progress, existing VST methods still face critical limitations. Some approaches rely on inefficient components, such as optical flow, to capture temporal consistency, which undermines the goals of fast style transfer. In contrast, frame-wise stylization methods without inter-frame information meet fast transfer requirements but struggle to mitigate flickering artifacts. To address these challenges, we propose a novel Continual Test-Time Fine-Tuning (C-TTFT) framework that dynamically adapts a pre-trained frame-based style transfer network during inference. Our method leverages lightweight inter-frame information by using the previously stylized output as guidance for the current frame, optimizing simple loss objectives in real time. Furthermore, we incorporate a parameter-efficient adaptation technique, Low-Rank Adaptation (LoRA), to reduce memory and computational overhead. We validate C-TTFT on MicroAST, a state-of-the-art, frame-based style transfer backbone, and demonstrate that our method reduces temporal consistency loss by up to 17% and enhances SSIM, CFSD, and ArtFID by up to 22%, 33%, and 8%, respectively, all while preserving real-time performance (24 FPS) with minimal added cost (10% more parameters, +48K; 4.9% more memory usage, +93KB). Unki Park, Jong Hwan Ko |
VCIP | 2 |
| 2025 | An FPGA-Based Energy-Efficient Real-Time Hand Pose Estimation System With an Integrated Image Signal Processor for Indirect 3-D Time-of-Flight SensorsabstractAs artificial intelligence (AI) technology advances, Internet of Things (IoT) devices, such as mobile phones and augmented reality devices, are increasingly becoming crucial enablers of user-device interactions. Among the various methods of interaction, hand pose recognition and analysis is a crucial method to understand the intentions of users and perform precise functions. However, to perform such functions, a substantial amount of computation and resources are required, making it challenging to implement them on small form-factor devices with low-power consumption. For this reason, improving energy efficiency is a crucial objective in real-time hand pose estimation (HPE) applied to low-power platforms with limited resources. In this article, we introduce an FPGA-based energy-efficient real-time HPE system with an integrated image signal processor (ISP). The proposed system uses several low-power design techniques, including a systolic array with dynamic on/off control per processing element (PE), to minimize power consumption and save energy when not in use. In addition, we improve area efficiency by reducing the buffer size in the systolic array using a half-size shift buffer stack. Furthermore, the use of parallel and pipelined structures improved operational efficiency, resulting in a reduction in both operational time and power consumption. The evaluation results on a KU115 FPGA board show that the system achieves an error of 7.78 mm and can process 52 fps, demonstrating its capability for real-time HPE. Moreover, this system achieves high-energy efficiency, up to 61.74 GOPs/W, making it suitable for energy-efficient and accurate HPE in low-power environments. Yongsoo Kim, Jaehyeon So, Chanwook Hwang, Wencan Cheng, Jaehyuk Choi 0001, Jong Hwan Ko |
IEEE Internet Things J. | 6 |
| 2025 | Input/mapping precision controllable digital CIM with adaptive adder tree architecture for flexible DNN inference
Juhong Park, Johnny Rhe, Chanwook Hwang, Jaehyeon So, Jong Hwan Ko |
J. Syst. Archit. | 5 |
| 2025 | GAROS: Genetic algorithm-aided row-skipping for shift and duplicate kernel mapping in processing-in-memory architectures
Johnny Rhe, Kang Eun Jeon, Jong Hwan Ko |
J. Syst. Archit. | 3 |
| 2024 | HandDiff: 3D Hand Pose Estimation with Diffusion on Image-Point CloudabstractExtracting keypoint locations from input hand frames, known as 3D hand pose estimation, is a critical task in various human-computer interaction applications. Essentially, the 3D hand pose estimation can be regarded as a 3D point subset generative problem conditioned on input frames. Thanks to the recent significant progress on diffusion-based generative models, hand pose estimation can also benefit from the diffusion model to estimate keypoint locations with high quality. However, directly deploying the existing diffusion models to solve hand pose estimation is non-trivial, since they cannot achieve the complex permutation mapping and precise localization. Based on this motivation, this paper proposes HandDiff, a diffusion-based hand pose estimation model that iteratively denoises accurate hand pose conditioned on hand-shaped image-point clouds. In order to recover keypoint permutation and accurate location, we further introduce Joint-wise condition and local detail condition. Experimental results demonstrate that the proposed HandDiff significantly outperforms the existing approaches on four challenging hand pose benchmark datasets. Codes and pre-trained models are publicly available at https://github.com/cwc1260/HandDiff. Wencan Cheng, Hao Tang 0005, Luc Van Gool, Jong Hwan Ko |
CVPR | 4 |
| 2024 | Compact 3D Gaussian Representation for Radiance FieldabstractNeural Radiance Fields (NeRFs) have demonstrated re-markable potential in capturing complex 3D scenes with high fidelity. However, one persistent challenge that hin-ders the widespread adoption of NeRFs is the computational bottleneck due to the volumetric rendering. On the other hand, 3D Gaussian splatting (3DGS) has recently emerged as an alternative representation that leverages a 3D Gaussisan-based representation and adopts the ras-terization pipeline to render the images rather than volumetric rendering, achieving very fast rendering speed and promising image quality. However, a significant draw-back arises as 3DGS entails a substantial number of 3D Gaussians to maintain the high fidelity of the rendered images, which requires a large amount of memory and stor-age. To address this critical issue, we place a specific emphasis on two key objectives: reducing the number of Gaussian points without sacrificing performance and compressing the Gaussian attributes, such as view-dependent color and covariance. To this end, we propose a learnable mask strategy that significantly reduces the number of Gaussians while preserving high performance. In addition, we propose a compact but effective representation of view-dependent color by employing a grid-based neural field rather than relying on spherical harmonics. Finally, we learn codebooks to compactly represent the geometric attributes of Gaussian by vector quantization. With model compression techniques such as quantization and entropy coding, we consistently show over 25× reduced storage and enhanced rendering speed, while maintaining the quality of the scene representation, compared to 3DGS. Our work provides a comprehensive framework for 3D scene representation, achieving high performance, fast training, compact-ness, and real-time rendering. Our project page is available at https://mainco/d2.github.io/c3dgs/. Joo Chan Lee, Daniel Rho, Jong Hwan Ko, Eunbyung Park |
CVPR | 4 |
| 2024 | TraiNDSim: A Simulation Framework for Comprehensive Performance Evaluation of Neuromorphic Devices for On-Chip TrainingabstractThe advancement of neuromorphic devices (NDs) for processing deep neural networks has narrowed the accuracy gap with software-trained models. To accurately assess ND performance, reliable simulation frameworks for on-chip training are crucial. However, existing frameworks encounter difficulties accurately reflecting the characteristics of NDs in training simulations. Consequently, we introduce TraiNDSim, a novel framework that comprehensively evaluates the performance of NDs to address these difficulties. Specifically, we propose an advanced conductance normalization strategy called layer-wise normalization, which limits the weight range by taking the initial weight distribution into account. Additionally, our framework integrates three conductance models, notably refining one of the conventional models to depend solely on nonlinearity. Moreover, it features a bi-directional weight representation method with a unique conductance compensation technique. Our comprehensive analysis using TraiNDSim demonstrates its effectiveness in accurately reflecting the impact of ND parameters on training, promising more precise device performance evaluations. Our framework is available at https://github.com/donghyeokheo/TraiNDSim. Donghyeok Heo, Hyeonsu Bang, Jong Hwan Ko |
DAC | 3 |
| 2024 | HandDAGT: A Denoising Adaptive Graph Transformer for 3D Hand Pose Estimation
Wencan Cheng, Jong Hwan Ko |
ECCV (88) | 3 |
| 2024 | Continuous Memory Representation for Anomaly Detection
Joo Chan Lee, Taejune Kim, Eunbyung Park, Simon S. Woo, Jong Hwan Ko |
ECCV (51) | 5 |
| 2024 | Coordinate-Aware Modulation for Neural FieldsabstractNeural fields, mapping low-dimensional input coordinates to corresponding signals, have shown promising results in representing various signals. Numerous methodologies have been proposed, and techniques employing MLPs and grid representations have achieved substantial success. MLPs allow compact and high expressibility, yet often suffer from spectral bias and slow convergence speed. On the other hand, methods using grids are free from spectral bias and achieve fast training speed, however, at the expense of high spatial complexity. In this work, we propose a novel way for exploiting both MLPs and grid representations in neural fields. Unlike the prevalent methods that combine them sequentially (extract features from the grids first and feed them to the MLP), we inject spectral bias-free grid representations into the intermediate features in the MLP. More specifically, we suggest a Coordinate-Aware Modulation (CAM), which modulates the intermediate features using scale and shift parameters extracted from the grid representations. This can maintain the strengths of MLPs while mitigating any remaining potential biases, facilitating the rapid learning of high-frequency components. In addition, we empirically found that the feature normalizations, which have not been successful in neural filed literature, proved to be effective when applied in conjunction with the proposed CAM. Experimental results demonstrate that CAM enhances the performance of neural representation and improves learning stability across a range of signals. Especially in the novel view synthesis task, we achieved state-of-the-art performance with the least number of parameters and fast training speed for dynamic scenes and the best performance under 1MB memory for static scenes. CAM also outperforms the best-performing video compression methods using neural fields by a large margin. Our project page is available at https://maincold2.github.io/cam/. Joo Chan Lee, Daniel Rho, Seungtae Nam, Jong Hwan Ko, Eunbyung Park |
ICLR | 4 |
| 2024 | DynaPP: A Dynamic Resolution Model with Patch Packing for Fast Online Video DetectionabstractOnline video detection becomes more challenging with higher resolution as computational costs increase proportionally with increasing resolution. To address this issue, we present a novel approach, DynaPP, which arranges object candidate regions into a compact form. DynaPP performs resource-intensive whole-image inference only on sparse key frames, employing reduced resolutions for inference on other frames. Additionally, we propose transforming a 1-stage detector into a dynamic resolution model to facilitate frame inference at reduced resolutions. Here, the dynamic resolution model signifies a model capable of inferring all resolutions, distinguishing itself from typical models by not having restricted inferable resolutions. Unlike prior studies introducing new model structures for multi-resolution models, our work demonstrates that slight modifications to existing models can convert them to dynamic resolution models. DynaPP showcases substantial acceleration in video detection across four representative video datasets: AU-AIR (5.5×), UAVDT (3.67×), VisDrone (2.73×), and ImageNet VID (3.69×), while maintaining a mean average precision with a small loss (≤2.2). Furthermore, we observed that our method achieves a detection acceleration of up to 8.84×, depending on the video clip. Changrok So, Simon S. Woo, Jong Hwan Ko |
IJCNN | 3 |
| 2024 | KARS: Kernel-Grouping Aided Row-Skipping for SDK-based Weight Compression in PIM ArraysabstractWith energy-efficient computation, processing-in-memory (PIM) architectures have been highlighted as one of the most viable candidates to substitute the traditional ones. Recently, shift and duplicate kernel (SDK) mapping method was proposed to enable efficient and fast convolutional neural networks (CNNs) inference in the PIM array. However, since its weight deployment to reuse input data, this method generates idle cells that do not involved in the computation, which leads to an increase of energy consumption. In this paper, we propose a novel weight mapping method called kernel-grouping aided row-skipping (KARS). KARS maximizes utilization by removing idle cells on a PIM array and reduces computing cycles. In comparison to the traditional methods, KARS achieves a speedup by up to 3× at Layer 2 of VGGNet-13 and ResNet-18. Juhong Park, Johnny Rhe, Jong Hwan Ko |
ISCAS | 3 |
| 2024 | F-3DGS: Factorized Coordinates and Representations for 3D Gaussian SplattingabstractThe neural radiance field (NeRF) has made significant strides in representing 3D scenes and synthesizing novel views. Despite its advancements, the high computational costs of NeRF have posed challenges for its deployment in resource-constrained environments and real-time applications. As an alternative to NeRF-like neural rendering methods, 3D Gaussian Splatting (3DGS) offers rapid rendering speeds while maintaining excellent image quality. However, as it represents objects and scenes using a myriad of Gaussians, it requires substantial storage to achieve high-quality representation. To mitigate the storage overhead, we propose Factorized 3D Gaussian Splatting (F-3DGS), a novel approach that drastically reduces storage requirements while preserving image quality. Inspired by classical matrix and tensor factorization techniques, our method represents and approximates dense clusters of Gaussians with significantly fewer Gaussians through efficient factorization. We aim to efficiently represent dense 3D Gaussians by approximating them with a limited amount of information for each axis and their combinations. This method allows us to encode a substantially large number of Gaussians along with their essential attributes'such as color, scale, and rotation-necessary for rendering using a relatively small number of elements. Extensive experimental results demonstrate that F-3DGS achieves a significant reduction in storage costs while maintaining comparable quality in rendered images. Our project page is available at https://xiangyu1sun.github.io/Factorize-3DGS/. Joo Chan Lee, Daniel Rho, Jong Hwan Ko, Usman Ali 0006, Eunbyung Park |
ACM Multimedia | 4 |
| 2024 | Energy-Efficient Video Streaming: A Study on Bit Depth and Color SubsamplingabstractAs video dimensions – including resolution, frame rate, and bit depth – increase, a larger bitrate is required to maintain a higher Quality of Experience (QoE). While videos are often optimized for resolution and frame rate to improve compression and energy efficiency, the impact of color space is often overlooked. Larger color spaces are essential for avoiding color banding and delivering High Dynamic Range (HDR) content with richer, more accurate colors, although this comes at the cost of higher processing energy. This paper investigates the effects of bit depth and color subsampling on video compression efficiency and energy consumption. By analyzing different bit depths and subsampling schemes, we aim to determine optimized settings that balance compression efficiency with energy consumption, ultimately contributing to more sustainable and high-quality video delivery. We evaluate both encoding and decoding energy consumption and assess the quality of videos using various metrics including PSNR, VMAF, ColorVideoVDP, and CAMBI. Our findings offer valuable insights for video codec developers and content providers aiming to improve the performance and environmental footprint of their video streaming services. Hadi Amirpour, Lingfeng Qu, Jong Hwan Ko, Cosmin Stejerean, Christian Timmerer |
VCIP | 3 |
| 2024 | A DNN partitioning framework with controlled lossy mechanisms for edge-cloud collaborative intelligence
Hyochan Kim, Ji Sub Choi, Jungrae Kim, Jong Hwan Ko |
Future Gener. Comput. Syst. | 4 |
| 2024 | KERNTROL: Kernel Shape Control Toward Ultimate Memory Utilization for In-Memory Convolutional Weight MappingabstractProcessing-in-memory (PIM) architectures have been highlighted as one of the most viable options for faster and more power-efficient computation. Paired with a convolutional weight mapping scheme, PIM arrays can accelerate various deep convolutional neural networks (CNNs) and the applications that adopt them. Recently, shift and duplicate kernel (SDK) convolutional weight mapping scheme was proposed, achieving up to 50% throughput improvement over the prior arts. However, the traditional pattern-based pruning methods, which were adopted for row-skipping and computing cycle reduction, are not optimal for the latest SDK mapping due to the loss of structural regularity caused by the shifted and duplicated kernels. To address this challenge, we propose kernel shape control (KERNTROL), a method where kernel shapes are controlled depending on their mapped columns with the purpose of fostering a structural regularity that is favorable in achieving a high row-skipping ratio and model accuracy. Instead of permanently pruning the weights, KERNTROL with an empty mask (KERNTROL-M) temporarily omits them in the underutilized row using a utilization threshold, thereby preserving important weight elements. However, a significant portion of the memory cells is still underutilized where the threshold is not enforced. To overcome this, we extend KERNTROL-M into KERNTROL with compensatory weights (KERNTROL-C). By populating idle cells with compensatory weights, KERNTROL-C can offset the accuracy drop from weight omission. In comparison to pattern-based pruning approaches, KERNTROL-C achieves simultaneous improvements of up to 36.4% improvement in the compression rate and 5% in model accuracy with up to 100% array utilization. Johnny Rhe, Kang Eun Jeon, Joo Chan Lee, Seongmoon Jeong, Jong Hwan Ko |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | Masked Wavelet Representation for Compact Neural Radiance FieldsabstractNeural radiance fields (NeRF) have demonstrated the potential of coordinate-based neural representation (neural fields or implicit neural representation) in neural rendering. However, using a multi-layer perceptron (MLP) to represent a 3D scene or object requires enormous computational resources and time. There have been recent studies on how to reduce these computational inefficiencies by using additional data structures, such as grids or trees. Despite the promising performance, the explicit data structure necessitates a substantial amount of memory. In this work, we present a method to reduce the size without compromising the advantages of having additional data structures. In detail, we propose using the wavelet transform on grid-based neural fields. Grid-based neural fields are for fast convergence, and the wavelet transform, whose efficiency has been demonstrated in high-performance standard codecs, is to improve the parameter efficiency of grids. Furthermore, in order to achieve a higher sparsity of grid coefficients while maintaining reconstruction quality, we present a novel trainable masking approach. Experimental results demonstrate that non-spatial grid coefficients, such as wavelet coefficients, are capable of attaining a higher level of sparsity than spatial grid coefficients, resulting in a more compact representation. With our proposed mask and compression pipeline, we achieved state-of-the-art performance within a memory budget of 2 MB. Our code is available at https://github.com/daniel03c1/masked_wavelet_nerf. Daniel Rho, Byeonghyeon Lee, Seungtae Nam, Joo Chan Lee, Jong Hwan Ko, Eunbyung Park |
CVPR | 5 |
| 2023 | Regression to Classification: Waveform Encoding for Neural Field-Based Audio Signal RepresentationabstractNeural fields, also known as coordinate-based representations, are an emerging signal representation framework. This approach has also been used to represent audio signals, but the generated audio often contains noise. To reduce noise and improve representation quality, we propose using waveform encoding in the neural field. Instead of yielding real numbers for each temporal coordinate, this involves using discrete integers as outputs, with waveform-encoded integers as target classes, and treating the representation problem as a classification task rather than a regression problem. The experimental results show that waveform encoding can improve the audio quality of neural fields across a variety of audio datasets. TaeSoo Kim, Daniel Rho, Gahui Lee, JaeHan Park, Jong Hwan Ko |
ICASSP | 5 |
| 2023 | Kernel Shape Control for Row-Efficient Convolution on Processing-In-Memory ArraysabstractProcessing-in-memory (PIM) architectures have been highlighted as one of the viable solutions for faster and more power-efficient convolutional neural networks (CNNs) inference. Recently, shift and duplicate kernel (SDK) convolutional weight mapping scheme was proposed, achieving up to 50% through-put improvement over the prior arts. However, the traditional pattern-based pruning methods, which were adopted for row-skipping and computing cycle reduction, are not optimal for the latest SDK mapping due to structural irregularity caused by the shifted and duplicated kernels. To address this issue, we propose a method called kernel shape control (KERNTROL) that aims to promote structural regularity for achieving a high row-skipping ratio and model accuracy. Instead of pruning certain weight elements permanently, KERNTROL controls the kernel shapes through the omission of certain weights based on their mapped columns. In comparison to the latest pattern-based pruning approaches, KERNTROL achieves up to 36.4% improvement in the compression rate, and 38.6% in array utilization with maintaining the original model accuracy. Johnny Rhe, Kang Eun Jeon, Joo Chan Lee, Seongmoon Jeong, Jong Hwan Ko |
ICCAD | 5 |
| 2023 | DCR: Decomposition-Aware Column Re-Mapping for Stuck-At-Fault Tolerance in ReRAM ArraysabstractThe ReRAM-based neuromorphic computing system (NCS) has been widely used as an energy-efficient platform for deep neural network (DNN) acceleration. However, ReRAM commonly suffers from stuck-at-fault (SAF), resulting in permanent device failure. SAF tolerance is an essential task to ensure the reliability of the system by minimizing the DNN inference accuracy degradation. Since hardware-based solutions incur additional overhead and power consumption, it is necessary to seek a solution that can be executed offline to mitigate the impact of SAF. In this work, we propose a decomposition-aware column re-mapping (DCR) for SAF tolerance in analog ReRAM arrays (RAs). Our DCR consists of the column re-mapping technique combined with fault-aware weight decomposition and an advanced sensitivity metric. As a result, it generates a final weight map optimized for the fault map. Our DCR achieves only about 1% loss of inference accuracy on CIFAR-10 and CIFAR-100 for the analog RAs with the SAF rate of 2% and 1%, respectively, without any hardware-based solution or re-training. Hyeonsu Bang, Kang Eun Jeon, Johnny Rhe, Jong Hwan Ko |
ICCD | 4 |
| 2023 | Multi-Scale Bidirectional Recurrent Network with Hybrid Correlation for Point Cloud Based Scene Flow EstimationabstractScene flow estimation provides the fundamental motion perception of a dynamic scene, which is of practical importance in many computer vision applications. In this paper, we propose a novel multi-scale bidirectional recurrent architecture that iteratively optimizes the coarse-tofine scene flow estimation. In each resolution scale of estimation, a novel bidirectional gated recurrent unit is proposed to bidirectionally and iteratively augment point features and produce progressively optimized scene flow. The optimization of each iteration is integrated with the hybrid correlation that captures not only local correlation but also semantic correlation for more accurate estimation. Experimental results indicate that our proposed architecture significantly outperforms the existing state-of-theart approaches on both FlyingThings3D and KITTI benchmarks while maintaining superior time efficiency. Codes and pre-trained models are publicly available at https://github.com/cwc1260/MSBRN. Wencan Cheng, Jong Hwan Ko |
ICCV | 2 |
| 2023 | HandR2N2: Iterative 3D Hand Pose Estimation Using a Residual Recurrent Neural Networkabstract3D hand pose estimation is a critical task in various human-computer interaction applications. Numerous deep learning based estimation models in this domain have been actively explored. However, the existing models follow a non-recurrent scheme and thus require complex architectures or redundant parameters in order to achieve acceptable model capacity. To tackle this limitation, this paper proposes HandR2N2, a compact neural network that iteratively regresses the hand pose using a novel residual recurrent unit. The recurrent design allows recursive exploitation of partial layers to gradually optimize previously estimated joint locations. In addition, we exploit graph reasoning to capture kinematic dependencies between joints for better performance. Experimental results show that the proposed model significantly outperforms the existing methods on three hand pose benchmark datasets in terms of both accuracy and efficiency. Codes and pre-trained models are publicly available at https://github.com/cwc1260/HandR2N2. Wencan Cheng, Jong Hwan Ko |
ICCV | 2 |
| 2023 | Weight-Aware Activation Mapping for Energy-Efficient Convolution on PIM ArraysabstractConvolutional weight mapping plays a stapling role in facilitating convolution operations on Processing-in-memory (PIM) architecture which is, at its essence, a matrix-vector multiplication (MVM) accelerator. Despite its importance, convolutional mapping methods are under-studied and existing mapping methods fail to exploit the sparse and redundant characteristics of heavily quantized convolutional weights, leading to low array utilization and ineffectual computations. To address these issues, this paper proposes a novel weight-aware activation mapping method where activations are mapped onto the memory cells instead of the weights. The proposed method significantly reduces the number of computing cycles by skipping zero-valued weights and merging those PIM array rows with the same weight values. Experimental results on ResNet-18 demonstrate that the proposed weight-aware activation mapping can achieve up to 90% energy saving and latency reduction compared to the conventional approaches. Kang Eun Jeon, Johnny Rhe, Hyeonsu Bang, Jong Hwan Ko |
ISLPED | 4 |
| 2023 | PAIRS: Pruning-AIded Row-Skipping for SDK-Based Convolutional Weight Mapping in Processing-In-Memory ArchitecturesabstractProcessing-in-memory (PIM) architecture is becoming a promising candidate for convolutional neural network (CNN) inference. A recent weight mapping method called shift and duplicate kernel (SDK) improves the utilization by the deployment of shifting the same kernels into idle columns. However, this method inevitably generates idle cells with an irregular distribution, which limits reducing the size of the weight matrix. To effectively compress the weight matrix in the PIM array, prior works have introduced a row-wise pruning scheme, one of the structured weight pruning schemes, that aims to skip the operation on a row by zeroing out all weight in the specific row (we call it row-skipping). However, due to the deployment of shifting kernels, SDK mapping complicates zeroing out all the weight in the same row. To address this issue, we propose pruning-aided row-skipping (PAIRS) that effectively reduces the number of rows of convolutional weights that are mapped with SDK mapping. By pairing the SDK mapping-aware pruning pattern design and row-wise pruning, PAIRS achieves a higher row-skipping ratio. In comparison to pruning methods, PAIRS achieves up to$1.95\times$rows skipped and$4\times$higher compression rate with similar or even better inference accuracy. Johnny Rhe, Kang Eun Jeon, Jong Hwan Ko |
ISLPED | 3 |
| 2023 | FFNeRV: Flow-Guided Frame-Wise Neural Representations for VideosabstractNeural fields, also known as coordinate-based or implicit neural representations, have shown a remarkable capability of representing, generating, and manipulating various forms of signals. For video representations, however, mapping pixel-wise coordinates to RGB colors has shown relatively low compression performance and slow convergence and inference speed. Frame-wise video representation, which maps a temporal coordinate to its entire frame, has recently emerged as an alternative method to represent videos, improving compression rates and encoding speed. While promising, it has still failed to reach the performance of state-of-the-art video compression algorithms. In this work, we propose FFNeRV, a novel method for incorporating flow information into frame-wise representations to exploit the temporal redundancy across the frames in videos inspired by the standard video codecs. Furthermore, we introduce a fully convolutional architecture, enabled by one-dimensional temporal grids, improving the continuity of spatial features. Experimental results show that FFNeRV yields the best performance for video compression and frame interpolation among the methods using frame-wise representations or neural fields. To reduce the model size even further, we devise a more compact convolutional architecture using the group and pointwise convolutions. With model compression techniques, including quantization-aware training and entropy coding, FFNeRV outperforms widely-used standard video codecs (H.264 and HEVC) and performs on par with state-of-the-art video compression algorithms. Joo Chan Lee, Daniel Rho, Jong Hwan Ko, Eunbyung Park |
ACM Multimedia | 3 |
| 2023 | Mip-Grid: Anti-aliased Grid Representations for Neural Radiance FieldsabstractDespite the remarkable achievements of neural radiance fields (NeRF) in representing 3D scenes and generating novel view images, the aliasing issue, rendering 'jaggies' or 'blurry' images at varying camera distances, remains unresolved in most existing approaches. The recently proposed mip-NeRF has effectively addressed this challenge by introducing integrated positional encodings (IPE). However, it relies on MLP architecture to represent the radiance fields, missing out on the fast training speed offered by the latest grid-based methods. In this work, we present mip-Grid, a novel approach that integrates anti-aliasing techniques into grid-based representations for radiance fields, mitigating the aliasing artifacts while enjoying fast training time. Notably, the proposed method uses a single-scale shared grid representation and a single-sampling approach, which only introduces minimal additions to the model parameters and computational costs. To handle scale ambiguity, mip-Grid generates multiple grids by applying simple convolution operations over the shared grid and uses the scale-aware coordinate to retrieve the appropriate features from the generated multiple grids. To test the effectiveness, we incorporated the proposed approach into the two recent representative grid-based methods, TensoRF and K-Planes. The experimental results demonstrated that mip-Grid greatly improved the rendering performance of both methods and showed comparable performance to mip-NeRF on multi-scale datasets while achieving significantly faster training time. Seungtae Nam, Daniel Rho, Jong Hwan Ko, Eunbyung Park |
NeurIPS | 3 |
| 2023 | An overhead-free region-based JPEG framework for task-driven image compression
Seonghye Jeong, Seongmoon Jeong, Simon S. Woo, Jong Hwan Ko |
Pattern Recognit. Lett. | 4 |
| 2023 | Scalable Color Quantization for Task-centric Image CompressionabstractConventional image compression techniques targeted for the perceptual quality are not generally optimized for classification tasks using deep neural networks (DNNs). To compress images for DNN inference tasks, recent studies have proposed task-centric image compression methods with quantization techniques optimized for DNN inference. Among them, color quantization was proposed to reduce the amount of data per pixel by limiting the number of distinct colors (color space) in an image. However, quantizing images into various color space sizes requires training and inference of multiple DNNs, each of which is dedicated to each color space. To overcome this limitation, we propose a scalable color quantization method, where images with variable color space sizes can be extracted from a master image generated by a single DNN model. This scalability is enabled by weighted color grouping that constructs a color palette using critical color components for the classification task. We also propose an adaptive training method that can jointly optimize images with various color-space sizes. The results show that the proposed method supports dynamic changes of the color space size between 1–6 bit color space per pixel, while even increasing the inference accuracy at a low bit precision up to 20.2% and 46.6% compared to other task- and human-centric color quantizations, respectively. Jaehyun Park 0012, Joo Chan Lee, Jong Hwan Ko |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Neural Residual Flow Fields for Efficient Video Representations
Daniel Rho, Junwoo Cho, Jong Hwan Ko, Eunbyung Park |
ACCV (2) | 3 |
| 2022 | VW-SDK: Efficient Convolutional Weight Mapping Using Variable Windows for Processing-In-Memory ArchitecturesabstractWith their high energy efficiency, processing-in-memory (PIM) arrays are increasingly used for convolutional neural network (CNN) inference. In PIM-based CNN inference, the computational latency and energy are dependent on how the CNN weights are mapped to the PIM array. A recent study proposed shifted and duplicated kernel (SDK) mapping that reuses the input feature maps with a unit of a parallel window, which is convolved with duplicated kernels to obtain multiple output elements in parallel. However, the existing SDK-based mapping algorithm does not always result in the minimum computing cycles because it only maps a square-shaped parallel window with the entire channels. In this paper, we introduce a novel mapping algorithm called variable-window SDK (VW-SDK), which adaptively determines the shape of the parallel window that leads to the minimum computing cycles for a given convolutional layer and PIM array. By allowing rectangular-shaped windows with partial channels, VW-SDK utilizes the PIM array more efficiently, thereby further reduces the number of computing cycles. The simulation with a$512\times 512$PIM array and Resnet-18 shows that VW-SDK improves the inference speed by$1.69\times$compared to the existing SDK-based algorithm. Johnny Rhe, Sungmin Moon, Jong Hwan Ko |
DATE | 3 |
| 2022 | Bi-PointFlowNet: Bidirectional Learning for Point Cloud Based Scene Flow Estimation
Wencan Cheng, Jong Hwan Ko |
ECCV (28) | 2 |
| 2022 | Streamable Neural Fields
Junwoo Cho, Seungtae Nam, Daniel Rho, Jong Hwan Ko, Eunbyung Park |
ECCV (20) | 4 |
| 2022 | ADA-VAD: Unpaired Adversarial Domain Adaptation for Noise-Robust Voice Activity DetectionabstractVoice Activity Detection (VAD) is becoming an essential front-end component in various speech processing systems. As those systems are commonly deployed in environments with diverse noise types and low signal-to-noise ratios (SNRs), an effective VAD method should perform robust detection of speech region out of noisy background signals. In this paper, we propose adversarial domain adaptive VAD (ADA-VAD), which is a deep neural network (DNN) based VAD method highly robust to audio samples with various noise types and low SNRs. The proposed method trains DNN models for a VAD task in a supervised manner. Simultaneously, to mitigate the performance degradation due to back-ground noises, the adversarial domain adaptation method is adopted to match the domain discrepancy between noisy and clean audio stream in an unsupervised manner. The results show that ADA-VAD achieves an average of 3.6%p and 7%p higher AUC than models trained with manually extracted features on the AVA-speech dataset and a speech database synthesized with an unseen noise database, respectively. Taesoo Kim, Jiho Chang, Jong Hwan Ko |
ICASSP | 3 |
| 2022 | NAS-VAD: Neural Architecture Search for Voice Activity DetectionabstractVarious neural network-based approaches have been proposed for more robust and accurate voice activity detection (VAD). Manual design of such neural architectures is an error-prone and time-consuming process, which prompted the development of neural architecture search (NAS) that automatically design and optimize network architectures. While NAS has been successfully applied to improve performance in a variety of tasks, it has not yet been exploited in the VAD domain. In this paper, we present the first work that utilizes NAS approaches on the VAD task. To effectively search architectures for the VAD task, we propose a modified macro structure and a new search space with a much broader range of operations that includes attention operations. The results show that the network structures found by the propose NAS framework outperform previous manually designed state-of-the-art VAD models in various noise-added and real-world-recorded datasets. We also show that the architectures searched on a particular dataset achieve improved generalization performance on unseen audio datasets. Our code and models are available at https://github.com/daniel03c1/NAS_VAD. Daniel Rho, Jinhyeok Park, Jong Hwan Ko |
INTERSPEECH | 3 |
| 2022 | A Reconfigurable Neural Architecture for Edge-Cloud Collaborative Real-Time Object DetectionabstractAlthough recent advances in deep neural networks (DNNs) have enabled remarkable performance on various computer vision tasks, it is challenging for edge devices to perform real-time inference of complex DNN models due to their stringent resource constraints. To enhance the inference throughput, recent studies have proposed collaborative intelligence (CI), which splits DNN computation into edge and cloud platforms, mostly for simple tasks, such as image classification. However, for general DNN-based object detectors with a branching architecture, CI is highly restricted because of a significant feature transmission overhead. To resolve this issue, in this study, we propose a reconfigurable DNN architecture for real-time object detection that can configure the optimal split point according to the edge–cloud CI environment. The proposed architecture allows the DNN model to be splittable by a feature reconstruction network and asymmetric scaling. Based on the splittable architecture, we integrate independent splittable models for each split point into a single-weight reconfigurable model that enables multipath inference by switchable quantization and distribution matching. Finally, we introduce an adaptive application procedure of the reconfigurable model for efficient CI, which includes asymmetric scale configuration and split point selection. The performance evaluation using YOLOv5 as the baseline showed that the proposed architecture achieved 30 frames/s ($2.6\times $and$1.6\times $higher than edge-only and cloud-only inference, respectively), on the NVIDIA Jetson TX2 platform in a WiFi environment. Joo Chan Lee, Sungtae Moon, Jong Hwan Ko |
IEEE Internet Things J. | 4 |
| 2022 | Adaptive weight-bit inversion for state error reduction for robust and efficient deep neural network inference using MLC NAND FlashabstractWhen Flash memory is used to store the weights of a deep neural network (DNN), the inference accuracy can degrade owing to the state errors of the Flash memory. To protect the weights from state errors, the existing methods rely on an error correction code (ECC) or parity, which can incur power/storage overhead. We propose a weight-bit inversion method that minimizes accuracy loss caused by state errors without using ECC or parity. First, the method applies weight-bit inversion for state elimination (WISE), which removes the most error-prone state from MLC NAND, thereby improving the error robustness and the most significant bit (MSB) page read speed. If the initial accuracy loss caused by the WISE is unacceptable, we apply weight-bit inversion for state error reduction (WISER), which reduces weight mapping to error-prone states with minimum changes in weight value. To further improve the read speed with minimum accuracy loss, we propose an adaptive weight-bit inversion scheme that selectively applies WISE or WISER to the unit of a weight group. The simulation results imply that after 16K program-erase cycles in NAND Flash, WISER reduces the CIFAR-100 accuracy loss by 1.33X for LeNet-5, 2.92X for VGG-16, and 2.74X for Resnet-20 compared with the existing methods. In addition, the adaptive inversion technique improves the read speed by 48.6% without accuracy loss, compared with the WISER-only scheme. Jaehun Jang, Jong Hwan Ko |
J. Syst. Archit. | 2 |
| 2022 | ORVAE: One-Class Residual Variational Autoencoder for Voice Activity Detection in Noisy Environment
Hasam Khalid, Shahroz Tariq, TaeSoo Kim, Jong Hwan Ko, Simon S. Woo |
Neural Process. Lett. | 4 |
| 2021 | A Splittable DNN-Based Object Detector for Edge-Cloud Collaborative Real-Time Video InferenceabstractWhile recent advances in deep neural networks (DNNs) enabled remarkable performance on various computer vision tasks, it is challenging for edge devices to perform real-time inference of complex DNN models due to their stringent resource constraint. To enhance the inference throughput, recent studies proposed collaborative intelligence (CI) that splits DNN computation into edge and cloud platforms, mostly for simple tasks such as image classification. However, for general DNN-based object detectors with a branching architecture, CI is highly restricted because of a significant feature transmission overhead. To solve this issue, this paper proposes a splittable object detector that enables edge-cloud collaborative real-time video inference. The proposed architecture includes a feature reconstruction network that can generate multiple features required for detection using a small-sized feature from the edge-side extractor. Asymmetric scaling on the feature extractor and reconstructor further reduces the transmitted feature size and edge inference latency, while maintaining detection accuracy. The performance evaluation using Yolov5 shows that the proposed model achieves 28 fps (2.45X and 1.56X higher than edge-only and cloud-only inference, respectively), on the NVIDIA Jetson TX2 platform in WiFi environment. Joo Chan Lee, Sungtae Moon, Jong Hwan Ko |
AVSS | 4 |
| 2021 | WISER: Deep Neural Network Weight-bit Inversion for State Error Reduction in MLC NAND FlashabstractWhen Flash memory is used to store the deep neural network (DNN) weights, inference accuracy can degrade due to the Flash memory state errors. To protect the weights from the state errors, the existing methods relied on ECC(Error Correction Code) or parity, which can incur power/storage overhead. In this study, we propose a weight bit inversion method that minimizes the accuracy loss due to the Flash memory state errors without using the ECC or parity. The method first applies WISE(Weight-bit Inversion for State Elimination) that removes the most error-prone state from MLC NAND, thereby improving both the error robustness and the MSB page read speed. If the initial accuracy loss due to weight inversion of WISE is unacceptable, we apply WISER(Weight-bit Inversion for State Error Reduction) that reduces weight mapping to the error-prone state with minimum weight value changes. The simulation results show that after 16K program-erase cycles in NAND Flash, WISER reduces CIFAR-100 accuracy loss by 2.92X for VGG-16 compared to the existing methods. Jaehun Jang, Jong Hwan Ko |
DATE | 2 |
| 2021 | HandFoldingNet: A 3D Hand Pose Estimation Network Using Multiscale-Feature Guided Folding of a 2D Hand SkeletonabstractWith increasing applications of 3D hand pose estimation in various human-computer interaction applications, convolution neural networks (CNNs) based estimation models have been actively explored. However, the existing models require complex architectures or redundant computational resources to trade with the acceptable accuracy. To tackle this limitation, this paper proposes HandFoldingNet, an accurate and efficient hand pose estimator that regresses the hand joint locations from the normalized 3D hand point cloud input. The proposed model utilizes a folding-based decoder that folds a given 2D hand skeleton into the corresponding joint coordinates. For higher estimation accuracy, folding is guided by multi-scale features, which include both global and joint-wise local features. Experimental results show that the proposed model outperforms the existing methods on three hand pose benchmark datasets with the lowest model parameter requirement. Code is available at https://github.com/cwc1260/HandFold. Wencan Cheng, Jaehyun Park 0012, Jong Hwan Ko |
ICCV | 3 |
| 2021 | A Charge-Domain Scalable-Weight In-Memory Computing Macro With Dual-SRAM Architecture for Precision-Scalable DNN AcceleratorsabstractThis paper presents a charge-domain in-memory computing (IMC) macro for precision-scalable deep neural network accelerators. The proposed Dual-SRAM cell structure with coupling capacitors enables charge-domain multiply and accumulate (MAC) operation with variable-precision signed weights. Unlike prior charge-domain IMC macros that only support binary neural networks or digitally compute weighted sums for MAC operation with multi-bit weights, the proposed macro implements analog weighted sums for energy-efficient bit-scalable MAC operations with a novel series-coupled merging scheme. A test chip with a 16-kb SRAM macro is fabricated in 28-nm FDSOI process, and the measured macro throughput is 125.2-876.5 GOPS for weight bit-precision varying from 2 to 8. The macro also achieves energy efficiency ranging from 18.4 TOPS/W for 8-b weight to 119.2 TOPS/W for 2-b weight. Eunyoung Lee, Taeyoung Han, Gicheol Shin, Jaerok Kim, Soyoun Jeong, Johnny Rhe, Jaehyun Park 0012, Jong Hwan Ko, Yoonmyung Lee |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2019 | A Camera with Brain - Embedding Machine Learning in 3D SensorsabstractThe cameras today are designed to capture signals with highest possible accuracy to most faithfully represent what it sees. However, many mission-critical autonomous applications ranging from traffic monitoring to disaster recovery to defense requires quality of information, where useful information depends on the tasks and is defined using complex features, rather than only changes in captured signal. Such applications require cameras that capture useful information from a scene with highest quality while meeting system constraints such as power, performance, and bandwidth. This paper will discuss the feasibility of a camera that learns how to capture task-dependent information with highest quality, paving the pathway to design a camera with brain. 3D integration of digital pixel sensors with massively parallel computing platform for machine learning creates a hardware architecture for such a camera. The paper will discuss embedded machine learning algorithms that can run on such platform to enhance quality of useful information by real-time control of the sensor parameters. We conclude by identifying critical challenges as well as opportunities for hardware and algorithmic innovations to enable machine learning in the feedback loop of a 3D image sensor based camera. Burhan Ahmad Mudassar, Priyabrata Saha, Mohammad Faisal Amir, Evan Gebhardt, Taesik Na, Jong Hwan Ko, Marilyn Wolf, Saibal Mukhopadhyay |
DATE | 7 |
| 2019 | Mixture of Pre-processing Experts Model for Noise Robust Deep Learning on Resource Constrained PlatformsabstractDeep learning on an edge device requires energy efficient operation due to ever diminishing power budget. Intentional low quality data during the data acquisition for longer battery life, and natural noise from the low cost sensor degrade the quality of target output which hinders adoption of deep learning on an edge device. To overcome these problems, we propose simple yet efficient mixture of pre-processing experts (MoPE) model to handle various image distortions including low resolution and noisy images. We also propose to use adversarially trained auto encoder as a pre-processing expert for the noisy images. We evaluate our proposed method for various machine learning tasks including object detection on MS-COCO 2014 dataset, multiple object tracking problem on MOT-Challenge dataset, and human activity classification on UCF 101 dataset. Experimental results show that the proposed method achieves better detection, tracking and activity classification accuracies under noise without sacrificing accuracies for the clean images. The overheads of our proposed MoPE are 0.67% and 0.17% in terms of memory and computation compared to the baseline object detection network. Taesik Na, Minah Lee, Burhan Ahmad Mudassar, Priyabrata Saha, Jong Hwan Ko, Saibal Mukhopadhyay |
IJCNN | 5 |
| 2019 | Energy Efficient and Side-Channel Secure Cryptographic Hardware for IoT-Edge NodesabstractDesign of ultralightweight but secure encryption engine is a key challenge for Internet-of-Things edge devices. This paper explores the system level design space for an ultralow power image sensor node for secure communication and proposes an optimized datapath architecture for 128-bit SIMON (SIMON128), a lightweight block cipher, for minimal performance, power, and area overheads with increased level of side-channel security. Various datapath architectures for SIMON are explored for simultaneously increasing energy-efficiency and resistance to power-based side-channel analysis (PSCA) attacks. Alternative datapath architectures are implemented on ASIC (15 nm CMOS) and field programmable gate array (FPGA) (Spartan-6, 45 nm) to perform power, performance, and area analysis. We show that, although a bitserial datapath minimizes area and power, a round unrolled datapath provides 80× higher energy-efficiency and 143× higher performance, compared to the baseline bitserial design. Moreover, the PSCA measurements performed using Sakura-G board with Spartan-6 FPGA, demonstrate that a 6-round unrolled datapath improves minimum-traces-to-disclosure for correlation power analysis (CPA) by at least 384× over baseline bitserial design with no successful CPA even with 500000 measurements. Finally, application to the image-sensor node demonstrates that optimized unrolled SIMON128 can provide equivalent performance to AES128 at lower area, higher energy efficiency, and improved side channel security. Nikhil Chawla, Jong Hwan Ko, Monodeep Kar, Saibal Mukhopadhyay |
IEEE Internet Things J. | 3 |
| 2019 | Design and Analysis of a Neural Network Inference Engine Based on Adaptive Weight CompressionabstractNeural networks generally require significant memory capacity/bandwidth to store/access a large number of synaptic weights. This paper presents design of an energy-efficient neural network inference engine based on adaptive weight compression using a JPEG image encoding algorithm. To maximize compression ratio with minimum accuracy loss, the quality factor of the JPEG encoder is adaptively controlled depending on the accuracy impact of each block. With 1% accuracy loss, the proposed approach achieves 63.4× compression for multilayer perceptron (MLP) and 31.3× for LeNet-5 with the MNIST dataset, and 15.3× for AlexNet and 10.2× for ResNet-50 with ImageNet. The reduced memory requirement leads to higher throughput and lower energy for neural network inference (3× effective memory bandwidth and 22× lower system energy for MLP). Jong Hwan Ko, Duckhwan Kim 0001, Taesik Na, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Edge-Host Partitioning of Deep Neural Networks with Feature Space Encoding for Resource-Constrained Internet-of-Things PlatformsabstractThis paper introduces partitioning an inference task of a deep neural network between an edge and a host platform in the IoT environment. We present a DNN as an encoding pipeline, and propose to transmit the output feature space of an intermediate layer to the host. Encoding of the feature space is proposed to enhance the maximum input rate supported by the edge platform and/or reduce the energy of the edge platform. Simulation results show that partitioning a DNN coupled with feature space encoding enables significant improvement in the energy-efficiency and throughput over the baseline configurations that perform the entire inference at the edge or at the host. Jong Hwan Ko, Taesik Na, Mohammad Faisal Amir, Saibal Mukhopadhyay |
AVSS | 1 |
| 2018 | Edge-cloud collaborative processing for intelligent internet of things: a case study on smart surveillanceabstractLimited processing power and memory prevent realization of state of the art algorithms on the edge level. Offloading computations to the cloud comes with tradeoffs as compression techniques employed to conserve transmission bandwidth and energy adversely impact accuracy of the algorithm. In this paper, we propose collaborative processing to actively guide the output of the sensor to improve performance on the end application. We apply this methodology to smart surveillance specifically the task of object detection from video. Perceptual quality and object detection performance is characterized and improved under a variety of channel conditions. Burhan Ahmad Mudassar, Jong Hwan Ko, Saibal Mukhopadhyay |
DAC | 2 |
| 2018 | The CAMEL approach to stacked sensor smart camerasabstractStacked image sensor systems combine an image sensor, memory, and processors using 3D technology. Stacking camera components that have traditionally been packaged separately provides several benefits: very high bandwidth out of the image sensor, allowing for higher frame rates; very low latency, providing opportunities for image processing and computer vision algorithms which can adapt at very high rates; and lower power consumption. This paper will review the characteristics of stacked image sensor systems and discuss novel algorithmic and systems concepts that are made possible by these stacked sensors. Saibal Mukhopadhyay, Marilyn Wolf, Mohammed Faisal Amir, Evan Gebhardt, Jong Hwan Ko, Jaeha Kung 0001, Burhan Ahmad Mudassar |
DATE | 5 |
| 2018 | Limiting Numerical Precision of Neural Networks to Achieve Real-Time Voice Activity DetectionabstractFast and robust voice-activity detection is critical to efficiently process speech. While deep-learning based methods to detect voice have shown competitive accuracies, the best models in the literature incur over a 100 ms latency on commodity processors. Such delays are unacceptable for real-time speech processing. In this paper, we study the impact of lowering the representation precision of the neural-network weights and neurons on both the accuracy and delay of voice-activity detection. Based on a design-space exploration, we not only determine the optimal scaling strategy but also adjust the network structure to accommodate the new quantization levels. Through experiments conducted with real user data, we demonstrate that optimized deep neural networks with lower bit precisions outperform the state-of-the-art WebRTC voice-activity detector with 87x lower delay and 6.8% lower error rate. Jong Hwan Ko, Josh Fromm, Matthai Philipose, Ivan Tashev, Shuayb Zarar |
ICASSP | 1 |
| 2018 | An Unsupervised Anomalous Event Detection Framework with Class Aware Source SeparationabstractThis paper presents a novel problem of detection and localization of anomalous events due to a certain class of objects in video data with applications to smart surveillance. A baseline system is proposed that uses a convolutional neural network (CNN) to generate pixel level masks corresponding to objects of a class of interest. A Restricted Boltzmann Machine (RBM) is then trained on the mask to learn patterns of normal behavior. The free energy of the RBM is used to detect the presence of an anomaly while the reconstruction error is used to localize the anomaly. Our approach is scalable to a low power and energy constrained setting with 1930.48 ms of latency and 4826 mJ energy consumed per frame on a mGPU. Burhan Ahmad Mudassar, Jong Hwan Ko, Saibal Mukhopadhyay |
ICASSP | 2 |
| 2018 | Cascade Adversarial Machine Learning Regularized with a Unified Embedding
Taesik Na, Jong Hwan Ko, Saibal Mukhopadhyay |
ICLR (Poster) | 2 |
| 2017 | Design of an Energy-Efficient Accelerator for Training of Convolutional Neural Networks using Frequency-Domain ComputationabstractConvolutional neural networks (CNNs) require high computation and memory demand for training. This paper presents the design of a frequency-domain accelerator for energy-efficient CNN training. With Fourier representations of parameters, we replace convolutions with simpler pointwise multiplications. To eliminate the Fourier transforms at every layer, we train the network entirely in the frequency domain using approximate frequency-domain nonlinear operations. We further reduce computation and memory requirements using sinc interpolation and Hermitian symmetry. The accelerator is designed and synthesized in 28nm CMOS, as well as prototyped in an FPGA. The simulation results show that the proposed accelerator significantly reduces training time and energy for a target recognition accuracy. Jong Hwan Ko, Burhan Ahmad Mudassar, Taesik Na, Saibal Mukhopadhyay |
DAC | 1 |
| 2017 | Adaptive weight compression for memory-efficient neural networksabstractNeural networks generally require significant memory capacity/bandwidth to store/access a large number of synaptic weights. This paper presents an application of JPEG image encoding to compress the weights by exploiting the spatial locality and smoothness of the weight matrix. To minimize the loss of accuracy due to JPEG encoding, we propose to adaptively control the quantization factor of the JPEG algorithm depending on the error-sensitivity (gradient) of each weight. With the adaptive compression technique, the weight blocks with higher sensitivity are compressed less for higher accuracy. The adaptive compression reduces memory requirement, which in turn results in higher performance and lower energy of neural network hardware. The simulation for inference hardware for multilayer perceptron with the MNIST dataset shows up to 42X compression with less than 1% loss of recognition accuracy, resulting in 3X higher effective memory bandwidth and ~19X lower system energy. Jong Hwan Ko, Duckhwan Kim 0001, Taesik Na, Jaeha Kung 0001, Saibal Mukhopadhyay |
DATE | 1 |
| 2017 | Clock data compensation aware clock tree synthesis in digital circuits with adaptive clock generationabstractAdaptive clock generation to track critical path delay enables lowering supply voltage with improved timing slack under supply noise. This paper presents how to synthesize clock tree in adaptive clocking to fully exploit the clock data compensation (CDC) effect in digital circuits. The paper first provides analytical proof of ideal CDC effect for ring oscillator based clock generation. Second, the paper analyzes non-ideal CDC effect in a gate dominated critical path and wire dominated clock tree design. The paper shows the delay sensitivity mismatch between clock tree and critical path can degrade CDC effect by analyzing timing slack under power supply noise (PSN). Finally, the paper proposes simple but efficient clock tree synthesis (CTS) technique to maximize timing slack under PSN in digital circuits with adaptive clock generation. Taesik Na, Jong Hwan Ko, Saibal Mukhopadhyay |
DATE | 2 |
| 2017 | On-chip training of recurrent neural networks with limited numerical precisionabstractTraining of neural network can be accelerated by limited numerical precision together with specialized low-precision hardware. This paper studies how low precision can impact on entire training of RNNs. We emulate low precision training for recently proposed gated recurrent unit (GRU) and use dynamic fixed point as a target numeric format. We first show that batch normalization on input sequences can help speed up training with low precision as well as high precision. We also show that the overflow rate should be carefully controlled for dynamic fixed point. We study low precision training with various rounding options including bit truncation, round to nearest, and stochastic rounding. Stochastic rounding shows superior results than the other options. The effect of fully low precision training is also analyzed by comparing partial low precision training. We show that the piecewise linear activation function with stochastic rounding can achieve comparable training results with floating point precision. Low precision multiplier and accumulator (MAC) with linear-feedback shift register (LFSR) is implemented with 28nm Synopsys PDK for energy and performance analysis. Implementation results show low precision hardware is 4.7× faster, and energy per task is up to 4.55× lower than that of floating point hardware. Taesik Na, Jong Hwan Ko, Jaeha Kung 0001, Saibal Mukhopadhyay |
IJCNN | 2 |
| 2016 | An energy-efficient wireless video sensor node with a region-of-interest based multi-parameter rate controller for moving object surveillanceabstractThis paper presents a lightweight video sensor node for moving object surveillance using region-of-interest (ROI) based coding and an on-line multi-parameter rate controller. The proposed ROI-based coding scheme determines ROI blocks, pre-processes non-ROI blocks using bit-truncation, and encodes all blocks using Motion JPEG. The on-line rate controller modulates the parameters of the ROI-based coding scheme to match the encoded data rate and transmission data rate under the variations in channel bandwidth and input video content. The low-complexity hardware of the ROI-based coding scheme reduces computation energy, and the on-line rate controller minimizes buffer requirement. The sensor node is designed in 130nm CMOS and prototyped in a Virtex-V FPGA. Simulations show that, under the same ROI quality, the proposed approach reduces system energy by 61% compared to H.264/AVC. Jong Hwan Ko, Taesik Na, Saibal Mukhopadhyay |
AVSS | 1 |
| 2016 | An Energy-Aware Approach to Noise-Robust Moving Object Detection for Low-Power Wireless Image Sensor PlatformsabstractThis paper presents an energy-aware approach to moving object detection that requires very low computation and memory while ensuring robust performance under noisy environments. The proposed approach is integrated into a wireless image sensor platform with a block-based processing unit and the motion JPEG encoder. The sensor platform is designed as an ASIC in 130nm CMOS for energy/area analysis, as well as prototyped into Virtex-5 FPGA for functional validation. The sensor platform designed with the proposed approach consumes less energy and area than the platforms with the existing methods such as Gaussian Mixture Model, while maintaining a reliable delivery of region-of-interest. Jong Hwan Ko, Saibal Mukhopadhyay |
ISLPED | 1 |
| 2015 | Exploring power attack protection of resource constrained encryption engines using integrated low-drop-out regulatorsabstractThe power attack protection of encryption engines often comes at the expense of area, power, and/or performance overheads making the design of a low-power and compact but secure encryption engine challenging. This paper explores the feasibility of using an on-chip low dropout regulator (LDO) as a countermeasure to power attack of low-power and compact encryption engine. We design an area minimized implementation of Advanced Encryption Standard (AES) using predictive 45nm node and show that lightweight implementations are more susceptible to power attack. Using behavioral modeling, we show that an on-chip LDO can enhance power attack resistance of this compact AES engine; however, the tradeoff between LDO performance and power attack protection is essential. Our analysis shows that LDO can increase power attack resistance of the compact AES by >800X with marginal area (1.4%) and power (5%) overheads. Monodeep Kar, Jong Hwan Ko, Saibal Mukhopadhyay |
ISLPED | 3 |