Eunhyeok Park

dblp:161/0829 · DBLP profile ↗
← Back
38ranked-venue papers
5as first author
29since 2021 · last 2026
0000-0002-7331-9819ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 3 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 3 first-author · 16 since 2021Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Fast and Energy-Efficient Support for Low-Precision LLMs on PIM
abstract
Processing-in-Memory (PIM) has gained momentum as a promising approach for mitigating memory bottlenecks, and it is particularly well suited to autoregressive decoding in large language models (LLMs), where memory-bound General Matrix-Vector Multiplication (GEMV) operations account for a large portion of the workload. Due to the large size of LLMs, it is challenging to support them in PIM environments with limited memory capacity. Applying group-wise weight-only quantization (GWQ), widely used in LLMs, can effectively reduce model size while minimizing accuracy degradation. However, the weights in GWQ-applied LLMs are typically dequantized using scales and zero-points before GEMV is performed, which can introduce non-trivial latency overhead. In this paper, we propose a method for DRAM-PIM to efficiently support GEMV operations in symmetric and asymmetric GWQ-based LLMs with INT2 and INT4 precision. Based on the Newton scheme, the proposed method incurs an area overhead of 20.6% compared to the PIM units used in a 16-bank, single-channel system. When asymmetric GWQ with a group size of 128 is applied, it achieves approximately a 4 reduction in storage at INT4, along with a 1.16× speedup and 1.41× energy efficiency compared to FP16 GEMV. At INT2, it achieves an 8× reduction in storage, along with a 1.27× speedup and 1.57× energy efficiency.
Byeori Kim, Eunhyeok Park
DATE3
2026 Stabilizing Direct Training of Spiking Neural Networks: Membrane Potential Initialization and Threshold-robust Surrogate Gradient
abstract
Recent advancements in the direct training of Spiking Neural Networks (SNNs) have demonstrated high-quality outputs even at early timesteps, paving the way for novel energy-efficient AI paradigms. However, the inherent non-linearity and temporal dependencies in SNNs introduce persistent challenges, such as temporal covariate shift (TCS) and unstable gradient flow with learnable neuron thresholds. In this paper, we present two key innovations: MP-Init (Membrane Potential Initialization) and TrSG (Threshold-robust Surrogate Gradient). MP-Init addresses TCS by aligning the initial membrane potential with its stationary distribution, while TrSG stabilizes gradient flow with respect to threshold voltage during training. Extensive experiments validate our approach, achieving state-of-the-art accuracy on both static and dynamic image datasets. The code is available at: https://github.com/kookhh0827/SNN-MP-Init-TRSG
Hyunho Kook, Byeongho Yu, Jeong Min Oh, Eunhyeok Park
WACV4
2026 DreamCatcher: Efficient Multi-Concept Customization via Representation Finetuning
abstract
Recent advances in customizing Text-to-Image models allow users to generate personalized images with just a few samples. As demand for multi-concept generation grows, methods using weight fusion and test-time optimization have emerged, integrating multiple concepts within a single image. However, these approaches inject concept knowledge into the parametric space, leading to high overhead in multi-concept generation. We introduce DreamCatcher, an efficient framework based on representation finetuning. Our key innovation embeds conceptual information into the feature space, achieving up to 5× faster multi-concept generation while reducing learnable storage per concept by 88%, all without quality loss. Besides, our method is highly versatile, enabling personalized inpainting without additional training.
Changhun Lee, Eunhyeok Park
WACV3
2025 SEAL: Scaling to Emphasize Attention for Long-Context Retrieval
abstract
While many advanced LLMs are designed to handle long sequence data, we can still observe notable quality degradation even within the sequence limit.In this work, we introduce a novel approach called Scaling to Emphasize Attention for Long-context retrieval (SEAL), which enhances the retrieval performance of large language models (LLMs) over long contexts.We observe that specific attention heads are closely tied to long-context retrieval, showing positive or negative correlation with retrieval scores, and adjusting the strength of these heads boosts the quality of LLMs in long context by a large margin.Built on this insight, we propose a learning-based mechanism that leverages generated data to emphasize these heads.By applying SEAL, we achieve significant improvements in long-context retrieval performance across various tasks and models.Additionally, when combined with existing training-free context extension techniques, SEAL extends the contextual limits of LLMs while maintaining highly reliable outputs.
Changhun Lee, Minsang Seok, Jungyu Jin, Younghyun Cho, Eunhyeok Park
ACL (1)5
2025 HOT: Hadamard-based Optimized Training
abstract
It has become increasingly important to optimize backpropagation to reduce memory usage and computational overhead. Achieving this goal is highly challenging, as multiple objectives must be considered jointly while maintaining training quality. In this paper, we focus on matrix multiplication, which accounts for the largest portion of training costs, and analyze its backpropagation in detail to identify lightweight techniques that offer the best benefits. Based on this analysis, we introduce a novel method, Hadamard-based Optimized Training (HOT). In this approach, we apply Hadamard-based optimizations, such as Hadamard quantization and Hadamard low-rank approximation, selectively and with awareness of the suitability of each optimization for different backward paths. Additionally, we introduce two enhancements: activation buffer compression and layer-wise quantizer selection. Our extensive analysis shows that HOT achieves up to 75% memory savings and a 2.6× acceleration on real GPUs, with negligible accuracy loss compared to FP32 precision. Our code is available at https://github.com/sungonuni/HOT.
Seonggon Kim, Juncheol Shin, Seung-taek Woo, Eunhyeok Park
CVPR4
2025 PCM : Picard Consistency Model for Fast Parallel Sampling of Diffusion Models
abstract
Recently, diffusion models have achieved significant advances in vision, text, and robotics. However, they still face slow generation speeds due to sequential denoising processes. To address this, a parallel sampling method based on Picard iteration was introduced, effectively reducing sequential steps while ensuring exact convergence to the original output. Nonetheless, Picard iteration does not guarantee faster convergence, which can still result in slow generation in practice. In this work, we propose a new parallelization scheme, the Picard Consistency Model (PCM), which significantly reduces the number of generation steps in Picard iteration. Inspired by the consistency model, PCM is directly trained to predict the fixed-point solution, or the final output, at any stage of the convergence trajectory. Additionally, we introduce a new concept called model switching, which addresses PCM’s limitations and ensures exact convergence. Extensive experiments demonstrate that PCM achieves up to a 2.71x speedup over sequential sampling and a 1.77x speedup over Picard iteration across various tasks, including image generation and robotic control.
Junhyuk So, Jiwoong Shin, Chaeyeon Jang, Eunhyeok Park
CVPR4
2025 AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models
abstract
To enable broader deployment of Large Language Models (LLMs), it is essential to identify the best-performing model under strict memory constraints.We present AMQ, Automated Mixed-Precision Weight-Only Quantization, a framework that assigns layer-wise quantization bit-widths to optimally balance model quality and memory usage.However, the combinatorial search space, with over 10 100 possible configurations, makes conventional black-box optimization infeasible.AMQ overcomes this challenge through four key innovations: (1) search space pruning using prior knowledge to exclude unpromising configurations, (2) quantization proxy to bypass costly format conversions during search, (3) quality predictor to minimize evaluation overhead, and (4) iterative search-and-update strategy for fast and stable convergence.By integrating these components, AMQ efficiently explores the qualityefficiency landscape, reaching the Pareto frontier and yielding LLMs that are both compact and high-performing.Our code is available at https://github.com/dlwns147/amq.
Sangjun Lee 0001, Seung-taek Woo, Jungyu Jin, Changhun Lee, Eunhyeok Park
EMNLP5
2025 PruneCD: Contrasting Pruned Self Model to Improve Decoding Factuality
abstract
To mitigate the hallucination problem in large language models, DoLa exploits early exit logits from the same model as a contrastive prior.However, we found that these early exit logits tend to be flat, low in magnitude, and fail to reflect meaningful contrasts.To address this, we propose PruneCD, a novel contrastive decoding method that constructs the amateur model via layer pruning rather than early exit.This design leads to more informative and well-aligned logits, enabling more effective contrastive decoding.Through qualitative and quantitative analyses, we demonstrate that PruneCD consistently improves factuality with minimal inference overhead, offering a robust and practical approach to mitigating hallucinations in LLMs.
Byeongho Yu, Changhun Lee, Jungyu Jin, Eunhyeok Park
EMNLP4
2025 Grouped Speculative Decoding for Autoregressive Image Generation
Junhyuk So, Juncheol Shin, Hyunho Kook, Eunhyeok Park
ICCV4
2025 Merge-Friendly Post-Training Quantization for Multi-Target Domain Adaptation
abstract
Model merging has emerged as a powerful technique for combining task-specific weights, achieving superior performance in multi-target domain adaptation. However, when applied to practical scenarios, such as quantized models, new challenges arise. In practical scenarios, quantization is often applied to target-specific data, but this process restricts the domain of interest and introduces discretization effects, making model merging highly non-trivial. In this study, we analyze the impact of quantization on model merging through the lens of error barriers. Leveraging these insights, we propose a novel post-training quantization, HDRQ - Hessian and distant regularizing quantization - that is designed to consider model merging for multi-target domain adaptation. Our approach ensures that the quantization process incurs minimal deviation from the source pre-trained model while flattening the loss surface to facilitate smooth model merging. To our knowledge, this is the first study on this challenge, and extensive experiments confirm its effectiveness.
Juncheol Shin, Minsang Seok, Seonggon Kim, Eunhyeok Park
ICML4
2025 Partial-Sum Quantization Based on Pseudo-Quantization Noise for Variation-Tolerant Analog In-Memory Computing
abstract
Analog Computing-In-Memory (ACiM) accelerators with multi-level cells (MLCs) offer high density and area benefits for DNNs. To enhance efficiency, ADC resolution needs to be minimized, but this introduces significant quantization errors, lowering accuracy. Additionally, device noise and ADC integral nonlinearity (INL) noise further degrade accuracy. To address these challenges, we propose a training method that reduces ADC resolution while compensating for noise generated in ACiM arrays. By incorporating pseudo-quantization noise into Partial-Sum Training (PST), our approach not only stabilizes PST but also trains the model to become robust to ACiM-specific noise effects. Experimental results based on an industry ReRAM technology show that our PST scheme demonstrates robust noise tolerance across various ACiM configurations and maintains accuracy degradation within 1% even in the presence of cell conductance variability and ADC INL noise, while enabling low-resolution ADCs that reduces area and energy consumption by up to 16× and 31×, respectively.
Nameun Kang, Eunhyeok Park, Sangsu Park, Jongil Kim, Jaeyun Yi, Jae-Joon Kim
ISLPED2
2025 GraLoRA: Granular Low-Rank Adaptation for Parameter-Efficient Fine-Tuning
abstract
Low-Rank Adaptation (LoRA) is a popular method for parameter-efficient fine-tuning (PEFT) of generative models, valued for its simplicity and effectiveness. Despite recent enhancements, LoRA still suffers from a fundamental limitation: overfitting when the bottleneck is widened. It performs best at ranks 32–64, yet its accuracy stagnates or declines at higher ranks, still falling short of full fine-tuning (FFT) performance. We identify the root cause as LoRA’s structural bottleneck, which introduces gradient entanglement to the unrelated input channels and distorts gradient propagation. To address this, we introduce a novel structure, Granular Low-Rank Adaptation (GraLoRA) that partitions weight matrices into sub-blocks, each with its own low-rank adapter. With negligible computational or storage cost, GraLoRA overcomes LoRA’s limitations, effectively increases the representational capacity, and more closely approximates FFT behavior. Experiments on code generation, commonsense reasoning, mathematical reasoning, general language understanding, and image generation benchmarks show that GraLoRA consistently outperforms LoRA and other baselines, achieving up to +8.5\% absolute gain in Pass@1 on HumanEval+. These improvements hold across model sizes and rank settings, making GraLoRA a scalable and robust solution for PEFT.
Yeonjoon Jung, Daehyun Ahn, Taesu Kim, Eunhyeok Park
NeurIPS5
2025 Improving Generative Behavior Cloning via Self-Guidance and Adaptive Chunking
abstract
Generative Behavior Cloning (GBC) is a simple yet effective framework for robot learning, particularly in multi-task settings. Recent GBC methods often employ diffusion policies with open-loop (OL) control, where actions are generated via a diffusion process and executed in multi-step chunks without replanning. While this approach has demonstrated strong success rates and generalization, its inherent stochasticity can result in erroneous action sampling, occasionally leading to unexpected task failures. Moreover, OL control suffers from delayed responses, which can degrade performance in noisy or dynamic environments. To address these limitations, we propose two novel techniques to enhance the consistency and reactivity of diffusion policies: (1) self-guidance, which improves action fidelity by leveraging past observations and implicitly promoting future-aware behavior; and (2) adaptive chunking, which selectively updates action sequences when the benefits of reactivity outweigh the need for temporal consistency. Extensive experiments show that our approach substantially improves GBC performance across a wide range of simulated and real-world robotic manipulation tasks.
Junhyuk So, Chiwoong Lee, Shinyoung Lee, Jungseul Ok, Eunhyeok Park
NeurIPS5
2025 PTQ4VM: Post-Training Quantization for Visual Mamba
abstract
Visual Mamba is an approach that extends the selective space state model, Mamba, to vision tasks. It processes image tokens sequentially in a fixed order, accumulating information to generate outputs. Despite its growing popularity for delivering high-quality outputs at a low computational cost across various tasks, Visual Mamba is highly susceptible to quantization, which makes further performance improvements challenging. Our analysis reveals that the fixed token access order in Visual Mamba introduces unique quantization challenges, which we categorize into three main issues: 1) token-wise variance, 2) channel-wise outliers, and 3) a long tail of activations. To address these challenges, we propose Post-Training Quantization for Visual Mamba (PTQ4VM), which introduces two key strategies: Per-Token Static (PTS) quantization and Joint Learning of Smoothing Scale and Step Size (JLSS). To the our best knowledge, this is the first quantization study on Visual Mamba. PTQ4VM can be applied to various Visual Mamba backbones, converting the pre-trained model to a quantized format in under 15 minutes without notable quality degradation. Extensive experiments on large-scale classification and regression tasks demonstrate its effectiveness, achieving up to 1.83x speedup on GPUs with negligible accuracy loss compared to FP16. Our code is available at https://github.com/YoungHyun197/ptq4vm.
Younghyun Cho, Changhun Lee, Seonggon Kim, Eunhyeok Park
WACV4
2025 Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
abstract
Recently, Video-Language Models (VideoLMs) have demonstrated remarkable capabilities, offering significant potential for flexible and powerful video query systems. These models typically rely on Vision Transformers (ViTs), which process video frames individually to extract visual embeddings. However, generating embeddings for large-scale videos requires ViT inferencing across numerous frames, posing a major hurdle to real-world deployment and necessitating solutions for integration into scalable video data management systems. This paper introduces Déjà Vu, a video-language query engine that accelerates ViT-based VideoLMs by reusing computations across consecutive frames. At its core is ReuseViT, a modified ViT model specifically designed for VideoLM tasks, which learns to detect inter-frame reuse opportunities, striking an effective balance between accuracy and reuse. Although ReuseViT significantly reduces computation, these savings do not directly translate into performance gains on GPUs. To overcome this, Déjà Vu integrates memory-compute joint compaction techniques that convert the FLOP savings into tangible performance gains. Evaluations on three VideoLM tasks show that Déjà Vu accelerates embedding generation by up to a 2.64× within a 2% error bound, dramatically enhancing the practicality of VideoLMs for large-scale video analytics.
Jinwoo Hwang, Yoonsung Kim, Guseul Heo, Hojoon Kim, Yunseok Jeong, Tadiwos Meaza, Eunhyeok Park, Jeongseob Ahn, Jongse Park
Proc. VLDB Endow.9
2024 OWQ: Outlier-Aware Weight Quantization for Efficient Fine-Tuning and Inference of Large Language Models
abstract
Large language models (LLMs) with hundreds of billions of parameters require powerful server-grade GPUs for inference, limiting their practical deployment. To address this challenge, we introduce the outlier-aware weight quantization (OWQ) method, which aims to minimize LLM's footprint through low-precision representation. OWQ prioritizes a small subset of structured weights sensitive to quantization, storing them in high-precision, while applying highly tuned quantization to the remaining dense weights. This sensitivity-aware mixed-precision scheme reduces the quantization error notably, and extensive experiments demonstrate that 3.1-bit models using OWQ perform comparably to 4-bit models optimized by OPTQ. Furthermore, OWQ incorporates a parameter-efficient fine-tuning for task-specific adaptation, called weak column tuning (WCT), enabling accurate task-specific LLM adaptation with minimal memory overhead in the optimized format. OWQ represents a notable advancement in the flexibility, efficiency, and practicality of LLM optimization literature. The source code is available at https://github.com/xvyaward/owq.
Changhun Lee, Jungyu Jin, Taesu Kim, Eunhyeok Park
AAAI5
2024 Diffusion Model Compression for Image-to-Image Translation
Geonung Kim, Eunhyeok Park, Sunghyun Cho
ACCV (5)3
2024 FRDiff : Feature Reuse for Universal Training-Free Acceleration of Diffusion Models
Junhyuk So, Eunhyeok Park
ECCV (73)3
2024 Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders
abstract
Emerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXL.mem protocol provides minimal latency overhead through an optimized protocol stack, frequent CXL memory accesses can result in significant slowdowns for memory-bound applications whether they are latency-sensitive or bandwidth-intensive. The near-data processing (NDP) in the CXL controller promises to overcome such limitations of passive CXL memory. However, prior work on NDP in CXL memory proposes application-specific units that are not suitable for practical CXL memory-based systems that should support various applications. On the other hand, existing CPU or GPU cores are not cost-effective for NDP because they are not optimized for memory-bound applications. In addition, the communication between the host processor and CXL controller for NDP offloading should achieve low latency, but existing CXL.io/PCIe-based mechanisms incur$\mu\mathbf{s}-\mathbf{scale}$latency and are not suitable for fine-grained NDP.
Hyungkyu Ham, Jeongmin Hong 0001, Geonwoo Park, Yunseon Shin, Okkyun Woo, Wonhyuk Yang, Jinhoon Bae, Eunhyeok Park, Hyojin Sung, Eui-Cheol Lim, Gwangsun Kim
MICRO8
2023 NIPQ: Noise proxy-based Integrated Pseudo-Quantization
abstract
Straight-through estimator (STE), which enables the gradient flow over the non-differentiable function via approximation, has been favored in studies related to quantization-aware training (QAT). However, STE incurs unstable convergence during QAT, resulting in notable quality degradation in low precision. Recently, pseudo-quantization training has been proposed as an alternative approach to updating the learnable parameters using the pseudo-quantization noise instead of STE. In this study, we propose a novel noise proxy-based integrated pseudo-quantization (NIPQ) that enables unified support of pseudo-quantization for both activation and weight by integrating the idea of truncation on the pseudo-quantization framework. NIPQ updates all of the quantization parameters (e.g., bit-width and truncation boundary) as well as the network parameters via gradient descent without STE instability. According to our extensive experiments, NIPQ outperforms existing quantization algorithms in various vision and language applications by a large margin.
Juncheol Shin, Junhyuk So, Sein Park, Seungyeop Kang, Sungjoo Yoo, Eunhyeok Park
CVPR6
2023 INSTA-BNN: Binary Neural Network with INSTAnce-aware Threshold
abstract
Binary Neural Networks (BNNs) have emerged as a promising solution for reducing the memory footprint and compute costs of deep neural networks, but they suffer from quality degradation due to the lack of freedom as activations and weights are constrained to the binary values. To compensate for the accuracy drop, we propose a novel BNN design called Binary Neural Network with INSTAnce-aware threshold (INSTA-BNN), which controls the quantization threshold dynamically in an input-dependent or instance-aware manner. According to our observation, higher-order statistics can be a representative metric to estimate the characteristics of the input distribution. INSTA-BNN is designed to adjust the threshold dynamically considering various information, including higher-order statistics, but it is also optimized judiciously to realize minimal overhead on a real device. Our extensive study shows that INSTA-BNN outperforms the baseline by 3.0% and 2.8% on the ImageNet classification task with comparable computing cost, achieving 68.5% and 72.2% top-1 accuracy on ResNet-18 and MobileNetV1 based models, respectively.
Changhun Lee, Eunhyeok Park, Jae-Joon Kim
ICCV3
2023 Temporal Dynamic Quantization for Diffusion Models
abstract
Diffusion model has gained popularity in vision applications due to its remarkable generative performance and versatility. However, its high storage and computation demands, resulting from the model size and iterative generation, hinder its use on mobile devices. Existing quantization techniques struggle to maintain performance even in 8-bit precision due to the diffusion model's unique property of temporal variation in activation. We introduce a novel quantization method that dynamically adjusts the quantization interval based on time step information, significantly improving output quality. Unlike conventional dynamic quantization techniques, our approach has no computational overhead during inference and is compatible with both post-training quantization (PTQ) and quantization-aware training (QAT). Our extensive experiments demonstrate substantial improvements in output quality with the quantized model across various configurations.
Junhyuk So, Daehyun Ahn, Eunhyeok Park
NeurIPS5
2023 Searching for Robust Binary Neural Networks via Bimodal Parameter Perturbation
abstract
Binary neural networks (BNNs) are advantageous in performance and memory footprint but suffer from low accuracy due to their limited expression capability. Recent works have tried to enhance the accuracy of BNNs via a gradient-based search algorithm and showed promising results. However, the mixture of architecture search and binarization induce the instability of the search process, resulting in convergence to the suboptimal point. To address this issue, we propose a BNN architecture search framework with bimodal parameter perturbation. The bimodal parameter perturbation can improve the stability of gradient-based architecture search by reducing the sharpness of the loss surface along both weight and architecture parameter axes. In addition, we refine the inverted bottleneck convolution block for having robustness with BNNs. The synergy of the refined space and the stabilized search process allows us to find out the accurate BNNs with high computation efficiency. Experimental results show that our framework finds the best architecture on CIFAR-100 and ImageNet datasets in the existing search space for BNNs. We also tested our framework on another search space based on the inverted bottleneck convolution block, and the selected BNN models using our approach achieved the highest accuracy on both datasets with a much smaller number of equivalent operations than previous works.
Daehyun Ahn, Taesu Kim, Eunhyeok Park, Jae-Joon Kim
WACV4
2022 One-shot tuner for deep learning compilers
abstract
Auto-tuning DL compilers are gaining ground as an optimizing back-end for DL frameworks. While existing work can generate deep learning models that exceed the performance of hand-tuned libraries, they still suffer from prohibitively long auto-tuning time due to repeated hardware measurements in large search spaces. In this paper, we take a neural-predictor inspired approach to reduce the auto-tuning overhead and show that a performance predictor model trained prior to compilation can produce optimized tensor operation codes without repeated search and hardware measurements. To generate a sample-efficient training dataset, we extend input representation to include task-specific information and to guide data sampling methods to focus on learning high-performing codes. We evaluated the resulting predictor model, One-Shot Tuner, against AutoTVM and other prior work, and the results show that One-Shot Tuner speeds up compilation by 2.81x to 67.7x compared to prior work while providing comparable or improved inference time for CNN and Transformer models.
Jaehun Ryu, Eunhyeok Park, Hyojin Sung
CC2
2022 BASQ: Branch-wise Activation-clipping Search Quantization for Sub-4-bit Neural Networks
Han-Byul Kim, Eunhyeok Park, Sungjoo Yoo
ECCV (12)2
2022 Symmetry Regularization and Saturating Nonlinearity for Robust Quantization
Sein Park, Yeongsang Jang, Eunhyeok Park
ECCV (11)3
2022 Online Hybrid Lightweight Representations Learning: Its Application to Visual Tracking
abstract
This paper presents a novel hybrid representation learning framework for streaming data, where an image frame in a video is modeled by an ensemble of two distinct deep neural networks; one is a low-bit quantized network and the other is a lightweight full-precision network. The former learns coarse primary information with low cost while the latter conveys residual information for high fidelity to original representations. The proposed parallel architecture is effective to maintain complementary information since fixed-point arithmetic can be utilized in the quantized network and the lightweight model provides precise representations given by a compact channel-pruned network. We incorporate the hybrid representation technique into an online visual tracking task, where deep neural networks need to handle temporal variations of target appearances in real-time. Compared to the state-of-the-art real-time trackers based on conventional deep neural networks, our tracking algorithm demonstrates competitive accuracy on the standard benchmarks with a small fraction of computational cost and memory footprint.
Ilchae Jung, Minji Kim 0002, Eunhyeok Park, Bohyung Han
IJCAI3
2021 Fine-grained Semantics-aware Representation Enhancement for Self-supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation has been widely studied, owing to its practical importance and recent promising improvements. However, most works suffer from limited supervision of photometric consistency, especially in weak texture regions and at object boundaries. To overcome this weakness, we propose novel ideas to improve self-supervised monocular depth estimation by leveraging cross-domain information, especially scene semantics. We focus on incorporating implicit semantic knowledge into geometric representation enhancement and suggest two ideas: a metric learning approach that exploits the semantics-guided local geometry to optimize intermediate depth representations and a novel feature fusion module that judiciously utilizes cross-modality between two heterogeneous feature representations. We comprehensively evaluate our methods on the KITTI dataset and demonstrate that our method outperforms state-of-the-art methods. The source code is available at https://github.com/hyBlue/FSRE-Depth.
Hyunyoung Jung 0001, Eunhyeok Park, Sungjoo Yoo
ICCV2
2021 FPGA Prototyping of Systolic Array-based Accelerator for Low-Precision Inference of Deep Neural Networks
abstract
In this study, we aim to design an energy-efficient computation system for deep neural networks on edge devices. To maximize energy efficiency, we design a novel hardware accelerator that supports low-precision computation and sparsity-aware structured zero-skipping on top of the well-known systolic-array structure. In addition, we introduce a full-stack software platform, including a model optimizer, instruction compiler, and host interface, to translate the pre-trained PyTorch model to the proposed accelerator and orchestrate it automatically. We validate the entire system by prototyping the accelerator on the Xilinx Alveo U250 FPGA board and demonstrating the inference of the 4-bit ResNet-50 model through the software stack. According to our experiment, our platform shows 317 GOPS inference speed and 51.96 GOPS/W energy efficiency for ResNet-50 on Xilinx Alveo U250 FPGA at 108 MHz, which is comparable to the advanced commercial acceleration system in terms of energy efficiency.
Soobeom Kim, Seunghwan Cho, Eunhyeok Park, Sungjoo Yoo
RSP3
2020 PROFIT: A Novel Training Method for sub-4-bit MobileNet Models
Eunhyeok Park, Sungjoo Yoo
ECCV (6)1
2020 MEANTIME: Mixture of Attention Mechanisms with Multi-temporal Embeddings for Sequential Recommendation
abstract
Recently, self-attention based models have achieved state-of-the-art performance in sequential recommendation task. Following the custom from language processing, most of these models rely on a simple positional embedding to exploit the sequential nature of the user’s history. However, there are some limitations regarding the current approaches. First, sequential recommendation is different from language processing in that timestamp information is available. Previous models have not made good use of it to extract additional contextual information. Second, using a simple embedding scheme can lead to information bottleneck since the same embedding has to represent all possible contextual biases. Third, since previous models use the same positional embedding in each attention head, they can wastefully learn overlapping patterns. To address these limitations, we propose MEANTIME (MixturE of AtteNTIon mechanisms with Multi-temporal Embeddings) which employs multiple types of temporal embeddings designed to capture various patterns from the user’s behavior sequence, and an attention structure that fully leverages such diversity. Experiments on real-world data show that our proposed method outperforms current state-of-the-art sequential recommendation methods, and we provide an extensive ablation study to analyze how the model gains from the diverse positional information.
Sung Min Cho, Eunhyeok Park, Sungjoo Yoo
RecSys2
2019 Tag2Pix: Line Art Colorization Using Text Tag With SECat and Changing Loss
abstract
Line art colorization is expensive and challenging to automate. A GAN approach is proposed, called Tag2Pix, of line art colorization which takes as input a grayscale line art and color tag information and produces a quality colored image. First, we present the Tag2Pix line art colorization dataset. A generator network is proposed which consists of convolutional layers to transform the input line art, a pre-trained semantic extraction network, and an encoder for input color information. The discriminator is based on an auxiliary classifier GAN to classify the tag information as well as genuineness. In addition, we propose a novel network structure called SECat, which makes the generator properly colorize even small features such as eyes, and also suggest a novel two-step training method where the generator and discriminator first learn the notion of object and shape and then, based on the learned notion, learn colorization, such as where and how to place which color. We present both quantitative and qualitative evaluations which prove the effectiveness of the proposed method.
Hyunsu Kim, Ho Young Jhoo, Eunhyeok Park, Sungjoo Yoo
ICCV3
2018 Value-Aware Quantization for Training and Inference of Neural Networks
Eunhyeok Park, Sungjoo Yoo, Peter Vajda
ECCV (4)1
2018 Energy-Efficient Neural Network Accelerator Based on Outlier-Aware Low-Precision Computation
abstract
Owing to the presence of large values, which we call outliers, conventional methods of quantization fail to achieve significantly low precision, e.g., four bits, for very deep neural networks, such as ResNet-101. In this study, we propose a hardware accelerator, called the outlier-aware accelerator (OLAccel). It performs dense and low-precision computations for a majority of data (weights and activations) while efficiently handling a small number of sparse and high-precision outliers (e.g., amounting to 3% of total data). The OLAccel is based on 4-bit multiply-accumulate (MAC) units and handles outlier weights and activations in a different manner. For outlier weights, it equips SIMD lanes of MAC units with an additional MAC unit, which helps avoid cycle overhead for the majority of outlier occurrences, i.e., a single occurrence in the SIMD lanes. The OLAccel performs computations using outlier activation on dedicated, high-precision MAC units. In order to avoid coherence problem due to updates from low- and high-precision computation units, both units update partial sums in a pipelined manner. Our experiments show that the OLAccel can reduce by 43.5% (27.0%), 56.7% (36.3%), and 62.2% (49.5%) energy consumption for AlexNet, VGG-16, and ResNet-18, respectively, compared with a 16-bit (8-bit) state-of-the-art zero-aware accelerator. The energy gain mostly comes from the memory components, the DRAM, and on-chip memory due to reduced precision.
Eunhyeok Park, Dongyoung Kim, Sungjoo Yoo
ISCA1
2018 McDRAM: Low Latency and Energy-Efficient Matrix Computations in DRAM
abstract
We propose a novel memory architecture for in-memory computation called McDRAM, where DRAM dies are equipped with a large number of multiply accumulate (MAC) units to perform matrix computation for neural networks. By exploiting high internal memory bandwidth and reducing offchip memory accesses, McDRAM realizes both low latency and energy efficient computation. In our experiments, we obtained the chip layout based on the state-of-the-art memory, LPDDR4 where McDRAM is equipped with 2048 MACs in a single chip package with a small area overhead (4.7%). Compared with the state-ofthe-art accelerator, TPU and the power-efficient GPU, Nvidia P4, McDRAM offers 9.5× and 14.4× speedup, respectively, in the case that the large-scale MLPs and RNNs adopt the batch size of 1. McDRAM also gives 2.1× and 3.7× better computational efficiency in TOPS/W than TPU and P4, respectively, for the large batches.
Hyunsung Shin, Dongyoung Kim, Eunhyeok Park, Yongsik Park, Sungjoo Yoo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2017 Weighted-Entropy-Based Quantization for Deep Neural Networks
abstract
Quantization is considered as one of the most effective methods to optimize the inference cost of neural network models for their deployment to mobile and embedded systems, which have tight resource constraints. In such approaches, it is critical to provide low-cost quantization under a tight accuracy loss constraint (e.g., 1%). In this paper, we propose a novel method for quantizing weights and activations based on the concept of weighted entropy. Unlike recent work on binary-weight neural networks, our approach is multi-bit quantization, in which weights and activations can be quantized by any number of bits depending on the target accuracy. This facilitates much more flexible exploitation of accuracy-performance trade-off provided by different levels of quantization. Moreover, our scheme provides an automated quantization flow based on conventional training algorithms, which greatly reduces the design-time effort to quantize the network. According to our extensive evaluations based on practical neural network models for image classification (AlexNet, GoogLeNet and ResNet-50/101), object detection (R-FCN with 50-layer ResNet), and language modeling (an LSTM network), our method achieves significant reductions in both the model size and the amount of computation with minimal accuracy loss. Also, compared to existing quantization schemes, ours provides higher accuracy with a similar resource constraint and requires much lower design effort.
Eunhyeok Park, Junwhan Ahn, Sungjoo Yoo
CVPR1
2015 Memory fast-forward: a low cost special function unit to enhance energy efficiency in GPU for big data processing
Eunhyeok Park, Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Sunggu Lee
DATE1
2015 Locality-aware vertex scheduling for GPU-based graph computation
abstract
Graph computation is becoming more and more popular in machine learning, big data analytics, etc. For such workloads, GPU is considered as an efficient execution platform since graph computation is characterized by massively parallel computation and high demand of memory bandwidth. In our investigation, existing GPU programming methods for graph computation do not fully exploit high memory bandwidth as well as high computing power in GPU. We propose a novel optimization called locality-aware vertex scheduling, which aims at minimizing memory requests by adjusting the order of vertex computations to improve temporal locality of vertex data stored in on-chip caches. Experiments with nine real-world graphs and three graph algorithms on the recent GPU platform show that the proposed method offers a significant speedup (average 46%) over the state-of-the-art graph algorithm implementation on GPUs.
Hyunsun Park, Junwhan Ahn, Eunhyeok Park, Sungjoo Yoo
VLSI-SoC3