Dongyoung Kim

dblp:58/6642 · DBLP profile ↗
← Back
27ranked-venue papers
13as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 8 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 9 since 2021Systems, architecture and hardware · 6 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
abstract
Cross-view image retrieval, particularly street-to-satellite matching, is a critical task for applications such as autonomous navigation, urban planning, and localization in GPS-denied environments. However, existing approaches often require supervised training on curated datasets and rely on panoramic or UAV-based images, which limits real-world deployment. In this paper, we present a simple yet effective cross-view image retrieval framework that leverages a pretrained vision encoder and a large language model (LLM), requiring no additional training. Given a monocular street-view image, our method extracts geographic cues through web-based image search and LLM-based location inference, generates a satellite query via geocoding API, and retrieves matching tiles using a pretrained vision encoder (e.g., DINOv2) with PCA-based whitening feature refinement. Despite not using ground-truth supervision or finetuning, our proposed method outperforms prior learning-based approaches on the benchmark dataset under zero-shot settings. Moreover, our pipeline enables automatic construction of semantically aligned street-to-satellite datasets, which is offering a scalable and cost-efficient alternative to manual annotation. All source codes will be made publicly available at https://jeonghomin.github.io/ street2orbit.github.io/.
Jeongho Min, Dongyoung Kim, Jaehyup Lee
WACV2
2026 PSG-free multi-view facial imaging and attention-based fusion for OSA severity classification
Dongyoung Kim, Yunhee Woo, Jae Min Jeong, Seunghun Oh, Young Woong Ko, Il-Hwan Lee, Jeong-Gun Lee
Expert Syst. Appl.1
2025 Pixel-Level Fire Origin Localization via Digital Twin Mapping for Wildfire Surveillance Framework
abstract
Wildfire monitoring systems play a critical role in minimizing environmental and societal damage. Recent advances in computer vision, particularly deep learning-based fire detection, have enabled more accurate and scalable solutions. However, conventional fire detection methods often struggle with wildfire scenarios due to wide spatial extent, the demand for precise localization, and the urgency of early response. To overcome these challenges, we propose a wildfire monitoring framework capable of pixel-level fire origin localization mapped onto a GPS-calibrated digital twin of mountainous terrain. Our system integrates visual fire detection with terrain-aware 3D projection, enabling accurate mapping of fire origins to real-world coordinates. Experimental results on wildfire datasets demonstrate that our method achieves high accuracy in both early fire detection and precise localization, offering a practical and scalable solution for real-world wildfire monitoring.
Dongyoung Kim, In-Su Jang, Kwang-Ju Kim, Kyoungoh Lee
AVSS1
2025 PETS2025: Multi-Authority Multi-Sensor Maritime Surveillance Challenge and Evaluation
abstract
This paper presents the outcomes of the PETS2025 challenge, held in conjunction with AVSS 2025 and sponsored by the EU-funded EURMARS project. The challenge introduces a novel maritime surveillance dataset comprising image sequences captured by diverse multi-altitude, multimodal sensors, reflecting the real-world multi-authority environment. The key tasks include: (1) object detection using various sensors across different platforms (ground-based and low-altitude aerial) and spectral ranges (visible, thermal, ultraviolet (UV), and short-wave infrared (SWIR)); (2) long-term tracking of targets in maritime environments spanning both sea and land; and (3) approximating target geolocations by using sensor imagery and telemetry data. Performance evaluations of results submitted by 12 international participants are discussed. The results show the effectiveness of these submissions and highlight ongoing challenges posed by heterogeneous sensors and complex environments. These challenges emphasise the need to further improve detection, tracking, and geolocation approximation for maritime and coastal surveillance.
Thanet Markchom, Jonathan N. Boyle, Lulu Chen, James M. Ferryman, Matteo Marturini, Stephan Veigl, Andreas Opitz, Andreas Kriechbaum-Zabini, Romaios Bratskas, Anastasios Gkamaris, Dimitris Papachristos, George Leventakis, Wenjun Fan, Hsiang-Wei Huang, Jeng-Neng Hwang, Pyong-Kun Kim, Kwangju Kim, Chung-I Huang, Kenta Saito, Shunta Kaneko, Kyoko Sudo, Nguyen Thanh Thien, Meng-Yu Kao, Jun-Wei Hsieh, Teepakorn Lilek, Tossapol Pomsuwan, Jinjie Gu, Tianyang Xu 0001, Xuefeng Zhu 0003, Xiaojun Wu 0001, Josef Kittler, Stephanie Stacy, Alfredo Gabaldon, Peter Tu, Dongyoung Kim, Kyoungoh Lee
AVSS36
2025 CCMNet: Leveraging Calibrated Color Correction Matrices for Cross-Camera Color Constancy
abstract
Computational color constancy, or white balancing, is a key module in a camera's image signal processor (ISP) that corrects color casts from scene lighting. Because this operation occurs in the camera-specific raw color space, white balance algorithms must adapt to different cameras. This paper introduces a learning-based method for cross-camera color constancy that generalizes to new cameras without retraining. Our method leverages pre-calibrated color correction matrices (CCMs) available on ISPs that map the camera's raw color space to a standard space (e.g., CIE XYZ). Our method uses these CCMs to transform predefined illumination colors (i.e., along the Planckian locus) into the test camera's raw space. The mapped illuminants are encoded into a compact camera fingerprint embedding (CFE) that enables the network to adapt to unseen cameras. To prevent overfitting due to limited cameras and CCMs during training, we introduce a data augmentation technique that interpolates between cameras and their CCMs. Experimental results across multiple datasets and backbones show that our method achieves state-of-the-art cross-camera color constancy while remaining lightweight and relying only on data readily available in camera ISPs.
Dongyoung Kim, Mahmoud Afifi, Dongyun Kim, Michael S. Brown, Seon Joo Kim
ICCV1
2025 Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
abstract
Aligning large language models (LLMs) with human preferences becomes a key component to obtaining state-of-the-art performance, but it yields a huge cost to construct a large human-annotated preference dataset. To tackle this problem, we propose a new framework, Spread Preference Annotation with direct preference judgment (SPA), that boosts the alignment of LLMs using only a very small amount of human-annotated preference data. Our key idea is leveraging the human prior knowledge within the small (seed) data and progressively improving the alignment of LLM, by iteratively generating the responses and learning from them with the self-annotated preference data. To be specific, we propose to derive the preference label from the logits of LLM to explicitly extract the model's inherent preference. Compared to the previous approaches using external reward models or implicit in-context learning, we observe that the proposed approach is significantly more effective. In addition, we introduce a noise-aware preference learning algorithm to mitigate the risk of low quality within generated preference data. Our experimental results demonstrate that the proposed framework significantly boosts the alignment of LLMs. For example, we achieve superior alignment performance on AlpacaEval 2.0 with only 3.3% of the ground-truth preference labels in the Ultrafeedback data compared to the cases using the entire data or state-of-the-art baselines.
Dongyoung Kim, Kimin Lee, Jinwoo Shin, Jaehyung Kim 0001
ICLR1
2025 Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
abstract
Large Vision-Language Models (LVLMs) have recently shown great promise in advancing robotics by combining embodied reasoning with robot control. A common approach involves training on embodied reasoning tasks related to robot control using Supervised Fine-Tuning (SFT). However, SFT datasets are often heuristically constructed and not explicitly optimized for improving robot control. Furthermore, SFT often leads to issues such as catastrophic forgetting and reduced generalization performance. To address these limitations, we introduce Robot-R1, a novel framework that leverages reinforcement learning to enhance embodied reasoning specifically for robot control. Robot-R1 learns to predict the next keypoint state required for task completion, conditioned on the current scene image and environment metadata derived from expert demonstrations. Inspired by the DeepSeek-R1 learning approach, Robot-R1 samples reasoning-based responses and reinforces those that lead to more accurate predictions. Our experiments show that models trained with Robot-R1 outperform SFT methods on embodied reasoning tasks. Despite having only 7B parameters, Robot-R1 even surpasses GPT-4o on reasoning tasks related to low-level action control, such as spatial and movement reasoning.
Dongyoung Kim, Huiwon Jang, Jinwoo Shin, Younggyo Seo
NeurIPS1
2024 MOVES: Motion-Oriented VidEo Sampling for Natural Language-Based Vehicle Retrieval
abstract
Retrieving the target vehicle through natural language descriptions plays a crucial role in intelligent transportation systems. Existing methods tackle this task by employing models that leverage the correlation between textual and visual representations, such as CLIP. However, these models struggle to capture the temporal characteristics of video data, and researchers enhance temporal understanding performance through various data augmentation and video encoders. Yet, conventional approaches in previous studies often overlook the detailed temporal characteristics of vehicles. To overcome this limitation, we introduce a MOVES: Motion-Oriented VidEo Sampling method to effectively utilize the motion information of the target vehicle. Furthermore, we construct a robust model by implementing a re-ranking algorithm to address a variety of vehicle attributes. As a result, our proposed model achieves state-of-the-art performance on the public vehicle retrieval dataset.
Dongyoung Kim, Kyoungoh Lee, In-Su Jang, Kwang-Ju Kim, Pyong-Kun Kim, Jaejun Yoo 0001
AVSS1
2024 Separated and Independent Contrastive Learning on Labeled and Unlabeled Samples: Boosting Performance on Long-tail Semi-supervised Learning
Dongyoung Kim, Jeong-Gun Lee
BMVC1
2024 Attentive Illumination Decomposition Model for Multi-Illuminant White Balancing
abstract
White balance (WB) algorithms in many commercial cameras assume single and uniform illumination, leading to undesirable results when multiple lighting sources with different chromaticities exist in the scene. Prior research on multi-illuminant WB typically predicts illumination at the pixel level without fully grasping the scene's actual lighting conditions, including the number and color of light sources. This often results in unnatural outcomes lacking in overall consistency. To handle this problem, we present a deep white balancing model that leverages the slot attention, where each slot is in charge of representing individual illuminants. This design enables the model to generate chromaticities and weight maps for individual illuminants, which are then fused to compose the final illumination map. Furthermore, we propose the centroid-matching loss, which regulates the activation of each slot based on the color range, thereby enhancing the model to separate illumination more effectively. Our method achieves the state-of-the-art performance on both single- and multi-illuminant WB benchmarks, and also offers additional information such as the number of illuminants in the scene and their chromaticity. This capability allows for illumination editing, an application not feasible with prior methods.
Dongyoung Kim, Jinwoo Kim 0007, Junsang Yu 0001, Seon Joo Kim
CVPR1
2024 Learning to Correct for QA Reasoning with Black-box LLMs
abstract
An open challenge in recent machine learning is about how to improve the reasoning capability of large language models (LLMs) in a blackbox setting, i.e., without access to detailed information such as output token probabilities.Existing approaches either rely on accessibility (which is often unrealistic) or involve significantly increased train-and inference-time costs.This paper addresses those limitations or shortcomings by proposing a novel approach, namely COBB (Correct for improving QA reasoning of Black-Box LLMs).It uses a trained adaptation model to perform a seq2seq mapping from the often-imperfect reasonings of the original black-box LLM to the correct or improved reasonings.Specifically, the adaptation model is initialized with a relatively small open-source LLM and adapted over a collection of sub-sampled training pairs.To select the representative pairs of correct and incorrect reasonings, we formulated the dataset construction as an optimization problem that minimizes the statistical divergence between the sampled subset and the entire collection, and solved it via a genetic algorithm.We then train the adaptation model over the sampled pairs by contrasting the likelihoods of correct and incorrect reasonings.Our experimental results demonstrate that COBB significantly improves reasoning accuracy across various QA benchmarks, compared to the best-performing adaptation baselines.1
Jaehyung Kim 0001, Dongyoung Kim, Yiming Yang 0002
EMNLP2
2024 Visual Representation Learning with Stochastic Frame Prediction
abstract
Self-supervised learning of image representations by predicting future frames is a promising direction but still remains a challenge. This is because of the under-determined nature of frame prediction; multiple potential futures can arise from a single current frame. To tackle this challenge, in this paper, we revisit the idea of stochastic video generation that learns to capture uncertainty in frame prediction and explore its effectiveness for representation learning. Specifically, we design a framework that trains a stochastic frame prediction model to learn temporal information between frames. Moreover, to learn dense information within each frame, we introduce an auxiliary masked image modeling objective along with a shared decoder architecture. We find this architecture allows for combining both objectives in a synergistic and compute-efficient manner. We demonstrate the effectiveness of our framework on a variety of tasks from video label propagation and vision-based robot learning domains, such as video segmentation, pose tracking, vision-based robotic locomotion, and manipulation tasks. Code is available on the project webpage: https://sites.google.com/view/2024rsp.
Huiwon Jang, Dongyoung Kim, Jinwoo Shin, Pieter Abbeel, Younggyo Seo
ICML2
2024 Bridging the Domain Gap: A Simple Domain Matching Method for Reference-Based Image Super-Resolution in Remote Sensing
abstract
Recently, reference-based image super-resolution (RefSR) has shown excellent performance in image super-resolution (SR) tasks. The main idea of RefSR is to utilize additional information from the reference (Ref) image to recover the high-frequency components in low-resolution (LR) images. By transferring relevant textures through feature matching, RefSR models outperform existing single-image SR (SISR) models. However, their performance significantly declines when a domain gap between Ref and LR images exists, which often occurs in real-world scenarios, such as satellite imaging. In this letter, we introduce a domain matching (DM) module that can be seamlessly integrated with existing RefSR models to enhance their performance in a plug-and-play manner. To the best of our knowledge, we are the first to explore DM-based RefSR in remote sensing image processing. Our analysis reveals that their domain gaps often occur in different satellites, and our model effectively addresses these challenges, whereas existing models struggle. Our experiments demonstrate that the proposed DM module improves SR performance both qualitatively and quantitatively for remote sensing SR tasks.
Jeongho Min, Yejun Lee, Dongyoung Kim, Jaejun Yoo 0001
IEEE Geosci. Remote. Sens. Lett.3
2023 Accelerating Reinforcement Learning with Value-Conditional State Entropy Exploration
abstract
A promising technique for exploration is to maximize the entropy of visited state distribution, i.e., state entropy, by encouraging uniform coverage of visited state space. While it has been effective for an unsupervised setup, it tends to struggle in a supervised setup with a task reward, where an agent prefers to visit high-value states to exploit the task reward. Such a preference can cause an imbalance between the distributions of high-value states and low-value states, which biases exploration towards low-value state regions as a result of the state entropy increasing when the distribution becomes more uniform. This issue is exacerbated when high-value states are narrowly distributed within the state space, making it difficult for the agent to complete the tasks. In this paper, we present a novel exploration technique that maximizes the value-conditional state entropy, which separately estimates the state entropies that are conditioned on the value estimates of each state, then maximizes their average. By only considering the visited states with similar value estimates for computing the intrinsic bonus, our method prevents the distribution of low-value states from affecting exploration around high-value states, and vice versa. We demonstrate that the proposed alternative to the state entropy baseline significantly accelerates various reinforcement learning algorithms across a variety of tasks within MiniGrid, DeepMind Control Suite, and Meta-World benchmarks. Source code is available at https://sites.google.com/view/rl-vcse.
Dongyoung Kim, Jinwoo Shin, Pieter Abbeel, Younggyo Seo
NeurIPS1
2022 ATTIQA: Generalizable Image Quality Feature Extractor Using Attribute-Aware Pretraining
Daekyu Kwon, Dongyoung Kim, Sehwan Ki, Younghyun Jo, Hyong-Euk Lee, Seon Joo Kim
ACCV (4)2
2021 Large Scale Multi-Illuminant (LSMI) Dataset for Developing White Balance Algorithm under Mixed Illumination
abstract
We introduce a Large Scale Multi-Illuminant (LSMI) Dataset that contains 7,486 images, captured with three different cameras on more than 2,700 scenes with two or three illuminants. For each image in the dataset, the new dataset provides not only the pixel-wise ground truth illumination but also the chromaticity of each illuminant in the scene and the mixture ratio of illuminants per pixel. Images in our dataset are mostly captured with illuminants existing in the scene, and the ground truth illumination is computed by taking the difference between the images with different illumination combination. Therefore, our dataset captures natural composition in the real-world setting with wide field-of-view, providing more extensive dataset compared to existing datasets for multi-illumination white balance. As conventional single illuminant white balance algorithms cannot be directly applied, we also apply per-pixel DNN-based white balance algorithm and show its effectiveness against using patch-wise white balancing. We validate the benefits of our dataset through extensive analysis including a user-study, and expect the dataset to make meaningful contribution for future work in white balancing.
Dongyoung Kim, Jinwoo Kim 0007, Seonghyeon Nam, Yeonkyung Lee, Nahyup Kang, Hyong-Euk Lee, ByungIn Yoo, Jae-Joon Han, Seon Joo Kim
ICCV1
2021 Sparsity-Aware and Re-configurable NPU Architecture for Samsung Flagship Mobile SoC
abstract
Of late, deep neural networks have become ubiquitous in mobile applications. As mobile devices generally require immediate response while maintaining user privacy, the demand for on-device machine learning technology is on the increase. Nevertheless, mobile devices suffer from restricted hardware resources, whereas deep neural networks involve considerable computation and communication. Therefore, the implementation of a neural-network specialized hardware accelerator, generally called neural processing unit (NPU), has started to gain attention for the mobile application processor (AP). However, NPUs for commercial mobile AP face two challenges that are difficult to realize simultaneously: execution of a wide range of applications and efficient performance.In this paper, we propose a flexible but efficient NPU architecture for a Samsung flagship mobile system-on-chip (SoC). To implement an efficient NPU, we design an energy-efficient inner-product engine that utilizes the input feature map sparsity. We propose a re-configurable MAC array to enhance the flexibility of the proposed NPU, dynamic internal memory port assignment to maximize on-chip memory bandwidth utilization, and efficient architecture to support mixed-precision arithmetic. We implement the proposed NPU using the Samsung 5nm library. Our silicon measurement experiments demonstrate that the proposed NPU achieves 290.7 FPS and 13.6 TOPS/W, when executing an 8-bit quantized Inception-v3 model [1] with a single NPU core. In addition, we analyze the proposed zero-skipping architecture in detail. Finally, we present the findings and lessons learned when implementing the commercial mobile NPU and interesting avenues for future work.
Jun-Woo Jang, Sehwan Lee, Dongyoung Kim, Hyunsun Park, Ali Shafiee, Yeongjae Choi, Channoh Kim, Yoojin Kim, Hyeongseok Yu, Hamzah Abdel-Aziz, Jun-Seok Park, Heonsoo Lee, Myeong Woo Kim, Hanwoong Jung, Heewoo Nam, Dongguen Lim, Seungwon Lee 0006, Joon-Ho Song, Suknam Kwon, Joseph Hassoun, Sukhwan Lim, Changkyu Choi
ISCA3
2020 MarioNETte: Few-Shot Face Reenactment Preserving Identity of Unseen Targets
abstract
When there is a mismatch between the target identity and the driver identity, face reenactment suffers severe degradation in the quality of the result, especially in a few-shot setting. The identity preservation problem, where the model loses the detailed information of the target leading to a defective output, is the most common failure mode. The problem has several potential sources such as the identity of the driver leaking due to the identity mismatch, or dealing with unseen large poses. To overcome such problems, we introduce components that address the mentioned problem: image attention block, target feature alignment, and landmark transformer. Through attending and warping the relevant features, the proposed architecture, called MarioNETte, produces high-quality reenactments of unseen identities in a few-shot setting. In addition, the landmark transformer dramatically alleviates the identity preservation problem by isolating the expression geometry through landmark disentanglement. Comprehensive experiments are performed to verify that the proposed framework can generate highly realistic faces, outperforming all other baselines, even under a significant mismatch of facial characteristics between the target and the driver.
Sungjoo Ha, Martin Kersner, Seokjun Seo, Dongyoung Kim
AAAI5
2020 Attentron: Few-Shot Text-to-Speech Utilizing Attention-Based Variable-Length Embedding
abstract
On account of growing demands for personalization, the need for a so-called few-shot TTS system that clones speakers with only a few data is emerging.To address this issue, we propose Attentron, a few-shot TTS model that clones voices of speakers unseen during training.It introduces two special encoders, each serving different purposes.A fine-grained encoder extracts variable-length style information via an attention mechanism, and a coarse-grained encoder greatly stabilizes the speech synthesis, circumventing unintelligible gibberish even for synthesizing speech of unseen speakers.In addition, the model can scale out to an arbitrary number of reference audios to improve the quality of the synthesized speech.According to our experiments, including a human evaluation, the proposed model significantly outperforms state-of-the-art models when generating speech for unseen speakers in terms of speaker similarity and quality.
Seungju Han 0002, Dongyoung Kim, Sungjoo Ha
INTERSPEECH3
2019 Temporal Convolution for Real-Time Keyword Spotting on Mobile Devices
abstract
Keyword spotting (KWS) plays a critical role in enabling speech-based user interactions on smart devices.Recent developments in the field of deep learning have led to wide adoption of convolutional neural networks (CNNs) in KWS systems due to their exceptional accuracy and robustness.The main challenge faced by KWS systems is the trade-off between high accuracy and low latency.Unfortunately, there has been little quantitative analysis of the actual latency of KWS models on mobile devices.This is especially concerning since conventional convolution-based KWS approaches are known to require a large number of operations to attain an adequate level of performance.In this paper, we propose a temporal convolution for real-time KWS on mobile devices.Unlike most of the 2D convolution-based KWS approaches that require a deep architecture to fully capture both low-and high-frequency domains, we exploit temporal convolutions with a compact ResNet architecture.In Google Speech Command Dataset, we achieve more than 385x speedup on Google Pixel 1 and surpass the accuracy compared to the state-of-the-art model.In addition, we release the implementation of the proposed and the baseline models including an end-to-end pipeline for training models and evaluating them on mobile devices.
Seokjun Seo, Beomjun Shin, Hyeongmin Byun, Martin Kersner, Dongyoung Kim, Sungjoo Ha
INTERSPEECH7
2018 Energy-Efficient Neural Network Accelerator Based on Outlier-Aware Low-Precision Computation
abstract
Owing to the presence of large values, which we call outliers, conventional methods of quantization fail to achieve significantly low precision, e.g., four bits, for very deep neural networks, such as ResNet-101. In this study, we propose a hardware accelerator, called the outlier-aware accelerator (OLAccel). It performs dense and low-precision computations for a majority of data (weights and activations) while efficiently handling a small number of sparse and high-precision outliers (e.g., amounting to 3% of total data). The OLAccel is based on 4-bit multiply-accumulate (MAC) units and handles outlier weights and activations in a different manner. For outlier weights, it equips SIMD lanes of MAC units with an additional MAC unit, which helps avoid cycle overhead for the majority of outlier occurrences, i.e., a single occurrence in the SIMD lanes. The OLAccel performs computations using outlier activation on dedicated, high-precision MAC units. In order to avoid coherence problem due to updates from low- and high-precision computation units, both units update partial sums in a pipelined manner. Our experiments show that the OLAccel can reduce by 43.5% (27.0%), 56.7% (36.3%), and 62.2% (49.5%) energy consumption for AlexNet, VGG-16, and ResNet-18, respectively, compared with a 16-bit (8-bit) state-of-the-art zero-aware accelerator. The energy gain mostly comes from the memory components, the DRAM, and on-chip memory due to reduced precision.
Eunhyeok Park, Dongyoung Kim, Sungjoo Yoo
ISCA2
2018 FPGA Prototyping of Low-Precision Zero-Skipping Accelerator for Neural Networks
abstract
Owing to the ever-increasing usage of neural networks in various applications from mobile devices to data centers, hardware accelerators for neural networks have been widely studied in recent years. Recently proposed accelerators mainly utilize sparsity and/or reduced precision approach to accelerate neural networks. Since neural networks go deeper to be applied to more complex applications, future hardware accelerators for neural networks may require to utilize both sparsity and very low-precision methods, which however, have not been analyzed quantitatively. In this paper, we first introduce an end-to-end FPGA prototyping flow and apply it to a neural network accelerator which supports both fine-grained zero-skipping and very low-precision. We report our analyses of resource usage and performance by varying bit-width and zero data ratio. We summarize the paper with our lessons learned from our prototyping study on future zero-aware very-low-precision accelerators.
Dongyoung Kim, Soobeom Kim, Sungjoo Yoo
RSP1
2018 McDRAM: Low Latency and Energy-Efficient Matrix Computations in DRAM
abstract
We propose a novel memory architecture for in-memory computation called McDRAM, where DRAM dies are equipped with a large number of multiply accumulate (MAC) units to perform matrix computation for neural networks. By exploiting high internal memory bandwidth and reducing offchip memory accesses, McDRAM realizes both low latency and energy efficient computation. In our experiments, we obtained the chip layout based on the state-of-the-art memory, LPDDR4 where McDRAM is equipped with 2048 MACs in a single chip package with a small area overhead (4.7%). Compared with the state-ofthe-art accelerator, TPU and the power-efficient GPU, Nvidia P4, McDRAM offers 9.5× and 14.4× speedup, respectively, in the case that the large-scale MLPs and RNNs adopt the batch size of 1. McDRAM also gives 2.1× and 3.7× better computational efficiency in TOPS/W than TPU and P4, respectively, for the large batches.
Hyunsung Shin, Dongyoung Kim, Eunhyeok Park, Yongsik Park, Sungjoo Yoo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 A novel zero weight/activation-aware hardware architecture of convolutional neural network
abstract
It is imperative to accelerate convolutional neural networks (CNNs) due to their ever-widening application areas from server, mobile to IoT devices. Based on the fact that CNNs can be characterized by a significant amount of zero values in both kernel weights and activations, we propose a novel hardware accelerator for CNNs exploiting zero weights and activations. We also report a zero-induced load imbalance problem, which exists in zero-aware parallel CNN hardware architectures, and present a zero-aware kernel allocation as a solution. According to our experiments with a cycle-accurate simulation model, RTL, and layout design of the proposed architecture running two real deep CNNs, pruned AlexNet [1] and VGG-16 [2], our architecture offers 4x/1.8x (AlexNet) and 5.2x/2.1x (VGG-16) speedup compared with state-of-the-art zero-agnostic/zero-activation-aware architectures.
Dongyoung Kim, Junwhan Ahn, Sungjoo Yoo
DATE1
2014 Coarse-grained Bubble Razor to exploit the potential of two-phase transparent latch designs
abstract
Timing margin to cover process variation is one of the most critical factors that limit the amount of supply voltage reduction thereby power consumption. To remove too conservative timing margin, Bubble Razor was introduced to dynamically detect and correct errors in two-phase transparent latch designs [13]. However, it does not fully exploit the potential of two-phase transparent latch design, e.g. time borrowing. Thus, especially at low supply voltage where the effect of process variation becomes significant, the existing Bubble Razor can suffer from significant overhead in performance and power consumption due to too frequent occurrence of bubble generations. We present a design methodology for coarse-grained Bubble Razor which exploits the time-borrowing characteristic of two-phase transparent latch design. By selectively inserting error checkpoints, i.e., shadow latches and error management logic, in the circuit, time borrowing can be applied between error checkpoints thereby avoiding bubbles which could occur in the existing Bubble Razor design with a checkpoint at every latch on the critical path. We present a methodology to choose the grain size (the number of stages between error checkpoints) based on 3-sigma delay distribution. We also verify the benefits of coarse-grained Bubble Razor with a real microprocessor, Core-A design [15] using 20nm Predictive Technology Model (PTM) [16]. The proposed methodology offers 62% improvement in performance (MIPS) and 49% less energy consumption (per instruction) at 0.6V operation (zero frequency margin) over the original Bubble Razor scheme. In addition, it gives 25% area reduction in core design.
Hayoung Kim, Dongyoung Kim, Jae-Joon Kim, Sungjoo Yoo, Sunggu Lee
DATE2
2013 MCMC particle filter-based vehicle tracking method using multiple hypotheses and appearance model
abstract
In this study, we propose a multiple vehicle tracking method using multiple hypotheses and the appearance model. The multiple hypotheses are associated with multiple tracks using track-to-multiple hypotheses association method. A target state is estimated using the maximum a posteriori probability estimation method. The posterior probability is proportional to the product of a priori probability and the likelihood that is calculated using similarities of multiple hypotheses and the appearance model. The posterior probability density function is estimated using the Markov chain Monte Carlo particle filter. An optimal posterior target state is determined using a sample with the maximum a posteriori probability. Our experimental results show that the proposed method can improve multiple objects tracking precision as well as multiple object tracking accuracy.
Young-Chul Lim, Dongyoung Kim, Chung-Hee Lee
Intelligent Vehicles Symposium2
2004 Fast broadcast by the divide-and-conquer algorithm
abstract
Collective communication functions including the broadcast in cluster computers usually take O(m log P) time in propagating the size-m message to P processors. We have devised a new O(m) broadcast algorithm, independent of the number of processors involved, by using divided-and-conquer algorithm. Details are given below.
Dongyoung Kim, Dongseung Kim
CLUSTER1