EDBT 2026 Demo / reviewers in the wild / expert
Joonho Song
dblp:59/8841
· DBLP profile ↗
9ranked-venue papers
0as first author
5since 2021 · last 2023
0009-0002-0718-6349ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Accelerating Deep Neural Networks on Mobile Multicore NPUsabstractNeural processing units (NPUs) have become indispensable parts of mobile SoCs. Furthermore, integrating multiple NPU cores into a single chip becomes a promising solution for ever-increasing computing power demands in mobile devices. This paper addresses techniques to maximize the utilization of NPU cores and reduce the latency of on-device inference. Mobile NPUs typically have a small amount of local memory (or scratch pad memory, SPM) that provides space only enough for input/output tensors and weights of one layer operation in deep neural networks (DNNs). Even in multicore NPUs, such local memories are distributed across the cores. In such systems, executing network layer operations in parallel is the primary vehicle to achieve performance. By partitioning a layer of DNNs into multiple sub-layers, we can execute them in parallel on multicore NPUs. Within a core, we can also employ pipelined execution to reduce the execution time of a sub-layer. In this execution model, synchronizing parallel execution and loading/storing intermediate tensors in global memory are the main bottlenecks. To alleviate these problems, we propose novel optimization techniques which carefully consider partitioning direction, execution order, synchronization, and global memory access. Using six popular convolutional neural networks (CNNs), we evaluate our optimization techniques in a flagship mobile SoC with three cores. Compared to the highest-performing partitioning approach, our techniques improve performance by 23%, achieving a speedup of 2.1x over single-core systems. Hanwoong Jung, Hexiang Ji, Alexey Pushchin, Maxim Ostapenko, Wenlong Niu, Ilya Palachev, Yutian Qu, Pavel Fedin, Yuri Gribov, Heewoo Nam, Dongguen Lim, Joonho Song, Hwansoo Han |
CGO | 13 |
| 2021 | A Novel Sensitivity Metric For Mixed-Precision Quantization With Synthetic Data GenerationabstractPost-training quantization is a representative technique for compressing neural networks, making them smaller and more efficient for deployment on edge devices. However, an inaccessible user dataset often makes it difficult to ensure the quality of the quantized neural network in practice. In addition, existing approaches may use a single uniform bit-width across the network, resulting in significant accuracy degradation at extremely low bit-widths. To utilize multiple bit-width, sensitivity metric plays a key role in balancing accuracy and compression. In this paper, we propose a novel sensitivity metric that considers the effect of quantization error on task loss and interaction with other layers. Moreover, we develop labeled data generation methods that are not dependent on a specific operation of the neural network. Our experiments show that the proposed metric better represents quantization sensitivity, and generated data are more feasible to apply to mixed-precision quantization. Minkyoung Cho, Joonho Song, Changkyu Choi |
ICIP | 4 |
| 2021 | An Automated Approach to Accelerate DNNs on Edge DevicesabstractDeployment of Deep Neural Networks (DNNs) on edge devices can significantly increase the utility of DNNs for a variety of applications. However executing DNN models on the edge device is still a major challenge, as heavy computation and memory bandwidth requirements of such models limit their adoption. Employing a highly optimized code for DNN model execution can easily enable many more use-cases than currently possible. However, current strategies are still based on manual optimization for efficient resource utilization. This is not only cumbersome but also requires a high level of expert intervention in the rapidly changing DNN Model landscape. In this work, we provide an automated way of optimizing Convolutional Neural Network (CNN) models using Deep Reinforcement Learning (DRL) algorithm. The experiments with our DRL technique demonstrate 1.85×, 1.58×, 1.64× speedup in execution time for MobileNetV1, MobileNetV2 and Efficientnet-lite0 CNN models respectively on Mobile CPU devices. Ujjawal Chugh, Arnab Mitra 0004, Ankur Deshwal, N. P. Swaroop, Aditi Saluja, Joonho Song |
ISCAS | 7 |
| 2021 | Transport Triggered near Memory Accelerator for Deep LearningabstractAs throughput of neural network accelerator datapaths have grown, memory has consistently fallen behind. Although attempts have been made to improve performance of neural networks via approaches such as batching, several layers often starve for memory bandwidth when used for tasks such as online inferencing. In order to mitigate memory bandwidth limitations, we propose a near memory accelerator for mobile devices based on Transport-Triggered Architecture (TTA) and evaluate its performance benefits compared to the existing approaches. Through experiments we demonstrate that our proposed accelerator achieves up to 4.3× speedup with respect to a non-near memory accelerator and the proposed TTA data-path achieves area and energy efficiency close to fixed-function datapaths while offering programmability similar to VLIW cores. Kavitha T. Madhu, Saptarsi Das, Abhishek Tyagi, Ankur Deshwal, Joonho Song |
ISCAS | 5 |
| 2021 | Autotuning LSTM for Accelerated Execution on EdgeabstractDeployment of Deep Neural Networks (DNNs) on edge devices is highly desirable to address user privacy concerns and minimize the turnaround time of AI applications. However, the execution of DNN models on a battery-operated device requires a highly optimized implementation specific to the target hardware. Moreover, as different layers of a DNN exhibit distinct computation and memory characteristics, it is imperative to optimize each layer separately. This is in contrast to the widely deployed library-based approach where all the configurations of DNN operations share the same implementation. In this paper, we address this issue by auto-tuning the implementation of Long Short Term Memory (LSTM) operations which are widely used in sequence based AI applications. To exhaustively search through the space of optimizations and its parameters, we develop a high-level autotuning framework based on Halide. We use grid search to find the parameters that lead to minimum runtime and further present TPE based search method to find the near-optimal runtime in a limited number of trials. We observe 2.2× -3.1× speedup in execution time for LSTM layers used in widely deployed GNMT and DeepSpeech2 models. Aditi Saluja, Arnab Mitra 0004, Ankur Deshwal, Kavitha T. Madhu, Ujjawal Chugh, Joonho Song |
ISCAS | 7 |
| 2020 | Sparse CNN Architecture Search (Scas)abstractAdvent of deep neural networks has revolutionized Computer Vision. However, designing of such models with high accuracy and low computation requirements is a difficult task and needs extensive human expertise. Recent advances in Neural Architecture Search use various methods like Deep Reinforcement Learning, Evolutionary methods, Gradient Descent, HyperNetworks etc. to automatically generate neural networks with high level of accuracy. However, large size of such generated models limit their practical use. Recent findings about lottery ticket hypothesis suggest the existence of sparse subnetworks (winning tickets) which can reach the accuracy comparable to that of original dense network. In this paper, we present a method for leveraging redundancies inherent to deep Convolutional Neural Networks (CNN) to guide the generation of sparse CNN models (to find the architectures with winning tickets) without significant loss in accuracy. We evaluate our proposed method with different NAS methods on CFAR-10,CIFAR-100 and MNIST datasets. Our results show a reduction ranging from $2\times \mathrm{to} 12\times$ in terms of model size and $2\times \mathrm{to}19\times$ in terms of number of MAC operations with less than 1% drop in accuracy. Yeshwanth V, Ankur Deshwal, Sundeep Krishnadasan, Joonho Song |
ICME | 5 |
| 2015 | DSP based programmable FHD HEVC decoder
Sangjo Lee, Joonho Song, Wonchang Lee, Doo Hyun Kim, Shihwa Lee |
DATE | 2 |
| 2011 | H.264/AVC UHD decoder implementation on multi-cluster platform using hybrid parallelization methodabstractWe propose a new hybrid parallelization method for H.264/AVC decoder for UHD (3840×2160) resolution on the proposed multi-clusters platform. We used 4 clusters to decode UHD video application, which cluster is composed of functional partitioned computing elements (1 DSP core and 3 hardware accelerators). The parallelizing efficiency is improved drastically by adopting frame-level parallelization between clusters. Due to the scalable architecture, we can control the number of cluster according to computing requirement. For example, we can use one cluster for FHD video application and four clusters for UHD application. The experimental results show that the parallelism is close to theoretical value by a 1.4~7.6% margin at 2~4 clusters and H.264/AVC UHD decoder runs under 650MHz. Sangjo Lee, Joonho Song, Do Hyung Kim 0002, Shihwa Lee |
ICIP | 2 |
| 2010 | H.264 decoder on embedded dual core with dynamically load-balanced functional paritioningabstractIn this paper, we address the problem of mapping H.264 main profile decoder on embedded dual core with dynamic load balancing. H.264 decoder is mapped to dual core system with a few hardware accelerators by proposed functional partitioning which enables simple interface with hardware accelerator and small memory usage for inter-core communication. We also propose dynamic load balancing method for the functional partitioning. The load balancing is done by mapping a few selected functions to each core dynamically at macroblock level. In this case, buffer level information is enough for making decision which core runs those functions. Because of this simple decision criterion and mechanism, performance loss for load balancing process can be negligible and it is also possible to extend the proposed load balancing method to multi-core systems easily. Experimental result shows that the proposed load balancing method reduces the waiting overhead dramatically and the reduced amount is 82.3% of the total waiting overhead. Joonho Song, Do Hyung Kim 0002, Shihwa Lee |
ICIP | 2 |