EDBT 2026 Demo / reviewers in the wild / expert
Jinho Lee 0001
dblp:62/757-1
· DBLP profile ↗
64ranked-venue papers
13as first author
41since 2021 · last 2026
0000-0003-4010-6611ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 10 first-author · 29 since 2021Artificial intelligence and machine learning · 11 · 9 since 2021Software engineering, systems software and programming languages · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 3Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMsabstractThe computational and memory demands of large language models for generative inference present significant challenges for practical deployment. One promising solution targeting offline inference is offloading-based batched inference, which extends the GPU's memory hierarchy with host memory and storage. However, it often suffers from substantial I/O overhead, primarily due to the large KV cache sizes that scale with batch size and context window length. Hongsun Jang, Jaeyong Song 0002, Changmin Shin 0002, Si Ung Noh, Jaewon Jung 0001, Jisung Park 0001, Jinho Lee 0001 |
ASPLOS (2) | 7 |
| 2026 | FlexiWalker: Extensible GPU Framework for Efficient Dynamic Random Walks with Runtime AdaptationabstractDynamic random walks are fundamental to various graph analysis applications, offering advantages by adapting to evolving graph properties. Their runtime-dependent transition probabilities break down the pre-computation strategy that underpins most existing CPU and GPU static random walk optimizations. This leaves practitioners suffering from suboptimal frameworks and having to write hand-tuned kernels that do not adapt to workload diversity. To handle this issue, we present FlexiWalker, the first GPU framework that delivers efficient, workload-generic support for dynamic random walks. Our design-space study shows that rejection sampling and reservoir sampling are more suitable than other sampling techniques under massive parallelism. Thus, we devise (i) new high-performance kernels for them that eliminate global reductions, redundant memory accesses, and random-number generation. Given the necessity of choosing the best-fitting sampling strategy at runtime, we adopt (ii) a lightweight first-order cost model that selects the faster kernel per node at runtime. To enhance usability, we introduce (iii) a compile-time component that automatically specializes user-supplied walk logic into optimized building blocks. On various dynamic random walk workloads with real-world graphs, FlexiWalker outperforms the best published CPU/GPU baselines by geometric means of 73.44× and 5.91×, respectively, while successfully executing workloads that prior systems cannot support. We open-source FlexiWalker in https://github.com/AIS-SNU/FlexiWalker. Seongyeon Park, Jaeyong Song 0002, Changmin Shin 0002, Sukjin Kim, Junguk Hong, Jinho Lee 0001 |
EuroSys | 6 |
| 2026 | LoCaLUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIMabstractLookup tables (LUTs) have recently gained attention as an alternative compute mechanism that maps input operands to precomputed results, eliminating the need for arithmetic logic. LUTs not only reduce logic complexity, but also naturally support diverse numerical precisions without requiring separate circuits for each bitwidth—an increasingly important feature in quantized DNNs. This creates a favorable tradeoff in PIM: memory capacity can be used in place of logic to increase computational throughput, aligning well with DRAM-PIM architectures that offer high bandwidth and easily available memory but limited logic density. In this work, we explore this capacity-computation tradeoff in LUT-based PIM designs, where memory capacity is traded for performance by packing multiple MAC operations into a single LUT lookup. Building on this insight, we propose LoCaLUT, a PIM-based design for efficient low-bit quantized DNN inference using operation-packed LUTs. First, we observe that these LUTs contain extensive redundancy and introduce LUT canonicalization, which eliminates duplicate entries to reduce LUT size. Second, we propose reordering LUT, a lightweight auxiliary LUT that remaps weight vectors to their canonical form required by LUT canonicalization with a simple LUT lookup. Third, we propose LUT slice streaming, a novel execution strategy that exploits the DRAM-buffer hierarchy by streaming only relevant LUT columns into the buffer and reusing them across multiple weight vectors. Evaluated on a real system based on UPMEM devices, we demonstrate a geometric mean speedup of$1.82 \times$across various numeric precisions and DNN models. We believe LoCaLUT opens a path toward scalable, low-logic PIM designs tailored for LUT-based DNN inference. Our implementation of LoCaLUT is available at https://github.com/AIS-SNU/LoCaLUT. Junguk Hong, Changmin Shin 0002, Sukjin Kim, Si Ung Noh, Taehee Kwon 0002, Seongyeon Park, Hanjun Kim 0001, Youngsok Kim, Jinho Lee 0001 |
HPCA | 9 |
| 2026 | DANCE++: Differentiable Accelerator/Network Co-Exploration With Hard Constraints and Data-Free Training for Real-World ScenariosabstractCo-exploration of neural architectures and hardware accelerators has emerged as a promising approach to address computational cost problems, especially in low-profile systems. However, existing co-exploration methods based on reinforcement learning or evolutionary search suffer from substantial search costs. To address this, this work presents DANCE++, a differentiable approach towards the co-exploration of hardware and network architecture design. At the heart of DANCE++ is a differentiable evaluator network that models hardware metrics with a neural network, enabling accelerator design through backpropagation. DANCE++ significantly reduces search time and enhances accuracy and hardware cost metrics compared to traditional approaches. To further address real-world scenarios, this work embodies two important practical topics: hard constraints and data dependency. To meet the constraints such as frame rates or area budget, this work proposes a gradient manipulation algorithm that guides differentiable optimization to find hard-constrained solutions. Also to consider cases where training dataset is inaccessible, this work proposes to use data-free training methods in both co-exploration and training phases. To the best of our knowledge, DANCE++ is the first co-exploration method that targets these real-world challenges, supported by extensive experiments demonstrating its effectiveness. Kanghyun Choi, Deokki Hong, Hyeyoon Lee, Joonsang Yu, Noseong Park, Youngsok Kim, Jinho Lee 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | DMO-DB: Mitigating the Data Movement Bottlenecks of GPU-Accelerated Relational OLAPabstractGraphics Processing Units (GPUs) offer high computational throughput and memory bandwidth, making them promising accelerators for relational OnLine Analytical Processing (OLAP). GPU-accelerated relational OLAP executes the relational operations of a Structured Query Language (SQL) query on a GPU instead of the host Central Processing Unit (CPU). Depending on where input columns and their values reside in, a GPU-accelerated SQL query execution can be classified into two scenarios: 1) in-GPU, in which all input columns fit in the GPU memory, or 2) in-host, in which the input columns reside in the host memory and get transferred to the GPU memory when needed. However, both scenarios incur significant intraGPU and host-to-GPU data movement overheads, respectively. In-GPU executions incur excessive GPU cache misses and thus frequent off-chip GPU memory accesses. In-host executions suffer from the limited host-to-GPU data transfer bandwidth. This paper presents DMO-DB, a Data Movement-Optimized GPU-accelerated relational OLAP engine. Since modern GPUaccelerated relational OLAP decomposes SQL queries into multiple pipelines-sequences of relational operations that can be executed on input columns from the same table, DMO-DB leverages inter-pipeline dependencies to overcome the two data movement bottlenecks. DMO-DB introduces two key ideas: cache-fit bloom filtering and Ahead-of-Time value Discarding (AoTD), which preemptively eliminate unnecessary input values before their movement across the memory hierarchies. For in-GPU execution, GPU L1 data cache-fit filters discard non-contributing values before triggering costly off-chip DRAM accesses. For in-host execution, host CPU last level cache-fit filters strategically prune unnecessary input values, minimizing PCIe transfer overhead. After that, AoTD exploits multiple inter-pipeline dependencies by collecting these cache-fit bloom filters to earlier pipeline execution stages. Our evaluation using NVIDIA RTX A4000 and TITAN RTX GPUs shows that DMO-DB achieves speedups of $\mathbf{1. 5 3 x}$ over in-GPU Crystal-Opt and 6.10x over in-host HeavyDB. Chaemin Lim, Suhyun Lee 0002, Jinwoo Choi 0003, Joonsung Kim 0001, Jinho Lee 0001, Youngsok Kim |
PACT | 5 |
| 2025 | MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention SimilarityabstractData-free quantization (DFQ) is a technique that creates a lightweight network from its full-precision counterpart without the original training data, often through a synthetic dataset. Although several DFQ methods have been proposed for vision transformer (ViT) architectures, they fail to achieve efficacy in low-bit settings. Examining the existing methods, we observe that their synthetic data produce misaligned attention maps, while those of the real samples are highly aligned. From this observation, we find that aligning attention maps of synthetic data helps improve the overall performance of quantized ViTs. Motivated by this finding, we devise MimiQ, a novel DFQ method designed for ViTs that enhances inter-head attention similarity. First, we generate synthetic data by aligning head-wise attention outputs from each spatial query patch. Then, we align the attention maps of the quantized network to those of the full-precision teacher by applying head-wise structural attention distillation. The experimental results show that the proposed method significantly outperforms baselines, setting a new state-of-the-art for ViT-DFQ. Kanghyun Choi, Hyeyoon Lee, Dain Kwon, Sunjong Park, Kyuyeun Kim, Noseong Park, Jinho Lee 0001 |
AAAI | 8 |
| 2025 | Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherabstractGraph processing requires irregular, fine-grained random access patterns incompatible with contemporary off-chip memory architecture, leading to inefficient data access. This inefficiency makes graph processing an extremely memory-bound application. Because of this, existing graph processing accelerators typically employ a graph tiling-based or processing-in-memory (PIM) approach to relieve the memory bottleneck. In the tiling-based approach, a graph is split into chunks that fit within the on-chip cache to maximize data reuse. In the PIM approach, arithmetic units are placed within memory to perform operations such as reduction or atomic addition. However, both approaches have several limitations, especially when implemented on current memory standards (i.e., DDR). Because the access granularity provided by DDR is much larger than that of the graph vertex property data, much of the bandwidth and cache capacity are wasted. PIM is meant to alleviate such issues, but it is difficult to use in conjunction with the tiling-based approach, resulting in a significant disadvantage. Furthermore, placing arithmetic units inside a memory chip is expensive, thereby supporting multiple types of operation is thought to be impractical. To address the above limitations, we present Piccolo, an end-to-end efficient graph processing accelerator with fine-grained in-memory random scatter-gather. Instead of placing expensive arithmetic units in off-chip memory, Piccolo focuses on reducing the off-chip traffic with non-arithmetic function-in-memory of random scatter-gather. To fully benefit from in-memory scatter-gather, Piccolo redesigns the cache and miss-handling architecture (MHA) of the accelerator such that it can enjoy both the advantage of tiling and in-memory operations. Piccolo achieves a maximum speedup of 3.28 × and a geometric mean speedup of 1.62 ×, along with up to 59.7% reduction in energy consumption across various and extensive benchmarks. Changmin Shin 0002, Jaeyong Song 0002, Hongsun Jang, Dogeun Kim, Jun Sung, Taehee Kwon 0002, Jae Hyung Ju, Frank Liu 0001, YeonKyu Choi, Jinho Lee 0001 |
HPCA | 10 |
| 2025 | G^3SA: A GPU-Accelerated Gold Standard Genomics Library for End-to-End Sequence AlignmentabstractSequence alignment is the first step of the bioinformatics pipelines for analyzing DNA sequences in genomics.As the DNA sequencer throughput is dramatically increasing, the pipeline bottleneck has now shifted to the massive amount of sequence alignment computation.GPU-based acceleration has emerged as a promising solution to this challenge, and several efforts have focused on accelerating core algorithmic components.However, when integrating existing GPUaccelerated components into BWA-MEM and Minimap2, two widely used tools for short and long reads respectively, we identify several unresolved issues.One could naively attempt to gather and/or fix those libraries, but such an approach would result in a severe slowdown despite non-trivial effort.To address these issues, we propose 𝐺 3 𝑆𝐴, the first GPU acceleration with efficient parallelization and optimization strategies for the end-to-end gold standard sequence alignment algorithms for short and long reads.We observe that the primary bottlenecks stem from irregular and redundant computation and memory access patterns.Therefore, we design our kernels to fully exploit the memory hierarchy and computational capabilities of GPUs, reducing costly computation and memory access.Furthermore, following the seed-chain-extend flow of standard sequence alignment algorithms, we propose techniques tailored to specific stages in the pipeline.Using these techniques, 𝐺 3 𝑆𝐴 achieves significant end-to-end speedup against the CPU and GPU baselines. Yeejoo Han, Seongyeon Park, Jinho Lee 0001 |
ICS | 4 |
| 2025 | CrossBit: Bitwise Computing in NAND Flash Memory with Inter-Bitline Data CommunicationabstractIn-flash processing (IFP), which involves performing data computation inside NAND flash memory, holds high potential for improving the performance and energy efficiency of data-intensive application by minimizing data movement.Recent research has introduced several IFP architectures enabling bulk bitwise operations inside NAND flash chips to demonstrate this potential.However, previous IFP designs were limited to performing bitwise operations on data sensed within the same bitline, thus lacking the capability to handle more complex functions requiring interactions between data from different bitlines.This paper presents CrossBit, a new IFP architecture that enables both intra-bitline and inter-bitline operations with minimal additional circuitry integrated into commodity NAND flash memory.With the capability for inter-bitline operations, CrossBit facilitates in-flash error correction code (IF-ECC) operations, thereby enabling reliable multi-level cell (MLC) IFP.Moreover, CrossBit efficiently processes fundamental database queries that were previously inefficient with existing work that only supports intra-bitline operations.Experimental results show that the implementation of IF-ECC in CrossBit results in a substantial reduction in bit error rate (BER) for MLC operations, leading to 1.8× increase in bit-density by using MLC compared to previous IFP designs which uses SLC only.When used for accelerating fundamental database queries, CrossBit achieves notable average speedup and energy efficiency improvement of 2.1× and 2.5× compared to the state-of-the-art (SOTA) IFP architecture.We further demonstrate the practicality by processing the full set of end-toend database queries from the widely used Star schema benchmark, where CrossBit achieves 1.7× speedup. Seunghwan Song, Sukhyun Choi, Jeongin Choe, Sanghyeok Han, Jisung Park 0001, Jinho Lee 0001, Jae-Joon Kim |
MICRO | 7 |
| 2025 | FALA: Locality-Aware PIM-Host Cooperation for Graph Processing with Fine-Grained Column Access
Changmin Shin 0002, Jaeyong Song 0002, Seongmin Na, Jun Sung, Hongsun Jang, Jinho Lee 0001 |
MICRO | 6 |
| 2025 | FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point ArithmeticabstractLow-bit floating-point (FP) formats, such as FP8, provide significant acceleration and memory savings in model training thanks to native hardware support on modern GPUs and NPUs. However, we analyze that FP8 quantization offers speedup primarily for large-dimensional matrix multiplications, while inherent quantization overheads diminish speedup when applied to low-rank adaptation (LoRA), which uses small-dimensional matrices for efficient fine-tuning of large language models (LLMs). To address this limitation, we propose FALQON, a novel framework that eliminates the quantization overhead from separate LoRA computational paths by directly merging LoRA adapters into an FP8-quantized backbone during fine-tuning. Furthermore, we reformulate the forward and backward computations for merged adapters to significantly reduce quantization overhead, and introduce a row-wise proxy update mechanism that efficiently integrates substantial updates into the quantized backbone. Experimental evaluations demonstrate that FALQON achieves approximately a 3$\times$ training speedup over existing quantized LoRA methods with a similar level of accuracy, providing a practical solution for efficient large-scale model fine-tuning. Moreover, FALQON’s end-to-end FP8 workflow removes the need for post-training quantization, facilitating efficient deployment. Code is available at https://github.com/iamkanghyunchoi/falqon. Kanghyun Choi, Hyeyoon Lee, Sunjong Park, Dain Kwon, Jinho Lee 0001 |
NeurIPS | 5 |
| 2025 | PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
Sukjin Kim, Seongyeon Park, Si Ung Noh, Junguk Hong, Taehee Kwon 0002, Hunseong Lim, Jinho Lee 0001 |
USENIX ATC | 7 |
| 2025 | AGIS: Fast Approximate Graph Pattern Mining with Structure-Informed Sampling
Seoyong Lee, Jinho Lee 0001 |
Proc. VLDB Endow. | 2 |
| 2024 | GraNNDis: Fast Distributed Graph Neural Network Training Framework for Multi-Server ClustersabstractGraph neural networks (GNNs) are one of the rapidly growing fields within deep learning. While many distributed GNN training frameworks have been proposed to increase the training throughput, they face three limitations when applied to multi-server clusters. 1) They suffer from an inter-server communication bottleneck because they do not consider the inter-/intra-server bandwidth gap, a representative characteristic of multi-server clusters. 2) Redundant memory usage and computation hinder the scalability of the distributed frameworks. 3) Sampling methods, de facto standard in mini-batch training, incur unnecessary errors in multi-server clusters. Jaeyong Song 0002, Hongsun Jang, Hunseong Lim, Jaewon Jung 0001, Youngsok Kim, Jinho Lee 0001 |
PACT | 6 |
| 2024 | PeerAiD: Improving Adversarial Distillation from a Specialized Peer TutorabstractAdversarial robustness of the neural network is a significant concern when it is applied to security-critical domains. In this situation, adversarial distillation is a promising option which aims to distill the robustness of the teacher network to improve the robustness of a small student network. Previous works pretrain the teacher network to make it robust against the adversarial examples aimed at itself. However, the adversarial examples are dependent on the parameters of the target network. The fixed teacher network inevitably degrades its robustness against the unseen transferred adversarial examples which target the parameters of the student network in the adversarial distillation process. We propose PeerAiD to make a peer network learn the adversarial examples of the student network instead of adversarial examples aimed at itself. PeerAiD is an adversarial distillation that trains the peer network and the student network simultaneously in order to specialize the peer network for defending the student network. We observe that such peer networks surpass the robustness of the pretrained robust teacher model against adversarial examples aimed at the student network. With this peer network and adversarial distillation, PeerAiD achieves significantly higher robustness of the student network with AutoAttack (AA) accuracy by up to 1.66%p and improves the natural accuracy of the student network by up to 4.72% p with ResNet-18 on TinyImageNet dataset. Code is available at https://github.com/jaewonalive/PeerAiD. Jaewon Jung 0001, Hongsun Jang, Jaeyong Song 0002, Jinho Lee 0001 |
CVPR | 4 |
| 2024 | Pipette: Automatic Fine-Grained Large Language Model Training Configurator for Real-World ClustersabstractTraining large language models (LLMs) is known to be challenging because of the huge computational and memory capacity requirements. To address these issues, it is common to use a cluster of GPUs with 3D parallelism, which splits a model along the data batch, pipeline stage, and intra-layer tensor dimensions. However, the use of 3D parallelism produces the additional challenge of finding the optimal number of ways on each dimension and mapping the split models onto the GPUs. Several previous studies have attempted to automatically find the optimal configuration, but many of these lacked several important aspects. For instance, the heterogeneous nature of the interconnect speeds is often ignored. While the peak bandwidths for the interconnects are usually made equal, the actual attained bandwidth varies per link in real-world clusters. Combined with the critical path modeling that does not properly consider the communication, they easily fall into sub-optimal configurations. In addition, they often fail to consider the memory requirement per GPU, often recommending solutions that could not be executed. To address these challenges, we propose Pipette, which is an automatic fine-grained LLM training configurator for real-world clusters. By devising better performance models along with the memory estimator and fine-grained individual GPU assignment, Pipette achieves faster configurations that satisfy the memory constraints. We evaluated Pipette on large clusters to show that it provides a significant speedup over the prior art. Jinkyu Yim, Jaeyong Song 0002, Yerim Choi, Jaebeen Lee, Jaewon Jung 0001, Hongsun Jang, Jinho Lee 0001 |
DATE | 7 |
| 2024 | Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real SystemabstractThe recent huge advance of Large Language Models (LLMs) is mainly driven by the increase in the number of parameters. This has led to substantial memory capacity requirements, necessitating the use of dozens of GPUs just to meet the capacity. One popular solution to this is storage-offloaded training, which uses host memory and storage as an extended memory hierarchy. However, this obviously comes at the cost of storage bandwidth bottleneck because storage devices have orders of magnitude lower bandwidth compared to that of GPU device memories. Our work, Smart-Infinity, addresses the storage bandwidth bottleneck of storage-offloaded LLM training using near-storage processing devices on a real system. The main component of Smart-Infinity is SmartUpdate, which performs parameter updates on custom near-storage accelerators. We identify that moving parameter updates to the storage side removes most of the storage traffic. In addition, we propose an efficient data transfer handler structure to address the system integration issues for Smart-Infinity. The handler allows overlapping data transfers with fixed memory consumption by reusing the device buffer. Lastly, we propose accelerator-assisted gradient compression/decompression to enhance the scalability of Smart-Infinity. When scaling to multiple near-storage processing devices, the write traffic on the shared channel becomes the bottleneck. To alleviate this, we compress the gradients on the GPU and decompress them on the accelerators. It provides further acceleration from reduced traffic. As a result, Smart-Infinity achieves a significant speedup compared to the baseline. Notably, SmartInfinity is a ready-to-use approach that is fully integrated into PyTorch on a real system. The implementation of Smart-Infinity is available at https://github.com/AIS-SNU/smart-infinity. Hongsun Jang, Jaeyong Song 0002, Jaewon Jung 0001, Youngsok Kim, Jinho Lee 0001 |
HPCA | 6 |
| 2024 | DataFreeShield: Defending Adversarial Attacks without Training DataabstractRecent advances in adversarial robustness rely on an abundant set of training data, where using external or additional datasets has become a common setting. However, in real life, the training data is often kept private for security and privacy issues, while only the pretrained weight is available to the public. In such scenarios, existing methods that assume accessibility to the original data become inapplicable. Thus we investigate the pivotal problem of data-free adversarial robustness, where we try to achieve adversarial robustness without accessing any real data. Through a preliminary study, we highlight the severity of the problem by showing that robustness without the original dataset is difficult to achieve, even with similar domain datasets. To address this issue, we propose DataFreeShield, which tackles the problem from two perspectives: surrogate dataset generation and adversarial training using the generated data. Through extensive validation, we show that DataFreeShield outperforms baselines, demonstrating that the proposed method sets the first entirely data-free solution for the adversarial robustness problem. Hyeyoon Lee, Kanghyun Choi, Dain Kwon, Sunjong Park, Mayoore Selvarasa Jaiswal, Noseong Park, Jinho Lee 0001 |
ICML | 8 |
| 2024 | PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM DevicesabstractRecent dual in-line memory modules (DIMMs) are starting to support processing-in-memory (PIM) by associating their memory banks with processing elements (PEs), allowing applications to overcome the data movement bottleneck by offloading memory-intensive operations to the PEs. Many highly parallel applications have been shown to benefit from these PIM-enabled DIMMs, but further speedup is often limited by the huge overhead of inter-PE collective communication. This mainly comes from the slow CPU-mediated inter-PE communication methods, making it difficult for PIM-enabled DIMMs to accelerate a wider range of applications. Prior studies have tried to alleviate the communication bottleneck, but they lack enough flexibility and performance to be used for a wide range of applications. In this paper, we present PID-Comm, a fast and flexible inter-PE collective communication framework for commodity PIM-enabled DIMMs. The key idea of PID-Comm is to abstract the PEs as a multi-dimensional hypercube and allow multiple instances of inter-PE collective communication between the PEs belonging to certain dimensions of the hypercube. Leveraging this abstraction, PID-Comm first defines eight interPE collective communication patterns that allow applications to easily express their complex communication patterns. Then, PIDComm provides high-performance implementations of the interPE collective communication patterns optimized for the DIMMs. Our evaluation using 16 UPMEM DIMMs and representative parallel algorithms shows that PID-Comm greatly improves the performance by up to $5.19 \times$ compared to the existing inter-PE communication implementations. The implementation of PIDComm is available at https://github.com/AIS-SNU/PID-Comm. Si Ung Noh, Junguk Hong, Chaemin Lim, Seongyeon Park, Jeehyun Kim, Hanjun Kim 0001, Youngsok Kim, Jinho Lee 0001 |
ISCA | 8 |
| 2024 | AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read MappingabstractWith the advance in genome sequencing technology, the lengths of deoxyribonucleic acid (DNA) sequencing results are rapidly increasing at lower prices than ever. However, the longer lengths come at the cost of a heavy computational burden on aligning them. For example, aligning sequences to a human reference genome can take tens or even hundreds of hours. The current de facto standard approach for alignment is based on the guided dynamic programming method. Although this takes a long time and could potentially benefit from high-throughput graphic processing units (GPUs), the existing GPU-accelerated approaches often compromise the algorithm's structure, due to the GPU-unfriendly nature of the computational pattern. Unfortunately, such compromise in the algorithm is not tolerable in the field, because sequence alignment is a part of complicated bioinformatics analysis pipelines. In such circumstances, we propose AGAThA, an exact and efficient GPU-based acceleration of guided sequence alignment. We diagnose and address the problems of the algorithm being unfriendly to GPUs, which comprises strided/redundant memory accesses and workload imbalances that are difficult to predict. According to the experiments on modern GPUs, AGAThA achieves 18.8× speedup against the CPU-based baseline, 9.6× against the best GPU-based baseline, and 3.6× against GPU-based algorithms with different heuristics. Seongyeon Park, Junguk Hong, Jaeyong Song 0002, Hajin Kim, Youngsok Kim, Jinho Lee 0001 |
PPoPP | 6 |
| 2023 | Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication CompressionabstractIn training of modern large natural language processing (NLP) models, it has become a common practice to split models using 3D parallelism to multiple GPUs. Such technique, however, suffers from a high overhead of inter-node communication. Compressing the communication is one way to mitigate the overhead by reducing the inter-node traffic volume; however, the existing compression techniques have critical limitations to be applied for NLP models with 3D parallelism in that 1) only the data parallelism traffic is targeted, and 2) the existing compression schemes already harm the model quality too much. Jaeyong Song 0002, Jinkyu Yim, Jaewon Jung 0001, Hongsun Jang, Youngsok Kim, Jinho Lee 0001 |
ASPLOS (2) | 7 |
| 2023 | Fast Adversarial Training with Dynamic Batch-level Attack ControlabstractDespite the fact that adversarial training provides an effective protection against adversarial attacks, it suffers from a huge computational overhead. To mitigate the overhead, we propose DBAC, a fast adversarial training with dynamic batch-level attack control. Based on a prior study where attack strength should gradually grow throughout the training, we control the number of samples attacked per batch for better throughput. Additionally, we collect samples from multiple batches to form a pseudo-batch and attack them simultaneously for higher GPU utilization. We implement DBAC using PyTorch to show its superior throughput with similar robust accuracy compared to the prior art. Jaewon Jung 0001, Jaeyong Song 0002, Hongsun Jang, Hyeyoon Lee, Kanghyun Choi, Noseong Park, Jinho Lee 0001 |
DAC | 7 |
| 2023 | Pipe-BD: Pipelined Parallel Blockwise DistillationabstractTraining large deep neural network models is highly challenging due to their tremendous computational and mem-ory requirements. Blockwise distillation provides one promising method towards faster convergence by splitting a large model into multiple smaller models. In state-of-the-art blockwise distillation methods, training is performed block-by-block in a data-parallel manner using multiple GPUs. To produce inputs for the student blocks, the teacher model is executed from the beginning until the current block under training. However, this results in a high overhead of redundant teacher execution, low GPU utilization, and extra data loading. To address these problems, we propose Pipe-BD, a novel parallelization method for blockwise distillation. Pipe-BD aggressively utilizes pipeline parallelism for blockwise distillation, eliminating redundant teacher block execution and increasing per-device batch size for better resource utilization. We also extend to hybrid parallelism for efficient workload balancing. As a result, Pipe-BD achieves significant acceleration without modifying the mathematical formulation of blockwise distillation. We implement Pipe-BD on PyTorch, and experiments reveal that Pipe-BD is effective on multiple scenarios, models, and datasets. Hongsun Jang, Jaewon Jung 0001, Jaeyong Song 0002, Joonsang Yu, Youngsok Kim, Jinho Lee 0001 |
DATE | 6 |
| 2023 | SGCN: Exploiting Compressed-Sparse Features in Deep Graph Convolutional Network AcceleratorsabstractGraph convolutional networks (GCNs) are becoming increasingly popular as they overcome the limited applicability of prior neural networks. One recent trend in GCNs is the use of deep network architectures. As opposed to the traditional GCNs, which only span only around two to five layers deep, modern GCNs now incorporate tens to hundreds of layers with the help of residual connections. From such deep GCNs, we find an important characteristic that they exhibit very high intermediate feature sparsity. This reveals a new opportunity for accelerators to exploit in GCN executions that was previously not present.In this paper, we propose SGCN, a fast and energy-efficient GCN accelerator which fully exploits the sparse intermediate features of modern GCNs. SGCN suggests several techniques to achieve significantly higher performance and energy efficiency than the existing accelerators. First, SGCN employs a GCN-friendly feature compression format. We focus on reducing the off-chip memory traffic, which often is the bottleneck for GCN executions. Second, we propose microarchitectures for seamlessly handling the compressed feature format. Specifically, we modify the aggregation phase of GCN to process compressed features, and design a combination engine that can output compressed features at no extra memory traffic cost. Third, to better handle locality in the existence of the varying sparsity, SGCN employs sparsity-aware cooperation. Sparsity-aware cooperation creates a pattern that exhibits multiple reuse windows, such that the cache can capture diverse sizes of working sets and therefore adapt to the varying level of sparsity. Through a thorough evaluation, we show that SGCN achieves 1.66× speedup and 44.1% higher energy efficiency compared to the existing accelerators in geometric mean. Mingi Yoo, Jaeyong Song 0002, Jounghoo Lee, Namhyung Kim, Youngsok Kim, Jinho Lee 0001 |
HPCA | 6 |
| 2023 | Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMsabstractModern dual in-line memory modules (DIMMs) support processing-in-memory (PIM) by implementing in-DIMM processors (IDPs) located near memory banks. PIM can greatly accelerate in-memory join, whose performance is frequently bounded by main-memory accesses, by offloading the operations of join from host central processing units (CPUs) to the IDPs. As real PIM hardware has not been available until very recently, the prior PIM-assisted join algorithms have relied on PIM hardware simulators which assume fast shared memory between the IDPs and fast inter-IDP communication; however, on commodity PIM-enabled DIMMs, the IDPs do not share memory and demand the CPUs to mediate inter-IDP communication. Such discrepancies in the architectural characteristics make the prior studies incompatible with the DIMMs. Thus, to exploit the high potential of PIM on commodity PIM-enabled DIMMs, we need a new join algorithm designed and optimized for the DIMMs and their architectural characteristics. In this paper, we design and analyze Processing-In-DIMM Join (PID-Join), a fast in-memory join algorithm which exploits UPMEM DIMMs, currently the only publicly-available PIM-enabled DIMMs. The DIMMs impose several key challenges on efficient acceleration of join including the shared-nothing nature and limited compute capabilities of the IDPs, the lack of hardware support for fast inter-IDP communication, and the slow IDP-wise data transfers between the IDPs and the main memory. PID-Join overcomes the challenges by prototyping and evaluating hash, sort-merge, and nested-loop algorithms optimized for the IDPs, enabling fast inter-IDP communication using host CPU cache streaming and vector instructions, and facilitating fast rank-wise data transfers between the IDPs and the main memory. Our evaluation using a real system equipped with eight UPMEM DIMMs and 1,024 IDPs shows that PID-Join greatly improves the performance of in-memory join over various CPU-based in-memory join algorithms. Chaemin Lim, Suhyun Lee 0002, Jinwoo Choi 0003, Jounghoo Lee, Seongyeon Park, Hanjun Kim 0001, Jinho Lee 0001, Youngsok Kim |
Proc. ACM Manag. Data | 7 |
| 2023 | Enabling Fine-Grained Spatial Multitasking on Systolic-Array NPUs Using Dataflow MirroringabstractNeural Processing Units (NPUs) frequently suffer from low hardware utilization as the efficiency of their systolic arrays heavily depends on the characteristics of a deep neural network (DNN). Spatial multitasking is a promising solution to overcome the low NPU hardware utilization; however, the state-of-the-art spatial-multitasking NPU architecture achieves sub-optimal performance due to its coarse-grained systolic-array distribution and incurs significant implementation costs. In this paper, we proposedataflow-mirroring NPU (DM-NPU), a novel spatial-multitasking NPU architecture supporting fine-grained systolic-array distribution. The key idea of DM-NPU is to reverse the dataflows of co-located DNNs in horizontal and/or vertical directions. DM-NPU can place allocation boundaries between any adjacent processing elements of a systolic array, both horizontally and vertically. We then proposeDM-Perf, an accurate systolic-array NPU performance model, to maximize the spatial-multitasking performance of DM-NPU. Utilizing the existing performance models achieves sub-optimal performance as they cannot accurately capture the resource contention caused by spatial multitasking. DM-Perf, on the other hand, exploits the per-layer performance profiles of a DNN to accurately capture the resource contention. Our evaluation using MLPerf DNNs shows that DM-NPU and DM-Perf can greatly improve the performance by up to 35.1% over the state-of-the-art NPU architecture and performance model. Jinwoo Choi 0003, Yeonan Ha, Jounghoo Lee, Jinho Lee 0001, Hanhwi Jang, Youngsok Kim |
IEEE Trans. Computers | 5 |
| 2022 | Decoupling Schedule, Topology Layout, and Algorithm to Easily Enlarge the Tuning Space of GPU Graph ProcessingabstractOnly with a right schedule and a right topology layout, a graph algorithm can be efficiently processed on GPUs. Existing GPU graph processing frameworks try to find an optimal schedule and topology layout for an algorithm via iterative search, but they fail to find the optimal configuration because their schedules and topology layouts are tightly coupled in their processing models. Moreover, their tightly coupled schedules and topology layouts make it difficult for developers to extend the tuning space. To easily enlarge the tuning space of GPU graph processing, this work proposes a new GPU graph processing abstraction scheme that fully decouples schedules, topology layouts, and algorithms from each other with abstraction interfaces. Moreover, this work proposes GRAssembler, a new GPU graph processing framework that efficiently integrates the decoupled schedule, topology layout, and algorithm without abstraction overhead. Thanks to the efficient decoupling and integration, GRAssembler increases the tuning space from 336 to 4,480 and achieves 30.4% higher performance on geomean average, compared to the state-of-the-art GPU graph processing framework. Shinnung Jeong, Yongwoo Lee 0001, Jaeho Lee 0005, Heelim Choi, Seungbin Song, Jinho Lee 0001, Youngsok Kim, Hanjun Kim 0001 |
PACT | 6 |
| 2022 | Slice-and-Forge: Making Better Use of Caches for Graph Convolutional Network AcceleratorsabstractGraph convolutional networks (GCNs) are becoming increasingly popular as they can process a wide variety of data formats that prior deep neural networks cannot easily support. One key challenge in designing hardware accelerators for GCNs is the vast size and randomness in their data access patterns which greatly reduces the effectiveness of the limited on-chip cache. Aimed at improving the effectiveness of the cache by mitigating the irregular data accesses, prior studies often employ the vertex tiling techniques used in traditional graph processing applications. While being effective at enhancing the cache efficiency, those approaches are often sensitive to the tiling configurations where the optimal setting heavily depends on target input datasets. Furthermore, the existing solutions require manual tuning through trial-and-error or rely on sub-optimal analytical models. Mingi Yoo, Jaeyong Song 0002, Hyeyoon Lee, Jounghoo Lee, Namhyung Kim, Youngsok Kim, Jinho Lee 0001 |
PACT | 7 |
| 2022 | Improving Gradient Paths for Binary Convolutional Neural Networks
Baozhou Zhu, H. Peter Hofstee, Jinho Lee 0001, Zaid Al-Ars |
BMVC | 3 |
| 2022 | It's All In the Teacher: Zero-Shot Quantization Brought Closer to the TeacherabstractModel quantization is considered as a promising method to greatly reduce the resource requirements of deep neural networks. To deal with the performance drop induced by quantization errors, a popular method is to use training data to fine-tune quantized networks. In real-world environments, however, such a method is frequently infeasible because training data is unavailable due to security, privacy, or confidentiality concerns. Zero-shot quantization addresses such problems, usually by taking information from the weights of a full-precision teacher network to compensate the performance drop of the quantized networks. In this paper, we first analyze the loss surface of state-of-the-art zero-shot quantization techniques and provide several findings. In contrast to usual knowledge distillation problems, zero-shot quantization often suffers from 1) the difficulty of optimizing multiple loss terms together, and 2) the poor generalization capability due to the use of synthetic samples. Furthermore, we observe that many weights fail to cross the rounding threshold during training the quantized networks even when it is necessary to do so for better performance. Based on the observations, we propose AIT, a simple yet powerful technique for zero-shot quantization, which addresses the aforementioned two problems in the following way: AIT i) uses a KL distance loss only without a cross-entropy loss, and ii) manipulates gradients to guarantee that a certain portion of weights are properly updated after crossing the rounding thresholds. Experiments show that AIT outperforms the performance of many existing methods by a great margin, taking over the overall state-of-the-art position in the field. Kanghyun Choi, Hyeyoon Lee, Deokki Hong, Joonsang Yu, Noseong Park, Youngsok Kim, Jinho Lee 0001 |
CVPR | 7 |
| 2022 | Enabling hard constraints in differentiable neural network and accelerator co-explorationabstractCo-exploration of an optimal neural architecture and its hardware accelerator is an approach of rising interest which addresses the computational cost problem, especially in low-profile systems. The large co-exploration space is often handled by adopting the idea of differentiable neural architecture search. However, despite the superior search efficiency of the differentiable co-exploration, it faces a critical challenge of not being able to systematically satisfy hard constraints such as frame rate. To handle the hard constraint problem of differentiable co-exploration, we propose HDX, which searches for hard-constrained solutions without compromising the global design objectives. By manipulating the gradients in the interest of the given hard constraint, high-quality solutions satisfying the constraint can be obtained. Deokki Hong, Kanghyun Choi, Hyeyoon Lee, Joonsang Yu, Noseong Park, Youngsok Kim, Jinho Lee 0001 |
DAC | 7 |
| 2022 | SALoBa: Maximizing Data Locality and Workload Balance for Fast Sequence Alignment on GPUsabstractSequence alignment forms an important backbone in many sequencing applications. A commonly used strategy for sequence alignment is an approximate string matching with a two-dimensional dynamic programming approach. Although some prior work has been conducted on GPU acceleration of a sequence alignment, we identify several shortcomings that limit exploiting the full computational capability of modern GPUs. This paper presents SALoBa, a GPU-accelerated sequence alignment library focused on seed extension. Based on the analysis of previous work with real-world sequencing data, we propose techniques to exploit the data locality and improve work-load balancing. The experimental results reveal that SALoBa significantly improves the seed extension kernel compared to state-of-the-art GPU-based methods. Seongyeon Park, Hajin Kim, Nauman Ahmed, Zaid Al-Ars, H. Peter Hofstee, Youngsok Kim, Jinho Lee 0001 |
IPDPS | 8 |
| 2022 | GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUsabstractAnalytical models can greatly help computer architects perform orders of magnitude faster early-stage design space exploration than using cycle-level simulators. To facilitate rapid design space exploration for graphics processing units (GPUs), prior studies have proposed GPU analytical models which capture first-order stall events causing performance degradation; however, the existing analytical models cannot accurately model modern GPUs due to their outdated and highly abstract GPU core microarchitecture assumptions. Therefore, to accurately evaluate the performance of modern GPUs, we need a new GPU analytical model which accurately captures the stall events incurred by the significant changes in the core microarchitectures of modern GPUs. Jounghoo Lee, Yeonan Ha, Suhyun Lee 0002, Jinyoung Woo, Jinho Lee 0001, Hanhwi Jang, Youngsok Kim |
ISCA | 5 |
| 2022 | GuardiaNN: Fast and Secure On-Device Inference in TrustZone Using Embedded SRAM and Cryptographic HardwareabstractAs more and more mobile/embedded applications employ Deep Neural Networks (DNNs) involving sensitive user data, mobile/embedded devices must provide a highly secure DNN execution environment to prevent privacy leaks. Aimed at securing DNN data, recent studies execute part of a DNN in a trusted execution environment (e.g., TrustZone) to isolate DNN execution from the other processes; however, as the trusted execution environments for mobile/embedded devices provide limited memory protection, DNN data remain unencrypted in DRAM and become vulnerable to physical attacks. The devices can prevent the physical attacks by keeping DNN data encrypted in DRAM; when DNN data get referenced during DNN execution, they get loaded to the SRAM and get decrypted by a CPU core. Unfortunately, using the SRAM with demand paging greatly increases DNN execution time due to the inefficient use of the SRAM and the high CPU consumption of data encryption/decryption. Jinwoo Choi 0003, Jaeyeon Kim, Chaemin Lim, Suhyun Lee 0002, Jinho Lee 0001, Dokyung Song, Youngsok Kim |
Middleware | 5 |
| 2022 | ComPreEND: Computation Pruning through Predictive Early Negative Detection for ReLU in a Deep Neural Network AcceleratorabstractA vast amount of activation values of DNNs are zeros due to ReLU (Rectified Linear Unit), which is one of the most common activation functions used in modern neural networks. Since ReLU outputs zero for all negative inputs, the inputs to ReLU do not need to be determined exactly as long as they are negative. However, many accelerators usually do not consider such aspects of DNNs, losing a huge amount of opportunities for speedups and energy savings. To exploit such opportunities, we propose early negative detection (END), a computation pruning technique that detects the negative results at an early stage. The key to the early negative detection is the adoption of inverted two's complement representation for filter parameters. This ensures that as soon as the intermediate results become negative, the final results are guaranteed to be negative. Upon detection, the remaining computation can be skipped and the following ReLU output can be simply set to zero. We also propose a DNN accelerator architecture (ComPreEND) that takes advantage of such skipping. ComPreEND with END significantly improves both the energy efficiency and the performance according to the evaluation. Compared to the baseline, we obtain 20.5 and 29.3 percent speedup with accurate mode and predictive mode, and energy savings by 28.4 and 41.4 percent, respectively. Namhyung Kim, Hanmin Park, Sungbum Kang, Jinho Lee 0001, Kiyoung Choi |
IEEE Trans. Computers | 5 |
| 2021 | DANCE: Differentiable Accelerator/Network Co-ExplorationabstractThis work presents DANCE, a differentiable approach towards the co-exploration of hardware accelerator and network architecture design. At the heart of DANCE is a differentiable evaluator network. By modeling the hardware evaluation software with a neural network, the relation between the accelerator design and the hardware metrics becomes differentiable, allowing the search to be performed with backpropagation. Compared to the naive existing approaches, our method performs co-exploration in a significantly shorter time, while achieving superior accuracy and hardware cost metrics. Kanghyun Choi, Deokki Hong, Hojae Yoon, Joonsang Yu, Youngsok Kim, Jinho Lee 0001 |
DAC | 6 |
| 2021 | Dataflow Mirroring: Architectural Support for Highly Efficient Fine-Grained Spatial Multitasking on Systolic-Array NPUsabstractWe present dataflow mirroring, architectural support for low-overhead fine-grained systolic array allocation which overcomes the limitations of prior coarse-grained spatial-multitasking Neural Processing Unit (NPU) architectures. The key idea of dataflow mirroring is to reverse the dataflows of co-located Neural Networks (NNs) in horizontal and/or vertical directions, allowing allocation boundaries to be set between any adjacent rows and columns of a systolic array and supporting up to four-way spatial multitasking. Our detailed experiments using MLPerf NNs and a dataflow-mirroring-augmented NPU prototype which extends Google’s TPU with dataflow mirroring shows that dataflow mirroring can significantly improve the multitasking performance by up to 46.4%. Jounghoo Lee, Jinwoo Choi 0003, Jaeyeon Kim, Jinho Lee 0001, Youngsok Kim |
DAC | 4 |
| 2021 | GradPIM: A Practical Processing-in-DRAM Architecture for Gradient DescentabstractIn this paper, we present GradPIM, a processingin-memory architecture which accelerates parameter updates of deep neural networks training. As one of processing-in-memory techniques that could be realized in the near future, we propose an incremental, simple architectural design that does not invade the existing memory protocol. Extending DDR4 SDRAM to utilize bank-group parallelism makes our operation designs in processing-in-memory (PIM) module efficient in terms of hardware cost and performance. Our experimental results show that the proposed architecture can improve the performance of DNN training and greatly reduce memory bandwidth requirement while posing only a minimal amount of overhead to the protocol and DRAM area. Heesu Kim, Hanmin Park, Kwanheum Cho, Eojin Lee, Soojung Ryu, Kiyoung Choi, Jinho Lee 0001 |
HPCA | 9 |
| 2021 | An Attention Module for Convolutional Neural Networks
Baozhou Zhu, H. Peter Hofstee, Jinho Lee 0001, Zaid Al-Ars |
ICANN (1) | 3 |
| 2021 | AutoReCon: Neural Architecture Search-based Reconstruction for Data-free CompressionabstractData-free compression raises a new challenge because the original training dataset for a pre-trained model to be compressed is not available due to privacy or transmission issues. Thus, a common approach is to compute a reconstructed training dataset before compression. The current reconstruction methods compute the reconstructed training dataset with a generator by exploiting information from the pre-trained model. However, current reconstruction methods focus on extracting more information from the pre-trained model but do not leverage network engineering. This work is the first to consider network engineering as an approach to design the reconstruction method. Specifically, we propose the AutoReCon method, which is a neural architecture search-based reconstruction method. In the proposed AutoReCon method, the generator architecture is designed automatically given the pre-trained model for reconstruction. Experimental results show that using generators discovered by the AutoRecon method always improve the performance of data-free compression. Baozhou Zhu, H. Peter Hofstee, Johan Peltenburg, Jinho Lee 0001, Zaid Al-Ars |
IJCAI | 4 |
| 2021 | Qimera: Data-free Quantization with Synthetic Boundary Supporting SamplesabstractModel quantization is known as a promising method to compress deep neural networks, especially for inferences on lightweight mobile or edge devices. However, model quantization usually requires access to the original training data to maintain the accuracy of the full-precision models, which is often infeasible in real-world scenarios for security and privacy issues.A popular approach to perform quantization without access to the original data is to use synthetically generated samples, based on batch-normalization statistics or adversarial learning.However, the drawback of such approaches is that they primarily rely on random noise input to the generator to attain diversity of the synthetic samples. We find that this is often insufficient to capture the distribution of the original data, especially around the decision boundaries.To this end, we propose Qimera, a method that uses superposed latent embeddings to generate synthetic boundary supporting samples.For the superposed embeddings to better reflect the original distribution, we also propose using an additional disentanglement mapping layer and extracting information from the full-precision model.The experimental results show that Qimera achieves state-of-the-art performances for various settings on data-free quantization. Code is available at https://github.com/iamkanghyunchoi/qimera. Kanghyun Choi, Deokki Hong, Noseong Park, Youngsok Kim, Jinho Lee 0001 |
NeurIPS | 5 |
| 2020 | FlexReduce: Flexible All-reduce for Distributed Deep Learning on Asymmetric Network TopologyabstractWe propose FlexReduce, an efficient and flexible all-reduce algorithm for distributed deep learning under irregular network hierarchies. With ever-growing deep neural networks, distributed learning over multiple nodes is becoming imperative for expedited training. There are several approaches leveraging the symmetric network structure to optimize the performance over different hierarchy levels of the network. However, the assumption of symmetric network does not always hold, especially in shared cloud environments. By allocating an uneven portion of gradients to each learner (GPU), FlexReduce outperforms conventional algorithms on asymmetric network structures, and still performs even or better on symmetric networks. Jinho Lee 0001, Inseok Hwang 0001, Soham Shah, Minsik Cho |
DAC | 1 |
| 2020 | In-memory database acceleration on FPGAs: a surveyabstractAbstract While FPGAs have seen prior use in database systems, in recent years interest in using FPGA to accelerate databases has declined in both industry and academia for the following three reasons. First, specifically for in-memory databases, FPGAs integrated with conventional I/O provide insufficient bandwidth, limiting performance. Second, GPUs, which can also provide high throughput, and are easier to program, have emerged as a strong accelerator alternative. Third, programming FPGAs required developers to have full-stack skills, from high-level algorithm design to low-level circuit implementations. The good news is that these challenges are being addressed. New interface technologies connect FPGAs into the system at main-memory bandwidth and the latest FPGAs provide local memory competitive in capacity and bandwidth with GPUs. Ease of programming is improving through support of shared coherent virtual memory between the host and the accelerator, support for higher-level languages, and domain-specific tools to generate FPGA designs automatically. Therefore, this paper surveys using FPGAs to accelerate in-memory database systems targeting designs that can operate at the speed of main memory. Jian Fang 0004, Yvo T. B. Mulder, Jan Hidders, Jinho Lee 0001, H. Peter Hofstee |
VLDB J. | 4 |
| 2019 | Refine and Recycle: A Method to Increase Decompression ParallelismabstractRapid increases in storage bandwidth, combined with a desire for operating on large datasets interactively, drives the need for improvements in high-bandwidth decompression. Existing designs either process only one token per cycle or process multiple tokens per cycle with low area efficiency and/or low clock frequency. We propose two techniques to achieve high single-decoder throughput at improved efficiency by keeping only a single copy of the history data across multiple BRAMs and operating on each BRAM independently. A first stage efficiently refines the tokens into commands that operate on a single BRAM and steers the commands to the appropriate one. In the second stage, a relaxed execution model is used where each BRAM command executes immediately and those with invalid data are recycled to avoid stalls caused by the read-after-write dependency. We apply these techniques to Snappy decompression and implement a Snappy decompression accelerator on a CAPI2-attached FPGA platform equipped with a Xilinx VU3P FPGA. Experimental results show that our proposed method achieves up to 7.2 GB/s output throughput per decompressor, with each decompressor using 14.2% of the logic and 7% of the BRAM resources of the device. Therefore, a single decompressor can easily keep pace with an NVMe device (PCIe Gen3 x4) on a small FPGA, while a larger device, integrated on a host bridge adapter and instantiating multiple decompressors, can keep pace with the full OpenCAPI 3.0 bandwidth of 25 GB/s. Jian Fang 0004, Jinho Lee 0001, Zaid Al-Ars, H. Peter Hofstee |
ASAP | 3 |
| 2019 | A Fine-Grained Parallel Snappy Decompressor for FPGAs Using a Relaxed Execution ModelabstractSnappy is a widely used (de) compression algorithm in many big data applications. Such a data compression technique has been proven to be successful to save storage space and to reduce the amount of data transmission from/to storage devices. In this paper, we present a fine-grained parallel Snappy decompressor on FPGAs running under a relaxed execution model that addresses the following main challenges in existing solutions. First, existing designs either can only process one token per cycle or can process multiple tokens per cycle with low area efficiency and/or low clock frequency. Second, the high read-after-write data dependency during decompression introduces stalls which pull down the throughput. Jian Fang 0004, Jinho Lee 0001, Zaid Al-Ars, H. Peter Hofstee |
FCCM | 3 |
| 2019 | Towards Peripheral Awareness of Remote Family Member's Context Using Self-mobile Robotic AvatarsabstractReal-time remote interaction has become easier and richer powered by recent advances in mobile computing and communication. A number of research have been explored on enriching family interaction by augmenting an interaction channel with asynchronous communication [6] or additional sensory stimuli [5]. However, it is still far from achieving a sense of living together for family members involuntarily living apart, especially in context-aware impromptu interaction. For families living together, it is trivial to naturally perceive behavioral and situational contexts of the other and initiate a relevant interaction intuitively. For example, a wife starts a casual chat with asking her husband what he is going to cook when she sees him going to the kitchen or hears a simmering sound. Bumsoo Kang, Inseok Hwang 0001, Jinho Lee 0001, Seungchul Lee, Taegyeong Lee, Youngjae Chang 0001, Min Kyung Lee |
MobiSys | 3 |
| 2019 | An Efficient Graph Compressor Based on Adaptive Prefix EncodingabstractIn this paper we introduce APEC, a graph compression/decompression framework. A key component of APEC is adaptive prefix code, a novel variable-length coding scheme which can adapt to varying characteristics of different vertices in the graph data. APEC also encompasses many software optimization techniques including compressed vertex indexing, bit counting and parallelization. The net outcome is that APEC not only achieves up to 20% improvement on compression ratio, which is equivalent to 2.28 bits/edge, but also as much as 9x faster in compression and up to 20x faster in decompression compared to the existing frameworks. Moreover, APEC is capable of random accessing compressed data and performing compression on extremely large graph datasets. Jinho Lee 0001, Frank Liu 0001 |
SSDBM | 1 |
| 2018 | My Being to Your Place, Your Being to My Place: Co-present Robotic Avatars Create Illusion of Living TogetherabstractPeople in work-separated families have been heavily relying on cutting-edge face-to-face communication services. Despite their ease of use and ubiquitous availability, experiences in living together are still far incomparable to those through remote face-to-face communication. We envision that enabling a remote person to be spatially superposed in one's living space would be a breakthrough to catalyze pseudo living-together interactivity. We propose HomeMeld, a zero-hassle self-mobile robotic system serving as a co-present avatar to create a persistent illusion of living together for those who are involuntarily living apart. The key challenges are 1) continuous spatial mapping between two heterogeneous floor plans and 2) navigating the robotic avatar to reflect the other's presence in real time under the limited maneuverability of the robot. We devise a notion of functionally equivalent location and orientation to translate a person's presence into another in a heterogeneous floor plan. We also develop predictive path warping to seamlessly synchronize the presence of the other. We conducted extensive experiments and deployment studies with real participants. Bumsoo Kang, Inseok Hwang 0001, Jinho Lee 0001, Seungchul Lee, Taegyeong Lee, Youngjae Chang 0001, Min Kyung Lee |
MobiSys | 3 |
| 2018 | HomeMeld: Co-present Robotic Avatar System for Illusion of Living TogetherabstractNo abstract available. Bumsoo Kang, Inseok Hwang 0001, Jinho Lee 0001, Seungchul Lee, Taegyeong Lee, Youngjae Chang 0001, Min Kyung Lee |
MobiSys | 3 |
| 2018 | Deep neural networks with weighted spikes
Heesu Kim, Subin Huh, Jinho Lee 0001, Kiyoung Choi |
Neurocomputing | 4 |
| 2018 | TEI-NoC: Optimizing Ultralow Power NoCs Exploiting the Temperature Effect InversionabstractThe era of the Internet of Things (IoT) is upon us. In this era, minimizing power consumption becomes a primary concern of system-on-chip designers. Ultralow power (ULP) very large-scale integration circuits have been receiving considerable interest from both academia and industry as the best-suited techniques for IoT devices, which can take full advantage of power-saving that voltage scaling potentially achieves. Consequently, research on ULP designs has begun to yield tangible outcomes, namely ULP circuits. However, little attention has been paid to ULP network-on-chip (NoC), although the NoC is an essential of the ULP chips, and its power consumption accounts for a significant portion of the total power. This paper focuses on ULP NoCs, and presents a new power management method that exploits delay versus temperature characteristics of ULP circuits. Recent studies on ULP circuits show that delay versus temperature characteristics are fundamentally different from normal circuits, i.e., the delay of the ULP circuits implemented in state-of-the-art bulk CMOS operating at low supply voltages or in FinFET technologies decreases with increasing temperature, a phenomenon known as the temperature effect inversion (TEI). Starting with an intuition that at a certain temperature point, power savings without performance penalty can be achieved by increasing the router frequency to create the opportunity to turn off some routers in ULP NoCs, or by decreasing the NoC supply voltage level, an optimization method is presented to maximize the power savings with minor performance penalty. To validate the proposed method, a concrete ULP NoC simulator, TEI-Noxim, has been developed. Experimental results demonstrate that TEI-aware NoC achieves an average of 36.0% power reduction over 21 applications. Kyuseung Han, Jae-Jin Lee, Jinho Lee 0001, Massoud Pedram |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2017 | Scalable time-versioning support for property graph databasesabstractWhen graphs change over time, it is important to make the changes trackable for many graph-based applications. We propose an implementation of OLTP-oriented graph database that supports time-versioning. There has been a few snapshot-based approaches for supporting time-versions, but they usually require the full-restoration of the graph, and lack the resolution of the time space. Using a B-tree as the datastructure for the backend storage, our database allow fast and scalable support for restoring the arbitrary part of the graph, without slowing down the normal accesses to the current graph. Experimental results show that our scheme is much efficient than the straightforward solutions, in terms of space and performance. Warut D. Vijitbenjaronk, Jinho Lee 0001, Toyotaro Suzumura, Ilie Gabriel Tanase |
IEEE BigData | 2 |
| 2017 | SCI-FII: Speculative Conversational Interface Framework for Incremental Inference on Modularized ServicesabstractWe propose Sci-Fii, a speculative conversational interface framework for incremental inference on modularized services. To build one's own conversational interface with existing business logic, cloud-based modularized services offer a suite of ready-to-use components to ease development, ensure cross-platform flexibility, and encapsulate computational complexities. However developing with the modularized services often results in a chain of discrete modules with limited inter-module data sharing, which yields unnecessarily long response times of the conversational interface, aggravates user experiences, and eventually harms user retention. Sci-Fii offers a uniform framework that enables existing serviced modules to benefit from intermediate data and early parallel execution. Transparent to developers, Sci-Fii helps the end-to-end conversational interface system work fluidly and exhibit faster and more natural response times. Jinho Lee 0001, Inseok Hwang 0001, Thomas Hubregtsen, Anne E. Gattiker, Christopher M. Durham |
MDM | 1 |
| 2017 | ExtraV: Boosting Graph Processing Near Storage with a Coherent AcceleratorabstractIn this paper, we propose ExtraV, a framework for near-storage graph processing. It is based on the novel concept of graph virtualization , which efficiently utilizes a cache-coherent hardware accelerator at the storage side to achieve performance and flexibility at the same time. ExtraV consists of four main components: 1) host processor, 2) main memory, 3) AFU (Accelerator Function Unit) and 4) storage. The AFU, a hardware accelerator, sits between the host processor and storage. Using a coherent interface that allows main memory accesses, it performs graph traversal functions that are common to various algorithms while the program running on the host processor (called the host program) manages the overall execution along with more application-specific tasks. Graph virtualization is a high-level programming model of graph processing that allows designers to focus on algorithm-specific functions. Realized by the accelerator, graph virtualization gives the host programs an illusion that the graph data reside on the main memory in a layout that fits with the memory access behavior of host programs even though the graph data are actually stored in a multi-level, compressed form in storage. We prototyped ExtraV on a Power8 machine with a CAPI-enabled FPGA. Our experiments on a real system prototype offer significant speedup compared to state-of-the-art software only implementations. Jinho Lee 0001, Heesu Kim, Sungjoo Yoo, Kiyoung Choi, H. Peter Hofstee, Gi-Joon Nam, Mark Nutter, Damir A. Jamsek |
Proc. VLDB Endow. | 1 |
| 2017 | Excavating the Hidden Parallelism Inside DRAM Architectures With Buffered ComparesabstractWe propose an approach called buffered compares, a less-invasive processing-in-memory solution that can be used with existing processor memory interfaces such as DDR3/4 with minimal changes. The approach is based on the observation that multibank architecture, a key feature of modern main memory DRAM devices, can be used to provide huge internal bandwidth without any major modification. We place a small buffer and a simple ALU per bank, define a set of new DRAM commands to fill the buffer and feed data to the ALU, and return the result for a set of commands (not for each command) to the host memory controller. By exploiting the under-utilized internal bandwidth using `compare-n-op' operations, which are frequently used in various applications, we not only reduce the amount of energy-inefficient processor-memory communication, but also accelerate the computation of big data processing applications by utilizing parallelism of the buffered compare units in DRAM banks. We present two versions of buffered compare architecture-full-scale architecture and reduced architecture-in trade of performance and energy. The experimental results show that our solution significantly improves the performance and efficiency of the system on the tested workloads. Jinho Lee 0001, Jongwook Chung, Jung Ho Ahn, Kiyoung Choi |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Buffered compares: Excavating the hidden parallelism inside DRAM architectures with lightweight logic
Jinho Lee 0001, Jung Ho Ahn, Kiyoung Choi |
DATE | 1 |
| 2015 | THOR: Orchestrated thermal management of cores and networks in 3D many-core architecturesabstractMost previous researches on thermal management of many-core architectures focus on the control of either core resources or network resources only, even though both have significant thermal impacts. This paper proposes a holistic thermal management that applies dynamic voltage/frequency scaling to cores and routers together to maximize system performance under temperature constraint. The proposed method first determines a power budget given in aggregate weighted power for every pillar of vertically adjacent tiles. Then it performs voltage/frequency assignment under the budget while exploiting the characteristics of the applications. Experiments show that our approach outperforms existing methods. Jinho Lee 0001, Junwhan Ahn, Kiyoung Choi, Kyungsu Kang |
ASP-DAC | 1 |
| 2015 | REDELF: An Energy-Efficient Deadlock-Free Routing for 3D NoCs with Partial Vertical Connectionsabstract3D integrated circuits (3D ICs) using through-silicon vias (TSVs) allow to envision the stacking of dies with different functions and technologies, using as an interconnect backbone a 3D network-on-chip (NoC). However, partial vertical connection in 3D NoCs seems unavoidable because of the large overhead of TSV itself (e.g., large footprint, low fabrication yield, additional fabrication processes) as well as the heterogeneity in dimension. This article proposes an energy-efficient deadlock-free routing algorithm for 3D mesh topologies where vertical connections partially exist. By introducing some rules for selecting elevators (i.e., vertical links between dies), the routing algorithm can eliminate the dedicated virtual channel requirement. In this article, the rules themselves as well as the proof of deadlock freedom are given. By eliminating the virtual channels for deadlock avoidance, the proposed routing algorithm reduces the energy consumption by 38.9% compared to a conventional routing algorithm. When the virtual channel is used for reducing the head-of-line blocking, the proposed routing algorithm increases performance by up to 23.1% and 6.9% on average. Jinho Lee 0001, Kyungsu Kang, Kiyoung Choi |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2013 | Deflection routing in 3D Network-on-Chip with TSV serializationabstractThis paper proposes a deflection routing for 3D NoC with serialized TSVs. Bufferless deflection routing provides area- and power-efficient communication under low to medium traffic load. Under 3D circumstances, the bufferless deflection routing can yield even better performance than buffered routing when key aspects are properly taken into account. Evaluation of the proposed scheme shows its effectiveness in throughput, latency, and energy consumption. Jinho Lee 0001, Sunwook Kim, Kiyoung Choi |
ASP-DAC | 1 |
| 2013 | A deadlock-free routing algorithm requiring no virtual channel on 3D-NoCs with partial vertical connectionsabstractElevator-first routing algorithm has been introduced for partially connected 3D network-on-chips, as a low-cost, distributed and deadlock-free routing algorithm using two virtual channels. This paper proposes Redelf, a modification of the elevator-first routing algorithm on a 3D mesh topology. The proposed algorithm requires no virtual channel to ensure deadlock-freedom. Jinho Lee 0001, Kiyoung Choi |
NOCS | 1 |
| 2013 | Mapping and Scheduling of Tasks and Communications on Many-Core SoC Under Local Memory ConstraintabstractThere has been extensive research on mapping and scheduling tasks on a many-core SoC. However, none considers the optimization of communication types, which can significantly affect performance, energy consumption, and local memory usage of the SoC. This paper presents an approach to automatic mapping and scheduling of tasks and communications on a many-core SoC. The key idea is to decide the type of each communication between message passing and shared memory when we do the mapping and scheduling. By assigning a proper type to each communication, we can optimize the energy consumption, performance, or energy-delay product. To solve the optimization problem, the approach adopts a probabilistic algorithm coupled with some heuristics. To enhance throughput of the system, it performs software pipelined scheduling of the tasks using a modified iterative modulo scheduling technique. Experiments show that our algorithm achieves on average 50.1% lower energy consumption, 21.0% higher throughput, and 64.9% lower energy- delay product, compared to shared memory only communication. Jinho Lee 0001, Moo-Kyoung Chung, Yeongon Cho, Soojung Ryu, Jung Ho Ahn, Kiyoung Choi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Deflection routing in 3D network-on-chip with limited vertical bandwidthabstractThis article proposes a deflection routing for 3D NoC with serialized TSVs for vertical links. Compared to buffered routing, deflection routing provides area- and power-efficient communication and little loss of performance under low to medium traffic load. Under 3D environments, the deflection routing can yield even better performance than buffered routing when key aspects are properly taken into account. However, the existing deflection routing technique cannot be directly applied because the serialized TSV links will take longer time to send data than ordinary planar links and cause many problems. A naive deflection through a TSV link can cause significantly longer latency and more energy consumption even for communications through planar links. This article proposes a method to mitigate the effect and also solve arising deadlock and livelock problems. Evaluation of the proposed scheme shows its effectiveness in throughput, latency, and energy consumption. Jinho Lee 0001, Sunwook Kim, Kiyoung Choi |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2012 | Memory-aware mapping and scheduling of tasks and communications on many-core SoCabstractThis paper presents an approach to automatic task mapping, scheduling, and communication routing on a many-core SoC, considering the trade-offs between two different communication types—message passing and shared memory—for the communication routing in order to optimize the energy consumption or performance. To solve the optimization problem, the approach uses the quantum-inspired evolutionary algorithm. For the scheduling of the tasks with backward dependencies, it uses the iterative modulo scheduling technique. Experiments with random task graphs as well as real applications show the effectiveness of the proposed approach. Jinho Lee 0001, Kiyoung Choi |
ASP-DAC | 1 |
| 2012 | An adaptive routing algorithm for 3D mesh NoC with limited vertical bandwidth
Mingyang Zhu, Jinho Lee 0001, Kiyoung Choi |
VLSI-SoC | 2 |