EDBT 2026 Demo / reviewers in the wild / expert
Youngsok Kim
dblp:147/1001
· DBLP profile ↗
48ranked-venue papers
3as first author
36since 2021 · last 2026
0000-0002-1015-9969ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 3 first-author · 29 since 2021Software engineering, systems software and programming languages · 11 · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LoCaLUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIMabstractLookup tables (LUTs) have recently gained attention as an alternative compute mechanism that maps input operands to precomputed results, eliminating the need for arithmetic logic. LUTs not only reduce logic complexity, but also naturally support diverse numerical precisions without requiring separate circuits for each bitwidth—an increasingly important feature in quantized DNNs. This creates a favorable tradeoff in PIM: memory capacity can be used in place of logic to increase computational throughput, aligning well with DRAM-PIM architectures that offer high bandwidth and easily available memory but limited logic density. In this work, we explore this capacity-computation tradeoff in LUT-based PIM designs, where memory capacity is traded for performance by packing multiple MAC operations into a single LUT lookup. Building on this insight, we propose LoCaLUT, a PIM-based design for efficient low-bit quantized DNN inference using operation-packed LUTs. First, we observe that these LUTs contain extensive redundancy and introduce LUT canonicalization, which eliminates duplicate entries to reduce LUT size. Second, we propose reordering LUT, a lightweight auxiliary LUT that remaps weight vectors to their canonical form required by LUT canonicalization with a simple LUT lookup. Third, we propose LUT slice streaming, a novel execution strategy that exploits the DRAM-buffer hierarchy by streaming only relevant LUT columns into the buffer and reusing them across multiple weight vectors. Evaluated on a real system based on UPMEM devices, we demonstrate a geometric mean speedup of$1.82 \times$across various numeric precisions and DNN models. We believe LoCaLUT opens a path toward scalable, low-logic PIM designs tailored for LUT-based DNN inference. Our implementation of LoCaLUT is available at https://github.com/AIS-SNU/LoCaLUT. Junguk Hong, Changmin Shin 0002, Sukjin Kim, Si Ung Noh, Taehee Kwon 0002, Seongyeon Park, Hanjun Kim 0001, Youngsok Kim, Jinho Lee 0001 |
HPCA | 8 |
| 2026 | SMTcheck: Accurate SMT Interference Prediction to Improve Scheduling Efficiency in DatacentersabstractSimultaneous multithreading (SMT) is widely used in modern x86 processors to improve core utilization by sharing hardware resources between co-located threads. However, such resource sharing often leads to severe performance interference, making efficient workload co-scheduling difficult, especially given the complexity and diversity of modern x86 CPUs. Our analysis reveals that SMT-aware workload scheduling can significantly improve system throughput and reduce tail latency for datacenter workloads, but identifying optimal thread combinations is challenging due to the lack of visibility into platform-specific resource sharing behaviors. In this paper, we present SMTcheck, a lightweight, accurate, and platformindependent methodology for predicting SMT interference for diverse x86 processors. SMTcheck uses carefully designed code snippets (Diags) to extract hidden microarchitectural features of performance-critical shared resources. With these extracted features, SMTcheck builds per-resource microbenchmarks (Injectors) to apply pinpoint pressure to specific target resources in order to capture workload-specific contention characteristics. SMTcheck then constructs a hardware-aware contention model to predict performance interference between arbitrary workload pairs without requiring exhaustive profiling. We evaluate SMTcheck on six x86 desktop processors and five x86 server processors from Intel and AMD across different generations and show that it achieves high prediction accuracy by up to 95.5 % (94.6 % on average). We further demonstrate its effectiveness by implementing a contention-aware scheduler in the Linux kernel. Compared to the default Linux scheduler, our contention-aware scheduler significantly reduces tail latency for latency-critical workloads (e.g., database, key-value store) by up to 36.09 %, and improves the overall system throughput by up to$1.072 \times$. Finally, using real-world cluster traces from Alibaba and Google, we demonstrate that SMTcheck incurs negligible profiling overheads ($\approx 0.113 \%$), making it practical for deployment in productionscale datacenter environments. Jinhyeok Oh, Gyutae Kim, Youngsok Kim, Jae-Hyun Hwang, Joonsung Kim 0001 |
HPCA | 5 |
| 2026 | FaScalSQL: A Fast and Scalable GPU-Accelerated SQL Query Engine for Out-of-Memory Tables
Chaemin Lim, Kwanghyun Park 0001, Youngsok Kim |
ICDE | 7 |
| 2026 | Peak-memory-aware partitioning and scheduling for multi-tenant DNN model inference
Jaeho Lee 0005, Ju Min Lee, Haeeun Jeong, Hyunho Kwon, Youngsok Kim, Yongjun Park 0001, Hanjun Kim 0001 |
J. Syst. Archit. | 5 |
| 2026 | DANCE++: Differentiable Accelerator/Network Co-Exploration With Hard Constraints and Data-Free Training for Real-World ScenariosabstractCo-exploration of neural architectures and hardware accelerators has emerged as a promising approach to address computational cost problems, especially in low-profile systems. However, existing co-exploration methods based on reinforcement learning or evolutionary search suffer from substantial search costs. To address this, this work presents DANCE++, a differentiable approach towards the co-exploration of hardware and network architecture design. At the heart of DANCE++ is a differentiable evaluator network that models hardware metrics with a neural network, enabling accelerator design through backpropagation. DANCE++ significantly reduces search time and enhances accuracy and hardware cost metrics compared to traditional approaches. To further address real-world scenarios, this work embodies two important practical topics: hard constraints and data dependency. To meet the constraints such as frame rates or area budget, this work proposes a gradient manipulation algorithm that guides differentiable optimization to find hard-constrained solutions. Also to consider cases where training dataset is inaccessible, this work proposes to use data-free training methods in both co-exploration and training phases. To the best of our knowledge, DANCE++ is the first co-exploration method that targets these real-world challenges, supported by extensive experiments demonstrating its effectiveness. Kanghyun Choi, Deokki Hong, Hyeyoon Lee, Joonsang Yu, Noseong Park, Youngsok Kim, Jinho Lee 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | DMO-DB: Mitigating the Data Movement Bottlenecks of GPU-Accelerated Relational OLAPabstractGraphics Processing Units (GPUs) offer high computational throughput and memory bandwidth, making them promising accelerators for relational OnLine Analytical Processing (OLAP). GPU-accelerated relational OLAP executes the relational operations of a Structured Query Language (SQL) query on a GPU instead of the host Central Processing Unit (CPU). Depending on where input columns and their values reside in, a GPU-accelerated SQL query execution can be classified into two scenarios: 1) in-GPU, in which all input columns fit in the GPU memory, or 2) in-host, in which the input columns reside in the host memory and get transferred to the GPU memory when needed. However, both scenarios incur significant intraGPU and host-to-GPU data movement overheads, respectively. In-GPU executions incur excessive GPU cache misses and thus frequent off-chip GPU memory accesses. In-host executions suffer from the limited host-to-GPU data transfer bandwidth. This paper presents DMO-DB, a Data Movement-Optimized GPU-accelerated relational OLAP engine. Since modern GPUaccelerated relational OLAP decomposes SQL queries into multiple pipelines-sequences of relational operations that can be executed on input columns from the same table, DMO-DB leverages inter-pipeline dependencies to overcome the two data movement bottlenecks. DMO-DB introduces two key ideas: cache-fit bloom filtering and Ahead-of-Time value Discarding (AoTD), which preemptively eliminate unnecessary input values before their movement across the memory hierarchies. For in-GPU execution, GPU L1 data cache-fit filters discard non-contributing values before triggering costly off-chip DRAM accesses. For in-host execution, host CPU last level cache-fit filters strategically prune unnecessary input values, minimizing PCIe transfer overhead. After that, AoTD exploits multiple inter-pipeline dependencies by collecting these cache-fit bloom filters to earlier pipeline execution stages. Our evaluation using NVIDIA RTX A4000 and TITAN RTX GPUs shows that DMO-DB achieves speedups of $\mathbf{1. 5 3 x}$ over in-GPU Crystal-Opt and 6.10x over in-host HeavyDB. Chaemin Lim, Suhyun Lee 0002, Jinwoo Choi 0003, Joonsung Kim 0001, Jinho Lee 0001, Youngsok Kim |
PACT | 6 |
| 2025 | PIM-CARE: A Compiler-Assisted Dynamic Resource Allocation Framework for Real-world DRAM PIMabstractProcessing-In-Memory (PIM) has recently emerged as a promising solution to alleviate the memory bottleneck by integrating computing capabilities into memory chips. Since PIM provides numerous Processing Elements (PEs) and high-bandwidth on-chip data transfers, full utilization of the PEs becomes a critical mission to maximize the performance of PIM applications. However, due to the diverse and complex characteristics of PIM applications, using more resources does not always improve performance. It is therefore important to find the suitable amount of resources to achieve the best performance and to fully utilize the PIM resources.To address this, we introduce PIM-CARE, a framework for dynamic resource allocation across multiple applications with compiler support on real-world PIM systems. PIM-CARE first determines the best amount of PIM resources to allocate for each application. To enable spatial multitasking, the PIM-CARE daemon monitors resource allocation and deallocation requests and estimates total PIM resource utilization at runtime. It then dynamically schedules applications using a priority-based out-of-order policy, considering both available PIM resources and resource requirements for best performance. Evaluation on real-world PIM systems shows that PIM-CARE improves throughput by 5.49x and average turnaround time by 5.71x compared to the baseline. Inyong Hwang, Donghyeon Kim 0001, Seokwon Kang, Taehyeong Park 0001, Taehoon Kim 0001, Jiwon Seo 0002, Hanjun Kim 0001, Youngsok Kim, Yongjun Park 0001 |
ICS | 8 |
| 2025 | GCStack+GCScaler: Fast and Accurate GPU Performance Analyses Using Fine-Grained Stall Cycle Accounting and Interval AnalysisabstractTo design next-generation Graphics Processing Units (GPUs), GPU architects rely on GPU performance analyses to identify key GPU performance bottlenecks and explore GPU design spaces.Unfortunately, the existing GPU performance analysis mechanisms make it difficult for GPU architects to conduct fast and accurate GPU performance analyses.The existing mechanisms can provide misleading Hanna Cha, Sungchul Lee, Jounghoo Lee, Yeonan Ha, Joonsung Kim 0001, Youngsok Kim |
ISCA | 6 |
| 2025 | LATPC: Accelerating GPU Address Translation Using Locality-Aware TLB Prefetching and MSHR Compression
Yeonan Ha, Hanna Cha, Jiwon Lee 0001, Joonsung Kim 0001, Won Woo Ro, Youngsok Kim |
MICRO | 7 |
| 2024 | GraNNDis: Fast Distributed Graph Neural Network Training Framework for Multi-Server ClustersabstractGraph neural networks (GNNs) are one of the rapidly growing fields within deep learning. While many distributed GNN training frameworks have been proposed to increase the training throughput, they face three limitations when applied to multi-server clusters. 1) They suffer from an inter-server communication bottleneck because they do not consider the inter-/intra-server bandwidth gap, a representative characteristic of multi-server clusters. 2) Redundant memory usage and computation hinder the scalability of the distributed frameworks. 3) Sampling methods, de facto standard in mini-batch training, incur unnecessary errors in multi-server clusters. Jaeyong Song 0002, Hongsun Jang, Hunseong Lim, Jaewon Jung 0001, Youngsok Kim, Jinho Lee 0001 |
PACT | 5 |
| 2024 | MPC-Wrapper: Fully Harnessing the Potential of Samsung Aquabolt-XL HBM2-PIM on FPGAsabstractProcessing-In-Memory (PIM) is an attractive solution for mitigating frequent and large data movement between computational units and memory devices. Among various PIM implementations, Samsung Aquabolt-XL is an HBM2 memory device which implements 16 PIM-enabled pseudo-channels and associates an In-Memory Processor (IMP) to each pair of the memory banks. Recent studies have shown that Aquabolt-XL can greatly accelerate various applications (e.g., deep learning) by offloading memory-intensive operations (e.g., matrix-vector multiplications) to the IMPs. However, the prior study fails to fully utilize Aquabolt-XL and achieves limited performance gains by offloading operations to the IMPs of only a single pseudo-channel. Ideally, utilizing all the 16 pseudo-channels of Aquabolt-XL can further accelerate the key operations by a factor of 16× compared to utilizing only a single pseudo-channel. To fully exploit Aquabolt-XL, therefore, memory-intensive operations should be offloaded to and concurrently executed on the IMPs of all the PIM-enabled pseudo-channels. This paper presents MPC-Wrapper, a multi-pseudo-channel wrapper interface which allows memory-intensive operations to be offloaded to and concurrently executed on the IMPs of all the 16 PIM-enabled pseudo-channels of Aquabolt-XL. First, MPC-Wrapper allows all the PIM-enabled pseudo-channels to operate independently and in parallel, thus achieving high scalability needed for fully utilizing all the PIM-enabled pseudo-channels of Aquabolt-XL. Second, MPC-Wrapper is highly flexible as it exposes the PIM-enabled pseudo-channels as separate ports and enables an FPGA logic to flexibly utilize any set of the PIM-enabled pseudo-channels according to its needs. Third, MPC-Wrapper achieves high usability by hiding the complex low-level interactions between the memory controller and Aquabolt-XL for initializing and invoking the PIM-enabled pseudo-channels from the other FPGA logics. Using an Aquabolt-XL-equipped Xilinx Alveo U280 FPGA and four memory-intensive benchmarks, we show that utilizing all the 16 PIM-enabled pseudo-channels of Aquabolt-XL with MPC-Wrapper achieves a geometric mean speedup of 13.66× over the baseline single PIM-enabled pseudo-channel implementations of the benchmarks. Jinwoo Choi 0003, Yeonan Ha, Hanna Cha, Seil Lee, Sungchul Lee, Jounghoo Lee, Shinhaeng Kang, Bongjun Kim, Hanwoong Jung, Hanjun Kim 0001, Youngsok Kim |
FCCM | 11 |
| 2024 | Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real SystemabstractThe recent huge advance of Large Language Models (LLMs) is mainly driven by the increase in the number of parameters. This has led to substantial memory capacity requirements, necessitating the use of dozens of GPUs just to meet the capacity. One popular solution to this is storage-offloaded training, which uses host memory and storage as an extended memory hierarchy. However, this obviously comes at the cost of storage bandwidth bottleneck because storage devices have orders of magnitude lower bandwidth compared to that of GPU device memories. Our work, Smart-Infinity, addresses the storage bandwidth bottleneck of storage-offloaded LLM training using near-storage processing devices on a real system. The main component of Smart-Infinity is SmartUpdate, which performs parameter updates on custom near-storage accelerators. We identify that moving parameter updates to the storage side removes most of the storage traffic. In addition, we propose an efficient data transfer handler structure to address the system integration issues for Smart-Infinity. The handler allows overlapping data transfers with fixed memory consumption by reusing the device buffer. Lastly, we propose accelerator-assisted gradient compression/decompression to enhance the scalability of Smart-Infinity. When scaling to multiple near-storage processing devices, the write traffic on the shared channel becomes the bottleneck. To alleviate this, we compress the gradients on the GPU and decompress them on the accelerators. It provides further acceleration from reduced traffic. As a result, Smart-Infinity achieves a significant speedup compared to the baseline. Notably, SmartInfinity is a ready-to-use approach that is fully integrated into PyTorch on a real system. The implementation of Smart-Infinity is available at https://github.com/AIS-SNU/smart-infinity. Hongsun Jang, Jaeyong Song 0002, Jaewon Jung 0001, Youngsok Kim, Jinho Lee 0001 |
HPCA | 5 |
| 2024 | CR2: Community-aware Compressed Regular Representation for Graph Processing on a GPUabstractThanks to its massive parallel resources, a GPU is a promising platform for graph processing. However, the increasing size and skewed characteristics of the real-world graphs limit the performance improvement. Prior work proposes locality-enhancing graph transformations and load balancing techniques to improve performance, but they still suffer from excessive memory usage and inefficient parallel resource utilization because their graph representations are not fully tailored for a GPU. To efficiently utilize the GPU resource with less memory, this work proposes a new graph representation, called CR2. First, CR2 extracts community-aware subgraphs from a graph by clustering densely-connected vertices together. For the community-aware subgraphs, CR2 decomposes a vertex ID into a cluster ID and a local ID and represents each vertex only with the local ID, thus reducing memory usage. Second, CR2 additionally partitions the graph into multiple degree-ordered subgraphs in which all the vertices have the same regularized number of edges, thus making parallel workload balanced across GPU warps. This work evaluates CR2 with four commonly used graph algorithms and shows that CR2 achieves 1.53 times performance speedup while using 32.1% less memory on the geomean average compared to the state-of-the-art techniques. Shinnung Jeong, Sungjun Cho, Yongwoo Lee 0001, Seonyeong Heo, Gwangsun Kim, Youngsok Kim, Hanjun Kim 0001 |
ICPP | 7 |
| 2024 | PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM DevicesabstractRecent dual in-line memory modules (DIMMs) are starting to support processing-in-memory (PIM) by associating their memory banks with processing elements (PEs), allowing applications to overcome the data movement bottleneck by offloading memory-intensive operations to the PEs. Many highly parallel applications have been shown to benefit from these PIM-enabled DIMMs, but further speedup is often limited by the huge overhead of inter-PE collective communication. This mainly comes from the slow CPU-mediated inter-PE communication methods, making it difficult for PIM-enabled DIMMs to accelerate a wider range of applications. Prior studies have tried to alleviate the communication bottleneck, but they lack enough flexibility and performance to be used for a wide range of applications. In this paper, we present PID-Comm, a fast and flexible inter-PE collective communication framework for commodity PIM-enabled DIMMs. The key idea of PID-Comm is to abstract the PEs as a multi-dimensional hypercube and allow multiple instances of inter-PE collective communication between the PEs belonging to certain dimensions of the hypercube. Leveraging this abstraction, PID-Comm first defines eight interPE collective communication patterns that allow applications to easily express their complex communication patterns. Then, PIDComm provides high-performance implementations of the interPE collective communication patterns optimized for the DIMMs. Our evaluation using 16 UPMEM DIMMs and representative parallel algorithms shows that PID-Comm greatly improves the performance by up to $5.19 \times$ compared to the existing inter-PE communication implementations. The implementation of PIDComm is available at https://github.com/AIS-SNU/PID-Comm. Si Ung Noh, Junguk Hong, Chaemin Lim, Seongyeon Park, Jeehyun Kim, Hanjun Kim 0001, Youngsok Kim, Jinho Lee 0001 |
ISCA | 7 |
| 2024 | Orchestrating Multiple Mixed Precision Models on a Shared Precision-Scalable NPUabstractMixed-precision quantization can reduce the computational requirements of Deep Neural Network (DNN) models with minimal loss of accuracy. As executing mixed-precision DNN models on Neural Processing Units (NPUs) incurs significant under-utilization of computational resources, Precision-Scalable NPUs (PSNPUs) which can process multiple low-precision layers simultaneously have been proposed. However, the under-utilization still remains significant due to the lack of adequate scheduling algorithms to support multiple mixed-precision models on PSNPUs. Therefore, in this paper, we propose a dynamic programming-based scheduling algorithm for the operations of multiple mixed-precision models. Our scheduling algorithm finds the optimal execution plan that exploits the precision-scalable MACs to improve the end-to-end inference latency of mixed-precision models. We evaluate the performance of this algorithm in terms of hardware utilization, inference latency, and schedule search time compared to baseline scheduling algorithms. The experimental results show 1.23 inference latency improvements over the baseline algorithms within the allowed minutes. Kiung Jung, Seok Namkoong, Hongjun Um, Hyejun Kim, Youngsok Kim, Yongjun Park 0001 |
LCTES | 5 |
| 2024 | AGAThA: Fast and Efficient GPU Acceleration of Guided Sequence Alignment for Long Read MappingabstractWith the advance in genome sequencing technology, the lengths of deoxyribonucleic acid (DNA) sequencing results are rapidly increasing at lower prices than ever. However, the longer lengths come at the cost of a heavy computational burden on aligning them. For example, aligning sequences to a human reference genome can take tens or even hundreds of hours. The current de facto standard approach for alignment is based on the guided dynamic programming method. Although this takes a long time and could potentially benefit from high-throughput graphic processing units (GPUs), the existing GPU-accelerated approaches often compromise the algorithm's structure, due to the GPU-unfriendly nature of the computational pattern. Unfortunately, such compromise in the algorithm is not tolerable in the field, because sequence alignment is a part of complicated bioinformatics analysis pipelines. In such circumstances, we propose AGAThA, an exact and efficient GPU-based acceleration of guided sequence alignment. We diagnose and address the problems of the algorithm being unfriendly to GPUs, which comprises strided/redundant memory accesses and workload imbalances that are difficult to predict. According to the experiments on modern GPUs, AGAThA achieves 18.8× speedup against the CPU-based baseline, 9.6× against the best GPU-based baseline, and 3.6× against GPU-based algorithms with different heuristics. Seongyeon Park, Junguk Hong, Jaeyong Song 0002, Hajin Kim, Youngsok Kim, Jinho Lee 0001 |
PPoPP | 5 |
| 2024 | SPID-Join: A Skew-resistant Processing-in-DIMM Join Algorithm Exploiting the Bank- and Rank-level Parallelisms of DIMMsabstractRecent advances in Dual In-line Memory Modules (DIMMs) allow DIMMs to support Processing-In-DIMM (PID) by placing In-DIMM Processors (IDPs) near their memory banks. Prior studies have shown that in-memory joins can benefit from PID by offloading their operations onto the IDPs and exploiting the high internal memory bandwidth of DIMMs. Aimed at evenly balancing the computational loads between the IDPs, the existing algorithms perform IDP-wise global partitioning on input tables and then make each IDP process a partition of the input tables. Unfortunately, we find that the existing PID join algorithms achieve low performance and scalability with skewed input tables. With skewed input tables, the IDP-wise global partitioning incurs imbalanced loads between the IDPs, making the IDPs remain idle until the heaviest-load IDP completes processing its partition. To fully exploit the IDPs for accelerating in-memory joins involving skewed input tables, therefore, we need a new PID join algorithm which achieves high skew resistance by mitigating the imbalanced inter-IDP loads. In this paper, we present SPID-Join, a skew-resistant PID join algorithm which exploits two parallelisms inherent in DIMM architectures, namely bank- and rank-level parallelisms. By replicating join keys across the banks within a rank and across ranks, SPID-Join significantly increases the internal memory bandwidth and computational throughput allocated to each join key, improving the load balance between the IDPs and accelerating join executions. SPID-Join exploits the bank- and the rank-level parallelisms to minimize join key replication overheads and support a wider range of join key replication ratios. Despite achieving high skew resistance, SPID-Join exhibits a trade-off between the join key replication ratio and the join execution latency, making the best-performing join key replication ratio depend on join and PID system configurations. We, therefore, augment SPID-Join with a cost model which identifies the best-performing join key replication ratio for given join and PID system configurations. By accurately modeling and scaling the IDPs' throughput and the inter-IDP communication bandwidth, the cost model accurately captures the impact of the join key replication ratio on SPID-Join. Our experimental results using eight UPMEM DIMMs, which collectively provide a total of 1,024 IDPs, show that SPID-Join achieves up to 10.38x faster join executions over PID-Join, the state-of-the-art PID join algorithm, with highly skewed input tables. Suhyun Lee 0002, Chaemin Lim, Jinwoo Choi 0003, Heelim Choi, Yongjun Park 0001, Kwanghyun Park 0001, Hanjun Kim 0001, Youngsok Kim |
Proc. ACM Manag. Data | 9 |
| 2023 | Virtual PIM: Resource-Aware Dynamic DPU Allocation and Workload Scheduling Framework for Multi-DPU PIM ArchitectureabstractProcessing-in-Memory (PIM) is an attractive device that can effectively satisfy the rapidly increasing demands for memory-intensive workloads in emerging application domains, such as deep learning and big data processing. Thanks to the integrated design of the main memory (MRAM) and multiple data processing units (DPUs) on a single chip, the PIM devices can provide massive parallelism from numerous DPUs and the substantial bandwidth between the MRAM and DPUs, thus achieving the high performance for the memory-intensive workloads. However, although the recent PIM architectures, including UPMEM, can efficiently execute a single memory-intensive application, they fail to efficiently orchestrate multiple applications on the multiple DPU resources due to the conservative resource allocation, without a resource monitoring system, and large scheduling granularity. To solve these problems, we propose a novel resource-aware dynamic DPU allocation and workload scheduling framework, called Virtual PIM, for multi-DPU PIM architectures such as UPMEM. The framework initially virtualizes the DPU and MRAM to ensure data consistency in multi-application environments. For dynamic DPU allocation, the Virtual PIM framework continuously gathers resource requests from multiple processes and current DPU occupancy information to estimate the dynamic DPU resource status, irrespective of PIM hardware support. Based on this information, the framework dynamically allocates DPUs and schedules workloads in fine-grained levels with minimum occupancy to maximize total DPU utilization. Our evaluations in real PIM environments demonstrate that Virtual PIM significantly improves system throughput and average normalized turnaround time by up to 4.83x and 3.45x, respectively, compared to the SLURM-based baseline. Donghyeon Kim 0001, Taehoon Kim 0001, Inyong Hwang, Taehyeong Park 0001, Hanjun Kim 0001, Youngsok Kim, Yongjun Park 0001 |
PACT | 6 |
| 2023 | Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication CompressionabstractIn training of modern large natural language processing (NLP) models, it has become a common practice to split models using 3D parallelism to multiple GPUs. Such technique, however, suffers from a high overhead of inter-node communication. Compressing the communication is one way to mitigate the overhead by reducing the inter-node traffic volume; however, the existing compression techniques have critical limitations to be applied for NLP models with 3D parallelism in that 1) only the data parallelism traffic is targeted, and 2) the existing compression schemes already harm the model quality too much. Jaeyong Song 0002, Jinkyu Yim, Jaewon Jung 0001, Hongsun Jang, Youngsok Kim, Jinho Lee 0001 |
ASPLOS (2) | 6 |
| 2023 | Occamy: Memory-efficient GPU Compiler for DNN InferenceabstractThis work proposes Occamy, a new memory-efficient DNN compiler that reduces the memory usage of a DNN model without affecting its accuracy. For each DNN operation, Occamy analyzes the dimensions of input and output tensors, and their liveness within the operation. Across all the operations, Occamy analyzes liveness of all the tensors, generates a memory pool after calculating the maximum required memory size, and schedules when and where to place each tensor in the memory pool. Compared to PyTorch, on an integrated embedded GPU for six DNNs, Occamy reduces the memory usage by 34.6% and achieves a geometric mean speedup of 1.25×. Jaeho Lee 0005, Shinnung Jeong, Seungbin Song, Kunwoo Kim, Heelim Choi, Youngsok Kim, Hanjun Kim 0001 |
DAC | 6 |
| 2023 | Pipe-BD: Pipelined Parallel Blockwise DistillationabstractTraining large deep neural network models is highly challenging due to their tremendous computational and mem-ory requirements. Blockwise distillation provides one promising method towards faster convergence by splitting a large model into multiple smaller models. In state-of-the-art blockwise distillation methods, training is performed block-by-block in a data-parallel manner using multiple GPUs. To produce inputs for the student blocks, the teacher model is executed from the beginning until the current block under training. However, this results in a high overhead of redundant teacher execution, low GPU utilization, and extra data loading. To address these problems, we propose Pipe-BD, a novel parallelization method for blockwise distillation. Pipe-BD aggressively utilizes pipeline parallelism for blockwise distillation, eliminating redundant teacher block execution and increasing per-device batch size for better resource utilization. We also extend to hybrid parallelism for efficient workload balancing. As a result, Pipe-BD achieves significant acceleration without modifying the mathematical formulation of blockwise distillation. We implement Pipe-BD on PyTorch, and experiments reveal that Pipe-BD is effective on multiple scenarios, models, and datasets. Hongsun Jang, Jaewon Jung 0001, Jaeyong Song 0002, Joonsang Yu, Youngsok Kim, Jinho Lee 0001 |
DATE | 5 |
| 2023 | SGCN: Exploiting Compressed-Sparse Features in Deep Graph Convolutional Network AcceleratorsabstractGraph convolutional networks (GCNs) are becoming increasingly popular as they overcome the limited applicability of prior neural networks. One recent trend in GCNs is the use of deep network architectures. As opposed to the traditional GCNs, which only span only around two to five layers deep, modern GCNs now incorporate tens to hundreds of layers with the help of residual connections. From such deep GCNs, we find an important characteristic that they exhibit very high intermediate feature sparsity. This reveals a new opportunity for accelerators to exploit in GCN executions that was previously not present.In this paper, we propose SGCN, a fast and energy-efficient GCN accelerator which fully exploits the sparse intermediate features of modern GCNs. SGCN suggests several techniques to achieve significantly higher performance and energy efficiency than the existing accelerators. First, SGCN employs a GCN-friendly feature compression format. We focus on reducing the off-chip memory traffic, which often is the bottleneck for GCN executions. Second, we propose microarchitectures for seamlessly handling the compressed feature format. Specifically, we modify the aggregation phase of GCN to process compressed features, and design a combination engine that can output compressed features at no extra memory traffic cost. Third, to better handle locality in the existence of the varying sparsity, SGCN employs sparsity-aware cooperation. Sparsity-aware cooperation creates a pattern that exhibits multiple reuse windows, such that the cache can capture diverse sizes of working sets and therefore adapt to the varying level of sparsity. Through a thorough evaluation, we show that SGCN achieves 1.66× speedup and 44.1% higher energy efficiency compared to the existing accelerators in geometric mean. Mingi Yoo, Jaeyong Song 0002, Jounghoo Lee, Namhyung Kim, Youngsok Kim, Jinho Lee 0001 |
HPCA | 5 |
| 2023 | McCore: A Holistic Management of High-Performance Heterogeneous MulticoresabstractHeterogeneous multicore systems have emerged as a promising approach to scale performance in high-end desktops within limited power and die size constraints. Despite their advantages, these systems face three major challenges: memory bandwidth limitation, shared cache contention, and heterogeneity. Small cores in these systems tend to occupy a significant portion of shared LLC and memory bandwidth, despite their lower computational capabilities, leading to performance degradation of up to 18% in memory-intensive workloads. Therefore, it is crucial to address these challenges holistically, considering shared resources and core heterogeneity while managing shared cache and bandwidth. Jaewon Kwon, Yongju Lee 0003, Hongju Kal, Minjae Kim 0010, Youngsok Kim, Won Woo Ro |
MICRO | 5 |
| 2023 | Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMsabstractModern dual in-line memory modules (DIMMs) support processing-in-memory (PIM) by implementing in-DIMM processors (IDPs) located near memory banks. PIM can greatly accelerate in-memory join, whose performance is frequently bounded by main-memory accesses, by offloading the operations of join from host central processing units (CPUs) to the IDPs. As real PIM hardware has not been available until very recently, the prior PIM-assisted join algorithms have relied on PIM hardware simulators which assume fast shared memory between the IDPs and fast inter-IDP communication; however, on commodity PIM-enabled DIMMs, the IDPs do not share memory and demand the CPUs to mediate inter-IDP communication. Such discrepancies in the architectural characteristics make the prior studies incompatible with the DIMMs. Thus, to exploit the high potential of PIM on commodity PIM-enabled DIMMs, we need a new join algorithm designed and optimized for the DIMMs and their architectural characteristics. In this paper, we design and analyze Processing-In-DIMM Join (PID-Join), a fast in-memory join algorithm which exploits UPMEM DIMMs, currently the only publicly-available PIM-enabled DIMMs. The DIMMs impose several key challenges on efficient acceleration of join including the shared-nothing nature and limited compute capabilities of the IDPs, the lack of hardware support for fast inter-IDP communication, and the slow IDP-wise data transfers between the IDPs and the main memory. PID-Join overcomes the challenges by prototyping and evaluating hash, sort-merge, and nested-loop algorithms optimized for the IDPs, enabling fast inter-IDP communication using host CPU cache streaming and vector instructions, and facilitating fast rank-wise data transfers between the IDPs and the main memory. Our evaluation using a real system equipped with eight UPMEM DIMMs and 1,024 IDPs shows that PID-Join greatly improves the performance of in-memory join over various CPU-based in-memory join algorithms. Chaemin Lim, Suhyun Lee 0002, Jinwoo Choi 0003, Jounghoo Lee, Seongyeon Park, Hanjun Kim 0001, Jinho Lee 0001, Youngsok Kim |
Proc. ACM Manag. Data | 8 |
| 2023 | Enabling Fine-Grained Spatial Multitasking on Systolic-Array NPUs Using Dataflow MirroringabstractNeural Processing Units (NPUs) frequently suffer from low hardware utilization as the efficiency of their systolic arrays heavily depends on the characteristics of a deep neural network (DNN). Spatial multitasking is a promising solution to overcome the low NPU hardware utilization; however, the state-of-the-art spatial-multitasking NPU architecture achieves sub-optimal performance due to its coarse-grained systolic-array distribution and incurs significant implementation costs. In this paper, we proposedataflow-mirroring NPU (DM-NPU), a novel spatial-multitasking NPU architecture supporting fine-grained systolic-array distribution. The key idea of DM-NPU is to reverse the dataflows of co-located DNNs in horizontal and/or vertical directions. DM-NPU can place allocation boundaries between any adjacent processing elements of a systolic array, both horizontally and vertically. We then proposeDM-Perf, an accurate systolic-array NPU performance model, to maximize the spatial-multitasking performance of DM-NPU. Utilizing the existing performance models achieves sub-optimal performance as they cannot accurately capture the resource contention caused by spatial multitasking. DM-Perf, on the other hand, exploits the per-layer performance profiles of a DNN to accurately capture the resource contention. Our evaluation using MLPerf DNNs shows that DM-NPU and DM-Perf can greatly improve the performance by up to 35.1% over the state-of-the-art NPU architecture and performance model. Jinwoo Choi 0003, Yeonan Ha, Jounghoo Lee, Jinho Lee 0001, Hanhwi Jang, Youngsok Kim |
IEEE Trans. Computers | 7 |
| 2022 | Decoupling Schedule, Topology Layout, and Algorithm to Easily Enlarge the Tuning Space of GPU Graph ProcessingabstractOnly with a right schedule and a right topology layout, a graph algorithm can be efficiently processed on GPUs. Existing GPU graph processing frameworks try to find an optimal schedule and topology layout for an algorithm via iterative search, but they fail to find the optimal configuration because their schedules and topology layouts are tightly coupled in their processing models. Moreover, their tightly coupled schedules and topology layouts make it difficult for developers to extend the tuning space. To easily enlarge the tuning space of GPU graph processing, this work proposes a new GPU graph processing abstraction scheme that fully decouples schedules, topology layouts, and algorithms from each other with abstraction interfaces. Moreover, this work proposes GRAssembler, a new GPU graph processing framework that efficiently integrates the decoupled schedule, topology layout, and algorithm without abstraction overhead. Thanks to the efficient decoupling and integration, GRAssembler increases the tuning space from 336 to 4,480 and achieves 30.4% higher performance on geomean average, compared to the state-of-the-art GPU graph processing framework. Shinnung Jeong, Yongwoo Lee 0001, Jaeho Lee 0005, Heelim Choi, Seungbin Song, Jinho Lee 0001, Youngsok Kim, Hanjun Kim 0001 |
PACT | 7 |
| 2022 | Slice-and-Forge: Making Better Use of Caches for Graph Convolutional Network AcceleratorsabstractGraph convolutional networks (GCNs) are becoming increasingly popular as they can process a wide variety of data formats that prior deep neural networks cannot easily support. One key challenge in designing hardware accelerators for GCNs is the vast size and randomness in their data access patterns which greatly reduces the effectiveness of the limited on-chip cache. Aimed at improving the effectiveness of the cache by mitigating the irregular data accesses, prior studies often employ the vertex tiling techniques used in traditional graph processing applications. While being effective at enhancing the cache efficiency, those approaches are often sensitive to the tiling configurations where the optimal setting heavily depends on target input datasets. Furthermore, the existing solutions require manual tuning through trial-and-error or rely on sub-optimal analytical models. Mingi Yoo, Jaeyong Song 0002, Hyeyoon Lee, Jounghoo Lee, Namhyung Kim, Youngsok Kim, Jinho Lee 0001 |
PACT | 6 |
| 2022 | It's All In the Teacher: Zero-Shot Quantization Brought Closer to the TeacherabstractModel quantization is considered as a promising method to greatly reduce the resource requirements of deep neural networks. To deal with the performance drop induced by quantization errors, a popular method is to use training data to fine-tune quantized networks. In real-world environments, however, such a method is frequently infeasible because training data is unavailable due to security, privacy, or confidentiality concerns. Zero-shot quantization addresses such problems, usually by taking information from the weights of a full-precision teacher network to compensate the performance drop of the quantized networks. In this paper, we first analyze the loss surface of state-of-the-art zero-shot quantization techniques and provide several findings. In contrast to usual knowledge distillation problems, zero-shot quantization often suffers from 1) the difficulty of optimizing multiple loss terms together, and 2) the poor generalization capability due to the use of synthetic samples. Furthermore, we observe that many weights fail to cross the rounding threshold during training the quantized networks even when it is necessary to do so for better performance. Based on the observations, we propose AIT, a simple yet powerful technique for zero-shot quantization, which addresses the aforementioned two problems in the following way: AIT i) uses a KL distance loss only without a cross-entropy loss, and ii) manipulates gradients to guarantee that a certain portion of weights are properly updated after crossing the rounding thresholds. Experiments show that AIT outperforms the performance of many existing methods by a great margin, taking over the overall state-of-the-art position in the field. Kanghyun Choi, Hyeyoon Lee, Deokki Hong, Joonsang Yu, Noseong Park, Youngsok Kim, Jinho Lee 0001 |
CVPR | 6 |
| 2022 | Enabling hard constraints in differentiable neural network and accelerator co-explorationabstractCo-exploration of an optimal neural architecture and its hardware accelerator is an approach of rising interest which addresses the computational cost problem, especially in low-profile systems. The large co-exploration space is often handled by adopting the idea of differentiable neural architecture search. However, despite the superior search efficiency of the differentiable co-exploration, it faces a critical challenge of not being able to systematically satisfy hard constraints such as frame rate. To handle the hard constraint problem of differentiable co-exploration, we propose HDX, which searches for hard-constrained solutions without compromising the global design objectives. By manipulating the gradients in the interest of the given hard constraint, high-quality solutions satisfying the constraint can be obtained. Deokki Hong, Kanghyun Choi, Hyeyoon Lee, Joonsang Yu, Noseong Park, Youngsok Kim, Jinho Lee 0001 |
DAC | 6 |
| 2022 | SALoBa: Maximizing Data Locality and Workload Balance for Fast Sequence Alignment on GPUsabstractSequence alignment forms an important backbone in many sequencing applications. A commonly used strategy for sequence alignment is an approximate string matching with a two-dimensional dynamic programming approach. Although some prior work has been conducted on GPU acceleration of a sequence alignment, we identify several shortcomings that limit exploiting the full computational capability of modern GPUs. This paper presents SALoBa, a GPU-accelerated sequence alignment library focused on seed extension. Based on the analysis of previous work with real-world sequencing data, we propose techniques to exploit the data locality and improve work-load balancing. The experimental results reveal that SALoBa significantly improves the seed extension kernel compared to state-of-the-art GPU-based methods. Seongyeon Park, Hajin Kim, Nauman Ahmed, Zaid Al-Ars, H. Peter Hofstee, Youngsok Kim, Jinho Lee 0001 |
IPDPS | 7 |
| 2022 | GCoM: a detailed GPU core model for accurate analytical modeling of modern GPUsabstractAnalytical models can greatly help computer architects perform orders of magnitude faster early-stage design space exploration than using cycle-level simulators. To facilitate rapid design space exploration for graphics processing units (GPUs), prior studies have proposed GPU analytical models which capture first-order stall events causing performance degradation; however, the existing analytical models cannot accurately model modern GPUs due to their outdated and highly abstract GPU core microarchitecture assumptions. Therefore, to accurately evaluate the performance of modern GPUs, we need a new GPU analytical model which accurately captures the stall events incurred by the significant changes in the core microarchitectures of modern GPUs. Jounghoo Lee, Yeonan Ha, Suhyun Lee 0002, Jinyoung Woo, Jinho Lee 0001, Hanhwi Jang, Youngsok Kim |
ISCA | 7 |
| 2022 | GuardiaNN: Fast and Secure On-Device Inference in TrustZone Using Embedded SRAM and Cryptographic HardwareabstractAs more and more mobile/embedded applications employ Deep Neural Networks (DNNs) involving sensitive user data, mobile/embedded devices must provide a highly secure DNN execution environment to prevent privacy leaks. Aimed at securing DNN data, recent studies execute part of a DNN in a trusted execution environment (e.g., TrustZone) to isolate DNN execution from the other processes; however, as the trusted execution environments for mobile/embedded devices provide limited memory protection, DNN data remain unencrypted in DRAM and become vulnerable to physical attacks. The devices can prevent the physical attacks by keeping DNN data encrypted in DRAM; when DNN data get referenced during DNN execution, they get loaded to the SRAM and get decrypted by a CPU core. Unfortunately, using the SRAM with demand paging greatly increases DNN execution time due to the inefficient use of the SRAM and the high CPU consumption of data encryption/decryption. Jinwoo Choi 0003, Jaeyeon Kim, Chaemin Lim, Suhyun Lee 0002, Jinho Lee 0001, Dokyung Song, Youngsok Kim |
Middleware | 7 |
| 2021 | Thread-Aware Area-Efficient High-Level Synthesis Compiler for Embedded DevicesabstractIn the embedded device market, custom hardware platforms such as an application specific integrated circuit (ASIC) and a field programmable gate array (FPGA) are attractive thanks to their high performance and power efficiency. However, its huge design costs make it challenging for manufacturers to timely launch new devices. High-level synthesis (HLS) helps significantly reduce the design costs by automating the translation of service algorithms into hardware logics; however, current HLS compilers do not fit well to embedded devices as they fail to produce area-efficient solutions while supporting concurrent events from diverse peripherals such as sensors, actuators and network modules. This paper proposes a new thread-aware HLS compiler named Duro that produces area-efficient embedded devices. Duro shares commonly-invoked functions and operators across different callers and threads with a new thread-aware area cost model, and thus effectively reduces the logic size. Moreover, Duro supports a variety of device peripherals by automatically integrating peripheral controllers and interfaces as peripheral drivers. The experiment results of six embedded devices with ten peripherals demonstrate that Duro reduces the area and energy dissipation of embedded devices by 28.5% and 25.3% compared with the designs generated by the state-of-the-art HLS compiler. This work also implements FPGA prototypes of the six devices using Duro, and the measurement results show 65.3% energy saving over Raspberry Pi Zero with slightly better computation performance. Changsu Kim 0004, Shinnung Jeong, Sungjun Cho, Yongwoo Lee 0001, William Song, Youngsok Kim, Hanjun Kim 0001 |
CGO | 6 |
| 2021 | DANCE: Differentiable Accelerator/Network Co-ExplorationabstractThis work presents DANCE, a differentiable approach towards the co-exploration of hardware accelerator and network architecture design. At the heart of DANCE is a differentiable evaluator network. By modeling the hardware evaluation software with a neural network, the relation between the accelerator design and the hardware metrics becomes differentiable, allowing the search to be performed with backpropagation. Compared to the naive existing approaches, our method performs co-exploration in a significantly shorter time, while achieving superior accuracy and hardware cost metrics. Kanghyun Choi, Deokki Hong, Hojae Yoon, Joonsang Yu, Youngsok Kim, Jinho Lee 0001 |
DAC | 5 |
| 2021 | Dataflow Mirroring: Architectural Support for Highly Efficient Fine-Grained Spatial Multitasking on Systolic-Array NPUsabstractWe present dataflow mirroring, architectural support for low-overhead fine-grained systolic array allocation which overcomes the limitations of prior coarse-grained spatial-multitasking Neural Processing Unit (NPU) architectures. The key idea of dataflow mirroring is to reverse the dataflows of co-located Neural Networks (NNs) in horizontal and/or vertical directions, allowing allocation boundaries to be set between any adjacent rows and columns of a systolic array and supporting up to four-way spatial multitasking. Our detailed experiments using MLPerf NNs and a dataflow-mirroring-augmented NPU prototype which extends Google’s TPU with dataflow mirroring shows that dataflow mirroring can significantly improve the multitasking performance by up to 46.4%. Jounghoo Lee, Jinwoo Choi 0003, Jaeyeon Kim, Jinho Lee 0001, Youngsok Kim |
DAC | 5 |
| 2021 | Qimera: Data-free Quantization with Synthetic Boundary Supporting SamplesabstractModel quantization is known as a promising method to compress deep neural networks, especially for inferences on lightweight mobile or edge devices. However, model quantization usually requires access to the original training data to maintain the accuracy of the full-precision models, which is often infeasible in real-world scenarios for security and privacy issues.A popular approach to perform quantization without access to the original data is to use synthetically generated samples, based on batch-normalization statistics or adversarial learning.However, the drawback of such approaches is that they primarily rely on random noise input to the generator to attain diversity of the synthetic samples. We find that this is often insufficient to capture the distribution of the original data, especially around the decision boundaries.To this end, we propose Qimera, a method that uses superposed latent embeddings to generate synthetic boundary supporting samples.For the superposed embeddings to better reflect the original distribution, we also propose using an additional disentanglement mapping layer and extracting information from the full-precision model.The experimental results show that Qimera achieves state-of-the-art performances for various settings on data-free quantization. Code is available at https://github.com/iamkanghyunchoi/qimera. Kanghyun Choi, Deokki Hong, Noseong Park, Youngsok Kim, Jinho Lee 0001 |
NeurIPS | 4 |
| 2020 | Real-Time Object Detection System with Multi-Path Neural NetworksabstractThanks to the recent advances in Deep Neural Networks (DNNs), DNN-based object detection systems become highly accurate and widely used in real-time environments such as autonomous vehicles, drones and security robots. Although the systems should detect objects within a certain time limit that can vary depending on their execution environments such as vehicle speeds, existing systems blindly execute the entire long-latency DNNs without reflecting the time-varying time limits, and thus they cannot guarantee real-time constraints. This work proposes a novel real-time object detection system that employs multipath neural networks based on a new worst-case execution time (WCET) model for DNNs on a GPU. This work designs the WCET model for a single DNN layer analyzing processor and memory contention on GPUs, and extends the WCET model to the end-to-end networks. This work also designs the multipath networks with three new operators such as skip, switch, and dynamic generate proposals that dynamically change their execution paths and the number of target objects. Finally, this work proposes a path decision model that chooses the optimal execution path at run-time reflecting dynamically changing environments and time constraints. Our detailed evaluation using widely-used driving datasets shows that the proposed real-time object detection system performs as good as a baseline object detection system without violating the time-varying time limits. Moreover, the WCET model predicts the worst-case execution latency of convolutional and group normalization layers with only 27% and 81% errors on average, respectively. Seonyeong Heo, Sungjun Cho, Youngsok Kim, Hanjun Kim 0001 |
RTAS | 3 |
| 2019 | μLayer: Low Latency On-Device Inference Using Cooperative Single-Layer Acceleration and Processor-Friendly QuantizationabstractEmerging mobile services heavily utilize Neural Networks (NNs) to improve user experiences. Such NN-assisted services depend on fast NN execution for high responsiveness, demanding mobile devices to minimize the NN execution latency by efficiently utilizing their underlying hardware resources. To better utilize the resources, existing mobile NN frameworks either employ various CPU-friendly optimizations (e.g., vectorization, quantization) or exploit data parallelism using heterogeneous processors such as GPUs and DSPs. However, their performance is still bounded by the performance of the single target processor, so that realtime services such as voice-driven search often fail to react to user requests in time. It is obvious that this problem will become more serious with the introduction of more demanding NN-assisted services. Youngsok Kim, Joonsung Kim 0001, Dongju Chae, Dae-Hyun Kim 0003, Jangwoo Kim |
EuroSys | 1 |
| 2019 | FlexLearn: Fast and Highly Efficient Brain Simulations Using Flexible On-Chip LearningabstractTo understand how the human brain works, neuroscientists heavily rely on brain simulations which incorporate the concept of time to their operating model. In the simulations, neurons transmit their signals through synapses whose weights change over time and by the activity of the associated neurons. Such changes in synaptic weights, known as learning, are thought to contribute to memory, and various learning rules exist to model different behaviors of the human brain. Due to the diverse neurons and learning rules, neuroscientists perform the simulations using highly programmable general-purpose processors. Unfortunately, the processors greatly suffer from the high computational overheads of the learning rules. As an alternative, brain simulation accelerators achieve orders of magnitude higher performance; however, they have limited flexibility and cannot support the diverse neurons and learning rules. Eunjin Baek, Hunjun Lee, Youngsok Kim, Jangwoo Kim |
MICRO | 3 |
| 2018 | Google Workloads for Consumer Devices: Mitigating Data Movement BottlenecksabstractWe are experiencing an explosive growth in the number of consumer devices, including smartphones, tablets, web-based computers such as Chromebooks, and wearable devices. For this class of devices, energy efficiency is a first-class concern due to the limited battery capacity and thermal power budget. We find that data movement is a major contributor to the total system energy and execution time in consumer devices. The energy and performance costs of moving data between the memory system and the compute units are significantly higher than the costs of computation. As a result, addressing data movement is crucial for consumer devices. In this work, we comprehensively analyze the energy and performance impact of data movement for several widely-used Google consumer workloads: (1) the Chrome web browser; (2) TensorFlow Mobile, Google's machine learning framework; (3) video playback, and (4) video capture, both of which are used in many video services such as YouTube and Google Hangouts. We find that processing-in-memory (PIM) can significantly reduce data movement for all of these workloads, by performing part of the computation close to memory. Each workload contains simple primitives and functions that contribute to a significant amount of the overall data movement. We investigate whether these primitives and functions are feasible to implement using PIM, given the limited area and power constraints of consumer devices. Our analysis shows that offloading these primitives to PIM logic, consisting of either simple cores or specialized accelerators, eliminates a large amount of data movement, and significantly reduces total system energy (by an average of 55.4% across the workloads) and execution time (by an average of 54.2%). Amirali Boroumand, Saugata Ghose, Youngsok Kim, Rachata Ausavarungnirun, Eric Shiu, Rahul Thakur, Dae-Hyun Kim 0003, Aki Kuusela, Allan Knies, Parthasarathy Ranganathan, Onur Mutlu |
ASPLOS | 3 |
| 2018 | DCS-ctrl: A Fast and Flexible Device-Control Mechanism for Device-Centric Server ArchitectureabstractModern high-performance servers leverage a large number of emerging peripheral devices (e.g., data processing accelerators, non-volatile memory storage, high-bandwidth network cards) to meet ever-increasing performance demands of server applications. However, as such servers experience severe kernel overhead due to frequently invoked device operations (e.g., buffer management and data copy), server architects have proposed various hardware and software approaches to enable direct communications among the devices. Unfortunately, existing direct device-to-device (D2D) communication schemes still suffer from low performance and the lack of flexibility. First, software-based schemes depend on complicated kernel routines and necessitate multiple hardware-software and user-kernel boundary crossings, which significantly limit the performance improvement opportunities from direct D2D communications. On the other hand, hardware-based schemes require tight integration and custom-built devices, preventing architects from flexibly adding off-the-shelf devices. In this paper, we propose DCS-ctrl, a novel Hardware-based Device-Control (HDC) mechanism for Device-Centric Server (DCS) architecture to provide fast and CPU-efficient direct D2D communications among a large number of off-the-shelf peripheral devices. The key idea of DCS-ctrl is to implement a low-cost and flexible device-control mechanism on an independent FPGA device called HDC Engine. As HDC Engine manages all data and control transfers among devices at the hardware level, the server achieves high performance, scalability, and flexibility. First, optimizing both data and control paths at the hardware level minimizes the latency of inter-device communications. Second, implementing FPGA-based reconfigurable device controllers enables direct D2D communications among commodity devices and thus improves per-device flexibility. Third, merging heterogeneous device operations with intermediate data processing supports creates more opportunities for direct inter-device communications in server applications. Our DCS-ctrl prototype reduces the latency of software-based direct D2D communications by 42% and the CPU utilization by 52%. Dongup Kwon, Jaehyung Ahn, Dongju Chae, Mohammadamin Ajdari, Suheon Bae, Youngsok Kim, Jangwoo Kim |
ISCA | 7 |
| 2018 | Flexon: A Flexible Digital Neuron for Efficient Spiking Neural Network SimulationsabstractSpiking Neural Networks (SNNs) play an important role in neuroscience as they help neuroscientists understand how the nervous system works. To model the nervous system, SNNs incorporate the concept of time into neurons and inter-neuron interactions called spikes; a neuron's internal state changes with respect to time and input spikes, and a neuron fires an output spike when its internal state satisfies certain conditions. As the neurons forming the nervous system behave differently, SNN simulation frameworks must be able to simulate the diverse behaviors of the neurons. To support any neuron models, some frameworks rely on general purpose processors at the cost of inefficiency in simulation speed and energy consumption. The other frameworks employ specialized accelerators to overcome the inefficiency; however, the accelerators support only a limited set of neuron models due to their model-driven designs, making accelerator-based frameworks unable to simulate target SNNs. In this paper, we present Flexon, a flexible digital neuron which exploits the biologically common features shared by diverse neuron models, to enable efficient SNN simulations. To design Flexon, we first collect SNNs from prior work in neuroscience research and analyze the neuron models the SNNs employ. From the analysis, we observe that the neuron models share a set of biologically common features, and that the features can be combined to simulate a significantly larger set of neuron behaviors than the existing model-driven designs. Furthermore, we find that the features share a small set of computational primitives which can be exploited to further reduce the chip area. The resulting digital neurons, Flexon and spatially folded Flexon, are flexible, highly efficient, and can be easily integrated with existing hardware. Our prototyping results using TSMC 45 nm standard cell library show that a 12-neuron Flexon array improves energy efficiency by 6,186x and 422x over CPU and GPU, respectively, in a small footprint of 9.26 mm2. The results also show that a 72-neuron spatially folded Flexon array incurs a smaller footprint of 7.62 mm2 and achieves geomean speedups of 122.45x and 9.83x over CPU and GPU, respectively. Dayeol Lee, Gwangmu Lee, Dongup Kwon, Sunghwa Lee, Youngsok Kim, Jangwoo Kim |
ISCA | 5 |
| 2017 | GPUpd: a fast and scalable multi-GPU architecture using cooperative projection and distributionabstractGraphics Processing Unit (GPU) vendors have been scaling single-GPU architectures to satisfy the ever-increasing user demands for faster graphics processing. However, as it gets extremely difficult to further scale single-GPU architectures, the vendors are aiming to achieve the scaled performance by simultaneously using multiple GPUs connected with newly developed, fast inter-GPU networks (e.g., NVIDIA NVLink, AMD XDMA). With fast inter-GPU networks, it is now promising to employ split frame rendering (SFR) which improves both frame rate and single-frame latency by assigning disjoint regions of a frame to different GPUs. Unfortunately, the scalability of current SFR implementations is seriously limited as they suffer from a large amount of redundant computation among GPUs. Youngsok Kim, Jae-Eon Jo, Hanhwi Jang, Minsoo Rhu, Hanjun Kim 0001, Jangwoo Kim |
MICRO | 1 |
| 2016 | CloudSwap: A Cloud-Assisted Swap Mechanism for Mobile DevicesabstractApplication caching is a key feature to enable fast application switches for mobile devices by caching the entire memory pages of applications in the device's physical memory. However, application caching requires a prohibitive amount of memory unless a swap feature is employed to maintain only the working sets of the applications in memory. Unfortunately, mobile devices often disable the invaluable swap feature as it can severely decrease the flash-based local storage device's already marginal lifespan due to the increased writes to the device. As a result, modern mobile devices suffering from the insufficient memory space end up killing memory-hungry applications and keeping only a few applications in the memory. In this paper, we propose CloudSwap, a fast and robust swap mechanism for mobile devices to enable the memory-oblivious application caching. The key idea of CloudSwap is to use the fast local storage as a cache of read-intensive swap pages, while storing prefetch-enabled, write-intensive swap pages in a cloud storage. To preserve the lifespan of the local storage, CloudSwap minimizes the number of writes to the local storage by storing the modified portions of the locally swapped pages in a cloud. To reduce the remote swap-in latency, CloudSwap exploits two cloud-assisted prefetch schemes, the app-aware read-ahead scheme and the access pattern-aware prefetch scheme. Our evaluation shows that the performance of CloudSwap is comparable to a fast, but lifespan-critical local swap system, with only 18% lifespan reduction, compared to the local swap system's 85% lifespan reduction. Dongju Chae, Joonsung Kim 0001, Youngsok Kim, Jangwoo Kim, Kyung-Ah Chang, Sang-Bum Suh, Hyogun Lee |
CCGrid | 3 |
| 2016 | Efficient footprint caching for Tagless DRAM CachesabstractEfficient cache tag management is a primary design objective for large, in-package DRAM caches. Recently, Tagless DRAM Caches (TDCs) have been proposed to completely eliminate tagging structures from both on-die SRAM and in-package DRAM, which are a major scalability bottleneck for future multi-gigabyte DRAM caches. However, TDC imposes a constraint on DRAM cache block size to be the same as OS page size (e.g., 4KB) as it takes a unified approach to address translation and cache tag management. Caching at a page granularity, or page-based caching, incurs significant off-package DRAM bandwidth waste by over-fetching blocks within a page that are not actually used. Footprint caching is an effective solution to this problem, which fetches only those blocks that will likely be touched during the page's lifetime in the DRAM cache, referred to as the page's footprint. In this paper we demonstrate TDC opens up unique opportunities to realize efficient footprint caching with higher prediction accuracy and a lower hardware cost than the original footprint caching scheme. Since there are no cache tags in TDC, the footprints of cached pages are tracked at TLB, instead of cache tag array, to incur much lower on-die storage overhead than the original design. Besides, when a cached page is evicted, its footprint will be stored in the corresponding page table entry, instead of an auxiliary on-die structure (i.e., Footprint History Table), to prevent footprint thrashing among different pages, thus yielding higher accuracy in footprint prediction. The resulting design, called Footprint-augmented Tagless DRAM Cache (F-TDC), significantly improves the bandwidth efficiency of TDC, and hence its performance and energy efficiency. Our evaluation with 3D Through-Silicon-Via-based in-package DRAM demonstrates an average reduction of off-package bandwidth by 32.0%, which, in turn, improves IPC and EDP by 17.7% and 25.4%, respectively, over the state-of-the-art TDC with no footprint caching. Hakbeom Jang, Yongjun Lee, Youngsok Kim, Jangwoo Kim, Jinkyu Jeong, Jae W. Lee |
HPCA | 4 |
| 2015 | DCS: a fast and scalable device-centric server architectureabstractConventional servers have achieved high performance by employing fast CPUs to run compute-intensive workloads, while making operating systems manage relatively slow I/O devices through memory accesses and interrupts. However, as the emerging workloads are becoming heavily data-intensive and the emerging devices (e.g., NVM storage, high-bandwidth NICs, and GPUs) come to enable low-latency and high-bandwidth device operations, the traditional host-centric server architectures fail to deliver high performance due to their inefficient device handling mechanisms. Furthermore, without resolving the architecture inefficiency, the performance loss will continue to increase as the emerging devices become faster. Jaehyung Ahn, Dongup Kwon, Youngsok Kim, Mohammadamin Ajdari, Jangwoo Kim |
MICRO | 3 |
| 2014 | GPUdmm: A high-performance and memory-oblivious GPU architecture using dynamic memory managementabstractGPU programmers suffer from programmer-managed GPU memory because both performance and programmability heavily depend on GPU memory allocation and CPU-GPU data transfer mechanisms. To improve performance and programmability, programmers should be able to place only the data frequently accessed by GPU on GPU memory while overlapping CPU-GPU data transfers and GPU executions as much as possible. However, current GPU architectures and programming models blindly place entire data on GPU memory, requiring a significantly large GPU memory size. Otherwise, they must trigger unnecessary CPU-GPU data transfers due to an insufficient GPU memory size. In this paper, we propose GPUdmm, a novel GPU architecture to enable high-performance and memory-oblivious GPU programming. First, GPUdmm uses GPU memory as a cache of CPU memory to provide programmers a view of the CPU memory-sized programming space. Second, GPUdmm achieves high performance by exploiting data locality and dynamically transferring data between CPU and GPU memories while effectively overlapping CPU-GPU data transfers and GPU executions. Third, GPUdmm can further reduce unnecessary CPU-GPU data transfers by exploiting simple programmer hints. Our carefully designed and validated experiments (e.g., PCIe/DMA timing) against representative benchmarks show that GPUdmm can achieve up to five times higher performance for the same GPU memory size, or reduce the GPU memory size requirement by up to 75% while maintaining the same performance. Youngsok Kim, Jae-Eon Jo, Jangwoo Kim |
HPCA | 1 |
| 2014 | Stealing Webpages Rendered on Your Browser by Exploiting GPU VulnerabilitiesabstractGraphics processing units (GPUs) are important components of modern computing devices for not only graphics rendering, but also efficient parallel computations. However, their security problems are ignored despite their importance and popularity. In this paper, we first perform an in-depth security analysis on GPUs to detect security vulnerabilities. We observe that contemporary, widely-used GPUs, both NVIDIA's and AMD's, do not initialize newly allocated GPU memory pages which may contain sensitive user data. By exploiting such vulnerabilities, we propose attack methods for revealing a victim program's data kept in GPU memory both during its execution and right after its termination. We further show the high applicability of the proposed attacks by applying them to the Chromium and Firefox web browsers which use GPUs for accelerating webpage rendering. We detect that both browsers leave rendered webpage textures in GPU memory, so that we can infer which web pages a victim user has visited by analyzing the remaining textures. The accuracy of our advanced inference attack that uses both pixel sequence matching and RGB histogram matching is up to 95.4%. Sangho Lee 0001, Youngsok Kim, Jangwoo Kim, Jong Kim 0001 |
IEEE Symposium on Security and Privacy | 2 |