VLDB 2026 Research / reviewers in the wild / expert
Shinnung Jeong
dblp:259/3732
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-8884-3851ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inside VOLT: Designing an Open-Source GPU Compiler (Tool)abstractRecent efforts in open-source GPU research are opening new avenues in a domain that has long been tightly coupled with a few commercial vendors. Emerging open GPU architectures define SIMT functionality through their own ISAs, but executing existing GPU programs and optimizing performance on these ISAs relies on a compiler framework that is technically complex and often undercounted in open-hardware development costs. Shinnung Jeong, Chihyo Ahn, Huanzhi Pu, Jisheng Zhao, Hyesoon Kim, Blaise-Pascal Tine |
CC | 1 |
| 2025 | SparseWeaver: Converting Sparse Operations as Dense Operations on GPUs for Graph WorkloadsabstractThanks to their scalable parallel processing capability, GPUs are promising computing resources for graph processing, in which identical operations are applied to a large number of edges and vertices. However, the sparsity and skewness of real-world graphs cause imbalanced workloads across GPU threads within the same warp, thus impeding efficient processing on the GPU. To mitigate this workload imbalance problem, existing works propose workload balancing hardware and software schemes. However, these solutions often suffer from additional memory overhead or increased computations and communication overheads during inter-warp and intra-warp synchronization. This work proposes a new hardware-software collaborative graph processing framework, SparseWeaver, that converts sparse operations in graph processing into dense operations using graph topology and makes the workloads balanced across GPU threads. Based on the analysis of common patterns in software schemes, we propose Weaver, a new lightweight GPU functional unit microarchitecture that fully leverages the benefits of the GPU architecture and exploits memory access locality. We prototype SparseWeaver on the open-source RISC-V Vortex GPU and demonstrate 2.36 times faster execution time compared to state-of-the-art schemes while incurring a low area overhead of 0.045% from increased dedicated logic registers. Shinnung Jeong, Liam Cooper, Ju Min Lee, Heelim Choi, Nicholas Parnenzini, Chihyo Ahn, Yongwoo Lee 0001, Hanjun Kim 0001, Hyesoon Kim |
HPCA | 1 |
| 2024 | CR2: Community-aware Compressed Regular Representation for Graph Processing on a GPUabstractThanks to its massive parallel resources, a GPU is a promising platform for graph processing. However, the increasing size and skewed characteristics of the real-world graphs limit the performance improvement. Prior work proposes locality-enhancing graph transformations and load balancing techniques to improve performance, but they still suffer from excessive memory usage and inefficient parallel resource utilization because their graph representations are not fully tailored for a GPU. To efficiently utilize the GPU resource with less memory, this work proposes a new graph representation, called CR2. First, CR2 extracts community-aware subgraphs from a graph by clustering densely-connected vertices together. For the community-aware subgraphs, CR2 decomposes a vertex ID into a cluster ID and a local ID and represents each vertex only with the local ID, thus reducing memory usage. Second, CR2 additionally partitions the graph into multiple degree-ordered subgraphs in which all the vertices have the same regularized number of edges, thus making parallel workload balanced across GPU warps. This work evaluates CR2 with four commonly used graph algorithms and shows that CR2 achieves 1.53 times performance speedup while using 32.1% less memory on the geomean average compared to the state-of-the-art techniques. Shinnung Jeong, Sungjun Cho, Yongwoo Lee 0001, Seonyeong Heo, Gwangsun Kim, Youngsok Kim, Hanjun Kim 0001 |
ICPP | 1 |
| 2024 | TeMCO: Tensor Memory Compiler Optimization across Tensor Decompositions in Deep Learning InferenceabstractSince the increasing complexity of deep learning models, tensor decomposition is one of the promising solutions that reduce computational complexity in deep learning models. By decomposing a convolution layer with a large weight tensor into multiple layers with smaller weight tensors, tensor decomposition can reduce the number of operations and weight memory spaces. However, existing tensor decomposition schemes face difficulties in reducing peak memory usage of the entire inference. The decomposed layers produce the reduced-sized tensors during inference, but the reduced tensors should be restored to their original sizes due to skip connections and non-decomposed activation layers between the decomposed layers. To reduce the peak memory usage of the end-to-end inference of the decomposed models, this work proposes a new tensor memory optimization scheme and its prototype compiler, called TeMCO. TeMCO replaces the original internal tensors used in the skip connections with reduced internal tensors derived by the decomposed layers. In addition, TeMCO fuses the decomposed layers and the non-decomposed activation layer and thus keeps the reduced internal tensors produced without restoring them. Thanks to the optimizations, this work reduces memory usage of internal tensors by 75.7% for 10 models of 5 deep learning architectures. Seungbin Song, Ju Min Lee, Haeeun Jeong, Hyunho Kwon, Shinnung Jeong, Jaeho Lee 0005, Hanjun Kim 0001 |
ICPP | 5 |
| 2023 | Occamy: Memory-efficient GPU Compiler for DNN InferenceabstractThis work proposes Occamy, a new memory-efficient DNN compiler that reduces the memory usage of a DNN model without affecting its accuracy. For each DNN operation, Occamy analyzes the dimensions of input and output tensors, and their liveness within the operation. Across all the operations, Occamy analyzes liveness of all the tensors, generates a memory pool after calculating the maximum required memory size, and schedules when and where to place each tensor in the memory pool. Compared to PyTorch, on an integrated embedded GPU for six DNNs, Occamy reduces the memory usage by 34.6% and achieves a geometric mean speedup of 1.25×. Jaeho Lee 0005, Shinnung Jeong, Seungbin Song, Kunwoo Kim, Heelim Choi, Youngsok Kim, Hanjun Kim 0001 |
DAC | 2 |
| 2022 | Decoupling Schedule, Topology Layout, and Algorithm to Easily Enlarge the Tuning Space of GPU Graph ProcessingabstractOnly with a right schedule and a right topology layout, a graph algorithm can be efficiently processed on GPUs. Existing GPU graph processing frameworks try to find an optimal schedule and topology layout for an algorithm via iterative search, but they fail to find the optimal configuration because their schedules and topology layouts are tightly coupled in their processing models. Moreover, their tightly coupled schedules and topology layouts make it difficult for developers to extend the tuning space. To easily enlarge the tuning space of GPU graph processing, this work proposes a new GPU graph processing abstraction scheme that fully decouples schedules, topology layouts, and algorithms from each other with abstraction interfaces. Moreover, this work proposes GRAssembler, a new GPU graph processing framework that efficiently integrates the decoupled schedule, topology layout, and algorithm without abstraction overhead. Thanks to the efficient decoupling and integration, GRAssembler increases the tuning space from 336 to 4,480 and achieves 30.4% higher performance on geomean average, compared to the state-of-the-art GPU graph processing framework. Shinnung Jeong, Yongwoo Lee 0001, Jaeho Lee 0005, Heelim Choi, Seungbin Song, Jinho Lee 0001, Youngsok Kim, Hanjun Kim 0001 |
PACT | 1 |
| 2022 | HECATE: Performance-Aware Scale Optimization for Homomorphic Encryption CompilerabstractDespite the benefit of Fully Homomorphic Encryption (FHE) that supports encrypted computation, writing an efficient FHE application is challenging due to magnitude scale management. Each FHE operation increases scales of ciphertext and leaving the scales high harms performance of the following FHE operations. Thus, rescaling ciphertext is inevitable to optimize an FHE application, but since FHE requires programmers to match the rescaling levels of operands of each FHE operation, programmers should rescale ciphertext reflecting the entire FHE application. Although recently proposed FHE compilers reduce the programming burden by automatically manipulating ciphertext scales, they fail to fully optimize the FHE application because they greedily rescale the ciphertext without considering their performance impacts throughout the entire application. This work proposes HECATE, a new FHE compiler framework that optimizes scales of ciphertext reflecting their rescaling levels and performance impact. With a new type system that embeds the scale and rescaling level, and a new rescaling operation called downscale, HECATE makes various scale management plans, analyzes their expected performance, and finds the optimal rescaling points throughout the entire FHE application. This work implements HECATE on top of the MLIR framework with a Python frontend and shows that HECATE achieves 27% speedup over the state-of-the-art approach for various FHE applications. Yongwoo Lee 0001, Seonyeong Heo, Seonyoung Cheon, Shinnung Jeong, Changsu Kim 0004, Eunkyung Kim 0002, Hanjun Kim 0001 |
CGO | 4 |
| 2022 | RTScale: Sensitivity-Aware Adaptive Image Scaling for Real-Time Object Detection
Seonyeong Heo, Shinnung Jeong, Hanjun Kim 0001 |
ECRTS | 2 |
| 2021 | Thread-Aware Area-Efficient High-Level Synthesis Compiler for Embedded DevicesabstractIn the embedded device market, custom hardware platforms such as an application specific integrated circuit (ASIC) and a field programmable gate array (FPGA) are attractive thanks to their high performance and power efficiency. However, its huge design costs make it challenging for manufacturers to timely launch new devices. High-level synthesis (HLS) helps significantly reduce the design costs by automating the translation of service algorithms into hardware logics; however, current HLS compilers do not fit well to embedded devices as they fail to produce area-efficient solutions while supporting concurrent events from diverse peripherals such as sensors, actuators and network modules. This paper proposes a new thread-aware HLS compiler named Duro that produces area-efficient embedded devices. Duro shares commonly-invoked functions and operators across different callers and threads with a new thread-aware area cost model, and thus effectively reduces the logic size. Moreover, Duro supports a variety of device peripherals by automatically integrating peripheral controllers and interfaces as peripheral drivers. The experiment results of six embedded devices with ten peripherals demonstrate that Duro reduces the area and energy dissipation of embedded devices by 28.5% and 25.3% compared with the designs generated by the state-of-the-art HLS compiler. This work also implements FPGA prototypes of the six devices using Duro, and the measurement results show 65.3% energy saving over Raspberry Pi Zero with slightly better computation performance. Changsu Kim 0004, Shinnung Jeong, Sungjun Cho, Yongwoo Lee 0001, William Song, Youngsok Kim, Hanjun Kim 0001 |
CGO | 2 |
| 2021 | Compiler-Assisted Semantic-Aware Encryption for Efficient and Secure Serverless ComputingabstractServerless computing like Function-as-a-Service (FaaS) is attractive for IoT service providers, liberating the providers from server maintenance. Since a data processing function is executed on the cloud instead of a dedicated server in the FaaS platform, the service users send their private data in their IoT devices to the third-party cloud, taking privacy leakage risks. Homomorphic encryption (HE) can preserve the privacy by enabling encrypted data processing on the cloud, but using HE for every data item incurs large computation and communication overheads. This work proposes SelectiveCrypt, a compiler-assisted semantic-aware encryption scheme that applies different cryptographic primitives depending on the operations on each data item. SelectiveCrypt homomorphically encrypts data items if arithmetic operations are applied to the data, while SelectiveCrypt encrypts data items with a symmetric key if the data are stored in the cloud without any arithmetic operation. The SelectiveCrypt framework consists of a compiler and its runtime system. The SelectiveCrypt compiler statically analyzes the data processing, determines an appropriate cryptographic primitive for each data item, and automatically transforms arithmetic operations into the homomorphic computation. The SelectiveCrypt runtime encrypts and decrypts the data items according to the static analysis result. This work evaluates the prototype SelectiveCrypt framework with five benchmarks that reflect real-world IoT scenarios. The evaluation results show that the SelectiveCrypt framework successfully reduces response time and communication overhead by 1.59 times and 9.61 times, respectively, compared with a HE scheme. Bongjun Kim, Seonyeong Heo, Jaeho Lee 0005, Shinnung Jeong, Yongwoo Lee 0001, Hanjun Kim 0001 |
IEEE Internet Things J. | 4 |
| 2020 | Pipeline-aware Logic Deduplication in High-Level Synthesis for Post-Quantum Cryptography AlgorithmsabstractWith the technical advance of quantum computers that can solve intractable problems for conventional computers, many of the currently used public-key cryptosystems become vulnerable. Recently proposed post-quantum cryptography (PQC) is secure against both classical and quantum computers, but existing embedded systems such as smart card can not easily support the PQC algorithms due to their much larger key sizes and more complex arithmetics. To accelerate the PQC algorithms, embedded systems have to embed the PQC hardware blocks, which can lead to huge hardware design costs. Although High-Level Synthesis (HLS) helps significantly reduce the design costs, current HLS frameworks produce inefficient hardware design for the PQC algorithms in terms of area and performance. This work analyzes common features of the PQC algorithms and proposes a new pipeline-aware logic deduplication method in HLS. The proposed method shares commonly invoked logic across hardware design while considering load balancing in pipeline and resolving dynamic memory accesses. This work implements FPGA hardware design of seven PQC algorithms in the round 2 candidates from the National Institute of Standards and Technology (NIST) PQC standardization process. Compared to commercial HLS framework, the proposed method achieves an area-delay-product reduction by 34.5%. Changsu Kim 0004, Yongwoo Lee 0001, Shinnung Jeong, Wen Wang 0007, Jakub Szefer, Hanjun Kim 0001 |
FPGA | 3 |