EDBT 2026 Demo / reviewers in the wild / expert
Yongwon Shin
dblp:97/5931
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-0481-172XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HYPERF: End-to-End Autotuning Framework for High-Performance Computing
Juseong Park, Yongwon Shin, Oh-Kyoung Kwon, Hyojin Sung |
HPDC | 2 |
| 2025 | ATiM: Autotuning Tensor Programs for Processing-in-DRAMabstractProcessing-in-DRAM (DRAM-PIM) has emerged as a promising technology for accelerating memory-intensive operations in modern applications, such as Large Language Models (LLMs).Despite its potential, current software stacks for DRAM-PIM face significant challenges, including reliance on hand-tuned libraries that hinder programmability, limited support for high-level abstractions, and the lack of systematic optimization frameworks.To address these limitations, we present ATiM, a search-based optimizing tensor compiler for UPMEM.Key features of ATiM include: (1) automated searches of the joint search space for host and kernel tensor programs, (2) PIM-aware optimizations for efficiently handling boundary conditions, and ( 3) improved search algorithms for the expanded search space of UPMEM systems.Our experimental results on UPMEM hardware demonstrate performance gains of up to 6.18× for various UPMEM benchmark kernels and 8.21× for GPT-J layers.To the best of our knowledge, ATiM is the first tensor compiler to provide fully automated, autotuning-integrated code generation support for a DRAM-PIM system.By bridging the gap between high-level tensor computation abstractions and low-level hardware-specific requirements, ATiM establishes a foundation for advancing DRAM-PIM programmability and enabling streamlined optimization. Yongwon Shin, Dookyung Kang, Hyojin Sung |
ISCA | 1 |
| 2023 | PIMFlow: Compiler and Runtime Support for CNN Models on Processing-in-Memory DRAMabstractProcessing-in-Memory (PIM) has evolved over decades into a feasible solution to addressing the exacerbating performance bottleneck with main memory by placing computational logic in or near memory. Recent proposals from DRAM manufacturers highlighted the HW constraint-aware design of PIM-enabled DRAM with specialized MAC logic, providing an order of magnitude speedup for memory-intensive operations in DL models. Although the main target for PIM acceleration did not initially include convolutional neural networks due to their high compute intensity, recent CNN models are increasingly adopting computationally lightweight implementation. Motivated by the potential for the software stack to enable CNN models on DRAM-PIM hardware without invasive changes, we propose PIMFlow, an end-to-end compiler and runtime support, to accelerate CNN models on a PIM-enabled GPU memory. PIMFlow transforms model graphs to create inter-node parallelism across GPU and PIM, explores possible task- and data-parallel execution scenarios for optimal execution time, and provides a code-generating back-end and execution engine for DRAM-PIM. PIMFlow achieves up to 82% end-to-end speedup and reduces energy consumption by 26% on average for CNN model inferences. Yongwon Shin, Juseong Park, Sungjun Cho, Hyojin Sung |
CGO | 1 |
| 2023 | PRIMO: A Full-Stack Processing-in-DRAM Emulation Framework for Machine Learning WorkloadsabstractRecently, the size of deep learning models has significantly increased, making the excessive memory access between the AI processor and DRAM a major bottleneck of the system. The processing-in-DRAM (DRAM-PIM) concept has emerged as a promising solution, which integrates computing logic within memory, thus saving abundant access to external memory. Although many simulators have been proposed to model and analyze the benefits of DRAM-PIM, they are often too slow to run an entire application. FPGA-based emulators have been introduced to overcome this limitation. However, none of the prior works include the full software stack from the model to DRAM-PIM hardware. This paper presents a full-stack processing-in-DRAM emulation framework named PRIMO, the first emulation framework that can model and analyze DRAM-PIM for end-to-end ML inference. PRIMO enables software developers to develop and test their customized software stacks on various ML workloads without requiring a real DRAM-PIM chip. Moreover, it allows designers to explore design space and monitor memory access patterns, facilitating software and hardware co-design for efficient DRAM-PIM architectures. To achieve these goals, we develop a real-time FPGA emulator that emulates DRAM-PIM architecture and generates experimental results such as predicted cycle information and computed output at incomparably high speeds compared to the CPU-based simulation. In addition, we propose a software stack comprising a PIM compiler that enables the execution of various ML workloads, including end-to-end inference, and a PIM driver that runs the workloads with high bandwidth utilization by leveraging virtual memory scatter-gather DMA. Finally, we demonstrate that PRIMO can successfully emulate DRAM-PIM 106.64-6093.56× faster than the CPU-based simulation framework for ML workloads ranging from small microbenchmarks to end-to-end inference of ResNets. Jaehoon Heo, Yongwon Shin, Sangjin Choi, Sungwoong Yune, Hyojin Sung, Youngjin Kwon, Joo-Young Kim 0001 |
ICCAD | 2 |
| 2023 | Multi-Objective Architecture Search and Optimization for Heterogeneous Neuromorphic ArchitectureabstractNeuro-inspired in-memory computing offers a promising solution to overcome the limitations of traditional von Neumann architectures by emulating brain activities. This approach takes advantage of parallel processing while minimizing power consumption and area overhead. However, maximizing the performance of neuromorphic hardware, particularly for increasingly deep and complex NN models, is a challenging task. Existing design space exploration methods focus primarily on layer placement and resource allocation, ignoring hardware-level configurations that directly influence performance, power, and area (PPA) trade-offs. Additionally, current tiled neuromorphic architectures lack support for size heterogeneity, making optimal resource utilization difficult. To address these challenges, we propose a multi-objective architecture search and optimization mechanism for neuromorphic architectures. Our approach introduces heterogeneous architectures with multiple tile/processing element/synaptic array sizes and provides a comprehensive end-to-end design automation tool to support them. We define a heterogeneous neuromorphic architecture, as exemplified by a “big-tile, little-tile” architecture with mesh interconnects. Our search mechanism expands the search space by considering candidates for both homogeneous and heterogeneous architectures and performs Pareto-front searches guided by user-defined weights or constraints on performance metrics. We also implement a hierarchical beam search technique to explore the vast search space of heterogeneous architecture candidates more effectively. Our mechanism identifies numerous heterogeneous architectures that outperform the baseline for different convolutional neural network (CNN) models. For EfficientNetB0, we achieve PPA improvements of 40.1%, 19.3%, and 4.4% over the baseline. Our tool is available at https://github.com/wntjd9805/hetero-neurosim-search. Juseong Park, Yongwon Shin, Hyojin Sung |
ICCAD | 2 |
| 2004 | Low-complexity predictive trellis coded quantization of wideband speech LSF parametersabstractIn this paper, low-complexity block-constrained trellis coded quantization (BC-TCQ) structures are introduced, and a predictive BC-TCQ encoding method is developed for quantization of line spectrum frequencies (LSF) parameters for wideband speech coding applications. The performance is compared to the linear predictive coding (LPC) vector quantizers used in the AMR-WB (ITU-G.722.2) speech coding standard, demonstrating reduction in spectral distortion and significant reduction in encoding complexity. Yongwon Shin, Sangwon Kang, Thomas R. Fischer, Changyong Son, Yongbeom Lee |
ICASSP (1) | 1 |