Hyojin Sung

dblp:37/7474 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-3036-6180ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 3 first-author · 9 since 2021Software engineering, systems software and programming languages · 9 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 FORTE: Online DataFrame Query Optimizer
abstract
DataFrame libraries are widely adopted in data science for their flexible, Pythonic interfaces, but their fragmented APIs and unstructured query patterns limit systematic optimization. Existing work has explored parallel execution or SQL-style logical rewrites, yet these approaches fall short in capturing DataFrame-specific semantics and Python control-flow context. We present FORTE, the first online, source-to-source query optimizer that unifies multiple DataFrame libraries under a shared intermediate representation (DFL). DFL makes DataFrame semantics explicit, enabling composable and portable rewriting rules such as user-defined function (UDF) lifting/lowering, loop lifting, and API tuning, alongside classical rewrites (e.g., predicate pushdown). FORTE employs a lightweight, learned cost model and greedy search to apply these rewrites with negligible overhead, while supporting both intra-library optimization and cross-library transpilation. Our evaluation on TPC-H workloads and real-world Kaggle/GitHub workloads shows that FORTE consistently delivers substantial speedups—up to 52.53× (3.7× on average) across Pandas, Modin, Polars, and Pandas-on-Spark—demonstrating that online, IR-guided rewriting can significantly outperform existing DataFrame engines and rewriters, while enabling cross-library retargetability.
Yoonho Choi, Kyoungtae Lee, Hyungsoo Jung 0003, Hyojin Sung
CGO5
2026 Aurora: A Disaggregated GPU-PNM-PIM System for High-Throughput Mixed-Length LLM Inference
Hyeonu Kim, Seunghyuk Yu, Jungsul Lee, Hyojin Sung, Eojin Lee
ICS5
2026 RAPO: Retrieval-Augmented Phase Ordering
abstract
Finding a high-quality ordering of compiler optimization passes for a given program, known as the phase-ordering problem, is NP-hard. Recent reinforcement learning (RL) methods can discover effective pass sequences, but their costly per-input exploration limits accessibility and practical deployment. We introduce RAPO, a retrieval-augmented phase-ordering framework that replaces online RL exploration with similarity-based retrieval. Offline, RAPO embeds LLVM IR using IR-BERT, clusters programs with k-means, and stores representative RL-discovered pass sequences for each cluster in a sequence cache. During compilation, RAPO embeds a new program, assigns it to the nearest cluster, and optimizes it by applying the corresponding cached sequence, thereby transforming per-program search into similarity-based retrieval. RAPO is model-agnostic, supporting pass sequences produced by different RL exploration strategies, and includes lightweight fallbacks for corner cases. RAPO reduces IR instruction count by up to ∼18.6% relative to -Oz, matching or outperforming RL baselines while reducing phase-ordering search overhead by up to ∼177×. These results suggest that RAPO achieves near-per-input optimization quality with deployment-grade efficiency by turning online phase ordering into fast, similarity-driven retrieval.
Jinwook Yang, Yeonsun-Hong, Hyojin Sung
LCTES4
2025 HYPERF: End-to-End Autotuning Framework for High-Performance Computing
Juseong Park, Yongwon Shin, Oh-Kyoung Kwon, Hyojin Sung
HPDC7
2025 ATiM: Autotuning Tensor Programs for Processing-in-DRAM
abstract
Processing-in-DRAM (DRAM-PIM) has emerged as a promising technology for accelerating memory-intensive operations in modern applications, such as Large Language Models (LLMs).Despite its potential, current software stacks for DRAM-PIM face significant challenges, including reliance on hand-tuned libraries that hinder programmability, limited support for high-level abstractions, and the lack of systematic optimization frameworks.To address these limitations, we present ATiM, a search-based optimizing tensor compiler for UPMEM.Key features of ATiM include: (1) automated searches of the joint search space for host and kernel tensor programs, (2) PIM-aware optimizations for efficiently handling boundary conditions, and ( 3) improved search algorithms for the expanded search space of UPMEM systems.Our experimental results on UPMEM hardware demonstrate performance gains of up to 6.18× for various UPMEM benchmark kernels and 8.21× for GPT-J layers.To the best of our knowledge, ATiM is the first tensor compiler to provide fully automated, autotuning-integrated code generation support for a DRAM-PIM system.By bridging the gap between high-level tensor computation abstractions and low-level hardware-specific requirements, ATiM establishes a foundation for advancing DRAM-PIM programmability and enabling streamlined optimization.
Yongwon Shin, Dookyung Kang, Hyojin Sung
ISCA3
2024 NavCim: Comprehensive Design Space Exploration for Analog Computing-in-Memory Architectures
abstract
Analog computing-in-memory (ACiM) technology has shown strong potential for neural network accelerators, addressing von-Neumann performance bottlenecks with in-memory data processing and computation. Understanding the ACiM design space, including its trade-offs and constraints, and systematically and effectively exploring it for optimal performance is essential to turn the promise into a viable product. Recent research demonstrated that multi-objective searches for ACiM architectures with heterogeneous tiles can simultaneously optimize power, performance, and area (PPA), outperforming existing tiled ACiM proposals. In this paper, we propose NavCim, a comprehensive ACiM design space exploration mechanism that advances the prior work in terms of search efficiency, search space coverage, and optimization metrics. NavCim introduces predictive modeling of ACiM hardware performance and uses the PPA prediction models instead of running simulators, significantly reducing search overheads. Faster searches enable NavCim to extend the architecture and model search spaces with an evolutionary search process to optimize architectures with more than two different tile sizes for multiple input models. With accuracy-aware searches, NavCim considers PPA and model accuracy together as optimization goals to achieve more balanced trade-offs. The experimental searches show that NavCim leverages predictive models to reduce search time by up to 7.3x without compromising the quality of search results. It also successfully identifies heterogeneous ACiM architectures that can efficiently execute multiple models on a single chip, improving accuracy by up to 19% over the prior work.
Juseong Park, Boseok Kim, Hyojin Sung
PACT3
2024 Low-Overhead General-Purpose Near-Data Processing in CXL Memory Expanders
abstract
Emerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXL.mem protocol provides minimal latency overhead through an optimized protocol stack, frequent CXL memory accesses can result in significant slowdowns for memory-bound applications whether they are latency-sensitive or bandwidth-intensive. The near-data processing (NDP) in the CXL controller promises to overcome such limitations of passive CXL memory. However, prior work on NDP in CXL memory proposes application-specific units that are not suitable for practical CXL memory-based systems that should support various applications. On the other hand, existing CPU or GPU cores are not cost-effective for NDP because they are not optimized for memory-bound applications. In addition, the communication between the host processor and CXL controller for NDP offloading should achieve low latency, but existing CXL.io/PCIe-based mechanisms incur$\mu\mathbf{s}-\mathbf{scale}$latency and are not suitable for fine-grained NDP.
Hyungkyu Ham, Jeongmin Hong 0001, Geonwoo Park, Yunseon Shin, Okkyun Woo, Wonhyuk Yang, Jinhoon Bae, Eunhyeok Park, Hyojin Sung, Eui-Cheol Lim, Gwangsun Kim
MICRO9
2023 PIMFlow: Compiler and Runtime Support for CNN Models on Processing-in-Memory DRAM
abstract
Processing-in-Memory (PIM) has evolved over decades into a feasible solution to addressing the exacerbating performance bottleneck with main memory by placing computational logic in or near memory. Recent proposals from DRAM manufacturers highlighted the HW constraint-aware design of PIM-enabled DRAM with specialized MAC logic, providing an order of magnitude speedup for memory-intensive operations in DL models. Although the main target for PIM acceleration did not initially include convolutional neural networks due to their high compute intensity, recent CNN models are increasingly adopting computationally lightweight implementation. Motivated by the potential for the software stack to enable CNN models on DRAM-PIM hardware without invasive changes, we propose PIMFlow, an end-to-end compiler and runtime support, to accelerate CNN models on a PIM-enabled GPU memory. PIMFlow transforms model graphs to create inter-node parallelism across GPU and PIM, explores possible task- and data-parallel execution scenarios for optimal execution time, and provides a code-generating back-end and execution engine for DRAM-PIM. PIMFlow achieves up to 82% end-to-end speedup and reduces energy consumption by 26% on average for CNN model inferences.
Yongwon Shin, Juseong Park, Sungjun Cho, Hyojin Sung
CGO4
2023 PRIMO: A Full-Stack Processing-in-DRAM Emulation Framework for Machine Learning Workloads
abstract
Recently, the size of deep learning models has significantly increased, making the excessive memory access between the AI processor and DRAM a major bottleneck of the system. The processing-in-DRAM (DRAM-PIM) concept has emerged as a promising solution, which integrates computing logic within memory, thus saving abundant access to external memory. Although many simulators have been proposed to model and analyze the benefits of DRAM-PIM, they are often too slow to run an entire application. FPGA-based emulators have been introduced to overcome this limitation. However, none of the prior works include the full software stack from the model to DRAM-PIM hardware. This paper presents a full-stack processing-in-DRAM emulation framework named PRIMO, the first emulation framework that can model and analyze DRAM-PIM for end-to-end ML inference. PRIMO enables software developers to develop and test their customized software stacks on various ML workloads without requiring a real DRAM-PIM chip. Moreover, it allows designers to explore design space and monitor memory access patterns, facilitating software and hardware co-design for efficient DRAM-PIM architectures. To achieve these goals, we develop a real-time FPGA emulator that emulates DRAM-PIM architecture and generates experimental results such as predicted cycle information and computed output at incomparably high speeds compared to the CPU-based simulation. In addition, we propose a software stack comprising a PIM compiler that enables the execution of various ML workloads, including end-to-end inference, and a PIM driver that runs the workloads with high bandwidth utilization by leveraging virtual memory scatter-gather DMA. Finally, we demonstrate that PRIMO can successfully emulate DRAM-PIM 106.64-6093.56× faster than the CPU-based simulation framework for ML workloads ranging from small microbenchmarks to end-to-end inference of ResNets.
Jaehoon Heo, Yongwon Shin, Sangjin Choi, Sungwoong Yune, Hyojin Sung, Youngjin Kwon, Joo-Young Kim 0001
ICCAD6
2023 Multi-Objective Architecture Search and Optimization for Heterogeneous Neuromorphic Architecture
abstract
Neuro-inspired in-memory computing offers a promising solution to overcome the limitations of traditional von Neumann architectures by emulating brain activities. This approach takes advantage of parallel processing while minimizing power consumption and area overhead. However, maximizing the performance of neuromorphic hardware, particularly for increasingly deep and complex NN models, is a challenging task. Existing design space exploration methods focus primarily on layer placement and resource allocation, ignoring hardware-level configurations that directly influence performance, power, and area (PPA) trade-offs. Additionally, current tiled neuromorphic architectures lack support for size heterogeneity, making optimal resource utilization difficult. To address these challenges, we propose a multi-objective architecture search and optimization mechanism for neuromorphic architectures. Our approach introduces heterogeneous architectures with multiple tile/processing element/synaptic array sizes and provides a comprehensive end-to-end design automation tool to support them. We define a heterogeneous neuromorphic architecture, as exemplified by a “big-tile, little-tile” architecture with mesh interconnects. Our search mechanism expands the search space by considering candidates for both homogeneous and heterogeneous architectures and performs Pareto-front searches guided by user-defined weights or constraints on performance metrics. We also implement a hierarchical beam search technique to explore the vast search space of heterogeneous architecture candidates more effectively. Our mechanism identifies numerous heterogeneous architectures that outperform the baseline for different convolutional neural network (CNN) models. For EfficientNetB0, we achieve PPA improvements of 40.1%, 19.3%, and 4.4% over the baseline. Our tool is available at https://github.com/wntjd9805/hetero-neurosim-search.
Juseong Park, Yongwon Shin, Hyojin Sung
ICCAD3
2022 One-shot tuner for deep learning compilers
abstract
Auto-tuning DL compilers are gaining ground as an optimizing back-end for DL frameworks. While existing work can generate deep learning models that exceed the performance of hand-tuned libraries, they still suffer from prohibitively long auto-tuning time due to repeated hardware measurements in large search spaces. In this paper, we take a neural-predictor inspired approach to reduce the auto-tuning overhead and show that a performance predictor model trained prior to compilation can produce optimized tensor operation codes without repeated search and hardware measurements. To generate a sample-efficient training dataset, we extend input representation to include task-specific information and to guide data sampling methods to focus on learning high-performing codes. We evaluated the resulting predictor model, One-Shot Tuner, against AutoTVM and other prior work, and the results show that One-Shot Tuner speeds up compilation by 2.81x to 67.7x compared to prior work while providing comparable or improved inference time for CNN and Transformer models.
Jaehun Ryu, Eunhyeok Park, Hyojin Sung
CC3
2019 POSTER: CogR: Exploiting Program Structures for Machine-Learning Based Runtime Solutions
abstract
We propose CogR, a machine-learning based runtime solution, that enables efficient and dynamic resource scheduling and performance optimization for high-level programming interfaces on heterogeneous systems. CogR tightly combines the structural information of programs and fine-grained static and dynamic statistics into sequenced input data. This structural and value-embedded representation of programs enables CogR to accurately model the runtime behaviors of nested loop-based constructs in the high-level parallel programs. The end-to-end CogR system consists of compiler and runtime support for feature collection and input generation, a machine learning model, and a runtime scheduler with online inference and prediction. The system provides 11% higher prediction accuracy than models simulated for prior work and improves kernel performance by 66% compared to the baseline runtime.
Hyojin Sung, Tong Chen 0001, Alexandre E. Eichenberger, Kevin O'Brien
PACT1
2017 Efficient Fork-Join on GPUs Through Warp Specialization
abstract
Graphics Processing Units (GPUs) are increasingly used to accelerate portions of general-purpose applications. Higher level language extensions have been proposed to help non-experts bridge the gap between a host and the GPU's threading model. Recent updates to the OpenMP standard allow a user to parallelize code on a GPU using the well known fork-join programming model for CPUs. Mapping this model to the architecturally visible threading model of typical GPUs has been challenging. In this work we propose a novel approach using the technique of Warp Specialization. We show how to specialize one warp (a unit of 32 GPU threads) to handle sequential code on a GPU. When this master warp reaches a user-specified parallel region, it awakens unused GPU warps to collectively execute the parallel code. Based on this method, we have implemented a Clang-based, OpenMP 4.5 compliant, open source compiler for GPUs. Our work achieves a 3.6x (and up to 32x) performance improvement over a baseline that does not exploit fork-join parallelism on an NVIDIA k40m GPU across a set of 25 kernels. Compared to state-of-the-art compilers (Clang-ykt, GCC-OpenMP, GCC-OpenACC) our work is 2.1 - 7.6x faster. Our proposed technique is simpler to implement, robust, and performant.
Arpith C. Jacob, Alexandre E. Eichenberger, Hyojin Sung, Samuel Antão, Gheorghe-Teodor Bercea, Carlo Bertolli, Alexey Bataev, Tong Chen 0001, Zehra Sura, Georgios Rokos, Kevin O'Brien
HiPC3
2015 DeNovoSync: Efficient Support for Arbitrary Synchronization without Writer-Initiated Invalidations
abstract
Current shared-memory hardware is complex and inefficient. Prior work on the DeNovo coherence protocol showed that disciplined shared-memory programming models can enable more complexity-, performance-, and energy-efficient hardware than the state-of-the-art MESI protocol. DeNovo, however, severely restricted the synchronization constructs an application can support. This paper proposes DeNovoSync, a technique to support arbitrary synchronization in DeNovo. The key challenge is that DeNovo exploits race-freedom to use reader-initiated local self-invalidations (instead of conventional writer-initiated remote cache invalidations) to ensure coherence. Synchronization accesses are inherently racy and not directly amenable to self-invalidations. DeNovoSync addresses this challenge using a novel combination of registration of all synchronization reads with a judicious hardware backoff to limit unnecessary registrations. For a wide variety of synchronization constructs and applications, compared to MESI, DeNovoSync shows comparable or up to 22% lower execution time and up to 58% lower network traffic, enabling DeNovo's advantages for a much broader class of software than previously possible.
Hyojin Sung, Sarita V. Adve
ASPLOS1
2015 Eliminating on-chip traffic waste: are we there yet?
abstract
While many techniques have been shown to be successful at reducing the amount of on-chip network traffic, no studies have shown how close a combined approach would come to eliminating all unnecessary data traffic, nor have any studies provided insight into where the remaining challenges are. This paper systematically analyzes the traffic inefficiencies of a directory-based MESI protocol and a more efficient hardware-software co-designed protocol, DeNovo. We categorize data waste into various categories and explore several simple optimizations extending DeNovo with the aim of eliminating all of the on-chip network traffic waste. With all the proposed optimizations, we are able to completely eliminate (100%) onchip network traffic waste at L2 for some of the applications (93.5% on average) compared to the previous DeNovo protocol.
Robert Smolinski, Rakesh Komuravelli, Hyojin Sung, Sarita V. Adve
ISPASS3
2013 DeNovoND: efficient hardware support for disciplined non-determinism
abstract
Recent work has shown that disciplined shared-memory programming models that provide deterministic-by-default semantics can simplify both parallel software and hardware. Specifically, the DeNovo hardware system has shown that the software guarantees of such models (e.g., data-race-freedom and explicit side-effects) can enable simpler, higher performance, and more energy-efficient hardware than the current state-of-the-art for deterministic programs. Many applications, however, contain non-deterministic parts; e.g., using lock synchronization. For commercial hardware to exploit the benefits of DeNovo, it is therefore necessary to extend DeNovo to support non-deterministic applications.
Hyojin Sung, Rakesh Komuravelli, Sarita V. Adve
ASPLOS1
2011 DeNovo: Rethinking the Memory Hierarchy for Disciplined Parallelism
abstract
For parallelism to become tractable for mass programmers, shared-memory languages and environments must evolve to enforce disciplined practices that ban "wild shared-memory behaviors;'' e.g., unstructured parallelism, arbitrary data races, and ubiquitous non-determinism. This software evolution is a rare opportunity for hardware designers to rethink hardware from the ground up to exploit opportunities exposed by such disciplined software models. Such a co-designed effort is more likely to achieve many-core scalability than a software-oblivious hardware evolution. This paper presents DeNovo, a hardware architecture motivated by these observations. We show how a disciplined parallel programming model greatly simplifies cache coherence and consistency, while enabling a more efficient communication and cache architecture. The DeNovo coherence protocol is simple because it eliminates transient states - verification using model checking shows 15X fewer reachable states than a state-of-the-art implementation of the conventional MESI protocol. The DeNovo protocol is also more extensible. Adding two sophisticated optimizations, flexible communication granularity and direct cache-to-cache transfers, did not introduce additional protocol states (unlike MESI). Finally, DeNovo shows better cache hit rates and network traffic, translating to better performance and energy. Overall, a disciplined shared-memory programming model allows DeNovo to seamlessly integrate message passing-like interactions within a global address space for improved design complexity, performance, and efficiency.
Byn Choi, Rakesh Komuravelli, Hyojin Sung, Robert Smolinski, Nima Honarmand, Sarita V. Adve, Vikram S. Adve, Nicholas P. Carter, Ching-Tsun Chou
PACT3
2009 A type and effect system for deterministic parallel Java
abstract
Today's shared-memory parallel programming models are complex and error-prone.While many parallel programs are intended to be deterministic, unanticipated thread interleavings can lead to subtle bugs and nondeterministic semantics. In this paper, we demonstrate that a practical type and effect system can simplify parallel programming by guaranteeing deterministic semantics with modular, compile-time type checking even in a rich, concurrent object-oriented language such as Java. We describe an object-oriented type and effect system that provides several new capabilities over previous systems for expressing deterministic parallel algorithms.We also describe a language called Deterministic Parallel Java (DPJ) that incorporates the new type system features, and we show that a core subset of DPJ is sound. We describe an experimental validation showing thatDPJ can express a wide range of realistic parallel programs; that the new type system features are useful for such programs; and that the parallel programs exhibit good performance gains (coming close to or beating equivalent, nondeterministic multithreaded programs where those are available).
Robert L. Bocchino Jr., Vikram S. Adve, Danny Dig, Sarita V. Adve, Stephen Heumann, Rakesh Komuravelli, Jeffrey Overbey, Patrick Simmons, Hyojin Sung, Mohsen Vakilian
OOPSLA9