VLDB 2026 Research / reviewers in the wild / expert
Jonghyun Bae
dblp:119/3995
· DBLP profile ↗
12ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalable training of trustworthy and energy-efficient predictive graph foundation models for atomistic materials modeling: a case study with HydraGNNabstractWe present our work on developing and training scalable, trustworthy, and energy-efficient predictive graph foundation models (GFMs) using HydraGNN, a multi-headed graph convolutional neural network architecture. HydraGNN expands the boundaries of graph neural network (GNN) computations in both training scale and data diversity. It abstracts over message passing algorithms, allowing both reproduction of and comparison across algorithmic innovations that define nearest-neighbor convolution in GNNs. This work discusses a series of optimizations that have allowed scaling up the GFMs training to tens of thousands of GPUs on datasets consisting of hundreds of millions of graphs. Our GFMs use multitask learning (MTL) to simultaneously learn graph-level and node-level properties of atomistic structures, such as energy and atomic forces. Using over 154 million atomistic structures for training, we illustrate the performance of our approach along with the lessons learned on two state-of-the-art US Department of Energy (US-DOE) supercomputers, namely the Perlmutter petascale system at the National Energy Research Scientific Computing Center and the Frontier exascale system at Oak Ridge Leadership Computing Facility. The HydraGNN architecture enables the GFM to achieve near-linear strong scaling performance using more than 2000 GPUs on Perlmutter and 16,000 GPUs on Frontier. Massimiliano Lupo Pasini, Jong Choi 0001, Kshitij Mehta, David M. Rogers 0001, Jonghyun Bae, Khaled Z. Ibrahim, Ashwin M. Aji, Karl W. Schulz, Jorda Polo, Prasanna Balaprakash |
J. Supercomput. | 6 |
| 2024 | ISP2DLA: Automated Deep Learning Accelerator Design for On-Sensor Image Signal ProcessingabstractDeep neural network-based image signal processing (ISP-DNN) improves image quality with techniques such as demosaicing, but these models pose substantial computational and memory challenges when implemented on CMOS image sensors, particularly due to the high-resolution inputs that increase memory requirements for activations. Layer fusion reduces memory usage by combining consecutive processing steps, yet it increases computational demands, a critical issue in resource-limited on-sensor environments. To address these challenges, we introduce ISP2DLA, an automated deep learning accelerator design framework that balances computational and memory demands for on-sensor ISP. This framework optimizes hardware designs by adjusting line buffer sizes and the number of MAC units, reducing gate counts by 14-79% across two ISP-DNN models, thus enabling efficient on-sensor ISP model inference within constrained resources. Dong-eon Won, Yeeun Kim, Janghwan Lee, Jonghyun Bae, Jongjoo Park, Jeongyong Song, Jungwook Choi |
ASAP | 5 |
| 2022 | L3: Accelerator-Friendly Lossless Image Format for High-Resolution, High-Throughput DNN Training
Jonghyun Bae, Woohyeon Baek, Tae Jun Ham, Jae W. Lee |
ECCV (11) | 1 |
| 2021 | FlashNeuron: SSD-Enabled Large-Batch Training of Very Deep Neural Networks
Jonghyun Bae, Jongsung Lee 0001, Yunho Jin, Sam Son, Shine Kim, Hakbeom Jang, Tae Jun Ham, Jae W. Lee |
FAST | 1 |
| 2021 | Behemoth: A Flash-centric Training Accelerator for Extreme-scale DNNs
Shine Kim, Yunho Jin, Gina Sohn, Jonghyun Bae, Tae Jun Ham, Jae W. Lee |
FAST | 4 |
| 2021 | Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise SchedulingabstractTo meet surging demands for deep learning inference services, many cloud computing vendors employ high-performance specialized accelerators, called neural processing units (NPUs). One important challenge for effective use of NPUs is to achieve high resource utilization over a wide spectrum of deep neural network (DNN) models with diverse arithmetic intensities. There is often an intrinsic mismatch between the compute-to-memory bandwidth ratio of an NPU and the arithmetic intensity of the model it executes, leading to under-utilization of either compute resources or memory bandwidth. Ideally, we want to saturate both compute TOP/s and DRAM bandwidth to achieve high system throughput. Thus, we propose Layerweaver, an inference serving system with a novel multi-model time-multiplexing scheduler for NPUs. Layerweaver reduces the temporal waste of computation resources by interweaving layer execution of multiple different models with opposing characteristics: compute-intensive and memory-intensive. Layerweaver hides the memory time of a memory-intensive model by overlapping it with the relatively long computation time of a compute-intensive model, thereby minimizing the idle time of the computation units waiting for off-chip data transfers. For a two-model serving scenario of batch 1 with 16 different pairs of compute- and memory-intensive models, Layerweaver improves the temporal utilization of computation units and memory channels by 44.0% and 28.7%, respectively, to increase the system throughput by 60.1% on average, over the baseline executing one model at a time. Young H. Oh, Seonghak Kim, Yunho Jin, Sam Son, Jonghyun Bae, Jongsung Lee 0001, Yeonhong Park, Dong Uk Kim, Tae Jun Ham, Jae W. Lee |
HPCA | 5 |
| 2021 | ASAP: Fast Mobile Application Switch via Adaptive Prepaging
Sam Son, Seung Yul Lee, Yunho Jin, Jonghyun Bae, Jinkyu Jeong, Tae Jun Ham, Jae W. Lee, Hongil Yoon |
USENIX ATC | 4 |
| 2020 | A Case for Hardware-Based Demand PagingabstractThe virtual memory system is pervasive in today's computer systems, and demand paging is the key enabling mechanism for it. At a page miss, the CPU raises an exception, and the page fault handler is responsible for fetching the requested page from the disk. The OS typically performs a context switch to run other threads as traditional disk access is slow. However, with the widespread adoption of high-performance storage devices, such as low-latency solid-state drives (SSDs), the traditional OS-based demand paging is no longer effective because a considerable portion of the demand paging latency is now spent inside the OS kernel. Thus, this paper makes a case for hardware-based demand paging that mostly eliminates OS involvement in page miss handling to provide a near-disk-access-time latency for demand paging. To this end, two architectural extensions are proposed: LBA-augmented page table that moves I/O stack operations to the control plane and Storage Management Unit that enables CPU to directly issue I/O commands without OS intervention in most cases. OS support is also proposed to detach tasks for memory resource management from the critical path. The evaluation results using both a cycle-level simulator and a real x86 machine with an ultra-low latency SSD show that the proposed scheme reduces the demand paging latency by 37.0%, and hence improves the performance of FIO read random benchmark by up to 57.1% and a NoSQL server by up to 27.3% with real-world workloads. As a side effect of eliminating OS intervention, the IPC of the user-level code is also increased by up to 7.0%. Gyusun Lee, Wenjing Jin 0001, Wonsuk Song, Jeonghun Gong, Jonghyun Bae, Tae Jun Ham, Jae W. Lee, Jinkyu Jeong |
ISCA | 5 |
| 2019 | Practical Erase Suspension for Modern Low-latency SSDs
Shine Kim, Jonghyun Bae, Hakbeom Jang, Wenjing Jin 0001, Jeonghun Gong, SeungYeon Lee, Tae Jun Ham, Jae W. Lee |
USENIX ATC | 2 |
| 2017 | Jointly optimizing task granularity and concurrency for in-memory mapreduce frameworksabstractRecently, in-memory big data processing frameworks have emerged, such as Apache Spark and Ignite, to accelerate workloads requiring frequent data reuse. With effective in-memory caching these frameworks eliminate most of I/O operations, which would otherwise be necessary for communication between producer and consumer tasks. However, this performance benefit is nullified if the memory footprint exceeds available memory size, due to excessive spill and garbage collection (GC) operations. To fit the working set in memory, two system parameters play an important role: number of data partitions (Npartitions) specifying task granularity, and number of tasks per each executor (Nthreads) specifying the degree of parallelism in execution. Existing approaches to optimizing these parameters either do not take into account workload characteristics, or optimize only one of the parameters in isolation, thus yielding suboptimal performance. This paper introduces WASP, a workload-aware task scheduler and partitioner, which jointly optimizes both parameters at runtime. To find an optimal setting, WASP first analyzes the DAG structure of a given workload, and uses an analytical model to predict optimal settings of Npartitionsand Nthreadsfor all stages based on their computation types. Taking this as input, the WASP scheduler employs a hill climbing algorithm to find an optimal Nthreadsfor each stage, thus maximizing concurrency while minimizing data spills and GCs. We prototype WASP on Spark and evaluate it using six workloads on three different parallel platforms. WASP improves performance by up to 3.22× and reduces the cluster operating cost on cloud by up to 40%, over the baseline following Spark Tuning Guidelines and provides robust performance for both shuffle-heavy and shuffle-light workloads. Jonghyun Bae, Hakbeom Jang, Wenjing Jin 0001, Jun Heo 0001, Jaeyoung Jang, Joo Young Hwang, Sangyeun Cho, Jae W. Lee |
IEEE BigData | 1 |
| 2015 | Exploring the Role of a Smartphone as a Motion Sensing and Control Device in the Wireless Networked Control of a Motor Test-bedabstractThe sensing, computing, and control potential of smartphones remains to be fully explored in automatic control applications. In this paper, we control the angular position of a motor test-bed using feedback from the embedded motion sensors of a smartphone while it is mounted to the test-bed. The smartphone hosts an interactive user interface which students and researchers can use to quickly and easily perform experiments with the test-bed and collect measurements using their own personal devices. Proportional-plus-derivative (PD) controllers designed using a sampled-data model of the system are compared for different sampling rates used on the smartphone. Results from simulations and experiments confirm the feasibility of utilizing mounted smartphones in the wireless networked control of systems with rotational degrees of freedom. Jared Alan Frank, Anthony Brill, Jonghyun Bae, Vikram Kapila |
ICINCO (2) | 3 |
| 2012 | A new edge directed interpolation algorithm using accurate estimation of edge directional covarianceabstractThis paper proposes an edge-directed interpolation algorithm to enhance the quality of natural images which are captured by low-resolution camera installed on car or CCTV. Based on the accurate estimation of edge directional covariance between low-resolution and high-resolution image, local covariance coefficients extracted from the low-resolution image has been adapted for the interpolation to obtain the high-resolution image. DCT (Discrete Cosine Transform) kernel function is used in order to reflect the multi-directional edge accurately without increasing of complexity. Simulation result shows that our new interpolation algorithm significantly improves the subjective quality of the interpolated images compared with conventional linear interpolation and NEDI one. It also demonstrates the improvements of objective metrics such as PSNR, SSIM(structural similarity index measurement) and WEA (Wiener filter coefficients Estimation Accuracy) which are used for the accuracy estimation of directionality. Jonghyun Bae, Yujin Yun, Kyungman Kim, Jaeseok Kim |
ISCAS | 1 |