EDBT 2026 Demo / reviewers in the wild / expert
Janghyeon Kim
dblp:320/3939
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PIMphony: Overcoming Bandwidth and Capacity Inefficiency in PIM-Based Long-Context LLM Inference SystemabstractThe expansion of long-context Large Language Models (LLMs) creates significant memory system challenges. While Processing-in-Memory (PIM) is a promising accelerator, we identify that it suffers from critical inefficiencies when scaled to long contexts: severe channel underutilization, performancelimiting I/O bottlenecks, and massive memory waste from static KV cache management. In this work, we propose PIMphony, a PIM orchestrator that systematically resolves these issues with three co-designed techniques. First, Token-Centric PIM Partitioning (TCP) ensures high channel utilization regardless of batch size. Second, Dynamic PIM Command Scheduling (DCS) mitigates the I/O bottleneck by overlapping data movement and computation. Finally, a Dynamic PIM Access (DPA) controller enables dynamic memory management to eliminate static memory waste. Implemented via an MLIR-based compiler and evaluated on a cycle-accurate simulator, PIMphony significantly improves throughput for long-context LLM inference (up to 72B parameters and 1M context length). Our evaluations show performance boosts of up to 11.3× on PIM-only systems and 8.4× on xPU+PIM systems, enabling more efficient deployment of LLMs in real-world long-context applications. Hyucksung Kwon, Kyungmo Koo, Janghyeon Kim, Woongkyu Lee, Gyeonggeun Jung, Hyungdeok Lee, Yousub Jung, Jaehan Park, Yosub Song, Byeongsu Yang, Haerang Choi, Guhyun Kim, Jongsoon Won, Woojae Shin, Gyeongcheol Shin, Yongkee Kwon, Ilkon Kim, Eui-Cheol Lim, John Kim 0001, Jungwook Choi |
HPCA | 3 |
| 2025 | Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU LimitsabstractThe expansion of context windows in large language models (LLMs) to multi-million tokens introduces severe memory and compute bottlenecks, particularly in managing the growing Key-Value (KV) cache. While Compute Express Link (CXL) enables non-eviction frameworks that offload the full KV-cache to scalable external memory, these frameworks still suffer from costly data transfers when recalling non-resident KV tokens to limited GPU memory as context lengths increase. This work proposes scalable Processing-NearMemory (PNM) for 1M-Token LLM Inference, a CXL-enabled KVcache management system that coordinates memory and computation beyond GPU limits. Our design offloads token page selection to a PNM accelerator within CXL memory, eliminating costly recalls and enabling larger GPU batch sizes. We further introduce a hybrid parallelization strategy and a steady-token selection mechanism to enhance compute efficiency and scalability. Implemented atop a state-of-the-art CXL-PNM system, our solution delivers consistent performance gains for LLMs with up to 405B parameters and 1Mtoken contexts. Our PNM-only offloading scheme (PNM-KV) and GPU-PNM hybrid with steady-token execution (PnG-KV) achieve up to $21.9 \times$ throughput improvement, up to $60 \times$ lower energy per token, and up to $7.3 \times$ better total cost efficiency than the baseline, demonstrating that CXL-enabled multi-PNM architectures can serve as a scalable backbone for future long-context LLM inference. Janghyeon Kim, Hyucksung Kwon, Hyeonggyu Jeong, Sang-Soo Park, Minyong Yoon, Si-Dong Roh, Yongsuk Kwon, Jinin So, Jungwook Choi |
PACT | 3 |
| 2023 | Range-Invariant Approximation of Non-Linear Operations for Efficient BERT Fine-TuningabstractThis paper proposes a range-invariant approximation of non-linear operations for training computations of Transformer-based large language models. The proposed method decomposes the approximation into the scaling and the range-invariant resolution for LUT approximation, covering diverse data ranges of non-linear operations with drastically reduced LUT entries during task-dependent BERT fine-tuning. We demonstrate that the proposed method robustly approximates all the non-linear operations of BERT without score degradation on challenging GLUE benchmarks using only a single-entry LUT, facilitating 52% area savings in hardware implementation. Janghyeon Kim, Janghwan Lee, Jungwook Choi |
DAC | 1 |
| 2023 | An Approach to Design a Biomechanically-Inspired Reward Function to Solve a Patience Cube Under Reinforcement Learning FrameworkabstractThis paper presents an approach to design a reward function by adopting both control theoretic and biomechanical perspectives. In reinforcement learning (RL), a reward function plays a crucial role for an RL agent training; especially, a task learning time and a task performance. Accordingly, designing a reward function becomes a key issue to train an RL agent generating human-like policy/strategy to perform dexterous manipulation. Since human beings are good at producing heuristic approaches to complete a given task, determining a set of basis functions as well as corresponding weights used not to be so straightforward. In this study, we consider solving a patience cube as an example of a dexterous manipulation task. In our approach, we first employed a quadratic regulator form as a backbone of a desired reward function. Next, the kinematic data of a controlled object and the sEMG data of a human expert were measured while performing a demonstration to solve a patience cube. Then, from the measured data, the weights of the basis functions were determined by utilizing muscle synergy extraction and inverse optimal control as two key tools. Finally, an RL agent was trained by the designed reward function and comparative analysis versus the other RL agents trained by prototypical weight settings was followed. The result showed that the RL agent trained by our approach yielded human-like learning curve as well as policy successfully and outperformed the others in terms of a task success rate and a task completion time. These findings substantiated the feasibility of extending our approach to an assistive robotic manipulator or prosthesis design to perform the activities of daily living. Janghyeon Kim, Ho Jin Jung, Dae Han Sim, Ji-Hyeon Yoo, Song Woo Kim |
IROS | 1 |