EDBT 2026 Demo / reviewers in the wild / expert
Yujeong Choi
dblp:248/8230
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Reliable Code De-Obfuscation with Large Language Models
Yujeong Choi, Dohwan Ji, Yujin Kwon |
SANER | 1 |
| 2026 | eBPF-VulnBench: A Benchmark of Real-World eBPF Malicious Bytecode
Yujin Kwon, Yujeong Choi, Dohwan Ji |
SANER | 2 |
| 2025 | Hera: A Heterogeneity-Aware Multi-Tenant Inference Server for Personalized RecommendationsabstractWhile providing low latency is a fundamental requirement in deploying recommendation services, achieving high resource utility is also crucial in cost-effectively maintaining the datacenter. Co-locating model workers is an effective way to maximize query-level parallelism and server throughput, but the interference caused by concurrent workers at shared resources can prevent server queries from meeting its SLA. Hera utilizes the heterogeneous memory requirement of multi-tenant recommendation models to intelligently determine a productive set of colocated models and its resource allocation, providing fast response time while achieving high throughput. Hera achieves an average 37.3% improvement in effective machine utilization, enabling 26% reduction in required servers, significantly improving upon the baseline recommendation inference server. Yujeong Choi, John Kim 0001, Minsoo Rhu |
PACT | 1 |
| 2024 | ElasticRec: A Microservice-based Model Serving Architecture Enabling Elastic Resource Scaling for Recommendation ModelsabstractWith the increasing popularity of recommendation systems (RecSys), the demand for compute resources in data-centers has surged. However, the model-wise resource allocation employed in current RecSys model serving architectures falls short in effectively utilizing resources, leading to sub-optimal total cost of ownership. We propose ElasticRec, a model serving architecture for RecSys providing resource elasticity and high memory efficiency. ElasticRec is based on a microservice-based software architecture for fine-grained resource allocation, tailored to the heterogeneous resource demands of RecSys. Additionally, ElasticRec achieves high memory efficiency via our utility-based resource allocation. Overall, ElasticRec achieves an average $3.3 \times$ reduction in memory allocation size and $8.1 \times$ increase in memory utility, resulting in an average $1.6 \times$ reduction in deployment cost compared to state-of-the-art RecSys inference serving system. Yujeong Choi, Jiin Kim, Minsoo Rhu |
ISCA | 1 |
| 2024 | vTrain: A Simulation Framework for Evaluating Cost-Effective and Compute-Optimal Large Language Model TrainingabstractAs large language models (LLMs) become widespread in various application domains, a critical challenge the AI community is facing is how to train these large AI models in a cost-effective manner. Existing LLM training plans typically employ a heuristic based parallel training strategy which is based on empirical observations rather than grounded upon a thorough examination of the search space of LLM parallelization. Such limitation renders existing systems to leave significant performance left on the table, wasting millions of dollars worth of training cost. This paper presents our profiling-driven simulator called vTrain, providing AI practitioners a fast yet accurate software framework to determine an efficient and cost-effective LLM training system configuration. We demonstrate vTrain's practicality through several case studies, e.g., effectively evaluating optimal training parallelization strategies that balances training time and its associated training cost, efficient multi-tenant GPU cluster schedulers targeting multiple LLM training jobs, and determining a compute-optimal LLM model architecture given a fixed compute budget. Jehyeon Bang, Yujeong Choi, Myeongwoo Kim, Yongdeok Kim, Minsoo Rhu |
MICRO | 2 |
| 2022 | PARIS and ELSA: an elastic scheduling algorithm for reconfigurable multi-GPU inference serversabstractProviding low latency to end-users while maximizing server utilization and system throughput is crucial for cloud ML servers. NVIDIA's recently announced Ampere GPU architecture provides features to "reconfigure" one large, monolithic GPU into multiple smaller "GPU partitions". Such feature provides cloud ML service providers the ability to utilize the reconfigurable GPU not only for large-batch training but also for small-batch inference with the potential to achieve high resource utilization. We study this emerging GPU architecture with reconfigurability to develop a high-performance multi-GPU ML inference server, presenting a sophisticated partitioning algorithm for reconfigurable GPUs combined with an elastic scheduling algorithm tailored for our heterogeneously partitioned GPU server. Yunseong Kim, Yujeong Choi, Minsoo Rhu |
DAC | 2 |
| 2021 | Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning InferenceabstractIn cloud ML inference systems, batching is an essential technique to increase throughput which helps optimize total-cost-of-ownership. Prior graph batching combines the individual DNN graphs into a single one, allowing multiple inputs to be concurrently executed in parallel. We observe that the coarse-grained graph batching becomes suboptimal in effectively handling the dynamic inference request traffic, leaving significant performance left on the table. This paper proposes LazyBatching, an SLA-aware batching system that considers both scheduling and batching in the granularity of individual graph nodes, rather than the entire graph for flexible batching. We show that LazyBatching can intelligently determine the set of nodes that can be efficiently batched together, achieving an average 15×, 1.5×, and 5.5 × improvement than graph batching in terms of average response time, throughput, and SLA satisfaction, respectively. Yujeong Choi, Yunseong Kim, Minsoo Rhu |
HPCA | 1 |
| 2020 | NeuMMU: Architectural Support for Efficient Address Translations in Neural Processing UnitsabstractTo satisfy the compute and memory demands of deep neural networks (DNNs), neural processing units (NPUs) are widely being utilized for accelerating DNNs. Similar to how GPUs have evolved from a slave device into a mainstream processor architecture, it is likely that NPUs will become first-class citizens in this fast-evolving heterogeneous architecture space. This paper makes a case for enabling address translation in NPUs to decouple the virtual and physical memory address space. Through a careful data-driven application characterization study, we root-cause several limitations of prior GPU-centric address translation schemes and propose a memory management unit (MMU) that is tailored for NPUs. Compared to an oracular MMU design point, our proposal incurs only an average 0.06% performance overhead. Bongjoon Hyun, Youngeun Kwon, Yujeong Choi, John Kim 0001, Minsoo Rhu |
ASPLOS | 3 |
| 2020 | PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsabstractTo amortize cost, cloud vendors providing DNN acceleration as a service to end-users employ consolidation and virtualization to share the underlying resources among multiple DNN service requests. This paper makes a case for a "preemptible" neural processing unit (NPU) and a "predictive" multi-task scheduler to meet the latency demands of high-priority inference while maintaining high throughput. We evaluate both the mechanisms that enable NPUs to be preemptible and the policies that utilize them to meet scheduling objectives. We show that preemptive NPU multi-tasking can achieve an average 7.8×, 1.4×, and 4.8× improvement in latency, throughput, and SLA satisfaction, respectively. Yujeong Choi, Minsoo Rhu |
HPCA | 1 |