EDBT 2026 Demo / reviewers in the wild / expert
Sam Son
dblp:290/9350
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0007-9268-9915ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CPC: Coordinated Page Cache for Serverless ComputingabstractVirtual machine-based serverless computing suffers from high memory overhead due to data duplication. This duplication occurs not only across function (or VM) images but also within in-memory caches shared between VMs. Such redundancy limits the scalability and efficiency of serverless computing infrastructure. While existing memory deduplication techniques can help, they are ill-suited to serverless computing due to their high CPU overhead and delayed deduplication, which often comes too late to provide meaningful memory savings given the short lifetime of a function. However, serverless platforms inherently know which functions (or VMs) will be executed in advance, presenting a unique opportunity for proactive memory optimization. In this paper, we propose CPC, a proactive deduplication scheme that enables efficient page sharing across microVMs, lightweight variants of traditional VMs, in serverless environments. Upon function submission, CPC deduplicates identical files across function images using existing techniques. During function execution, CPC leverages this deduplication information to further deduplicate redundant cache data across VMs and the hypervisor. Comprehensive experiments using serverless function benchmarks show that CPC reduces memory consumption by up to 32.0% compared to AWS Lambda when hosting heterogeneous functions on a shared worker node. Furthermore, with Azure Functions traces, CPC achieves a 41.2% reduction in 99th percentile tail latency by reducing unnecessary instance evictions through more efficient memory utilization. Keun Soo Lim, Yunjay Hong, Jongheon Jeong, Sam Son, Yeonhong Park, Jae W. Lee, Jinkyu Jeong |
PACT | 4 |
| 2024 | Efficient Microsecond-scale Blind Scheduling with Tiny QuantaabstractA longstanding performance challenge in datacenter-based applications is how to efficiently handle incoming client requests that spawn many very short (μs scale) jobs that must be handled with high throughput and low tail latency. When no assumptions are made about the duration of individual jobs, or even about the distribution of their durations, this requires blind scheduling with frequent and efficient preemption, which is not scalably supported for μs-level tasks. We present Tiny Quanta (TQ), a system that enables efficient blind scheduling of μs-level workloads. TQ performs fine-grained preemptive scheduling and does so with high performance via a novel combination of two mechanisms: forced multitasking and two-level scheduling. Evaluations with a wide variety of μs-level workloads show that TQ achieves low tail latency while sustaining 1.2x to 6.8x the throughput of prior blind scheduling systems. Zhihong Luo, Sam Son, Dev Bali, Emmanuel Amaro, Amy Ousterhout, Sylvia Ratnasamy, Scott Shenker |
ASPLOS (2) | 2 |
| 2024 | Harvesting Memory-bound CPU Stall Cycles in Software with MSH
Zhihong Luo, Sam Son, Sylvia Ratnasamy, Scott Shenker |
OSDI | 2 |
| 2021 | FlashNeuron: SSD-Enabled Large-Batch Training of Very Deep Neural Networks
Jonghyun Bae, Jongsung Lee 0001, Yunho Jin, Sam Son, Shine Kim, Hakbeom Jang, Tae Jun Ham, Jae W. Lee |
FAST | 4 |
| 2021 | Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise SchedulingabstractTo meet surging demands for deep learning inference services, many cloud computing vendors employ high-performance specialized accelerators, called neural processing units (NPUs). One important challenge for effective use of NPUs is to achieve high resource utilization over a wide spectrum of deep neural network (DNN) models with diverse arithmetic intensities. There is often an intrinsic mismatch between the compute-to-memory bandwidth ratio of an NPU and the arithmetic intensity of the model it executes, leading to under-utilization of either compute resources or memory bandwidth. Ideally, we want to saturate both compute TOP/s and DRAM bandwidth to achieve high system throughput. Thus, we propose Layerweaver, an inference serving system with a novel multi-model time-multiplexing scheduler for NPUs. Layerweaver reduces the temporal waste of computation resources by interweaving layer execution of multiple different models with opposing characteristics: compute-intensive and memory-intensive. Layerweaver hides the memory time of a memory-intensive model by overlapping it with the relatively long computation time of a compute-intensive model, thereby minimizing the idle time of the computation units waiting for off-chip data transfers. For a two-model serving scenario of batch 1 with 16 different pairs of compute- and memory-intensive models, Layerweaver improves the temporal utilization of computation units and memory channels by 44.0% and 28.7%, respectively, to increase the system throughput by 60.1% on average, over the baseline executing one model at a time. Young H. Oh, Seonghak Kim, Yunho Jin, Sam Son, Jonghyun Bae, Jongsung Lee 0001, Yeonhong Park, Dong Uk Kim, Tae Jun Ham, Jae W. Lee |
HPCA | 4 |
| 2021 | ASAP: Fast Mobile Application Switch via Adaptive Prepaging
Sam Son, Seung Yul Lee, Yunho Jin, Jonghyun Bae, Jinkyu Jeong, Tae Jun Ham, Jae W. Lee, Hongil Yoon |
USENIX ATC | 1 |