EDBT 2026 Demo / reviewers in the wild / expert
Taeksang Song
dblp:170/1775
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0001-2958-6303ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Clone: A Collaborative Multi-device System for Retrieval-Augmented Generation over CXLabstractAs vector databases scale in Retrieval-Augmented Generation (RAG), the retrieval phase increasingly bottlenecks end-to-end latency. While Compute Express Link (CXL) offers scalable memory expansion, naïve CXL deployments suffer from intra-device bandwidth saturation and inter-device load imbalance, which collectively hinder system responsiveness. Seoyoung Ko, Wanju Doh, Eojin Na, Hyunjeong Shim, Sungmin Yun 0001, Jinin So, Yongsuk Kwon, Sang-Soo Park, Si-Dong Roh, Minyong Yoon, Taeksang Song, Eojin Lee, Jung Ho Ahn |
ICS | 11 |
| 2026 | S-Tiering: A Unified HW/SW Solution for Memory Tiering Based on the Standard CXL Hotness Monitoring UnitabstractIn this paper, we proposeS-Tiering, a unified hardware and software solution for memory tiering based on the standard CXL Hotness Monitoring Unit (CHMU).S-Tieringconsists of hardware components that comply with CHMU hardware specification defined in the CXL 3.2 Specification, and software components that control hardware components. Based on these components,S-Tieringminimizes the access to CXL memory by properly steering page migration.We evaluate various probabilistic data structure algorithms and adopt a Count-Min Sketch-based Hot Page Tracker that achieves 99% accuracy with only 0.3% tracking buffer overhead compared to assigning a dedicated counter for every 4KB page. We implement hardware components ofS-Tieringon an Field-Programmable Gate Array board and software components ofS-Tieringon Ubuntu 22.04 with Linux-v6.8 kernel. We evaluate the performance impact ofS-Tieringon benchmarks representative of real applications (e.g., High Performance Computing, Graph-processing, In-Memory Database).S-Tieringachieves a performance improvement of up to 193% compared to first-touch allocation and outperforms AutoNUMA memory tiering by 184%p.S-Tieringminimizes the memory access to CXL memory and increases the bandwidth utilization of DDR memory up to ×11. Seunghak Lee, Wonjae Lee 0001, Hojin Nam, Jehoon Park, Youngshin Park, Junhyeok Im, Jinin So, Raghu Vamsi Krishna Talanki, Praful Ramesh O, Rajeev Verma, Taeksang Song, Wonhwa Shin, Sangjoon Hwang 0001 |
IEEE Trans. Computers | 12 |
| 2026 | Pangaea v2: CXL-Based Disaggregated Memory System Architecture for Cloud-Native OrchestrationabstractToday’s data centers suffer from CPU and memory resource stranding because they often over-provision resources when deploying servers for worst-case scenarios. This problem gives rise to a disaggregated system architecture allowing each type of resource to be allocated, utilized and freed separately as required. In particular, research on disaggregated memory systems over the past few years has focused primarily on achieving low remote memory access latency over Ethernet, which is known as the RDMA optimization approach.In this paper, we introduce a dynamic rack-scale disaggregated memory system architecture, so called Pangaea v2 using ASIC-CXL H/W and memory orchestration S/W designed to increase the memory utilization of worker nodes between containerized applications execution in a Kubernetes, a major process container platform in the data center. In our evaluation with in-memory database application, disaggregated CXL memory system shows significantly better throughput improved by up to 10.2x/6.7x and 99th tail latency reduced to 96%/93% compared to RDMA with RoCEv2/InfiniBand. Han Deok Lee, Jehoon Park, Younghyun Lee, Junhyeok Im, Jin Jung, Jinin So, Siamak Tavallaei, Woo Taek Shim, Chin-Hua Chang, Sungwook Ryu, Taeksang Song, Wonhwa Shin, Sangjoon Hwang 0001 |
IEEE Trans. Computers | 13 |
| 2025 | Efficient Caching with A Tag-enhanced DRAMabstractAs SRAM-based caches are hitting a scaling wall, manufacturers are integrating DRAM-based caches into system designs to continue increasing cache sizes. While DRAM caches can improve the performance of memory systems, existing DRAM cache designs suffer from high miss penalties, wasted data movement, and interference between misses and demands. In this paper, we propose TDRAM, a novel DRAM microarchitecture tailored for caching. TDRAM enhances existing DRAM, such as HBM3, by adding small, low-latency mats to store tags and metadata on the same die as the data mats. These mats enable tag and data access in lockstep, in-DRAM tag comparison, and conditional data response based on the comparison result (reducing wasted data transfers), akin to SRAM cache mechanisms. TDRAM further optimizes hit and miss latencies through opportunistic early tag probing. Moreover, TDRAM introduces a flush buffer to store conflicting dirty data on write misses, eliminating data bus turnaround delays on write demands. We evaluate TDRAM in a full-system simulation using a set of HPC workloads with large memory footprints, showing that TDRAM, on average, provides $2.65 \times$ faster tag checks, $1.23 \times$ speedup, and 21% less energy consumption compared to state-of-the-art commercial and research designs. Maryam Babaie, Ayaz Akram, Wendy Elsasser, Brent Haukness, Michael R. Miller, Taeksang Song, Thomas Vogelsang, Steven C. Woo, Jason Lowe-Power |
HPCA | 6 |