EDBT 2026 Demo / reviewers in the wild / expert
Seungcheol Baek
dblp:55/10504
· DBLP profile ↗
8ranked-venue papers
4as first author
2since 2021 · last 2024
0009-0007-1624-4586ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 51% Hardware accelerators and domain-specific architectures · 42% Processor architecture and microarchitecture · 5% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.8 | 1 | 2024 | IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System · ASPLOS (3) 2024 |
Memory systems
processing-in-memory |
0.8 | 1 | 2024 | IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System · ASPLOS (3) 2024 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
transformer inference accelerator |
0.8 | 1 | 2024 | IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System · ASPLOS (3) 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization · EMNLP 2023 |
Machine learning › Efficient and distributed learning › model compression › quantization
post-training quantization |
0.7 | 1 | 2023 | Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization · EMNLP 2023 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.7 | 1 | 2023 | Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization · EMNLP 2023 |
Machine learning › Efficient and distributed learning › model compression › quantization
weight-activation quantization |
0.7 | 1 | 2023 | Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization · EMNLP 2023 |
Memory systems › memory hierarchy › cache hierarchy
last-level cache |
0.4 | 2 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 ECM: Effective Capacity Maximizer for high-performance compressed caching · HPCA 2013 |
Hardware accelerators and domain-specific architectures
accelerator integration |
0.2 | 1 | 2024 | IANUS: Integrated Accelerator based on NPU-PIM Unified Memory System · ASPLOS (3) 2024 |
Memory systems
cache |
0.2 | 1 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 |
Memory systems › memory compression
cache compression |
0.2 | 1 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 |
Memory systems › cache management
cache replacement |
0.2 | 1 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 |
Processor architecture and microarchitecture
arithmetic unit |
0.2 | 1 | 2023 | Enhancing Computation Efficiency in Large Language Models through Weight and Activation Quantization · EMNLP 2023 |
Memory systems
cache management |
0.2 | 1 | 2013 | ECM: Effective Capacity Maximizer for high-performance compressed caching · HPCA 2013 |
Energy-efficient computing › power management › memory power management
cache energy reduction |
0.1 | 1 | 2015 | Size-Aware Cache Management for Compressed Cache Architectures · IEEE Trans. Computers 2015 |
Performance modeling and evaluation › simulation
cache simulation |
0.0 | 1 | 2013 | ECM: Effective Capacity Maximizer for high-performance compressed caching · HPCA 2013 |
Methods — techniques the papers use, named apart from their topics
sequence-length-aware calibration · 1.3dINT format · 1.3activation-quantization-aware scaling · 1.3simulation · 0.8trace-driven simulation · 0.2full-system simulation · 0.2size-aware replacement · 0.2size-aware insertion · 0.2dynamic threshold adjustment · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | IANUS: Integrated Accelerator based on NPU-PIM Unified Memory SystemabstractAccelerating end-to-end inference of transformer-based large language models (LLMs) is a critical component of AI services in datacenters. However, the diverse compute characteristics of LLMs' end-to-end inference present challenges as previously proposed accelerators only address certain operations or stages (e.g., self-attention, generation stage, etc.). To address the unique challenges of accelerating end-to-end inference, we propose IANUS - Integrated Accelerator based on NPU-PIM Unified Memory System. IANUS is a domain-specific system architecture that combines a Neural Processing Unit (NPU) with a Processing-in-Memory (PIM) to leverage both the NPU's high computation throughput and the PIM's high effective memory bandwidth. In particular, IANUS employs a unified main memory system where the PIM memory is used both for PIM operations and for NPU's main memory. The unified main memory system ensures that memory capacity is efficiently utilized and the movement of shared data between NPU and PIM is minimized. However, it introduces new challenges since normal memory accesses and PIM computations cannot be performed simultaneously. Thus, we propose novel PIM Access Scheduling that manages not only the scheduling of normal memory accesses and PIM computations but also workload mapping across the PIM and the NPU. Our detailed simulation evaluations show that IANUS improves the performance of GPT-2 by 6.2× and 3.2×, on average, compared to the NVIDIA A100 GPU and the state-of-the-art accelerator. As a proof-of-concept, we develop a prototype of IANUS with a commercial PIM, NPU, and an FPGA-based PIM controller to demonstrate the feasibility of IANUS. Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon, Guhyun Kim, Chanwook Park, Ilkon Kim, Jaehan Park, Jeongbin Kim 0001, Woojae Shin, Jongsoon Won, Haerang Choi, Kyuyoung Kim, Daehan Kwon, Chunseok Jeong, Yongseok Choi, Wooseok Byun, Seungcheol Baek, John Kim 0001 |
ASPLOS (3) | 19 |
| 2023 | Enhancing Computation Efficiency in Large Language Models through Weight and Activation QuantizationabstractLarge Language Models (LLMs) are proficient in natural language processing tasks, but their deployment is often restricted by extensive parameter sizes and computational demands.This paper focuses on post-training quantization (PTQ) in LLMs, specifically 4-bit weight and 8-bit activation (W4A8) quantization, to enhance computational efficiency-a topic less explored compared to weight-only quantization.We present two innovative techniques: activation-quantization-aware scaling (AQAS) and sequence-length-aware calibration (SLAC) to enhance PTQ by considering the combined effects on weights and activations and aligning calibration sequence lengths to target tasks.Moreover, we introduce dINT, a hybrid data format combining integer and denormal representations, to address the underflow issue in W4A8 quantization, where small values are rounded to zero.Through rigorous evaluations of LLMs, including OPT and LLaMA, we demonstrate that our techniques significantly boost task accuracies to levels comparable with full-precision models.By developing arithmetic units compatible with dINT, we further confirm that our methods yield a 2× hardware efficiency improvement compared to 8-bit integer MAC unit. Janghwan Lee, Seungcheol Baek, Seok Joong Hwang, Wonyong Sung, Jungwook Choi |
EMNLP | 3 |
| 2017 | HoPE: Hot-Cacheline Prediction for Dynamic Early Decompression in Compressed LLCsabstractData compression plays a pivotal role in improving system performance and reducing energy consumption, because it increases the logical effective capacity of a compressed memory system without physically increasing the memory size. However, data compression techniques incur some cost, such as non-negligible compression and decompression overhead. This overhead becomes more severe if compression is used in the cache. In this article, we aim to minimize the read-hit decompression penalty in compressed Last-Level Caches (LLCs) by speculatively decompressing frequently used cachelines. To this end, we propose a Hot-cacheline Prediction and Early decompression (HoPE) mechanism that consists of three synergistic techniques: Hot-cacheline Prediction (HP), Early Decompression (ED), and Hit-history-based Insertion (HBI). HP and HBI efficiently identify the hot compressed cachelines, while ED selectively decompresses hot cachelines, based on their size information. Unlike previous approaches, the HoPE framework considers the performance balance/tradeoff between the increased effective cache capacity and the decompression penalty. To evaluate the effectiveness of the proposed HoPE mechanism, we run extensive simulations on memory traces obtained from multi-threaded benchmarks running on a full-system simulation framework. We observe significant performance improvements over compressed cache schemes employing the conventional Least-Recently Used (LRU) replacement policy, the Dynamic Re-Reference Interval Prediction (DRRIP) scheme, and the Effective Capacity Maximizer (ECM) compressed cache management mechanism. Specifically, HoPE exhibits system performance improvements of approximately 11%, on average, over LRU, 8% over DRRIP, and 7% over ECM by reducing the read-hit decompression penalty by around 65%, over a wide range of applications. Jaehyun Park 0005, Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Vinson Young, Junghee Lee 0004, Jongman Kim |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2015 | Size-Aware Cache Management for Compressed Cache ArchitecturesabstractA practical way to increase the effective capacity of a microprocessor's cache, without physically increasing the cache size, is to employ data compression. Last-Level Caches (LLC) are particularly amenable to such compression schemes, since the primary purpose of the LLC is to minimize the miss rate, i.e., it directly benefits from a larger logical capacity. In compressed LLCs, the cacheline size varies depending on the achieved compression ratio. Our observations indicate that this size information gives useful hints when managing the cache (e.g., when selecting a victim), which can lead to increased cache performance. However, there are currently no replacement policies tailored to compressed LLCs; existing techniques focus primarily on locality information. This article introduces the concept of size-aware cache management as a way to maximize the performance of compressed caches. Upon analyzing the benefits of considering size information in the management of compressed caches, we propose a novel mechanism-called Effective Capacity Maximizer (ECM)-to further enhance the performance and energy consumption of compressed LLCs. The proposed technique revolves around four fundamental principles: ECM Insertion (ECM-I), ECM Promotion (ECM-P), ECM Eviction Scheduling (ECM-ES), and ECM Replacement (ECM-R). Extensive simulations with memory traces from real applications running on a full-system simulator demonstrate significant improvements compared to compressed cache schemes employing conventional locality-aware cache replacement policies. Specifically, our ECM shows an average effective capacity increase of 18.4 percent over the Least-Recently Used (LRU) policy, and 23.9 percent over the Dynamic Re-Reference Interval Prediction (DRRIP) [1] scheme. This translates into average system performance improvements of 7.2 percent over LRU and 4.2 percent over DRRIP. Moreover, the average energy consumption is also reduced by 5.9 percent over LRU and 3.8 percent over DRRIP. Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Junghee Lee 0004, Jongman Kim |
IEEE Trans. Computers | 1 |
| 2014 | Designing Hybrid DRAM/PCM Main Memory Systems Utilizing Dual-Phase CompressionabstractThe last few years have witnessed the emergence of a promising new memory technology, namely Phase-Change Memory (PCM). Due to its inherent ability to scale deeply into the nanoscale regime and its low power consumption, PCM is increasingly viewed as an attractive alternative for the memory subsystem of future microprocessor architectures. However, PCM is marred by a duo of potentially show-stopping deficiencies, that is, poor write performance (especially when compared to the prevalent and ubiquitous DRAM technology) and limited durability. These weaknesses have urged designers to develop various supporting architectural techniques to aid and complement the operation of the PCM while mitigating its innate flaws. One promising such solution is the deployment of hybridized memory architectures that fuse DRAM and PCM, in order to combine the best attributes of each technology. In this article, we introduce a novel Dual-Phase Compression (DPC) scheme and its architectural design aimed at DRAM/PCM hybrids, which caters to the limitations of PCM technology while optimizing memory performance. The DPC technique is specifically optimized for PCM-based environments and is transparent to the operation of the remaining components of the memory subsystem. Furthermore, the proposed architecture is imbued with a multifaceted wear-leveling technique to enhance the durability and prolong the lifetime of the PCM. Extensive simulations with traces from real applications running on a full-system simulator demonstrate 20.4% performance improvement and 46.9% energy reduction, on average, as compared to a baseline DRAM/PCM hybrid implementation. Additionally, the multifaceted wear-leveling technique is shown to significantly prolong the lifetime of the PCM. Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Jongman Kim |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2013 | ECM: Effective Capacity Maximizer for high-performance compressed cachingabstractCompressed Last-Level Cache (LLC) architectures have been proposed to enhance system performance by efficiently increasing the effective capacity of the cache, without physically increasing the cache size. In a compressed cache, the cacheline size varies depending on the achieved compression ratio. We observe that this size information gives a useful hint when selecting a victim, which can lead to increased cache performance. However, no replacement policy tailored to compressed LLCs has been investigated so far. This paper introduces the notion of size-aware compressed cache management as a way to maximize the performance of compressed caches. Toward this end, the Effective Capacity Maximizer (ECM) scheme is introduced, which targets compressed LLCs. The proposed mechanism revolves around three fundamental principles: Size-Aware Insertion (SAI), a Dynamically Adjustable Threshold Scheme (DATS), and Size-Aware Replacement (SAR). By adjusting the eviction criteria, based on the compressed data size, one may increase the effective cache capacity and minimize the miss penalty. Extensive simulations with memory traces from real applications running on a full-system simulator demonstrate significant improvements compared to compressed cache schemes employing the conventional Least-Recently Used (LRU) and Dynamic Re-Reference Interval Prediction (DRRIP) [11] replacement policies. Specifically, ECM shows an average effective capacity increase of 15% over LRU and 18.8% over DRRIP, an average cache miss reduction of 9.4% over LRU and 3.9% over DRRIP, and an average system performance improvement of 6.2% over LRU and 3.3% over DRRIP. Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Junghee Lee 0004, Jongman Kim |
HPCA | 1 |
| 2012 | A dual-phase compression mechanism for hybrid DRAM/PCM main memory architecturesabstractPhase-Change Memory (PCM) is emerging as a promising new memory technology, due to its inherent ability to scale deeply into the nanoscale regime. However, PCM is still marred by a duet of potentially show-stopping deficiencies: poor write performance and limited durability. These weaknesses have urged designers to develop various supporting architectural techniques to aid and complement the operation of the PCM, while mitigating its innate flaws. One promising such solution is the deployment of hybridized memory architectures that fuse DRAM and PCM, in order to combine the best attributes of each technology. In this paper, we introduce a Dual-Phase Compression (DPC) scheme specifically optimized for DRAM/PCM hybrid environments. Extensive simulations with traces from real applications running on a full-system simulator of a multicore system demonstrate 35.1% performance improvement and 29.3% energy reduction, on average, as compared to a baseline DRAM/PCM hybrid implementation. Seungcheol Baek, Hyung Gyu Lee, Chrysostomos Nicopoulos, Jongman Kim |
ACM Great Lakes Symposium on VLSI | 1 |
| 2011 | An energy- and performance-aware DRAM cache architecture for hybrid DRAM/PCM main memory systemsabstractThe last few years have witnessed the emergence of a promising new memory technology. Phase-Change Memory (PCM) is increasingly viewed as an attractive alternative for the memory sub-system of future microprocessor architectures, mainly because of its inherent ability to scale deeply into the nanoscale regime, and its low power consumption. However, PCM's write performance is its Achilles' heel, especially when compared to the prevalent DRAM technology. This weakness necessitates the deployment of hybridized solutions that fuse DRAM and PCM, in order to attain high overall system performance. In this paper, we set out to explore how various DRAM/PCM hybrid configurations affect system performance and energy consumption, and then proceed with the presentation of a novel architecture that maximizes performance without adversely affecting power efficiency. An energy-delay product improvement of 42.2%, on average, over conventional hybrid structures, is demonstrated. Hyung Gyu Lee, Seungcheol Baek, Chrysostomos Nicopoulos, Jongman Kim |
ICCD | 2 |