Guseul Heo

dblp:371/9017 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0008-3179-8541ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 46% Memory systems · 23% Reconfigurable computing and FPGAs · 23%
Databases, data mining, and information retrieval
1 paper
Indexing and storage engines · 50% Machine learning and data management · 50%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer graphics and multimedia
1 paper
Multimedia systems and quality of experience · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
0.912025
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse · Proc. VLDB Endow. 2025
Machine learning and data management › continual learning
incremental learning
0.812024
Accelerating String-key Learned Index Structures via Memoization-based Incremental Training · Proc. VLDB Endow. 2024
Indexing and storage engines
learned index
0.812024
Accelerating String-key Learned Index Structures via Memoization-based Incremental Training · Proc. VLDB Endow. 2024
Reconfigurable computing and FPGAs
FPGA accelerator
0.812024
Accelerating String-key Learned Index Structures via Memoization-based Incremental Training · Proc. VLDB Endow. 2024
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator
0.812024
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing · ASPLOS (3) 2024
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.812024
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing · ASPLOS (3) 2024
Memory systems
processing-in-memory
0.812024
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing · ASPLOS (3) 2024

Methods — techniques the papers use, named apart from their topics

vision transformer · 1.7memory-compute joint compaction · 1.7learning-based reuse detection · 1.7memoization · 1.5matrix decomposition · 1.5QR factorization · 1.5GEMV · 0.8GEMM · 0.8
YearPublicationVenuePosition
2026 LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
Jaehong Cho, Hyunmin Choi, Guseul Heo, Jongse Park
ISPASS3
2025 Throughput-Oriented LLM Inference via KV-Activation Hybrid Caching with A Single GPU
Hongbeen Kim, Soojin Hwang, Guseul Heo, Minwoo Noh, Jaehyuk Huh 0001
ICCD4
2025 Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
abstract
Recently, Video-Language Models (VideoLMs) have demonstrated remarkable capabilities, offering significant potential for flexible and powerful video query systems. These models typically rely on Vision Transformers (ViTs), which process video frames individually to extract visual embeddings. However, generating embeddings for large-scale videos requires ViT inferencing across numerous frames, posing a major hurdle to real-world deployment and necessitating solutions for integration into scalable video data management systems. This paper introduces Déjà Vu, a video-language query engine that accelerates ViT-based VideoLMs by reusing computations across consecutive frames. At its core is ReuseViT, a modified ViT model specifically designed for VideoLM tasks, which learns to detect inter-frame reuse opportunities, striking an effective balance between accuracy and reuse. Although ReuseViT significantly reduces computation, these savings do not directly translate into performance gains on GPUs. To overcome this, Déjà Vu integrates memory-compute joint compaction techniques that convert the FLOP savings into tangible performance gains. Evaluations on three VideoLM tasks show that Déjà Vu accelerates embedding generation by up to a 2.64× within a 2% error bound, dramatically enhancing the practicality of VideoLMs for large-scale video analytics.
Jinwoo Hwang, Yoonsung Kim, Guseul Heo, Hojoon Kim, Yunseok Jeong, Tadiwos Meaza, Eunhyeok Park, Jeongseob Ahn, Jongse Park
Proc. VLDB Endow.5
2024 NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
abstract
Modern transformer-based Large Language Models (LLMs) are constructed with a series of decoder blocks. Each block comprises three key components: (1) QKV generation, (2) multi-head attention, and (3) feed-forward networks. In batched processing, QKV generation and feed-forward networks involve compute-intensive matrix-matrix multiplications (GEMM), while multi-head attention requires bandwidth-heavy matrix-vector multiplications (GEMV). Machine learning accelerators like TPUs or NPUs are proficient in handling GEMM but are less efficient for GEMV computations. Conversely, Processing-in-Memory (PIM) technology is tailored for efficient GEMV computation, while it lacks the computational power to handle GEMM effectively.
Guseul Heo, Jaehong Cho, Hyunmin Choi, Hyungkyu Ham, Gwangsun Kim, Divya Mahajan 0001, Jongse Park
ASPLOS (3)1
2024 Accelerating String-key Learned Index Structures via Memoization-based Incremental Training
abstract
Learned indexes use machine learning models to learn the mappings between keys and their corresponding positions in key-value indexes. These indexes use the mapping information as training data. Learned indexes require frequent retrainings of their models to incorporate the changes introduced by update queries. To efficiently retrain the models, existing learned index systems often harness a linear algebraic QR factorization technique that performs matrix decomposition. This factorization approach processes all key-position pairs during each retraining, resulting in compute operations that grow linearly with the total number of keys and their lengths. Consequently, the retrainings create a severe performance bottleneck, especially for variable-length string keys, while the retrainings are crucial for maintaining high prediction accuracy and in turn, ensuring low query service latency. To address this performance problem, we develop an algorithm-hardware co-designed string-key learned index system, dubbed SIA. In designing SIA, we leverage a unique algorithmic property of the matrix decomposition-based training method. Exploiting the property, we develop a memoization-based incremental training scheme, which only requires computation over updated keys, while decomposition results of non-updated keys from previous computations can be reused. We further enhance SIA to offload a portion of this training process to an FPGA accelerator to not only relieve CPU resources for serving index queries (i.e., inference), but also accelerate the training itself. Our evaluation shows that compared to ALEX, LIPP, and SIndex, a state-of-the-art learned index systems, SIA-accelerated learned indexes offer 2.6× and 3.4× higher throughput on the two real-world benchmark suites, YCSB and Twitter cache trace, respectively.
Minsu Kim 0004, Jinwoo Hwang, Guseul Heo, Seiyeon Cho, Divya Mahajan 0001, Jongse Park
Proc. VLDB Endow.3