VLDB 2026 Research / reviewers in the wild / expert
Junkyum Kim
dblp:343/5336
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2026
0009-0003-2519-2379ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Hardware accelerators and domain-specific architectures · 45% Cloud and datacenter computing · 27% Electronic design automation · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Artificial intelligence
3 papers |
Efficient and distributed learning · 57% Language models and text generation · 26% Optimization for machine learning · 17% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
1.3 | 2 | 2023 | MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERT · IEEE Trans. Computers 2023 OptimStore: In-Storage Optimization of Large Scale DNNs with On-Die Processing · HPCA 2023 |
Information retrieval
retrieval-augmented generation |
1.0 | 1 | 2026 | VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG · HPCA 2026 |
Information retrieval › similarity search
vector similarity search |
1.0 | 1 | 2026 | VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG · HPCA 2026 |
Cloud and datacenter computing
resource management |
1.0 | 1 | 2026 | VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG · HPCA 2026 |
Cloud and datacenter computing › resource allocation › resource allocation policy
resource partitioning |
1.0 | 1 | 2026 | VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG · HPCA 2026 |
Electronic design automation › physical design › circuit partitioning
timing-driven partitioning |
1.0 | 1 | 2026 | VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG · HPCA 2026 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.7 | 1 | 2023 | MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERT · IEEE Trans. Computers 2023 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator |
0.7 | 1 | 2023 | OptimStore: In-Storage Optimization of Large Scale DNNs with On-Die Processing · HPCA 2023 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
in-memory computing accelerator |
0.7 | 1 | 2023 | MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERT · IEEE Trans. Computers 2023 |
Storage systems › computational storage
in-storage computing |
0.7 | 1 | 2023 | OptimStore: In-Storage Optimization of Large Scale DNNs with On-Die Processing · HPCA 2023 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN accelerator
ReRAM-based DNN accelerator |
0.7 | 1 | 2023 | MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERT · IEEE Trans. Computers 2023 |
Natural language and speech › Language models and text generation
large language model inference |
0.3 | 1 | 2026 | VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAG · HPCA 2026 |
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent |
0.2 | 1 | 2023 | OptimStore: In-Storage Optimization of Large Scale DNNs with On-Die Processing · HPCA 2023 |
Memory systems
processing-in-memory |
0.2 | 1 | 2023 | MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERT · IEEE Trans. Computers 2023 |
Memory systems › non-volatile memory
resistive memory |
0.2 | 1 | 2023 | MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERT · IEEE Trans. Computers 2023 |
Methods — techniques the papers use, named apart from their topics
performance estimation · 3.0index partitioning · 3.0task scheduling · 1.3model generation · 1.3adam optimizer · 1.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VectorLiteRAG: Latency-Aware and Fine-Grained Resource Partitioning for Efficient RAGabstractRetrieval-Augmented Generation leverages vector similarity search to enhance large language models with up-to-date, external knowledge, enabling accurate and reliable responses. While CPU-only vector search incurs high latency on large, high-dimensional indices, co-locating the retriever and the LLM on the GPU leads to resource sharing that can create resource contention. Specifically, vector search is memory and I/O intensive, placing it in direct conflict with LLM inference, which demands memory for KV cache and compute for higher throughput. We present VectorLiterAG, a latency-aware RAG serving system that explicitly orchestrates data placement and execution across retrieval and inference to meet strict end-to-end SLOs. VectorLiterAG is driven by access-pattern analysis and performance estimation to regulate how retrieval variability can be mitigated and managed in the system with LLM inference, enabling SLO-compliant execution under skewed and dynamic workloads. By jointly modeling search latency and query hit-rate distributions, VectorLiterAG identifies an optimal index partitioning point across CPU and GPU that minimizes contention and stabilizes batching behavior, thereby maximizing sustained throughput under skewed access patterns. A low-overhead online index update mechanism allows VectorLiterag to continuously adapt to evolving request distributions, preserving batching efficiency and throughput as access patterns evolve. Our evaluations demonstrate that VecTORLITERAG consistently expands the range of SLO-compliant request rate across all tested configurations. Without increasing the generation latency or requiring additional hardware, VecTORLITERAG outperforms both naive and existing alternative frameworks, improving attainable SLO-bound throughput by up to$\mathbf{1. 5} \times$. Junkyum Kim, Divya Mahajan 0001 |
HPCA | 1 |
| 2023 | OptimStore: In-Storage Optimization of Large Scale DNNs with On-Die ProcessingabstractTraining deep neural network (DNN) models is a resource-intensive, iterative process. For this reason, nowadays, complex optimizers like Adam are widely adopted as it increases the speed and efficiency of training. These optimizers, however, employ additional variables and raise the memory demand 2× to 3× of model parameters, worsening the memory capacity bottleneck. Moreover, as the size of DNN models is projected to grow even further, it is not practical to assume that the future models will fit in accelerator memory. This has triggered various efforts to offload models to flash-based storage. However, when the model, especially the optimizer, is offloaded to flash, the limited I/O bandwidth severely slows down the overall training process. To this end, we present OptimStore, a solid-state drive (SSD) system with on-die processing (ODP) architectures for gradient descent-based machine learning models. OptimStore accelerates the training process of such large-scale models by processing model optimization in the storage device, specifically inside the flash dies. ODP capability of OptimStore eliminates the heavy data movement over external interconnect and internal flash channels. Overall, OptimStore achieves, on average, a 2.8× speedup and a 3.6× improved energy efficiency in the weight update stage over baseline SSD offloading. Junkyum Kim, Myeonggu Kang, Yunki Han, Yanggon Kim 0001, Lee-Sup Kim |
HPCA | 1 |
| 2023 | MGen: A Framework for Energy-Efficient In-ReRAM Acceleration of Multi-Task BERTabstractRecently, multiple transformer models, such as BERT, have been utilized together to support multiple natural language processing (NLP) tasks in a system, also known as multi-task BERT. Multi-task BERT with very high weight parameters increases the area requirement of a processing in resistive memory (ReRAM) architecture, and several works have attempted to address this model size issue. Despite the reduced parameters, the number of multi-task BERT computations remains the same, leading to massive energy consumption in ReRAM-based deep neural network (DNN) accelerators. Therefore, we suggest a framework for better energy efficiency during the ReRAM acceleration of multi-task BERT. First, we analyze the inherent redundancies of multi-task BERT and the computational properties of the ReRAM-based DNN accelerator, after which we propose what is termed the model generator, which produces optimal BERT models supporting multiple tasks. The model generator reduces multi-task BERT computations while maintaining the algorithmic performance. Furthermore, we present task scheduler, which adjusts the execution order of multiple tasks, to run the produced models efficiently. As a result, the proposed framework achieves maximally 4.4× higher energy efficiency over the baseline, and it can also be combined with the previous multi-task BERT works to achieve both a smaller area and higher energy efficiency. Myeonggu Kang, Hyein Shin, Junkyum Kim, Lee-Sup Kim |
IEEE Trans. Computers | 3 |