VLDB 2026 Research / reviewers in the wild / expert
Heejun Lee
dblp:194/8770
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Training-Free Sub-quadratic Cost Transformer Model Serving Framework with Hierarchically Pruned AttentionabstractIn modern large language models (LLMs), increasing the context length is crucial for improving comprehension and coherence in long-context, multi-modal, and retrieval-augmented language generation.
While many recent transformer models attempt to extend their context length over a million tokens, they remain impractical due to the quadratic time and space complexities.
Although recent works on linear and sparse attention mechanisms can achieve this goal, their real-world applicability is often limited by the need to re-train from scratch and significantly worse performance. In response, we propose a novel approach, Hierarchically Pruned Attention (HiP), which reduces the time complexity of the attention mechanism to $O(T \log T)$ and the space complexity to $O(T)$, where $T$ is the sequence length.
We notice a pattern in the attention scores of pretrained LLMs where tokens close together tend to have similar scores, which we call "attention locality". Based on this observation, we utilize a novel tree-search-like algorithm that estimates the top-$k$ key tokens for a given query on the fly, which is mathematically guaranteed to have better performance than random attention pruning. In addition to improving the time complexity of the attention mechanism, we further optimize GPU memory usage by implementing KV cache offloading, which stores only $O(\log T)$ tokens on the GPU while maintaining similar decoding throughput. Experiments on benchmarks show that HiP, with its training-free nature, significantly reduces both prefill and decoding latencies, as well as memory usage, while maintaining high-quality generation with minimal degradation.
HiP enables pretrained LLMs to scale up to millions of tokens on commodity GPUs, potentially unlocking long-context LLM applications previously deemed infeasible. Heejun Lee, Geon Park, Youngwan Lee, Jaduk Suh, Wonyong Jeong, Bumsik Kim, Hyemin Lee, Myeongjae Jeon, Sung Ju Hwang |
ICLR | 1 |
| 2025 | Training Free Exponential Context Extension via Cascading KV CacheabstractThe transformer's context window is vital for tasks such as few-shot learning and conditional generation as it preserves previous tokens for active memory. However, as the context lengths increase, the computational costs grow quadratically, hindering the deployment of large language models (LLMs) in real-world, long sequence scenarios. Although some recent key-value caching (KV Cache) methods offer linear inference complexity, they naively manage the stored context, prematurely evicting tokens and losing valuable information. Moreover, they lack an optimized prefill/prompt stage strategy, resulting in higher latency than even quadratic attention for realistic context sizes. In response, we introduce a novel mechanism that leverages cascading sub-cache buffers to selectively retain the most relevant tokens, enabling the model to maintain longer context histories without increasing the cache size. Our approach outperforms linear caching baselines across key benchmarks, including streaming perplexity, question answering, book summarization, and passkey retrieval, where it retains better retrieval accuracy at 1M tokens after four doublings of the cache size of 65K. Additionally, our method reduces prefill stage latency by a factor of 6.8 when compared to flash attention on 1M tokens. These innovations not only enhance the computational efficiency of LLMs but also pave the way for their effective deployment in resource-constrained environments, enabling large-scale, real-time applications with significantly reduced latency. Jeffrey Willette, Heejun Lee, Youngwan Lee, Myeongjae Jeon, Sung Ju Hwang |
ICLR | 2 |
| 2025 | Delta Attention: Fast and Accurate Sparse Attention Inference by Delta CorrectionabstractThe attention mechanism of a transformer has a quadratic complexity, leading to high inference costs and latency for long sequences. However, attention matrices are mostly sparse, which implies that many entries may be omitted from computation for efficient inference. Sparse attention inference methods aim to reduce this computational burden; however, they also come with a troublesome performance degradation. We discover that one reason for this degradation is that the sparse calculation induces a distributional shift in the attention outputs. The distributional shift causes decoding-time queries to fail to align well with the appropriate keys from the prefill stage, leading to a drop in performance. We propose a simple, novel, and effective procedure for correcting this distributional shift, bringing the distribution of sparse attention outputs closer to that of quadratic attention. Our method can be applied on top of any sparse attention method, and results in an average 36\%pt performance increase, recovering 88\% of quadratic attention accuracy on the 131K RULER benchmark when applied on top of sliding window attention with sink tokens while only adding a small overhead. Our method can maintain approximately 98.5\% sparsity over full quadratic attention, making our model 32 times faster than Flash Attention 2 when processing 1M token prefills. Jeffrey Willette, Heejun Lee, Sung Ju Hwang |
NeurIPS | 2 |
| 2024 | BoneStory: Visual Storytelling in 3D Virtual Surgical Planning for Bone Fracture Reduction
Heejun Lee, Peter A. J. Pijpker, Joep Kraeima, Lorenzo Amabili, Fokie Cnossen, Jos B. T. M. Roerdink, Peter M. A. van Ooijen, Jirí Kosinka |
CGI (2) | 1 |
| 2024 | SEA: Sparse Linear Attention with Estimated Attention MaskabstractThe transformer architecture has driven breakthroughs in recent years on tasks
which require modeling pairwise relationships between sequential elements, as
is the case in natural language understanding. However, long seqeuences pose a
problem due to the quadratic complexity of the attention operation. Previous re-
search has aimed to lower the complexity by sparsifying or linearly approximating
the attention matrix. Yet, these approaches cannot straightforwardly distill knowl-
edge from a teacher’s attention matrix, and often require complete retraining from
scratch. Furthermore, previous sparse and linear approaches lose interpretability
if they cannot produce full attention matrices. To address these challenges, we
propose SEA: Sparse linear attention with an Estimated Attention mask. SEA
estimates the attention matrix with linear complexity via kernel-based linear at-
tention, then subsequently creates a sparse attention matrix with a top-k̂ selection
to perform a sparse attention operation. For language modeling tasks (Wikitext2),
previous linear and sparse attention methods show roughly two-fold worse per-
plexity scores over the quadratic OPT-1.3B baseline, while SEA achieves better
perplexity than OPT-1.3B, using roughly half the memory of OPT-1.3B. More-
over, SEA maintains an interpretable attention matrix and can utilize knowledge
distillation to lower the complexity of existing pretrained transformers. We be-
lieve that our work will have a large practical impact, as it opens the possibility of
running large transformers on resource-limited devices with less memory.
Code: https://github.com/gmlwns2000/sea-attention Heejun Lee, Jeffrey Willette, Sung Ju Hwang |
ICLR | 1 |
| 2023 | Sparse Token Transformer with Attention Back Tracking
Heejun Lee, Minki Kang, Youngwan Lee, Sung Ju Hwang |
ICLR | 1 |
| 2023 | The Rise of Chatbots in Political Campaigns: The Effects of Conversational Agents on Voting IntentionabstractThis study investigates the role and effectiveness of political chatbots’ anthropomorphism in influencing voting intentions. Our findings reveal that participants show higher voting intention when they were exposed to the political chatbot with higher (vs. lower) levels of anthropomorphic visual cues. This research also investigated two moderators underlying this relationship: message type (factual vs. emotional message) and message interactivity (high vs. low interactivity). We found that compared to the less anthropomorphized political chatbot, participants’ voting intention was enhanced when the highly anthropomorphic chatbot delivered emotional (vs. factual) messages. In terms of message interactivity, the participants exhibited higher voting intentions when the political chatbot had lower levels of anthropomorphism and higher (lower) message interactivity. Theoretical and practical implications of these findings are discussed. Yunju Kim, Heejun Lee |
Int. J. Hum. Comput. Interact. | 2 |
| 2023 | Humanizing Chatbots for Political Campaigns: How Do Voters Respond to Feasibility and Desirability Appeals from Political Chatbots?abstractAbstract Informed by the construal level theory (CLT) and accounting for anthropomorphism, we investigated the effectiveness of political chatbots in influencing voting intentions. This study employed a three-way analysis of variance test with a 2 (anthropomorphism: anthropomorphism vs. non-anthropomorphism) × 2 (message type: feasibility vs. desirability appeal) × 2 (political ideology: conservatives vs. liberals) between-subjects experiment (n = 360). The findings reveal that participants showed higher voting intention after conversing with a highly anthropomorphic chatbot (vs. non-anthropomorphic chatbot) and when the chatbot delivered desirability (vs. feasibility) appeals. Participants also exhibited a higher voting intention when the chatbot was less anthropomorphic and it delivered feasibility (vs. desirability) messages. Moreover, we identified the three-way interaction effects of anthropomorphism, message appeal type and political ideology on voting intention. These findings are discussed in terms of their theoretical and practical implications. Yunju Kim, Heejun Lee |
Interact. Comput. | 2 |
| 2022 | K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News CommentabstractOnline hate speech detection has become an important issue due to the growth of online content, but resources in languages other than English are extremely limited. We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns. The dataset consists of 109k utterances from news comments and provides a multi-label classification using 1 to 4 labels, and handles subjectivity and intersectionality. We evaluate strong baselines on K-MHaS. KR-BERT with a sub-character tokenizer outperforms others, recognizing decomposed characters in each hate speech class. Jean Lee, Taejun Lim, Heejun Lee, Bogeun Jo, Yang Sok Kim, Hee-Geun Yoon, Soyeon Caren Han |
COLING | 3 |
| 2022 | Falling in Love with Virtual Reality Art: A New Perspective on 3D Immersive Virtual Reality for Future Sustaining Art ConsumptionabstractMany art activities have been trialed in 360-degree virtual reality (VR) and this has become an increasingly representative way toward sustainable consumption of art. However, understanding audience responses to the VR art remains a matter of debate and investigation. By employing the uses and gratifications (U&G) approach, this study explores motives for and consequences of watching 360-degree VR art. The identified motives are three dimensions: pursuing learning from entertainment, pursuing social conformity, and pursuing convenience. Among the three use motives, it was found that pursuing learning from entertainment is the most significant factor influencing audiences’ transportation experience while watching 360-degree VR art, and audiences’ innovativeness functions as a moderator in the relationship between transportation and the attitude toward the VR content. The theoretical and practical implications of this study are discussed. Yunju Kim, Heejun Lee |
Int. J. Hum. Comput. Interact. | 2 |
| 2018 | Clothing Attribute Extraction Using Convolutional Neural Networks
Sangmin Jo, Heejun Lee, Mijin Noh, Yang Sok Kim |
PKAW | 3 |