Haejun Lee

dblp:95/6562 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-6284-2903ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Deep learning architectures and training · 29% Question answering and dialogue systems · 25% Language models and text generation · 25%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 62% Information retrieval · 38%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
open-domain question answering
1.632022
You Only Need One Model for Open-domain Question Answering · EMNLP 2022
FiE: Building a Global Probability Space by Leveraging Early Fusion in Encoder for Open-Domain Question Answering · EMNLP 2022
Answering Open-Domain Questions of Varying Reasoning Steps from Text · EMNLP (1) 2021
Machine learning › Deep learning architectures and training
transformer
1.632024
Transformers Get Stable: An End-to-End Signal Propagation Theory for Language Models · ICML 2024
Span-Selective Linear Attention Transformers for Effective and Robust Schema-Guided Dialogue State Tracking · ACL (1) 2023
You Only Need One Model for Open-domain Question Answering · EMNLP 2022
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.912025
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs · ACL (1) 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
LLM distillation
0.912025
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs · ACL (1) 2025
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
0.712023
Span-Selective Linear Attention Transformers for Effective and Robust Schema-Guided Dialogue State Tracking · ACL (1) 2023
Machine learning › Deep learning architectures and training › attention mechanism › efficient attention
linear attention
0.712023
Span-Selective Linear Attention Transformers for Effective and Robust Schema-Guided Dialogue State Tracking · ACL (1) 2023
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning
0.512021
Answering Open-Domain Questions of Varying Reasoning Steps from Text · EMNLP (1) 2021
Natural language and speech › Language models and text generation › natural language understanding › question answering
answerability prediction
0.412020
NeurQuRI: Neural Question Requirement Inspector for Answerability Prediction in Machine Reading Comprehension · ICLR 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
discourse representation
0.412020
SLM: Learning a Discourse Language Representation with Sentence Unshuffling · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.412020
NeurQuRI: Neural Question Requirement Inspector for Answerability Prediction in Machine Reading Comprehension · ICLR 2020
Natural language and speech › Language models and text generation › large language model training
pre-training objectives
0.412020
SLM: Learning a Discourse Language Representation with Sentence Unshuffling · EMNLP (1) 2020
Recommender systems › interactive recommendation
conversational recommendation
0.212016
Quote Recommendation in Dialogue using Deep Neural Network · SIGIR 2016
Recommender systems › content recommendation
quote recommendation
0.212016
Quote Recommendation in Dialogue using Deep Neural Network · SIGIR 2016
Information retrieval
multi-hop retrieval
0.112021
Answering Open-Domain Questions of Varying Reasoning Steps from Text · EMNLP (1) 2021
Information retrieval
retrieval models
0.112021
Answering Open-Domain Questions of Varying Reasoning Steps from Text · EMNLP (1) 2021
Machine learning › Deep learning architectures and training › transformer
hierarchical transformer
0.112020
SLM: Learning a Discourse Language Representation with Sentence Unshuffling · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

sparse sampling · 0.9knowledge distillation · 0.9signal propagation theory · 0.8initialization scheme · 0.8schema-guided modeling · 0.7linear attention · 0.7passage-wise attention · 0.6global attention · 0.6end-to-end training · 0.6early fusion · 0.6transformer · 0.5multi-task learning · 0.5recurrent neural network · 0.2deep learning · 0.2convolutional neural network · 0.2
YearPublicationVenuePosition
2025 Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
abstract
Anshumann Anshumann, Mohd Abbas Zaidi, Akhil Kedia, Jinwoo Ahn, Taehwak Kwon, Kangwook Lee, Haejun Lee, Joohyung Lee. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Anshumann, Mohd Abbas Zaidi, Akhil Kedia, Jinwoo Ahn, Taehwak Kwon, Kangwook Lee 0004, Haejun Lee
ACL (1)7
2024 Transformers Get Stable: An End-to-End Signal Propagation Theory for Language Models
abstract
In spite of their huge success, transformer models remain difficult to scale in depth. In this work, we develop a unified signal propagation theory and provide formulae that govern the moments of the forward and backward signal through the transformer model. Our framework can be used to understand and mitigate vanishing/exploding gradients, rank collapse, and instability associated with high attention scores. We also propose DeepScaleLM, an initialization and scaling scheme that conserves unit output/gradient moments throughout the model, enabling the training of very deep models with 1000 layers. We find that transformer models could be much deeper - our deep models with fewer parameters outperform shallow models in Language Modeling, Speech Translation, and Image Classification, across encoder-only, decoder-only and encoder-decoder variants, for both Pre-LN and Post-LN transformers, for multiple datasets and model sizes. These improvements also translate into improved performance on downstream Question Answering tasks and improved robustness for Image Classification.
Akhil Kedia, Mohd Abbas Zaidi, Sushil Khyalia, Jungho Jung, Harshith Goka, Haejun Lee
ICML6
2023 Span-Selective Linear Attention Transformers for Effective and Robust Schema-Guided Dialogue State Tracking
abstract
In schema-guided dialogue state tracking models estimate the current state of a conversation using natural language descriptions of the service schema for generalization to unseen services.Prior generative approaches which decode slot values sequentially do not generalize well to variations in schema, while discriminative approaches separately encode history and schema and fail to account for inter-slot and intent-slot dependencies.We introduce SPLAT, a novel architecture which achieves better generalization and efficiency than prior approaches by constraining outputs to a limited prediction space.At the same time, our model allows for rich attention among descriptions and history while keeping computation costs constrained by incorporating linear-time attention.We demonstrate the effectiveness of our model on the Schema-Guided Dialogue (SGD) and Mul-tiWOZ datasets.Our approach significantly improves upon existing models achieving 85.3 JGA on the SGD dataset.Further, we show increased robustness on the SGD-X benchmark: our model outperforms the more than 30× larger D3ST-XXL model by 5.0 points.
Björn Bebensee, Haejun Lee
ACL (1)2
2022 FiE: Building a Global Probability Space by Leveraging Early Fusion in Encoder for Open-Domain Question Answering
abstract
Generative models have recently started to outperform extractive models in Open Domain Question Answering, largely by leveraging their decoder to attend over multiple encoded passages and combining their information.However, generative models tend to be larger than extractive models due to the need for a decoder, run slower during inference due to auto-regressive decoder beam search, and their generated output often suffers from hallucinations.We propose to extend transformer encoders with the ability to fuse information from multiple passages, using global representation to provide cross-sample attention over all tokens across samples.Furthermore, we propose an alternative answer span probability calculation to better aggregate answer scores in the global space of all samples.Using our proposed method, we outperform the current stateof-the-art method by 2.5 Exact Match score on the Natural Question dataset while using only 25% of parameters and 35% of the latency during inference, and 4.4 Exact Match on We-bQuestions dataset.When coupled with synthetic data augmentation, we outperform larger models on the TriviaQA dataset as well.The latency and parameter savings of our method make it particularly attractive for open-domain question answering, as these models are often compute-intensive.Global Tokens Global Encoder Layer Question + Passage 1 Passage Encoder Layer Question + Passage 2 Passage Encoder Layer Span Classifier 𝑃𝑃 𝑠𝑠 (𝑠𝑠 1 ) Span Classifier 𝑃𝑃 𝑠𝑠 (𝑠𝑠 2 ) 𝑃𝑃 𝐴𝐴 ("𝑁𝑁𝑁𝑁𝑁𝑁 𝑌𝑌𝑌𝑌𝑌𝑌𝑌𝑌") Σ Multi-Head Attention (Global Encoder Layer)x N Encoder Layers Q KV
Akhil Kedia, Mohd Abbas Zaidi, Haejun Lee
EMNLP3
2022 You Only Need One Model for Open-domain Question Answering
abstract
Recent approaches to Open-domain Question Answering refer to an external knowledge base using a retriever model, optionally rerank passages with a separate reranker model and generate an answer using another reader model.Despite performing related tasks, the models have separate parameters and are weaklycoupled during training.We propose casting the retriever and the reranker as internal passage-wise attention mechanisms applied sequentially within the transformer architecture and feeding computed representations to the reader, with the hidden representations progressively refined at each stage.This allows us to use a single question answering model trained end-to-end, which is a more efficient use of model capacity and also leads to better gradient flow.We present a pre-training method to effectively train this architecture and evaluate our model on the Natural Questions and TriviaQA open datasets.For a fixed parameter budget, our model outperforms the previous state-of-the-art model by 1.0 and 0.7 exact match scores.Reading Layer Reranking Layer Retrieval Layer Encoder Query Passage 1 Passage 2 Passage 3 Passage 4 Passage n Encoder Encoder Encoder Encoder Encoder Encoder … … Passage-wise Disjoint Attention Encoder Encoder Encoder Encoder Passage-wise Joint Attention Concat Query and Passage Answer Concat All Decoder Encoder Encoder Encoder … Decoder
Haejun Lee, Akhil Kedia, Ashwin Paranjape, Christopher D. Manning, Kyoung-Gu Woo
EMNLP1
2021 Answering Open-Domain Questions of Varying Reasoning Steps from Text
abstract
We develop a unified system to answer directly from text open-domain questions that may require a varying number of retrieval steps.We employ a single multi-task transformer model to perform all the necessary subtasks-retrieving supporting facts, reranking them, and predicting the answer from all retrieved documents-in an iterative fashion.We avoid crucial assumptions of previous work that do not transfer well to real-world settings, including exploiting knowledge of the fixed number of retrieval steps required to answer each question or using structured metadata like knowledge bases or web links that have limited availability.Instead, we design a system that can answer open-domain questions on any text collection without prior knowledge of reasoning complexity.To emulate this setting, we construct a new benchmark, called BeerQA, by combining existing one-and twostep datasets with a new collection of 530 questions that require three Wikipedia pages to answer, unifying Wikipedia corpora versions in the process.We show that our model demonstrates competitive performance on both existing benchmarks and this new benchmark.We make the new benchmark available at https: //beerqa.github.io/.
Peng Qi 0003, Haejun Lee, Tg Sido, Christopher D. Manning
EMNLP (1)2
2020 SLM: Learning a Discourse Language Representation with Sentence Unshuffling
abstract
We introduce Sentence-level Language Modeling, a new pre-training objective for learning a discourse language representation in a fully self-supervised manner.Recent pre-training methods in NLP focus on learning either bottom or top-level language representations: contextualized word representations derived from language model objectives at one extreme and a whole sequence representation learned by order classification of two given textual segments at the other.However, these models are not directly encouraged to capture representations of intermediate-size structures that exist in natural languages such as sentences and the relationships among them.To that end, we propose a new approach to encourage learning of a contextualized sentence-level representation by shuffling the sequence of input sentences and training a hierarchical transformer model to reconstruct the original ordering.Through experiments on downstream tasks such as GLUE, SQuAD, and DiscoEval, we show that this feature of our model improves the performance of the original BERT by large margins.
Haejun Lee, Drew A. Hudson, Kangwook Lee 0004, Christopher D. Manning
EMNLP (1)1
2020 NeurQuRI: Neural Question Requirement Inspector for Answerability Prediction in Machine Reading Comprehension
Seohyun Back, Sai Chetan Chinthakindi, Akhil Kedia, Haejun Lee, Jaegul Choo
ICLR4
2016 Quote Recommendation in Dialogue using Deep Neural Network
abstract
Quotes, or quotations, are well known phrases or sentences that we use for various purposes such as emphasis, elaboration, and humor. In this paper, we introduce a task of recommending quotes which are suitable for given dialogue context and we present a deep learning recommender system which combines recurrent neural network and convolutional neural network in order to learn semantic representation of each utterance and construct a sequence model for the dialog thread. We collected a large set of twitter dialogues with quote occurrences in order to evaluate proposed recommender system. Experimental results show that our approach outperforms not only the other state-of-the-art algorithms in quote recommendation task, but also other neural network based methods built for similar tasks.
Hanbit Lee, Yeonchan Ahn, Haejun Lee, Seungdo Ha, Sang-goo Lee
SIGIR3