VLDB 2026 Research / reviewers in the wild / expert
Dohyeon Lee
dblp:297/3811
· DBLP profile ↗
16ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Markovian Forgetfulness: Episodic Memory for Reasoning-Intensive RetrievalabstractReasoning-intensive information retrieval uses large language models to solve complex queries via multi-step reasoning.However, existing methods have critical limitations.Chain-of-Thought (CoT) approaches suffer from inefficiency, while state-based methods, despite better token efficiency, often fall into reasoning cycles that trap the query refinement process.To address these issues, we propose Episodic Memory for Retrieval (EMR), which enhances the state-based framework with an episodic memory.This module stores the full history of prior states for a query, allowing the model to avoid repetition of such cycles.Experiments on the BRIGHT benchmark show that EMR consistently outperforms both CoT and state-based baselines.Moreover, it is highly token-efficient, reducing token usage by 72% on average.Our results show that episodic memory is an effective and tokenefficient mechanism for reasoning-intensive retrieval.The gains also generalize across different base models and stay efficient in terms of end-to-end latency.The code is available in https://github.com/ldilab/EMR. Dohyeon Lee, Yeonseok Jeong, Seung-won Hwang |
ACL (1) | 1 |
| 2025 | Query-focused Referentiability Learning for Zero-shot RetrievalabstractJaeyoung Kim, Dohyeon Lee, Seung-won Hwang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dohyeon Lee, Seung-won Hwang |
NAACL (Long Papers) | 2 |
| 2025 | tRAG: Term-level Retrieval-Augmented Generation for Domain-Adaptive RetrievalabstractDohyeon Lee, Jongyoon Kim, Jihyuk Kim, Seung-won Hwang, Joonsuk Park. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dohyeon Lee, Jongyoon Kim, Jihyuk Kim, Seung-won Hwang, Joonsuk Park |
NAACL (Long Papers) | 1 |
| 2024 | Colorful Intersections and Tverberg Partitions
Michael Gene Dobbins, Andreas F. Holmsen, Dohyeon Lee |
SoCG | 3 |
| 2024 | Chaining Event Spans for Temporal Relation GroundingabstractAccurately understanding temporal relations between events is a critical building block of diverse tasks, such as temporal reading comprehension (TRC) and relation extraction (TRE).For example in TRC, we need to understand the temporal semantic differences between the following two questions that are lexically nearidentical: "What finished right before the decision?"or "What finished right after the decision?".To discern the two questions, existing solutions have relied on answer overlaps as a proxy label to contrast similar and dissimilar questions.However, we claim that answer overlap can lead to unreliable results, due to spurious overlaps of two dissimilar questions with coincidentally identical answers.To address the issue, we propose a novel approach that elicits proper reasoning behaviors through a module for predicting time spans of events.We introduce the Timeline Reasoning Network (TRN) operating in a two-step inductive reasoning process: In the first step model initially answers each question with semantic and syntactic information.The next step chains multiple questions on the same event to predict a timeline, which is then used to ground the answers.Results on the TORQUE and TB-dense, TRC and TRE tasks respectively, demonstrate that TRN outperforms previous methods by effectively resolving the spurious overlaps using the predicted timeline 1 . Dohyeon Lee, Seung-won Hwang |
EACL (1) | 2 |
| 2024 | Interventional Speech Noise Injection for ASR Generalizable Spoken Language UnderstandingabstractRecently, pre-trained language models (PLMs) have been increasingly adopted in spoken language understanding (SLU).However, automatic speech recognition (ASR) systems frequently produce inaccurate transcriptions, leading to noisy inputs for SLU models, which can significantly degrade their performance.To address this, our objective is to train SLU models to withstand ASR errors by exposing them to noises commonly observed in ASR systems, referred to as ASR-plausible noises.Speech noise injection (SNI) methods have pursued this objective by introducing ASR-plausible noises, but we argue that these methods are inherently biased towards specific ASR systems, or ASR-specific noises.In this work, we propose a novel and less biased augmentation method of introducing the noises that are plausible to any ASR system, by cutting off the non-causal effect of noises.Experimental results and analyses demonstrate the effectiveness of our proposed methods in enhancing the robustness and generalizability of SLU models against unseen ASR systems by introducing more diverse and plausible ASR noises in advance. YeonJoon Jung, Jaeseong Lee 0002, Seungtaek Choi, Dohyeon Lee, Seung-won Hwang |
EMNLP | 4 |
| 2024 | ScriptMix: Mixing Scripts for Low-resource Language ParsingabstractJaeseong Lee, Dohyeon Lee, Seung-won Hwang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jaeseong Lee 0002, Dohyeon Lee, Seung-won Hwang |
NAACL-HLT | 2 |
| 2024 | HIL: Hybrid Isotropy Learning for Zero-shot Performance in Dense retrievalabstractJaeyoung Kim, Dohyeon Lee, Seung-won Hwang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Dohyeon Lee, Seung-won Hwang |
NAACL-HLT | 2 |
| 2023 | Script, Language, and Labels: Overcoming Three Discrepancies for Low-Resource Language SpecializationabstractAlthough multilingual pretrained models (mPLMs) enabled support of various natural language processing in diverse languages, its limited coverage of 100+ languages lets 6500+ languages remain ‘unseen’. One common approach for an unseen language is specializing the model for it as target, by performing additional masked language modeling (MLM) with the target language corpus. However, we argue that, due to the discrepancy from multilingual MLM pretraining, a naive specialization as such can be suboptimal. Specifically, we pose three discrepancies to overcome. Script and linguistic discrepancy of the target language from the related seen languages, hinder a positive transfer, for which we propose to maximize representation similarity, unlike existing approaches maximizing overlaps. In addition, label space for MLM prediction can vary across languages, for which we propose to reinitialize top layers for a more effective adaptation. Experiments over four different language families and three tasks shows that our method improves the task performance of unseen languages with statistical significance, while previous approach fails to. Jaeseong Lee 0002, Dohyeon Lee, Seung-won Hwang |
AAAI | 2 |
| 2023 | On Complementarity Objectives for Hybrid RetrievalabstractDense retrieval has shown promising results in various information retrieval tasks, and hybrid retrieval, combined with the strength of sparse retrieval, has also been actively studied.A key challenge in hybrid retrieval is to make sparse and dense complementary to each other.Existing models have focused on dense models to capture "residual" features neglected in the sparse models.Our key distinction is to show how this notion of residual complementarity is limited, and propose a new objective, denoted as RoC (Ratio of Complementarity), which captures a fuller notion of complementarity.We propose a two-level orthogonality designed to improve RoC, then show that the improved RoC of our model, in turn, improves the performance of hybrid retrieval.Our method outperforms all state-of-the-art methods on three representative IR benchmarks: MSMARCO-Passage, Natural Questions, and TREC Ro-bust04, with statistical significance.Our finding is also consistent in various adversarial settings. Dohyeon Lee, Seung-won Hwang, Kyungjae Lee 0002, Seungtaek Choi, Sunghyun Park 0005 |
ACL (1) | 1 |
| 2023 | C2LIR: Continual Cross-Lingual Transfer for Low-Resource Information Retrieval
Jaeseong Lee 0002, Dohyeon Lee, Seung-won Hwang |
ECIR (2) | 2 |
| 2023 | A Highly Maneuverable Flying Squirrel Drone with Controllable Foldable WingsabstractTypical drones with multi rotors are generally less maneuverable due to unidirectional thrust, which may be unfavorable to agile flight in very narrow and confined spaces. This paper suggests a new bio-inspired drone that is empowered with high maneuverability in a lightweight and easy-to-carry way. The proposed flying squirrel inspired drone has controllable foldable wings to cover a wider range of flight attitudes and provide more maneuverable flight capability with stable tracking performance. The wings of a drone are fabricated with silicone membranes and sophisticatedly controlled by reinforcement learning based on human-demonstrated data. Specially, such learning based wing control serves to capture even the complex aerodynamics that are often impossible to model mathematically. It is shown through experiment that the proposed flying squirrel drone intentionally induces aerodynamic drag and hence provides the desired additional repulsive force even under saturated mechanical thrust. This work is very meaningful in demonstrating the potential of biomimicry and machine learning for realizing an animal-like agile drone. Jun-Gill Kang, Dohyeon Lee, Soohee Han |
IROS | 2 |
| 2022 | PLM-based World Models for Text-based GamesabstractWorld models have improved the ability of reinforcement learning agents to operate in a sample efficient manner, by being trained to predict plausible changes in the underlying environment.As the core tasks of world models are future prediction and commonsense understanding, our claim is that pre-trained language models (PLMs) already provide a strong base upon which to build world models.Worldformer is a recently proposed world model for text-based game environments, based only partially on PLM and transformers.Our distinction is to fully leverage PLMs as actionable world models in text-based game environments, by reformulating generation as constrained decoding which decomposes actions into verb templates and objects.We show that our model improves future valid action prediction and graph change prediction. 1 Additionally, we show that our model better reflects commonsense than standard PLM. YeonJoon Jung, Dohyeon Lee, Seung-won Hwang |
EMNLP | 3 |
| 2022 | FastqCLS: a FASTQ compressor for long-read sequencing via read reordering using a novel scoring modelabstractMOTIVATION: Over the past decades, vast amounts of genome sequencing data have been produced, requiring an enormous level of storage capacity. The time and resources needed to store and transfer such data cause bottlenecks in genome sequencing analysis. To resolve this issue, various compression techniques have been proposed to reduce the size of original FASTQ raw sequencing data, but these remain suboptimal. Long-read sequencing has become dominant in genomics, whereas most existing compression methods focus on short-read sequencing only. RESULTS: We designed a compression algorithm based on read reordering using a novel scoring model for reducing FASTQ file size with no information loss. We integrated all data processing steps into a software package called FastqCLS and provided it as a Docker image for ease of installation and execution to help users easily install and run. We compared our method with existing major FASTQ compression tools using benchmark datasets. We also included new long-read sequencing data in this validation. As a result, FastqCLS outperformed in terms of compression ratios for storing long-read sequencing data. AVAILABILITY AND IMPLEMENTATION: FastqCLS can be downloaded from https://github.com/krlucete/FastqCLS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dohyeon Lee, Giltae Song |
Bioinform. | 1 |
| 2021 | Robustifying Multi-hop QA through Pseudo-Evidentiality TrainingabstractKyungjae Lee, Seung-won Hwang, Sang-eun Han, Dohyeon Lee. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kyungjae Lee 0002, Seung-won Hwang, Sang-eun Han, Dohyeon Lee |
ACL/IJCNLP (1) | 4 |
| 2021 | SCOPA: Soft Code-Switching and Pairwise Alignment for Zero-Shot Cross-lingual TransferabstractThe recent advent of cross-lingual embeddings, such as multilingual BERT (mBERT), provides a strong baseline for zero-shot cross-lingual transfer. There also exists increasing research attention to reduce the alignment discrepancy of cross-lingual embeddings between source and target languages, via generating code-switched sentences by substituting randomly selected words in the source languages with their counterparts of the target languages. Although these approaches improve the performance, naively code-switched sentences can have inherent limitations. In this paper, we propose SCOPA, a novel technique to improve the performance of zero-shot cross-lingual transfer. Instead of using the embeddings of code-switched sentences directly, SCOPA mixes them softly with the embeddings of original sentences. In addition, SCOPA utilizes an additional pairwise alignment objective, which aligns the vector differences of word pairs instead of word-level embeddings, in order to transfer contextualized information between different languages while preserving language-specific information. Experiments on the PAWS-X and MLDoc dataset show the effectiveness of SCOPA. Dohyeon Lee, Jaeseong Lee 0002, Gyewon Lee, Byung-Gon Chun, Seung-won Hwang |
CIKM | 1 |