EDBT 2026 Demo / reviewers in the wild / expert
Howard Yen
dblp:348/5988
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 89% Trustworthy machine learning · 8% Question answering and dialogue systems · 2% | |
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 100% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model |
1.7 | 2 | 2025 | Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking · EMNLP 2025 How to Train Long-Context Language Models (Effectively) · ACL (1) 2025 |
Information retrieval
retrieval models |
1.7 | 2 | 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025 Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking · EMNLP 2025 |
Information retrieval
retrieval-augmented generation |
1.1 | 2 | 2025 | HELMET: How to Evaluate Long-context Models Effectively and Thoroughly · ICLR 2025 Enabling Large Language Models to Generate Text with Citations · EMNLP 2023 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.0 | 2 | 2025 | Long-Context Language Modeling with Parallel Context Encoding · ACL (1) 2024 Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Natural language and speech › Language models and text generation
faithful generation |
0.9 | 1 | 2025 | Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Natural language and speech › Language models and text generation
hallucination mitigation |
0.9 | 1 | 2025 | Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | How to Train Long-Context Language Models (Effectively) · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
long-context language model evaluation |
0.9 | 1 | 2025 | HELMET: How to Evaluate Long-context Models Effectively and Thoroughly · ICLR 2025 |
Natural language and speech › Language models and text generation › text generation
long-form text generation |
0.9 | 1 | 2025 | Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › language model interpretability
retrieval heads |
0.9 | 1 | 2025 | Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking · EMNLP 2025 |
Information retrieval › document retrieval › text search
long-context retrieval |
0.9 | 1 | 2025 | Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking · EMNLP 2025 |
Information retrieval
ranking |
0.9 | 1 | 2025 | Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking · EMNLP 2025 |
Information retrieval › document retrieval
reasoning-intensive retrieval |
0.9 | 1 | 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025 |
Information retrieval
reranking |
0.9 | 1 | 2025 | Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking · EMNLP 2025 |
Natural language and speech › Language models and text generation › language modeling
long-context language modeling |
0.8 | 1 | 2024 | Long-Context Language Modeling with Parallel Context Encoding · ACL (1) 2024 |
Natural language and speech › Language models and text generation › text generation › scientific text generation
citation generation |
0.7 | 1 | 2023 | Enabling Large Language Models to Generate Text with Citations · EMNLP 2023 |
Natural language and speech › Language models and text generation › text generation
factual text generation |
0.3 | 1 | 2025 | Precise Information Control in Long-Form Text Generation · NeurIPS 2025 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
long-context question answering |
0.3 | 1 | 2025 | Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking · EMNLP 2025 |
Information retrieval
evaluation |
0.3 | 1 | 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025 |
Information retrieval › evaluation › test collection
retrieval benchmark |
0.3 | 1 | 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
model-based evaluation · 1.7few-shot prompting · 1.7dense retrieval · 1.7attention analysis · 1.7weakly supervised preference learning · 0.9supervised fine-tuning · 0.9post-training · 0.9data mixing · 0.9continued pretraining · 0.9cross-attention · 0.8prompting strategies · 0.7automatic evaluation metrics · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How to Train Long-Context Language Models (Effectively)abstractWe study continued training and supervised fine-tuning (SFT) of a language model (LM) to make effective use of long-context information.We first establish a reliable evaluation protocol to guide model development-instead of perplexity or simple needle-in-a-haystack (NIAH) tests, we use a broad set of long-context downstream tasks, and we evaluate models after SFT as this better reveals long-context abilities.Supported by our robust evaluations, we run thorough experiments to decide the data mix for continued pre-training, the instruction tuning dataset, and many other design choices such as position extrapolation.We find that (1) code repositories and books are excellent sources of long data, but it is crucial to combine them with high-quality short-context data;(2) training with a sequence length beyond the evaluation length boosts long-context performance; (3) for SFT, using only short instruction datasets yields strong performance on long-context tasks.Our final model, ProLong-8B, which is initialized from Llama-3 and trained on 40B tokens, demonstrates state-ofthe-art long-context performance among similarly sized models at a length of 128K.Pro-Long outperforms Llama-3.1-8B-Instruct on the majority of long-context tasks despite using only 5% as many tokens during long-context training.Additionally, ProLong can effectively process up to 512K tokens, one of the longest context windows of publicly available LMs. Tianyu Gao 0001, Alexander Wettig, Howard Yen, Danqi Chen 0001 |
ACL (1) | 3 |
| 2025 | Improving Interpersonal Communication by Simulating Audiences with Large Language Models
Ryan Liu 0001, Howard Yen, Raja Marjieh, Thomas L. Griffiths 0001, Ranjay Krishna |
CogSci | 2 |
| 2025 | Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-rankingabstractRecent work has identified retrieval heads (Wu et al., 2025b), a subset of attention heads responsible for retrieving salient information in long-context language models (LMs), as measured by their copy-paste behavior in Needlein-a-Haystack tasks.In this paper, we introduce QRHEAD (Query-Focused Retrieval Head), an improved set of attention heads that enhance retrieval from long context.We identify QRHEAD by aggregating attention scores with respect to the input query, using a handful of examples from real-world tasks (e.g., long-context QA).We further introduce QR-RETRIEVER, an efficient and effective retriever that uses the accumulated attention mass of QRHEAD as retrieval scores.We use QR-RETRIEVER for long-context reasoning by selecting the most relevant parts with the highest retrieval scores.On multi-hop reasoning tasks LongMemEval and CLIPPER, this yields over 10% performance gains over full context and outperforms strong dense retrievers.We also evaluate QRRETRIEVER as a re-ranker on the BEIR benchmark and find that it achieves strong zero-shot performance, outperforming other LLM-based re-rankers such as RankGPT.Further analysis shows that both the querycontext attention scoring and task selection are crucial for identifying QRHEAD with strong downstream utility.Overall, our work contributes a general-purpose retriever and offers interpretability insights into the long-context capabilities of LMs. 1 Wuwei Zhang, Fangcong Yin, Howard Yen, Danqi Chen 0001, Xi Ye 0003 |
EMNLP | 3 |
| 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalabstractExisting retrieval benchmarks primarily consist of information-seeking queries (e.g., aggregated questions from search engines) where keyword or semantic-based retrieval is usually sufficient. However, many complex real-world queries require in-depth reasoning to identify relevant documents that go beyond surface form matching. For example, finding documentation for a coding question requires understanding the logic and syntax of the functions involved. To better benchmark retrieval on such challenging queries, we introduce BRIGHT, the first text retrieval benchmark that requires intensive reasoning to retrieve relevant documents. Our dataset consists of 1,398 real-world queries spanning diverse domains such as economics, psychology, mathematics, coding, and more. These queries are drawn from naturally occurring or carefully curated human data. Extensive evaluation reveals that even state-of-the-art retrieval models perform poorly on BRIGHT. The leading model on the MTEB leaderboard (Muennighoff et al., 2023), which achieves a score of 59.0 nDCG@10,1 produces a score of nDCG@10 of 18.0 on BRIGHT. We show that incorporating explicit reasoning about the query improves retrieval performance by up to 12.2 points. Moreover, incorporating retrieved documents from the top-performing retriever boosts question answering performance by over 6.6 points. We believe that BRIGHT paves the way for future research on retrieval systems in more realistic and challenging settings. Hongjin Su, Howard Yen, Mengzhou Xia, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Zachary S. Siegel, Michael Tang, Ruoxi Sun 0002, Jinsung Yoon, Sercan Ö. Arik, Danqi Chen 0001, Tao Yu 0009 |
ICLR | 2 |
| 2025 | HELMET: How to Evaluate Long-context Models Effectively and ThoroughlyabstractMany benchmarks exist for evaluating long-context language models (LCLMs), yet developers often rely on synthetic tasks such as needle-in-a-haystack (NIAH) or an arbitrary subset of tasks. However, it remains unclear whether these benchmarks reflect the diverse downstream applications of LCLMs, and such inconsistencies further complicate model comparison. We investigate the underlying reasons behind these practices and find that existing benchmarks often provide noisy signals due to limited coverage of applications, insufficient context lengths, unreliable metrics, and incompatibility with base models. In this work, we introduce HELMET (How to Evaluate Long-context Models Effectively and Thoroughly), a comprehensive benchmark encompassing seven diverse, application-centric categories. We also address several issues in previous benchmarks by adding controllable lengths up to 128K tokens, model-based evaluation for reliable metrics, and few-shot prompting for robustly evaluating base models. Consequently, we demonstrate that HELMET offers more reliable and consistent rankings of frontier LCLMs. Through a comprehensive study of 59 LCLMs, we find that (1) synthetic tasks like NIAH do not reliably predict downstream performance; (2) the diverse categories in HELMET exhibit distinct trends and low correlations with each other; and (3) while most LCLMs achieve perfect NIAH scores, open-source models significantly lag behind closed ones when tasks require full-context reasoning or following complex instructions---the gap widens as length increases. Finally, we recommend using our RAG tasks for fast model development, as they are easy to run and better predict other downstream performance; ultimately, we advocate for a holistic evaluation across diverse tasks. Howard Yen, Tianyu Gao 0001, Minmin Hou, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, Danqi Chen 0001 |
ICLR | 1 |
| 2025 | Precise Information Control in Long-Form Text GenerationabstractA central challenge in language models (LMs) is faithfulness hallucination: the generation of information unsubstantiated by input context. To study this problem, we propose Precise Information Control (PIC), a new task formulation that requires models to generate long-form outputs grounded in a provided set of short self-contained statements, without adding any unsupported ones. PIC includes a full setting that tests a model’s ability to include exactly all input claims, and a partial setting that requires the model to selectively incorporate only relevant claims. We present PIC-Bench, a benchmark of eight long-form generation tasks (e.g., summarization, biography generation) adapted to the PIC setting, where LMs are supplied with well-formed, verifiable input claims. Our evaluation of a range of open and proprietary LMs on PIC-Bench reveals that, surprisingly, state-of-the-art LMs still hallucinate against user-provided input in over 70% of generations. To alleviate this lack of faithfulness, we introduce a post-training framework that uses a weakly supervised preference data construction method to train an 8B PIC-LM with stronger PIC ability—improving from 69.1% to 91.0% F1 in the full PIC setting. When integrated into end-to-end factual generation pipelines, PIC-LM improves exact match recall by 17.1% on ambiguous QA with retrieval, and factual precision by 30.5% on a birthplace fact-checking task, underscoring the potential of precisely grounded generation. Jacqueline He, Howard Yen, Margaret Li, Shuyue Stella Li, Yulia Tsvetkov, Danqi Chen 0001, Pang Wei W. Koh, Luke Zettlemoyer |
NeurIPS | 2 |
| 2024 | Long-Context Language Modeling with Parallel Context EncodingabstractExtending large language models (LLMs) to process longer inputs is crucial for a wide range of applications.However, the substantial computational cost of transformers and limited generalization of positional encoding restrict the size of their context window.We introduce Context Expansion with Parallel Encoding (CEPE ), a framework that can be applied to any existing decoder-only LLMs to extend their context window.CEPE employs a small encoder to process long inputs chunk by chunk, enabling the frozen decoder to utilize additional contexts via cross-attention.CEPE is efficient, generalizable, and versatile: trained with 8K-token documents, it extends the context window of LLAMA-2 to 128K tokens, offering 10× the throughput with only 1/6 of the memory.CEPE yields strong performance on language modeling and in-context learning.CEPE also excels in retrieval-augmented applications, while existing long-context models degenerate with retrieved contexts.We further introduce a CEPE variant that can extend the context window of instruction-tuned models using only unlabeled data, and showcase its effectiveness on LLAMA-2-CHAT, leading to a strong instruction-following model that can leverage very long contexts on downstream tasks. 1Chapter 01: Dune ... Howard Yen, Tianyu Gao 0001, Danqi Chen 0001 |
ACL (1) | 1 |
| 2023 | Enabling Large Language Models to Generate Text with CitationsabstractLarge language models (LLMs) have emerged as a widely-used tool for information seeking, but their generated outputs are prone to hallucination.In this work, our aim is to allow LLMs to generate text with citations, improving their factual correctness and verifiability.Existing work mainly relies on commercial search engines and human evaluation, making it challenging to reproduce and compare different modeling approaches.We propose ALCE, the first benchmark for Automatic LLMs' Citation Evaluation.ALCE collects a diverse set of questions and retrieval corpora and requires building end-to-end systems to retrieve supporting evidence and generate answers with citations.We develop automatic metrics along three dimensions-fluency, correctness, and citation quality-and demonstrate their strong correlation with human judgements.Our experiments with state-of-the-art LLMs and novel prompting strategies show that current systems have considerable room for improvement-For example, on the ELI5 dataset, even the best models lack complete citation support 50% of the time.Our analyses further highlight promising future directions, including developing better retrievers, advancing long-context LLMs, and improving the ability to synthesize information from multiple sources. 1 When did the US break away from England? Question Short answers (from the dataset) Tianyu Gao 0001, Howard Yen, Jiatong Yu, Danqi Chen 0001 |
EMNLP | 2 |