Nelson F. Liu

dblp:203/9152 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Question answering and dialogue systems · 72% Language models and text generation · 22% Information extraction and text analysis · 7%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 70% Knowledge graphs · 30%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval-augmented generation
0.912025
Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models · EMNLP 2025
Natural language and speech › Question answering and dialogue systems
question answering evaluation
0.712023
Do Question Answering Modeling Improvements Hold Across Benchmarks? · ACL (1) 2023
Natural language and speech › Language models and text generation › large language model › knowledge in language models
knowledge-grounded language model
0.412019
Barack's Wife Hillary: Using Knowledge Graphs for Fact-Aware Language Modeling · ACL (1) 2019
Natural language and speech › Question answering and dialogue systems
machine reading comprehension
0.412019
Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning · EMNLP/IJCNLP (1) 2019
Natural language and speech › Question answering and dialogue systems
open-domain question answering
0.212023
Do Question Answering Modeling Improvements Hold Across Benchmarks? · ACL (1) 2023
Natural language and speech › Information extraction and text analysis
coreference resolution
0.112019
Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning · EMNLP/IJCNLP (1) 2019

Methods — techniques the papers use, named apart from their topics

embedding model · 1.7long-context language models · 0.9long-context language model · 0.9neural language model · 0.8knowledge graph fact selection and copying · 0.8model ranking analysis · 0.7benchmark comparison · 0.7
YearPublicationVenuePosition
2025 Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models
abstract
With the rise of long-context language models (LMs) capable of processing tens of thousands of tokens in a single context window, do multi-stage retrieval-augmented generation (RAG) pipelines still offer measurable benefits over simpler, single-stage approaches?To assess this question, we conduct a controlled evaluation for QA tasks under systematically scaled token budgets, comparing two recent multi-stage pipelines, ReadAgent and RAPTOR, against three baselines, including DOS RAG (Document's Original Structure RAG), a simple retrieve-then-read method that preserves original passage order.Despite its straightforward design, DOS RAG consistently matches or outperforms more intricate methods on multiple long-context QA benchmarks.We trace this strength to a combination of maintaining source fidelity and document structure, prioritizing recall within effective context windows, and favoring simplicity over added pipeline complexity.We recommend establishing DOS RAG as a simple yet strong baseline for future RAG evaluations, paired with stateof-the-art embedding and language models, and benchmarked under matched token budgets, to ensure that added pipeline complexity is justified by clear performance gains as models continue to improve. 1
Alex Laitenberger, Christopher D. Manning, Nelson F. Liu
EMNLP3
2024 Lost in the Middle: How Language Models Use Long Contexts
abstract
Abstract While recent language models have the ability to take long contexts as input, relatively little is known about how well they use longer context. We analyze the performance of language models on two tasks that require identifying relevant information in their input contexts: multi-document question answering and key-value retrieval. We find that performance can degrade significantly when changing the position of relevant information, indicating that current language models do not robustly make use of information in long input contexts. In particular, we observe that performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models. Our analysis provides a better understanding of how language models use their input context and provides new evaluation protocols for future long-context language models.
Nelson F. Liu, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang
Trans. Assoc. Comput. Linguistics1
2023 Do Question Answering Modeling Improvements Hold Across Benchmarks?
abstract
Do question answering (QA) modeling improvements (e.g., choice of architecture and training procedure) hold consistently across the diverse landscape of QA benchmarks?To study this question, we introduce the notion of concurrence-two benchmarks have high concurrence on a set of modeling approaches if they rank the modeling approaches similarly.We measure the concurrence between 32 QA benchmarks on a set of 20 diverse modeling approaches and find that human-constructed benchmarks have high concurrence amongst themselves, even if their passage and question distributions are very different.Surprisingly, even downsampled human-constructed benchmarks (i.e., collecting less data) and programmatically-generated benchmarks (e.g., cloze-formatted examples) have high concurrence with human-constructed benchmarks.These results indicate that, despite years of intense community focus on a small number of benchmarks, the modeling improvements studied hold broadly.
Nelson F. Liu, Robin Jia, Percy Liang
ACL (1)1
2019 Barack's Wife Hillary: Using Knowledge Graphs for Fact-Aware Language Modeling
abstract
Modeling human language requires the ability to not only generate fluent text but also encode factual knowledge.However, traditional language models are only capable of remembering facts seen at training time, and often have difficulty recalling them.To address this, we introduce the knowledge graph language model (KGLM), a neural language model with mechanisms for selecting and copying facts from a knowledge graph that are relevant to the context.These mechanisms enable the model to render information it has never seen before, as well as generate out-of-vocabulary tokens.We also introduce the Linked WikiText-2 dataset, 1 a corpus of annotated text aligned to the Wikidata knowledge graph whose contents (roughly) match the popular WikiText-2 benchmark (Merity et al., 2017).In experiments, we demonstrate that the KGLM achieves significantly better performance than a strong baseline language model.We additionally compare different language models' ability to complete sentences requiring factual knowledge, and show that the KGLM outperforms even very large language models in generating facts.
Robert L. Logan IV, Nelson F. Liu, Matthew E. Peters, Matt Gardner 0001, Sameer Singh 0001
ACL (1)2
2019 Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning
abstract
Pradeep Dasigi, Nelson F. Liu, Ana Marasović, Noah A. Smith, Matt Gardner. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Pradeep Dasigi, Nelson F. Liu, Ana Marasovic, Noah A. Smith, Matt Gardner 0001
EMNLP/IJCNLP (1)2