Nathaniel Weir

dblp:218/5179 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Language models and text generation · 45% Knowledge representation and reasoning · 26% Question answering and dialogue systems · 12%
Databases, data mining, and information retrieval
3 papers
Data models and query languages · 50% Machine learning and data management · 28% Information retrieval · 22%
Theoretical computer science
2 papers
Automated reasoning and model checking · 82% Logic in computer science · 18%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 24 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning
1.522024
NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning · IJCAI 2024
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning · EMNLP 2024
Data models and query languages › natural language interface
natural language interface to database
1.132020
DBPal: A Fully Pluggable NL2SQL Training Pipeline · SIGMOD Conference 2020
Bootstrapping an End-to-End Natural Language Interface for Databases · SIGMOD Conference 2019
DBPal: A Learned NL-Interface for Databases · SIGMOD Conference 2018
Automated reasoning and model checking › automated reasoning › mathematical reasoning
autoformalization
1.012026
A Neurosymbolic Approach to Natural Language Formalization and Verification · CAV (2) 2026
Natural language and speech › Language models and text generation › natural language understanding › question answering
grounded question answering
0.912025
From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-Answering · ICLR 2025
Machine learning › Efficient and distributed learning › distillation
knowledge distillation from language models
0.912025
From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-Answering · ICLR 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses · AAAI 2025
Machine learning and data management › training data management
training data generation
0.822020
DBPal: A Fully Pluggable NL2SQL Training Pipeline · SIGMOD Conference 2020
Bootstrapping an End-to-End Natural Language Interface for Databases · SIGMOD Conference 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning
automated reasoning
0.812024
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning · EMNLP 2024
Natural language and speech › Question answering and dialogue systems
dialogue generation
0.812024
Ontologically Faithful Generation of Non-Player Character Dialogues · EMNLP 2024
Natural language and speech › Question answering and dialogue systems › dialogue generation
knowledge-grounded dialogue generation
0.812024
Ontologically Faithful Generation of Non-Player Character Dialogues · EMNLP 2024
Natural language and speech › Language models and text generation
large language model reasoning
0.812024
Learning to Reason via Program Generation, Emulation, and Search · NeurIPS 2024
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference
0.812024
Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic · EMNLP 2024
Computer vision › Video understanding and tracking
video question answering
0.812024
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning · EMNLP 2024
Program synthesis and code generation › neural program synthesis
LLM-based program synthesis
0.812024
Learning to Reason via Program Generation, Emulation, and Search · NeurIPS 2024
Natural language and speech › Language models and text generation › text generation
conditional text generation
0.412020
COD3S: Diverse Generation with Discrete Semantic Signatures · EMNLP (1) 2020
Natural language and speech › Language models and text generation › text generation
diverse text generation
0.412020
COD3S: Diverse Generation with Discrete Semantic Signatures · EMNLP (1) 2020
Information retrieval › query suggestion
query auto-completion
0.312018
DBPal: A Learned NL-Interface for Databases · SIGMOD Conference 2018
Information retrieval
query formulation
0.312018
DBPal: A Learned NL-Interface for Databases · SIGMOD Conference 2018
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL
0.312018
DBPal: A Learned NL-Interface for Databases · SIGMOD Conference 2018
Natural language and speech › Language models and text generation › large language model safety
LLM guardrail
0.312026
A Neurosymbolic Approach to Natural Language Formalization and Verification · CAV (2) 2026
Natural language and speech › Language models and text generation
text generation
0.312025
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses · AAAI 2025
Machine learning › Trustworthy machine learning › interpretability
explainable reasoning
0.212024
NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning · IJCAI 2024
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity
0.112020
COD3S: Diverse Generation with Discrete Semantic Signatures · EMNLP (1) 2020
Machine learning › Transfer learning and domain adaptation
cross-domain learning
0.112019
Bootstrapping an End-to-End Natural Language Interface for Databases · SIGMOD Conference 2019

Methods — techniques the papers use, named apart from their topics

in-context learning · 2.3large language model · 2.0automated reasoning · 2.0p-relevance · 0.9knowledge distillation · 0.9generative evaluation · 0.9entailment · 0.9discriminative evaluation · 0.9deep learning · 0.8program search · 0.8neuro-symbolic reasoning · 0.8informal logic formalization · 0.8entailment tree search · 0.8decompositional inference · 0.8synthetic data generation · 0.4neural machine translation · 0.4multi-task learning · 0.4cross-domain learning · 0.4
YearPublicationVenuePosition
2026 A Neurosymbolic Approach to Natural Language Formalization and Verification
abstract
Abstract Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits their adoption in regulated industries like finance and healthcare that operate under strict policies. To address this limitation, we launched Automated Reasoning checks (ARc) : a public service that (1) uses LLMs with optional human guidance to formalize natural language policies, allowing fine-grained control of the formalization process, and (2) uses inference-time autoformalization to validate logical correctness of natural language statements against those policies. ARc performs multiple redundant formalization steps at inference time, checking the formalizations for semantic equivalence. Our benchmarks show that ARc exceeds 99% soundness and achieves a near-zero false positive rate in identifying logical validity. Our approach produces auditable artifacts that substantiate the verification outcomes and can be used to improve the original text. ARc is the first commercial offering from a major cloud provider to integrate automated reasoning into a generative AI guardrail.
Chenyang An, Sam Bayless, Stefano Buliani, Darion Cassel, Byron Cook, Duncan Clough, Rémi Delmas, Nafi Diallo, Ferhat Erata, Nick Feng, Dimitra Giannakopoulou, Aman Goel, Aditya Gokhale, Joe Hendrix, Victor Heorhiadi, Marc Hudak, Dejan Jovanovic, Andrew M. Kent, Benjamin Kiesl-Reiter, Jeffrey J. Kuna, Nadia Labai, Joe Lilien, Divya Raghunathan, Zvonimir Rakamaric, Niloofar Razavi, Michael Tautschnig, Ali Torkamani, Nathaniel Weir, Michael W. Whalen, Jianan Yao
CAV (2)28
2025 SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
abstract
Can LLMs consistently improve their previous outputs for better results? For this to be true, LLMs would need to be better at discriminating among previously-generated alternatives, than generating initial responses. We explore the validity of this hypothesis in practice. We first formulate a unified framework that allows us to compare the generative and discriminative capability of any model on any task. In our resulting experimental analysis of several open-source and industrial LLMs, we observe that model’s are not reliably better at discriminating among previously-generated alternatives than generating initial responses. This finding challenges the notion that LLMs may be able to enhance their performance only through their own judgment.
Dongwei Jiang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, Daniel Khashabi
AAAI4
2025 From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-Answering
abstract
Recent reasoning methods (e.g., chain-of-thought) help users understand how language models (LMs) answer a single question, but they do little to reveal the LM’s overall understanding, or “theory,” about the question’s topic, making it still hard to trust the model. Our goal is to materialize such theories - here called microtheories (a linguistic analog of logical microtheories) - as a set of sentences encapsulating an LM’s core knowledge about a topic. These statements systematically work together to entail answers to a set of questions to both engender trust and improve performance. Our approach is to first populate a knowledge store with (model-generated) sentences that entail answers to training questions, and then distill those down to a core microtheory which is concise, general, and non-redundant. We show that, when added to a general corpus (e.g., Wikipedia), microtheories can supply critical information not necessarily present in the corpus, improving both a model’s ability to ground its answers to verifiable knowledge (i.e., show how answers are systematically entailed by documents in the corpus, grounding up to +8% more answers), and the accuracy of those grounded answers (up to +8% absolute). We also show that, in a human evaluation in the medical domain, our distilled microtheories contain a significantly higher concentration of topically critical facts than the non-distilled knowledge store. Finally, we show we can quantify the coverage of a microtheory for a topic (characterized by a dataset) using a notion of p-relevance. Together, these suggest that microtheories are an efficient distillation of an LM’s topic-relevant knowledge, that they can usefully augment existing corpora, and can provide both performance gains and an interpretable, verifiable window into the model’s knowledge of a topic.
Nathaniel Weir, Bhavana Dalvi, Orion Weller, Oyvind Tafjord, Sam Hornstein, Alexander Sabol, Peter A. Jansen, Benjamin Van Durme, Peter Clark
ICLR1
2024 "According to . . . ": Prompting Language Models Improves Quoting from Pre-Training Data
abstract
Orion Weller, Marc Marone, Nathaniel Weir, Dawn Lawrie, Daniel Khashabi, Benjamin Van Durme. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Orion Weller, Marc Marone, Nathaniel Weir, Dawn J. Lawrie, Daniel Khashabi, Benjamin Van Durme
EACL (1)3
2024 TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
abstract
It is challenging for models to understand complex, multimodal content such as television clips, and this is in part because video-language models often rely on single-modality reasoning and lack interpretability.To combat these issues we propose TV-TREES, the first multimodal entailment tree generator.TV-TREES serves as an approach to video understanding that promotes interpretable joint-modality reasoning by searching for trees of entailment relationships between simple text-video evidence and higher-level conclusions that prove question-answer pairs.We also introduce the task of multimodal entailment tree generation to evaluate reasoning quality.Our method's performance on the challenging TVQA benchmark demonstrates interpretable, state-of-theart zero-shot performance on full clips, illustrating that multimodal entailment tree generation can be a best-of-both-worlds alternative to black-box systems.
Kate Sanders 0002, Nathaniel Weir, Benjamin Van Durme
EMNLP2
2024 Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
abstract
Nathaniel Weir, Kate Sanders, Orion Weller, Shreya Sharma, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Jansen, Peter Clark, Benjamin Van Durme. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Nathaniel Weir, Kate Sanders 0002, Orion Weller, Shreya Sharma 0010, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi, Oyvind Tafjord, Peter A. Jansen, Peter Clark, Benjamin Van Durme
EMNLP1
2024 Ontologically Faithful Generation of Non-Player Character Dialogues
abstract
We introduce a language generation dataset grounded in a popular video game.KNUDGE (KNowledge Constrained User-NPC Dialogue GEneration) requires models to produce trees of dialogue between video game characters that accurately reflect quest and entity specifications stated in natural language.KNUDGE is constructed from side quest dialogues drawn directly from game data of Obsidian Entertainment's The Outer Worlds, leading to real-world complexities in generation: (1) utterances must remain faithful to the game lore, including character personas and backstories; (2) a dialogue must accurately reveal new quest details to the human player; and (3) dialogues are large trees as opposed to linear chains of utterances.We report results for a set of neural generation models using supervised and in-context learning techniques; we find competent performance but room for future work addressing the challenges of creating realistic, game-quality dialogues.
Nathaniel Weir, Ryan Thomas, Randolph D'Amore, Kellie Hill, Benjamin Van Durme, Harsh Jhamtani
EMNLP1
2024 NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning
Nathaniel Weir, Peter Clark, Benjamin Van Durme
IJCAI1
2024 Learning to Reason via Program Generation, Emulation, and Search
abstract
Program synthesis with language models (LMs) has unlocked a large set of reasoning abilities; code-tuned LMs have proven adept at generating programs that solve a wide variety of algorithmic symbolic manipulation tasks (e.g. word concatenation). However, not all reasoning tasks are easily expressible as code, e.g. tasks involving commonsense reasoning, moral decision-making, and sarcasm understanding. Our goal is to extend a LM’s program synthesis skills to such tasks and evaluate the results via pseudo-programs, namely Python programs where some leaf function calls are left undefined. To that end, we propose, Code Generation and Emulated EXecution (COGEX). COGEX works by (1) training LMs to generate pseudo-programs and (2) teaching them to emulate their generated program’s execution, including those leaf functions, allowing the LM’s knowledge to fill in the execution gaps; and (3) using them to search over many programs to find an optimal one. To adapt the COGEX model to a new task, we introduce a method for performing program search to find a single program whose pseudo-execution yields optimal performance when applied to all the instances of a given dataset. We show that our approach yields large improvements compared to standard in-context learning approaches on a battery of tasks, both algorithmic and soft reasoning. This result thus demonstrates that code synthesis can be applied to a much broader class of problems than previously considered.
Nathaniel Weir, Muhammad Khalifa, Linlu Qiu, Orion Weller, Peter Clark
NeurIPS1
2020 Probing Neural Language Models for Human Tacit Assumptions
Nathaniel Weir, Adam Poliak, Benjamin Van Durme
CogSci1
2020 COD3S: Diverse Generation with Discrete Semantic Signatures
abstract
We present COD3S, a novel method for generating semantically diverse sentences using neural sequence-to-sequence (seq2seq) models.Conditioned on an input, seq2seq models typically produce semantically and syntactically homogeneous sets of sentences and thus perform poorly on one-to-many sequence generation tasks.Our two-stage approach improves output diversity by conditioning generation on locality-sensitive hash (LSH)-based semantic sentence codes whose Hamming distances highly correlate with human judgments of semantic textual similarity.Though it is generally applicable, we apply COD3S to causal generation, the task of predicting a proposition's plausible causes or effects.We demonstrate through automatic and human evaluation that responses produced using our method exhibit improved diversity without degrading task performance.
Nathaniel Weir, João Sedoc, Benjamin Van Durme
EMNLP (1)1
2020 DBPal: A Fully Pluggable NL2SQL Training Pipeline
abstract
Natural language is a promising alternative interface to DBMSs because it enables non-technical users to formulate complex questions in a more concise manner than SQL. Recently, deep learning has gained traction for translating natural language to SQL, since similar ideas have been successful in the related domain of machine translation. However, the core problem with existing deep learning approaches is that they require an enormous amount of training data in order to provide accurate translations. This training data is extremely expensive to curate, since it generally requires humans to manually annotate natural language examples with the corresponding SQL queries (or vice versa). Based on these observations, we propose DBPal, a new approach that augments existing deep learning techniques in order to improve the performance of models for natural language to SQL translation. More specifically, we present a novel training pipeline that automatically generates synthetic training data in order to (1) improve overall translation accuracy, (2) increase robustness to linguistic variation, and (3) specialize the model for the target database. As we show, our DBPal training pipeline is able to improve both the accuracy and linguistic robustness of state-of-the-art natural language to SQL translation models.
Nathaniel Weir, Prasetya Ajie Utama, Alex Galakatos, Andrew Crotty, Amir Ilkhechi, Shekar Ramaswamy, Rohin Bhushan, Nadja Geisler, Benjamin Hättasch, Steffen Eger, Ugur Çetintemel, Carsten Binnig
SIGMOD Conference1
2019 Bootstrapping an End-to-End Natural Language Interface for Databases
abstract
The ability to extract insights from data is critical for decision making. Intuitive natural language interfaces to databases provide non-technical users with an effective way to formulate complex questions and information needs efficiently and effectively. A recent trend in the area of Natural Language Interfaces for Databases (NLIDBs) has been the use of neural machine translation models to synthesize executable Structured Query Language (SQL) queries from natural language utterances. The main bottleneck in this type of approach is the acquisition of examples for training the model. Recent work has assumed access to a rich manually-curated training set for a given target database. However, this assumption ignores the large manual overhead required to curate the training set for any new database. As a result, NLIDB systems that can simply 'plug in' to any new database and perform effectively for naive users have yet to make their way into commercial products. Here we present DBPal, an end-to-end NLIDB framework in which a neural translation model is trained for any new database schema with minimal manual overhead. In addition to being the first off-the-shelf, neural machine translationbased system of its kind, the contributions of our project are 1) its use of a synthetic training set generation pipeline used to bootstrap a translation model without requiring manually curated data, and 2) its use of state-of-the-art multi-task and cross-domain learning techniques that increases the robustness of the translation model towards unseen linguistic phenomena in new domains. In experiments we show that our system can achieve competitive performance on the recently released benchmarks for nl-to-sql translation. Through ablation experiments we show the benefit of using cross-domain learning techniques on the performance of the system. In a user study we show that DBPal outperforms a well-known rule-based NLIDB and performs comparably to an approach using a similar neural model that relies on manually curated data.
Nathaniel Weir, Prasetya Ajie Utama
SIGMOD Conference1
2018 DBPal: A Learned NL-Interface for Databases
abstract
In this demo, we present DBPal, a novel data exploration tool with a natural language interface. DBPal leverages recent advances in deep models to make query understanding more robust in the following ways: First, DBPal uses novel machine translation models to translate natural language statements to SQL, making the translation process more robust to paraphrasing and linguistic variations. Second, to support the users in phrasing questions without knowing the database schema and the query features, DBPal provides a learned auto-completion model that suggests to users partial query extensions during query formulation and thus helps to write complex queries.
Fuat Basik, Benjamin Hättasch, Amir Ilkhechi, Arif Usta, Shekar Ramaswamy, Prasetya Ajie Utama, Nathaniel Weir, Carsten Binnig, Ugur Çetintemel
SIGMOD Conference7