VLDB 2026 Research / reviewers in the wild / expert
Nathaniel Weir
dblp:218/5179
· DBLP profile ↗
14ranked-venue papers
9as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
Language models and text generation · 45% Knowledge representation and reasoning · 26% Question answering and dialogue systems · 12% | |
| Databases, data mining, and information retrieval
3 papers |
Data models and query languages · 50% Machine learning and data management · 28% Information retrieval · 22% | |
| Theoretical computer science
2 papers |
Automated reasoning and model checking · 82% Logic in computer science · 18% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 24 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning |
1.5 | 2 | 2024 | NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning · IJCAI 2024 TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning · EMNLP 2024 |
Data models and query languages › natural language interface
natural language interface to database |
1.1 | 3 | 2020 | DBPal: A Fully Pluggable NL2SQL Training Pipeline · SIGMOD Conference 2020 Bootstrapping an End-to-End Natural Language Interface for Databases · SIGMOD Conference 2019 DBPal: A Learned NL-Interface for Databases · SIGMOD Conference 2018 |
Automated reasoning and model checking › automated reasoning › mathematical reasoning
autoformalization |
1.0 | 1 | 2026 | A Neurosymbolic Approach to Natural Language Formalization and Verification · CAV (2) 2026 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
grounded question answering |
0.9 | 1 | 2025 | From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-Answering · ICLR 2025 |
Machine learning › Efficient and distributed learning › distillation
knowledge distillation from language models |
0.9 | 1 | 2025 | From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-Answering · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.9 | 1 | 2025 | SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses · AAAI 2025 |
Machine learning and data management › training data management
training data generation |
0.8 | 2 | 2020 | DBPal: A Fully Pluggable NL2SQL Training Pipeline · SIGMOD Conference 2020 Bootstrapping an End-to-End Natural Language Interface for Databases · SIGMOD Conference 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
automated reasoning |
0.8 | 1 | 2024 | TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems
dialogue generation |
0.8 | 1 | 2024 | Ontologically Faithful Generation of Non-Player Character Dialogues · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems › dialogue generation
knowledge-grounded dialogue generation |
0.8 | 1 | 2024 | Ontologically Faithful Generation of Non-Player Character Dialogues · EMNLP 2024 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.8 | 1 | 2024 | Learning to Reason via Program Generation, Emulation, and Search · NeurIPS 2024 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.8 | 1 | 2024 | Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic · EMNLP 2024 |
Computer vision › Video understanding and tracking
video question answering |
0.8 | 1 | 2024 | TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning · EMNLP 2024 |
Program synthesis and code generation › neural program synthesis
LLM-based program synthesis |
0.8 | 1 | 2024 | Learning to Reason via Program Generation, Emulation, and Search · NeurIPS 2024 |
Natural language and speech › Language models and text generation › text generation
conditional text generation |
0.4 | 1 | 2020 | COD3S: Diverse Generation with Discrete Semantic Signatures · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation › text generation
diverse text generation |
0.4 | 1 | 2020 | COD3S: Diverse Generation with Discrete Semantic Signatures · EMNLP (1) 2020 |
Information retrieval › query suggestion
query auto-completion |
0.3 | 1 | 2018 | DBPal: A Learned NL-Interface for Databases · SIGMOD Conference 2018 |
Information retrieval
query formulation |
0.3 | 1 | 2018 | DBPal: A Learned NL-Interface for Databases · SIGMOD Conference 2018 |
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL |
0.3 | 1 | 2018 | DBPal: A Learned NL-Interface for Databases · SIGMOD Conference 2018 |
Natural language and speech › Language models and text generation › large language model safety
LLM guardrail |
0.3 | 1 | 2026 | A Neurosymbolic Approach to Natural Language Formalization and Verification · CAV (2) 2026 |
Natural language and speech › Language models and text generation
text generation |
0.3 | 1 | 2025 | SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses · AAAI 2025 |
Machine learning › Trustworthy machine learning › interpretability
explainable reasoning |
0.2 | 1 | 2024 | NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning · IJCAI 2024 |
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity |
0.1 | 1 | 2020 | COD3S: Diverse Generation with Discrete Semantic Signatures · EMNLP (1) 2020 |
Machine learning › Transfer learning and domain adaptation
cross-domain learning |
0.1 | 1 | 2019 | Bootstrapping an End-to-End Natural Language Interface for Databases · SIGMOD Conference 2019 |
Methods — techniques the papers use, named apart from their topics
in-context learning · 2.3large language model · 2.0automated reasoning · 2.0p-relevance · 0.9knowledge distillation · 0.9generative evaluation · 0.9entailment · 0.9discriminative evaluation · 0.9deep learning · 0.8program search · 0.8neuro-symbolic reasoning · 0.8informal logic formalization · 0.8entailment tree search · 0.8decompositional inference · 0.8synthetic data generation · 0.4neural machine translation · 0.4multi-task learning · 0.4cross-domain learning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Neurosymbolic Approach to Natural Language Formalization and VerificationabstractAbstract Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits their adoption in regulated industries like finance and healthcare that operate under strict policies. To address this limitation, we launched Automated Reasoning checks (ARc) : a public service that (1) uses LLMs with optional human guidance to formalize natural language policies, allowing fine-grained control of the formalization process, and (2) uses inference-time autoformalization to validate logical correctness of natural language statements against those policies. ARc performs multiple redundant formalization steps at inference time, checking the formalizations for semantic equivalence. Our benchmarks show that ARc exceeds 99% soundness and achieves a near-zero false positive rate in identifying logical validity. Our approach produces auditable artifacts that substantiate the verification outcomes and can be used to improve the original text. ARc is the first commercial offering from a major cloud provider to integrate automated reasoning into a generative AI guardrail. Chenyang An, Sam Bayless, Stefano Buliani, Darion Cassel, Byron Cook, Duncan Clough, Rémi Delmas, Nafi Diallo, Ferhat Erata, Nick Feng, Dimitra Giannakopoulou, Aman Goel, Aditya Gokhale, Joe Hendrix, Victor Heorhiadi, Marc Hudak, Dejan Jovanovic, Andrew M. Kent, Benjamin Kiesl-Reiter, Jeffrey J. Kuna, Nadia Labai, Joe Lilien, Divya Raghunathan, Zvonimir Rakamaric, Niloofar Razavi, Michael Tautschnig, Ali Torkamani, Nathaniel Weir, Michael W. Whalen, Jianan Yao |
CAV (2) | 28 |
| 2025 | SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated ResponsesabstractCan LLMs consistently improve their previous outputs for better results? For this to be true, LLMs would need to be better at discriminating among previously-generated alternatives, than generating initial responses. We explore the validity of this hypothesis in practice. We first formulate a unified framework that allows us to compare the generative and discriminative capability of any model on any task. In our resulting experimental analysis of several open-source and industrial LLMs, we observe that model’s are not reliably better at discriminating among previously-generated alternatives than generating initial responses. This finding challenges the notion that LLMs may be able to enhance their performance only through their own judgment. Dongwei Jiang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, Daniel Khashabi |
AAAI | 4 |
| 2025 | From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-AnsweringabstractRecent reasoning methods (e.g., chain-of-thought) help users understand how language models (LMs) answer a single question, but they do little to reveal the LM’s overall understanding, or “theory,” about the question’s topic, making it still hard to trust the model. Our goal is to materialize such theories - here called microtheories (a linguistic analog of logical microtheories) - as a set of sentences encapsulating an LM’s core knowledge about a topic. These statements systematically work together to entail answers to a set of questions to both engender trust and improve performance. Our approach is to first populate a knowledge store with (model-generated) sentences that entail answers to training questions, and then distill those down to a core microtheory which is concise, general, and non-redundant. We show that, when added to a general corpus (e.g., Wikipedia), microtheories can supply critical information not necessarily present in the corpus, improving both a model’s ability to ground its answers to verifiable knowledge (i.e., show how answers are systematically entailed by documents in the corpus, grounding up to +8% more answers), and the accuracy of those grounded answers (up to +8% absolute). We also show that, in a human evaluation in the medical domain, our distilled microtheories contain a significantly higher concentration of topically critical facts than the non-distilled knowledge store. Finally, we show we can quantify the coverage of a microtheory for a topic (characterized by a dataset) using a notion of p-relevance. Together, these suggest that microtheories are an efficient distillation of an LM’s topic-relevant knowledge, that they can usefully augment existing corpora, and can provide both performance gains and an interpretable, verifiable window into the model’s knowledge of a topic. Nathaniel Weir, Bhavana Dalvi, Orion Weller, Oyvind Tafjord, Sam Hornstein, Alexander Sabol, Peter A. Jansen, Benjamin Van Durme, Peter Clark |
ICLR | 1 |
| 2024 | "According to . . . ": Prompting Language Models Improves Quoting from Pre-Training DataabstractOrion Weller, Marc Marone, Nathaniel Weir, Dawn Lawrie, Daniel Khashabi, Benjamin Van Durme. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Orion Weller, Marc Marone, Nathaniel Weir, Dawn J. Lawrie, Daniel Khashabi, Benjamin Van Durme |
EACL (1) | 3 |
| 2024 | TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video ReasoningabstractIt is challenging for models to understand complex, multimodal content such as television clips, and this is in part because video-language models often rely on single-modality reasoning and lack interpretability.To combat these issues we propose TV-TREES, the first multimodal entailment tree generator.TV-TREES serves as an approach to video understanding that promotes interpretable joint-modality reasoning by searching for trees of entailment relationships between simple text-video evidence and higher-level conclusions that prove question-answer pairs.We also introduce the task of multimodal entailment tree generation to evaluate reasoning quality.Our method's performance on the challenging TVQA benchmark demonstrates interpretable, state-of-theart zero-shot performance on full clips, illustrating that multimodal entailment tree generation can be a best-of-both-worlds alternative to black-box systems. Kate Sanders 0002, Nathaniel Weir, Benjamin Van Durme |
EMNLP | 2 |
| 2024 | Enhancing Systematic Decompositional Natural Language Inference Using Informal LogicabstractNathaniel Weir, Kate Sanders, Orion Weller, Shreya Sharma, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Jansen, Peter Clark, Benjamin Van Durme. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Nathaniel Weir, Kate Sanders 0002, Orion Weller, Shreya Sharma 0010, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi, Oyvind Tafjord, Peter A. Jansen, Peter Clark, Benjamin Van Durme |
EMNLP | 1 |
| 2024 | Ontologically Faithful Generation of Non-Player Character DialoguesabstractWe introduce a language generation dataset grounded in a popular video game.KNUDGE (KNowledge Constrained User-NPC Dialogue GEneration) requires models to produce trees of dialogue between video game characters that accurately reflect quest and entity specifications stated in natural language.KNUDGE is constructed from side quest dialogues drawn directly from game data of Obsidian Entertainment's The Outer Worlds, leading to real-world complexities in generation: (1) utterances must remain faithful to the game lore, including character personas and backstories; (2) a dialogue must accurately reveal new quest details to the human player; and (3) dialogues are large trees as opposed to linear chains of utterances.We report results for a set of neural generation models using supervised and in-context learning techniques; we find competent performance but room for future work addressing the challenges of creating realistic, game-quality dialogues. Nathaniel Weir, Ryan Thomas, Randolph D'Amore, Kellie Hill, Benjamin Van Durme, Harsh Jhamtani |
EMNLP | 1 |
| 2024 | NELLIE: A Neuro-Symbolic Inference Engine for Grounded, Compositional, and Explainable Reasoning
Nathaniel Weir, Peter Clark, Benjamin Van Durme |
IJCAI | 1 |
| 2024 | Learning to Reason via Program Generation, Emulation, and SearchabstractProgram synthesis with language models (LMs) has unlocked a large set of reasoning abilities; code-tuned LMs have proven adept at generating programs that solve a wide variety of algorithmic symbolic manipulation tasks (e.g. word concatenation). However, not all reasoning tasks are easily expressible as code, e.g. tasks involving commonsense reasoning, moral decision-making, and sarcasm understanding. Our goal is to extend a LM’s program synthesis skills to such tasks and evaluate the results via pseudo-programs, namely Python programs where some leaf function calls are left undefined. To that end, we propose, Code Generation and Emulated EXecution (COGEX). COGEX works by (1) training LMs to generate pseudo-programs and (2) teaching them to emulate their generated program’s execution, including those leaf functions, allowing the LM’s knowledge to fill in the execution gaps; and (3) using them to search over many programs to find an optimal one. To adapt the COGEX model to a new task, we introduce a method for performing program search to find a single program whose pseudo-execution yields optimal performance when applied to all the instances of a given dataset. We show that our approach yields large improvements compared to standard in-context learning approaches on a battery of tasks, both algorithmic and soft reasoning. This result thus demonstrates that code synthesis can be applied to a much broader class of problems than previously considered. Nathaniel Weir, Muhammad Khalifa, Linlu Qiu, Orion Weller, Peter Clark |
NeurIPS | 1 |
| 2020 | Probing Neural Language Models for Human Tacit Assumptions
Nathaniel Weir, Adam Poliak, Benjamin Van Durme |
CogSci | 1 |
| 2020 | COD3S: Diverse Generation with Discrete Semantic SignaturesabstractWe present COD3S, a novel method for generating semantically diverse sentences using neural sequence-to-sequence (seq2seq) models.Conditioned on an input, seq2seq models typically produce semantically and syntactically homogeneous sets of sentences and thus perform poorly on one-to-many sequence generation tasks.Our two-stage approach improves output diversity by conditioning generation on locality-sensitive hash (LSH)-based semantic sentence codes whose Hamming distances highly correlate with human judgments of semantic textual similarity.Though it is generally applicable, we apply COD3S to causal generation, the task of predicting a proposition's plausible causes or effects.We demonstrate through automatic and human evaluation that responses produced using our method exhibit improved diversity without degrading task performance. Nathaniel Weir, João Sedoc, Benjamin Van Durme |
EMNLP (1) | 1 |
| 2020 | DBPal: A Fully Pluggable NL2SQL Training PipelineabstractNatural language is a promising alternative interface to DBMSs because it enables non-technical users to formulate complex questions in a more concise manner than SQL. Recently, deep learning has gained traction for translating natural language to SQL, since similar ideas have been successful in the related domain of machine translation. However, the core problem with existing deep learning approaches is that they require an enormous amount of training data in order to provide accurate translations. This training data is extremely expensive to curate, since it generally requires humans to manually annotate natural language examples with the corresponding SQL queries (or vice versa). Based on these observations, we propose DBPal, a new approach that augments existing deep learning techniques in order to improve the performance of models for natural language to SQL translation. More specifically, we present a novel training pipeline that automatically generates synthetic training data in order to (1) improve overall translation accuracy, (2) increase robustness to linguistic variation, and (3) specialize the model for the target database. As we show, our DBPal training pipeline is able to improve both the accuracy and linguistic robustness of state-of-the-art natural language to SQL translation models. Nathaniel Weir, Prasetya Ajie Utama, Alex Galakatos, Andrew Crotty, Amir Ilkhechi, Shekar Ramaswamy, Rohin Bhushan, Nadja Geisler, Benjamin Hättasch, Steffen Eger, Ugur Çetintemel, Carsten Binnig |
SIGMOD Conference | 1 |
| 2019 | Bootstrapping an End-to-End Natural Language Interface for DatabasesabstractThe ability to extract insights from data is critical for decision making. Intuitive natural language interfaces to databases provide non-technical users with an effective way to formulate complex questions and information needs efficiently and effectively. A recent trend in the area of Natural Language Interfaces for Databases (NLIDBs) has been the use of neural machine translation models to synthesize executable Structured Query Language (SQL) queries from natural language utterances. The main bottleneck in this type of approach is the acquisition of examples for training the model. Recent work has assumed access to a rich manually-curated training set for a given target database. However, this assumption ignores the large manual overhead required to curate the training set for any new database. As a result, NLIDB systems that can simply 'plug in' to any new database and perform effectively for naive users have yet to make their way into commercial products. Here we present DBPal, an end-to-end NLIDB framework in which a neural translation model is trained for any new database schema with minimal manual overhead. In addition to being the first off-the-shelf, neural machine translationbased system of its kind, the contributions of our project are 1) its use of a synthetic training set generation pipeline used to bootstrap a translation model without requiring manually curated data, and 2) its use of state-of-the-art multi-task and cross-domain learning techniques that increases the robustness of the translation model towards unseen linguistic phenomena in new domains. In experiments we show that our system can achieve competitive performance on the recently released benchmarks for nl-to-sql translation. Through ablation experiments we show the benefit of using cross-domain learning techniques on the performance of the system. In a user study we show that DBPal outperforms a well-known rule-based NLIDB and performs comparably to an approach using a similar neural model that relies on manually curated data. Nathaniel Weir, Prasetya Ajie Utama |
SIGMOD Conference | 1 |
| 2018 | DBPal: A Learned NL-Interface for DatabasesabstractIn this demo, we present DBPal, a novel data exploration tool with a natural language interface. DBPal leverages recent advances in deep models to make query understanding more robust in the following ways: First, DBPal uses novel machine translation models to translate natural language statements to SQL, making the translation process more robust to paraphrasing and linguistic variations. Second, to support the users in phrasing questions without knowing the database schema and the query features, DBPal provides a learned auto-completion model that suggests to users partial query extensions during query formulation and thus helps to write complex queries. Fuat Basik, Benjamin Hättasch, Amir Ilkhechi, Arif Usta, Shekar Ramaswamy, Prasetya Ajie Utama, Nathaniel Weir, Carsten Binnig, Ugur Çetintemel |
SIGMOD Conference | 7 |