VLDB 2026 Research / reviewers in the wild / expert
Willis Guo
dblp:371/8797
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0003-5409-7768ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Knowledge representation and reasoning · 62% Question answering and dialogue systems · 32% Language models and text generation · 5% | |
| Theoretical computer science
1 paper |
Automated reasoning and model checking · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
2.4 | 3 | 2025 | CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge · SIGIR 2025 Verifiable, Debuggable, and Repairable Commonsense Logical Reasoning via LLM-based Theory Resolution · EMNLP 2024 Right for Right Reasons: Large Language Models for Verifiable Commonsense Knowledge Graph Question Answering · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
1.6 | 2 | 2025 | CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge · SIGIR 2025 Right for Right Reasons: Large Language Models for Verifiable Commonsense Knowledge Graph Question Answering · EMNLP 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
logic-based reasoning |
0.8 | 1 | 2024 | Verifiable, Debuggable, and Repairable Commonsense Logical Reasoning via LLM-based Theory Resolution · EMNLP 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2025 | CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail Knowledge · SIGIR 2025 |
Methods — techniques the papers use, named apart from their topics
large language model · 3.1theory resolution · 1.5knowledge graph grounding · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CoLoTa: A Dataset for Entity-based Commonsense Reasoning over Long-Tail KnowledgeabstractThe rise of Large Language Models (LLMs) has redefined the AI landscape, particularly due to their ability to encode factual and commonsense knowledge, and their outstanding performance in tasks requiring reasoning. Despite these advances, hallucinations and reasoning errors remain a significant barrier to their deployment in high-stakes settings. In this work, we observe that even the most prominent LLMs, such as OpenAI-o1, suffer from high rates of reasoning errors and hallucinations on tasks requiring commonsense reasoning over obscure, long-tail entities. To investigate this limitation, we present a new dataset for Commonsense reasoning over Long-Tail entities (CoLoTa), that consists of 3,300 queries from question answering and claim verification tasks and covers a diverse range of commonsense reasoning skills. We remark that CoLoTa can also serve as a Knowledge Graph Question Answering (KGQA) dataset since the support of knowledge required to answer its queries is present in the Wikidata knowledge graph. However, as opposed to existing KGQA benchmarks that merely focus on factoid questions, our CoLoTa queries also require commonsense reasoning. Our experiments with strong LLM-based KGQA methodologies indicate their severe inability to answer queries involving commonsense reasoning. Hence, we propose CoLoTa as a novel benchmark for assessing both (i) LLM commonsense reasoning capabilities and their robustness to hallucinations on long-tail entities and (ii) the commonsense reasoning capabilities of KGQA methods. Armin Toroghi, Willis Guo, Scott Sanner |
SIGIR | 2 |
| 2024 | Right for Right Reasons: Large Language Models for Verifiable Commonsense Knowledge Graph Question AnsweringabstractKnowledge Graph Question Answering (KGQA) methods seek to answer Natural Language questions using the relational information stored in Knowledge Graphs (KGs).With the recent advancements of Large Language Models (LLMs) and their remarkable reasoning abilities, there is a growing trend to leverage them for KGQA.However, existing methodologies have only focused on answering factual questions, e.g., "In which city was Silvio Berlusconi's first wife born?", leaving questions involving commonsense reasoning that real-world users may pose more often, e.g., "Do I need separate visas to see the Venus of Willendorf and attend the Olympics this summer?"unaddressed.In this work, we first observe that existing LLM-based methods for KGQA struggle with hallucination on such questions, especially on queries targeting long-tail entities (e.g., non-mainstream and recent entities), thus hindering their applicability in real-world applications especially since their reasoning processes are not easily verifiable.In response, we propose Right for Right Reasons (R 3 ), a commonsense KGQA methodology that allows for a verifiable reasoning procedure by axiomatically surfacing intrinsic commonsense knowledge of LLMs and grounding every factual reasoning step on KG triples.Through experimental evaluations across three different tasks-question answering, claim verification, and preference matching-our findings showcase R 3 as a superior approach, outperforming existing methodologies and notably reducing instances of hallucination and reasoning errors. P o in t in ti m Armin Toroghi, Willis Guo, Mohammad Mahdi Abdollah Pour, Scott Sanner |
EMNLP | 2 |
| 2024 | Verifiable, Debuggable, and Repairable Commonsense Logical Reasoning via LLM-based Theory ResolutionabstractRecent advances in Large Language Models (LLM) have led to substantial interest in their application to commonsense reasoning tasks.Despite their potential, LLMs are susceptible to reasoning errors and hallucinations that may be harmful in use cases where accurate reasoning is critical.This challenge underscores the need for verifiable, debuggable, and repairable LLM reasoning.Recent works have made progress toward verifiable reasoning with LLMs by using them as either (i) a reasoner over an axiomatic knowledge base, or (ii) a semantic parser for use in existing logical inference systems.However, both settings are unable to extract commonsense axioms from the LLM that are not already formalized in the knowledge base, and also lack a reliable method to repair missed commonsense inferences.In this work, we present LLM-TRes, a logical reasoning framework based on the notion of "theory resolution" that allows for seamless integration of the commonsense knowledge from LLMs with a verifiable logical reasoning framework that mitigates hallucinations and facilitates debugging of the reasoning procedure as well as repair.We crucially prove that repaired axioms are theoretically guaranteed to be given precedence over flawed ones in our theory resolution inference process.We conclude by evaluating on three diverse language-based reasoning tasks -preference reasoning, deductive reasoning, and causal commonsense reasoning -and demonstrate the superior performance of LLM-TRes vs. state-of-the-art LLM-based reasoning methods in terms of both accuracy and reasoning correctness. Armin Toroghi, Willis Guo, Ali Pesaranghader, Scott Sanner |
EMNLP | 2 |