VLDB 2026 Research / reviewers in the wild / expert
Karthik Valmeekam
dblp:279/2957
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Planning, search and constraint satisfaction · 38% Language models and text generation · 36% Trustworthy machine learning · 10% | |
| Theoretical computer science
1 paper |
Automated reasoning and model checking · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 20 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › LLM agents
large language model planning |
1.5 | 2 | 2025 | On the self-verification limitations of large language models on reasoning and planning tasks · ICLR 2025 On the Planning Abilities of Large Language Models - A Critical Investigation · NeurIPS 2023 |
Natural language and speech › Language models and text generation
large language model evaluation |
1.3 | 2 | 2023 | On the Planning Abilities of Large Language Models - A Critical Investigation · NeurIPS 2023 PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change · NeurIPS 2023 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.9 | 1 | 2025 | On the self-verification limitations of large language models on reasoning and planning tasks · ICLR 2025 |
Machine learning › Trustworthy machine learning › verification
self-verification |
0.9 | 1 | 2025 | On the self-verification limitations of large language models on reasoning and planning tasks · ICLR 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › classical planning
blocks world |
0.8 | 1 | 2024 | Chain of Thoughtlessness? An Analysis of CoT in Planning · NeurIPS 2024 |
Natural language and speech › Language models and text generation › prompting
chain-of-thought prompting |
0.8 | 1 | 2024 | Chain of Thoughtlessness? An Analysis of CoT in Planning · NeurIPS 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
classical planning |
0.8 | 1 | 2024 | Chain of Thoughtlessness? An Analysis of CoT in Planning · NeurIPS 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning |
0.8 | 1 | 2024 | Position: LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks · ICML 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
domain model learning |
0.7 | 1 | 2023 | Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning · NeurIPS 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › goal reasoning
goal specification |
0.7 | 1 | 2023 | Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human Preferences · ICLR 2023 |
Machine learning › Reinforcement learning › preference learning
human preference learning |
0.7 | 1 | 2023 | Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human Preferences · ICLR 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › planning evaluation
planning benchmarks |
0.7 | 1 | 2023 | PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change · NeurIPS 2023 |
Machine learning › Reinforcement learning
reward learning |
0.7 | 1 | 2023 | Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human Preferences · ICLR 2023 |
Machine learning › Trustworthy machine learning › interpretability › local explanation
contrastive explanation |
0.5 | 1 | 2021 | RADAR-X: An Interactive Interface Pairing Contrastive Explanations with Revised Plan Suggestions · AAAI 2021 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
explicable planning |
0.5 | 1 | 2021 | RADAR-X: An Interactive Interface Pairing Contrastive Explanations with Revised Plan Suggestions · AAAI 2021 |
Human-AI interaction
decision support |
0.5 | 1 | 2021 | RADAR-X: An Interactive Interface Pairing Contrastive Explanations with Revised Plan Suggestions · AAAI 2021 |
Natural language and speech › Language models and text generation
in-context learning |
0.2 | 1 | 2024 | Chain of Thoughtlessness? An Analysis of CoT in Planning · NeurIPS 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.2 | 1 | 2023 | Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning · NeurIPS 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search |
0.2 | 1 | 2023 | On the Planning Abilities of Large Language Models - A Critical Investigation · NeurIPS 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
plan generation |
0.2 | 1 | 2023 | PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
self-critique · 1.7iterative prompting · 1.7external verification · 1.7large language model · 1.4neuro-symbolic integration · 0.8model-based verification · 0.8chain-of-thought prompting · 0.8benchmark suite design · 0.7back-prompting · 0.7PDDL validator · 0.7preference elicitation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the self-verification limitations of large language models on reasoning and planning tasksabstractThere has been considerable divergence of opinion on the reasoning abilities of Large Language Models (LLMs).
While the initial optimism that reasoning might emerge automatically with scale has been tempered thanks to a slew of counterexamples--ranging from multiplication to simple planning--there persists a wide spread belief that LLMs can self-critique and improve their own solutions in an iterative fashion.
This belief seemingly rests on the assumption that verification of correctness should be easier than generation--a rather classical argument from computational complexity--which should be irrelevant to LLMs to the extent that what they are doing is approximate retrieval.
In this paper, we set out to systematically investigate the effectiveness of iterative prompting in the context of reasoning and planning.
We present a principled empirical study of the performance of GPT-4 in three domains: Game of 24, Graph Coloring, and STRIPS planning.
We experiment both with the model critiquing its own answers and with an external correct reasoner verifying proposed solutions.
In each case, we analyze whether the content of criticisms actually affects bottom line performance, and whether we can ablate elements of the augmented system without losing performance. We observe significant performance collapse
with self-critique and significant performance gains with sound external verification.
We also note that merely re-prompting with a sound verifier maintains most of the benefits of more involved setups. Kaya Stechly, Karthik Valmeekam, Subbarao Kambhampati |
ICLR | 2 |
| 2024 | Position: LLMs Can't Plan, But Can Help Planning in LLM-Modulo FrameworksabstractWe argue that auto-regressive LLMs cannot, by themselves, do planning or self-verification (which is after all a form of reasoning), and shed some light on the reasons for misunderstandings in the literature. We will also argue that LLMs should be viewed as universal approximate knowledge sources that have much more meaningful roles to play in planning/reasoning tasks beyond simple front-end/back-end format translators. We present a vision of LLM-Modulo Frameworks that combine the strengths of LLMs with external model-based verifiers in a tighter bi-directional interaction regime. We will show how the models driving the external verifiers themselves can be acquired with the help of LLMs. We will also argue that rather than simply pipelining LLMs and symbolic components, this LLM-Modulo Framework provides a better neuro-symbolic approach that offers tighter integration between LLMs and symbolic components, and allows extending the scope of model-based planning/reasoning regimes towards more flexible knowledge, problem and preference specifications. Subbarao Kambhampati, Karthik Valmeekam, Lin Guan 0003, Mudit Verma, Kaya Stechly, Siddhant Bhambri, Lucas Saldyt, Anil Murthy |
ICML | 2 |
| 2024 | Chain of Thoughtlessness? An Analysis of CoT in PlanningabstractLarge language model (LLM) performance on reasoning problems typically does not generalize out of distribution. Previous work has claimed that this can be mitigated with chain of thought prompting--a method of demonstrating solution procedures--with the intuition that it is possible to in-context teach an LLM an algorithm for solving the problem.
This paper presents a case study of chain of thought on problems from Blocksworld, a classical planning domain, and examines the performance of two state-of-the-art LLMs across two axes: generality of examples given in prompt, and complexity of problems queried with each prompt. While our problems are very simple, we only find meaningful performance improvements from chain of thought prompts when those prompts are exceedingly specific to their problem class, and that those improvements quickly deteriorate as the size n of the query-specified stack grows past the size of stacks shown in the examples.
We also create scalable variants of three domains commonly studied in previous CoT papers and demonstrate the existence of similar failure modes.
Our results hint that, contrary to previous claims in the literature, CoT's performance improvements do not stem from the model learning general algorithmic procedures via demonstrations but depend on carefully engineering highly problem specific prompts. This spotlights drawbacks of chain of thought, especially the sharp tradeoff between possible performance gains and the amount of human labor necessary to generate examples with correct reasoning traces. Kaya Stechly, Karthik Valmeekam, Subbarao Kambhampati |
NeurIPS | 2 |
| 2023 | Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human Preferences
Lin Guan 0003, Karthik Valmeekam, Subbarao Kambhampati |
ICLR | 2 |
| 2023 | Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task PlanningabstractThere is a growing interest in applying pre-trained large language models (LLMs) to planning problems. However, methods that use LLMs directly as planners are currently impractical due to several factors, including limited correctness of plans, strong reliance on feedback from interactions with simulators or even the actual environment, and the inefficiency in utilizing human feedback. In this work, we introduce a novel alternative paradigm that constructs an explicit world (domain) model in planning domain definition language (PDDL) and then uses it to plan with sound domain-independent planners. To address the fact that LLMs may not generate a fully functional PDDL model initially, we employ LLMs as an interface between PDDL and sources of corrective feedback, such as PDDL validators and humans. For users who lack a background in PDDL, we show that LLMs can translate PDDL into natural language and effectively encode corrective feedback back to the underlying domain model. Our framework not only enjoys the correctness guarantee offered by the external planners but also reduces human involvement by allowing users to correct domain models at the beginning, rather than inspecting and correcting (through interactive prompting) every generated plan as in previous work. On two IPC domains and a Household domain that is more complicated than commonly used benchmarks such as ALFWorld, we demonstrate that GPT-4 can be leveraged to produce high-quality PDDL models for over 40 actions, and the corrected PDDL models are then used to successfully solve 48 challenging planning tasks. Resources, including the source code, are released at: https://guansuns.github.io/pages/llm-dm. Lin Guan 0003, Karthik Valmeekam, Sarath Sreedharan, Subbarao Kambhampati |
NeurIPS | 2 |
| 2023 | PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about ChangeabstractGenerating plans of action, and reasoning about change have long been considered a core competence of intelligent agents. It is thus no surprise that evaluating the planning and reasoning capabilities of large language models (LLMs) has become a hot topic of research. Most claims about LLM planning capabilities are however based on common sense tasks–where it becomes hard to tell whether LLMs are planning or merely retrieving from their vast world knowledge. There is a strong need for systematic and extensible planning benchmarks with sufficient diversity to evaluate whether LLMs have innate planning capabilities. Motivated by this, we propose PlanBench, an extensible benchmark suite based on the kinds of domains used in the automated planning community, especially in the International Planning Competition, to test the capabilities of LLMs in planning or reasoning about actions and change. PlanBench provides sufficient diversity in both the task domains and the specific planning capabilities. Our studies also show that on many critical capabilities–including plan generation–LLM performance falls quite short, even with the SOTA models. PlanBench can thus function as a useful marker of progress of LLMs in planning and reasoning. Karthik Valmeekam, Matthew Marquez, Alberto Olmo Hernandez, Sarath Sreedharan, Subbarao Kambhampati |
NeurIPS | 1 |
| 2023 | On the Planning Abilities of Large Language Models - A Critical InvestigationabstractIntrigued by the claims of emergent reasoning capabilities in LLMs trained on general web corpora, in this paper, we set out to investigate their planning capabilities. We aim to evaluate (1) the effectiveness of LLMs in generating plans autonomously in commonsense planning tasks and (2) the potential of LLMs as a source of heuristic guidance for other agents (AI planners) in their planning tasks. We conduct a systematic study by generating a suite of instances on domains similar to the ones employed in the International Planning Competition and evaluate LLMs in two distinct modes: autonomous and heuristic. Our findings reveal that LLMs’ ability to generate executable plans autonomously is rather limited, with the best model (GPT-4) having an average success rate of ~12% across the domains. However, the results in the heuristic mode show more promise. In the heuristic mode, we demonstrate that LLM-generated plans can improve the search process for underlying sound planners and additionally show that external verifiers can help provide feedback on the generated plans and back-prompt the LLM for better plan generation. Karthik Valmeekam, Matthew Marquez, Sarath Sreedharan, Subbarao Kambhampati |
NeurIPS | 1 |
| 2021 | RADAR-X: An Interactive Interface Pairing Contrastive Explanations with Revised Plan SuggestionsabstractAutomated Planning techniques can be leveraged to build effective decision support systems that assist the human-in-the-loop. Such systems must provide intuitive explanations when the suggestions made by these systems seem inexplicable to the human. In this regard, we consider scenarios where the user questions the system's suggestion by providing alternatives (referred to as foils). In response, we empower existing decision support technologies to engage in an interactive explanatory dialogue with the user and provide contrastive explanations based on user-specified foils to reach a consensus on proposed decisions. To provide contrastive explanations, we adapt existing techniques in Explainable AI Planning (XAIP). Furthermore, we use this dialog to elicit the user's latent preferences and propose three modes of interaction that use these preferences to provide revised plan suggestions. Finally, we showcase a decision support system that provides all these capabilities. Karthik Valmeekam, Sarath Sreedharan, Sailik Sengupta, Subbarao Kambhampati |
AAAI | 1 |