VLDB 2026 Research / reviewers in the wild / expert
Pouya Pezeshkpour
dblp:159/1696
· DBLP profile ↗
8ranked-venue papers
4as first author
7since 2021 · last 2026
0009-0002-7055-0035ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 59% Trustworthy machine learning · 20% Knowledge representation and reasoning · 13% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › interpretability
faithful reasoning |
1.0 | 1 | 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation
mathematical reasoning |
1.0 | 1 | 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models |
1.0 | 1 | 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › agentic language model › tool-augmented language models
tool-augmented reasoning |
1.0 | 1 | 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language Models · ACL (1) 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph reasoning
knowledge base completion |
0.3 | 1 | 2018 | Embedding Multimodal Relational Data for Knowledge Base Completion · EMNLP 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
knowledge graph completion |
0.3 | 1 | 2018 | Embedding Multimodal Relational Data for Knowledge Base Completion · EMNLP 2018 |
Machine learning › Graph learning
link prediction |
0.3 | 1 | 2018 | Embedding Multimodal Relational Data for Knowledge Base Completion · EMNLP 2018 |
Machine learning › Generative modeling
multimodal generation |
0.1 | 1 | 2018 | Embedding Multimodal Relational Data for Knowledge Base Completion · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
preference optimization · 1.0benchmark construction · 1.0relational embedding models · 0.3neural encoder · 0.3neural decoder · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Proof to Program: Characterizing Tool-Induced Reasoning Hallucinations in Large Language ModelsabstractTool-augmented Language Models (TaLMs) can invoke external tools to solve problems beyond their parametric capacity.However, it remains unclear whether these tool-enabled gains reflect trustworthy reasoning.Focusing on the Code Interpreter tool, we show that even when tools are selected and executed correctly, TaLMs treat tool outputs as substitutes for reasoning, producing solutions that appear correct but lack coherent justification.We term this failure mode Tool-Induced Myopia (TIM), and study it using PYMATH, a benchmark of 1,679 competition-level mathematical problems for which Python code is helpful but not sufficient.We further develop a multi-dimensional evaluation suite to quantify reasoning degradation in TaLMs relative to their non-tool counterparts.Our findings reveal that while TaLMs achieve up to a 19.3 percentage point gain in final-answer accuracy, their reasoning behavior consistently deteriorates (e.g., non-tool language models win up to 41.5% more often in pairwise comparisons of reasoning processes).This degradation intensifies with tool use; the more frequently a model invokes tools, the less coherent its reasoning becomes.Moreover, tool use shifts errors from arithmetic mistakes toward global reasoning failures (logic, assumption, creativity).Finally, we propose a preference-optimization-based framework that realigns TaLMs to use tool outputs as assistive evidence, improving both final-answer accuracy and the reasoning depth under tool use. 1 Farima Fatahi Bayat, Pouya Pezeshkpour, Estevam Hruschka |
ACL (1) | 2 |
| 2025 | LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMsabstractArash Gholami Davoodi, Seyed Pouyan Mousavi Davoudi, Pouya Pezeshkpour. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Arash Gholami Davoodi, Seyed Pouyan Mousavi Davoudi, Pouya Pezeshkpour |
NAACL (Long Papers) | 3 |
| 2025 | Multi-Conditional Ranking with Large Language ModelsabstractPouya Pezeshkpour, Estevam Hruschka. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Pouya Pezeshkpour, Estevam Hruschka |
NAACL (Long Papers) | 1 |
| 2023 | Measuring and Modifying Factual Knowledge in Large Language ModelsabstractLarge Language Models (LLMs) store an extensive amount of factual knowledge obtained from vast collections of text. To effectively utilize these models for downstream tasks, it is crucial to have reliable methods for measuring their knowledge. However, existing approaches for knowledge measurement have certain limitations, and despite recent efforts, they fail to provide accurate measurements and the necessary insights for modifying the knowledge within LLMs. In this work, we employ information theory-based measurements to provide a framework estimating the factual knowledge contained within large language models. More specifically, we measure knowledge by analyzing the LLM's prediction probability distribution before and after instilling the target knowledge, employing metrics such as entropy and KL-divergence. Introducing our metrics, we first assess their accuracy in comparison to previous ranking-based methods, surpassing them by around 30% in a synthetic experiment. Then, we explore two prominent methods of knowledge instillation, discovering that LLMs exhibit limitations in capturing new knowledge under specific circumstances for one of these methods. Lastly, we demonstrate the applicability of our methods in extracting unlearned and mislearned facts in LLMs through their application to in-context learning. Pouya Pezeshkpour |
ICMLA | 1 |
| 2023 | The Extremal GDoF Gain of Optimal Versus Binary Power Control in K User Interference Networks is Θ (√K)abstractUsing ideas from Generalized Degrees of Freedom (GDoF) analyses and extremal network theory, this work studies the extremal gain of optimal power control over binary (on/off) power control, especially in large interference networks, in search of new theoretical insights. Whereas numerical studies have already established that in most practical settings binary power control is close to optimal, the extremal analysis shows not only that there exist settings where the gain from optimal power control can be quite significant, but also bounds the extremal values of such gains from a GDoF perspective. As its main contribution, this work explicitly characterizes the extremal GDoF gain of optimal over binary power control as$\Theta (\sqrt {K})$for all$K$. In particular, the extremal gain is bounded between$\lfloor \sqrt {K}\rfloor $and$2.5\sqrt {K}$for every$K$. For$K=2,3,4,5,6$users, the precise extremal gain is found to be 1, 3/2, 2, 9/4 and 41/16, respectively. Networks shown to achieve the extremal gain may be interpreted as multi-tier heterogeneous networks. It is worthwhile to note that because of their focus on asymptotic analysis, the sharp characterizations of extremal gains are valuable primarily from a theoretical perspective, and not as contradictions to the conventional wisdom that binary power control is generally close to optimal in practical, non-asymptotic settings. Yao-Chia Chan, Pouya Pezeshkpour, Chunhua Geng, Syed Ali Jafar |
IEEE Trans. Wirel. Commun. | 2 |
| 2021 | An Empirical Comparison of Instance Attribution Methods for NLPabstractPouya Pezeshkpour, Sarthak Jain, Byron Wallace, Sameer Singh. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Pouya Pezeshkpour, Byron C. Wallace, Sameer Singh 0001 |
NAACL-HLT | 1 |
| 2021 | ParsiNLU: A Suite of Language Understanding Challenges for PersianabstractAbstract Despite the progress made in recent years in addressing natural language understanding (NLU) challenges, the majority of this progress remains to be concentrated on resource-rich languages like English. This work focuses on Persian language, one of the widely spoken languages in the world, and yet there are few NLU datasets available for this language. The availability of high-quality evaluation datasets is a necessity for reliable assessment of the progress on different NLU tasks and domains. We introduce ParsiNLU, the first benchmark in Persian language that includes a range of language understanding tasks—reading comprehension, textual entailment, and so on. These datasets are collected in a multitude of ways, often involving manual annotations by native speakers. This results in over 14.5k new instances across 6 distinct NLU tasks. Additionally, we present the first results on state-of-the-art monolingual and multilingual pre-trained language models on this benchmark and compare them with human performance, which provides valuable insights into our ability to tackle natural language understanding challenges in Persian. We hope ParsiNLU fosters further research and advances in Persian language understanding.1 Daniel Khashabi, Arman Cohan, Siamak Shakeri, Pedram Hosseini, Pouya Pezeshkpour, Malihe Alikhani, Moin Aminnaseri, Marzieh Bitaab, Faeze Brahman, Sarik Ghazarian, Mozhdeh Gheini, Arman Kabiri, Rabeeh Karimi Mahabadi, Omid Memarrast, Ahmadreza Mosallanezhad, Erfan Noury, Shahab Raji, Mohammad Sadegh Rasooli, Sepideh Sadeghi, Erfan Sadeqi Azer, Niloofar Safi Samghabadi, Mahsa Shafaei, Saber Sheybani, Ali Tazarv, Yadollah Yaghoobzadeh |
Trans. Assoc. Comput. Linguistics | 5 |
| 2018 | Embedding Multimodal Relational Data for Knowledge Base CompletionabstractRepresenting entities and relations in an embedding space is a well-studied approach for machine learning on relational data.Existing approaches, however, primarily focus on simple link structure between a finite set of entities, ignoring the variety of data types that are often used in knowledge bases, such as text, images, and numerical values.In this paper, we propose multimodal knowledge base embeddings (MKBE) that use different neural encoders for this variety of observed data, and combine them with existing relational models to learn embeddings of the entities and multimodal data.Further, using these learned embedings and different neural decoders, we introduce a novel multimodal imputation model to generate missing multimodal values, like text and images, from information in the knowledge base.We enrich existing relational datasets to create two novel benchmarks that contain additional information such as textual descriptions and images of the original entities.We demonstrate that our models utilize this additional information effectively to provide more accurate link prediction, achieving state-of-the-art results with a considerable gap of 5-7% over existing methods.Further, we evaluate the quality of our generated multimodal values via a user study.We have release the datasets and the opensource implementation of our models at https: //github.com/pouyapez/mkbe. Pouya Pezeshkpour, Sameer Singh 0001 |
EMNLP | 1 |