VLDB 2026 Research / reviewers in the wild / expert
Abulhair Saparov
dblp:117/6287
· DBLP profile ↗
15ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-0450-9125ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex ProofsabstractLarge language models (LLMs) can solve arithmetic word problems with high accuracy, but little is known about how well they generalize to more complex problems. This is difficult to study, as (i) much of the available evaluation data has already been seen by the most capable models during training, and (ii) existing benchmarks do not capture how problem proofs may be arbitrarily complex in various ways. In this paper, we present a data-generation framework for evaluating LLMs on problems with arbitrarily complex arithmetic proofs, called MathGAP. MathGAP generates problem statements and chain-of-thought reasoning traces according to specifications about their arithmetic proof structure, enabling systematic studies on easy-to-hard generalization with respect to complexity of proof trees. Using MathGAP, we find that LLMs show a significant decrease in performance as proofs get deeper and wider. This effect is more pronounced in complex, nonlinear proof structures, which are challenging even for the most capable models. The models are also sensitive to simple changes in sentence ordering. However, they remain capable of solving some complex problems, suggesting that reasoning generalization is noisy. Andreas Opedal, Haruki Shirakami, Bernhard Schölkopf, Abulhair Saparov, Mrinmaya Sachan |
ICLR | 4 |
| 2025 | Transformers Struggle to Learn to SearchabstractSearch is an ability foundational in many important tasks, and recent studies have shown that large language models (LLMs) struggle to perform search robustly. It is unknown whether this inability is due to a lack of data, insufficient model parameters, or fundamental limitations of the transformer architecture. In this work, we use the foundational graph connectivity problem as a testbed to generate effectively limitless high-coverage data to train small transformers and test whether they can learn to perform search. We find that, when given the right training distribution, the transformer is able to learn to search.
We analyze the algorithm that the transformer has learned through a novel mechanistic interpretability technique that enables us to extract the computation graph from the trained model. We find that for each vertex in the input graph, transformers compute the set of vertices reachable from that vertex. Each layer then progressively expands these sets, allowing the model to search over a number of vertices exponential in the number of layers.
However, we find that as the input graph size increases, the transformer has greater difficulty in learning the task. This difficulty is not resolved even as the number of parameters is increased, suggesting that increasing model scale will not lead to robust search abilities. We also find that performing search in-context (i.e., chain-of-thought) does not resolve this inability to learn to search on larger graphs. Abulhair Saparov, Srushti Pawar, Shreyas Pimpalgaonkar, Nitish Joshi, Richard Yuanzhe Pang, Vishakh Padmakumar, Mehran Kazemi, Najoung Kim, He He 0001 |
ICLR | 1 |
| 2025 | Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain PerspectiveabstractReinforcement learning (RL) has shown promise in enhancing large language model (LLM) reasoning, yet progress towards broader capabilities is limited by the availability of high-quality, multi-domain datasets. This work introduces \ours, a 92K RL-for-reasoning dataset designed to address this gap, covering six reasoning domains: Math, Code, Science, Logic, Simulation, and Tabular, each with corresponding verifiers. We build \ours via a careful data-curation pipeline, including sourcing, deduplication, reward design, and domain-specific and difficulty-based filtering, to facilitate the systematic investigation of cross-domain RL generalization. Our study using \ours suggests the efficacy of a simple mixed-domain RL training approach and reveals several key aspects affecting cross-domain transferability. We further train two models {\ours}-7B and {\ours}-32B purely with RL on our curated data and observe largely improved performance over leading open RL reasoning model baselines, with gains of 7.3\% and 7.8\% respectively on an extensive 17-task, six-domain evaluation suite. We are releasing our dataset, code, and evaluation suite to the community, aiming to support further research and development of more general RL-enhanced reasoning models. Jorge (Zhoujun) Cheng, Shibo Hao, Tianyang Liu 0003, Yuexin Bian, Nilabjo Dey, Yonghao Zhuang 0001, Yuheng Zha, Yi Gu 0002, Kun Zhou 0002, Yuan Li 0032, Richard Fan, Jianshu She, Chengqian Gao, Abulhair Saparov, Taylor W. Killian, Haonan Li 0002, Mikhail Yurochkin, Eric P. Xing, Zhengzhong Liu 0001, Zhiting Hu |
NeurIPS | 18 |
| 2025 | Are Language Models Efficient Reasoners? A Perspective from Logic ProgrammingabstractModern language models (LMs) exhibit strong deductive reasoning capabilities, yet standard evaluations emphasize correctness while overlooking a key aspect of reasoning: *efficiency*. In real-world reasoning scenarios, much of the available information is irrelevant, and effective deductive inference requires identifying and ignoring such distractions. We propose a framework for assessing LM reasoning efficiency through the lens of logic programming, introducing a simple method to align proofs written in natural language---as generated by an LM---with shortest proofs found by executing the logic program. Efficiency is quantified by measuring how well a model avoids unnecessary inference. Empirically, we construct a dataset of math word problems injected with various number of irrelevant axioms that vary in semantic overlap with the goal theorem. We find that current LMs show marked accuracy declines under such conditions---even with minimal, domain-consistent distractions---and the proofs they generate frequently exhibit detours through irrelevant inferences. Andreas Opedal, Yanick Zengaffinen, Haruki Shirakami, Clemente Pasti, Mrinmaya Sachan, Abulhair Saparov, Ryan Cotterell, Bernhard Schölkopf |
NeurIPS | 6 |
| 2024 | Personas as a Way to Model Truthfulness in Language ModelsabstractLarge language models (LLMs) are trained on vast amounts of text from the internet, which contains both factual and misleading information about the world.While unintuitive from a classic view of language models, recent work has shown that the truth value of a statement can be elicited from the model's representations.This paper presents an explanation, persona hypothesis, for why LLMs appear to know the truth despite not being trained with truth labels.We hypothesize that the pretraining data is generated by groups of (un)truthful agents whose outputs share common features, and they form a (un)truthful persona.By training on this data, LMs can infer and represent the persona in its activation space.This allows the model to separate truth from falsehoods and controls the truthfulness of its generation.We show evidence for the persona hypothesis via two observations: (1) we can probe whether a model's answer will be truthful before it is generated; (2) finetuning a model on a set of true facts improves its truthfulness on unseen topics.Next, using arithmetics as a synthetic environment, we show that structures of the pretraining data are crucial for the model to infer the truthful persona.Overall, our findings suggest that models can exploit hierarchical structures in the data to learn abstract concepts like truthfulness. Nitish Joshi, Javier Rando, Abulhair Saparov, Najoung Kim, He He 0001 |
EMNLP | 3 |
| 2024 | LLMs Are Prone to Fallacies in Causal InferenceabstractRecent work shows that causal facts can be effectively extracted from LLMs through prompting, facilitating the creation of causal graphs for causal inference tasks.However, it is unclear if this success is limited to explicitly-mentioned causal facts in the pretraining data which the model can memorize.Thus, this work investigates: Can LLMs infer causal relations from other relational data in text?To disentangle the role of memorized causal facts vs inferred causal relations, we finetune LLMs on synthetic data containing temporal, spatial and counterfactual relations, and measure whether the LLM can then infer causal relations.We find that: (a) LLMs are susceptible to inferring causal relations from the order of two entity mentions in text (e.g.X mentioned before Y implies X causes Y); (b) if the order is randomized, LLMs still suffer from the post hoc fallacy, i.e.X occurs before Y (temporal relation) implies X causes Y.We also find that while LLMs can correctly deduce the absence of causal relations from temporal and spatial relations, they have difficulty inferring causal relations from counterfactuals, questioning their understanding of causality. Nitish Joshi, Abulhair Saparov, He He 0001 |
EMNLP | 2 |
| 2024 | Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?abstractThere is increasing interest in employing large language models (LLMs) as cognitive models. For such purposes, it is central to understand which properties of human cognition are well-modeled by LLMs, and which are not. In this work, we study the biases of LLMs in relation to those known in children when solving arithmetic word problems. Surveying the learning science literature, we posit that the problem-solving process can be split into three distinct steps: text comprehension, solution planning and solution execution. We construct tests for each one in order to understand whether current LLMs display the same cognitive biases as children in these steps. We generate a novel set of word problems for each of these tests, using a neuro-symbolic approach that enables fine-grained control over the problem features. We find evidence that LLMs, with and without instruction-tuning, exhibit human-like biases in both the text-comprehension and the solution-planning steps of the solving process, but not in the final step, in which the arithmetic expressions are executed to obtain the answer. Andreas Opedal, Alessandro Stolfo, Haruki Shirakami, Ying Jiao, Ryan Cotterell, Bernhard Schölkopf, Abulhair Saparov, Mrinmaya Sachan |
ICML | 7 |
| 2023 | Noisy Exemplars Make Large Language Models More Robust: A Domain-Agnostic Behavioral AnalysisabstractRecent advances in prompt engineering enable large language models (LLMs) to solve multihop logical reasoning problems with impressive accuracy.However, there is little existing work investigating the robustness of LLMs with few-shot prompting techniques.Therefore, we introduce a systematic approach to test the robustness of LLMs in multi-hop reasoning tasks via domain-agnostic perturbations.We include perturbations at multiple levels of abstractions (e.g.lexical perturbations such as typos, and semantic perturbations such as the inclusion of intermediate reasoning steps in the questions) to conduct behavioral analysis on the LLMs.Throughout our experiments, we find that models are more sensitive to certain perturbations such as replacing words with their synonyms.We also demonstrate that increasing the proportion of perturbed exemplars in the prompts improves the robustness of few-shot prompting methods. Hongyi Zheng, Abulhair Saparov |
EMNLP | 2 |
| 2023 | Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
Abulhair Saparov, He He 0001 |
ICLR | 1 |
| 2023 | Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesabstractGiven the intractably large size of the space of proofs, any model that is capable of general deductive reasoning must generalize to proofs of greater complexity. Recent studies have shown that large language models (LLMs) possess some abstract deductive reasoning ability given chain-of-thought prompts. However, they have primarily been tested on proofs using modus ponens or of a specific size, and from the same distribution as the in-context examples. To measure the general deductive reasoning ability of LLMs, we test on a broad set of deduction rules and measure their ability to generalize to more complex proofs from simpler demonstrations from multiple angles: depth-, width-, and compositional generalization. To facilitate systematic exploration, we construct a new synthetic and programmable reasoning dataset that enables control over deduction rules and proof complexity. Our experiments on four LLMs of various sizes and training objectives show that they are able to generalize to compositional proofs. However, they have difficulty generalizing to longer proofs, and they require explicit demonstrations to produce hypothetical subproofs, specifically in proof by cases and proof by contradiction. Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Mehran Kazemi, Najoung Kim, He He 0001 |
NeurIPS | 1 |
| 2022 | Towards General Natural Language Understanding with Probabilistic WorldbuildingabstractAbstract We introduce the Probabilistic Worldbuilding Model (PWM), a new fully symbolic Bayesian model of semantic parsing and reasoning, as a first step in a research program toward more domain- and task-general NLU and AI. Humans create internal mental models of their observations that greatly aid in their ability to understand and reason about a large variety of problems. In PWM, the meanings of sentences, acquired facts about the world, and intermediate steps in reasoning are all expressed in a human-readable formal language, with the design goal of interpretability. PWM is Bayesian, designed specifically to be able to generalize to new domains and new tasks. We derive and implement an inference algorithm that reads sentences by parsing and abducing updates to its latent world model that capture the semantics of those sentences, and evaluate it on two out-of-domain question-answering datasets: (1) ProofWriter and (2) a new dataset we call FictionalGeoQA, designed to be more representative of real language but still simple enough to focus on evaluating reasoning ability, while being robust against heuristics. Our method outperforms baselines on both, thereby demonstrating its value as a proof-of-concept. Abulhair Saparov, Tom M. Mitchell |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Jelly Bean World: A Testbed for Never-Ending Learning
Emmanouil A. Platanios, Abulhair Saparov, Tom M. Mitchell |
ICLR | 2 |
| 2017 | A Probabilistic Generative Grammar for Semantic ParsingabstractWe present a generative model of natural language sentences and demonstrate its application to semantic parsing.In the generative process, a logical form sampled from a prior, and conditioned on this logical form, a grammar probabilistically generates the output sentence.Grammar induction using MCMC is applied to learn the grammar given a set of labeled sentences with corresponding logical forms.We develop a semantic parser that finds the logical form with the highest posterior probability exactly.We obtain strong results on the GeoQuery dataset and achieve state-of-the-art F1 on Jobs. Abulhair Saparov, Vijay A. Saraswat, Tom M. Mitchell |
CoNLL | 1 |
| 2015 | Never-Ending LearningabstractWhereas people learn many different types of knowledge from diverse experiences over many years, most current machine learning systems acquire just a single function or data model from just a single data set. We propose a never-ending learning paradigm for machine learning, to better reflect the more ambitious and encompassing type of learning performed by humans. As a case study, we describe the Never-Ending Language Learner (NELL), which achieves some of the desired properties of a never-ending learner, and we discuss lessons learned. NELL has been learning to read the web 24 hours/day since January 2010, and so far has acquired a knowledge base with over 80 million confidence-weighted beliefs (e.g., servedWith(tea, biscuits)). NELL has also learned millions of features and parameters that enable it to read these beliefs from the web. Additionally, it has learned to reason over these beliefs to infer new beliefs, and is able to extend its ontology by synthesizing new relational predicates. NELL can be tracked online at http://rtw.ml.cmu.edu, and followed on Twitter at @CMUNELL. Tom M. Mitchell, William W. Cohen, Estevam Hruschka, Partha P. Talukdar, Justin Betteridge, Andrew Carlson, Bhavana Dalvi, Matt Gardner 0001, Bryan Kisiel, Jayant Krishnamurthy, Ni Lao, Kathryn Mazaitis, Thahir Mohamed, Ndapandula Nakashole, Emmanouil A. Platanios, Alan Ritter, Mehdi Samadi, Burr Settles, Richard C. Wang, Derry Wijaya, Abhinav Gupta 0001, Xinlei Chen, Abulhair Saparov, Malcolm Greaves, Joel Welling |
AAAI | 23 |
| 2012 | Schelling points on 3D surface meshesabstractThis paper investigates "Schelling points" on 3D meshes, feature points selected by people in a pure coordination game due to their salience. To collect data for this investigation, we designed an online experiment that asked people to select points on 3D surfaces that they expect will be selected by other people. We then analyzed properties of the selected points, finding that: 1) Schelling point sets are usually highly symmetric, and 2) local curvature properties (e.g., Gauss curvature) are most helpful for identifying obvious Schelling points (tips of protrusions), but 3) global properties (e.g., segment centeredness, proximity to a symmetry axis, etc.) are required to explain more subtle features. Based on these observations, we use regression analysis to combine multiple properties into an analytical model that predicts where Schelling points are likely to be on new meshes. We find that this model benefits from a variety of surface properties, particularly when training data comes from examples in the same object class. Xiaobai Chen, Abulhair Saparov, Bill Pang, Thomas A. Funkhouser |
ACM Trans. Graph. | 2 |