VLDB 2026 Research / reviewers in the wild / expert
Soumya Sanyal 0001
dblp:86/1950-1
· DBLP profile ↗
13ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mixing Inference-time Experts for Enhancing LLM ReasoningabstractLarge Language Models (LLMs) have demonstrated impressive reasoning abilities, but their generated rationales often suffer from issues such as reasoning inconsistency and factual errors, undermining their reliability.Prior work has explored improving rationale quality via multi-reward fine-tuning or reinforcement learning (RL), where models are optimized for diverse objectives.While effective, these approaches train the model in a fixed manner and do not have any inference-time adaptability, nor can they generalize reasoning requirements for new test-time inputs.Another approach is to train specialized reasoning experts using reward signals and use them to improve generation at inference time.Existing methods in this paradigm are limited to using only a single expert and cannot improve upon multiple reasoning aspects.To address this, we propose MIXIE, a novel inference-time expertmixing framework that dynamically determines mixing proportions for each expert, enabling contextualized and flexible fusion.We demonstrate the effectiveness of MIXIE on improving chain-of-thought reasoning in LLMs by merging commonsense and entailment reasoning experts finetuned on reward-filtered data.Our approach outperforms existing baselines on three question-answering datasets: StrategyQA, CommonsenseQA, and ARC, highlighting its potential to enhance LLM reasoning with efficient, adaptable expert integration. Soumya Sanyal 0001, Tianyi Xiao, Xiang Ren 0001 |
EMNLP | 1 |
| 2024 | PlaSma: Procedural Knowledge Models for Language-based Planning and Re-PlanningabstractProcedural planning, which entails decomposing a high-level goal into a sequence of temporally ordered steps, is an important yet intricate task for machines. It involves integrating common-sense knowledge to reason about complex and often contextualized situations, e.g. ``scheduling a doctor's appointment without a phone''. While current approaches show encouraging results using large language models (LLMs), they are hindered by drawbacks such as costly API calls and reproducibility issues. In this paper, we advocate planning using smaller language models. We present PlaSma, a novel two-pronged approach to endow small language models with procedural knowledge and (constrained) language-based planning capabilities. More concretely, we develop *symbolic procedural knowledge distillation* to enhance the commonsense knowledge in small language models and an *inference-time algorithm* to facilitate more structured and accurate reasoning. In addition, we introduce a new related task, *Replanning*, that requires a revision of a plan to cope with a constrained situation. In both the planning and replanning settings, we show that orders-of-magnitude smaller models (770M-11B parameters) can compete and often surpass their larger teacher models' capabilities. Finally, we showcase successful application of PlaSma in an embodied environment, VirtualHome. Faeze Brahman, Chandra Bhagavatula, Valentina Pyatkin, Jena D. Hwang, Xiang Li 0069, Hirona Jacqueline Arai, Soumya Sanyal 0001, Keisuke Sakaguchi, Xiang Ren 0001, Yejin Choi 0001 |
ICLR | 7 |
| 2023 | APOLLO: A Simple Approach for Adaptive Pretraining of Language Models for Logical ReasoningabstractSoumya Sanyal, Yichong Xu, Shuohang Wang, Ziyi Yang, Reid Pryzant, Wenhao Yu, Chenguang Zhu, Xiang Ren. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Soumya Sanyal 0001, Yichong Xu, Shuohang Wang, Ziyi Yang 0011, Reid Pryzant, Wenhao Yu 0002, Chenguang Zhu 0001, Xiang Ren 0001 |
ACL (1) | 1 |
| 2023 | Generate rather than Retrieve: Large Language Models are Strong Context Generators
Wenhao Yu 0002, Dan Iter, Shuohang Wang, Yichong Xu, Mingxuan Ju, Soumya Sanyal 0001, Chenguang Zhu 0001, Michael Zeng 0001, Meng Jiang 0001 |
ICLR | 6 |
| 2023 | Faith and Fate: Limits of Transformers on CompositionalityabstractTransformer large language models (LLMs) have sparked admiration for their exceptional performance on tasks that demand intricate multi-step reasoning. Yet, these models simultaneously show failures on surprisingly trivial problems.
This begs the question: Are these errors incidental, or do they signal more substantial limitations?
In an attempt to demystify transformer LLMs, we investigate the limits of these models across three representative compositional tasks---multi-digit multiplication, logic grid puzzles, and a classic dynamic programming problem. These tasks require breaking problems down into sub-steps and synthesizing these steps into a precise answer. We formulate compositional tasks as computation graphs to systematically quantify the level of complexity, and break down reasoning steps into intermediate sub-procedures.
Our empirical findings suggest that transformer LLMs solve compositional tasks by reducing multi-step compositional reasoning into linearized subgraph matching, without necessarily developing systematic problem-solving skills. To round off our empirical study, we provide theoretical arguments on abstract multi-step reasoning problems that highlight how autoregressive generations' performance can rapidly decay with increased task complexity. Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Li 0069, Bill Y. Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras 0001, Jena D. Hwang, Soumya Sanyal 0001, Xiang Ren 0001, Allyson Ettinger, Zaïd Harchaoui, Yejin Choi 0001 |
NeurIPS | 12 |
| 2022 | FaiRR: Faithful and Robust Deductive Reasoning over Natural LanguageabstractTransformers have been shown to be able to perform deductive reasoning on a logical rulebase containing rules and statements written in natural language.Recent works show that such models can also produce the reasoning steps (i.e., the proof graph) that emulate the model's logical reasoning process.Currently, these black-box models generate both the proof graph and intermediate inferences within the same model and thus may be unfaithful.In this work, we frame the deductive logical reasoning task by defining three modular components: rule selection, fact selection, and knowledge composition.The rule and fact selection steps select the candidate rule and facts to be used and then the knowledge composition combines them to generate new inferences.This ensures model faithfulness by assured causal relation from the proof step to the inference reasoning.To test our framework, we propose FAIRR (Faithful and Robust Reasoner) where the above three components are independently modeled by transformers.We observe that FAIRR is robust to novel language perturbations, and is faster at inference than previous works on existing reasoning datasets.Additionally, in contrast to black-box generative models, the errors made by FAIRR are more interpretable due to the modular approach.1 Soumya Sanyal 0001, Harman Singh, Xiang Ren 0001 |
ACL (1) | 1 |
| 2022 | RobustLR: A Diagnostic Benchmark for Evaluating Logical Robustness of Deductive ReasonersabstractTransformers have been shown to be able to perform deductive reasoning on inputs containing rules and statements written in English natural language.However, it is unclear if these models indeed follow rigorous logical reasoning to arrive at the prediction, or rely on spurious correlation patterns in making decision.A strong deductive reasoning model should consistently understand the semantics of different logical operators.To this end, we present ROBUSTLR, a deductive reasoning-based diagnostic benchmark that evaluates the robustness of language models to minimal logical edits in the inputs and different logical equivalence conditions.In our experiments with RoBERTa, T5, and GPT3, we show that the models trained on deductive reasoning datasets with various logical operations do not perform consistently on the RO-BUSTLR test set, thus showing that the models are not robust to our proposed logical perturbations.Further, we observe that the models find it especially hard to learn logical negation operator.Our results demonstrate the shortcomings of current language models in logical reasoning, and call for the development of better inductive biases to teach the logical semantics to language models.All the datasets and code base have been made publicly available.1 Soumya Sanyal 0001, Zeyi Liao, Xiang Ren 0001 |
EMNLP | 1 |
| 2021 | Discretized Integrated Gradients for Explaining Language ModelsabstractAs a prominent attribution-based explanation algorithm, Integrated Gradients (IG) is widely adopted due to its desirable explanation axioms and the ease of gradient computation.It measures feature importance by averaging the model's output gradient interpolated along a straight-line path in the input data space.However, such straight-line interpolated points are not representative of text data due to the inherent discreteness of the word embedding space.This questions the faithfulness of the gradients computed at the interpolated points and consequently, the quality of the generated explanations.Here we propose Discretized Integrated Gradients (DIG), which allows effective attribution along non-linear interpolation paths.We develop two interpolation strategies for the discrete word embedding space that generates interpolation points that lie close to actual words in the embedding space, yielding more faithful gradient computation.We demonstrate the effectiveness of DIG over IG through experimental and human evaluations on multiple sentiment classification datasets.We provide the source code of DIG to encourage reproducible research 1 . Soumya Sanyal 0001, Xiang Ren 0001 |
EMNLP (1) | 1 |
| 2021 | SalKG: Learning From Knowledge Graph Explanations for Commonsense ReasoningabstractAugmenting pre-trained language models with knowledge graphs (KGs) has achieved success on various commonsense reasoning tasks. However, for a given task instance, the KG, or certain parts of the KG, may not be useful. Although KG-augmented models often use attention to focus on specific KG components, the KG is still always used, and the attention mechanism is never explicitly taught which KG components should be used. Meanwhile, saliency methods can measure how much a KG feature (e.g., graph, node, path) influences the model to make the correct prediction, thus explaining which KG features are useful. This paper explores how saliency explanations can be used to improve KG-augmented models' performance. First, we propose to create coarse (Is the KG useful?) and fine (Which nodes/paths in the KG are useful?) saliency explanations. Second, to motivate saliency-based supervision, we analyze oracle KG-augmented models which directly use saliency explanations as extra inputs for guiding their attention. Third, we propose SalKG, a framework for KG-augmented models to learn from coarse and/or fine saliency explanations. Given saliency explanations created from a task's training set, SalKG jointly trains the model to predict the explanations, then solve the task by attending to KG features highlighted by the predicted explanations. On three popular commonsense QA benchmarks (CSQA, OBQA, CODAH) and a range of KG-augmented models, we show that SalKG can yield considerable performance gains --- up to 2.76% absolute improvement on CSQA. Aaron Chan, Boyuan Long, Soumya Sanyal 0001, Tanishq Gupta, Xiang Ren 0001 |
NeurIPS | 4 |
| 2020 | ASAP: Adaptive Structure Aware Pooling for Learning Hierarchical Graph RepresentationsabstractGraph Neural Networks (GNN) have been shown to work effectively for modeling graph structured data to solve tasks such as node classification, link prediction and graph classification. There has been some recent progress in defining the notion of pooling in graphs whereby the model tries to generate a graph level representation by downsampling and summarizing the information present in the nodes. Existing pooling methods either fail to effectively capture the graph substructure or do not easily scale to large graphs. In this work, we propose ASAP (Adaptive Structure Aware Pooling), a sparse and differentiable pooling method that addresses the limitations of previous graph pooling architectures. ASAP utilizes a novel self-attention network along with a modified GNN formulation to capture the importance of each node in a given graph. It also learns a sparse soft cluster assignment for nodes at each layer to effectively pool the subgraphs to form the pooled graph. Through extensive experiments on multiple datasets and theoretical analysis, we motivate our choice of the components used in ASAP. Our experimental results show that combining existing GNN architectures with ASAP leads to state-of-the-art results on multiple graph classification benchmarks. ASAP has an average improvement of 4%, compared to current sparse hierarchical state-of-the-art method. We make the source code of ASAP available to encourage reproducible research 1. Ekagra Ranjan, Soumya Sanyal 0001, Partha P. Talukdar |
AAAI | 2 |
| 2020 | InteractE: Improving Convolution-Based Knowledge Graph Embeddings by Increasing Feature InteractionsabstractMost existing knowledge graphs suffer from incompleteness, which can be alleviated by inferring missing links based on known facts. One popular way to accomplish this is to generate low-dimensional embeddings of entities and relations, and use these to make inferences. ConvE, a recently proposed approach, applies convolutional filters on 2D reshapings of entity and relation embeddings in order to capture rich interactions between their components. However, the number of interactions that ConvE can capture is limited. In this paper, we analyze how increasing the number of these interactions affects link prediction performance, and utilize our observations to propose InteractE. InteractE is based on three key ideas – feature permutation, a novel feature reshaping, and circular convolution. Through extensive experiments, we find that InteractE outperforms state-of-the-art convolutional link prediction baselines on FB15k-237. Further, InteractE achieves an MRR score that is 9%, 7.5%, and 23% better than ConvE on the FB15k-237, WN18RR and YAGO3-10 datasets respectively. The results validate our central hypothesis – that increasing feature interaction is beneficial to link prediction performance. We make the source code of InteractE available to encourage reproducible research. Shikhar Vashishth, Soumya Sanyal 0001, Vikram Nitin, Nilesh Agrawal, Partha P. Talukdar |
AAAI | 2 |
| 2020 | A Re-evaluation of Knowledge Graph Completion MethodsabstractKnowledge Graph Completion (KGC) aims at automatically predicting missing links for large-scale knowledge graphs.A vast number of state-of-the-art KGC techniques have got published at top conferences in several research fields, including data mining, machine learning, and natural language processing.However, we notice that several recent papers report very high performance, which largely outperforms previous state-of-the-art methods.In this paper, we find that this can be attributed to the inappropriate evaluation protocol used by them and propose a simple evaluation protocol to address this problem.The proposed protocol is robust to handle bias in the model, which can substantially affect the final results.We conduct extensive experiments and report performance of several existing methods using our protocol.The reproducible code has been made publicly available. Zhiqing Sun, Shikhar Vashishth, Soumya Sanyal 0001, Partha P. Talukdar, Yiming Yang 0002 |
ACL | 3 |
| 2020 | Composition-based Multi-Relational Graph Convolutional Networks
Shikhar Vashishth, Soumya Sanyal 0001, Vikram Nitin, Partha P. Talukdar |
ICLR | 2 |