VLDB 2026 Research / reviewers in the wild / expert
Shrimai Prabhumoye
dblp:203/8169
· DBLP profile ↗
16ranked-venue papers
6as first author
10since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MIND: Math Informed syNthetic Dialogues for Pretraining LLMsabstractThe utility of synthetic data to enhance pretraining data quality and hence to improve downstream task accuracy has been widely explored in recent large language models (LLMs). Yet, these approaches fall inadequate in complex, multi-hop and mathematical reasoning tasks as the synthetic data typically fails to add complementary knowledge to the existing raw corpus. In this work, we propose a novel large-scale and diverse Math Informed syNthetic Dialogue (MIND) generation method that improves the mathematical reasoning ability of LLMs. Specifically, using MIND, we generate synthetic conversations based on OpenWebMath (OWM), resulting in a new math corpus, MIND-OWM. Our experiments with different conversational settings reveal that incorporating knowledge gaps between dialog participants is essential for generating high-quality math data. We further identify an effective way to format and integrate synthetic and raw data during pretraining to maximize the gain in mathematical reasoning, emphasizing the need to restructure raw data rather than use it as-is. Compared to pretraining just on raw data, a model pretrained on MIND-OWM shows significant boost in mathematical reasoning (GSM8K: +13.42%, MATH: +2.30%), including superior performance in specialized knowledge (MMLU: +4.55%, MMLU-STEM: +4.28%) and general
purpose reasoning tasks (GENERAL REASONING: +2.51%). Syeda Nahida Akter, Shrimai Prabhumoye, John Kamalu, Sanjeev Satheesh, Eric Nyberg, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro |
ICLR | 2 |
| 2025 | Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM ReasoningabstractData diversity is crucial for training a strong language model. Yet metrics of diversity often diverge from this goal, measuring variations in heuristic features—like n-grams or embeddings—that are detached from how the model actually performs on a target task. This motivates us to ask: *Can we redefine data diversity—beyond measuring variations in heuristic features—in a way that better predicts model generalization?* Through large-scale empirical analyses spanning over 300 training runs, carefully controlled for data scale and quality, we show that data diversity can be a strong predictor of generalization in LLM reasoning—as measured by average model performance on unseen out-of-distribution benchmarks. We introduce **G-Vendi**, a metric that quantifies diversity via the entropy of model-induced loss gradients. G-Vendi scales to million-sample datasets and yet consistently outperforms heuristic alternatives, achieving strong correlation ($\text{Spearman's } \rho \approx 0.9$) with out-of-distribution (OOD) performance across both natural language inference (NLI) and math reasoning tasks. Building on this insight, we present **Prismatic Synthesis**, a framework for generating diverse synthetic data by targeting underrepresented regions in gradient space. Experimental results show that Prismatic Synthesis consistently improves model performance as we scale synthetic data—not just on in-distribution test but across unseen, out-of-distribution benchmarks—significantly outperforming state-of-the-art models in both domains. For example, PrismMath-7B, our model distilled from a 32B LLM without human verification, outperforms R1-Distill-Qwen-7B—trained on proprietary data generated by 671B R1—on 6 out of 7 challenging math benchmarks. Jaehun Jung, Seungju Han 0002, Ximing Lu, Skyler Hallinan, David Acuna, Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Yejin Choi 0001 |
NeurIPS | 6 |
| 2024 | Data, Data Everywhere: A Guide for Pretraining Dataset ConstructionabstractJupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Bo Liu, Aastha Jhunjhunwala, Zhilin Wang, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Aastha Jhunjhunwala, Zhilin Wang, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro |
EMNLP | 2 |
| 2024 | LLM-Evolve: Evaluation for LLM's Evolving Capability on BenchmarksabstractThe advancement of large language models (LLMs) has extended their use to dynamic and interactive real-world applications, where models engage continuously with their environment and potentially enhance their performance over time.Most existing LLM benchmarks evaluate LLMs on i.i.d.tasks, overlooking their ability to learn iteratively from past experiences.Our paper bridges this evaluation gap by proposing a novel framework, LLM-Evolve, which extends established benchmarks to sequential problem-solving settings.LLM-Evolve evaluates LLMs over multiple rounds, providing feedback after each round to build a demonstration memory that the models can query in future tasks.We applied LLM-Evolve to the MMLU, GSM8K, and AgentBench benchmarks, testing 8 state-of-the-art open-source and closed-source models.Results show that LLMs can achieve performance improvements of up to 17% by learning from past interactions, with the quality of retrieval algorithms and feedback significantly influencing this capability.These insights advocate for more understanding and benchmarks for LLMs' performance in evolving interactive scenarios. Jiaxuan You, Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro |
EMNLP | 3 |
| 2023 | Adding Instructions during Pretraining: Effective way of Controlling Toxicity in Language ModelsabstractPretrained large language models have become indispensable for solving various natural language processing (NLP) tasks.However, safely deploying them in real world applications is challenging because they generate toxic content.To address this challenge, we propose two novel pretraining data augmentation strategies that significantly reduce model toxicity without compromising its utility.Our two strategies are: (1) MEDA: adds raw toxicity score as meta-data to the pretraining samples, and (2) INST: adds instructions to those samples indicating their toxicity.Our results indicate that our best performing strategy (INST) substantially reduces the toxicity probability up to 61% while preserving the accuracy on five benchmark NLP tasks as well as improving AUC scores on four bias detection tasks by 1.3%.We also demonstrate the generalizability of our techniques by scaling the number of training samples and the number of model parameters. Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro |
EACL | 1 |
| 2023 | Self-Refine: Iterative Refinement with Self-FeedbackabstractLike humans, large language models (LLMs) do not always generate the best output on their first try. Motivated by how humans refine their written text, we introduce Self-Refine, an approach for improving initial outputs from LLMs through iterative feedback and refinement. The main idea is to generate an initial output using an LLMs; then, the same LLMs provides *feedback* for its output and uses it to *refine* itself, iteratively. Self-Refine does not require any supervised training data, additional training, or reinforcement learning, and instead uses a single LLM as the generator, refiner and the feedback provider. We evaluate Self-Refine across 7 diverse tasks, ranging from dialog response generation to mathematical reasoning, using state-of-the-art (GPT-3.5, ChatGPT, and GPT-4) LLMs. Across all evaluated tasks, outputs generated with Self-Refine are preferred by humans and automatic metrics over those generated with the same LLM using conventional one-step generation, improving by $\sim$20\% absolute on average in task performance. Our work demonstrates that even state-of-the-art LLMs like GPT-4 can be further improved at test-time using our simple, standalone approach. Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon 0002, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang 0002, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, Peter Clark |
NeurIPS | 9 |
| 2023 | SPRING: Studying Papers and Reasoning to play GamesabstractOpen-world survival games pose significant challenges for AI algorithms due to their multi-tasking, deep exploration, and goal prioritization requirements. Despite reinforcement learning (RL) being popular for solving games, its high sample complexity limits its effectiveness in complex open-world games like Crafter or Minecraft. We propose a novel approach, SPRING, to read Crafter's original academic paper and use the knowledge learned to reason and play the game through a large language model (LLM).
Prompted with the LaTeX source as game context and a description of the agent's current observation, our SPRING framework employs a directed acyclic graph (DAG) with game-related questions as nodes and dependencies as edges. We identify the optimal action to take in the environment by traversing the DAG and calculating LLM responses for each node in topological order, with the LLM's answer to final node directly translating to environment actions.
In our experiments, we study the quality of in-context "reasoning" induced by different forms of prompts under the setting of the Crafter environment. Our experiments suggest that LLMs, when prompted with consistent chain-of-thought, have great potential in completing sophisticated high-level trajectories. Quantitatively, SPRING with GPT-4 outperforms all state-of-the-art RL baselines, trained for 1M steps, without any training.
Finally, we show the potential of Crafter as a test bed for LLMs. Code at github.com/holmeswww/SPRING Yue Wu 0001, So Yeon Min, Shrimai Prabhumoye, Yonatan Bisk, Ruslan Salakhutdinov, Amos Azaria, Tom M. Mitchell, Yuanzhi Li |
NeurIPS | 3 |
| 2022 | Evaluating Parameter Efficient Learning for GenerationabstractPeng Xu, Mostofa Patwary, Shrimai Prabhumoye, Virginia Adams, Ryan Prenger, Wei Ping, Nayeon Lee, Mohammad Shoeybi, Bryan Catanzaro. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Peng Xu 0008, Mostofa Patwary, Shrimai Prabhumoye, Virginia Adams, Ryan Prenger, Wei Ping, Nayeon Lee, Mohammad Shoeybi, Bryan Catanzaro |
EMNLP | 3 |
| 2021 | Case Study: Deontological Ethics in NLPabstractShrimai Prabhumoye, Brendon Boldt, Ruslan Salakhutdinov, Alan W Black. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shrimai Prabhumoye, Brendon Boldt, Ruslan Salakhutdinov, Alan W. Black |
NAACL-HLT | 1 |
| 2021 | Focused Attention Improves Document-Grounded GenerationabstractShrimai Prabhumoye, Kazuma Hashimoto, Yingbo Zhou, Alan W Black, Ruslan Salakhutdinov. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shrimai Prabhumoye, Kazuma Hashimoto, Yingbo Zhou 0002, Alan W. Black, Ruslan Salakhutdinov |
NAACL-HLT | 1 |
| 2020 | Generating Interactive Worlds with TextabstractProcedurally generating cohesive and interesting game environments is challenging and time-consuming. In order for the relationships between the game elements to be natural, common-sense has to be encoded into arrangement of the elements. In this work, we investigate a machine learning approach for world creation using content from the multi-player text adventure game environment LIGHT (Urbanek et al. 2019). We introduce neural network based models to compositionally arrange locations, characters, and objects into a coherent whole. In addition to creating worlds based on existing elements, our models can generate new game content. Humans can also leverage our models to interactively aid in worldbuilding. We show that the game environments created with our approach are cohesive, diverse, and preferred by human evaluators compared to other machine learning based world construction algorithms. Angela Fan, Jack Urbanek, Pratik Ringshia, Emily Dinan, Emma Qian, Siddharth Karamcheti, Shrimai Prabhumoye, Douwe Kiela, Tim Rocktäschel, Arthur Szlam, Jason Weston |
AAAI | 7 |
| 2020 | Politeness Transfer: A Tag and Generate ApproachabstractAman Madaan, Amrith Setlur, Tanmay Parekh, Barnabas Poczos, Graham Neubig, Yiming Yang, Ruslan Salakhutdinov, Alan W Black, Shrimai Prabhumoye. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Aman Madaan, Amrith Setlur, Tanmay Parekh, Barnabás Póczos, Graham Neubig, Yiming Yang 0002, Ruslan Salakhutdinov, Alan W. Black, Shrimai Prabhumoye |
ACL | 9 |
| 2020 | Topological Sort for Sentence OrderingabstractSentence ordering is the task of arranging the sentences of a given text in the correct order.Recent work using deep neural networks for this task has framed it as a sequence prediction problem.In this paper, we propose a new framing of this task as a constraint solving problem and introduce a new technique to solve it.Additionally, we propose a human evaluation for this task.The results on both automatic and human metrics across four different datasets show that this new technique is better at capturing coherence in documents. Shrimai Prabhumoye, Ruslan Salakhutdinov, Alan W. Black |
ACL | 1 |
| 2020 | Exploring Controllable Text Generation TechniquesabstractNeural controllable text generation is an important area gaining attention due to its plethora of applications.Although there is a large body of prior work in controllable text generation, there is no unifying theme.In this work, we provide a new schema of the pipeline of the generation process by classifying it into five modules.The control of attributes in the generation process requires modification of these modules.We present an overview of different techniques used to perform the modulation of these modules.We also provide an analysis on the advantages and disadvantages of these techniques.We further pave ways to develop new architectures based on the combination of the modules described in this paper. Shrimai Prabhumoye, Alan W. Black, Ruslan Salakhutdinov |
COLING | 1 |
| 2018 | Style Transfer Through Back-TranslationabstractStyle transfer is the task of rephrasing the text to contain specific stylistic properties without changing the intent or affect within the context.This paper introduces a new method for automatic style transfer.We first learn a latent representation of the input sentence which is grounded in a language translation model in order to better preserve the meaning of the sentence while reducing stylistic properties.Then adversarial generation techniques are used to make the output match the desired style.We evaluate this technique on three different style transformations: sentiment, gender and political slant.Compared to two state-of-the-art style transfer modeling techniques we show improvements both in automatic evaluation of style transfer and in manual evaluation of meaning preservation and fluency. Shrimai Prabhumoye, Yulia Tsvetkov, Ruslan Salakhutdinov, Alan W. Black |
ACL (1) | 1 |
| 2018 | A Dataset for Document Grounded ConversationsabstractThis paper introduces a document grounded dataset for conversations.We define "Document Grounded Conversations" as conversations that are about the contents of a specified document.In this dataset the specified documents were Wikipedia articles about popular movies.The dataset contains 4112 conversations with an average of 21.43 turns per conversation.This positions this dataset to not only provide a relevant chat history while generating responses but also provide a source of information that the models could use.We describe two neural architectures that provide benchmark performance on the task of generating the next response.We also evaluate our models for engagement and fluency, and find that the information from the document helps in generating more engaging and fluent responses. Kangyan Zhou, Shrimai Prabhumoye, Alan W. Black |
EMNLP | 2 |