VLDB 2026 Research / reviewers in the wild / expert
Vishakh Padmakumar
dblp:285/5184
· DBLP profile ↗
14ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0002-3396-3589ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User PersonasabstractNishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng, Rachel Rudinger, Jordan Lee Boyd-Graber. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nishant Balepur, Vishakh Padmakumar, Fumeng Yang, Shi Feng 0005, Rachel Rudinger, Jordan L. Boyd-Graber |
ACL (1) | 2 |
| 2025 | Principled Content Selection to Generate Diverse and Personalized Multi-Document SummariesabstractWhile large language models (LLMs) are increasingly capable of handling longer contexts, recent work has demonstrated that they exhibit the "lost in the middle" phenomenon (Liu et al., 2024) of unevenly attending to different parts of the provided context.This hinders their ability to cover diverse source material in multidocument summarization, as noted in the DI-VERSESUMM benchmark (Huang et al., 2024).In this work, we contend that principled content selection is a simple way to increase source coverage on this task.As opposed to prompting an LLM to perform the summarization in a single step, we explicitly divide the task into three steps-(1) reducing document collections to atomic key points, (2) using determinantal point processes (DPP) to perform select key points that prioritize diverse content, and (3) rewriting to the final summary.By combining prompting steps, for extraction and rewriting, with principled techniques, for content selection, we consistently improve source coverage on the DIVERSESUMM benchmark across various LLMs.Finally, we also show that by incorporating relevance to a provided user intent into the DPP kernel, we can generate personalized summaries that cover relevant source information while retaining coverage. Vishakh Padmakumar, Zichao Wang 0001, David T. Arbour, Jennifer A. Healey |
ACL (1) | 1 |
| 2025 | Transformers Struggle to Learn to SearchabstractSearch is an ability foundational in many important tasks, and recent studies have shown that large language models (LLMs) struggle to perform search robustly. It is unknown whether this inability is due to a lack of data, insufficient model parameters, or fundamental limitations of the transformer architecture. In this work, we use the foundational graph connectivity problem as a testbed to generate effectively limitless high-coverage data to train small transformers and test whether they can learn to perform search. We find that, when given the right training distribution, the transformer is able to learn to search.
We analyze the algorithm that the transformer has learned through a novel mechanistic interpretability technique that enables us to extract the computation graph from the trained model. We find that for each vertex in the input graph, transformers compute the set of vertices reachable from that vertex. Each layer then progressively expands these sets, allowing the model to search over a number of vertices exponential in the number of layers.
However, we find that as the input graph size increases, the transformer has greater difficulty in learning the task. This difficulty is not resolved even as the number of parameters is increased, suggesting that increasing model scale will not lead to robust search abilities. We also find that performing search in-context (i.e., chain-of-thought) does not resolve this inability to learn to search on larger graphs. Abulhair Saparov, Srushti Pawar, Shreyas Pimpalgaonkar, Nitish Joshi, Richard Yuanzhe Pang, Vishakh Padmakumar, Mehran Kazemi, Najoung Kim, He He 0001 |
ICLR | 6 |
| 2024 | Creativity Support in the Age of Large Language Models: An Empirical Study Involving Professional WritersabstractThe development of large language models (LLMs) capable of following instructions and engaging in conversational interactions has led to increased interest in their use across various support tools. We investigate the effectiveness of contemporary LLMs in assisting professional writers via an empirical user study (n=30). The design of our collaborative writing interface is grounded in the cognitive process model of writing [17]. This allows writers to obtain model help in each of the three non-linear cognitive activities in the writing process: planning, translating and reviewing. Participants write short fiction/non-fiction with model help and are subsequently asked to submit a post-completion survey to provide qualitative feedback on the potential and pitfalls of LLMs as writing collaborators. Upon analyzing the writer-LLM interactions, we find that while seeking help across all three types of cognitive activities, writers find LLMs more helpful in translation and reviewing. Our findings from analyzing both the interactions and the survey responses highlight future research directions in creative writing assistance using LLMs. Tuhin Chakrabarty, Vishakh Padmakumar, Faeze Brahman, Smaranda Muresan |
Creativity & Cognition | 2 |
| 2024 | Does Writing with Language Models Reduce Content Diversity?abstractLarge language models (LLMs) have led to a surge in collaborative writing with model assistance. As different users incorporate suggestions from the same model, there is a risk of decreased diversity in the produced content, potentially limiting diverse perspectives in public discourse. In this work, we measure the impact of co-writing on diversity via a controlled experiment, where users write argumentative essays in three setups---using a base LLM (GPT3), a feedback-tuned LLM (InstructGPT), and writing without model help. We develop a set of diversity metrics and find that writing with InstructGPT (but not the GPT3) results in a statistically significant reduction in diversity. Specifically, it increases the similarity between the writings of different authors and reduces the overall lexical and content diversity. We additionally find that this effect is mainly attributable to InstructGPT contributing less diverse text to co-written essays. In contrast, the user-contributed text remains unaffected by model collaboration. This suggests that the recent improvement in generation quality from adapting models to human feedback might come at the cost of more homogeneous and less diverse content. Vishakh Padmakumar, He He 0001 |
ICLR | 1 |
| 2023 | Reward Gaming in Conditional Text GenerationabstractTo align conditional text generation model outputs with desired behaviors, there has been an increasing focus on training the model using reinforcement learning (RL) with reward functions learned from human annotations.Under this framework, we identify three common cases where high rewards are incorrectly assigned to undesirable patterns: noise-induced spurious correlation, naturally occurring spurious correlation, and covariate shift.We show that even though learned metrics achieve high performance on the distribution of the data used to train the reward function, the undesirable patterns may be amplified during RL training of the text generation model.While there has been discussion about reward gaming in the RL or safety community, in this discussion piece, we would like to highlight reward gaming in the natural language generation (NLG) community using concrete conditional text generation examples and discuss potential fixes and areas for future work. Richard Yuanzhe Pang, Vishakh Padmakumar, Thibault Sellam, Ankur P. Parikh, He He 0001 |
ACL (1) | 2 |
| 2023 | Extrapolative Controlled Sequence Generation via Iterative RefinementabstractWe study the problem of extrapolative controlled generation, i.e., generating sequences with attribute values beyond the range seen in training. This task is of significant importance in automated design, especially drug discovery, where the goal is to design novel proteins that are better (e.g., more stable) than existing sequences. Thus, by definition the target sequences and their attribute values are out of the training distribution, posing challenges to existing methods that aim to directly generate the target sequence. Instead, in this work, we propose Iterative Controlled Extrapolation (ICE) which iteratively makes local edits to a sequence to enable extrapolation. We train the model on synthetically generated sequence pairs that demonstrate small improvement in the attribute value. Results on one natural language task (sentiment analysis) and two protein engineering tasks (ACE2 stability and AAV fitness) show that ICE outperforms state-of-the-art approaches despite its simplicity. Vishakh Padmakumar, Richard Yuanzhe Pang, He He 0001, Ankur P. Parikh |
ICML | 1 |
| 2023 | Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesabstractGiven the intractably large size of the space of proofs, any model that is capable of general deductive reasoning must generalize to proofs of greater complexity. Recent studies have shown that large language models (LLMs) possess some abstract deductive reasoning ability given chain-of-thought prompts. However, they have primarily been tested on proofs using modus ponens or of a specific size, and from the same distribution as the in-context examples. To measure the general deductive reasoning ability of LLMs, we test on a broad set of deduction rules and measure their ability to generalize to more complex proofs from simpler demonstrations from multiple angles: depth-, width-, and compositional generalization. To facilitate systematic exploration, we construct a new synthetic and programmable reasoning dataset that enables control over deduction rules and proof complexity. Our experiments on four LLMs of various sizes and training objectives show that they are able to generalize to compositional proofs. However, they have difficulty generalizing to longer proofs, and they require explicit demonstrations to produce hypothetical subproofs, specifically in proof by cases and proof by contradiction. Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Mehran Kazemi, Najoung Kim, He He 0001 |
NeurIPS | 3 |
| 2023 | Investigating the Representation of Open Domain Dialogue Context for Transformer ModelsabstractVishakh Padmakumar, Behnam Hedayatnia, Di Jin, Patrick Lange, Seokhwan Kim, Nanyun Peng, Yang Liu, Dilek Hakkani-Tur. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Vishakh Padmakumar, Behnam Hedayatnia, Di Jin 0005, Patrick Lange, Seokhwan Kim, Nanyun Peng 0001, Yang Liu 0004, Dilek Hakkani-Tür |
SIGDIAL | 1 |
| 2022 | Help me write a Poem - Instruction Tuning as a Vehicle for Collaborative Poetry WritingabstractRecent work in training large language models (LLMs) to follow natural language instructions has opened up exciting opportunities for natural language interface design.Building on the prior success of LLMs in the realm of computerassisted creativity, we aim to study if LLMs can improve the quality of user-generated content through collaboration.We present CoPoet, a collaborative poetry writing system.In contrast to auto-completing a user's text, CoPoet is controlled by user instructions that specify the attributes of the desired text, such as Write a sentence about 'love' or Write a sentence ending in 'fly'.The core component of our system is a language model fine-tuned on a diverse collection of instructions for poetry writing.Our model is not only competitive with publicly available LLMs trained on instructions (InstructGPT), but is also capable of satisfying unseen compositional instructions.A study with 15 qualified crowdworkers shows that users successfully write poems with CoPoet on diverse topics ranging from Monarchy to Climate change.Further, the collaboratively written poems are preferred by third-party evaluators over those written without the system. 1 Tuhin Chakrabarty, Vishakh Padmakumar, He He 0001 |
EMNLP | 2 |
| 2022 | Machine-in-the-Loop Rewriting for Creative Image CaptioningabstractMachine-in-the-loop writing aims to build models that assist humans to accomplish their writing tasks more effectively.Prior work has found that providing users a machine-written draft or sentence-level continuations has limited success since the generated text tends to deviate from users' intention.To allow the user to retain control over the content, we train a rewriting model that, when prompted, modifies specified spans of text within the user's original draft to introduce descriptive and figurative elements in the text.We evaluate the model on its ability to collaborate with humans on the task of creative image captioning.On a user study through Amazon Mechanical Turk, our model is rated to be more helpful by users than a baseline infilling language model.In addition, third-party evaluation shows that users write more descriptive and figurative captions when collaborating with our model compared to completing the task alone.However, the improvement is not uniform across user groups: the model is more helpful to skilled users, which risks widening the gap between skilled and novice users, highlighting a need for careful, user-centric evaluation of interactive systems. 1 Vishakh Padmakumar, He He 0001 |
NAACL-HLT | 1 |
| 2022 | Exploring the Role of Task Transferability in Large-Scale Multi-Task LearningabstractVishakh Padmakumar, Leonard Lausen, Miguel Ballesteros, Sheng Zha, He He, George Karypis. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Vishakh Padmakumar, Leonard Lausen, Miguel Ballesteros, Sheng Zha, He He 0001, George Karypis |
NAACL-HLT | 1 |
| 2022 | QuALITY: Question Answering with Long Input Texts, Yes!abstractRichard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, Samuel Bowman. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He 0001, Samuel R. Bowman |
NAACL-HLT | 7 |
| 2021 | Unsupervised Extractive Summarization using Pointwise Mutual InformationabstractUnsupervised approaches to extractive summarization usually rely on a notion of sentence importance defined by the semantic similarity between a sentence and the document.We propose new metrics of relevance and redundancy using pointwise mutual information (PMI) between sentences, which can be easily computed by a pre-trained language model.Intuitively, a relevant sentence allows readers to infer the document content (high PMI with the document), and a redundant sentence can be inferred from the summary (high PMI with the summary).We then develop a greedy sentence selection algorithm to maximize relevance and minimize redundancy of extracted sentences.We show that our method outperforms similarity-based methods on datasets in a range of domains including news, medical journal articles, and personal anecdotes. Vishakh Padmakumar, He He 0001 |
EACL | 1 |