Aman Madaan

dblp:138/1043 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
15since 2021 · last 2024
0009-0009-5604-9755ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Learning Performance-Improving Code Edits
abstract
With the decline of Moore's law, optimizing program performance has become a major focus of software research. However, high-level optimizations such as API and algorithm changes remain elusive due to the difficulty of understanding the semantics of code. Simultaneously, pretrained large language models (LLMs) have demonstrated strong capabilities at solving a wide range of programming tasks. To that end, we introduce a framework for adapting LLMs to high-level program optimization. First, we curate a dataset of performance-improving edits made by human programmers of over 77,000 competitive C++ programming submission pairs, accompanied by extensive unit tests. A major challenge is the significant variability of measuring performance on commodity hardware, which can lead to spurious "improvements." To isolate and reliably evaluate the impact of program optimizations, we design an environment based on the gem5 full system simulator, the de facto simulator used in academia and industry. Next, we propose a broad range of adaptation strategies for code optimization; for prompting, these include retrieval-based few-shot prompting and chain-of-thought, and for finetuning, these include performance-conditioned generation and synthetic data augmentation based on self-play. A combination of these techniques achieves a mean speedup of 6.86$\times$ with eight generations, higher than average optimizations from individual programmers (3.66$\times$). Using our model's fastest generations, we set a new upper limit on the fastest speedup possible for our dataset at 9.64$\times$ compared to using the fastest human submissions available (9.56$\times$).
Alexander Shypula, Aman Madaan, Yimeng Zeng, Uri Alon 0002, Jacob R. Gardner, Yiming Yang 0002, Milad Hashemi, Graham Neubig, Parthasarathy Ranganathan, Osbert Bastani, Amir Yazdanbakhsh
ICLR2
2024 In-Context Principle Learning from Mistakes
abstract
In-context learning (ICL, also known as few-shot prompting) has been the standard method of adapting LLMs to downstream tasks, by learning from a few input-output examples. Nonetheless, all ICL-based approaches only learn from correct input-output pairs. In this paper, we revisit this paradigm, by learning more from the few given input-output examples. We introduce Learning Principles (LEAP): First, we intentionally induce the model to make mistakes on these few examples; then we reflect on these mistakes, and learn explicit task-specific “principles” from them, which help solve similar problems and avoid common mistakes; finally, we prompt the model to answer unseen test questions using the original few-shot examples and these learned general principles. We evaluate LEAP on a wide range of benchmarks, including multi-hop question answering (Hotpot QA), textual QA (DROP), Big-Bench Hard reasoning, and math problems (GSM8K and MATH); in all these benchmarks, LEAP improves the strongest available LLMs such as GPT-3.5-turbo, GPT-4, GPT-4-turbo and Claude-2.1. For example, LEAP improves over the standard few-shot prompting using GPT-4 by 7.5% in DROP, and by 3.3% in HotpotQA. Importantly, LEAP does not require any more input or examples than the standard few-shot prompting settings.
Tianjun Zhang, Aman Madaan, Luyu Gao, Steven Zheng, Swaroop Mishra, Yiming Yang 0002, Niket Tandon, Uri Alon 0002
ICML2
2024 Program-Aided Reasoners (Better) Know What They Know
abstract
Anubha Kabra, Sanketh Rangreji, Yash Mathur, Aman Madaan, Emmy Liu, Graham Neubig. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Anubha Kabra, Sanketh Rangreji, Yash Mathur, Aman Madaan, Emmy Liu, Graham Neubig
NAACL-HLT4
2024 AutoMix: Automatically Mixing Language Models
abstract
Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively leveraging the options to optimize computational cost and performance remains challenging. In this work, we present AutoMix, an approach that strategically routes queries to larger LMs, based on the approximate correctness of outputs from a smaller LM. Central to AutoMix are two key technical contributions. First, it has a few-shot self-verification mechanism, which estimates the reliability of its own outputs without requiring extensive training. Second, given that self-verification can be noisy, it employs a POMDP based router that can effectively select an appropriately sized model, based on answer confidence. Experiments across five language models and five challenging datasets show that Automix consistently surpasses strong baselines, reducing computational cost by over 50\% for comparable performance.
Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Aditya Gupta 0001, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang 0002, Shyam Upadhyay, Manaal Faruqui, Mausam
NeurIPS2
2024 Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale
abstract
LLMs can now act as autonomous agents that interact with digital environments and complete specific objectives (e.g., arranging an online meeting). However, accuracy is still far from satisfactory, partly due to a lack of large-scale, direct demonstrations for digital tasks. Obtaining supervised data from humans is costly, and automatic data collection through exploration or reinforcement learning relies on complex environmental and content setup, resulting in datasets that lack comprehensive coverage of various scenarios. On the other hand, there is abundant knowledge that may indirectly assist task completion, such as online tutorials that were created for human consumption. In this work, we present Synatra, an approach that effectively transforms this indirect knowledge into direct supervision at scale. We define different types of indirect knowledge, and carefully study the available sources to obtain it, methods to encode the structure of direct demonstrations, and finally methods to transform indirect knowledge into direct demonstrations. We use 100k such synthetically-created demonstrations to finetune a 7B CodeLlama, and demonstrate that the resulting agent surpasses all comparably sized models on three web-based task benchmarks Mind2Web, MiniWoB++ and WebArena, as well as surpassing GPT-3.5 on WebArena and Mind2Web. In addition, while synthetic demonstrations prove to be only 3% the cost of human demonstrations (at $0.031 each), we show that the synthetic demonstrations can be more effective than an identical number of human demonstrations collected from limited domains.
Tianyue Ou, Frank F. Xu, Aman Madaan, Jiarui Liu 0004, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth 0001, Graham Neubig, Shuyan Zhou
NeurIPS3
2023 Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs
abstract
A popular approach for improving the correctness of output from large language models (LLMs) is Self-Consistency -poll the LLM multiple times and output the most frequent solution.Existing Self-Consistency techniques always generate a constant number of samples per question, where a better approach will be to non-uniformly distribute the available budget based on the amount of agreement in the samples generated so far.In response, we introduce Adaptive-Consistency, a cost-efficient, model-agnostic technique that dynamically adjusts the number of samples per question using a lightweight stopping criterion.Our experiments over 17 reasoning and code generation datasets and three LLMs demonstrate that Adaptive-Consistency reduces sample budget by up to 7.9 times with an average accuracy drop of less than 0.1%. 1
Pranjal Aggarwal, Aman Madaan, Yiming Yang 0002, Mausam
EMNLP2
2023 PAL: Program-aided Language Models
abstract
Large language models (LLMs) have demonstrated an impressive ability to perform arithmetic and symbolic reasoning tasks, when provided with a few examples at test time ("few-shot prompting"). Much of this success can be attributed to prompting methods such as "chain-of-thought", which employ LLMs for both understanding the problem description by decomposing it into steps, as well as solving each step of the problem. While LLMs seem to be adept at this sort of step-by-step decomposition, LLMs often make logical and arithmetic mistakes in the solution part, even when the problem is decomposed correctly. In this paper, we present Program-Aided Language models (PAL): a novel approach that uses the LLM to read natural language problems and generate programs as the intermediate reasoning steps, but offloads the solution step to a runtime such as a Python interpreter. With PAL, decomposing the natural language problem into runnable steps remains the only learning task for the LLM, while solving is delegated to the interpreter. We demonstrate this synergy between a neural LLM and a symbolic interpreter across 13 mathematical, symbolic, and algorithmic reasoning tasks from BIG-Bench Hard and others. In all these natural language reasoning tasks, generating code using an LLM and reasoning using a Python interpreter leads to more accurate results than much larger models. For example, PAL using Codex achieves state-of-the-art few-shot accuracy on GSM8K, surpassing PaLM which uses chain-of-thought by absolute 15% top-1.
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 0002, Pengfei Liu 0003, Yiming Yang 0002, Jamie Callan, Graham Neubig
ICML2
2023 Self-Refine: Iterative Refinement with Self-Feedback
abstract
Like humans, large language models (LLMs) do not always generate the best output on their first try. Motivated by how humans refine their written text, we introduce Self-Refine, an approach for improving initial outputs from LLMs through iterative feedback and refinement. The main idea is to generate an initial output using an LLMs; then, the same LLMs provides *feedback* for its output and uses it to *refine* itself, iteratively. Self-Refine does not require any supervised training data, additional training, or reinforcement learning, and instead uses a single LLM as the generator, refiner and the feedback provider. We evaluate Self-Refine across 7 diverse tasks, ranging from dialog response generation to mathematical reasoning, using state-of-the-art (GPT-3.5, ChatGPT, and GPT-4) LLMs. Across all evaluated tasks, outputs generated with Self-Refine are preferred by humans and automatic metrics over those generated with the same LLM using conventional one-step generation, improving by $\sim$20\% absolute on average in task performance. Our work demonstrates that even state-of-the-art LLMs like GPT-4 can be further improved at test-time using our simple, standalone approach.
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon 0002, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang 0002, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, Peter Clark
NeurIPS1
2023 Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation
abstract
Abstract Natural language generation has witnessed significant advancements due to the training of large language models on vast internet-scale datasets. Despite these advancements, there exists a critical challenge: These models can inadvertently generate content that is toxic, inaccurate, and unhelpful, and existing automatic evaluation metrics often fall short of identifying these shortcomings. As models become more capable, human feedback is an invaluable signal for evaluating and improving models. This survey aims to provide an overview of recent research that has leveraged human feedback to improve natural language generation. First, we introduce a taxonomy distilled from existing research to categorize and organize the varied forms of feedback. Next, we discuss how feedback can be described by its format and objective, and cover the two approaches proposed to use feedback (either for training or decoding): directly using feedback or training feedback models. We also discuss existing datasets for human-feedback data collection, and concerns surrounding feedback collection. Finally, we provide an overview of the nascent field of AI feedback, which uses large language models to make judgments based on a set of principles and minimize the need for human intervention. We also release a website of this survey at feedback-gap-survey.info.
Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José Guilherme Camargo de Souza, Shuyan Zhou, Sherry Tongshuang Wu, Graham Neubig, André F. T. Martins
Trans. Assoc. Comput. Linguistics2
2022 Conditional set generation using Seq2seq models
abstract
Conditional set generation learns a mapping from an input sequence of tokens to a set.Several NLP tasks, such as entity typing and dialogue emotion tagging, are instances of set generation.SEQ2SEQ models, a popular choice for set generation, treat a set as a sequence and do not fully leverage its key properties, namely order-invariance and cardinality.We propose a novel algorithm for effectively sampling informative orders over the combinatorial space of label orders.We jointly model the set cardinality and output by prepending the set size and taking advantage of the autoregressive factorization used by SEQ2SEQ models.Our method is a model-independent data augmentation approach that endows any SEQ2SEQ model with the signals of order-invariance and cardinality.Training a SEQ2SEQ model on this augmented data (without any additional annotations) gets an average relative improvement of 20% on four benchmark datasets across various models: BART-base, T5-11B, and GPT3-175B. 1
Aman Madaan, Dheeraj Rajagopal, Niket Tandon, Yiming Yang 0002, Antoine Bosselut
EMNLP1
2022 Memory-assisted prompt editing to improve GPT-3 after deployment
abstract
Large LMs such as GPT-3 are powerful, but can commit mistakes that are obvious to humans.For example, GPT-3 would mistakenly interpret "What word is similar to good?" to mean a homophone, while the user intended a synonym.Our goal is to effectively correct such errors via user interactions with the system but without retraining, which will be prohibitively costly.We pair GPT-3 with a growing memory of recorded cases where the model misunderstood the user's intents, along with user feedback for clarification.Such a memory allows our system to produce enhanced prompts for any new query based on the user feedback for error correction on similar cases in the past.On four tasks (two lexical tasks, two advanced ethical reasoning tasks), we show how a (simulated) user can interactively teach a deployed GPT-3, substantially increasing its accuracy over the queries with different kinds of misunderstandings by the GPT-3.Our approach is a step towards the low-cost utility enhancement for very large pre-trained LMs. 1 * Equal Contribution 1 Code, data, and instructions to implement MemPrompt for a new task at https://www.memprompt.com/
Aman Madaan, Niket Tandon, Peter Clark, Yiming Yang 0002
EMNLP1
2022 Language Models of Code are Few-Shot Commonsense Learners
abstract
We address the general task of structured commonsense reasoning: given a natural language input, the goal is to generate a graph such as an event or a reasoning-graph.To employ large language models (LMs) for this task, existing approaches "serialize" the output graph as a flat list of nodes and edges.Although feasible, these serialized graphs strongly deviate from the natural language corpora that LMs were pre-trained on, hindering LMs from generating them correctly.In this paper, we show that when we instead frame structured commonsense reasoning tasks as code generation tasks, pre-trained LMs of code are better structured commonsense reasoners than LMs of natural language, even when the downstream task does not involve source code at all.We demonstrate our approach across three diverse structured commonsense reasoning tasks.In all these natural language tasks, we show that using our approach, a code generation LM (CODEX) outperforms natural-LMs that are fine-tuned on the target task (e.g., T5) and other strong LMs such as GPT-3 in the few-shot setting.Our code and data are available at https: //github.com/madaan/CoCoGen .
Aman Madaan, Shuyan Zhou, Uri Alon 0002, Yiming Yang 0002, Graham Neubig
EMNLP1
2021 Towards Using Heterogeneous Relation Graphs for End-to-End TTS
abstract
Neural models for end-to-end text-to-speech (TTS) synthe-sis are increasingly outperforming traditional approaches in statistical parametric speech synthesis. Speech generation in these neural models predominantly relies on using free-form text as the input modality. However, the earlier statistical parametric models were built on encoded phonetic and syn-tactic features. In this work, we explore the possibility of explicitly feeding deterministic linguistic structure to a neural TTS system in the form of Heterogeneous Relational Graphs (HRGs), an expressive formalism capable of representing pho-netic and syntactic information. Specifically, we use Graph Convolutional Networks to learn structurally informed contin-uous representations of the HRGs, which can be seamlessly passed to the encoders of popular neural TTS models like TransformerTTS or Tacotron. Furthermore, our simple HRG based text-to-speech synthesis leverages the syntactic bias in HRGs as demonstrated by improvements in automated met-rics and human evaluation on i) the single speaker dataset LJSpeech; ii) the multi-speaker dataset Arctic; and iii) out-of-domain test sets from the Blizzard challenge.11The code, trained models, and our dataset of HRGs will be released at https://github.com/ars22/GraphNeuralTTS/.
Amrith Setlur, Aman Madaan, Tanmay Parekh, Yiming Yang 0002, Alan W. Black
ASRU2
2021 Think about it! Improving defeasible reasoning by first modeling the question scenario
abstract
Defeasible reasoning is the mode of reasoning where conclusions can be overturned by taking into account new evidence.Existing cognitive science literature on defeasible reasoning suggests that a person forms a mental model of the problem scenario before answering questions.Our research goal asks whether neural models can similarly benefit from envisioning the question scenario before answering a defeasible query.Our approach is, given a question, to have a model first create a graph of relevant influences, and then leverage that graph as an additional input when answering the question.Our system, CURIOUS, achieves a new stateof-the-art on three different defeasible reasoning datasets.This result is significant as it illustrates that performance can be improved by guiding a system to "think about" a question and explicitly model the scenario, rather than answering reflexively. 1
Aman Madaan, Niket Tandon, Dheeraj Rajagopal, Peter Clark, Yiming Yang 0002, Eduard H. Hovy
EMNLP (1)1
2021 Neural Language Modeling for Contextualized Temporal Graph Generation
abstract
This paper presents the first study on using large-scale pre-trained language models for automated generation of an event-level temporal graph for a document.Despite the huge success of neural pre-training methods in NLP tasks, its potential for temporal reasoning over event graphs has not been sufficiently explored.Part of the reason is the difficulty in obtaining large training corpora with humanannotated events and temporal links.We address this challenge by using existing IE/NLP tools to automatically generate a large quantity (89,000) of system-produced document-graph pairs, and propose a novel formulation of the contextualized graph generation problem as a sequence-to-sequence mapping task.These strategies enable us to leverage and fine-tune pre-trained language models on the systeminduced training data for the graph generation task.Our experiments show that our approach is highly effective in generating structurally and semantically valid graphs.Further, evaluation on a challenging hand-labeled, out-ofdomain corpus shows that our method outperforms the closest existing method by a large margin on several metrics.We also show a downstream application of our approach by adapting it to answer open-ended temporal questions in a reading comprehension setting. 1
Aman Madaan, Yiming Yang 0002
NAACL-HLT1
2020 Politeness Transfer: A Tag and Generate Approach
abstract
Aman Madaan, Amrith Setlur, Tanmay Parekh, Barnabas Poczos, Graham Neubig, Yiming Yang, Ruslan Salakhutdinov, Alan W Black, Shrimai Prabhumoye. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Aman Madaan, Amrith Setlur, Tanmay Parekh, Barnabás Póczos, Graham Neubig, Yiming Yang 0002, Ruslan Salakhutdinov, Alan W. Black, Shrimai Prabhumoye
ACL1
2016 Numerical Relation Extraction with Minimal Supervision
abstract
We study a novel task of numerical relation extraction with the goal of extracting relations where one of the arguments is a number or a quantity ( e.g., atomic_number(Aluminium, 13), inflation_rate(India, 10.9%)). This task presents peculiar challenges not found in standard IE, such as the difficulty of matching numbers in distant supervision and the importance of units. We design two extraction systems that require minimal human supervision per relation: (1) NumberRule, a rule based extractor, and (2) NumberTron, a probabilistic graphical model. We find that both systems dramatically outperform MultiR, a state-of-the-art non-numerical IE model, obtaining up to 25 points F-score improvement.
Aman Madaan, Ashish R. Mittal, Mausam, Ganesh Ramakrishnan, Sunita Sarawagi
AAAI1