Srinivasan Iyer 0001

dblp:78/4928-1 · also Srini Iyer 0001 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0002-6186-2603ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 5 first-author · 13 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Byte Latent Transformer: Patches Scale Better Than Tokens
abstract
Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srini Iyer. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez 0001, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srinivasan Iyer 0001
ACL (1)14
2024 Instruction-tuned Language Models are Better Knowledge Learners
abstract
Zhengbao Jiang, Zhiqing Sun, Weijia Shi, Pedro Rodriguez, Chunting Zhou, Graham Neubig, Xi Lin, Wen-tau Yih, Srini Iyer. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zhengbao Jiang, Zhiqing Sun, Pedro Rodríguez 0001, Chunting Zhou, Graham Neubig, Xi Victoria Lin, Scott Yih, Srinivasan Iyer 0001
ACL (1)9
2023 Methods for Measuring, Updating, and Visualizing Factual Beliefs in Language Models
abstract
Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Peter Hase, Mona T. Diab, Asli Celikyilmaz, Xian Li 0003, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, Srinivasan Iyer 0001
EACL8
2023 LEVER: Learning to Verify Language-to-Code Generation with Execution
abstract
The advent of large language models trained on code (code LLMs) has led to significant progress in language-to-code generation. State-of-the-art approaches in this area combine LLM decoding with sample pruning and reranking using test cases or heuristics based on the execution results. However, it is challenging to obtain test cases for many real-world language-to-code applications, and heuristics cannot well capture the semantic features of the execution results, such as data type and value range, which often indicates the correctness of the program. In this work, we propose LEVER, a simple approach to improve language-to-code generation by learning to verify the generated programs with their execution results. Specifically, we train verifiers to determine whether a program sampled from the LLMs is correct or not based on the natural language input, the program itself and its execution results. The sampled programs are reranked by combining the verification score with the LLM generation probability, and marginalizing over programs with the same execution results. On four datasets across the domains of table QA, math QA and basic Python programming, LEVER consistently improves over the base code LLMs (4.6% to 10.9% with code-davinci-002) and achieves new state-of-the-art results on all of them.
Ansong Ni, Srinivasan Iyer 0001, Dragomir R. Radev, Veselin Stoyanov, Scott Yih, Sida I. Wang, Xi Victoria Lin
ICML2
2023 LIMA: Less Is More for Alignment
abstract
Large language models are trained in two stages: (1) unsupervised pretraining from raw text, to learn general-purpose representations, and (2) large scale instruction tuning and reinforcement learning, to better align to end tasks and user preferences. We measure the relative importance of these two stages by training LIMA, a 65B parameter LLaMa language model fine-tuned with the standard supervised loss on only 1,000 carefully curated prompts and responses, without any reinforcement learning or human preference modeling. LIMA demonstrates remarkably strong performance, learning to follow specific response formats from only a handful of examples in the training data, including complex queries that range from planning trip itineraries to speculating about alternate history. Moreover, the model tends to generalize well to unseen tasks that did not appear in the training data. In a controlled human study, responses from LIMA are either equivalent or strictly preferred to GPT-4 in 43\% of cases; this statistic is as high as 58\% when compared to Bard and 65\% versus DaVinci003, which was trained with human feedback. Taken together, these results strongly suggest that almost all knowledge in large language models is learned during pretraining, and only limited instruction tuning data is necessary to teach models to produce high quality output.
Chunting Zhou, Puxin Xu, Srinivasan Iyer 0001, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Lili Yu, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, Omer Levy
NeurIPS4
2022 ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech Detection
abstract
Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer, Veselin Stoyanov, Zornitsa Kozareva, Xian Li, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona Diab. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Badr AlKhamissi, Faisal Ladhak, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva, Xian Li 0003, Pascale Fung, Lambert Mathias, Asli Celikyilmaz, Mona T. Diab
EMNLP3
2022 Efficient Large Scale Language Modeling with Mixtures of Experts
abstract
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O’Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, Veselin Stoyanov. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Mikel Artetxe, Shruti Bhosale, Naman Goyal 0001, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer 0001, Ramakanth Pasunuru, Giri Anantharaman, Xian Li 0003, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Punit Singh Koura, Brian O'Horo, Jeffrey Wang, Luke Zettlemoyer, Mona T. Diab, Zornitsa Kozareva, Veselin Stoyanov
EMNLP9
2022 Improving In-Context Few-Shot Learning via Self-Supervised Training
abstract
Mingda Chen, Jingfei Du, Ramakanth Pasunuru, Todor Mihaylov, Srini Iyer, Veselin Stoyanov, Zornitsa Kozareva. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Mingda Chen, Jingfei Du, Ramakanth Pasunuru, Todor Mihaylov, Srinivasan Iyer 0001, Veselin Stoyanov, Zornitsa Kozareva
NAACL-HLT5
2022 AnswerSumm: A Manually-Curated Dataset and Pipeline for Answer Summarization
abstract
Alexander Fabbri, Xiaojian Wu, Srini Iyer, Haoran Li, Mona Diab. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Alexander R. Fabbri, Xiaojian Wu, Srinivasan Iyer 0001, Haoran Li 0007, Mona T. Diab
NAACL-HLT3
2022 QUASER: Question Answering with Scalable Extractive Rationalization
abstract
Designing natural language processing (NLP) models that produce predictions by first extracting a set of relevant input sentences, i.e., rationales, is gaining importance for improving model interpretability and producing supporting evidence for users. Current unsupervised approaches are designed to extract rationales that maximize prediction accuracy, which is invariably obtained by exploiting spurious correlations in datasets, and leads to unconvincing rationales. In this paper, we introduce unsupervised generative models to extract dual-purpose rationales, which must not only be able to support a subsequent answer prediction, but also support a reproduction of the input query. We show that such models can produce more meaningful rationales, that are less influenced by dataset artifacts, and as a result, also achieve the state-of-the-art on rationale extraction metrics on four datasets from the ERASER benchmark, significantly improving upon previous unsupervised methods. Our multi-task model is scalable and enables using state-of-the-art pretrained language models to design explainable question answering systems.
Asish Ghoshal, Srinivasan Iyer 0001, Bhargavi Paranjape, Kushal Lakhotia, Scott Yih, Yashar Mehdad
SIGIR2
2021 FiD-Ex: Improving Sequence-to-Sequence Models for Extractive Rationale Generation
abstract
Natural language (NL) explanations of model predictions are gaining popularity as a means to understand and verify decisions made by large black-box pre-trained models, for tasks such as Question Answering (QA) and Fact Verification.Recently, pre-trained sequence to sequence (seq2seq) models have proven to be very effective in jointly making predictions, as well as generating NL explanations.However, these models have many shortcomings; they can fabricate explanations even for incorrect predictions, they are difficult to adapt to long input documents, and their training requires a large amount of labeled data.In this paper, we develop FiD-Ex 1 , which addresses these shortcomings for seq2seq models by: 1) introducing sentence markers to eliminate explanation fabrication by encouraging extractive generation, 2) using the fusion-in-decoder architecture to handle long input contexts, and 3) intermediate fine-tuning on re-structured open domain QA datasets to improve few-shot performance.FiD-Ex significantly improves over prior work in terms of explanation metrics and task accuracy on five tasks from the ERASER explainability benchmark in both fully supervised and few-shot settings.
Kushal Lakhotia, Bhargavi Paranjape, Asish Ghoshal, Scott Yih, Yashar Mehdad, Srinivasan Iyer 0001
EMNLP (1)6
2021 DeLighT: Deep and Light-weight Transformer
Sachin Mehta, Marjan Ghazvininejad, Srinivasan Iyer 0001, Luke Zettlemoyer, Hannaneh Hajishirzi
ICLR3
2021 Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval
Wenhan Xiong, Xiang Li 0069, Srinivasan Iyer 0001, Jingfei Du, Patrick S. H. Lewis, William Yang Wang, Yashar Mehdad, Scott Yih, Sebastian Riedel 0001, Douwe Kiela, Barlas Oguz
ICLR3
2021 RECONSIDER: Improved Re-Ranking using Span-Focused Cross-Attention for Open Domain Question Answering
abstract
Srinivasan Iyer, Sewon Min, Yashar Mehdad, Wen-tau Yih. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Srinivasan Iyer 0001, Sewon Min, Yashar Mehdad, Scott Yih
NAACL-HLT1
2020 Efficient One-Pass End-to-End Entity Linking for Questions
abstract
We present ELQ, a fast end-to-end entity linking model for questions, which uses a biencoder to jointly perform mention detection and linking in one pass.Evaluated on WebQSP and GraphQuestions with extended annotations that cover multiple entities per question, ELQ outperforms the previous state of the art by a large margin of +12.7% and +19.6% F1, respectively.With a very fast inference time (1.57examples/s on a single CPU), ELQ can be useful for downstream question answering systems.In a proof-of-concept experiment, we demonstrate that using ELQ significantly improves the downstream QA performance of GraphRetriever (Min et al., 2019). 1
Belinda Z. Li, Sewon Min, Srinivasan Iyer 0001, Yashar Mehdad, Scott Yih
EMNLP (1)3
2019 JuICe: A Large Scale Distantly Supervised Dataset for Open Domain Context-based Code Generation
abstract
Rajas Agashe, Srinivasan Iyer, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Rajas Agashe, Srinivasan Iyer 0001, Luke Zettlemoyer
EMNLP/IJCNLP (1)2
2019 Learning Programmatic Idioms for Scalable Semantic Parsing
abstract
Srinivasan Iyer, Alvin Cheung, Luke Zettlemoyer. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Srinivasan Iyer 0001, Alvin Cheung, Luke Zettlemoyer
EMNLP/IJCNLP (1)1
2018 Mapping Language to Code in Programmatic Context
abstract
Source code is rarely written in isolation.It depends significantly on the programmatic context, such as the class that the code would reside in.To study this phenomenon, we introduce the task of generating class member functions given English documentation and the programmatic context provided by the rest of the class.This task is challenging because the desired code can vary greatly depending on the functionality the class provides (e.g., a sort function may or may not be available when we are asked to "return the smallest element" in a particular member variable list).We introduce CONCODE, a new large dataset with over 100,000 examples consisting of Java classes from online code repositories, and develop a new encoder-decoder architecture that models the interaction between the method documentation and the class environment.We also present a detailed error analysis suggesting that there is significant room for future work on this task.
Srinivasan Iyer 0001, Ioannis Konstas, Alvin Cheung, Luke Zettlemoyer
EMNLP1
2018 Learning to Map Context-Dependent Sentences to Executable Formal Queries
abstract
Alane Suhr, Srinivasan Iyer, Yoav Artzi. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Alane Suhr, Srinivasan Iyer 0001, Yoav Artzi
NAACL-HLT2
2017 Learning a Neural Semantic Parser from User Feedback
abstract
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, Luke Zettlemoyer. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Srinivasan Iyer 0001, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, Luke Zettlemoyer
ACL (1)1
2017 Neural AMR: Sequence-to-Sequence Models for Parsing and Generation
abstract
Sequence-to-sequence models have shown strong performance across a broad range of applications.However, their application to parsing and generating text using Abstract Meaning Representation (AMR) has been limited, due to the relatively limited amount of labeled data and the nonsequential nature of the AMR graphs.We present a novel training procedure that can lift this limitation using millions of unlabeled sentences and careful preprocessing of the AMR graphs.For AMR parsing, our model achieves competitive results of 62.1 SMATCH, the current best score reported without significant use of external semantic resources.For AMR generation, our model establishes a new state-of-the-art performance of BLEU 33.8.We present extensive ablative and qualitative analysis including strong evidence that sequencebased AMR models are robust against ordering variations of graph-to-sequence conversions.
Ioannis Konstas, Srinivasan Iyer 0001, Mark Yatskar, Yejin Choi 0001, Luke Zettlemoyer
ACL (1)2
2016 Summarizing Source Code using a Neural Attention Model
abstract
High quality source code is often paired with high level summaries of the computation it performs, for example in code documentation or in descriptions posted in online forums.Such summaries are extremely useful for applications such as code search but are expensive to manually author, hence only done for a small fraction of all code that is produced.In this paper, we present the first completely datadriven approach for generating high level summaries of source code.Our model, CODE-NN , uses Long Short Term Memory (LSTM) networks with attention to produce sentences that describe C# code snippets and SQL queries.CODE-NN is trained on a new corpus that is automatically collected from StackOverflow, which we release.Experiments demonstrate strong performance on two tasks: (1) code summarization, where we establish the first end-to-end learning results and outperform strong baselines, and (2) code retrieval, where our learned model improves the state of the art on a recently introduced C# benchmark by a large margin.
Srinivasan Iyer 0001, Ioannis Konstas, Alvin Cheung, Luke Zettlemoyer
ACL (1)1