EDBT 2026 Demo / reviewers in the wild / expert
Suchin Gururangan
dblp:217/1570
· DBLP profile ↗
20ranked-venue papers
4as first author
17since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 16 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BTS: Harmonizing Specialized Experts into a Generalist LLMabstractQizhen Zhang, Prajjwal Bhargava, Chloe Bi, Chris X. Cai, Jakob Nicolaus Foerster, Jeremy Fu, Punit Singh Koura, Ruan Silva, Sheng Shen, Emily Dinan, Suchin Gururangan, Mike Lewis. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Qizhen Zhang 0002, Prajjwal Bhargava, Chloe Bi, Chris X. Cai, Jakob N. Foerster, Jeremy Fu, Punit Singh Koura, Ruan Silva, Sheng Shen 0016, Emily Dinan, Suchin Gururangan, Mike Lewis |
EMNLP | 11 |
| 2025 | Language models scale reliably with over-training and on downstream tasksabstractScaling laws are useful guides for derisking expensive training runs, as they predict performance of large models using cheaper, small-scale experiments. However, there remain gaps between current scaling studies and how language models are ultimately trained and evaluated. For instance, scaling is usually studied in the compute-optimal training regime (i.e., "Chinchilla optimal" regime). In contrast, models are often over-trained to reduce inference costs. Moreover, scaling laws mostly predict loss on next-token prediction, but models are usually compared on downstream task performance. To address both shortcomings, we create a testbed of 104 models with 0.011B to 6.9B parameters trained with various numbers of tokens on three data distributions. First, we fit scaling laws that extrapolate in both the amount of over-training and the number of model parameters. This enables us to predict the validation loss of a 1.4B parameter, 900B token run (i.e., 32$\times$ over-trained) and a 6.9B parameter, 138B token run (i.e., a compute-optimal run)––each from experiments that take 300$\times$ less compute. Second, we relate the perplexity of a language model to its downstream task performance by proposing a power law. We use this law to predict top-1 error averaged over downstream tasks for the two aforementioned models, using experiments that take 20$\times$ less compute. Samir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan, Mitchell Wortsman, Rulin Shao, Jean Mercat, Alex Fang, Jeffrey Li, Sedrick Keh, Marianna Nezhurina, Igor Vasiljevic, Luca Soldaini, Jenia Jitsev, Alexandros G. Dimakis, Gabriel Ilharco, Pang Wei Koh, Shuran Song, Thomas Kollar |
ICLR | 4 |
| 2025 | Self-Generated Critiques Boost Reward Modeling for Language ModelsabstractYue Yu, Zhengxing Chen, Aston Zhang, Liang Tan, Chenguang Zhu, Richard Yuanzhe Pang, Yundi Qian, Xuewei Wang, Suchin Gururangan, Chao Zhang, Melanie Kambadur, Dhruv Mahajan, Rui Hou. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yue Yu 0009, Zhengxing Chen, Aston Zhang, Chenguang Zhu 0001, Richard Yuanzhe Pang, Yundi Qian, Suchin Gururangan, Melanie Kambadur, Dhruv Mahajan 0001 |
NAACL (Long Papers) | 9 |
| 2024 | AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data FiltersabstractLi Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren Klein, Jesse Dodge. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Li Lucy, Suchin Gururangan, Luca Soldaini, Emma Strubell, David Bamman, Lauren F. Klein, Jesse Dodge |
ACL (1) | 2 |
| 2024 | Time is Encoded in the Weights of Finetuned Language ModelsabstractWe present time vectors, a simple tool to customize language models to new time periods.Time vectors are created by finetuning a language model on data from a single time (e.g., a year or month), and then subtracting the weights of the original pretrained model.This vector specifies a direction in weight space that, as our experiments show, improves performance on text from that time period.Time vectors specialized to adjacent time periods appear to be positioned closer together in a manifold.Using this structure, we interpolate between time vectors to induce new models that perform better on intervening and future time periods, without any additional training.We demonstrate the consistency of our findings across different tasks, domains, model sizes, and time scales.Our results suggest that time is encoded in the weight space of finetuned models. Kai Nylund, Suchin Gururangan, Noah A. Smith |
ACL (1) | 2 |
| 2024 | Breaking the Curse of Multilinguality with Cross-lingual Expert Language ModelsabstractTerra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, Luke Zettlemoyer. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Terra Blevins, Tomasz Limisiewicz, Suchin Gururangan, Margaret Li, Hila Gonen, Noah A. Smith, Luke Zettlemoyer |
EMNLP | 3 |
| 2024 | SILO Language Models: Isolating Legal Risk In a Nonparametric DatastoreabstractThe legality of training language models (LMs) on copyrighted or otherwise restricted data is under intense debate. However, as we show, model performance significantly degrades if trained only on low-risk text (e.g., out-of-copyright books or government documents), due to its limited size and domain coverage. We present SILO, a new language model that manages this risk-performance tradeoff during inference. SILO is built by (1) training a parametric LM on the Open License Corpus (OLC), a new corpus we curate with 228B tokens of public domain and permissively licensed text and (2) augmenting it with a more general and easily modifiable nonparametric datastore (e.g., containing copyrighted books or news) that is only queried during inference. The datastore allows use of high-risk data without training on it, supports sentence-level data attribution, and enables data producers to opt out from the model by removing content from the store. These capabilities can foster compliance with data-use regulations such as the fair use doctrine in the United States and the GDPR in the European Union. Our experiments show that the parametric LM struggles on its own with domains not covered by OLC. However, access to the datastore greatly improves out of domain performance, closing 90% of the performance gap with an LM trained on the Pile, a more diverse corpus with mostly high-risk text. We also analyze which nonparametric approach works best, where the remaining errors lie, and how performance scales with datastore size. Our results suggest that it is possible to build high quality language models while mitigating legal risk. Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A. Smith, Luke Zettlemoyer |
ICLR | 2 |
| 2024 | LESS: Selecting Influential Data for Targeted Instruction TuningabstractInstruction tuning has unlocked powerful capabilities in large language models (LLMs), using combined datasets to develop general-purpose chatbots. However, real-world applications often require a specialized suite of skills (e.g., reasoning). The challenge lies in identifying the most relevant data from these extensive datasets to effectively develop specific capabilities, a setting we frame as targeted instruction tuning. We propose LESS, an optimizer-aware and practically efficient algorithm to estimate data influences and perform Low-rank gradiEnt Similarity Search for instruction data selection. Crucially, LESS adapts existing influence formulations to work with the Adam optimizer and variable-length instruction data. LESS first constructs a highly reusable and transferable gradient datastore with low-dimensional gradient features and then selects examples based on their similarity to few-shot examples embodying a specific capability. Experiments show that training on a LESS-selected 5% of the data can often outperform training on the full dataset across diverse downstream tasks. Furthermore, the selected data is highly transferable: smaller models can be leveraged to select useful data for larger models and models from different families. Our qualitative analysis shows that our method goes beyond surface form cues to identify data that exemplifies the necessary reasoning skills for the intended downstream application. To facilitate future work, we release code and data at princeton-nlp/LESS. Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, Danqi Chen 0001 |
ICML | 3 |
| 2024 | DataComp-LM: In search of the next generation of training sets for language modelsabstractWe introduce DataComp for Language Models, a testbed for controlled dataset experiments with the goal of improving language models.As part of DCLM, we provide a standardized corpus of 240T tokens extracted from Common Crawl, effective pretraining recipes based on the OpenLM framework, and a broad suite of 53 downstream evaluations.Participants in the DCLM benchmark can experiment with data curation strategies such as deduplication, filtering, and data mixing atmodel scales ranging from 412M to 7B parameters.As a baseline for DCLM, we conduct extensive experiments and find that model-based filtering is key to assembling a high-quality training set.The resulting dataset, DCLM-Baseline, enables training a 7B parameter language model from scratch to 63% 5-shot accuracy on MMLU with 2T training tokens.Compared to MAP-Neo, the previous state-of-the-art in open-data language models, DCLM-Baseline represents a 6 percentage point improvement on MMLU while being trained with half the compute.Our results highlight the importance of dataset design for training language models and offer a starting point for further research on data curation. We release the \dclm benchmark, framework, models, and datasets at https://www.datacomp.ai/dclm/ Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Kumar Guha, Sedrick Keh, Kushal Arora, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee F. Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, Cheng-Yu Hsieh, Dhruba Ghosh, Josh Gardner 0001, Maciej Kilian, Hanlin Zhang 0002, Rulin Shao, Sarah M. Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang 0001, Khyathi Raghavi Chandu, Igor Vasiljevic, Sham M. Kakade, Shuran Song, Sujay Sanghavi, Fartash Faghri, Sewoong Oh, Luke Zettlemoyer, Kyle Lo, Alaaeldin El-Nouby, Hadi Pouransari, Alexander Toshev, Stephanie Wang, Dirk Groeneveld, Luca Soldaini, Pang Wei Koh, Jenia Jitsev, Thomas Kollar, Alexandros G. Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, Vaishaal Shankar |
NeurIPS | 17 |
| 2024 | Information Flow Control in Machine Learning through Modular Model Architecture
Trishita Tiwari, Suchin Gururangan, Chuan Guo 0001, Weizhe Hua, Sanjay Kariyappa, Udit Gupta 0001, Wenjie Xiong 0001, Kiwan Maeng, Hsien-Hsin S. Lee, G. Edward Suh |
USENIX Security Symposium | 2 |
| 2022 | Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data SelectionabstractSuchin Gururangan, Dallas Card, Sarah Dreier, Emily Gade, Leroy Wang, Zeyu Wang, Luke Zettlemoyer, Noah A. Smith. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Suchin Gururangan, Dallas Card, Sarah K. Dreier, Emily K. Gade, Leroy Z. Wang, Luke Zettlemoyer, Noah A. Smith |
EMNLP | 1 |
| 2022 | M2D2: A Massively Multi-Domain Language Modeling DatasetabstractWe present M2D2, a fine-grained, massively multi-domain corpus for studying domain adaptation in language models (LMs).M2D2 consists of 8.5B tokens and spans 145 domains extracted from Wikipedia and Semantic Scholar.Using ontologies derived from Wikipedia and ArXiv categories, we organize the domains in each data source into 22 groups.This two-level hierarchy enables the study of relationships between domains and their effects on in-and out-of-domain performance after adaptation.We also present a number of insights into the nature of effective domain adaptation in LMs, as examples of the new types of studies M2D2 enables.To improve in-domain performance, we show the benefits of adapting the LM along a domain hierarchy; adapting to smaller amounts of fine-grained domainspecific data can lead to larger in-domain performance gains than larger amounts of weakly relevant data.We further demonstrate a tradeoff between in-domain specialization and outof-domain generalization within and across ontologies, as well as a strong correlation between out-of-domain performance and lexical overlap between domains. Machel Reid, Victor Zhong, Suchin Gururangan, Luke Zettlemoyer |
EMNLP | 3 |
| 2022 | Nearest Neighbor Zero-Shot InferenceabstractRetrieval-augmented language models (LMs) use non-parametric memory to substantially outperform their non-retrieval counterparts on perplexity-based evaluations, but it is an open question whether they achieve similar gains in few- and zero-shot end-task accuracy. We extensively study one such model, the k-nearest neighbor LM (kNN-LM), showing that the gains marginally transfer. The main challenge is to achieve coverage of the verbalizer tokens that define the different end-task class labels. To address this challenge, we also introduce kNN-Prompt, a simple and effective kNN-LM with automatically expanded fuzzy verbalizers (e.g. to expand "terrible" to also include "silly" and other task-specific synonyms for sentiment classification). Across nine diverse end-tasks, using kNN-Prompt with GPT-2 large yields significant performance boosts over strong zeroshot baselines (13.4% absolute improvement over the base LM on average). We also show that other advantages of non-parametric augmentation hold for end tasks; kNN-Prompt is effective for domain adaptation with no further training, and gains increase with the size of the retrieval model. Julian Michael, Suchin Gururangan, Luke Zettlemoyer |
EMNLP | 3 |
| 2022 | DEMix Layers: Disentangling Domains for Modular Language ModelingabstractSuchin Gururangan, Mike Lewis, Ari Holtzman, Noah Smith, Luke Zettlemoyer. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, Luke Zettlemoyer |
NAACL-HLT | 1 |
| 2022 | Time Waits for No One! Analysis and Challenges of Temporal MisalignmentabstractKelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, Noah Smith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, Noah A. Smith |
NAACL-HLT | 3 |
| 2021 | All That's 'Human' Is Not Gold: Evaluating Human Evaluation of Generated TextabstractElizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, Noah A. Smith. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, Noah A. Smith |
ACL/IJCNLP (1) | 5 |
| 2021 | Detoxifying Language Models Risks Marginalizing Minority VoicesabstractAlbert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, Dan Klein. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, Daniel Klein 0001 |
NAACL-HLT | 4 |
| 2020 | Don't Stop Pretraining: Adapt Language Models to Domains and TasksabstractLanguage models pretrained on text from a wide variety of sources form the foundation of today's NLP. In light of the success of these broad-coverage models, we investigate whether it is still helpful to tailor a pretrained model to the domain of a target task. We present a study across four domains (biomedical and computer science publications, news, and reviews) and eight classification tasks, showing that a second phase of pretraining in-domain (domain-adaptive pretraining) leads to performance gains, under both high- and low-resource settings. Moreover, adapting to the task's unlabeled data (task-adaptive pretraining) improves performance even after domain-adaptive pretraining. Finally, we show that adapting to a task corpus augmented using simple data selection strategies is an effective alternative, especially when resources for domain-adaptive pretraining might be unavailable. Overall, we consistently find that multi-phase adaptive pretraining offers large gains in task performance. Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, Noah A. Smith |
ACL | 1 |
| 2019 | Variational Pretraining for Semi-supervised Text ClassificationabstractWe introduce VAMPIRE, 1 a lightweight pretraining framework for effective text classification when data and computing resources are limited.We pretrain a unigram document model as a variational autoencoder on in-domain, unlabeled data and use its internal states as features in a downstream classifier.Empirically, we show the relative strength of VAMPIRE against computationally expensive contextual embeddings and other popular semi-supervised baselines under low resource settings.We also find that fine-tuning to indomain data is crucial to achieving decent performance from contextual embeddings when working with limited supervision.We accompany this paper with code to pretrain and use VAMPIRE embeddings in downstream tasks. Suchin Gururangan, Tam Dang, Dallas Card, Noah A. Smith |
ACL (1) | 1 |
| 2019 | Show Your Work: Improved Reporting of Experimental ResultsabstractJesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, Noah A. Smith. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz 0001, Noah A. Smith |
EMNLP/IJCNLP (1) | 2 |