VLDB 2026 Research / reviewers in the wild / expert
Sneha Reddy Kudugunta
dblp:248/7500 · also Sneha Kudugunta
· DBLP profile ↗
11ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0002-0186-2433ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Machine translation · 36% Deep learning architectures and training · 23% Efficient and distributed learning · 16% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
1.5 | 3 | 2023 | MADLAD-400: A Multilingual And Document-Level Large Audited Dataset · NeurIPS 2023 Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation · ACL 2020 Investigating Multilingual NMT Representations at Scale · EMNLP/IJCNLP (1) 2019 |
Performance modeling and evaluation
benchmarking |
0.9 | 1 | 2025 | (Mis)Fitting Scaling Laws: A Survey of Scaling Law Fitting Techniques in Deep Learning · ICLR 2025 |
Performance modeling and evaluation › parallel system performance › speedup modeling
scaling laws |
0.9 | 1 | 2025 | (Mis)Fitting Scaling Laws: A Survey of Scaling Law Fitting Techniques in Deep Learning · ICLR 2025 |
Machine learning › Efficient and distributed learning › adaptive computation
elastic inference |
0.8 | 1 | 2024 | MatFormer: Nested Transformer for Elastic Inference · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | MiTTenS: A Dataset for Evaluating Gender Mistranslation · EMNLP 2024 |
Natural language and speech › Machine translation
gender bias in machine translation |
0.8 | 1 | 2024 | MiTTenS: A Dataset for Evaluating Gender Mistranslation · EMNLP 2024 |
Machine learning › Efficient and distributed learning › inference acceleration
speculative decoding |
0.8 | 1 | 2024 | MatFormer: Nested Transformer for Elastic Inference · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | MatFormer: Nested Transformer for Elastic Inference · NeurIPS 2024 |
Natural language and speech › Language models and text generation
multilingual language models |
0.7 | 1 | 2023 | MADLAD-400: A Multilingual And Document-Level Large Audited Dataset · NeurIPS 2023 |
Machine learning › Deep learning architectures and training
loss landscape |
0.6 | 1 | 2022 | A Loss Curvature Perspective on Training Instabilities of Deep Learning Models · ICLR 2022 |
Machine learning › Deep learning architectures and training › training dynamics
training instability |
0.6 | 1 | 2022 | A Loss Curvature Perspective on Training Instabilities of Deep Learning Models · ICLR 2022 |
Natural language and speech › Machine translation
monolingual data augmentation |
0.4 | 1 | 2020 | Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation · ACL 2020 |
Machine learning › Representation and self-supervised learning
representation analysis |
0.4 | 1 | 2019 | Investigating Multilingual NMT Representations at Scale · EMNLP/IJCNLP (1) 2019 |
Machine learning › Deep learning architectures and training › foundation model
foundation model training |
0.3 | 1 | 2025 | (Mis)Fitting Scaling Laws: A Survey of Scaling Law Fitting Techniques in Deep Learning · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
power law fitting · 1.7neural machine translation · 0.8nested feed-forward network · 0.8knowledge distillation · 0.8foundation model evaluation · 0.8large-scale pretraining · 0.7few-shot learning · 0.7loss curvature analysis · 0.6self-supervision · 0.4back-translation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | (Mis)Fitting Scaling Laws: A Survey of Scaling Law Fitting Techniques in Deep LearningabstractModern foundation models rely heavily on using scaling laws to guide crucial training decisions. Researchers often extrapolate the optimal architecture and hyper parameters settings from smaller training runs by describing the relationship between, loss, or task performance, and scale. All components of this process vary, from the specific equation being fit, to the training setup, to the optimization method. Each of these factors may affect the fitted law, and therefore, the conclusions of a given study. We discuss discrepancies in the conclusions that several prior works reach, on questions such as the optimal token to parameter ratio. We augment this discussion with our own analysis of the critical impact that changes in specific details may effect in a scaling study, and the resulting altered conclusions. Additionally, we survey over 50 papers that study scaling trends: while 45 of these papers quantify these trends using a power law, most under-report crucial details needed to reproduce their findings. To mitigate this, we we propose a checklist for authors to consider while contributing to scaling law research. Margaret Li, Sneha Reddy Kudugunta, Luke Zettlemoyer |
ICLR | 2 |
| 2024 | MiTTenS: A Dataset for Evaluating Gender MistranslationabstractTranslation systems, including foundation models capable of translation, can produce errors that result in gender mistranslations, and such errors create potential for harm.To measure the extent of such potential harms when translating into and out of English, we introduce a dataset, MiTTenS 1 , covering 26 languages from a variety of language families and scripts, including several traditionally underrepresented in digital resources.The dataset is constructed with handcrafted passages that target known failure patterns, longer synthetically generated passages, and natural passages sourced from multiple domains.We demonstrate the usefulness of the dataset by evaluating both neural machine translation systems and foundation models, and show that all systems exhibit gender mistranslation and potential harm, even in high resource languages. Kevin Robinson, Sneha Reddy Kudugunta, Romi Stella, Sunipa Dev, Jasmijn Bastings |
EMNLP | 2 |
| 2024 | BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual TransferabstractAkari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Akari Asai, Sneha Reddy Kudugunta, Xinyan Yu 0001, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, Hannaneh Hajishirzi |
NAACL-HLT | 2 |
| 2024 | MatFormer: Nested Transformer for Elastic InferenceabstractFoundation models are applied in a broad spectrum of settings with different inference constraints, from massive multi-accelerator clusters to resource-constrained standalone mobile devices. However, the substantial costs associated with training these models often limit the number of unique model sizes that can be offered. Consequently, practitioners are compelled to select a model that may not be optimally aligned with their specific latency and cost requirements. We present MatFormer, a novel Transformer architecture designed to provide elastic inference across diverse deployment constraints. MatFormer achieves this by incorporating a nested Feed Forward Network (FFN) block structure within a standard Transformer model. During training, we optimize the parameters of multiple nested FFN blocks with varying sizes, enabling the extraction of hundreds of accurate smaller models without incurring additional computational costs. We empirically validate the efficacy of MatFormer across different model classes (decoders and encoders) and modalities (language and vision), demonstrating its potential for real-world deployment. We show that a 850M decoder-only MatFormer language model (MatLM) allows us to extract multiple smaller models spanning from 582M to 850M parameters, each exhibiting better validation loss and one-shot downstream evaluations than independently trained counterparts. Furthermore, we observe that smaller encoders extracted from a universal MatFormer-based ViT (MatViT) encoder preserve the metric-space structure for adaptive large-scale retrieval. Finally, we showcase that speculative decoding with the accurate and consistent submodels extracted from MatFormer can lead to significant reduction in inference latency. Devvrit, Sneha Reddy Kudugunta, Aditya Kusupati, Tim Dettmers, Kaifeng Chen, Inderjit S. Dhillon, Yulia Tsvetkov, Hannaneh Hajishirzi, Sham M. Kakade, Ali Farhadi, Prateek Jain 0002 |
NeurIPS | 2 |
| 2023 | MADLAD-400: A Multilingual And Document-Level Large Audited DatasetabstractWe introduce MADLAD-400, a manually audited, general domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. We discuss the limitations revealed by self-auditing MADLAD-400, and the role data auditing had in the dataset creation process. We then train and release a 10.7B-parameter multilingual machine translation model on 250 billion tokens covering over 450 languages using publicly available data, and find that it is competitive with models that are significantly larger, and report the results on different domains. In addition, we train a 8B-parameter language model, and assess the results on few-shot translation. We make the baseline models available to the research community. Sneha Reddy Kudugunta, Isaac Caswell, Biao Zhang 0006, Xavier Garcia, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, Orhan Firat |
NeurIPS | 1 |
| 2022 | A Loss Curvature Perspective on Training Instabilities of Deep Learning Models
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Reddy Kudugunta, Behnam Neyshabur, David Cardoze, George E. Dahl, Zachary Nado, Orhan Firat |
ICLR | 4 |
| 2022 | Quality at a Glance: An Audit of Web-Crawled Multilingual DatasetsabstractAbstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with five major public datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, mC4). Lower-resource corpora have systematic issues: At least 15 corpora have no usable text, and a significant fraction contains less than 50% sentences of acceptable quality. In addition, many are mislabeled or use nonstandard/ambiguous language codes. We demonstrate that these issues are easy to detect even for non-proficient speakers, and supplement the human audit with automatic analyses. Finally, we recommend techniques to evaluate and improve multilingual corpora and discuss potential risks that come with low-quality data releases. Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov 0001, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias Müller 0002, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Reddy Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Balli, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofe Adeyemi |
Trans. Assoc. Comput. Linguistics | 33 |
| 2020 | Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine TranslationabstractAditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, Yonghui Wu. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Aditya Siddhant, Ankur Bapna, Yuan Cao 0007, Orhan Firat, Mia Xu Chen, Sneha Reddy Kudugunta, Naveen Arivazhagan |
ACL | 6 |
| 2020 | DANTE: Deep alternations for training neural networks
Vaibhav B. Sinha, Sneha Reddy Kudugunta, Adepu Ravi Sankar, Surya Teja Chavali, Vineeth N. Balasubramanian |
Neural Networks | 2 |
| 2019 | Investigating Multilingual NMT Representations at ScaleabstractSneha Kudugunta, Ankur Bapna, Isaac Caswell, Orhan Firat. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Sneha Reddy Kudugunta, Ankur Bapna, Isaac Caswell, Orhan Firat |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Deep neural networks for bot detection
Sneha Reddy Kudugunta, Emilio Ferrara |
Inf. Sci. | 1 |