VLDB 2026 Research / reviewers in the wild / expert
Vedanuj Goswami
dblp:156/5885
· DBLP profile ↗
16ranked-venue papers
0as first author
12since 2021 · last 2025
0009-0005-5027-1452ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Llama 3 Training with Efficient Parallelism StrategiesabstractLlama is a widely used open-source large language model.This paper presents the design and implementation of the parallelism techniques used in Llama 3 pre-training.To achieve efficient training on tens of thousands of GPUs, Llama 3 employs a combination of four-dimensional parallelism: fully sharded data parallelism, tensor parallelism, pipeline parallelism, and context parallelism.Beyond achieving efficiency through parallelism and model co-design, we Weiwei Chu, Xinfeng Xie, Jiecao Yu, Jie Wang 0022, Amar Phanishayee, Chunqiang Tang, Yuchen Hao, Muhammet Mustafa Ozdal, Vedanuj Goswami, Naman Goyal 0001, Abhishek Kadian, Andrew Gu, Chris Cai, Xiaodong Wang 0020, Min Si, Pavan Balaji, Ching-Hsiang Chu, Jongsoo Park |
ISCA | 11 |
| 2023 | Causes and Cures for Interference in Multilingual TranslationabstractMultilingual machine translation models can benefit from synergy between different language pairs, but also suffer from interference.While there is a growing number of sophisticated methods that aim to eliminate interference, our understanding of interference as a phenomenon is still limited.This work identifies the main factors that contribute to interference in multilingual machine translation.Through systematic experimentation, we find that interference (or synergy) are primarily determined by model size, data size, and the proportion of each language pair within the total dataset.We observe that substantial interference occurs mainly when the model is very small with respect to the available training data, and that using standard transformer configurations with less than one billion parameters largely alleviates interference and promotes synergy.Moreover, we show that tuning the sampling temperature to control the proportion of each language pair in the data is key to balancing the amount of interference between low and high resource language pairs effectively, and can lead to superior performance overall. Uri Shaham 0002, Maha Elbayad, Vedanuj Goswami, Omer Levy, Shruti Bhosale |
ACL (1) | 3 |
| 2023 | SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech TranslationsabstractPaul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, Holger Schwenk. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Paul-Ambroise Duquenne, Hongyu Gong, Jingfei Du, Ann Lee 0001, Vedanuj Goswami, Changhan Wang, Juan Pino 0001, Benoît Sagot, Holger Schwenk |
ACL (1) | 6 |
| 2023 | Small Data, Big Impact: Leveraging Minimal Data for Effective Machine TranslationabstractJean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzman. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan, Vedanuj Goswami, Philipp Koehn, Angela Fan, Francisco Guzmán |
ACL (1) | 5 |
| 2023 | Revisiting Machine Translation for Cross-lingual ClassificationabstractMachine Translation (MT) has been widely used for cross-lingual classification, either by translating the test set into English and running inference with a monolingual model (translatetest), or translating the training set into the target languages and finetuning a multilingual model (translate-train).However, most research in the area focuses on the multilingual models rather than the MT component.We show that, by using a stronger MT system and mitigating the mismatch between training on original text and running inference on machine translated text, translate-test can do substantially better than previously assumed.The optimal approach, however, is highly task dependent, as we identify various sources of cross-lingual transfer gap that affect different tasks and approaches differently.Our work calls into question the dominance of multilingual models for cross-lingual classification, and prompts to pay more attention to MTbased baselines. Mikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan, Luke Zettlemoyer |
EMNLP | 2 |
| 2023 | MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation
Mohamed Anwar, Bowen Shi 0002, Vedanuj Goswami, Wei-Ning Hsu, Juan Pino 0001, Changhan Wang |
INTERSPEECH | 3 |
| 2022 | FLAVA: A Foundational Language And Vision Alignment ModelabstractState-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks. Generally, such models are often either cross-modal (contrastive) or multi-modal (with earlier fusion) but not both; and they often only target specific modalities or tasks. A promising direction would be to use a single holistic universal model, as a “foundation”, that targets all modalities at once-a true vision and language foundation model should be good at vision tasks, language tasks, and cross- and multi-modal vision and language tasks. We introduce FLAVA as such a model and demonstrate impressive performance on a wide range of 35 tasks spanning these target modalities. Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, Douwe Kiela |
CVPR | 3 |
| 2022 | Tricks for Training Sparse Translation ModelsabstractDheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross, Mike Lewis, Angela Fan. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Dheeru Dua, Shruti Bhosale, Vedanuj Goswami, James Cross 0003, Mike Lewis, Angela Fan |
NAACL-HLT | 3 |
| 2021 | Creative Sketch Generation
Songwei Ge, Vedanuj Goswami, C. Lawrence Zitnick, Devi Parikh |
ICLR | 2 |
| 2021 | MoVie: Revisiting Modulated Convolutions for Visual Counting and Beyond
Duy-Kien Nguyen, Vedanuj Goswami, Xinlei Chen |
ICLR | 2 |
| 2021 | Human-Adversarial Visual Question AnsweringabstractPerformance on the most commonly used Visual Question Answering dataset (VQA v2) is starting to approach human accuracy. However, in interacting with state-of-the-art VQA models, it is clear that the problem is far from being solved. In order to stress test VQA models, we benchmark them against human-adversarial examples. Human subjects interact with a state-of-the-art VQA model, and for each image in the dataset, attempt to find a question where the model’s predicted answer is incorrect. We find that a wide range of state-of-the-art models perform poorly when evaluated on these examples. We conduct an extensive analysis of the collected adversarial examples and provide guidance on future research directions. We hope that this Adversarial VQA (AdVQA) benchmark can help drive progress in the field and advance the state of the art. Sasha Sheng, Amanpreet Singh, Vedanuj Goswami, José Alberto López Magaña, Tristan Thrush, Wojciech Galuba, Devi Parikh, Douwe Kiela |
NeurIPS | 3 |
| 2021 | Only Time Can Tell: Discovering Temporal Data for Temporal ModelingabstractUnderstanding temporal information and how the visual world changes over time, is a fundamental ability of intelligent systems. In video understanding, temporal information is at the core of many current challenges, including compression, efficient inference, motion estimation or summarization. However, in current video datasets it has been observed that action classes can often be recognized without any temporal information, from a single frame of video. As a result, both benchmarking and training in these datasets may give an unintentional advantage to models with strong image understanding capabilities, as opposed to those with strong temporal understanding. In other words, current datasets may not reward good temporal understanding, potentially hindering progress. In this paper we address this problem head on by identifying action classes where temporal information is actually necessary to recognize them and call these "temporal classes". Selecting temporal classes using a computational method would bias the process. Instead, we propose a methodology based on a simple and effective human annotation experiment. We remove just the temporal information, by shuffling frames in time, and measure if the action can still be recognized. Classes that cannot be recognized when frames are not in order, are included in the temporal set. We observe that this set is statistically different from other static classes, and that performance in it correlates with a network's ability to capture temporal information. Thus we use it as a benchmark on current popular networks, which reveals a series of interesting facts, like inflated convolutions bias networks towards classes where motion is not important. We also explore the effect of training on the temporal set, and observe that this leads to better generalization in unseen classes, demonstrating the need for more temporal data. We hope that the proposed dataset of temporal categories will help guide future research in temporal modeling for better video understanding. Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan 0001, Vedanuj Goswami, Matt Feiszli, Lorenzo Torresani |
WACV | 4 |
| 2020 | 12-in-1: Multi-Task Vision and Language Representation LearningabstractMuch of vision-and-language research focuses on a small but diverse set of independent tasks and supporting datasets often studied in isolation; however, the visually-grounded language understanding skills required for success at these tasks overlap significantly. In this work, we investigate these relationships between vision-and-language tasks by developing a large-scale, multi-task model. Our approach culminates in a single model on 12 datasets from four broad categories of task including visual question answering, caption-based image retrieval, grounding referring expressions, and multimodal verification. Compared to independently trained single-task models, this represents a reduction from approximately 3 billion parameters to 270 million while simultaneously improving performance by 2.05 points on average across tasks. We use our multi-task framework to perform in-depth analysis of the effect of joint training diverse tasks. Further, we show that finetuning task-specific models from our single multi-task model can lead to further improvements, achieving performance at or above the state-of-the-art. Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, Stefan Lee |
CVPR | 2 |
| 2020 | Building Recommender Systems with PyTorchabstractIn this tutorial we show how to build deep learning recommendation systems and resolve the associated interpretability, integrity and privacy challenges. We start with an overview of the PyTorch framework, features that it offers and a brief review of the evolution of recommendation models. We delineate their typical components and build a proxy deep learning recommendation model (DLRM) in PyTorch. Then, we discuss how to interpret recommendation system results as well as how to address the corresponding integrity and quality challenges. Dheevatsa Mudigere, Maxim Naumov, Joe Spisak, Geeta Chauhan, Narine Kokhlikyan, Amanpreet Singh, Vedanuj Goswami |
KDD | 7 |
| 2020 | The Hateful Memes Challenge: Detecting Hate Speech in Multimodal MemesabstractThis work proposes a new challenge set for multimodal classification, focusing on detecting hate speech in multimodal memes. It is constructed such that unimodal models struggle and only multimodal models can succeed: difficult examples (“benign confounders”) are added to the dataset to make it hard to rely on unimodal signals. The task requires subtle reasoning, yet is straightforward to evaluate as a binary classification problem. We provide baseline performance numbers for unimodal models, as well as for multimodal models with various degrees of sophistication. We find that state-of-the-art methods perform poorly compared to humans, illustrating the difficulty of the task and highlighting the challenge that this important problem poses to the community. Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, Davide Testuggine |
NeurIPS | 4 |
| 2016 | Knowledge Extraction and Annotation for Cross-Domain Textual Case-Based Reasoning in Biologically Inspired Design
Spencer Rugaber, Shruti Bhati, Vedanuj Goswami, Evangelia Spiliopoulou, Sasha Azad, Sridevi Koushik, Rishikesh Kulkarni, Mithun Kumble, Sriya Sarathy, Ashok K. Goel 0001 |
ICCBR | 3 |