Aditya Siddhant

dblp:211/7727 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 44% Machine translation · 16% Transfer learning and domain adaptation · 15%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.922020
XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation · ICML 2020
Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine Translation · AAAI 2020
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.922020
Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation · ACL 2020
Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine Translation · AAAI 2020
Natural language and speech › Language models and text generation › text summarization
summarization evaluation
0.712023
SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation · EMNLP 2023
Natural language and speech › Language models and text generation
text generation evaluation
0.712023
Dialect-robust Evaluation of Generated Text · ACL (1) 2023
Natural language and speech › Language models and text generation
multilingual language models
0.622020
XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation · ICML 2020
Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine Translation · AAAI 2020
Natural language and speech › Language models and text generation › evaluation of language models › multilingual evaluation
multilingual language model evaluation
0.512021
XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation · EMNLP (1) 2021
Natural language and speech › Language models and text generation › natural language understanding
multilingual language understanding
0.512021
XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation · EMNLP (1) 2021
Natural language and speech › Language models and text generation › multilingual language models
cross-lingual generalization
0.412020
XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation · ICML 2020
Natural language and speech › Machine translation
monolingual data augmentation
0.412020
Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation · ACL 2020
Natural language and speech › Speech recognition and synthesis
spoken language understanding
0.412019
Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents · AAAI 2019
Machine learning › Transfer learning and domain adaptation › knowledge transfer
unsupervised transfer learning
0.412019
Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents · AAAI 2019
Machine learning › Efficient and distributed learning
active learning
0.312018
Deep Bayesian Active Learning for Natural Language Processing: Results of a Large-Scale Empirical Study · EMNLP 2018
Machine learning › Efficient and distributed learning › active learning › deep active learning
deep bayesian active learning
0.312018
Deep Bayesian Active Learning for Natural Language Processing: Results of a Large-Scale Empirical Study · EMNLP 2018
Machine learning › Trustworthy machine learning
fairness
0.212023
Dialect-robust Evaluation of Generated Text · ACL (1) 2023
Natural language and speech › Language models and text generation › text summarization
summarization datasets
0.212023
SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation · EMNLP 2023
Machine learning › Representation and self-supervised learning › word representation
contextualized word representation
0.112019
Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents · AAAI 2019
Machine learning › Trustworthy machine learning
uncertainty estimation
0.112018
Deep Bayesian Active Learning for Natural Language Processing: Results of a Large-Scale Empirical Study · EMNLP 2018

Methods — techniques the papers use, named apart from their topics

evaluation metrics · 0.7dialect perturbation · 0.7zero-shot evaluation · 0.5few-shot evaluation · 0.5benchmarking · 0.5zero-shot transfer · 0.4self-supervision · 0.4encoder representation · 0.4back-translation · 0.4ELMo · 0.4
YearPublicationVenuePosition
2023 Dialect-robust Evaluation of Generated Text
abstract
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, Sebastian Gehrmann
ACL (1)7
2023 SEAHORSE: A Multilingual, Multifaceted Dataset for Summarization Evaluation
abstract
Elizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das, Ankur Parikh. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Elizabeth Clark, Shruti Rijhwani, Sebastian Gehrmann, Joshua Maynez, Roee Aharoni, Vitaly Nikolaev, Thibault Sellam, Aditya Siddhant, Dipanjan Das 0001, Ankur P. Parikh
EMNLP8
2021 XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation
abstract
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, Melvin Johnson. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Sebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu 0003, Junjie Hu 0001, Dan Garrette, Graham Neubig, Melvin Johnson
EMNLP (1)4
2021 Harnessing Multilinguality in Unsupervised Machine Translation for Rare Languages
abstract
Xavier Garcia, Aditya Siddhant, Orhan Firat, Ankur Parikh. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xavier Garcia, Aditya Siddhant, Orhan Firat, Ankur P. Parikh
NAACL-HLT2
2021 Explicit Alignment Objectives for Multilingual Bidirectional Encoders
abstract
Junjie Hu, Melvin Johnson, Orhan Firat, Aditya Siddhant, Graham Neubig. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Junjie Hu 0001, Melvin Johnson, Orhan Firat, Aditya Siddhant, Graham Neubig
NAACL-HLT4
2021 mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer
abstract
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, Colin Raffel
NAACL-HLT6
2020 Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine Translation
abstract
The recently proposed massively multilingual neural machine translation (NMT) system has been shown to be capable of translating over 100 languages to and from English within a single model (Aharoni, Johnson, and Firat 2019). Its improved translation performance on low resource languages hints at potential cross-lingual transfer capability for downstream tasks. In this paper, we evaluate the cross-lingual effectiveness of representations from the encoder of a massively multilingual NMT model on 5 downstream classification and sequence labeling tasks covering a diverse set of over 50 languages. We compare against a strong baseline, multilingual BERT (mBERT) (Devlin et al. 2018), in different cross-lingual transfer learning scenarios and show gains in zero-shot transfer in 4 out of these 5 tasks.
Aditya Siddhant, Melvin Johnson, Henry Tsai, Naveen Ari, Jason Riesa, Ankur Bapna, Orhan Firat, Karthik Raman 0001
AAAI1
2020 Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation
abstract
Aditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, Yonghui Wu. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Aditya Siddhant, Ankur Bapna, Yuan Cao 0007, Orhan Firat, Mia Xu Chen, Sneha Reddy Kudugunta, Naveen Arivazhagan
ACL1
2020 XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation
abstract
Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We will release the benchmark to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks.
Junjie Hu 0001, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, Melvin Johnson
ICML3
2019 Unsupervised Transfer Learning for Spoken Language Understanding in Intelligent Agents
abstract
User interaction with voice-powered agents generates large amounts of unlabeled utterances. In this paper, we explore techniques to efficiently transfer the knowledge from these unlabeled utterances to improve model performance on Spoken Language Understanding (SLU) tasks. We use Embeddings from Language Model (ELMo) to take advantage of unlabeled data by learning contextualized word representations. Additionally, we propose ELMo-Light (ELMoL), a faster and simpler unsupervised pre-training method for SLU. Our findings suggest unsupervised pre-training on a large corpora of unlabeled utterances leads to significantly better SLU performance compared to training from scratch and it can even outperform conventional supervised transfer. Additionally, we show that the gains from unsupervised transfer techniques can be further improved by supervised transfer. The improvements are more pronounced in low resource settings and when using only 1000 labeled in-domain samples, our techniques match the performance of training from scratch on 10-15x more labeled in-domain data.
Aditya Siddhant, Anuj Kumar Goyal, Angeliki Metallinou
AAAI1
2018 Deep Bayesian Active Learning for Natural Language Processing: Results of a Large-Scale Empirical Study
abstract
Several recent papers investigate Active Learning (AL) for mitigating the datadependence of deep learning for natural language processing.However, the applicability of AL to real-world problems remains an open question.While in supervised learning, practitioners can try many different methods, evaluating each against a validation set before selecting a model, AL affords no such luxury.Over the course of one AL run, an agent annotates its dataset exhausting its labeling budget.Thus, given a new task, an active learner has no opportunity to compare models and acquisition functions.This paper provides a largescale empirical study of deep active learning, addressing multiple tasks and, for each, multiple datasets, multiple models, and a full suite of acquisition functions.We find that across all settings, Bayesian active learning by disagreement, using uncertainty estimates provided either by Dropout or Bayes-by-Backprop significantly improves over i.i.d.baselines and usually outperforms classic uncertainty sampling.
Aditya Siddhant, Zachary C. Lipton
EMNLP1
2017 Leveraging native language speech for accent identification using deep Siamese networks
abstract
The problem of automatic accent identification is important for several applications like speaker profiling and recognition as well as for improving speech recognition systems. The accented nature of speech can be primarily attributed to the influence of the speaker's native language on the given speech recording. In this paper, we propose a novel accent identification system whose training exploits speech in native languages along with the accented speech. Specifically, we develop a deep Siamese network based model which learns the association between accented speech recordings and the native language speech recordings. The Siamese networks are trained with i-vector features extracted from the speech recordings using either an unsupervised Gaussian mixture model (GMM) or a supervised deep neural network (DNN) model. We perform several accent identification experiments using the CSLU Foreign Accented English (FAE) corpus. In these experiments, our proposed approach using deep Siamese networks yield significant relative performance improvements of 15.4% on a 10-class accent identification task, over a baseline DNN-based classification system that uses GMM i-vectors. Furthermore, we present a detailed error analysis of the proposed accent identification system.
Aditya Siddhant, Preethi Jyothi, Sriram Ganapathy
ASRU1