Orhan Firat

dblp:120/2225 · DBLP profile ↗
← Back
50ranked-venue papers
5as first author
32since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 4 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 BIG-Bench Extra Hard
abstract
Mehran Kazemi, Bahare Fatemi, Hritik Bansal, John Palowitch, Chrysovalantis Anastasiou, Sanket Vaibhav Mehta, Lalit K Jain, Virginia Aglietti, Disha Jindal, Peter Chen, Nishanth Dikkala, Gladys Tyen, Xin Liu, Uri Shalit, Silvia Chiappa, Kate Olszewska, Yi Tay, Vinh Q. Tran, Quoc V Le, Orhan Firat. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Mehran Kazemi, Bahare Fatemi, Hritik Bansal, John Palowitch, Chrysovalantis Anastasiou, Sanket Vaibhav Mehta, Lalit K. Jain, Virginia Aglietti, Disha Jindal, Peter Chen, Nishanth Dikkala, Gladys Tyen, Xin Liu 0034, Uri Shalit, Silvia Chiappa, Kate Olszewska, Yi Tay, Vinh Q. Tran 0002, Quoc V. Le, Orhan Firat
ACL (1)20
2024 When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method
abstract
While large language models (LLMs) often adopt finetuning to unlock their capabilities for downstream applications, our understanding on the inductive biases (especially the scaling properties) of different finetuning methods is still limited. To fill this gap, we conduct systematic experiments studying whether and how different scaling factors, including LLM model size, pretraining data size, new finetuning parameter size and finetuning data size, affect the finetuning performance. We consider two types of finetuning – full-model tuning (FMT) and parameter efficient tuning (PET, including prompt tuning and LoRA), and explore their scaling behaviors in the data-limited regime where the LLM model size substantially outweighs the finetuning data size. Based on two sets of pretrained bilingual LLMs from 1B to 16B and experiments on bilingual machine translation and multilingual summarization benchmarks, we find that 1) LLM finetuning follows a powerbased multiplicative joint scaling law between finetuning data size and each other scaling factor; 2) LLM finetuning benefits more from LLM model scaling than pretraining data scaling, and PET parameter scaling is generally ineffective; and 3) the optimal finetuning method is highly task- and finetuning data-dependent. We hope our findings could shed light on understanding, selecting and developing LLM finetuning methods.
Biao Zhang 0006, Zhongtao Liu, Colin Cherry, Orhan Firat
ICLR4
2024 Scaling Sign Language Translation
abstract
Sign language translation (SLT) addresses the problem of translating information from a sign language in video to a spoken language in text. Existing studies, while showing progress, are often limited to narrow domains and/or few sign languages and struggle with open-domain tasks. In this paper, we push forward the frontier of SLT by scaling pretraining data, model size, and number of translation directions. We perform large-scale SLT pretraining on different data including 1) noisy multilingual Youtube SLT data, 2) parallel text corpora, and 3) SLT data augmented by translating video captions to other languages with off-the-shelf machine translation models. We unify different pretraining tasks with task-specific prompts under the encoder-decoder architecture, and initialize the SLT model with pretrained (m/By)T5 models across model sizes. SLT pretraining results on How2Sign and FLEURS-ASL\#0 (ASL to 42 spoken languages) demonstrate the significance of data/model scaling and cross-lingual cross-modal transfer, as well as the feasibility of zero-shot SLT. We finetune the pretrained SLT models on 5 downstream open-domain SLT benchmarks covering 5 sign languages. Experiments show substantial quality improvements over the vanilla baselines, surpassing the previous state-of-the-art (SOTA) by wide margins.
Biao Zhang 0006, Garrett Tanzer, Orhan Firat
NeurIPS3
2023 GATITOS: Using a New Multilingual Lexicon for Low-resource Machine Translation
abstract
Modern machine translation models and language models are able to translate without having been trained on parallel data, greatly expanding the set of languages that they can serve.However, these models still struggle in a variety of predictable ways, a problem that cannot be overcome without at least some trusted bilingual data.This work expands on a cheap and abundant resource to combat this problem: bilingual lexica (BILEXs).We test the efficacy of bilingual lexica in a real-world setup, on 200-language translation models trained on web-crawled text.We present several findings: (1) using lexical data augmentation, we demonstrate sizable performance gains for unsupervised translation; (2) we compare several families of data augmentation, demonstrating that they yield similar improvements, and can be combined for even greater improvements;(3) we demonstrate the importance of carefully curated lexica over larger, noisier ones, especially with larger models; and (4) we compare the efficacy of multilingual lexicon data versus human-translated parallel data.Based on results from (3), we develop and open-source GATITOS, a high-quality, curated dataset covering 170 mostly low-resource languages at the time of this submission, one of the first humantranslated resources to support many of these languages 1 .
Isaac Caswell, Orhan Firat, Ishank Saxena
EMNLP3
2023 UniMax: Fairer and More Effective Language Sampling for Large-Scale Multilingual Pretraining
Hyung Won Chung, Xavier Garcia, Adam Roberts, Yi Tay, Orhan Firat, Sharan Narang, Noah Constant
ICLR5
2023 Scaling Laws for Multilingual Neural Machine Translation
abstract
In this work, we provide a large-scale empirical study of the scaling properties of multilingual neural machine translation models. We examine how increases in the model size affect the model performance and investigate the role of the individual language pair weights on the scaling behavior. We find that these weights only affect the multiplicative factor of the scaling law, and in particular, the scaling exponent is unaffected by them. Through a novel joint scaling law formulation, we compute the effective number of parameters allocated to each language pair and examine the role of language similarity in the scaling behavior of our models. We find little evidence that language similarity has any impact. In contrast, ``direction'' of the multilinguality plays a significant role, with models translating from multiple languages into English having a larger number of effective parameters per task than their reversed counterparts. Finally, we leverage our observations to predict the performance of multilingual models trained with any language weighting at any scale, greatly reducing efforts required for language balancing in large multilingual models. Our findings apply to both in-domain and out-of-domain test sets and to multiple evaluation metrics, such as ChrF and BLEURT.
Patrick Fernandes, Behrooz Ghorbani, Xavier Garcia, Markus Freitag, Orhan Firat
ICML5
2023 The Unreasonable Effectiveness of Few-shot Learning for Machine Translation
abstract
We demonstrate the potential of few-shot translation systems, trained with unpaired language data, for both high and low-resource language pairs. We show that with only 5 examples of high-quality translation data shown at inference, a transformer decoder-only model trained solely with self-supervised learning, is able to match specialized supervised state-of-the-art models as well as more general commercial translation systems. In particular, we outperform the best performing system on the WMT'21 English-Chinese news translation task by only using five examples of English-Chinese parallel data at inference. Furthermore, the resulting models are two orders of magnitude smaller than state-of-the-art language models. We then analyze the factors which impact the performance of few-shot translation systems, and highlight that the quality of the few-shot demonstrations heavily determines the quality of the translations generated by our models. Finally, we show that the few-shot paradigm also provides a way to control certain attributes of the translation --- we show that we are able to control for regional varieties and formality using only a five examples at inference, paving the way towards controllable machine translation systems.
Xavier Garcia, Yamini Bansal, Colin Cherry, George F. Foster, Maxim Krikun, Melvin Johnson, Orhan Firat
ICML7
2023 Interactive-Chain-Prompting: Ambiguity Resolution for Crosslingual Conditional Generation with Interaction
abstract
Jonathan Pilault, Xavier Garcia, Arthur Bražinskas, Orhan Firat. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jonathan Pilault, Xavier Garcia, Arthur Brazinskas, Orhan Firat
IJCNLP (1)4
2023 Binarized Neural Machine Translation
abstract
The rapid scaling of language models is motivating research using low-bitwidth quantization. In this work, we propose a novel binarization technique for Transformers applied to machine translation (BMT), the first of its kind. We identify and address the problem of inflated dot-product variance when using one-bit weights and activations. Specifically, BMT leverages additional LayerNorms and residual connections to improve binarization quality. Experiments on the WMT dataset show that a one-bit weight-only Transformer can achieve the same quality as a float one, while being 16$\times$ smaller in size. One-bit activations incur varying degrees of quality drop, but mitigated by the proposed architectural changes. We further conduct a scaling law study using production-scale translation datasets, which shows that one-bit weight Transformers scale and generalize well in both in-domain and out-of-domain settings. Implementation in JAX/Flax will be open sourced.
Yichi Zhang 0006, Ankush Garg, Lukasz Lew, Behrooz Ghorbani, Zhiru Zhang, Orhan Firat
NeurIPS7
2023 Order Matters in the Presence of Dataset Imbalance for Multilingual Learning
abstract
In this paper, we empirically study the optimization dynamics of multi-task learning, particularly focusing on those that govern a collection of tasks with significant data imbalance. We present a simple yet effective method of pre-training on high-resource tasks, followed by fine-tuning on a mixture of high/low-resource tasks. We provide a thorough empirical study and analysis of this method's benefits showing that it achieves consistent improvements relative to the performance trade-off profile of standard static weighting. We analyze under what data regimes this method is applicable and show its improvements empirically in neural machine translation (NMT) and multi-lingual language modeling.
Dami Choi, Derrick Xin, Hamid Dadkhahi, Justin Gilmer, Ankush Garg, Orhan Firat, Chih-Kuan Yeh, Andrew M. Dai, Behrooz Ghorbani
NeurIPS6
2023 MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
abstract
We introduce MADLAD-400, a manually audited, general domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. We discuss the limitations revealed by self-auditing MADLAD-400, and the role data auditing had in the dataset creation process. We then train and release a 10.7B-parameter multilingual machine translation model on 250 billion tokens covering over 450 languages using publicly available data, and find that it is competitive with models that are significantly larger, and report the results on different domains. In addition, we train a 8B-parameter language model, and assess the results on few-shot translation. We make the baseline models available to the research community.
Sneha Reddy Kudugunta, Isaac Caswell, Biao Zhang 0006, Xavier Garcia, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, Orhan Firat
NeurIPS9
2023 Block-State Transformers
abstract
State space models (SSMs) have shown impressive results on tasks that require modeling long-range dependencies and efficiently scale to long sequences owing to their subquadratic runtime complexity. Originally designed for continuous signals, SSMs have shown superior performance on a plethora of tasks, in vision and audio; however, SSMs still lag Transformer performance in Language Modeling tasks. In this work, we propose a hybrid layer named Block-State Transformer (*BST*), that internally combines an SSM sublayer for long-range contextualization, and a Block Transformer sublayer for short-term representation of sequences. We study three different, and completely *parallelizable*, variants that integrate SSMs and block-wise attention. We show that our model outperforms similar Transformer-based architectures on language modeling perplexity and generalizes to longer sequences. In addition, the Block-State Transformer demonstrates a more than *tenfold* increase in speed at the layer level compared to the Block-Recurrent Transformer when model parallelization is employed.
Jonathan Pilault, Mahan Fathi, Orhan Firat, Christopher Joseph Pal, Pierre-Luc Bacon, Ross Goroshin
NeurIPS3
2023 PaLM: Scaling Language Modeling with Pathways
abstract
Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM). We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of-the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state-of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Adam Roberts, Paul Barham 0001, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du 0002, Ben Hutchinson, Reiner Pope, Jacob Austin, Michael Isard, Guy Gur-Ari, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, William Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang 0002, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeffrey Dean, Slav Petrov, Noah Fiedel
J. Mach. Learn. Res.60
2023 FRMT: A Benchmark for Few-Shot Region-Aware Machine Translation
abstract
Abstract We present FRMT, a new dataset and evaluation benchmark for Few-shot Region-aware Machine Translation, a type of style-targeted translation. The dataset consists of professional translations from English into two regional variants each of Portuguese and Mandarin Chinese. Source documents are selected to enable detailed analysis of phenomena of interest, including lexically distinct terms and distractor terms. We explore automatic evaluation metrics for FRMT and validate their correlation with expert human evaluation across both region-matched and mismatched rating scenarios. Finally, we present a number of baseline models for this task, and offer guidelines for how researchers can train, evaluate, and compare their own models. Our dataset and evaluation code are publicly available: https://bit.ly/frmt-task.
Parker Riley, Timothy Dozat, Jan A. Botha, Xavier Garcia, Dan Garrette, Jason Riesa, Orhan Firat, Noah Constant
Trans. Assoc. Comput. Linguistics7
2022 Multilingual Mix: Example Interpolation Improves Multilingual Neural Machine Translation
abstract
Multilingual neural machine translation models are trained to maximize the likelihood of a mix of examples drawn from multiple language pairs.The dominant inductive bias applied to these models is a shared vocabulary and a shared set of parameters across languages; the inputs and labels corresponding to examples drawn from different language pairs might still reside in distinct subspaces.In this paper, we introduce multilingual crossover encoder-decoder (mXEncDec) to fuse language pairs at an instance level.Our approach interpolates instances from different language pairs into joint 'crossover examples' in order to encourage sharing input and output spaces across languages.To ensure better fusion of examples in multilingual settings, we propose several techniques to improve example interpolation across dissimilar languages under heavy data imbalance.Experiments on a large-scale WMT multilingual dataset demonstrate that our approach significantly improves quality on English-to-Many, Many-to-English and zero-shot translation tasks (from +0.5 BLEU up to +5.5 BLEU points).Results on code-switching sets demonstrate the capability of our approach to improve model generalization to out-of-distribution multilingual examples.We also conduct qualitative and quantitative representation comparisons to analyze the advantages of our approach at the representation level.
Yong Cheng 0003, Ankur Bapna, Orhan Firat, Yuan Cao 0007, Pidong Wang, Wolfgang Macherey
ACL (1)3
2022 Multilingual Document-Level Translation Enables Zero-Shot Transfer From Sentences to Documents
abstract
Biao Zhang, Ankur Bapna, Melvin Johnson, Ali Dabirmoghaddam, Naveen Arivazhagan, Orhan Firat. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Biao Zhang 0006, Ankur Bapna, Melvin Johnson, Ali Dabirmoghaddam, Naveen Arivazhagan, Orhan Firat
ACL (1)6
2022 Scaling Laws for Neural Machine Translation
Behrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna, Maxim Krikun, Xavier Garcia, Ciprian Chelba, Colin Cherry
ICLR2
2022 A Loss Curvature Perspective on Training Instabilities of Deep Learning Models
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Reddy Kudugunta, Behnam Neyshabur, David Cardoze, George E. Dahl, Zachary Nado, Orhan Firat
ICLR9
2022 Data Scaling Laws in NMT: The Effect of Noise and Architecture
abstract
In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that the test loss of encoder-decoder transformer models scales as a power law in the number of training samples, with a dependence on the model size. Then, we systematically vary aspects of the training setup to understand how they impact the data scaling laws. In particular, we change the following (1) Architecture and task setup: We compare to a transformer-LSTM hybrid, and a decoder-only transformer with a language modeling loss (2) Noise level in the training distribution: We experiment with filtering, and adding iid synthetic noise. In all the above cases, we find that the data scaling exponents are minimally impacted, suggesting that marginally worse architectures or training data can be compensated for by adding more data. Lastly, we find that using back-translated data instead of parallel data, can significantly degrade the scaling exponent.
Yamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang 0006, Colin Cherry, Behnam Neyshabur, Orhan Firat
ICML7
2022 GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
abstract
Scaling language models with more data, compute and parameters has driven significant progress in natural language processing. For example, thanks to scaling, GPT-3 was able to achieve strong results on in-context learning tasks. However, training these large dense models requires significant amounts of computing resources. In this paper, we propose and develop a family of language models named \glam (\textbf{G}eneralist \textbf{La}nguage \textbf{M}odel), which uses a sparsely activated mixture-of-experts architecture to scale the model capacity while also incurring substantially less training cost compared to dense variants. The largest \glam has 1.2 trillion parameters, which is approximately 7x larger than GPT-3. It consumes only 1/3 of the energy used to train GPT-3 and requires half of the computation flops for inference, while still achieving better overall fewshot performance across 29 NLP tasks.
Nan Du 0002, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, William Fedus, Maarten Bosma, Zongwei Zhou, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathy Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang 0043, Quoc V. Le, Claire Cui
ICML10
2022 Examining Scaling and Transfer of Language Model Architectures for Machine Translation
abstract
Natural language understanding and generation models follow one of the two dominant architectural paradigms: language models (LMs) that process concatenated sequences in a single stack of layers, and encoder-decoder models (EncDec) that utilize separate layer stacks for input and output processing. In machine translation, EncDec has long been the favoured approach, but with few studies investigating the performance of LMs. In this work, we thoroughly examine the role of several architectural design choices on the performance of LMs on bilingual, (massively) multilingual and zero-shot translation tasks, under systematic variations of data conditions and model sizes. Our results show that: (i) Different LMs have different scaling properties, where architectural differences often have a significant impact on model performance at small scales, but the performance gap narrows as the number of parameters increases, (ii) Several design choices, including causal masking and language-modeling objectives for the source sequence, have detrimental effects on translation quality, and (iii) When paired with full-visible masking for source sequences, LMs could perform on par with EncDec on supervised bilingual and multilingual translation tasks, and improve greatly on zero-shot directions by facilitating the reduction of off-target translations.
Biao Zhang 0006, Behrooz Ghorbani, Ankur Bapna, Yong Cheng 0003, Xavier Garcia, Jonathan Shen, Orhan Firat
ICML7
2022 XTREME-S: Evaluating Cross-lingual Speech Representations
abstract
We introduce XTREME-S, a new benchmark to evaluate universal cross-lingual speech representations in many languages.XTREME-S covers four task families: speech recognition, classification, speech-to-text translation and retrieval.Covering 102 languages from 10+ language families, 3 different domains and 4 task families, XTREME-S aims to simplify multilingual speech representation evaluation, as well as catalyze research in "universal" speech representation learning.This paper describes the new benchmark and establishes the first speech-only and speechtext baselines using XLS-R and mSLAM on all downstream tasks.We motivate the design choices and detail how to use the benchmark.Datasets and fine-tuning scripts are made easily accessible through the HuggingFace platform. 1
Alexis Conneau, Ankur Bapna, Yu Zhang 0033, Patrick von Platen, Anton Lozhkov, Colin Cherry, Ye Jia, Clara Rivera, Mihir Kale, Daan van Esch, Vera Axelrod, Simran Khanuja, Jonathan H. Clark, Orhan Firat, Michael Auli, Sebastian Ruder, Jason Riesa, Melvin Johnson
INTERSPEECH15
2022 Do Current Multi-Task Optimization Methods in Deep Learning Even Help?
abstract
Recent research has proposed a series of specialized optimization algorithms for deep multi-task models. It is often claimed that these multi-task optimization (MTO) methods yield solutions that are superior to the ones found by simply optimizing a weighted average of the task losses. In this paper, we perform large-scale experiments on a variety of language and vision tasks to examine the empirical validity of these claims. We show that, despite the added design and computational complexity of these algorithms, MTO methods do not yield any performance improvements beyond what is achievable via traditional optimization approaches. We highlight alternative strategies that consistently yield improvements to the performance profile and point out common training pitfalls that might cause suboptimal results. Finally, we outline challenges in reliably evaluating the performance of MTO algorithms and discuss potential solutions.
Derrick Xin, Behrooz Ghorbani, Justin Gilmer, Ankush Garg, Orhan Firat
NeurIPS5
2022 Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
abstract
Abstract With the success of large-scale pre-training and multilingual modeling in Natural Language Processing (NLP), recent years have seen a proliferation of large, Web-mined text datasets covering hundreds of languages. We manually audit the quality of 205 language-specific corpora released with five major public datasets (CCAligned, ParaCrawl, WikiMatrix, OSCAR, mC4). Lower-resource corpora have systematic issues: At least 15 corpora have no usable text, and a significant fraction contains less than 50% sentences of acceptable quality. In addition, many are mislabeled or use nonstandard/ambiguous language codes. We demonstrate that these issues are easy to detect even for non-proficient speakers, and supplement the human audit with automatic analyses. Finally, we recommend techniques to evaluate and improve multilingual corpora and discuss potential risks that come with low-quality data releases.
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov 0001, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Rubungo Andre Niyongabo, Toan Q. Nguyen, Mathias Müller 0002, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Reddy Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Balli, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofe Adeyemi
Trans. Assoc. Comput. Linguistics36
2021 A Large-Scale Study of Machine Translation in Turkic Languages
abstract
Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta Jr., Bekhzodbek Moydinboyev, Esra Onal, Shaxnoza Pulatova, Ahsan Wahab, Orhan Firat, Sriram Chellappan. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Jamshidbek Mirzakhalov, Anoop Babu, Duygu Ataman, Sherzod Kariev, Francis M. Tyers, Otabek Abduraufov, Mammad Hajili, Sardana Ivanova, Abror Khaytbaev, Antonio Laverghetta, Behzodbek Moydinboyev, Esra Onal, Shaxnoza Pulatova, Ahsan Wahab, Orhan Firat, Sriram Chellappan
EMNLP (1)15
2021 XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation
abstract
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, Melvin Johnson. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Sebastian Ruder, Noah Constant, Jan A. Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu 0003, Junjie Hu 0001, Dan Garrette, Graham Neubig, Melvin Johnson
EMNLP (1)5
2021 GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer
ICLR5
2021 Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual Models
Yulia Tsvetkov, Orhan Firat, Yuan Cao 0007
ICLR3
2021 Share or Not? Learning to Schedule Language-Specific Capacity for Multilingual Translation
Biao Zhang 0006, Ankur Bapna, Rico Sennrich, Orhan Firat
ICLR4
2021 Towards Continual Learning for Multilingual Machine Translation via Vocabulary Substitution
abstract
Xavier Garcia, Noah Constant, Ankur Parikh, Orhan Firat. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xavier Garcia, Noah Constant, Ankur P. Parikh, Orhan Firat
NAACL-HLT4
2021 Harnessing Multilinguality in Unsupervised Machine Translation for Rare Languages
abstract
Xavier Garcia, Aditya Siddhant, Orhan Firat, Ankur Parikh. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Xavier Garcia, Aditya Siddhant, Orhan Firat, Ankur P. Parikh
NAACL-HLT3
2021 Explicit Alignment Objectives for Multilingual Bidirectional Encoders
abstract
Junjie Hu, Melvin Johnson, Orhan Firat, Aditya Siddhant, Graham Neubig. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Junjie Hu 0001, Melvin Johnson, Orhan Firat, Aditya Siddhant, Graham Neubig
NAACL-HLT3
2020 Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine Translation
abstract
The recently proposed massively multilingual neural machine translation (NMT) system has been shown to be capable of translating over 100 languages to and from English within a single model (Aharoni, Johnson, and Firat 2019). Its improved translation performance on low resource languages hints at potential cross-lingual transfer capability for downstream tasks. In this paper, we evaluate the cross-lingual effectiveness of representations from the encoder of a massively multilingual NMT model on 5 downstream classification and sequence labeling tasks covering a diverse set of over 50 languages. We compare against a strong baseline, multilingual BERT (mBERT) (Devlin et al. 2018), in different cross-lingual transfer learning scenarios and show gains in zero-shot transfer in 4 out of these 5 tasks.
Aditya Siddhant, Melvin Johnson, Henry Tsai, Naveen Ari, Jason Riesa, Ankur Bapna, Orhan Firat, Karthik Raman 0001
AAAI7
2020 Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation
abstract
Aditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, Yonghui Wu. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Aditya Siddhant, Ankur Bapna, Yuan Cao 0007, Orhan Firat, Mia Xu Chen, Sneha Reddy Kudugunta, Naveen Arivazhagan
ACL4
2020 XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalisation
abstract
Much recent progress in applications of machine learning models to NLP has been driven by benchmarks that evaluate models across a wide variety of tasks. However, these broad-coverage benchmarks have been mostly limited to English, and despite an increasing interest in multilingual models, a benchmark that enables the comprehensive evaluation of such methods on a diverse range of languages and tasks is still missing. To this end, we introduce the Cross-lingual TRansfer Evaluation of Multilingual Encoders (XTREME) benchmark, a multi-task benchmark for evaluating the cross-lingual generalization capabilities of multilingual representations across 40 languages and 9 tasks. We demonstrate that while models tested on English reach human performance on many tasks, there is still a sizable gap in the performance of cross-lingually transferred models, particularly on syntactic and sentence retrieval tasks. There is also a wide spread of results across languages. We will release the benchmark to encourage research on cross-lingual learning methods that transfer linguistic knowledge across a diverse and representative set of languages and tasks.
Junjie Hu 0001, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, Melvin Johnson
ICML5
2019 Simple, Scalable Adaptation for Neural Machine Translation
abstract
Ankur Bapna, Orhan Firat. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ankur Bapna, Orhan Firat
EMNLP/IJCNLP (1)2
2019 Investigating Multilingual NMT Representations at Scale
abstract
Sneha Kudugunta, Ankur Bapna, Isaac Caswell, Orhan Firat. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Sneha Reddy Kudugunta, Ankur Bapna, Isaac Caswell, Orhan Firat
EMNLP/IJCNLP (1)4
2019 GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
abstract
Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or infrastructure. These solutions are often architecture-specific and do not transfer to other machine learning tasks. To address the need for efficient and task-independent model parallelism, we introduce TensorPipe, a pipeline parallelism library that allows scaling any network that can be expressed as a sequence of layers. By pipelining different sub-sequences of layers on separate accelerators, TensorPipe provides the flexibility of scaling a variety of different networks to gigantic sizes efficiently. Moreover, TensorPipe utilizes a novel batch-splitting pipelining algorithm, resulting in almost linear speedup when a model is partitioned across multiple accelerators. We demonstrate the advantages of TensorPipe by training large-scale neural networks on two different tasks with distinct network architectures: (i)Image Classification: We train a 557-million-parameter AmoebaNet model and attain a top-1 accuracy of 84.4% on ImageNet-2012, (ii)Multilingual Neural Machine Translation: We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models.
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le
NeurIPS4
2018 The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation
abstract
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, Macduff Hughes. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George F. Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Macduff Hughes
ACL (1)2
2018 Training Deeper Neural Machine Translation Models with Transparent Attention
abstract
While current state-of-the-art NMT models, such as RNN seq2seq and Transformers, possess a large number of parameters, they are still shallow in comparison to convolutional models used for both text and vision applications.In this work we attempt to train significantly (2-3x) deeper Transformer and Bi-RNN encoders for machine translation.We propose a simple modification to the attention mechanism that eases the optimization of deeper models, and results in consistent gains of 0.7-1.1 BLEU on the benchmark WMT'14 English-German and WMT'15 Czech-English tasks for both architectures.
Ankur Bapna, Mia Xu Chen, Orhan Firat, Yuan Cao 0007
EMNLP3
2018 Revisiting Character-Based Neural Machine Translation with Capacity and Compression
abstract
Translating characters instead of words or word-fragments has the potential to simplify the processing pipeline for neural machine translation (NMT), and improve results by eliminating hyper-parameters and manual feature engineering.However, it results in longer sequences in which each symbol contains less information, creating both modeling and computational challenges.In this paper, we show that the modeling problem can be solved by standard sequence-to-sequence architectures of sufficient depth, and that deep models operating at the character level outperform identical models operating over word fragments.This result implies that alternative architectures for handling character input are better viewed as methods for reducing computation time than as improved ways of modeling longer sequences.From this perspective, we evaluate several techniques for characterlevel NMT, verify that they do not match the performance of our deep character baseline model, and evaluate the performance versus computation time tradeoffs they offer.Within this framework, we also perform the first evaluation for NMT of conditional computation over time, in which the model learns which timesteps can be skipped, rather than having them be dictated by a fixed schedule specified before training begins.
Colin Cherry, George F. Foster, Ankur Bapna, Orhan Firat, Wolfgang Macherey
EMNLP4
2017 Multi-way, multilingual neural machine translation
Orhan Firat, Kyunghyun Cho, Baskaran Sankaran, Fatos T. Yarman-Vural, Yoshua Bengio
Comput. Speech Lang.1
2017 On integrating a language model into neural machine translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Yoshua Bengio
Comput. Speech Lang.2
2016 Zero-Resource Translation with Multi-Lingual Neural Machine Translation
abstract
In this paper, we propose a novel finetuning algorithm for the recently introduced multiway, multilingual neural machine translate that enables zero-resource machine translation.When used together with novel manyto-one translation strategies, we empirically show that this finetuning algorithm allows the multi-way, multilingual model to translate a zero-resource language pair (1) as well as a single-pair neural translation model trained with up to 1M direct parallel sentences of the same language pair and (2) better than pivotbased translation strategy, while keeping only one additional copy of attention-related parameters.
Orhan Firat, Baskaran Sankaran, Yaser Al-Onaizan, Fatos T. Yarman-Vural, Kyunghyun Cho
EMNLP1
2016 Multi-Way, Multilingual Neural Machine Translation with a Shared Attention Mechanism
abstract
We propose multi-way, multilingual neural machine translation.The proposed approach enables a single neural translation model to translate between multiple languages, with a number of parameters that grows only linearly with the number of languages.This is made possible by having a single attention mechanism that is shared across all language pairs.We train the proposed multiway, multilingual model on ten language pairs from WMT'15 simultaneously and observe clear performance improvements over models trained on only one language pair.In particular, we observe that the proposed model significantly improves the translation quality of low-resource language pairs.
Orhan Firat, Kyunghyun Cho, Yoshua Bengio
HLT-NAACL1
2014 Deep learning for brain decoding
abstract
Learning low dimensional embedding spaces (manifolds) for efficient feature representation is crucial for complex and high dimensional input spaces. Functional magnetic resonance imaging (fMRI) produces high dimensional input data and with a less then ideal number of labeled samples for a classification task. In this study, we explore deep learning methods for fMRI classification tasks in order to reduce dimensions of feature space, along with improving classification performance for brain decoding. We employ sparse autoencoders for unsupervised feature learning, leveraging unlabeled fMRI data to learn efficient, non-linear representations as the building blocks of a deep learning architecture by stacking them. Proposed method is tested on a memory encoding/retrieval experiment with ten classes. The results support the efficiency compared to the baseline multi-voxel pattern analysis techniques.
Orhan Firat, Like Oztekin, Fatos T. Yarman-Vural
ICIP1
2014 Representation Learning for Contextual Object and Region Detection in Remote Sensing
abstract
The performance of object recognition and classification on remote sensing imagery is highly dependent on the quality of extracted features, amount of labelled data and the priors defined for contextual models. In this study, we examine the representation learning opportunities for remote sensing. First we attacked localization of contextual cues for complex object detection using disentangling factors learnt from a small amount of labelled data. The complex object, which consists of several sub-parts is further represented under the Conditional Markov Random Fields framework. As a second task, end-to-end target detection using convolutional sparse auto-encoders (CSA) using large amount of unlabelled data is analysed. Proposed methodologies are tested on complex airfield detection problem using Conditional Random Fields and recognition of dispersal areas, park areas, taxi routes, airplanes using CSA. The method is also tested on the detection of the dry docks in harbours. Performance of the proposed method is compared with standard feature engineering methods and found competitive with currently used rule-based and supervised methods.
Orhan Firat, Gulcan Can, Fatos T. Yarman-Vural
ICPR1
2014 Modeling the Brain Connectivity for Pattern Analysis
abstract
An information theoretic approach is proposed to estimate the degree of connectivity for each voxel with its neighboring voxels. The neighborhood system is defined by spatial and functional connectivity metrics. Then, a local mesh of variable size is formed around each voxel using spatial or functional neighborhood. The mesh arc weights, called Mesh Arc Descriptors (MAD), are estimated by a linear regression model fitted to the voxel intensity values of the functional Magnetic Resonance Images (fMRI). Finally, the error term of the linear regression equation is used to estimate the mesh size for a voxel by optimizing Akaike's information Criterion, Bayesian Information Criterion and Rissanen's Minimum Description Length. fMRI measurements are obtained during a memory encoding and retrieval experiment performed on a subject who is exposed to the stimuli from 10 semantic categories. For each sample, a k-NN classifier is trained using the Mesh Arc Descriptors (MAD) having the variable mesh sizes. The classification performances reflect that the suggested variable-size Mesh Arc Descriptors represents the mental states better than the classical multi-voxel pattern representation. Moreover, we observe that the degree of connectivities in the brain greatly varies for each voxel.
Itir Önal, Emre Aksan, Burak Velioglu, Orhan Firat, Mete Ozay, Ilke Öztekin, Fatos T. Yarman-Vural
ICPR4
2013 An information theoretic approach to classify cognitive states using fMRI
abstract
In this study, an information theoretic approach is proposed to model brain connectivity during a cognitive processing task, measured by functional Magnetic Resonance Imaging (fMRI). For this purpose, a local mesh of varying size is formed around each voxel. The arc weights of each mesh are estimated using a linear regression model by minimizing the squared error. Then, the optimal mesh size for each sample, that represents the information distribution in the brain, is estimated by minimizing various information criteria which employ the mean square error of linear regression model. The estimated mesh size shows the degree of locality or degree of connectivity of the voxels for the underlying cognitive process. The samples are generated during an fMRI experiment employing item recognition (IR) and judgment of recency (JOR) tasks. For each sample, estimated arc weights of the local mesh with optimal size are used to classify whether it belongs to IR or JOR tasks. Results indicate that the suggested connectivity model with optimal mesh size for each sample represent the information distribution in the brain better than the state-of-the art methods.
Itir Önal, Mete Ozay, Orhan Firat, Ilke Öztekin, Fatos T. Yarman-Vural
BIBE3
2013 Mesh learning for object classification using fMRI measurements
abstract
Machine learning algorithms have been widely used as reliable methods for modeling and classifying cognitive processes using functional Magnetic Resonance Imaging (fMRI) data. In this study, we aim to classify fMRI measurements recorded during an object recognition experiment. Previous studies focus on Multi Voxel Pattern Analysis (MVPA) which feeds a set of active voxels in a concatenated vector form to a machine learning algorithm to train and classify the cognitive processes. In most of the MVPA methods, after an image preprocessing step, the voxel intensity values are fed to a classifier to train and recognize the underlying cognitive process. Sometimes, the fMRI data is further processed for de-noising or feature selection where techniques, such as Generalized Linear Model (GLM), Independent Component Analysis (ICA) or Principal Component Analysis are employed. Although these techniques are proved to be useful in MVPA, they do not model the spatial connectivity among the voxels. In this study, we attempt to represent the local relations among the voxel intensity values by forming a mesh network around each voxel to model the relationship of a voxel and its surroundings. The degree of connectivity of a voxel to its surroundings is represented by the arc weights of each mesh. The arc weights, which are estimated by a linear regression model, are fed to a classifier to discriminate the brain states during an object recognition task. This approach, called Mesh Learning, provides a powerful tool to analyze various cognitive states using fMRI data. Compared to traditional studies which focus either merely on multi-voxel pattern vectors or their reduced-dimension versions, the suggested Mesh Learning provides a better representation of object recognition task. Various machine learning algorithms are tested to compare the suggested Mesh Learning to the state-of-the art MVPA techniques. The performance of the Mesh Learning is shown to be higher than that of the available MVPA techniques.
Omer Ekmekci, Orhan Firat, Mete Ozay, Ilke Öztekin, Fatos T. Yarman-Vural, Uygar Öztekin
ICIP2