Derrick Xin

dblp:329/6502 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 28% Learning paradigms · 27% Transfer learning and domain adaptation · 19%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
multilingual language models
1.322023
MADLAD-400: A Multilingual And Document-Level Large Audited Dataset · NeurIPS 2023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Machine learning › Learning paradigms
multi-task learning
1.222023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Do Current Multi-Task Optimization Methods in Deep Learning Even Help? · NeurIPS 2022
Machine learning › Learning paradigms
imbalanced learning
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Natural language and speech › Language models and text generation
multilingual learning
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.712023
MADLAD-400: A Multilingual And Document-Level Large Audited Dataset · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › cross-lingual transfer
multilingual transfer
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Machine learning › Transfer learning and domain adaptation › pre-training and adaptation
pre-training and fine-tuning
0.712023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023
Machine learning › Optimization for machine learning
multi-task optimization
0.612022
Do Current Multi-Task Optimization Methods in Deep Learning Even Help? · NeurIPS 2022
Natural language and speech › Machine translation
neural machine translation
0.212023
Order Matters in the Presence of Dataset Imbalance for Multilingual Learning · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

task weighting · 0.7pre-training · 0.7large-scale pretraining · 0.7fine-tuning · 0.7few-shot learning · 0.7weighted average of task losses · 0.6large-scale empirical study · 0.6
YearPublicationVenuePosition
2023 Order Matters in the Presence of Dataset Imbalance for Multilingual Learning
abstract
In this paper, we empirically study the optimization dynamics of multi-task learning, particularly focusing on those that govern a collection of tasks with significant data imbalance. We present a simple yet effective method of pre-training on high-resource tasks, followed by fine-tuning on a mixture of high/low-resource tasks. We provide a thorough empirical study and analysis of this method's benefits showing that it achieves consistent improvements relative to the performance trade-off profile of standard static weighting. We analyze under what data regimes this method is applicable and show its improvements empirically in neural machine translation (NMT) and multi-lingual language modeling.
Dami Choi, Derrick Xin, Hamid Dadkhahi, Justin Gilmer, Ankush Garg, Orhan Firat, Chih-Kuan Yeh, Andrew M. Dai, Behrooz Ghorbani
NeurIPS2
2023 MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
abstract
We introduce MADLAD-400, a manually audited, general domain 3T token monolingual dataset based on CommonCrawl, spanning 419 languages. We discuss the limitations revealed by self-auditing MADLAD-400, and the role data auditing had in the dataset creation process. We then train and release a 10.7B-parameter multilingual machine translation model on 250 billion tokens covering over 450 languages using publicly available data, and find that it is competitive with models that are significantly larger, and report the results on different domains. In addition, we train a 8B-parameter language model, and assess the results on few-shot translation. We make the baseline models available to the research community.
Sneha Reddy Kudugunta, Isaac Caswell, Biao Zhang 0006, Xavier Garcia, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, Orhan Firat
NeurIPS5
2022 Do Current Multi-Task Optimization Methods in Deep Learning Even Help?
abstract
Recent research has proposed a series of specialized optimization algorithms for deep multi-task models. It is often claimed that these multi-task optimization (MTO) methods yield solutions that are superior to the ones found by simply optimizing a weighted average of the task losses. In this paper, we perform large-scale experiments on a variety of language and vision tasks to examine the empirical validity of these claims. We show that, despite the added design and computational complexity of these algorithms, MTO methods do not yield any performance improvements beyond what is achievable via traditional optimization approaches. We highlight alternative strategies that consistently yield improvements to the performance profile and point out common training pitfalls that might cause suboptimal results. Finally, we outline challenges in reliably evaluating the performance of MTO algorithms and discuss potential solutions.
Derrick Xin, Behrooz Ghorbani, Justin Gilmer, Ankush Garg, Orhan Firat
NeurIPS1