Mia Xu Chen

dblp:83/6331-27 · also Xu Chen 0027 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Machine translation · 36% Efficient and distributed learning · 26% Representation and self-supervised learning · 18%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
neural machine translation
0.722018
Training Deeper Neural Machine Translation Models with Transparent Attention · EMNLP 2018
The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation · ACL (1) 2018
Natural language and speech › Machine translation
monolingual data augmentation
0.412020
Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation · ACL 2020
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation
0.412020
Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation · ACL 2020
Machine learning › Efficient and distributed learning
distributed training
0.412019
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training
0.412019
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019
Natural language and speech › Language models and text generation
neural language model
0.412019
Gmail Smart Compose: Real-Time Assisted Writing · KDD 2019
Machine learning › Efficient and distributed learning › distributed training › model parallelism
pipeline parallelism
0.412019
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019
Natural language and speech › Language models and text generation
text generation
0.412019
Gmail Smart Compose: Real-Time Assisted Writing · KDD 2019
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning
deep encoder training
0.312018
Training Deeper Neural Machine Translation Models with Transparent Attention · EMNLP 2018
Machine learning › Deep learning architectures and training
attention mechanism
0.112018
Training Deeper Neural Machine Translation Models with Transparent Attention · EMNLP 2018

Methods — techniques the papers use, named apart from their topics

transformer · 0.7self-supervision · 0.4back-translation · 0.4neural language model · 0.4model parallelism · 0.4large-scale serving infrastructure · 0.4batch-splitting pipelining · 0.4neural machine translation · 0.3attention · 0.3Bi-RNN · 0.3
YearPublicationVenuePosition
2020 Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation
abstract
Aditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, Yonghui Wu. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Aditya Siddhant, Ankur Bapna, Yuan Cao 0007, Orhan Firat, Mia Xu Chen, Sneha Reddy Kudugunta, Naveen Arivazhagan
ACL5
2019 Gmail Smart Compose: Real-Time Assisted Writing
abstract
In this paper, we present Smart Compose, a novel system for generating interactive, real-time suggestions in Gmail that assists users in writing mails by reducing repetitive typing. In the design and deployment of such a large-scale and complicated system, we faced several challenges including model selection, performance evaluation, serving and other practical issues. At the core of Smart Compose is a large-scale neural language model. We leveraged state-of-the-art machine learning techniques for language model training which enabled high-quality suggestion prediction, and constructed novel serving infrastructure for high-throughput and real-time inference. Experimental results show the effectiveness of our proposed system design and deployment approach. This system is currently being served in Gmail.
Mia Xu Chen, Benjamin N. Lee, Gagan Bansal, Yuan Cao 0007, Shuyuan Zhang 0002, Justin Lu, Jackie Tsay, Andrew M. Dai, Timothy Sohn
KDD1
2019 GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
abstract
Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or infrastructure. These solutions are often architecture-specific and do not transfer to other machine learning tasks. To address the need for efficient and task-independent model parallelism, we introduce TensorPipe, a pipeline parallelism library that allows scaling any network that can be expressed as a sequence of layers. By pipelining different sub-sequences of layers on separate accelerators, TensorPipe provides the flexibility of scaling a variety of different networks to gigantic sizes efficiently. Moreover, TensorPipe utilizes a novel batch-splitting pipelining algorithm, resulting in almost linear speedup when a model is partitioned across multiple accelerators. We demonstrate the advantages of TensorPipe by training large-scale neural networks on two different tasks with distinct network architectures: (i)Image Classification: We train a 557-million-parameter AmoebaNet model and attain a top-1 accuracy of 84.4% on ImageNet-2012, (ii)Multilingual Neural Machine Translation: We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models.
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le
NeurIPS6
2018 The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation
abstract
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, Macduff Hughes. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George F. Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Macduff Hughes
ACL (1)1
2018 Training Deeper Neural Machine Translation Models with Transparent Attention
abstract
While current state-of-the-art NMT models, such as RNN seq2seq and Transformers, possess a large number of parameters, they are still shallow in comparison to convolutional models used for both text and vision applications.In this work we attempt to train significantly (2-3x) deeper Transformer and Bi-RNN encoders for machine translation.We propose a simple modification to the attention mechanism that eases the optimization of deeper models, and results in consistent gains of 0.7-1.1 BLEU on the benchmark WMT'14 English-German and WMT'15 Czech-English tasks for both architectures.
Ankur Bapna, Mia Xu Chen, Orhan Firat, Yuan Cao 0007
EMNLP2
2014 Collaborative representation, sparsity or nonlinearity: What is key to dictionary based classification?
abstract
Recent studies have suggested that the critical aspect of sparse representation-based classification (SRC) is collaborative representation, rather than sparsity. This has given rise to fast collaborative representation-based classification using 2-norm regularized least squares (CRC-RLS). This paper digs deeper into the difference between SRC and CRC-RLS. We show that linear coding schemes such as CRC-RLS share a common pairwise boundary class B. Moreover, the corresponding pairwise classifiers can be realized by quadratic SVMs. Using three datasets, we show empirically that collaborative representations are not always required, and that a quadratic SVM has superior generalization over CRC-RLS, with fast classification times. However, SRC exhibits the best prediction accuracy. This leads us to posit that the nonlinear coding of SRC is a key attribute.
Mia Xu Chen, Peter J. Ramadge
ICASSP1