EDBT 2026 Demo / reviewers in the wild / expert
Mia Xu Chen
dblp:83/6331-27 · also Xu Chen 0027
· DBLP profile ↗
6ranked-venue papers
3as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Machine translation · 36% Efficient and distributed learning · 26% Representation and self-supervised learning · 18% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
neural machine translation |
0.7 | 2 | 2018 | Training Deeper Neural Machine Translation Models with Transparent Attention · EMNLP 2018 The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation · ACL (1) 2018 |
Natural language and speech › Machine translation
monolingual data augmentation |
0.4 | 1 | 2020 | Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation · ACL 2020 |
Natural language and speech › Machine translation › neural machine translation
multilingual neural machine translation |
0.4 | 1 | 2020 | Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine Translation · ACL 2020 |
Machine learning › Efficient and distributed learning
distributed training |
0.4 | 1 | 2019 | GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019 |
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training |
0.4 | 1 | 2019 | GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019 |
Natural language and speech › Language models and text generation
neural language model |
0.4 | 1 | 2019 | Gmail Smart Compose: Real-Time Assisted Writing · KDD 2019 |
Machine learning › Efficient and distributed learning › distributed training › model parallelism
pipeline parallelism |
0.4 | 1 | 2019 | GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism · NeurIPS 2019 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 1 | 2019 | Gmail Smart Compose: Real-Time Assisted Writing · KDD 2019 |
Machine learning › Representation and self-supervised learning › representation learning › neural network representation learning
deep encoder training |
0.3 | 1 | 2018 | Training Deeper Neural Machine Translation Models with Transparent Attention · EMNLP 2018 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.1 | 1 | 2018 | Training Deeper Neural Machine Translation Models with Transparent Attention · EMNLP 2018 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.7self-supervision · 0.4back-translation · 0.4neural language model · 0.4model parallelism · 0.4large-scale serving infrastructure · 0.4batch-splitting pipelining · 0.4neural machine translation · 0.3attention · 0.3Bi-RNN · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Leveraging Monolingual Data with Self-Supervision for Multilingual Neural Machine TranslationabstractAditya Siddhant, Ankur Bapna, Yuan Cao, Orhan Firat, Mia Chen, Sneha Kudugunta, Naveen Arivazhagan, Yonghui Wu. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Aditya Siddhant, Ankur Bapna, Yuan Cao 0007, Orhan Firat, Mia Xu Chen, Sneha Reddy Kudugunta, Naveen Arivazhagan |
ACL | 5 |
| 2019 | Gmail Smart Compose: Real-Time Assisted WritingabstractIn this paper, we present Smart Compose, a novel system for generating interactive, real-time suggestions in Gmail that assists users in writing mails by reducing repetitive typing. In the design and deployment of such a large-scale and complicated system, we faced several challenges including model selection, performance evaluation, serving and other practical issues. At the core of Smart Compose is a large-scale neural language model. We leveraged state-of-the-art machine learning techniques for language model training which enabled high-quality suggestion prediction, and constructed novel serving infrastructure for high-throughput and real-time inference. Experimental results show the effectiveness of our proposed system design and deployment approach. This system is currently being served in Gmail. Mia Xu Chen, Benjamin N. Lee, Gagan Bansal, Yuan Cao 0007, Shuyuan Zhang 0002, Justin Lu, Jackie Tsay, Andrew M. Dai, Timothy Sohn |
KDD | 1 |
| 2019 | GPipe: Efficient Training of Giant Neural Networks using Pipeline ParallelismabstractScaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or infrastructure. These solutions are often architecture-specific and do not transfer to other machine learning tasks. To address the need for efficient and task-independent model parallelism, we introduce TensorPipe, a pipeline parallelism library that allows scaling any network that can be expressed as a sequence of layers. By pipelining different sub-sequences of layers on separate accelerators, TensorPipe provides the flexibility of scaling a variety of different networks to gigantic sizes efficiently. Moreover, TensorPipe utilizes a novel batch-splitting pipelining algorithm, resulting in almost linear speedup when a model is partitioned across multiple accelerators. We demonstrate the advantages of TensorPipe by training large-scale neural networks on two different tasks with distinct network architectures: (i)Image Classification: We train a 557-million-parameter AmoebaNet model and attain a top-1 accuracy of 84.4% on ImageNet-2012, (ii)Multilingual Neural Machine Translation: We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models. Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le |
NeurIPS | 6 |
| 2018 | The Best of Both Worlds: Combining Recent Advances in Neural Machine TranslationabstractMia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Zhifeng Chen, Yonghui Wu, Macduff Hughes. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Mia Xu Chen, Orhan Firat, Ankur Bapna, Melvin Johnson, Wolfgang Macherey, George F. Foster, Llion Jones, Mike Schuster, Noam Shazeer, Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Macduff Hughes |
ACL (1) | 1 |
| 2018 | Training Deeper Neural Machine Translation Models with Transparent AttentionabstractWhile current state-of-the-art NMT models, such as RNN seq2seq and Transformers, possess a large number of parameters, they are still shallow in comparison to convolutional models used for both text and vision applications.In this work we attempt to train significantly (2-3x) deeper Transformer and Bi-RNN encoders for machine translation.We propose a simple modification to the attention mechanism that eases the optimization of deeper models, and results in consistent gains of 0.7-1.1 BLEU on the benchmark WMT'14 English-German and WMT'15 Czech-English tasks for both architectures. Ankur Bapna, Mia Xu Chen, Orhan Firat, Yuan Cao 0007 |
EMNLP | 2 |
| 2014 | Collaborative representation, sparsity or nonlinearity: What is key to dictionary based classification?abstractRecent studies have suggested that the critical aspect of sparse representation-based classification (SRC) is collaborative representation, rather than sparsity. This has given rise to fast collaborative representation-based classification using 2-norm regularized least squares (CRC-RLS). This paper digs deeper into the difference between SRC and CRC-RLS. We show that linear coding schemes such as CRC-RLS share a common pairwise boundary class B. Moreover, the corresponding pairwise classifiers can be realized by quadratic SVMs. Using three datasets, we show empirically that collaborative representations are not always required, and that a quadratic SVM has superior generalization over CRC-RLS, with fast classification times. However, SRC exhibits the best prediction accuracy. This leads us to posit that the nonlinear coding of SRC is a key attribute. Mia Xu Chen, Peter J. Ramadge |
ICASSP | 1 |