VLDB 2026 Research / reviewers in the wild / expert
Zihang Dai
dblp:182/1934
· DBLP profile ↗
26ranked-venue papers
7as first author
8since 2021 · last 2023
0000-0002-6509-5640ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
23 papers |
Deep learning architectures and training · 31% Language models and text generation · 19% Representation and self-supervised learning · 11% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% |
Topics — the 30 heaviest of 55, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
transformer |
1.7 | 4 | 2022 | Transformer Quality in Linear Time · ICML 2022 Pay Attention to MLPs · NeurIPS 2021 Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing · NeurIPS 2020 |
Machine learning › Learning paradigms
semi-supervised learning |
1.2 | 3 | 2021 | Meta Pseudo Labels · CVPR 2021 Unsupervised Data Augmentation for Consistency Training · NeurIPS 2020 Good Semi-supervised Learning That Requires a Bad GAN · NIPS 2017 |
Machine learning › Deep learning architectures and training › transformer
efficient transformer |
1.0 | 2 | 2022 | Transformer Quality in Linear Time · ICML 2022 Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing · NeurIPS 2020 |
Natural language and speech › Language models and text generation
language modeling |
1.0 | 2 | 2022 | Transformer Quality in Linear Time · ICML 2022 Transformer-XL: Attentive Language Models beyond a Fixed-Length Context · ACL (1) 2019 |
Machine learning › Representation and self-supervised learning
pre-training |
1.0 | 2 | 2022 | SimVLM: Simple Visual Language Model Pretraining with Weak Supervision · ICLR 2022 XLNet: Generalized Autoregressive Pretraining for Language Understanding · NeurIPS 2019 |
Machine learning › Deep learning architectures and training
data augmentation |
0.8 | 2 | 2020 | Unsupervised Data Augmentation for Consistency Training · NeurIPS 2020 SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation · EMNLP 2018 |
Natural language and speech › Language models and text generation
neural language model |
0.7 | 2 | 2019 | Fast and Simple Mixture of Softmaxes with BPE and Hybrid-LightRNN for Language Generation · AAAI 2019 Breaking the Softmax Bottleneck: A High-Rank RNN Language Model · ICLR 2018 |
Natural language and speech › Machine translation
neural machine translation |
0.7 | 2 | 2019 | Fast and Simple Mixture of Softmaxes with BPE and Hybrid-LightRNN for Language Generation · AAAI 2019 SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine Translation · EMNLP 2018 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.7 | 2 | 2021 | CoAtNet: Marrying Convolution and Attention for All Data Sizes · NeurIPS 2021 Pay Attention to MLPs · NeurIPS 2021 |
Machine learning › Generative modeling
generative adversarial network |
0.6 | 2 | 2017 | Good Semi-supervised Learning That Requires a Bad GAN · NIPS 2017 Calibrating Energy-based Generative Adversarial Networks · ICLR (Poster) 2017 |
Computer vision › Vision and language
vision-language pretraining |
0.6 | 1 | 2022 | SimVLM: Simple Visual Language Model Pretraining with Weak Supervision · ICLR 2022 |
Natural language and speech › Language models and text generation › neural language model
autoregressive language model |
0.5 | 1 | 2021 | Searching for Efficient Transformers for Language Modeling · NeurIPS 2021 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.5 | 1 | 2021 | CoAtNet: Marrying Convolution and Attention for All Data Sizes · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention |
0.5 | 1 | 2021 | Combiner: Full Attention Transformer with Sparse Computation Cost · NeurIPS 2021 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
0.5 | 1 | 2021 | Searching for Efficient Transformers for Language Modeling · NeurIPS 2021 |
Machine learning › Learning paradigms › semi-supervised learning
pseudo-labeling |
0.5 | 1 | 2021 | Meta Pseudo Labels · CVPR 2021 |
Machine learning › Deep learning architectures and training
teacher-student framework |
0.5 | 1 | 2021 | Meta Pseudo Labels · CVPR 2021 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
transformer architecture search |
0.5 | 1 | 2021 | Searching for Efficient Transformers for Language Modeling · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › attention mechanism
transformer attention |
0.5 | 1 | 2021 | Combiner: Full Attention Transformer with Sparse Computation Cost · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › regularization
consistency training |
0.4 | 1 | 2020 | Unsupervised Data Augmentation for Consistency Training · NeurIPS 2020 |
Machine learning › Representation and self-supervised learning › representation learning
language representation learning |
0.4 | 1 | 2020 | A Mutual Information Maximization Perspective of Language Representation Learning · ICLR 2020 |
Machine learning › Representation and self-supervised learning
mutual information maximization |
0.4 | 1 | 2020 | A Mutual Information Maximization Perspective of Language Representation Learning · ICLR 2020 |
Machine learning › Representation and self-supervised learning › pre-training
autoregressive pre-training |
0.4 | 1 | 2019 | XLNet: Generalized Autoregressive Pretraining for Language Understanding · NeurIPS 2019 |
Natural language and speech › Language models and text generation › language modeling
long-context language modeling |
0.4 | 1 | 2019 | Transformer-XL: Attentive Language Models beyond a Fixed-Length Context · ACL (1) 2019 |
Machine learning › Transfer learning and domain adaptation
negative transfer |
0.4 | 1 | 2019 | Characterizing and Avoiding Negative Transfer · CVPR 2019 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.4 | 1 | 2019 | Re-examination of the Role of Latent Variables in Sequence Modeling · NeurIPS 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
sequential latent variable model |
0.4 | 1 | 2019 | Re-examination of the Role of Latent Variables in Sequence Modeling · NeurIPS 2019 |
Machine learning › Transfer learning and domain adaptation › multi-source learning
source data selection |
0.4 | 1 | 2019 | Characterizing and Avoiding Negative Transfer · CVPR 2019 |
Natural language and speech › Language models and text generation › natural language understanding
cloze task |
0.3 | 1 | 2018 | Large-scale Cloze Test Dataset Created by Teachers · EMNLP 2018 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
credit assignment |
0.3 | 1 | 2018 | From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence Prediction · ACL (1) 2018 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.3recurrent neural network · 0.6weak supervision · 0.6linear approximation · 0.6gated attention unit · 0.6teacher-student network · 0.5pseudo labels · 0.5meta-learning · 0.5hybrid architecture · 0.5depthwise convolution · 0.5language model · 0.3baseline models · 0.3sparse attention · 0.3knowledge transfer · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Combined scaling for zero-shot transfer learningabstractRecent developments in multimodal training methodologies, including CLIP and ALIGN, obviate the necessity for individual data labeling. These approaches utilize pairs of data and corresponding textual information found online as a form of weak supervision signal. However, models employing this kind of weak supervision are not as competitive as their supervised and semi-supervised counterparts when sufficient labeled data is accessible. This performance gap constrains the applicability of weekly supervised models. In this paper, we narrow the gap by proposing a combined scaling method, named BASIC, that achieves 85.7% top-1 accuracy on the ImageNet ILSVRC-2012 validation set without learning from any labeled ImageNet example. This accuracy surpasses best-published similar models, CLIP and ALIGN, by 9.3%. Our BASIC model also shows significant improvements in robustness benchmarks. For instance, on 5 test sets with natural distribution shifts such as ImageNet-{A,R,V2,Sketch} and ObjectNet, our model achieves 84.3% top-1 average accuracy, only a small drop from its original ImageNet accuracy. To achieve these results, we first develop a theoretical framework which shows that larger contrastive batch sizes lead to smaller generalization gaps for image-text models such as CLIP and ALIGN. Based on this theoretical result, we scale up the contrastive learning framework of CLIP and ALIGN in three dimensions (data size, model size, and batch size) by proposing a new method using gradient checkpointing and model parallelism. As a result, our dataset has 6.6B noisy image-text pairs, which is 4x larger than ALIGN, and 16x larger than CLIP. Our largest model has 3B weights, which is 3.75x larger in parameters and 8x larger in FLOPs than ALIGN and CLIP. Finally, our batch size is 65536 which is 2x more than CLIP and 4x more than ALIGN. Hieu Pham 0001, Zihang Dai, Golnaz Ghiasi, Kenji Kawaguchi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, Quoc V. Le |
Neurocomputing | 2 |
| 2022 | SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, Yuan Cao 0007 |
ICLR | 4 |
| 2022 | Transformer Quality in Linear TimeabstractWe revisit the design choices in Transformers, and propose methods to address their weaknesses in handling long sequences. First, we propose a simple layer named gated attention unit, which allows the use of a weaker single-head attention with minimal quality loss. We then propose a linear approximation method complementary to this new layer, which is accelerator-friendly and highly competitive in quality. The resulting model, named FLASH, matches the perplexity of improved Transformers over both short (512) and long (8K) context lengths, achieving training speedups of up to 4.9x on Wiki-40B and 12.1x on PG-19 for auto-regressive language modeling, and 4.8x on C4 for masked language modeling. Weizhe Hua, Zihang Dai, Hanxiao Liu, Quoc V. Le |
ICML | 2 |
| 2021 | Meta Pseudo LabelsabstractWe present Meta Pseudo Labels, a semi-supervised learning method that achieves a new state-of-the-art top-1 accuracy of 90.2% on ImageNet, which is 1.6% better than the existing state-of-the-art [16]. Like Pseudo Labels, Meta Pseudo Labels has a teacher network to generate pseudo labels on unlabeled data to teach a student network. However, unlike Pseudo Labels where the teacher is fixed, the teacher in Meta Pseudo Labels is constantly adapted by the feedback of the student’s performance on the labeled dataset. As a result, the teacher generates better pseudo labels to teach the student.1 Hieu Pham 0001, Zihang Dai, Qizhe Xie, Quoc V. Le |
CVPR | 2 |
| 2021 | CoAtNet: Marrying Convolution and Attention for All Data SizesabstractTransformers have attracted increasing interests in computer vision, but they still fall behind state-of-the-art convolutional networks. In this work, we show that while Transformers tend to have larger model capacity, their generalization can be worse than convolutional networks due to the lack of the right inductive bias. To effectively combine the strengths from both architectures, we present CoAtNets(pronounced "coat" nets), a family of hybrid models built from two key insights: (1) depthwise Convolution and self-Attention can be naturally unified via simple relative attention; (2) vertically stacking convolution layers and attention layers in a principled way is surprisingly effective in improving generalization, capacity and efficiency. Experiments show that our CoAtNets achieve state-of-the-art performance under different resource constraints across various datasets: Without extra data, CoAtNet achieves 86.0% ImageNet top-1 accuracy; When pre-trained with 13M images from ImageNet-21K, our CoAtNet achieves 88.56% top-1 accuracy, matching ViT-huge pre-trained with 300M images from JFT-300M while using 23x less data; Notably, when we further scale up CoAtNet with JFT-3B, it achieves 90.88% top-1 accuracy on ImageNet, establishing a new state-of-the-art result. Zihang Dai, Hanxiao Liu, Quoc V. Le, Mingxing Tan |
NeurIPS | 1 |
| 2021 | Pay Attention to MLPsabstractTransformers have become one of the most important architectural innovations in deep learning and have enabled many breakthroughs over the past few years. Here we propose a simple network architecture, gMLP, based solely on MLPs with gating, and show that it can perform as well as Transformers in key language and vision applications. Our comparisons show that self-attention is not critical for Vision Transformers, as gMLP can achieve the same accuracy. For BERT, our model achieves parity with Transformers on pretraining perplexity and is better on some downstream NLP tasks. On finetuning tasks where gMLP performs worse, making the gMLP model substantially larger can close the gap with Transformers. In general, our experiments show that gMLP can scale as well as Transformers over increased data and compute. Hanxiao Liu, Zihang Dai, David R. So, Quoc V. Le |
NeurIPS | 2 |
| 2021 | Combiner: Full Attention Transformer with Sparse Computation CostabstractTransformers provide a class of expressive architectures that are extremely effective for sequence modeling. However, the key limitation of transformers is their quadratic memory and time complexity $\mathcal{O}(L^2)$ with respect to the sequence length in attention layers, which restricts application in extremely long sequences. Most existing approaches leverage sparsity or low-rank assumptions in the attention matrix to reduce cost, but sacrifice expressiveness. Instead, we propose Combiner, which provides full attention capability in each attention head while maintaining low computation and memory complexity. The key idea is to treat the self-attention mechanism as a conditional expectation over embeddings at each location, and approximate the conditional distribution with a structured factorization. Each location can attend to all other locations, either via direct attention, or through indirect attention to abstractions, which are again conditional expectations of embeddings from corresponding local regions. We show that most sparse attention patterns used in existing sparse transformers are able to inspire the design of such factorization for full attention, resulting in the same sub-quadratic cost ($\mathcal{O}(L\log(L))$ or $\mathcal{O}(L\sqrt{L})$). Combiner is a drop-in replacement for attention layers in existing transformers and can be easily implemented in common frameworks. An experimental evaluation on both autoregressive and bidirectional sequence tasks demonstrates the effectiveness of this approach, yielding state-of-the-art results on several image and text modeling tasks. Hongyu Ren, Hanjun Dai, Zihang Dai, Sherry Yang 0001, Jure Leskovec, Dale Schuurmans, Bo Dai 0001 |
NeurIPS | 3 |
| 2021 | Searching for Efficient Transformers for Language ModelingabstractLarge Transformer models have been central to recent advances in natural language processing. The training and inference costs of these models, however, have grown rapidly and become prohibitively expensive. Here we aim to reduce the costs of Transformers by searching for a more efficient variant. Compared to previous approaches, our search is performed at a lower level, over the primitives that define a Transformer TensorFlow program. We identify an architecture, named Primer, that has a smaller training cost than the original Transformer and other variants for auto-regressive language modeling. Primer’s improvements can be mostly attributed to two simple modifications: squaring ReLU activations and adding a depthwise convolution layer after each Q, K, and V projection in self-attention.Experiments show Primer’s gains over Transformer increase as compute scale grows and follow a power law with respect to quality at optimal model sizes. We also verify empirically that Primer can be dropped into different codebases to significantly speed up training without additional tuning. For example, at a 500M parameter size, Primer improves the original T5 architecture on C4 auto-regressive language modeling, reducing the training cost by 4X. Furthermore, the reduced training cost means Primer needs much less compute to reach a target one-shot performance. For instance, in a 1.9B parameter configuration similar to GPT-3 XL, Primer uses 1/3 of the training compute to achieve the same one-shot performance as Transformer. We open source our models and several comparisons in T5 to help with reproducibility. David R. So, Wojciech Manke, Hanxiao Liu, Zihang Dai, Noam Shazeer, Quoc V. Le |
NeurIPS | 4 |
| 2020 | A Mutual Information Maximization Perspective of Language Representation Learning
Lingpeng Kong, Cyprien de Masson d'Autume, Lei Yu 0008, Wang Ling, Zihang Dai, Dani Yogatama |
ICLR | 5 |
| 2020 | Wiki-40B: Multilingual Language Model DatasetabstractWe propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families. With around 40 billion characters, we hope this new resource will accelerate the research of multilingual modeling. We train monolingual causal language models using a state-of-the-art model (Transformer-XL) establishing baselines for many languages. We also introduce the task of multilingual causal language modeling where we train our model on the combined text of 40+ languages from Wikipedia with different vocabulary sizes and evaluate on the languages individually. We released the cleaned-up text of 40+ Wikipedia language editions, the corresponding trained monolingual language models, and several multilingual language models with different fixed vocabulary sizes. Mandy Guo, Zihang Dai, Denny Vrandecic, Rami Al-Rfou |
LREC | 2 |
| 2020 | Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language ProcessingabstractWith the success of language pretraining, it is highly desirable to develop more efficient architectures of good scalability that can exploit the abundant unlabeled data at a lower cost. To improve the efficiency, we examine the much-overlooked redundancy in maintaining a full-length token-level presentation, especially for tasks that only require a single-vector presentation of the sequence. With this intuition, we propose Funnel-Transformer which gradually compresses the sequence of hidden states to a shorter one and hence reduces the computation cost. More importantly, by re-investing the saved FLOPs from length reduction in constructing a deeper or wider model, we further improve the model capacity. In addition, to perform token-level predictions as required by common pretraining objectives, Funnel-Transformer is able to recover a deep representation for each token from the reduced hidden sequence via a decoder. Empirically, with comparable or fewer FLOPs, Funnel-Transformer outperforms the standard Transformer on a wide variety of sequence-level prediction tasks, including text classification, language understanding, and reading comprehension. Zihang Dai, Guokun Lai, Yiming Yang 0002, Quoc V. Le |
NeurIPS | 1 |
| 2020 | Unsupervised Data Augmentation for Consistency TrainingabstractSemi-supervised learning lately has shown much promise in improving deep learning models when labeled data is scarce. Common among recent approaches is the use of consistency training on a large amount of unlabeled data to constrain model predictions to be invariant to input noise. In this work, we present a new perspective on how to effectively noise unlabeled examples and argue that the quality of noising, specifically those produced by advanced data augmentation methods, plays a crucial role in semi-supervised learning. By substituting simple noising operations with advanced data augmentation methods such as RandAugment and back-translation, our method brings substantial improvements across six language and three vision tasks under the same consistency training framework. On the IMDb text classification dataset, with only 20 labeled examples, our method achieves an error rate of 4.20, outperforming the state-of-the-art model trained on 25,000 labeled examples. On a standard semi-supervised learning benchmark, CIFAR-10, our method outperforms all previous approaches and achieves an error rate of 5.43 with only 250 examples. Our method also combines well with transfer learning, e.g., when finetuning from BERT, and yields improvements in high-data regime, such as ImageNet, whether when there is only 10% labeled data or when a full labeled set with 1.3M extra unlabeled examples is used. Code is available at https://github.com/google-research/uda. Qizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong, Quoc V. Le |
NeurIPS | 2 |
| 2019 | Fast and Simple Mixture of Softmaxes with BPE and Hybrid-LightRNN for Language GenerationabstractMixture of Softmaxes (MoS) has been shown to be effective at addressing the expressiveness limitation of Softmax-based models. Despite the known advantage, MoS is practically sealed by its large consumption of memory and computational time due to the need of computing multiple Softmaxes. In this work, we set out to unleash the power of MoS in practical applications by investigating improved word coding schemes, which could effectively reduce the vocabulary size and hence relieve the memory and computation burden. We show both BPE and our proposed Hybrid-LightRNN lead to improved encoding mechanisms that can halve the time and memory consumption of MoS without performance losses. With MoS, we achieve an improvement of 1.5 BLEU scores on IWSLT 2014 German-to-English corpus and an improvement of 0.76 CIDEr score on image captioning. Moreover, on the larger WMT 2014 machine translation dataset, our MoSboosted Transformer yields 29.6 BLEU score for English-toGerman and 42.1 BLEU score for English-to-French, outperforming the single-Softmax Transformer by 0.9 and 0.4 BLEU scores respectively and achieving the state-of-the-art result on WMT 2014 English-to-German task. Xiang Kong, Qizhe Xie, Zihang Dai, Eduard H. Hovy |
AAAI | 3 |
| 2019 | Transformer-XL: Attentive Language Models beyond a Fixed-Length ContextabstractTransformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling.We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence.It consists of a segment-level recurrence mechanism and a novel positional encoding scheme.Our method not only enables capturing longer-term dependency, but also resolves the context fragmentation problem.As a result, Transformer-XL learns dependency that is 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800+ times faster than vanilla Transformers during evaluation.Notably, we improve the state-ofthe-art results of bpc/perplexity to 0.99 on en-wiki8, 1.08 on text8, 18.3 on WikiText-103, 21.8 on One Billion Word, and 54.5 on Penn Treebank (without finetuning).When trained only on WikiText-103, Transformer-XL manages to generate reasonably coherent, novel text articles with thousands of tokens.Our code, pretrained models, and hyperparameters are available in both Tensorflow and PyTorch 1 . Zihang Dai, Zhilin Yang 0001, Yiming Yang 0002, Jaime G. Carbonell, Quoc V. Le, Ruslan Salakhutdinov |
ACL (1) | 1 |
| 2019 | Characterizing and Avoiding Negative TransferabstractWhen labeled data is scarce for a specific target task, transfer learning often offers an effective solution by utilizing data from a related source task. However, when transferring knowledge from a less related source, it may inversely hurt the target performance, a phenomenon known as negative transfer. Despite its pervasiveness, negative transfer is usually described in an informal manner, lacking rigorous definition, careful analysis, or systematic treatment. This paper proposes a formal definition of negative transfer and analyzes three important aspects thereof. Stemming from this analysis, a novel technique is proposed to circumvent negative transfer by filtering out unrelated source data. Based on adversarial networks, the technique is highly generic and can be applied to a wide range of transfer learning algorithms. The proposed approach is evaluated on six state-of-the-art deep transfer methods via experiments on four benchmark datasets with varying levels of difficulty. Empirically, the proposed method consistently improves the performance of all baseline methods and largely avoids negative transfer, even when the source data is degenerate. Zihang Dai, Barnabás Póczos, Jaime G. Carbonell |
CVPR | 2 |
| 2019 | Re-examination of the Role of Latent Variables in Sequence ModelingabstractWith latent variables, stochastic recurrent models have achieved state-of-the-art performance in modeling sound-wave sequence. However, opposite results are also observed in other domains, where standard recurrent networks often outperform stochastic models. To better understand this discrepancy, we re-examine the roles of latent variables in stochastic recurrent models for speech density estimation. Our analysis reveals that under the restriction of fully factorized output distribution in previous evaluations, the stochastic variants were implicitly leveraging intra-step correlation but the deterministic recurrent baselines were prohibited to do so, resulting in an unfair comparison. To correct the unfairness, we remove such restriction in our re-examination, where all the models can explicitly leverage intra-step correlation with an auto-regressive structure. Over a diverse set of univariate and multivariate sequential data, including human speech, MIDI music, handwriting trajectory, and frame-permuted speech, our results show that stochastic recurrent models fail to deliver the performance advantage claimed in previous work. %exhibit any practical advantage despite the claimed theoretical superiority. In contrast, standard recurrent models equipped with an auto-regressive output distribution consistently perform better, dramatically advancing the state-of-the-art results on three speech datasets. Guokun Lai, Zihang Dai, Yiming Yang 0002, Shinjae Yoo |
NeurIPS | 2 |
| 2019 | XLNet: Generalized Autoregressive Pretraining for Language UnderstandingabstractWith the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency between the masked positions and suffers from a pretrain-finetune discrepancy. In light of these pros and cons, we propose XLNet, a generalized autoregressive pretraining method that (1) enables learning bidirectional contexts by maximizing the expected likelihood over all permutations of the factorization order and (2) overcomes the limitations of BERT thanks to its autoregressive formulation. Furthermore, XLNet integrates ideas from Transformer-XL, the state-of-the-art autoregressive model, into pretraining. Empirically, under comparable experiment setting, XLNet outperforms BERT on 20 tasks, often by a large margin, including question answering, natural language inference, sentiment analysis, and document ranking. Zhilin Yang 0001, Zihang Dai, Yiming Yang 0002, Jaime G. Carbonell, Ruslan Salakhutdinov, Quoc V. Le |
NeurIPS | 2 |
| 2018 | From Credit Assignment to Entropy Regularization: Two New Algorithms for Neural Sequence PredictionabstractIn this work, we study the credit assignment problem in reward augmented maximum likelihood (RAML) learning, and establish a theoretical equivalence between the token-level counterpart of RAML and the entropy regularized reinforcement learning.Inspired by the connection, we propose two sequence prediction algorithms, one extending RAML with fine-grained credit assignment and the other improving Actor-Critic with a systematic entropy regularization.On two benchmark datasets, we show the proposed algorithms outperform RAML and Actor-Critic respectively, providing new alternatives to sequence prediction. Zihang Dai, Qizhe Xie, Eduard H. Hovy |
ACL (1) | 1 |
| 2018 | SwitchOut: an Efficient Data Augmentation Algorithm for Neural Machine TranslationabstractIn this work, we examine methods for data augmentation for text-based tasks such as neural machine translation (NMT).We formulate the design of a data augmentation policy with desirable properties as an optimization problem, and derive a generic analytic solution.This solution not only subsumes some existing augmentation schemes, but also leads to an extremely simple data augmentation strategy for NMT: randomly replacing words in both the source sentence and the target sentence with other random words from their corresponding vocabularies.We name this method SwitchOut.Experiments on three translation datasets of different scales show that SwitchOut yields consistent improvements of about 0.5 BLEU, achieving better or comparable performances to strong alternatives such as word dropout (Sennrich et al., 2016a).Code to implement this method is included in the appendix. Xinyi Wang 0001, Zihang Dai, Graham Neubig |
EMNLP | 3 |
| 2018 | Large-scale Cloze Test Dataset Created by TeachersabstractCloze tests are widely adopted in language exams to evaluate students' language proficiency.In this paper, we propose the first large-scale human-created cloze test dataset CLOTH 1 2 , containing questions used in middle-school and high-school language exams.With missing blanks carefully created by teachers and candidate choices purposely designed to be nuanced, CLOTH requires a deeper language understanding and a wider attention span than previously automaticallygenerated cloze datasets.We test the performance of dedicatedly designed baseline models including a language model trained on the One Billion Word Corpus and show humans outperform them by a significant margin.We investigate the source of the performance gap, trace model deficiencies to some distinct properties of CLOTH, and identify the limited ability of comprehending the long-term context to be the key bottleneck. Qizhe Xie, Guokun Lai, Zihang Dai, Eduard H. Hovy |
EMNLP | 3 |
| 2018 | Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
Zhilin Yang 0001, Zihang Dai, Ruslan Salakhutdinov, William W. Cohen |
ICLR | 2 |
| 2017 | An Interpretable Knowledge Transfer Model for Knowledge Base CompletionabstractKnowledge bases are important resources for a variety of natural language processing tasks but suffer from incompleteness.We propose a novel embedding model, ITransF, to perform knowledge base completion.Equipped with a sparse attention mechanism, ITransF discovers hidden concepts of relations and transfer statistical strength through the sharing of concepts.Moreover, the learned associations between relations and concepts, which are represented by sparse attention vectors, can be interpreted easily.We evaluate ITransF on two benchmark datasets-WN18 and FB15k for knowledge base completion and obtains improvements on both the mean rank and Hits@10 metrics, over all baselines that do not use additional information. Qizhe Xie, Xuezhe Ma, Zihang Dai, Eduard H. Hovy |
ACL (1) | 3 |
| 2017 | Calibrating Energy-based Generative Adversarial Networks
Zihang Dai, Amjad Almahairi, Philip Bachman, Eduard H. Hovy, Aaron C. Courville |
ICLR (Poster) | 1 |
| 2017 | Good Semi-supervised Learning That Requires a Bad GANabstractSemi-supervised learning methods based on generative adversarial networks (GANs) obtained strong empirical results, but it is not clear 1) how the discriminator benefits from joint training with a generator, and 2) why good semi-supervised classification performance and a good generator cannot be obtained at the same time. Theoretically we show that given the discriminator objective, good semi-supervised learning indeed requires a bad generator, and propose the definition of a preferred generator. Empirically, we derive a novel formulation based on our analysis that substantially improves over feature matching GANs, obtaining state-of-the-art results on multiple benchmark datasets. Zihang Dai, Zhilin Yang 0001, Fan Yang 0058, William W. Cohen, Ruslan Salakhutdinov |
NIPS | 1 |
| 2017 | Controllable Invariance through Adversarial Feature LearningabstractLearning meaningful representations that maintain the content necessary for a particular task while filtering away detrimental variations is a problem of great interest in machine learning. In this paper, we tackle the problem of learning representations invariant to a specific factor or trait of data. The representation learning process is formulated as an adversarial minimax game. We analyze the optimal equilibrium of such a game and find that it amounts to maximizing the uncertainty of inferring the detrimental factor given the representation while maximizing the certainty of making task-specific predictions. On three benchmark tasks, namely fair and bias-free classification, language-independent generation, and lighting-independent image classification, we show that the proposed framework induces an invariant representation, and leads to better generalization evidenced by the improved performance. Qizhe Xie, Zihang Dai, Yulun Du, Eduard H. Hovy, Graham Neubig |
NIPS | 2 |
| 2016 | CFO: Conditional Focused Neural Question Answering with Large-scale Knowledge BasesabstractHow can we enable computers to automatically answer questions like "Who created the character Harry Potter"? Carefully built knowledge bases provide rich sources of facts.However, it remains a challenge to answer factoid questions raised in natural language due to numerous expressions of one question.In particular, we focus on the most common questions -ones that can be answered with a single fact in the knowledge base.We propose CFO, a Conditional Focused neuralnetwork-based approach to answering factoid questions with knowledge bases.Our approach first zooms in a question to find more probable candidate subject mentions, and infers the final answers with a unified conditional probabilistic framework.Powered by deep recurrent neural networks and neural embeddings, our proposed CFO achieves an accuracy of 75.7% on a dataset of 108k questions -the largest public one to date.It outperforms the current state of the art by an absolute margin of 11.8%. Zihang Dai, Lei Li 0005, Wei Xu 0017 |
ACL (1) | 1 |