VLDB 2026 Research / reviewers in the wild / expert
Zhilin Yang 0001
dblp:54/6349-1
· DBLP profile ↗
31ranked-venue papers
14as first author
10since 2021 · last 2023
0009-0008-0681-9603ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 14 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A Universal Discriminator for Zero-Shot GeneralizationabstractGenerative modeling has been the dominant approach for large-scale pretraining and zeroshot generalization.In this work, we challenge this convention by showing that discriminative approaches perform substantially better than generative ones on a large number of NLP tasks.Technically, we train a single discriminator to predict whether a text sample comes from the true data distribution, similar to GANs.Since many NLP tasks can be formulated as selecting from a few options, we use this discriminator to predict the concatenation of input and which option has the highest probability of coming from the true data distribution.This simple formulation achieves state-of-theart zero-shot results on the T0 benchmark, outperforming T0 by 16.0%, 7.8%, and 11.5% respectively on different scales.In the finetuning setting, our approach also achieves new stateof-the-art results on a wide range of NLP tasks, with only 1/4 parameters of previous methods.Meanwhile, our approach requires minimal prompting efforts, which largely improves robustness and is essential for real-world applications.Furthermore, we also jointly train a generalized UD in combination with generative tasks, which maintains its advantage on discriminative tasks and simultaneously works on generative tasks. Haike Xu, Zongyu Lin, Yanan Zheng, Zhilin Yang 0001 |
ACL (1) | 5 |
| 2023 | Compositional Task Representations for Large Language Models
Zefan Cai, Hanwei Xu, Chonghua Liao, Yanan Zheng, Zhilin Yang 0001 |
ICLR | 6 |
| 2023 | Not All Tasks Are Born Equal: Understanding Zero-Shot Generalization
Zongyu Lin, Yanan Zheng, Jian Li 0015, Zhilin Yang 0001 |
ICLR | 5 |
| 2023 | CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-XabstractLarge pre-trained code generation models, such as OpenAI Codex, can generate syntax-and function-correct code, making the coding of programmers more productive. In this paper, we introduce CodeGeeX, a multilingual model with 13 billion parameters for code generation. CodeGeeX is pre-trained on 850 billion tokens of 23 programming languages as of June 2022. Our extensive experiments suggest that CodeGeeX outperforms multilingual code models of similar scale for both the tasks of code generation and translation on HumanEval-X. Building upon HumanEval (Python only), we develop the HumanEval-X benchmark for evaluating multilingual models by hand-writing the solutions in C++, Java, JavaScript, and Go. In addition, we build CodeGeeX-based extensions on Visual Studio Code, JetBrains, and Cloud Studio, generating 8 billion tokens for tens of thousands of active users per week. Our user study demonstrates that CodeGeeX can help to increase coding efficiency for 83.4% of its users. Finally, CodeGeeX is publicly accessible since Sep. 2022, we open-sourced its code, model weights, API, extensions, and HumanEval-X at https://github.com/THUDM/CodeGeeX. Qinkai Zheng, Xu Zou 0001, Yuxiao Dong, Shan Wang 0023, Lei Shen 0002, Andi Wang 0003, Yang Li 0074, Teng Su, Zhilin Yang 0001, Jie Tang 0001 |
KDD | 12 |
| 2022 | GLM: General Language Model Pretraining with Autoregressive Blank InfillingabstractZhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, Jie Tang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Zhengxiao Du, Yujie Qian, Xiao Liu 0036, Ming Ding 0004, Jiezhong Qiu, Zhilin Yang 0001, Jie Tang 0001 |
ACL (1) | 6 |
| 2022 | FewNLU: Benchmarking State-of-the-Art Methods for Few-Shot Natural Language UnderstandingabstractYanan Zheng, Jing Zhou, Yujie Qian, Ming Ding, Chonghua Liao, Li Jian, Ruslan Salakhutdinov, Jie Tang, Sebastian Ruder, Zhilin Yang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yanan Zheng, Yujie Qian, Ming Ding 0004, Chonghua Liao, Li Jian, Ruslan Salakhutdinov, Jie Tang 0001, Sebastian Ruder, Zhilin Yang 0001 |
ACL (1) | 10 |
| 2022 | FlipDA: Effective and Robust Data Augmentation for Few-Shot LearningabstractMost previous methods for text data augmentation are limited to simple tasks and weak baselines.We explore data augmentation on hard tasks (i.e., few-shot natural language understanding) and strong baselines (i.e., pretrained models with over one billion parameters).Under this setting, we reproduced a large number of previous augmentation methods and found that these methods bring marginal gains at best and sometimes degrade the performance much.To address this challenge, we propose a novel data augmentation method FlipDA that jointly uses a generative model and a classifier to generate label-flipped data.Central to the idea of FlipDA is the discovery that generating labelflipped data is more crucial to the performance than generating label-preserved data.Experiments show that FlipDA achieves a good tradeoff between effectiveness and robustness-it substantially improves many tasks while not negatively affecting the others. 1 Yanan Zheng, Jie Tang 0001, Li Jian, Zhilin Yang 0001 |
ACL (1) | 5 |
| 2022 | NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient FrameworkabstractPretrained language models have become the standard approach for many NLP tasks due to strong performance, but they are very expensive to train. We propose a simple and efficient learning framework, TLM, that does not rely on large-scale pretraining. Given some labeled task data and a large general corpus, TLM uses task data as queries to retrieve a tiny subset of the general corpus and jointly optimizes the task objective and the language modeling objective from scratch. On eight classification datasets in four domains, TLM achieves results better than or similar to pretrained language models (e.g., RoBERTa-Large) while reducing the training FLOPs by two orders of magnitude. With high accuracy and efficiency, we hope TLM will contribute to democratizing NLP and expediting its development. Xingcheng Yao, Yanan Zheng, Xiaocong Yang, Zhilin Yang 0001 |
ICML | 4 |
| 2021 | The International Workshop on Pretraining: Algorithms, Architectures, and Applications ([email protected] 2021)abstractThe International Workshop on Pretraining: Algorithms, Architectures, and Applications ([email protected] 2021) presents interdisciplinary contributions in pretraining. The workshop is related to machine learning, deep learning, representation learning, natural language processing, computer vision, graph learning, and knowledge discovery. The program of the workshop will focus on presenting and discussing the state-of-the-art, open problems, challenges and latest models, techniques and algorithms in the field of pretraining, covering aspects of algorithms, architectures and applications. Ming Ding 0004, Yuxiao Dong, Xiao Liu 0036, Jiezhong Qiu, Jie Tang 0001, Zhilin Yang 0001 |
KDD | 6 |
| 2021 | Controllable Generation from Pre-trained Language Models via Inverse PromptingabstractLarge-scale pre-trained language models have demonstrated strong capabilities of generating realistic texts. However, it remains challenging to control the generation results. Previous approaches such as prompting are far from sufficient, and lack of controllability limits the usage of language models. To tackle this challenge, we propose an innovative method, inverse prompting, to better control text generation. The core idea of inverse prompting is to use generated text to inversely predict the prompt during beam search, which enhances the relevance between the prompt and the generated text and thus improves controllability. Empirically, we pre-train a large-scale Chinese language model to perform a systematic study using human evaluation on the tasks of open-domain poem generation and open-domain long-form question answering. Results demonstrate that our proposed method substantially outperforms the baselines and that our generation quality is close to human performance on some of the tasks. Xu Zou 0001, Da Yin, Qingyang Zhong, Hongxia Yang, Zhilin Yang 0001, Jie Tang 0001 |
KDD | 5 |
| 2019 | Transformer-XL: Attentive Language Models beyond a Fixed-Length ContextabstractTransformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling.We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence.It consists of a segment-level recurrence mechanism and a novel positional encoding scheme.Our method not only enables capturing longer-term dependency, but also resolves the context fragmentation problem.As a result, Transformer-XL learns dependency that is 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800+ times faster than vanilla Transformers during evaluation.Notably, we improve the state-ofthe-art results of bpc/perplexity to 0.99 on en-wiki8, 1.08 on text8, 18.3 on WikiText-103, 21.8 on One Billion Word, and 54.5 on Penn Treebank (without finetuning).When trained only on WikiText-103, Transformer-XL manages to generate reasonably coherent, novel text articles with thousands of tokens.Our code, pretrained models, and hyperparameters are available in both Tensorflow and PyTorch 1 . Zihang Dai, Zhilin Yang 0001, Yiming Yang 0002, Jaime G. Carbonell, Quoc V. Le, Ruslan Salakhutdinov |
ACL (1) | 2 |
| 2019 | XLNet: Generalized Autoregressive Pretraining for Language UnderstandingabstractWith the capability of modeling bidirectional contexts, denoising autoencoding based pretraining like BERT achieves better performance than pretraining approaches based on autoregressive language modeling. However, relying on corrupting the input with masks, BERT neglects dependency between the masked positions and suffers from a pretrain-finetune discrepancy. In light of these pros and cons, we propose XLNet, a generalized autoregressive pretraining method that (1) enables learning bidirectional contexts by maximizing the expected likelihood over all permutations of the factorization order and (2) overcomes the limitations of BERT thanks to its autoregressive formulation. Furthermore, XLNet integrates ideas from Transformer-XL, the state-of-the-art autoregressive model, into pretraining. Empirically, under comparable experiment setting, XLNet outperforms BERT on 20 tasks, often by a large margin, including question answering, natural language inference, sentiment analysis, and document ranking. Zhilin Yang 0001, Zihang Dai, Yiming Yang 0002, Jaime G. Carbonell, Ruslan Salakhutdinov, Quoc V. Le |
NeurIPS | 1 |
| 2019 | Mixtape: Breaking the Softmax Bottleneck EfficientlyabstractThe softmax bottleneck has been shown to limit the expressiveness of neural lan- guage models. Mixture of Softmaxes (MoS) is an effective approach to address such a theoretical limitation, but are expensive compared to softmax in terms of both memory and time. We propose Mixtape, an output layer that breaks the softmax bottleneck more efficiently with three novel techniques—logit space vector gating, sigmoid tree decomposition, and gate sharing. On four benchmarks including language modeling and machine translation, the Mixtape layer substantially improves the efficiency over the MoS layer by 3.5x to 10.5x while obtaining similar performance. A network equipped with Mixtape is only 20% to 34% slower than a softmax-based network with 10-30K vocabulary sizes, and outperforms softmax in perplexity and translation quality. Zhilin Yang 0001, Thang Luong, Ruslan Salakhutdinov, Quoc V. Le |
NeurIPS | 1 |
| 2018 | Neural Cross-lingual Named Entity Recognition with Minimal ResourcesabstractFor languages with no annotated resources, unsupervised transfer of natural language processing models such as named-entity recognition (NER) from resource-rich languages would be an appealing capability.However, differences in words and word order across languages make it a challenging problem.To improve mapping of lexical items across languages, we propose a method that finds translations based on bilingual word embeddings.To improve robustness to word order differences, we propose to use self-attention, which allows for a degree of flexibility with respect to word order.We demonstrate that these methods achieve state-of-the-art or competitive NER performance on commonly tested languages under a cross-lingual setting, with much lower resource requirements than past approaches.We also evaluate the challenges of applying these methods to Uyghur, a lowresource language.1 Jiateng Xie, Zhilin Yang 0001, Graham Neubig, Noah A. Smith, Jaime G. Carbonell |
EMNLP | 2 |
| 2018 | HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question AnsweringabstractZhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, Christopher D. Manning. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. Zhilin Yang 0001, Peng Qi 0003, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, Christopher D. Manning |
EMNLP | 1 |
| 2018 | Breaking the Softmax Bottleneck: A High-Rank RNN Language Model
Zhilin Yang 0001, Zihang Dai, Ruslan Salakhutdinov, William W. Cohen |
ICLR | 1 |
| 2018 | Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent
Zhilin Yang 0001, Saizheng Zhang, Jack Urbanek, Will Feng, Alexander H. Miller, Arthur Szlam, Douwe Kiela, Jason Weston |
ICLR (Poster) | 1 |
| 2018 | GLoMo: Unsupervised Learning of Transferable Relational GraphsabstractModern deep transfer learning approaches have mainly focused on learning generic feature vectors from one task that are transferable to other tasks, such as word embeddings in language and pretrained convolutional features in vision. However, these approaches usually transfer unary features and largely ignore more structured graphical representations. This work explores the possibility of learning generic latent relational graphs that capture dependencies between pairs of data units (e.g., words or pixels) from large-scale unlabeled data and transferring the graphs to downstream tasks. Our proposed transfer learning framework improves performance on various tasks including question answering, natural language inference, sentiment analysis, and image classification. We also show that the learned graphs are generic enough to be transferred to different embeddings on which the graphs have not been trained (including GloVe embeddings, ELMo embeddings, and task-specific RNN hidden units), or embedding-free units such as image pixels. Zhilin Yang 0001, Junbo Jake Zhao, Bhuwan Dhingra, Kaiming He, William W. Cohen, Ruslan Salakhutdinov, Yann LeCun |
NeurIPS | 1 |
| 2017 | Gated-Attention Readers for Text ComprehensionabstractIn this paper we study the problem of answering cloze-style questions over documents.Our model, the Gated-Attention (GA) Reader 1 , integrates a multi-hop architecture with a novel attention mechanism, which is based on multiplicative interactions between the query embedding and the intermediate states of a recurrent neural network document reader.This enables the reader to build query-specific representations of tokens in the document for accurate answer selection.The GA Reader obtains state-of-the-art results on three benchmarks for this task-the CNN & Daily Mail news stories and the Who Did What dataset.The effectiveness of multiplicative interaction is demonstrated by an ablation study, and by comparing to alternative compositional operators for implementing the gated-attention. Bhuwan Dhingra, Hanxiao Liu, Zhilin Yang 0001, William W. Cohen, Ruslan Salakhutdinov |
ACL (1) | 3 |
| 2017 | Semi-Supervised QA with Generative Domain-Adaptive NetsabstractWe study the problem of semi-supervised question answering--utilizing unlabeled text to boost the performance of question answering models.We propose a novel training framework, the Generative Domain-Adaptive Nets.In this framework, we train a generative model to generate questions based on the unlabeled text, and combine model-generated questions with human-generated questions for training question answering models.We develop novel domain adaptation algorithms, based on reinforcement learning, to alleviate the discrepancy between the modelgenerated data distribution and the humangenerated data distribution.Experiments show that our proposed framework obtains substantial improvement from unlabeled text. Zhilin Yang 0001, Junjie Hu 0001, Ruslan Salakhutdinov, William W. Cohen |
ACL (1) | 1 |
| 2017 | Words or Characters? Fine-grained Gating for Reading Comprehension
Zhilin Yang 0001, Bhuwan Dhingra, Junjie Hu 0001, William W. Cohen, Ruslan Salakhutdinov |
ICLR (Poster) | 1 |
| 2017 | Transfer Learning for Sequence Tagging with Hierarchical Recurrent Networks
Zhilin Yang 0001, Ruslan Salakhutdinov, William W. Cohen |
ICLR (Poster) | 1 |
| 2017 | Good Semi-supervised Learning That Requires a Bad GANabstractSemi-supervised learning methods based on generative adversarial networks (GANs) obtained strong empirical results, but it is not clear 1) how the discriminator benefits from joint training with a generator, and 2) why good semi-supervised classification performance and a good generator cannot be obtained at the same time. Theoretically we show that given the discriminator objective, good semi-supervised learning indeed requires a bad generator, and propose the definition of a preferred generator. Empirically, we derive a novel formulation based on our analysis that substantially improves over feature matching GANs, obtaining state-of-the-art results on multiple benchmark datasets. Zihang Dai, Zhilin Yang 0001, Fan Yang 0058, William W. Cohen, Ruslan Salakhutdinov |
NIPS | 2 |
| 2017 | Differentiable Learning of Logical Rules for Knowledge Base ReasoningabstractWe study the problem of learning probabilistic first-order logical rules for knowledge base reasoning. This learning problem is difficult because it requires learning the parameters in a continuous space as well as the structure in a discrete space. We propose a framework, Neural Logic Programming, that combines the parameter and structure learning of first-order logical rules in an end-to-end differentiable model. This approach is inspired by a recently-developed differentiable logic called TensorLog [5], where inference tasks can be compiled into sequences of differentiable operations. We design a neural controller system that learns to compose these operations. Empirically, our method outperforms prior work on multiple knowledge base benchmark datasets, including Freebase and WikiMovies. Fan Yang 0058, Zhilin Yang 0001, William W. Cohen |
NIPS | 2 |
| 2016 | Revisiting Semi-Supervised Learning with Graph EmbeddingsabstractWe present a semi-supervised learning framework based on graph embeddings. Given a graph between instances, we train an embedding for each instance to jointly predict the class label and the neighborhood context in the graph. We develop both transductive and inductive variants of our method. In the transductive variant of our method, the class labels are determined by both the learned embeddings and input feature vectors, while in the inductive variant, the embeddings are defined as a parametric function of the feature vectors, so predictions can be made on instances not seen during training. On a large and diverse set of benchmark tasks, including text classification, distantly supervised entity extraction, and entity classification, we show improved performance over many of the existing models. Zhilin Yang 0001, William W. Cohen, Ruslan Salakhutdinov |
ICML | 1 |
| 2016 | Multi-Modal Bayesian Embeddings for Learning Social Knowledge Graphs
Zhilin Yang 0001, Jie Tang 0001, William W. Cohen |
IJCAI | 1 |
| 2016 | Review Networks for Caption GenerationabstractWe propose a novel extension of the encoder-decoder framework, called a review network. The review network is generic and can enhance any existing encoder- decoder model: in this paper, we consider RNN decoders with both CNN and RNN encoders. The review network performs a number of review steps with attention mechanism on the encoder hidden states, and outputs a thought vector after each review step; the thought vectors are used as the input of the attention mechanism in the decoder. We show that conventional encoder-decoders are a special case of our framework. Empirically, we show that our framework improves over state-of- the-art encoder-decoder systems on the tasks of image captioning and source code captioning. Zhilin Yang 0001, Yuexin Wu, William W. Cohen, Ruslan Salakhutdinov |
NIPS | 1 |
| 2015 | COSNET: Connecting Heterogeneous Social Networks with Local and Global ConsistencyabstractMore often than not, people are active in more than one social network. Identifying users from multiple heterogeneous social networks and integrating the different networks is a fundamental issue in many applications. The existing methods tackle this problem by estimating pairwise similarity between users in two networks. However, those methods suffer from potential inconsistency of matchings between multiple networks. Jie Tang 0001, Zhilin Yang 0001, Jian Pei 0001, Philip S. Yu |
KDD | 3 |
| 2014 | Active Learning for Streaming Networked DataabstractMining high-speed data streams has become an important topic due to the rapid growth of online data. In this paper, we study the problem of active learning for streaming networked data. The goal is to train an accurate model for classifying networked data that arrives in a streaming manner by querying as few labels as possible. The problem is extremely challenging, as both the data distribution and the network structure may change over time. The query decision has to be made for each data instance sequentially, by considering the dynamic network structure. Zhilin Yang 0001, Jie Tang 0001 |
CIKM | 1 |
| 2014 | Active learning for networked data based on non-progressive diffusion modelabstractWe study the problem of active learning for networked data, where samples are connected with links and their labels are correlated with each other. We particularly focus on the setting of using the probabilistic graphical model to model the networked data, due to its effectiveness in capturing the dependency between labels of linked samples. We propose a novel idea of connecting the graphical model to the information diffusion process, and precisely define the active learning problem based on the non-progressive diffusion model. We show the NP-hardness of the problem and propose a method called MaxCo to solve it. We derive the lower bound for the optimal solution for the active learning setting, and develop an iterative greedy algorithm with provable approximation guarantees. We also theoretically prove the convergence and correctness of MaxCo. Zhilin Yang 0001, Jie Tang 0001, Bin Xu 0001, Chunxiao Xing |
WSDM | 1 |
| 2013 | SAE: social analytic engine for large networksabstractOnline social networks become a bridge to connect our physical daily life and the virtual Web space, which not only provides rich data for mining, but also brings many new challenges. In this paper, we present a novel Social Analytic Engine (SAE) for large online social networks. The key issues we pursue in the analytic engine are concerned with the following problems: 1) at the micro-level, how do people form different types of social ties and how people influence each other? 2) at the meso-level, how do people group into communities? 3) at the macro-level, what are the hottest topics in a social network and how the topics evolve over time? Yang Yang 0009, Wei Chen 0013, Jing Zhang 0001, Honglei Zhuang, Zhilin Yang 0001, Zhanpeng Fang, Sen Wu 0001, Debing Liu, Jie Tang 0001 |
KDD | 7 |