Tianyu Gao 0001

dblp:207/8893-1 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0002-5178-0866ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
YearPublicationVenuePosition
2025 How to Train Long-Context Language Models (Effectively)
abstract
We study continued training and supervised fine-tuning (SFT) of a language model (LM) to make effective use of long-context information.We first establish a reliable evaluation protocol to guide model development-instead of perplexity or simple needle-in-a-haystack (NIAH) tests, we use a broad set of long-context downstream tasks, and we evaluate models after SFT as this better reveals long-context abilities.Supported by our robust evaluations, we run thorough experiments to decide the data mix for continued pre-training, the instruction tuning dataset, and many other design choices such as position extrapolation.We find that (1) code repositories and books are excellent sources of long data, but it is crucial to combine them with high-quality short-context data;(2) training with a sequence length beyond the evaluation length boosts long-context performance; (3) for SFT, using only short instruction datasets yields strong performance on long-context tasks.Our final model, ProLong-8B, which is initialized from Llama-3 and trained on 40B tokens, demonstrates state-ofthe-art long-context performance among similarly sized models at a length of 128K.Pro-Long outperforms Llama-3.1-8B-Instruct on the majority of long-context tasks despite using only 5% as many tokens during long-context training.Additionally, ProLong can effectively process up to 512K tokens, one of the longest context windows of publicly available LMs.
Tianyu Gao 0001, Alexander Wettig, Howard Yen, Danqi Chen 0001
ACL (1)1
2025 HELMET: How to Evaluate Long-context Models Effectively and Thoroughly
abstract
Many benchmarks exist for evaluating long-context language models (LCLMs), yet developers often rely on synthetic tasks such as needle-in-a-haystack (NIAH) or an arbitrary subset of tasks. However, it remains unclear whether these benchmarks reflect the diverse downstream applications of LCLMs, and such inconsistencies further complicate model comparison. We investigate the underlying reasons behind these practices and find that existing benchmarks often provide noisy signals due to limited coverage of applications, insufficient context lengths, unreliable metrics, and incompatibility with base models. In this work, we introduce HELMET (How to Evaluate Long-context Models Effectively and Thoroughly), a comprehensive benchmark encompassing seven diverse, application-centric categories. We also address several issues in previous benchmarks by adding controllable lengths up to 128K tokens, model-based evaluation for reliable metrics, and few-shot prompting for robustly evaluating base models. Consequently, we demonstrate that HELMET offers more reliable and consistent rankings of frontier LCLMs. Through a comprehensive study of 59 LCLMs, we find that (1) synthetic tasks like NIAH do not reliably predict downstream performance; (2) the diverse categories in HELMET exhibit distinct trends and low correlations with each other; and (3) while most LCLMs achieve perfect NIAH scores, open-source models significantly lag behind closed ones when tasks require full-context reasoning or following complex instructions---the gap widens as length increases. Finally, we recommend using our RAG tasks for fast model development, as they are easy to run and better predict other downstream performance; ultimately, we advocate for a holistic evaluation across diverse tasks.
Howard Yen, Tianyu Gao 0001, Minmin Hou, Daniel Fleischer, Peter Izsak, Moshe Wasserblat, Danqi Chen 0001
ICLR2
2025 Metadata Conditioning Accelerates Language Model Pre-training
abstract
The vast diversity of styles, domains, and quality levels present in language model pre-training corpora is essential in developing general model capabilities, but efficiently learning and deploying the correct behaviors exemplified in each of these heterogeneous data sources is challenging. To address this, we propose a new method, termed Metadata Conditioning then Cooldown (MeCo), to incorporate additional learning cues during pre-training. MeCo first provides metadata (e.g., URLs like en.wikipedia.org) alongside the text during training and later uses a cooldown phase with only the standard text, thereby enabling the model to function normally even without metadata. MeCo significantly accelerates pre-training across different model scales (600M to 8B parameters) and training sources (C4, RefinedWeb, and DCLM). For instance, a 1.6B language model trained with MeCo matches the downstream task performance of standard pre-training while using 33% less data. Additionally, MeCo enables us to steer language models by conditioning the inference prompt on either real or fabricated metadata that encodes the desired properties of the output: for example, prepending wikipedia.org to reduce harmful generations or factquizmaster.com (fabricated) to improve common knowledge task performance. We also demonstrate that MeCo is compatible with different types of metadata, such as model-generated topics. MeCo is remarkably simple, adds no computational overhead, and demonstrates promise in producing more capable and steerable language models.
Tianyu Gao 0001, Alexander Wettig, Luxi He, Yihe Dong, Sadhika Malladi, Danqi Chen 0001
ICML1
2024 Long-Context Language Modeling with Parallel Context Encoding
abstract
Extending large language models (LLMs) to process longer inputs is crucial for a wide range of applications.However, the substantial computational cost of transformers and limited generalization of positional encoding restrict the size of their context window.We introduce Context Expansion with Parallel Encoding (CEPE ), a framework that can be applied to any existing decoder-only LLMs to extend their context window.CEPE employs a small encoder to process long inputs chunk by chunk, enabling the frozen decoder to utilize additional contexts via cross-attention.CEPE is efficient, generalizable, and versatile: trained with 8K-token documents, it extends the context window of LLAMA-2 to 128K tokens, offering 10× the throughput with only 1/6 of the memory.CEPE yields strong performance on language modeling and in-context learning.CEPE also excels in retrieval-augmented applications, while existing long-context models degenerate with retrieved contexts.We further introduce a CEPE variant that can extend the context window of instruction-tuned models using only unlabeled data, and showcase its effectiveness on LLAMA-2-CHAT, leading to a strong instruction-following model that can leverage very long contexts on downstream tasks. 1Chapter 01: Dune ...
Howard Yen, Tianyu Gao 0001, Danqi Chen 0001
ACL (1)2
2024 LitSearch: A Retrieval Benchmark for Scientific Literature Search
abstract
Literature search questions, such as "Where can I find research on the evaluation of consistency in generated summaries?"pose significant challenges for modern search engines and retrieval systems.These questions often require a deep understanding of research concepts and the ability to reason across entire articles.In this work, we introduce LitSearch, a retrieval benchmark comprising 597 realistic literature search queries about recent ML and NLP papers.Lit-Search is constructed using a combination of (1) questions generated by GPT-4 based on paragraphs containing inline citations from research papers and (2) questions manually written by authors about their recently published papers.All LitSearch questions were manually examined or edited by experts to ensure high quality.We extensively benchmark state-ofthe-art retrieval models and also evaluate two LLM-based reranking pipelines.We find a significant performance gap between BM25 and state-of-the-art dense retrievers, with a 24.8% absolute difference in [email protected] LLMbased reranking strategies further improve the best-performing dense retriever by 4.4%.Additionally, commercial search engines and research tools like Google Search perform poorly on LitSearch, lagging behind the best dense retriever by up to 32 recall points.Taken together, these results show that LitSearch is an informative new testbed for retrieval systems while catering to a real-world use case.
Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen 0001, Tianyu Gao 0001
EMNLP6
2024 Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
abstract
The popularity of LLaMA (Touvron et al., 2023a;b) and other recently emerged moderate-sized large language models (LLMs) highlights the potential of building smaller yet powerful LLMs. Regardless, the cost of training such models from scratch on trillions of tokens remains high. In this work, we study structured pruning as an effective means to develop smaller LLMs from pre-trained, larger models. Our approach employs two key techniques: (1) targeted structured pruning, which prunes a larger model to a specified target shape by removing layers, heads, and intermediate and hidden dimensions in an end-to-end manner, and (2) dynamic batch loading, which dynamically updates the composition of sampled data in each training batch based on varying losses across different domains. We demonstrate the efficacy of our approach by presenting the Sheared-LLaMA series, pruning the LLaMA2-7B model down to 1.3B and 2.7B parameters. Sheared-LLaMA models outperform state-of-the-art open-source models of equivalent sizes, such as Pythia, INCITE, OpenLLaMA and the concurrent TinyLlama models, on a wide range of downstream and instruction tuning evaluations, while requiring only 3% of compute compared to training such models from scratch. This work provides compelling evidence that leveraging existing LLMs with structured pruning is a far more cost-effective approach for building competitive small-scale LLMs
Mengzhou Xia, Tianyu Gao 0001, Danqi Chen 0001
ICLR2
2024 Evaluating Large Language Models at Evaluating Instruction Following
abstract
As research in large language models (LLMs) continues to accelerate, LLM-based evaluation has emerged as a scalable and cost-effective alternative to human evaluations for comparing the ever increasing list of models. This paper investigates the efficacy of these “LLM evaluators”, particularly in using them to assess instruction following, a metric that gauges how closely generated text adheres to the given instruction. We introduce a challenging meta-evaluation benchmark, LLMBar, designed to test the ability of an LLM evaluator in discerning instruction-following outputs. The authors manually curated 419 pairs of outputs, one adhering to instructions while the other diverging, yet may possess deceptive qualities that mislead an LLM evaluator, e.g., a more engaging tone. Contrary to existing meta-evaluation, we discover that different evaluators (i.e., combinations of LLMs and prompts) exhibit distinct performance on LLMBar and even the highest-scoring ones have substantial room for improvement. We also present a novel suite of prompting strategies that further close the gap between LLM and human evaluators. With LLMBar, we hope to offer more insight into LLM evaluators and foster future research in developing better instruction-following models.
Jiatong Yu, Tianyu Gao 0001, Yu Meng 0001, Tanya Goyal, Danqi Chen 0001
ICLR3
2023 Should You Mask 15% in Masked Language Modeling?
abstract
Masked language models (MLMs) conventionally mask 15% of tokens due to the belief that more masking would leave insufficient context to learn good representations; this masking rate has been widely used, regardless of model sizes or masking strategies.In this work, we revisit this important choice of MLM pre-training.We first establish that 15% is not universally optimal, and larger models should adopt a higher masking rate.Specifically, we find that masking 40% outperforms 15% for BERT-large size models on GLUE and SQuAD.Interestingly, an extremely high masking rate of 80% can still preserve 95% fine-tuning performance and most of the accuracy in linguistic probing, challenging the conventional wisdom about the role of the masking rate.We then examine the interplay between masking rates and masking strategies and find that uniform masking requires a higher masking rate compared to sophisticated masking strategies such as span or PMI masking.Finally, we argue that increasing the masking rate has two distinct effects: it leads to more corruption, which makes the prediction task harder; it also enables more predictions, which benefits optimization.Using this framework, we revisit BERT's 80-10-10 corruption strategy.Together, our results contribute to a better understanding of MLM pre-training. 1
Alexander Wettig, Tianyu Gao 0001, Zexuan Zhong, Danqi Chen 0001
EACL2
2023 Enabling Large Language Models to Generate Text with Citations
abstract
Large language models (LLMs) have emerged as a widely-used tool for information seeking, but their generated outputs are prone to hallucination.In this work, our aim is to allow LLMs to generate text with citations, improving their factual correctness and verifiability.Existing work mainly relies on commercial search engines and human evaluation, making it challenging to reproduce and compare different modeling approaches.We propose ALCE, the first benchmark for Automatic LLMs' Citation Evaluation.ALCE collects a diverse set of questions and retrieval corpora and requires building end-to-end systems to retrieve supporting evidence and generate answers with citations.We develop automatic metrics along three dimensions-fluency, correctness, and citation quality-and demonstrate their strong correlation with human judgements.Our experiments with state-of-the-art LLMs and novel prompting strategies show that current systems have considerable room for improvement-For example, on the ELI5 dataset, even the best models lack complete citation support 50% of the time.Our analyses further highlight promising future directions, including developing better retrievers, advancing long-context LLMs, and improving the ability to synthesize information from multiple sources. 1 When did the US break away from England? Question Short answers (from the dataset)
Tianyu Gao 0001, Howard Yen, Jiatong Yu, Danqi Chen 0001
EMNLP1
2023 Fine-Tuning Language Models with Just Forward Passes
abstract
Fine-tuning language models (LMs) has yielded success on diverse downstream tasks, but as LMs grow in size, backpropagation requires a prohibitively large amount of memory. Zeroth-order (ZO) methods can in principle estimate gradients using only two forward passes but are theorized to be catastrophically slow for optimizing large models. In this work, we propose a memory-efficient zerothorder optimizer (MeZO), adapting the classical ZO-SGD method to operate in-place, thereby fine-tuning LMs with the same memory footprint as inference. For example, with a single A100 80GB GPU, MeZO can train a 30-billion parameter model, whereas fine-tuning with backpropagation can train only a 2.7B LM with the same budget. We conduct comprehensive experiments across model types (masked and autoregressive LMs), model scales (up to 66B), and downstream tasks (classification, multiple-choice, and generation). Our results demonstrate that (1) MeZO significantly outperforms in-context learning and linear probing; (2) MeZO achieves comparable performance to fine-tuning with backpropagation across multiple tasks, with up to 12× memory reduction and up to 2× GPU-hour reduction in our implementation; (3) MeZO is compatible with both full-parameter and parameter-efficient tuning techniques such as LoRA and prefix tuning; (4) MeZO can effectively optimize non-differentiable objectives (e.g., maximizing accuracy or F1). We support our empirical findings with theoretical insights, highlighting how adequate pre-training and task prompts enable MeZO to fine-tune huge models, despite classical ZO analyses suggesting otherwise.
Sadhika Malladi, Tianyu Gao 0001, Eshaan Nichani, Alex Damian, Jason D. Lee, Danqi Chen 0001, Sanjeev Arora
NeurIPS2
2022 Ditch the Gold Standard: Re-evaluating Conversational Question Answering
abstract
Conversational question answering aims to provide natural-language answers to users in information-seeking conversations. Existing conversational QA benchmarks compare models with pre-collected human-human conversations, using ground-truth answers provided in conversational history. It remains unclear whether we can rely on this static evaluation for model development and whether current systems can well generalize to real-world human-machine conversations. In this work, we conduct the first large-scale human evaluation of state-of-the-art conversational QA systems, where human evaluators converse with models and judge the correctness of their answers. We find that the distribution of human machine conversations differs drastically from that of human-human conversations, and there is a disagreement between human and gold-history evaluation in terms of model ranking. We further investigate how to improve automatic evaluations, and propose a question rewriting mechanism based on predicted history, which better correlates with human judgments. Finally, we analyze the impact of various modeling strategies and discuss future directions towards building better conversational question answering systems.
Huihan Li 0001, Tianyu Gao 0001, Manan Goenka, Danqi Chen 0001
ACL (1)2
2022 Automatic Label Sequence Generation for Prompting Sequence-to-sequence Models
abstract
Prompting, which casts downstream applications as language modeling tasks, has shown to be sample efficient compared to standard fine-tuning with pre-trained models. However, one pitfall of prompting is the need of manually-designed patterns, whose outcome can be unintuitive and requires large validation sets to tune. To tackle the challenge, we propose AutoSeq, a fully automatic prompting method: (1) We adopt natural language prompts on sequence-to-sequence models, enabling free-form generation and larger label search space; (2) We propose label sequences – phrases with indefinite lengths to verbalize the labels – which eliminate the need of manual templates and are more expressive than single label words; (3) We use beam search to automatically generate a large amount of label sequence candidates and propose contrastive re-ranking to get the best combinations. AutoSeq significantly outperforms other no-manual-design methods, such as soft prompt tuning, adapter tuning, and automatic search on single label words; the generated label sequences are even better than curated manual ones on a variety of tasks. Our method reveals the potential of sequence-to-sequence models in few-shot learning and sheds light on a path to generic and automatic prompting. The source code of this paper can be obtained from https://github.com/thunlp/Seq2Seq-Prompt.
Zichun Yu, Tianyu Gao 0001, Zhengyan Zhang, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016
COLING2
2022 Recovering Private Text in Federated Learning of Language Models
abstract
Federated learning allows distributed users to collaboratively train a model while keeping each user’s data private. Recently, a growing body of work has demonstrated that an eavesdropping attacker can effectively recover image data from gradients transmitted during federated learning. However, little progress has been made in recovering text data. In this paper, we present a novel attack method FILM for federated learning of language models (LMs). For the first time, we show the feasibility of recovering text from large batch sizes of up to 128 sentences. Unlike image-recovery methods that are optimized to match gradients, we take a distinct approach that first identifies a set of words from gradients and then directly reconstructs sentences based on beam search and a prior-based reordering strategy. We conduct the FILM attack on several large-scale datasets and show that it can successfully reconstruct single sentences with high fidelity for large batch sizes and even multiple sentences if applied iteratively.We evaluate three defense methods: gradient pruning, DPSGD, and a simple approach to freeze word embeddings that we propose. We show that both gradient pruning and DPSGD lead to a significant drop in utility. However, if we fine-tune a public pre-trained LM on private text without updating word embeddings, it can effectively defend the attack with minimal data utility loss. Together, we hope that our results can encourage the community to rethink the privacy concerns of LM training and its standard practices in the future. Our code is publicly available at https://github.com/Princeton-SysML/FILM .
Samyak Gupta, Yangsibo Huang, Zexuan Zhong, Tianyu Gao 0001, Kai Li 0001, Danqi Chen 0001
NeurIPS4
2021 Making Pre-trained Language Models Better Few-shot Learners
abstract
Tianyu Gao, Adam Fisch, Danqi Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Tianyu Gao 0001, Adam Fisch, Danqi Chen 0001
ACL/IJCNLP (1)1
2021 SimCSE: Simple Contrastive Learning of Sentence Embeddings
abstract
This paper presents SimCSE, a simple contrastive learning framework that greatly advances the state-of-the-art sentence embeddings.We first describe an unsupervised approach, which takes an input sentence and predicts itself in a contrastive objective, with only standard dropout used as noise.This simple method works surprisingly well, performing on par with previous supervised counterparts.We find that dropout acts as minimal data augmentation and removing it leads to a representation collapse.Then, we propose a supervised approach, which incorporates annotated pairs from natural language inference datasets into our contrastive learning framework, by using "entailment" pairs as positives and "contradiction" pairs as hard negatives.We evaluate SimCSE on standard semantic textual similarity (STS) tasks, and our unsupervised and supervised models using BERT base achieve an average of 76.3% and 81.6% Spearman's correlation respectively, a 4.2% and 2.2% improvement compared to previous best results.We also show-both theoretically and empirically-that contrastive learning objective regularizes pre-trained embeddings' anisotropic space to be more uniform, and it better aligns positive pairs when supervised signals are available.1
Tianyu Gao 0001, Xingcheng Yao, Danqi Chen 0001
EMNLP (1)1
2021 KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation
abstract
Abstract Pre-trained language representation models (PLMs) cannot well capture factual knowledge from text. In contrast, knowledge embedding (KE) methods can effectively represent the relational facts in knowledge graphs (KGs) with informative entity embeddings, but conventional KE models cannot take full advantage of the abundant textual information. In this paper, we propose a unified model for Knowledge Embedding and Pre-trained LanguagERepresentation (KEPLER), which can not only better integrate factual knowledge into PLMs but also produce effective text-enhanced KE with the strong PLMs. In KEPLER, we encode textual entity descriptions with a PLM as their embeddings, and then jointly optimize the KE and language modeling objectives. Experimental results show that KEPLER achieves state-of-the-art performances on various NLP tasks, and also works remarkably well as an inductive KE model on KG link prediction. Furthermore, for pre-training and evaluating KEPLER, we construct Wikidata5M1 , a large-scale KG dataset with aligned entity descriptions, and benchmark state-of-the-art KE methods on it. It shall serve as a new KE benchmark and facilitate the research on large KG, inductive KE, and KG with text. The source code can be obtained from https://github.com/THU-KEG/KEPLER.
Xiaozhi Wang, Tianyu Gao 0001, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu 0001, Juan-Zi Li, Jian Tang 0005
Trans. Assoc. Comput. Linguistics2
2020 Neural Snowball for Few-Shot Relation Learning
abstract
Knowledge graphs typically undergo open-ended growth of new relations. This cannot be well handled by relation extraction that focuses on pre-defined relations with sufficient training data. To address new relations with few-shot instances, we propose a novel bootstrapping approach, Neural Snowball, to learn new relations by transferring semantic knowledge about existing relations. More specifically, we use Relational Siamese Networks (RSN) to learn the metric of relational similarities between instances based on existing relations and their labeled data. Afterwards, given a new relation and its few-shot instances, we use RSN to accumulate reliable instances from unlabeled corpora; these instances are used to train a relation classifier, which can further identify new facts of the new relation. The process is conducted iteratively like a snowball. Experiments show that our model can gather high-quality instances for better few-shot relation learning and achieves significant improvement compared to baselines. Codes and datasets are released on https://github.com/thunlp/Neural-Snowball.
Tianyu Gao 0001, Xu Han 0007, Ruobing Xie, Zhiyuan Liu 0001, Fen Lin 0002, Leyu Lin, Maosong Sun 0001
AAAI1
2020 Continual Relation Learning via Episodic Memory Activation and Reconsolidation
abstract
Continual relation learning aims to continually train a model on new data to learn incessantly emerging novel relations while avoiding catastrophically forgetting old relations.Some pioneering work has proved that storing a handful of historical relation examples in episodic memory and replaying them in subsequent training is an effective solution for such a challenging problem.However, these memorybased methods usually suffer from overfitting the few memorized examples of old relations, which may gradually cause inevitable confusion among existing relations.Inspired by the mechanism in human long-term memory formation, we introduce episodic memory activation and reconsolidation (EMAR) to continual relation learning.Every time neural models are activated to learn both new and memorized data, EMAR utilizes relation prototypes for memory reconsolidation exercise to keep a stable understanding of old relations.The experimental results show that EMAR could get rid of catastrophically forgetting old relations and outperform the state-of-the-art continual learning models.The code and datasets are released on https://github.com/thunlp/ ContinualRE.
Xu Han 0007, Tianyu Gao 0001, Yankai Lin 0001, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
ACL3
2020 Meta-Information Guided Meta-Learning for Few-Shot Relation Classification
abstract
Few-shot classification requires classifiers to adapt to new classes with only a few training instances.State-of-the-art meta-learning approaches such as MAML learn how to initialize and fast adapt parameters from limited instances, which have shown promising results in few-shot classification.However, existing meta-learning models solely rely on implicit instance-based statistics, and thus suffer from instance unreliability and weak interpretability.To solve this problem, we propose a novel meta-information guided meta-learning (MIML) framework, where semantic concepts of classes provide strong guidance for meta-learning in both initialization and adaptation.In effect, our model can establish connections between instance-based information and semantic-based information, which enables more effective initialization and faster adaptation.Comprehensive experimental results on few-shot relation classification demonstrate the effectiveness of the proposed framework.Notably, MIML achieves comparable or superior performance to humans with only one shot on FewRel evaluation.The source code and experiment details of this paper can be obtained from https://github.com/thunlp/MIML.
Bowen Dong 0005, Yuan Yao 0013, Ruobing Xie, Tianyu Gao 0001, Xu Han 0007, Zhiyuan Liu 0001, Fen Lin 0002, Leyu Lin, Maosong Sun 0001
COLING4
2020 Learning from Context or Names? An Empirical Study on Neural Relation Extraction
abstract
Neural models have achieved remarkable success on relation extraction (RE) benchmarks.However, there is no clear understanding which type of information affects existing RE models to make decisions and how to further improve the performance of these models.To this end, we empirically study the effect of two main information sources in text: textual context and entity mentions (names).We find that (i) while context is the main source to support the predictions, RE models also heavily rely on the information from entity mentions, most of which is type information, and (ii) existing datasets may leak shallow heuristics via entity mentions and thus contribute to the high performance on RE benchmarks.Based on the analyses, we propose an entity-masked contrastive pre-training framework for RE to gain a deeper understanding on both textual context and type information while avoiding rote memorization of entities or use of superficial cues in mentions.We carry out extensive experiments to support our views, and show that our framework can improve the effectiveness and robustness of neural models in different RE scenarios.All the code and datasets are released at https://github.com/thunlp/
Hao Peng 0015, Tianyu Gao 0001, Xu Han 0007, Yankai Lin 0001, Peng Li 0030, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016
EMNLP (1)2
2020 Few-shot Relation Extraction via Bayesian Meta-learning on Relation Graphs
abstract
This paper studies few-shot relation extraction, which aims at predicting the relation for a pair of entities in a sentence by training with a few labeled examples in each relation. To more effectively generalize to new relations, in this paper we study the relationships between different relations and propose to leverage a global relation graph. We propose a novel Bayesian meta-learning approach to effectively learn the posterior distribution of the prototype vectors of relations, where the initial prior of the prototype vectors is parameterized with a graph neural network on the global relation graph. Moreover, to effectively optimize the posterior distribution of the prototype vectors, we propose to use the stochastic gradient Langevin dynamics, which is related to the MAML algorithm but is able to handle the uncertainty of the prototype vectors. The whole framework can be effectively and efficiently optimized in an end-to-end fashion. Experiments on two benchmark datasets prove the effectiveness of our proposed approach against competitive baselines in both the few-shot and zero-shot settings.
Meng Qu, Tianyu Gao 0001, Louis-Pascal A. C. Xhonneux, Jian Tang 0005
ICML2
2019 Hybrid Attention-Based Prototypical Networks for Noisy Few-Shot Relation Classification
abstract
The existing methods for relation classification (RC) primarily rely on distant supervision (DS) because large-scale supervised training datasets are not readily available. Although DS automatically annotates adequate amounts of data for model training, the coverage of this data is still quite limited, and meanwhile many long-tail relations still suffer from data sparsity. Intuitively, people can grasp new knowledge by learning few instances. We thus provide a different view on RC by formalizing RC as a few-shot learning (FSL) problem. However, the current FSL models mainly focus on low-noise vision tasks, which makes them hard to directly deal with the diversity and noise of text. In this paper, we propose hybrid attention-based prototypical networks for the problem of noisy few-shot RC. We design instancelevel and feature-level attention schemes based on prototypical networks to highlight the crucial instances and features respectively, which significantly enhances the performance and robustness of RC models in a noisy FSL scenario. Besides, our attention schemes accelerate the convergence speed of RC models. Experimental results demonstrate that our hybrid attention-based models require fewer training iterations and outperform the state-of-the-art baseline models. The code and datasets are released on https://github.com/thunlp/ HATT-Proto.
Tianyu Gao 0001, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001
AAAI1
2019 FewRel 2.0: Towards More Challenging Few-Shot Relation Classification
abstract
Tianyu Gao, Xu Han, Hao Zhu, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tianyu Gao 0001, Xu Han 0007, Hao Zhu 0006, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016
EMNLP/IJCNLP (1)1