Timo Schick

dblp:203/9176 · DBLP profile ↗
← Back
16ranked-venue papers
12as first author
12since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 12 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
YearPublicationVenuePosition
2024 Self-Alignment with Instruction Backtranslation
abstract
We present a scalable method to build a high quality instruction following language model by automatically labelling human-written text with corresponding instructions. Our approach, named instruction backtranslation, starts with a language model finetuned on a small amount of seed data, and a given web corpus. The seed model is used to construct training examples by generating instruction prompts for web documents (self-augmentation), and then selecting high quality examples from among these candidates (self-curation). This data is then used to finetune a stronger model. Finetuning LLaMa on two iterations of our approach yields a model that outperforms all other LLaMa-based models on the Alpaca leaderboard not relying on distillation data, demonstrating highly effective self-alignment.
Xian Li 0003, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason Weston, Mike Lewis
ICLR4
2023 Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor
abstract
Instruction tuning enables pretrained language models to perform new tasks from inferencetime natural language descriptions.These approaches rely on vast amounts of human supervision in the form of crowdsourced datasets or user interactions.In this work, we introduce Unnatural Instructions: a large dataset of creative and diverse instructions, collected with virtually no human labor.We collect 64,000 examples by prompting a language model with three seed examples of instructions and eliciting a fourth.This set is then expanded by prompting the model to rephrase each instruction, creating a total of approximately 240,000 examples of instructions, inputs, and outputs.Experiments show that despite containing a fair amount of noise, training on Unnatural Instructions rivals the effectiveness of training on open-source manually-curated datasets, surpassing the performance of models such as T0++ and Tk-Instruct across various benchmarks.These results demonstrate the potential of model-generated data as a cost-effective alternative to crowdsourcing for dataset expansion and diversification. Example 1Instruction: You are given a science question (easy-level) and four answer options (associated with "A", "B", "C", "D").Your task is to find the correct answer based on scientific facts, knowledge, and reasoning.Do not generate anything else apart from one of the following characters: 'A', 'B, 'C', 'D'.There is only one correct answer for each question. Input: Which part of a bicycle BEST moves in a circle? (A) Seat (B) Frame (C) Foot pedal (D) KickstandConstraints: The output should be one of the following characters: 'A', 'B, 'C', 'D'. Example 2Instruction: You are given a negative review and your task is to convert it to a positive review by one or more making minimal changes.Avoid changing the context of the review.Input: we stood there in shock, because we never expected this. Constraints: None.Example 3 Instruction: In this task, you are given two sentences taken from a conversation, and your job is to classify whether these given sentences are sequential or not.We will mark the given sentence pair as 'True' if it's sequential, otherwise 'False'.The two sentences are spoken by two different people.
Or Honovich, Thomas Scialom, Omer Levy, Timo Schick
ACL (1)4
2023 PEER: A Collaborative Language Model
Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick S. H. Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, Sebastian Riedel 0001
ICLR1
2023 Toolformer: Language Models Can Teach Themselves to Use Tools
abstract
Language models (LMs) exhibit remarkable abilities to solve new tasks from just a few examples or textual instructions, especially at scale. They also, paradoxically, struggle with basic functionality, such as arithmetic or factual lookup, where much simpler and smaller specialized models excel. In this paper, we show that LMs can teach themselves to *use external tools* via simple APIs and achieve the best of both worlds. We introduce *Toolformer*, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction. This is done in a self-supervised way, requiring nothing more than a handful of demonstrations for each API. We incorporate a range of tools, including a calculator, a Q&A system, a search engine, a translation system, and a calendar. Toolformer achieves substantially improved zero-shot performance across a variety of downstream tasks, often competitive with much larger models, without sacrificing its core language modeling abilities.
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, Thomas Scialom
NeurIPS1
2023 Atlas: Few-shot Learning with Retrieval Augmented Language Models
abstract
Large language models have shown impressive few-shot results on a wide range of tasks. However, when knowledge is key for such results, as is the case for tasks such as question answering and fact checking, massive parameter counts to store knowledge seem to be needed. Retrieval-augmented models are known to excel at knowledge intensive tasks without the need for as many parameters, but it is unclear whether they work in few-shot settings. In this work we present Atlas, a carefully designed and pre-trained retrieval-augmented language model able to learn knowledge intensive tasks with very few training examples. We perform evaluations on a wide range of tasks, including MMLU, KILT and Natural Questions, and study the impact of the content of the document index, showing that it can easily be updated. Notably, Atlas reaches over 42% accuracy on Natural Questions using only 64 examples, outperforming a 540B parameter model by 3% despite having 50x fewer parameters.
Gautier Izacard, Patrick S. H. Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel 0001, Edouard Grave
J. Mach. Learn. Res.6
2022 Leveraging QA Datasets to Improve Generative Data Augmentation
abstract
The ability of generative language models (GLMs) to generate text has improved considerably in the last few years, enabling their use for generative data augmentation.In this work, we propose CONDA, an approach to further improve GLMs' ability to generate synthetic data by reformulating data generation as context generation for a given question-answer (QA) pair and leveraging QA datasets for training context generators.Then, we cast downstream tasks into the same question answering format and adapt the fine-tuned context generators to the target task domain.Finally, we use the fine-tuned GLM to generate relevant contexts, which are in turn used as synthetic training data for their corresponding tasks.We perform extensive experiments on multiple classification datasets and demonstrate substantial improvements in performance for both few-and zeroshot settings.Our analysis reveals that QA datasets that require high-level reasoning abilities (e.g., abstractive and common-sense QA datasets) tend to give the best boost in performance in both few-shot and zero-shot settings.
Dheeraj Mekala, Tu Vu, Timo Schick, Jingbo Shang
EMNLP3
2022 True Few-Shot Learning With Prompts - A Real-World Perspective
abstract
Abstract Prompt-based approaches excel at few-shot learning. However, Perez et al. (2021) recently cast doubt on their performance as they had difficulty getting good results in a “true” few-shot setting in which prompts and hyperparameters cannot be tuned on a dev set. In view of this, we conduct an extensive study of Pet, a method that combines textual instructions with example-based finetuning. We show that, if correctly configured, Pet performs strongly in true few-shot settings without a dev set. Crucial for this strong performance is a number of design choices, including Pet’s ability to intelligently handle multiple prompts. We put our findings to a real-world test by running Pet on RAFT, a benchmark of tasks taken from realistic NLP applications for which no labeled dev or test sets are available. Pet achieves a new state of the art on RAFT and performs close to non-expert humans for 7 out of 11 tasks. These results demonstrate that prompt-based learners can successfully be applied in true few-shot settings and underpin our belief that learning from instructions will play an important role on the path towards human-like few-shot learning capabilities.
Timo Schick, Hinrich Schütze
Trans. Assoc. Comput. Linguistics1
2021 Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference
abstract
Some NLP tasks can be solved in a fully unsupervised fashion by providing a pretrained language model with "task descriptions" in natural language (e.g., Radford et al., 2019).While this approach underperforms its supervised counterpart, we show in this work that the two ideas can be combined: We introduce Pattern-Exploiting Training (PET), a semi-supervised training procedure that reformulates input examples as cloze-style phrases to help language models understand a given task.These phrases are then used to assign soft labels to a large set of unlabeled examples.Finally, standard supervised training is performed on the resulting training set.For several tasks and languages, PET outperforms supervised training and strong semi-supervised approaches in lowresource settings by a large margin. 1
Timo Schick, Hinrich Schütze
EACL1
2021 Few-Shot Text Generation with Natural Language Instructions
abstract
Providing pretrained language models with simple task descriptions in natural language enables them to solve some tasks in a fully unsupervised fashion.Moreover, when combined with regular learning from examples, this idea yields impressive few-shot results for a wide range of text classification tasks.It is also a promising direction to improve data efficiency in generative settings, but there are several challenges to using a combination of task descriptions and example-based learning for text generation.In particular, it is crucial to find task descriptions that are easy to understand for the pretrained model and to ensure that it actually makes good use of them; furthermore, effective measures against overfitting have to be implemented.In this paper, we show how these challenges can be tackled: We introduce GENPET, a method for text generation that is based on pattern-exploiting training, a recent approach for combining textual instructions with supervised learning that only works for classification tasks.On several summarization and headline generation datasets, GENPET gives consistent improvements over strong baselines in few-shot settings.1
Timo Schick, Hinrich Schütze
EMNLP (1)1
2021 Generating Datasets with Pretrained Language Models
abstract
To obtain high-quality sentence embeddings from pretrained language models (PLMs), they must either be augmented with additional pretraining objectives or finetuned on a large set of labeled text pairs.While the latter approach typically outperforms the former, it requires great human effort to generate suitable datasets of sufficient size.In this paper, we show how PLMs can be leveraged to obtain high-quality sentence embeddings without the need for labeled data, finetuning or modifications to the pretraining objective: We utilize the generative abilities of large and high-performing PLMs to generate entire datasets of labeled text pairs from scratch, which we then use for finetuning much smaller and more efficient models.Our fully unsupervised approach outperforms strong baselines on several semantic textual similarity datasets.1
Timo Schick, Hinrich Schütze
EMNLP (1)1
2021 It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
abstract
When scaled to hundreds of billions of parameters, pretrained language models such as GPT-3 (Brown et al., 2020) achieve remarkable few-shot performance. However, enormous amounts of compute are required for training and applying such big models, resulting in a large carbon footprint and making it difficult for researchers and practitioners to use them. We show that performance similar to GPT-3 can be obtained with language models that are much "greener" in that their parameter count is several orders of magnitude smaller. This is achieved by converting textual inputs into cloze questions that contain a task description, combined with gradient-based optimization; exploiting unlabeled data gives further improvements. We identify key factors required for successful natural language understanding with small language models.
Timo Schick, Hinrich Schütze
NAACL-HLT1
2021 Self-Diagnosis and Self-Debiasing: A Proposal for Reducing Corpus-Based Bias in NLP
abstract
Abstract ⚠ This paper contains prompts and model outputs that are offensive in nature. When trained on large, unfiltered crawls from the Internet, language models pick up and reproduce all kinds of undesirable biases that can be found in the data: They often generate racist, sexist, violent, or otherwise toxic language. As large models require millions of training examples to achieve good performance, it is difficult to completely prevent them from being exposed to such content. In this paper, we first demonstrate a surprising finding: Pretrained language models recognize, to a considerable degree, their undesirable biases and the toxicity of the content they produce. We refer to this capability as self-diagnosis. Based on this finding, we then propose a decoding algorithm that, given only a textual description of the undesired behavior, reduces the probability of a language model producing problematic text. We refer to this approach as self-debiasing. Self-debiasing does not rely on manually curated word lists, nor does it require any training data or changes to the model’s parameters. While we by no means eliminate the issue of language models generating biased text, we believe our approach to be an important step in this direction.1
Timo Schick, Sahana Udupa, Hinrich Schütze
Trans. Assoc. Comput. Linguistics1
2020 Rare Words: A Major Problem for Contextualized Embeddings and How to Fix it by Attentive Mimicking
abstract
Pretraining deep neural network architectures with a language modeling objective has brought large improvements for many natural language processing tasks. Exemplified by BERT, a recently proposed such architecture, we demonstrate that despite being trained on huge amounts of data, deep language models still struggle to understand rare words. To fix this problem, we adapt Attentive Mimicking, a method that was designed to explicitly learn embeddings for rare words, to deep language models. In order to make this possible, we introduce one-token approximation, a procedure that enables us to use Attentive Mimicking even when the underlying language model uses subword-based tokenization, i.e., it does not assign embeddings to all words. To evaluate our method, we create a novel dataset that tests the ability of language models to capture semantic properties of words without any task-specific fine-tuning. Using this dataset, we show that adding our adapted version of Attentive Mimicking to BERT does substantially improve its understanding of rare words.
Timo Schick, Hinrich Schütze
AAAI1
2020 BERTRAM: Improved Word Embeddings Have Big Impact on Contextualized Model Performance
abstract
Pretraining deep language models has led to large performance gains in NLP.Despite this success, Schick and Schütze (2020) recently showed that these models struggle to understand rare words.For static word embeddings, this problem has been addressed by separately learning representations for rare words.In this work, we transfer this idea to pretrained language models: We introduce BERTRAM, a powerful architecture based on BERT that is capable of inferring high-quality embeddings for rare words that are suitable as input representations for deep language models.This is achieved by enabling the surface form and contexts of a word to interact with each other in a deep architecture.Integrating BERTRAM into BERT leads to large performance increases due to improved representations of rare and medium frequency words on both a rare word probing task and three downstream tasks. 1
Timo Schick, Hinrich Schütze
ACL1
2020 Automatically Identifying Words That Can Serve as Labels for Few-Shot Text Classification
abstract
A recent approach for few-shot text classification is to convert textual inputs to cloze questions that contain some form of task description, process them with a pretrained language model and map the predicted words to labels.Manually defining this mapping between words and labels requires both domain expertise and an understanding of the language model's abilities.To mitigate this issue, we devise an approach that automatically finds such a mapping given small amounts of training data.For a number of tasks, the mapping found by our approach performs almost as well as hand-crafted label-to-word mappings.
Timo Schick, Helmut Schmid, Hinrich Schütze
COLING1
2019 Learning Semantic Representations for Novel Words: Leveraging Both Form and Context
abstract
Word embeddings are a key component of high-performing natural language processing (NLP) systems, but it remains a challenge to learn good representations for novel words on the fly, i.e., for words that did not occur in the training data. The general problem setting is that word embeddings are induced on an unlabeled training corpus and then a model is trained that embeds novel words into this induced embedding space. Currently, two approaches for learning embeddings of novel words exist: (i) learning an embedding from the novel word’s surface-form (e.g., subword n-grams) and (ii) learning an embedding from the context in which it occurs. In this paper, we propose an architecture that leverages both sources of information – surface-form and context – and show that it results in large increases in embedding quality. Our architecture obtains state-of-the-art results on the Definitional Nonce and Contextual Rare Words datasets. As input, we only require an embedding set and an unlabeled corpus for training our architecture to produce embeddings appropriate for the induced embedding space. Thus, our model can easily be integrated into any existing NLP system and enhance its capability to handle novel words.
Timo Schick, Hinrich Schütze
AAAI1