Jason Weston

dblp:29/6977 · also Jason Aaron Edward Weston, Jason E. Weston · DBLP profile ↗
← Back
148ranked-venue papers
23as first author
31since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 128 · 17 first-author · 31 since 2021Databases, data management, data science and information retrieval · 15 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 12 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author
YearPublicationVenuePosition
2025 Byte Latent Transformer: Patches Scale Better Than Tokens
abstract
Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodriguez, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason E Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srini Iyer. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Artidoro Pagnoni, Ramakanth Pasunuru, Pedro Rodríguez 0001, John Nguyen, Benjamin Muller, Margaret Li, Chunting Zhou, Lili Yu, Jason Weston, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Ari Holtzman, Srinivasan Iyer 0001
ACL (1)9
2025 Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
abstract
Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason E Weston, Sainbayar Sukhbaatar. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Tianhao Wu 0002, Weizhe Yuan, Olga Golovneva, Jing Xu 0014, Yuandong Tian, Jiantao Jiao, Jason Weston, Sainbayar Sukhbaatar
EMNLP7
2025 Following Length Constraints in Instructions
abstract
Aligned instruction following models can better fulfill user requests than their unaligned counterparts.However, it has been shown that there is a length bias in evaluation of such models, and that training algorithms tend to exploit this bias by learning longer responses.In this work we show how to train models that can be controlled at inference time with instructions containing desired length constraints.Such models are superior in length instructed evaluations, outperforming standard instruction following models such as GPT4, Llama 3 and Mixtral.
Weizhe Yuan, Ilia Kulikov, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014
EMNLP6
2025 Backtracking Improves Generation Safety
abstract
Text generation has a fundamental limitation almost by definition: there is no taking back tokens that have been generated, even when they are clearly problematic. In the context of language model safety, when a partial unsafe generation is produced, language models by their nature tend to happily keep on generating similarly unsafe additional text. This is in fact how safety alignment of frontier models gets circumvented in the wild, despite great efforts in improving their safety. Deviating from the paradigm of approaching safety alignment as prevention (decreasing the probability of harmful responses), we propose backtracking, a technique that allows language models to "undo" and recover from their own unsafe generation through the introduction of a special [RESET] token. Our method can be incorporated into either SFT or DPO training to optimize helpfulness and harmlessness. We show that models trained to backtrack are consistently safer than baseline models: backtracking Llama-3-8B is four times more safe than the baseline model (6.1\% $\to$ 1.5\%) in our evaluations without regression in helpfulness. Our method additionally provides protection against four adversarial attacks including an adaptive attack, despite not being trained to do so.
Jianfeng Chi, Hailey Nguyen, Kartikeya Upasani, Dan Bikel, Jason Weston, Eric Michael Smith
ICLR6
2025 Thinking LLMs: General Instruction Following with Thought Generation
abstract
LLMs are typically trained to answer user questions or follow instructions similarly to how human experts respond. However, in the standard alignment framework they lack the basic ability of explicit thinking before answering. Thinking is important for complex questions that require reasoning and planning – but can be applied to any task. We propose a training method for equipping existing LLMs with such thinking abilities for general instruction following without use of additional human data. We achieve this by an iterative search and optimization procedure that explores the space of possible thought generations, allowing the model to learn how to think without direct supervision. For each instruction, the thought candidates are scored using a judge model to evaluate their responses only, and then optimized via preference optimization. We show that this procedure leads to superior performance on AlpacaEval and Arena-Hard, and shows gains from thinking on non-reasoning categories such as marketing, health and general knowledge, in addition to more traditional reasoning & problem-solving tasks.
Tianhao Wu 0002, Janice Lan, Weizhe Yuan, Jiantao Jiao, Jason Weston, Sainbayar Sukhbaatar
ICML5
2025 Self-Consistency Preference Optimization
abstract
Self-alignment, whereby models learn to improve themselves without human annotation, is a rapidly growing research area. However, existing techniques often fail to improve complex reasoning tasks due to the difficulty of assigning correct rewards. An orthogonal approach that is known to improve correctness is self-consistency, a method applied at inference time based on multiple sampling in order to find the most consistent answer. In this work, we extend the self-consistency concept to help train models. We thus introduce self-consistency preference optimization (ScPO), which iteratively trains consistent answers to be preferred over inconsistent ones on unsupervised new problems. We show ScPO leads to large improvements over conventional reward model training on reasoning tasks such as GSM8K and MATH, closing the gap with supervised training with gold answers or preferences, and that combining ScPO with standard supervised learning improves results even further. On ZebraLogic, ScPO finetunes Llama-3 8B to be superior to Llama-3 70B, Gemma-2 27B, and Claude-3 Haiku.
Archiki Prasad, Weizhe Yuan, Richard Yuanzhe Pang, Jing Xu 0014, Maryam Fazel-Zarandi, Mohit Bansal, Sainbayar Sukhbaatar, Jason Weston, Jane Dwivedi-Yu
ICML8
2025 Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
abstract
LLM-as-a-Judge models generate chain-of-thought (CoT) sequences intended to capture the step-by-step reasoning process that underlies the final evaluation of a response. However, due to the lack of human-annotated CoTs for evaluation, the required components and structure of effective reasoning traces remain understudied. Consequently, previous approaches often (1) constrain reasoning traces to hand-designed components, such as a list of criteria, reference answers, or verification questions and (2) structure them such that planning is intertwined with the reasoning for evaluation. In this work, we propose EvalPlanner, a preference optimization algorithm for Thinking-LLM-as-a-Judge that first generates an unconstrained evaluation plan, followed by its execution, and then the final judgment. In a self-training loop, EvalPlanner iteratively optimizes over synthetically constructed evaluation plans and executions, leading to better final verdicts. Our method achieves a new state-of-the-art performance for generative reward models on RewardBench and PPE, despite being trained on fewer amount of, and synthetically generated, preference pairs. Additional experiments on other benchmarks like RM-Bench, JudgeBench, and FollowBenchEval further highlight the utility of both planning and reasoning for building robust LLM-as-a-Judge reasoning models.
Swarnadeep Saha, Xian Li 0003, Marjan Ghazvininejad, Jason Weston
ICML4
2025 R.I.P.: Better Models by Survival of the Fittest Prompts
abstract
Training data quality is one of the most important drivers of final model quality. In this work, we introduce a method for evaluating data integrity based on the assumption that low-quality input prompts result in high variance and low quality responses. This is achieved by measuring the rejected response quality and the reward gap between the chosen and rejected preference pair. Our method, Rejecting Instruction Preferences (RIP) can be used to filter prompts from existing training sets, or to make high quality synthetic datasets, yielding large performance gains across various benchmarks compared to unfiltered data. Using Llama 3.1-8B-Instruct, RIP improves AlpacaEval2 LC Win Rate by 9.4%, Arena-Hard by 8.7%, and WildBench by 9.9%. Using Llama 3.3-70B-Instruct, RIP improves Arena-Hard from 67.5 to 82.9, from 18th place to 6th overall in the leaderboard.
Weizhe Yuan, Olga Golovneva, Tianhao Wu 0002, Sainbayar Sukhbaatar, Jason Weston, Jing Xu 0014
ICML6
2025 Meta CLIP 2: A Worldwide Scaling Recipe
abstract
Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully trained on billion-scale image-text pairs from the English world, scaling CLIP's training further to learning from the worldwide web data is still challenging: (1) no curation method is available to handle data points from non-English world; (2) the English performance from existing multilingual CLIP is worse than its English-only counterpart, i.e., "curse of multilinguality" that is common in LLMs. Here, we present Meta CLIP 2, the first recipe training CLIP from scratch on worldwide web-scale image-text pairs. To generalize our findings, we conduct rigorous ablations with minimal changes that are necessary to address the above challenges and present a recipe enabling mutual benefits from English and non-English world data. In zero-shot ImageNet classification, Meta CLIP 2 ViT-H/14 surpasses its English-only counterpart by 0.8% and mSigLIP by 0.7%, and surprisingly sets new state-of-the-art without system-level confounding factors (e.g., translation, bespoke architecture changes) on multilingual benchmarks, such as CVQA with 57.4%, Babel-ImageNet with 50.2% and XM3600 with 64.3% on image-to-text retrieval. Code and model are available at https://github.com/facebookresearch/MetaCLIP.
Yung-Sung Chuang, Ching-Feng Yeh, Kehan Lyu, Ramya Raghavendra, James R. Glass, Lifei Huang, Jason Weston, Luke Zettlemoyer, Xinlei Chen, Zhuang Liu 0003, Saining Xie, Scott Yih, Shang-Wen Li 0001, Hu Xu 0001
NeurIPS9
2025 NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions
abstract
Scaling reasoning capabilities beyond traditional domains such as math and coding is hindered by the lack of diverse and high-quality questions. To overcome this limitation, we introduce a scalable approach for generating diverse and challenging reasoning questions, accompanied by reference answers. We present NaturalReasoning, a comprehensive dataset comprising 2.8 million questions that span multiple domains, including STEM fields (e.g., Physics, Computer Science), Economics, Social Sciences, and more. We demonstrate the utility of the questions in NaturalReasoning through knowledge distillation experiments which show that NaturalReasoning can effectively elicit and transfer reasoning capabilities from a strong teacher model. Furthermore, we demonstrate that NaturalReasoning is also effective for unsupervised self-training using external reward models or self-rewarding.
Weizhe Yuan, Jane Dwivedi-Yu, Karthik Padthe, Ilia Kulikov, Kyunghyun Cho, Yuandong Tian, Jason Weston, Xian Li 0003
NeurIPS10
2025 Self-Challenging Language Model Agents
abstract
Large language models are quickly becoming the foundation for intelligent agents that are capable of using tools. However, training such agents is challenging because it requires human creation and annotation of a diverse set of tasks, tools, and evaluation criteria. In this paper, we propose the Self-Challenging Agent framework for training an agent on high-quality tasks that are generated by itself. The agent first plays the role of challenger and generates a task after interacting with the given tools. The tasks take the form of a novel general class of problems termed Code-as-Task, which are defined by an instruction, a verification function and solution and failure cases which serve as tests, allowing to filter only for high-quality tasks. The agent then takes an executor role and trains on those tasks with reinforcement learning using the evaluation feedback as a reward. We show our method improves the performance of Llama-3.1-8B-Instruct on two existing multi-turn tool-use agent benchmarks, M$^3$ToolEval and TauBench, with a two-fold average success rate increase, despite using only self-generated training data.
Sergey Levine, Jason Weston, Xian Li 0003, Sainbayar Sukhbaatar
NeurIPS3
2024 System-Level Natural Language Feedback
abstract
Natural language (NL) feedback offers rich insights into user experience.While existing studies focus on an instance-level approach, where feedback is used to refine specific examples, we introduce a framework for system-level use of NL feedback.We show how to use feedback to formalize system-level design decisions in a human-in-the-loop-process -in order to produce better models.In particular this is done through: (i) metric design for tasks; and (ii) language model prompt design for refining model responses.We conduct two case studies of this approach for improving search query and dialog response generation, demonstrating the effectiveness of system-level feedback.We show the combination of system-level and instancelevel feedback brings further gains, and that human written instance-level feedback results in more grounded refinements than GPT-3.5 written ones, underlying the importance of human feedback for building systems.We release our code and data at https://github.com/ yyy-Apple/Sys-NL-Feedback.
Weizhe Yuan, Kyunghyun Cho, Jason Weston
EACL (1)3
2024 Self-Alignment with Instruction Backtranslation
abstract
We present a scalable method to build a high quality instruction following language model by automatically labelling human-written text with corresponding instructions. Our approach, named instruction backtranslation, starts with a language model finetuned on a small amount of seed data, and a given web corpus. The seed model is used to construct training examples by generating instruction prompts for web documents (self-augmentation), and then selecting high quality examples from among these candidates (self-curation). This data is then used to finetune a stronger model. Finetuning LLaMa on two iterations of our approach yields a model that outperforms all other LLaMa-based models on the Alpaca leaderboard not relying on distillation data, demonstrating highly effective self-alignment.
Xian Li 0003, Chunting Zhou, Timo Schick, Omer Levy, Luke Zettlemoyer, Jason Weston, Mike Lewis
ICLR7
2024 Self-Rewarding Language Models
abstract
We posit that to achieve superhuman agents, future models require superhuman feedback in order to provide an adequate training signal. Current approaches commonly train reward models from human preferences, which may then be bottlenecked by human performance level, and secondly these reward models require additional human preferences data to further improve.In this work, we study Self-Rewarding Language Models, where the language model itself is used via LLM-as-a-Judge prompting to provide its own rewards during training. We show that during Iterative DPO training, not only does instruction following ability improve, but also the ability to provide high-quality rewards to itself. Fine-tuning Llama 2 70B on three iterations of our approach yields a model that outperforms many existing systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro, and GPT-4 0613. While there is much left still to explore, this work opens the door to the possibility of models that can continually improve in both axes.
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 0003, Sainbayar Sukhbaatar, Jing Xu 0014, Jason Weston
ICML7
2024 Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
abstract
Swarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, Xian Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Swarnadeep Saha, Omer Levy, Asli Celikyilmaz, Mohit Bansal, Jason Weston, Xian Li 0003
NAACL-HLT5
2024 When Life Gives You Lemons, Make Cherryade: Converting Feedback from Bad Responses into Good Labels
abstract
Weiyan Shi, Emily Dinan, Kurt Shuster, Jason Weston, Jing Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Weiyan Shi 0001, Emily Dinan, Kurt Shuster 0001, Jason Weston, Jing Xu 0014
NAACL-HLT4
2024 The ART of LLM Refinement: Ask, Refine, and Trust
abstract
Kumar Shridhar, Koustuv Sinha, Andrew Cohen, Tianlu Wang, Ping Yu, Ramakanth Pasunuru, Mrinmaya Sachan, Jason Weston, Asli Celikyilmaz. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Kumar Shridhar, Koustuv Sinha, Andrew Cohen, Ramakanth Pasunuru, Mrinmaya Sachan, Jason Weston, Asli Celikyilmaz
NAACL-HLT8
2024 Iterative Reasoning Preference Optimization
abstract
Iterative preference optimization methods have recently been shown to perform well for general instruction tuning tasks, but typically make little improvement on reasoning tasks. In this work we develop an iterative approach that optimizes the preference between competing generated Chain-of-Thought (CoT) candidates by optimizing for winning vs. losing reasoning steps. We train using a modified DPO loss with an additional negative log-likelihood term, which we find to be crucial. We show reasoning improves across repeated iterations of this scheme. While only relying on examples in the training set, our approach results in increasing accuracy on GSM8K, MATH, and ARC-Challenge for Llama-2-70B-Chat, outperforming other Llama-2-based models not relying on additionally sourced datasets. For example, we see a large improvement from 55.6% to 81.6% on GSM8K and an accuracy of 88.7% with majority voting out of 32 samples.
Richard Yuanzhe Pang, Weizhe Yuan, He He 0001, Kyunghyun Cho, Sainbayar Sukhbaatar, Jason Weston
NeurIPS6
2023 The CRINGE Loss: Learning what language not to model
abstract
Leonard Adolphs, Tianyu Gao, Jing Xu, Kurt Shuster, Sainbayar Sukhbaatar, Jason Weston. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Leonard Adolphs, Jing Xu 0014, Kurt Shuster 0001, Sainbayar Sukhbaatar, Jason Weston
ACL (1)6
2023 Learning New Skills after Deployment: Improving open-domain internet-driven dialogue with human feedback
abstract
Frozen models trained to mimic static datasets can never improve their performance.Models that can employ internet-retrieval for up-to-date information and obtain feedback from humans during deployment provide the promise of both adapting to new information, and improving their performance.In this work we study how to improve internet-driven conversational skills in such a learning framework.We collect deployment data, which we make publicly available, of human interactions, and collect various types of human feedback -including binary quality measurements, free-form text feedback, and fine-grained reasons for failure.We then study various algorithms for improving from such feedback, including standard supervised learning, rejection sampling, modelguiding and reward-based learning, in order to make recommendations on which type of feedback and algorithms work best.We find the recently introduced DIRECTOR model (Arora et al., 2022) shows significant improvements over other existing approaches.
Jing Xu 0014, Megan Ung, Mojtaba Komeili, Kushal Arora, Y-Lan Boureau, Jason Weston
ACL (1)6
2023 Learning to Reason and Memorize with Self-Notes
abstract
Large language models have been shown to struggle with multi-step reasoning, and do not retain previous reasoning steps for future use. We propose a simple method for solving both of these problems by allowing the model to take Self-Notes. Unlike recent chain-of-thought or scratchpad approaches, the model can deviate from the input context at any time to explicitly think and write down its thoughts. This allows the model to perform reasoning on the fly as it reads the context and even integrate previous reasoning steps, thus enhancing its memory with useful information and enabling multi-step reasoning. Experiments across a wide variety of tasks demonstrate that our method can outperform chain-of-thought and scratchpad methods by taking Self-Notes that interleave the input text.
Jack Lanchantin, Shubham Toshniwal, Jason Weston, Arthur Szlam, Sainbayar Sukhbaatar
NeurIPS3
2022 Internet-Augmented Dialogue Generation
abstract
The largest store of continually updating knowledge on our planet can be accessed via internet search.In this work we study giving access to this information to conversational agents.Large language models, even though they store an impressive amount of knowledge within their weights, are known to hallucinate facts when generating dialogue (Shuster et al., 2021); moreover, those facts are frozen in time at the point of model training.In contrast, we propose an approach that learns to generate an internet search query based on the context, and then conditions on the search results to finally generate a response, a method that can employ up-to-the-minute relevant information.We train and evaluate such models on a newly collected dataset of human-human conversations whereby one of the speakers is given access to internet search during knowledge-driven discussions in order to ground their responses.We find that search-query based access of the internet in conversation provides superior performance compared to existing approaches that either use no augmentation or FAISS-based retrieval (Lewis et al., 2020b).
Mojtaba Komeili, Kurt Shuster 0001, Jason Weston
ACL (1)3
2022 Beyond Goldfish Memory: Long-Term Open-Domain Conversation
abstract
Despite recent improvements in open-domain dialogue models, state-of-the-art models are trained and evaluated on short conversations with little context.In contrast, the long-term conversation setting has hardly been studied.In this work we collect and release a humanhuman dataset consisting of multiple chat sessions whereby the speaking partners learn about each other's interests and discuss the things they have learnt from past sessions.We show how existing models trained on existing datasets perform poorly in this long-term conversation setting in both automatic and human evaluations, and we study long-context models that can perform much better.In particular, we find retrieval-augmented methods and methods with an ability to summarize and recall previous conversations outperform the standard encoder-decoder architectures currently considered state-of-the-art.* We use this term colloquially, see Agranoff et al. (1965) for evidence of goldfish long-term memory.
Jing Xu 0014, Arthur Szlam, Jason Weston
ACL (1)3
2022 Staircase Attention for Recurrent Processing of Sequences
abstract
Attention mechanisms have become a standard tool for sequence modeling tasks, in particular by stacking self-attention layers over the entire input sequence as in the Transformer architecture. In this work we introduce a novel attention procedure called staircase attention that, unlike self-attention, operates across the sequence (in time) recurrently processing the input by adding another step of processing. A step in the staircase comprises of backward tokens (encoding the sequence so far seen) and forward tokens (ingesting a new part of the sequence). Thus our model can trade off performance and compute, by increasing the amount of recurrence through time and depth. Staircase attention is shown to be able to solve tasks that involve tracking that conventional Transformers cannot, due to this recurrence. Further, it is shown to provide improved modeling power for the same size model (number of parameters) compared to self-attentive Transformers on large language modeling and dialogue tasks, yielding significant perplexity gains.
Da Ju, Stephen Roller, Sainbayar Sukhbaatar, Jason Weston
NeurIPS4
2021 I like fish, especially dolphins: Addressing Contradictions in Dialogue Modeling
abstract
Yixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, Jason Weston. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yixin Nie, Mary Williamson, Mohit Bansal, Douwe Kiela, Jason Weston
ACL/IJCNLP (1)5
2021 Recipes for Building an Open-Domain Chatbot
abstract
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, Jason Weston. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021.
Stephen Roller, Emily Dinan, Naman Goyal 0001, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu 0014, Myle Ott, Eric Michael Smith, Y-Lan Boureau, Jason Weston
EACL11
2021 Multi-Modal Open-Domain Dialogue
abstract
Recent work in open-domain conversational agents has demonstrated that significant improvements in humanness and user preference can be achieved via massive scaling in both pre-training data and model size (Adiwardana et al., 2020;Roller et al., 2020).However, if we want to build agents with human-like abilities, we must expand beyond handling just text.A particularly important topic is the ability to see images and communicate about what is perceived.With the goal of getting humans to engage in multi-modal dialogue, we investigate combining components from state-of-the-art open-domain dialogue agents with those from state-of-the-art vision models.We study incorporating different image fusion schemes and domain-adaptive pre-training and fine-tuning strategies, and show that our best resulting model outperforms strong existing models in multi-modal dialogue while simultaneously performing as well as its predecessor (text-only) BlenderBot (Roller et al., 2020) in text-based conversation.We additionally investigate and incorporate safety components in our final model, and show that such efforts do not diminish model performance with respect to human preference.
Kurt Shuster 0001, Eric Michael Smith, Da Ju, Jason Weston
EMNLP (1)4
2021 Not All Memories are Created Equal: Learning to Forget by Expiring
abstract
Attention mechanisms have shown promising results in sequence modeling tasks that require long-term memory. Recent work investigated mechanisms to reduce the computational cost of preserving and storing memories. However, not all content in the past is equally important to remember. We propose Expire-Span, a method that learns to retain the most important information and expire the irrelevant information. This forgetting of memories enables Transformers to scale to attend over tens of thousands of previous timesteps efficiently, as not all states from previous timesteps are preserved. We demonstrate that Expire-Span can help models identify and retain critical information and show it can achieve strong performance on reinforcement learning tasks specifically designed to challenge this functionality. Next, we show that Expire-Span can scale to memories that are tens of thousands in size, setting a new state of the art on incredibly long context tasks such as character-level language modeling and a frame-by-frame moving objects task. Finally, we analyze the efficiency of Expire-Span compared to existing approaches and demonstrate that it trains faster and uses less memory.
Sainbayar Sukhbaatar, Da Ju, Spencer Poff, Stephen Roller, Arthur Szlam, Jason Weston, Angela Fan
ICML6
2021 How to Motivate Your Dragon: Teaching Goal-Driven Agents to Speak and Act in Fantasy Worlds
abstract
Prithviraj Ammanabrolu, Jack Urbanek, Margaret Li, Arthur Szlam, Tim Rocktäschel, Jason Weston. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Prithviraj Ammanabrolu, Jack Urbanek, Margaret Li, Arthur Szlam, Tim Rocktäschel, Jason Weston
NAACL-HLT6
2021 Bot-Adversarial Dialogue for Safe Conversational Agents
abstract
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, Emily Dinan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Jing Xu 0014, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, Emily Dinan
NAACL-HLT5
2021 Hash Layers For Large Sparse Models
abstract
We investigate the training of sparse layers that use different parameters for different inputs based on hashing in large Transformer models. Specifically, we modify the feedforward layer to hash to different sets of weights depending on the current token, over all tokens in the sequence. We show that this procedure either outperforms or is competitive with learning-to-route mixture-of-expert methods such as Switch Transformers and BASE Layers, while requiring no routing parameters or extra terms in the objective function such as a load balancing loss, and no sophisticated assignment algorithm. We study the performance of different hashing techniques, hash sizes and input features, and show that balanced and random hashes focused on the most local features work best, compared to either learning clusters or using longer-range context. We show our approach works well both on large language modeling and dialogue tasks, and on downstream fine-tuning tasks.
Stephen Roller, Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston
NeurIPS4
2020 Generating Interactive Worlds with Text
abstract
Procedurally generating cohesive and interesting game environments is challenging and time-consuming. In order for the relationships between the game elements to be natural, common-sense has to be encoded into arrangement of the elements. In this work, we investigate a machine learning approach for world creation using content from the multi-player text adventure game environment LIGHT (Urbanek et al. 2019). We introduce neural network based models to compositionally arrange locations, characters, and objects into a coherent whole. In addition to creating worlds based on existing elements, our models can generate new game content. Humans can also leverage our models to interactively aid in worldbuilding. We show that the game environments created with our approach are cohesive, diverse, and preferred by human evaluators compared to other machine learning based world construction algorithms.
Angela Fan, Jack Urbanek, Pratik Ringshia, Emily Dinan, Emma Qian, Siddharth Karamcheti, Shrimai Prabhumoye, Douwe Kiela, Tim Rocktäschel, Arthur Szlam, Jason Weston
AAAI11
2020 Don't Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training
abstract
Generative dialogue models currently suffer from a number of problems which standard maximum likelihood training does not address.They tend to produce generations that (i) rely too much on copying from the context, (ii) contain repetitions within utterances, (iii) overuse frequent words, and (iv) at a deeper level, contain logical flaws.In this work we show how all of these problems can be addressed by extending the recently introduced unlikelihood loss (Welleck et al., 2019a) to these cases.We show that appropriate loss functions which regularize generated outputs to match human distributions are effective for the first three issues.For the last important general issue, we show applying unlikelihood to collected data of what a model should not do is effective for improving logical consistency, potentially paving the way to generative models with greater reasoning ability.We demonstrate the efficacy of our approach across several dialogue tasks.
Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, Jason Weston
ACL7
2020 Adversarial NLI: A New Benchmark for Natural Language Understanding
abstract
We introduce a new large-scale NLI benchmark dataset, collected via an iterative, adversarial human-and-model-in-the-loop procedure.We show that training models on this new dataset leads to state-of-the-art performance on a variety of popular NLI benchmarks, while posing a more difficult challenge with its new test set.Our analysis sheds light on the shortcomings of current state-of-theart models, and shows that non-expert annotators are successful at finding their weaknesses.The data collection method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate.
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, Douwe Kiela
ACL5
2020 Image-Chat: Engaging Grounded Conversations
abstract
To achieve the long-term goal of machines being able to engage humans in conversation, our models should captivate the interest of their speaking partners.Communication grounded in images, whereby a dialogue is conducted based on a given photo, is a setup naturally appealing to humans (Hu et al., 2014).In this work we study large-scale architectures and datasets for this goal.We test a set of neural architectures using state-of-the-art image and text representations, considering various ways to fuse the components.To test such models, we collect a dataset of grounded human-human conversations, where speakers are asked to play roles given a provided emotional mood or style, as the use of such traits is also a key factor in engagingness (Guo et al., 2019).Our dataset, Image-Chat, consists of 202k dialogues over 202k images using 215 possible style traits.Automatic metrics and human evaluations of engagingness show the efficacy of our approach; in particular, we obtain state-of-the-art performance on the existing IGC task, and our best performing model is almost on par with humans on the Image-Chat test set (preferred 47.7% of the time).
Kurt Shuster 0001, Samuel Humeau 0001, Antoine Bordes, Jason Weston
ACL4
2020 The Dialogue Dodecathlon: Open-Domain Knowledge and Image Grounded Conversational Agents
abstract
We introduce dodecaDialogue: a set of 12 tasks that measures if a conversational agent can communicate engagingly with personality and empathy, ask questions, answer questions by utilizing knowledge resources, discuss topics and situations, and perceive and converse about images.By multi-tasking on such a broad large-scale set of data, we hope to both move towards and measure progress in producing a single unified agent that can perceive, reason and converse with humans in an open-domain setting.We show that such multi-tasking improves over a BERT pretrained baseline, largely due to multi-tasking with very large dialogue datasets in a similar domain, and that the multi-tasking in general provides gains to both text and image-based tasks using several metrics in both the finetune and task transfer settings.We obtain stateof-the-art results on many of the tasks, providing a strong baseline for this challenge.
Kurt Shuster 0001, Da Ju, Stephen Roller, Emily Dinan, Y-Lan Boureau, Jason Weston
ACL6
2020 Can You Put it All Together: Evaluating Conversational Agents' Ability to Blend Skills
abstract
Being engaging, knowledgeable, and empathetic are all desirable general qualities in a conversational agent.Previous work has introduced tasks and datasets that aim to help agents to learn those qualities in isolation and gauge how well they can express them.But rather than being specialized in one single quality, a good open-domain conversational agent should be able to seamlessly blend them all into one cohesive conversational flow.In this work, we investigate several ways to combine models trained towards isolated capabilities, ranging from simple model aggregation schemes that require minimal additional training, to various forms of multi-task training that encompass several skills at all training stages.We further propose a new dataset, Blended-SkillTalk, to analyze how these capabilities would mesh together in a natural conversation, and compare the performance of different architectures and training schemes.Our experiments show that multi-tasking over several tasks that focus on particular capabilities results in better blended conversation performance compared to models trained on a single skill, and that both unified or two-stage approaches perform well if they are constructed to avoid unwanted bias in skill selection or are fine-tuned on our new task.
Eric Michael Smith, Mary Williamson, Kurt Shuster 0001, Jason Weston, Y-Lan Boureau
ACL4
2020 Queens are Powerful too: Mitigating Gender Bias in Dialogue Generation
abstract
Models often easily learn biases present in the training data, and their predictions directly reflect this bias.We analyze gender bias in dialogue data, and examine how this bias is actually amplified in subsequent generative chit-chat dialogue models.We measure gender bias in six existing dialogue datasets, and focus on the most biased one, the multiplayer text-based fantasy adventure dataset LIGHT (Urbanek et al., 2019), as a testbed for our bias mitigation techniques.The LIGHT dataset is highly imbalanced with respect to gender, containing predominantly male characters, likely because it is entirely collected by crowdworkers and reflects common biases that exist in fantasy or medieval settings.We consider three techniques to mitigate gender bias: counterfactual data augmentation, targeted data collection, and bias controlled training.We show that our proposed techniques mitigate gender bias in LIGHT by balancing the genderedness of generated dialogue utterances and are particularly effective in combination.We quantify performance using various evaluation methods-such as quantity of gendered words, a dialogue safety classifier, and human studies-all of which show that our models generate less gendered, but equally engaging chit-chat responses.
Emily Dinan, Angela Fan, Adina Williams, Jack Urbanek, Douwe Kiela, Jason Weston
EMNLP (1)6
2020 Multi-Dimensional Gender Bias Classification
abstract
Machine learning models are trained to find patterns in data.NLP models can inadvertently learn socially undesirable patterns when training on gender biased text.In this work, we propose a novel, general framework that decomposes gender bias in text along several pragmatic and semantic dimensions: bias from the gender of the person being spoken about, bias from the gender of the person being spoken to, and bias from the gender of the speaker.Using this fine-grained framework, we automatically annotate eight large scale datasets with gender information.In addition, we collect a new, crowdsourced evaluation benchmark.Distinguishing between gender bias along multiple dimensions enables us to train better and more fine-grained gender bias classifiers.We show our classifiers are valuable for a variety of applications, like controlling for gender bias in generative models, detecting gender bias in arbitrary text, and classifying text as offensive based on its genderedness.
Emily Dinan, Angela Fan, Ledell Wu, Jason Weston, Douwe Kiela, Adina Williams
EMNLP (1)4
2020 Poly-encoders: Architectures and Pre-training Strategies for Fast and Accurate Multi-sentence Scoring
Samuel Humeau 0001, Kurt Shuster 0001, Marie-Anne Lachaux, Jason Weston
ICLR4
2020 Neural Text Generation With Unlikelihood Training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, Jason Weston
ICLR6
2019 ELI5: Long Form Question Answering
abstract
We introduce the first large-scale corpus for long-form question answering, a task requiring elaborate and in-depth answers to openended questions.The dataset comprises 270K threads from the Reddit forum "Explain Like I'm Five" (ELI5) where an online community provides answers to questions which are comprehensible by five year olds.Compared to existing datasets, ELI5 comprises diverse questions requiring multi-sentence answers.We provide a large set of web documents to help answer the question.Automatic and human evaluations show that an abstractive model trained with a multi-task objective outperforms conventional Seq2Seq, language modeling, as well as a strong extractive baseline.However, our best model is still far from human performance since raters prefer gold responses in over 86% of cases, leaving ample opportunity for future improvement.1
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, Michael Auli
ACL (1)5
2019 Learning from Dialogue after Deployment: Feed Yourself, Chatbot!
abstract
The majority of conversations a dialogue agent sees over its lifetime occur after it has already been trained and deployed, leaving a vast store of potential training signal untapped.In this work, we propose the self-feeding chatbot, a dialogue agent with the ability to extract new training examples from the conversations it participates in.As our agent engages in conversation, it also estimates user satisfaction in its responses.When the conversation appears to be going well, the user's responses become new training examples to imitate.When the agent believes it has made a mistake, it asks for feedback; learning to predict the feedback that will be given improves the chatbot's dialogue abilities further.On the PERSONACHAT chitchat dataset with over 131k training examples, we find that learning from dialogue with a selffeeding chatbot significantly improves performance, regardless of the amount of traditional supervision.
Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazaré, Jason Weston
ACL (1)4
2019 Dialogue Natural Language Inference
abstract
Consistency is a long standing issue faced by dialogue models. In this paper, we frame the consistency of dialogue agents as natural language inference (NLI) and create a new natural language inference dataset called Dialogue NLI. We propose a method which demonstrates that a model trained on Dialogue NLI can be used to improve the consistency of a dialogue model, and evaluate the method with human evaluation and with automatic metrics on a suite of evaluation sets designed to measure a dialogue model’s consistency.
Sean Welleck, Jason Weston, Arthur Szlam, Kyunghyun Cho
ACL (1)2
2019 Engaging Image Captioning via Personality
abstract
Standard image captioning tasks such as COCO and Flickr30k are factual, neutral in tone and (to a human) state the obvious (e.g., “a man playing a guitar”). While such tasks are useful to verify that a machine understands the content of an image, they are not engaging to humans as captions. With this in mind we define a new task, PERSONALITY-CAPTIONS, where the goal is to be as engaging to humans as possible by incorporating controllable style and personality traits. We collect and release a large dataset of 241,858 of such captions conditioned over 215 possible traits. We build models that combine existing work from (i) sentence representations [36] with Transformers trained on 1.7 billion dialogue examples; and (ii) image representations [32] with ResNets trained on 3.5 billion social media images. We obtain state-of-the-art performance on Flickr30k and COCO, and strong performance on our new task. Finally, online evaluations validate that our task and models are engaging to humans, with our best model close to human performance.
Kurt Shuster 0001, Samuel Humeau 0001, Hexiang Hu, Antoine Bordes, Jason Weston
CVPR5
2019 Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack
abstract
Emily Dinan, Samuel Humeau, Bharath Chintagunta, Jason Weston. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Emily Dinan, Samuel Humeau 0001, Bharath Chintagunta, Jason Weston
EMNLP/IJCNLP (1)4
2019 Recommendation as a Communication Game: Self-Supervised Bot-Play for Goal-oriented Dialogue
abstract
Dongyeop Kang, Anusha Balakrishnan, Pararth Shah, Paul Crook, Y-Lan Boureau, Jason Weston. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Dongyeop Kang, Anusha Balakrishnan, Pararth Shah, Paul A. Crook, Y-Lan Boureau, Jason Weston
EMNLP/IJCNLP (1)6
2019 Finding Generalizable Evidence by Learning to Convince Q&A Models
abstract
Ethan Perez, Siddharth Karamcheti, Rob Fergus, Jason Weston, Douwe Kiela, Kyunghyun Cho. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Ethan Perez, Siddharth Karamcheti, Rob Fergus, Jason Weston, Douwe Kiela, Kyunghyun Cho
EMNLP/IJCNLP (1)4
2019 Learning to Speak and Act in a Fantasy Text Adventure Game
abstract
Jack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, Jason Weston. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jack Urbanek, Angela Fan, Siddharth Karamcheti, Saachi Jain, Samuel Humeau 0001, Emily Dinan, Tim Rocktäschel, Douwe Kiela, Arthur Szlam, Jason Weston
EMNLP/IJCNLP (1)10
2019 Wizard of Wikipedia: Knowledge-Powered Conversational Agents
Emily Dinan, Stephen Roller, Kurt Shuster 0001, Angela Fan, Michael Auli, Jason Weston
ICLR (Poster)6
2019 Importance of Search and Evaluation Strategies in Neural Dialogue Modeling
abstract
We investigate the impact of search strategies in neural dialogue modeling.We first compare two standard search algorithms, greedy and beam search, as well as our newly proposed iterative beam search which produces a more diverse set of candidate responses.We evaluate these strategies in realistic full conversations with humans and propose a modelbased Bayesian calibration to address annotator bias.These conversations are analyzed using two automatic metrics: log-probabilities assigned by the model and utterance diversity.Our experiments reveal that better search algorithms lead to higher rated conversations.However, finding the optimal selection mechanism to choose from a more diverse set of candidates is still an open question.
Ilia Kulikov, Alexander H. Miller, Kyunghyun Cho, Jason Weston
INLG4
2018 StarSpace: Embed All The Things!
abstract
We present StarSpace, a general-purpose neural embedding model that can solve a wide variety of problems: labeling tasks such as text classification,ranking tasks such as information retrieval/web search,collaborative filtering-based or content-based recommendation,embedding of multi-relational graphs, and learning word, sentence or document level embeddings.In each case the model works by embedding those entities comprised of discrete features and comparing them against each other -- learning similarities dependent on the task.Empirical results on a number of tasks show that StarSpace is highly competitive with existing methods, whilst also being generally applicable to new cases where those methods are not.
Ledell Wu, Adam Fisch, Sumit Chopra, Keith Adams, Antoine Bordes, Jason Weston
AAAI6
2018 Personalizing Dialogue Agents: I have a dog, do you have pets too?
abstract
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, Jason Weston. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018.
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, Jason Weston
ACL (1)6
2018 Emergent Translation in Multi-Agent Communication
Jason Lee 0002, Kyunghyun Cho, Jason Weston, Douwe Kiela
ICLR (Poster)3
2018 Mastering the Dungeon: Grounded Language Learning by Mechanical Turker Descent
Zhilin Yang 0001, Saizheng Zhang, Jack Urbanek, Will Feng, Alexander H. Miller, Arthur Szlam, Douwe Kiela, Jason Weston
ICLR (Poster)8
2017 Reading Wikipedia to Answer Open-Domain Questions
abstract
This paper proposes to tackle open-domain question answering using Wikipedia as the unique knowledge source: the answer to any factoid question is a text span in a Wikipedia article. This task of machine reading at scale combines the challenges of document retrieval (finding the relevant articles) with that of machine comprehension of text (identifying the answer spans from those articles). Our approach combines a search component based on bigram hashing and TF-IDF matching with a multi-layer recurrent neural network model trained to detect answers in Wikipedia paragraphs. Our experiments on multiple existing QA datasets indicate that (1) both modules are highly competitive with respect to existing counterparts and (2) multitask learning using distant supervision on their combination is an effective complete system on this challenging task.
Danqi Chen 0001, Adam Fisch, Jason Weston, Antoine Bordes
ACL (1)3
2017 Learning End-to-End Goal-Oriented Dialog
Antoine Bordes, Y-Lan Boureau, Jason Weston
ICLR3
2017 Tracking the World State with Recurrent Entity Networks
Mikael Henaff, Jason Weston, Arthur Szlam, Antoine Bordes, Yann LeCun
ICLR (Poster)2
2017 Dialogue Learning With Human-in-the-Loop
Alexander H. Miller, Sumit Chopra, Marc'Aurelio Ranzato, Jason Weston
ICLR (Poster)5
2017 Learning through Dialogue Interactions by Asking Questions
Alexander H. Miller, Sumit Chopra, Marc'Aurelio Ranzato, Jason Weston
ICLR (Poster)5
2017 Memory Networks for Recommendation
abstract
Memory networks are a recently introduced model that combines reasoning, attention and memory for solving tasks in the areas of language understanding and dialogue -- where one exciting direction is the use of these models for dialogue-based recommendation. In this talk we describe these models and how they can learn to discuss, answer questions about, and recommend sets of items to a user. The ultimate goal of this research is to produce a full dialogue-based recommendation assistant. We will discuss recent datasets and evaluation tasks that have been built to assess these models abilities to see how far we have come.
Jason Weston
RecSys1
2016 Key-Value Memory Networks for Directly Reading Documents
abstract
Directly reading documents and being able to answer questions from them is an unsolved challenge.To avoid its inherent difficulty, question answering (QA) has been directed towards using Knowledge Bases (KBs) instead, which has proven effective.Unfortunately KBs often suffer from being too restrictive, as the schema cannot support certain types of answers, and too sparse, e.g.Wikipedia contains much more information than Freebase.In this work we introduce a new method, Key-Value Memory Networks, that makes reading documents more viable by utilizing different encodings in the addressing and output stages of the memory read operation.To compare using KBs, information extraction or Wikipedia documents directly in a single framework we construct an analysis tool, WIKIMOVIES, a QA dataset that contains raw text alongside a preprocessed KB, in the domain of movies.Our method reduces the gap between all three settings.It also achieves state-of-the-art results on the existing WIKIQA benchmark.
Alexander H. Miller, Adam Fisch, Jesse Dodge, Antoine Bordes, Jason Weston
EMNLP6
2016 Dialog-based Language Learning
abstract
A long-term goal of machine learning research is to build an intelligent dialog agent. Most research in natural language understanding has focused on learning from fixed training sets of labeled data, with supervision either at the word level (tagging, parsing tasks) or sentence level (question answering, machine translation). This kind of supervision is not realistic of how humans learn, where language is both learned by, and used for, communication. In this work, we study dialog-based language learning, where supervision is given naturally and implicitly in the response of the dialog partner during the conversation. We study this setup in two domains: the bAbI dataset of (Weston et al., 2015) and large-scale question answering from (Dodge et al., 2015). We evaluate a set of baseline learning strategies on these tasks, and show that a novel model incorporating predictive lookahead is a promising approach for learning from a teacher's response. In particular, a surprising result is that it can learn to answer questions correctly without any reward-based supervision at all.
Jason Weston
NIPS1
2015 Learning Anaphoricity and Antecedent Ranking Features for Coreference Resolution
abstract
Sam Wiseman, Alexander M. Rush, Stuart Shieber, Jason Weston. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Sam Wiseman, Alexander M. Rush, Stuart M. Shieber, Jason Weston
ACL (1)4
2015 A Neural Attention Model for Abstractive Sentence Summarization
abstract
Summarization based on text extraction is inherently limited, but generation-style abstractive methods have proven challenging to build.In this work, we propose a fully data-driven approach to abstractive sentence summarization.Our method utilizes a local attention-based model that generates each word of the summary conditioned on the input sentence.While the model is structurally simple, it can easily be trained end-to-end and scales to a large amount of training data.The model shows significant performance gains on the DUC-2004 shared task compared with several strong baselines.
Alexander M. Rush, Sumit Chopra, Jason Weston
EMNLP3
2015 User Conditional Hashtag Prediction for Images
abstract
Understanding the content of user's image posts is a particularly interesting problem in social networks and web settings. Current machine learning techniques focus mostly on curated training sets of image-label pairs, and perform image classification given the pixels within the image. In this work we instead leverage the wealth of information available from users: firstly, we employ user hashtags to capture the description of image content; and secondly, we make use of valuable contextual information about the user. We show how user metadata (age, gender, etc.) combined with image features derived from a convolutional neural network can be used to perform hashtag prediction. We explore two ways of combining these heterogeneous features into a learning framework: (i) simple concatenation; and (ii) a 3-way multiplicative gating, where the image model is conditioned on the user metadata. We apply these models to a large dataset of de-identified Facebook posts and demonstrate that modeling the user can significantly improve the tag prediction quality over current state-of-the-art methods.
Remi Denton, Jason Weston, Manohar Paluri, Lubomir D. Bourdev, Rob Fergus
KDD2
2015 End-To-End Memory Networks
abstract
We introduce a neural network with a recurrent attention model over a possibly large external memory. The architecture is a form of Memory Network (Weston et al., 2015) but unlike the model in that work, it is trained end-to-end, and hence requires significantly less supervision during training, making it more generally applicable in realistic settings. It can also be seen as an extension of RNNsearch to the case where multiple computational steps (hops) are performed per output symbol. The flexibility of the model allows us to apply it to tasks as diverse as (synthetic) question answering and to language modeling. For the former our approach is competitive with Memory Networks, but with less supervision. For the latter, on the Penn TreeBank and Text8 datasets our approach demonstrates comparable performance to RNNs and LSTMs. In both cases we show that the key concept of multiple computational hops yields improved results.
Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, Rob Fergus
NIPS3
2014 Semantic Frame Identification with Distributed Word Representations
abstract
A computer-implemented technique can include receiving, at a server, labeled training data including a plurality of groups of words, each group of words having a predicate word, each word having generic word embeddings. The technique can include extracting, at the server, the plurality of groups of words in a syntactic context of their predicate words. The technique can include concatenating, at the server, the generic word embeddings to create a high dimensional vector space representing features for each word. The technique can include obtaining, at the server, a model having a learned mapping from the high dimensional vector space to a low dimensional vector space and learned embeddings for each possible semantic frame in the low dimensional vector space. The technique can also include outputting, by the server, the model for storage, the model being configured to identify a specific semantic frame for an input.
Karl Moritz Hermann, Dipanjan Das 0001, Jason Weston, Kuzman Ganchev
ACL (1)3
2014 Deep Learning for Character-Based Information Extraction
Yanjun Qi, Sujatha G. Das, Ronan Collobert, Jason Weston
ECIR4
2014 Question Answering with Subgraph Embeddings
abstract
This paper presents a system which learns to answer questions on a broad range of topics from a knowledge base using few hand-crafted features.Our model learns low-dimensional embeddings of words and knowledge base constituents; these representations are used to score natural language questions against candidate answers.Training our system using pairs of questions and structured representations of their answers, and pairs of question paraphrases, yields competitive results on a recent benchmark of the literature.
Antoine Bordes, Sumit Chopra, Jason Weston
EMNLP3
2014 #TagSpace: Semantic Embeddings from Hashtags
abstract
We describe a convolutional neural net-work that learns feature representations for short textual posts using hashtags as a su-pervised signal. The proposed approach is trained on up to 5.5 billion words predict-ing 100,000 possible hashtags. As well as strong performance on the hashtag predic-tion task itself, we show that its learned representation of text (ignoring the hash-tag labels) is useful for other tasks as well. To that end, we present results on a docu-ment recommendation task, where it also outperforms a number of baselines. 1
Jason Weston, Sumit Chopra, Keith Adams
EMNLP1
2014 Affinity Weighted Embedding
abstract
Supervised linear embedding models like Wsabie (Weston et al., 2011) and supervised semantic indexing (Bai et al., 2010) have proven successful at ranking, recommendation and annotation tasks. However, despite being scalable to large datasets they do not take full advantage of the extra data due to their linear nature, and we believe they typically underfit. We propose a new class of models which aim to provide improved performance while retaining many of the benefits of the existing class of embedding models. Our approach works by reweighting each component of the embedding of features and labels with a potentially nonlinear affinity function. We describe several variants of the family, and show its usefulness on several datasets.
Jason Weston, Ron J. Weiss, Hector Yee
ICML1
2014 Open Question Answering with Weakly Supervised Embedding Models
Antoine Bordes, Jason Weston, Nicolas Usunier
ECML/PKDD (1)2
2014 Training highly multiclass classifiers
Maya R. Gupta, Samy Bengio, Jason Weston
J. Mach. Learn. Res.3
2014 Introduction to the special issue on learning semantics
Antoine Bordes, Léon Bottou, Ronan Collobert, Dan Roth 0001, Jason Weston, Luke Zettlemoyer
Mach. Learn.5
2014 A semantic matching energy function for learning with multi-relational data - Application to word-sense disambiguation
Antoine Bordes, Xavier Glorot, Jason Weston, Yoshua Bengio
Mach. Learn.3
2014 Learning semantic representations of objects and their parts
Grégoire Mesnil, Antoine Bordes, Jason Weston, Gal Chechik, Yoshua Bengio
Mach. Learn.3
2013 Connecting Language and Knowledge Bases with Embedding Models for Relation Extraction
abstract
This paper proposes a novel approach for relation extraction from free text which is trained to jointly use information from the text and from existing knowledge.Our model is based on scoring functions that operate by learning low-dimensional embeddings of words, entities and relationships from a knowledge base.We empirically show on New York Times articles aligned with Freebase relations that our approach is able to efficiently use the extra information provided by a large subset of Freebase data (4M entities, 23k relationships) to improve over methods that rely on text features alone.
Jason Weston, Antoine Bordes, Oksana Yakhnenko, Nicolas Usunier
EMNLP1
2013 Label Partitioning For Sublinear Ranking
abstract
We consider the case of ranking a very large set of labels, items, or documents, which is common to information retrieval, recommendation, and large-scale annotation tasks. We present a general approach for converting an algorithm which has linear time in the size of the set to a sublinear one via label partitioning. Our method consists of learning an input partition and a label assignment to each partition of the space such that precision at k is optimized, which is the loss function of interest in this setting. Experiments on large-scale ranking and recommendation tasks show that our method not only makes the original linear time algorithm computationally tractable, but can also improve its performance.
Jason Weston, Ameesh Makadia, Hector Yee
ICML (2)1
2013 Translating Embeddings for Modeling Multi-relational Data
abstract
We consider the problem of embedding entities and relationships of multi-relational data in low-dimensional vector spaces. Our objective is to propose a canonical model which is easy to train, contains a reduced number of parameters and can scale up to very large databases. Hence, we propose, TransE, a method which models relationships by interpreting them as translations operating on the low-dimensional embeddings of the entities. Despite its simplicity, this assumption proves to be powerful since extensive experiments show that TransE significantly outperforms state-of-the-art methods in link prediction on two knowledge bases. Besides, it can be successfully trained on a large scale data set with 1M entities, 25k relationships and more than 17M training samples.
Antoine Bordes, Nicolas Usunier, Alberto García-Durán, Jason Weston, Oksana Yakhnenko
NIPS4
2013 Nonlinear latent factorization by embedding multiple user interests
abstract
Classical matrix factorization approaches to collaborative filtering learn a latent vector for each user and each item, and recommendations are scored via the similarity between two such vectors, which are of the same dimension. In this work, we are motivated by the intuition that a user is a much more complicated entity than any single item, and cannot be well described by the same representation. Hence, the variety of a user's interests could be better captured by a more complex representation. We propose to model the user with a richer set of functions, specifically via a set of latent vectors, where each vector captures one of the user's latent interests or tastes. The overall recommendation model is then nonlinear where the matching score between a user and a given item is the maximum matching score over each of the user's latent interests with respect to the item's latent representation. We describe a simple, general and efficient algorithm for learning such a model, and apply it to large scale, real-world datasets from YouTube and Google Music, where our approach outperforms existing techniques.
Jason Weston, Ron J. Weiss, Hector Yee
RecSys1
2013 Learning to rank recommendations with the k-order statistic loss
abstract
Making recommendations by learning to rank is becoming an increasingly studied area. Approaches that use stochastic gradient descent scale well to large collaborative filtering datasets, and it has been shown how to approximately optimize the mean rank, or more recently the top of the ranked list. In this work we present a family of loss functions, the k-order statistic loss, that includes these previous approaches as special cases, and also derives new ones that we show to be useful. In particular, we present (i) a new variant that more accurately optimizes precision at k, and (ii) a novel procedure of optimizing the mean maximum rank, which we hypothesize is useful to more accurately cover all of the user's tastes. The general approach works by sampling N positive items, ordering them by the score assigned by the model, and then weighting the example as a function of this ordered set. Our approach is studied in two real-world systems, Google Music and YouTube video recommendations, where we obtain improvements for computable metrics, and in the YouTube case, increased user click through and watch duration when deployed live on www.youtube.com.
Jason Weston, Hector Yee, Ron J. Weiss
RecSys1
2012 Joint Image and Word Sense Discrimination for Image Retrieval
Aurélien Lucchi, Jason Weston
ECCV (1)2
2012 Learning improved linear transforms for speech recognition
abstract
This paper explores a novel large margin approach to learning a linear transform for dimensionality reduction in speech recognition. The method assumes a trained Gaussian mixture model for each class to be discriminated and trains a dimensionality-reducing linear transform with respect to the fixed model, optimizing a hinge loss on the difference between the distance to the nearest in- and out-of-class Gaussians using stochastic gradient descent. Results are presented showing that the learnt transform improves state classification for individual frames and reduces word error rate compared to Linear Discriminant Analysis (LDA) in a large vocabulary speech recognition problem even after discriminative training.
Andrew W. Senior, Youngmin Cho, Jason Weston
ICASSP3
2012 Latent Collaborative Retrieval
Jason Weston, Ron J. Weiss, Adam Berenzweig
ICML1
2012 Latent Structured Ranking
Jason Weston, John Blitzer
UAI1
2011 Learning Structured Embeddings of Knowledge Bases
abstract
Many Knowledge Bases (KBs) are now readily available and encompass colossal quantities of information thanks to either a long-term funding effort (e.g. WordNet, OpenCyc) or a collaborative process (e.g. Freebase, DBpedia). However, each of them is based on a different rigorous symbolic framework which makes it hard to use their data in other systems. It is unfortunate because such rich structured knowledge might lead to a huge leap forward in many other areas of AI like nat- ural language processing (word-sense disambiguation, natural language understanding, ...), vision (scene classification, image semantic annotation, ...) or collaborative filtering. In this paper, we present a learning process based on an innovative neural network architecture designed to embed any of these symbolic representations into a more flexible continuous vector space in which the original knowledge is kept and enhanced. These learnt embeddings would allow data from any KB to be easily used in recent machine learning meth- ods for prediction and information retrieval. We illustrate our method on WordNet and Freebase and also present a way to adapt it to knowledge extraction from raw text.
Antoine Bordes, Jason Weston, Ronan Collobert, Yoshua Bengio
AAAI2
2011 WSABIE: Scaling Up to Large Vocabulary Image Annotation
Jason Weston, Samy Bengio, Nicolas Usunier
IJCAI1
2011 Natural Language Processing (Almost) from Scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, Pavel P. Kuksa
J. Mach. Learn. Res.2
2011 Detecting Remote Evolutionary Relationships among Proteins by Large-Scale Semantic Embedding
abstract
Virtually every molecular biologist has searched a protein or DNA sequence database to find sequences that are evolutionarily related to a given query. Pairwise sequence comparison methods--i.e., measures of similarity between query and target sequences--provide the engine for sequence database search and have been the subject of 30 years of computational research. For the difficult problem of detecting remote evolutionary relationships between protein sequences, the most successful pairwise comparison methods involve building local models (e.g., profile hidden Markov models) of protein sequences. However, recent work in massive data domains like web search and natural language processing demonstrate the advantage of exploiting the global structure of the data space. Motivated by this work, we present a large-scale algorithm called ProtEmbed, which learns an embedding of protein sequences into a low-dimensional "semantic space." Evolutionarily related proteins are embedded in close proximity, and additional pieces of evidence, such as 3D structural similarity or class labels, can be incorporated into the learning process. We find that ProtEmbed achieves superior accuracy to widely used pairwise sequence methods like PSI-BLAST and HHSearch for remote homology detection; it also outperforms our previous RankProp algorithm, which incorporates global structure in the form of a protein similarity network. Finally, the ProtEmbed embedding space can be visualized, both at the global level and local to a given query, yielding intuition about the structure of protein sequence space.
Iain Melvin, Jason Weston, William Stafford Noble, Christina S. Leslie
PLoS Comput. Biol.2
2010 Label Ranking under Ambiguous Supervision for Learning Semantic Correspondences
Antoine Bordes, Nicolas Usunier, Jason Weston
ICML3
2010 Label Embedding Trees for Large Multi-Class Tasks
abstract
Multi-class classification becomes challenging at test time when the number of classes is very large and testing against every possible class can become computationally infeasible. This problem can be alleviated by imposing (or learning) a structure over the set of classes. We propose an algorithm for learning a tree-structure of classifiers which, by optimizing the overall tree loss, provides superior accuracy to existing tree labeling methods. We also propose a method that learns to embed labels in a low dimensional space that is faster than non-embedding approaches and has superior accuracy to existing embedding approaches. Finally we combine the two ideas resulting in the label embedding tree that outperforms alternative methods including One-vs-Rest while being orders of magnitude faster.
Samy Bengio, Jason Weston, David Grangier
NIPS2
2010 Semi-supervised Abstraction-Augmented String Kernel for Multi-level Bio-Relation Extraction
Pavel P. Kuksa, Yanjun Qi, Ronan Collobert, Jason Weston, Vladimir Pavlovic 0001, Xia Ning
ECML/PKDD (2)5
2010 Semi-supervised multi-task learning for predicting interactions between HIV-1 and human proteins
abstract
MOTIVATION: Protein-protein interactions (PPIs) are critical for virtually every biological function. Recently, researchers suggested to use supervised learning for the task of classifying pairs of proteins as interacting or not. However, its performance is largely restricted by the availability of truly interacting proteins (labeled). Meanwhile, there exists a considerable amount of protein pairs where an association appears between two partners, but not enough experimental evidence to support it as a direct interaction (partially labeled). RESULTS: We propose a semi-supervised multi-task framework for predicting PPIs from not only labeled, but also partially labeled reference sets. The basic idea is to perform multi-task learning on a supervised classification task and a semi-supervised auxiliary task. The supervised classifier trains a multi-layer perceptron network for PPI predictions from labeled examples. The semi-supervised auxiliary task shares network layers of the supervised classifier and trains with partially labeled examples. Semi-supervision could be utilized in multiple ways. We tried three approaches in this article, (i) classification (to distinguish partial positives with negatives); (ii) ranking (to rate partial positive more likely than negatives); (iii) embedding (to make data clusters get similar labels). We applied this framework to improve the identification of interacting pairs between HIV-1 and human proteins. Our method improved upon the state-of-the-art method for this task indicating the benefits of semi-supervised multi-task learning using auxiliary information. AVAILABILITY: http://www.cs.cmu.edu/~qyj/HIVsemi.
Yanjun Qi, Öznur Tastan, Jaime G. Carbonell, Judith Klein-Seetharaman, Jason Weston
Bioinform.5
2010 Learning to rank with (a lot of) word features
Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Olivier Chapelle, Kilian Q. Weinberger
Inf. Retr.2
2010 Large scale image annotation: learning to rank with joint word-image embeddings
Jason Weston, Samy Bengio, Nicolas Usunier
Mach. Learn.1
2010 Semisupervised Neural Networks for Efficient Hyperspectral Image Classification
abstract
A framework for semisupervised remote sensing image classification based on neural networks is presented. The methodology consists of adding a flexible embedding regularizer to the loss function used for training neural networks. Training is done using stochastic gradient descent with additional balancing constraints to avoid falling into local minima. The method constitutes a generalization of both supervised and unsupervised methods and can handle millions of unlabeled samples. Therefore, the proposed approach gives rise to an operational classifier, as opposed to previously presented transductive or Laplacian support vector machines (TSVM or LapSVM, respectively). The proposed methodology constitutes a general framework for building computationally efficient semisupervised methods. The method is compared with LapSVM and TSVM in semisupervised scenarios, to SVM in supervised settings, and to online and batchk-means for unsupervised learning. Results demonstrate the improved classification accuracy and scalability of this approach on several hyperspectral image classification problems.
Frédéric Ratle, Gustau Camps-Valls, Jason Weston
IEEE Trans. Geosci. Remote. Sens.3
2009 Supervised semantic indexing
abstract
In this article we propose Supervised Semantic Indexing (SSI), an algorithm that is trained on (query, document) pairs of text documents to predict the quality of their match. Like Latent Semantic Indexing (LSI), our models take account of correlations between words (synonymy, polysemy). However, unlike LSI our models are trained with a supervised signal directly on the ranking task of interest, which we argue is the reason for our superior results. As the query and target texts are modeled separately, our approach is easily generalized to different retrieval tasks, such as online advertising placement. Dealing with models on all pairs of words features is computationally challenging. We propose several improvements to our basic model for addressing this issue, including low rank (but diagonal preserving) representations, and correlated feature hashing (CFH). We provide an empirical study of all these methods on retrieval tasks based on Wikipedia documents as well as an Internet advertisement task. We obtain state-of-the-art performance while providing realistically scalable methods.
Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Olivier Chapelle, Kilian Q. Weinberger
CIKM2
2009 Combining labeled and unlabeled data with word-class distribution learning
abstract
We describe a novel simple and highly scalable semi-supervised method called Word-Class Distribution Learning (WCDL), and apply it task of information extraction (IE) by utilizing unlabeled sentences to improve supervised classification methods. WCDL iteratively builds class label distributions for each word in the dictionary by averaging predicted labels over all cases in the unlabeled corpus, and re-training a base classifier adding these distributions as word features. In contrast, traditional self-training or co-training methods self-labeled examples (rather than features) which can degrade performance due to incestuous learning bias. WCDL exhibits robust behavior, and has no difficult parameters to tune. We applied our method on German and English name entity recognition (NER) tasks. WCDL shows improvements over self-training, multi-task semi-supervision or supervision alone, in particular yielding a state-of-the art 75.72 F1 score on the German NER task.
Yanjun Qi, Ronan Collobert, Pavel P. Kuksa, Koray Kavukcuoglu, Jason Weston
CIKM5
2009 Supervised Semantic Indexing
Jason Weston, Ronan Collobert, David Grangier
ECIR2
2009 Semi-Supervised Sequence Labeling with Self-Learned Features
abstract
Typical information extraction (IE) systems can be seen as tasks assigning labels to words in a natural language sequence. The performance is restricted by the availability of labeled words. To tackle this issue, we propose a semi-supervised approach to improve the sequence labeling procedure in IE through a class of algorithms with self-learned features (SLF). A supervised classifier can be trained with annotated text sequences and used to classify each word in a large set of unannotated sentences. By averaging predicted labels over all cases in the unlabeled corpus, SLF training builds class label distribution patterns for each word (or word attribute) in the dictionary and re-trains the current model iteratively adding these distributions as extra word features. Basic SLF models how likely a word could be assigned to target class types. Several extensions are proposed, such as learning words' class boundary distributions. SLF exhibits robust and scalable behaviour and is easy to tune. We applied this approach on four classical IE tasks: named entity recognition (German and English), part-of-speech tagging (English) and one gene name recognition corpus. Experimental results show effective improvements over the supervised baselines on all tasks. In addition, when compared with the closely related self-training idea, this approach shows favorable advantages.
Yanjun Qi, Pavel P. Kuksa, Ronan Collobert, Kunihiko Sadamasa, Koray Kavukcuoglu, Jason Weston
ICDM6
2009 Curriculum learning
abstract
Humans and animals learn much better when the examples are not randomly presented but organized in a meaningful order which illustrates gradually more concepts, and gradually more complex ones. Here, we formalize such training strategies in the context of machine learning, and call them "curriculum learning". In the context of recent research studying the difficulty of training in the presence of non-convex training criteria (for deep deterministic and stochastic neural networks), we explore curriculum learning in various set-ups. The experiments show that significant improvements in generalization can be achieved. We hypothesize that curriculum learning has both an effect on the speed of convergence of the training process to a minimum and, in the case of non-convex criteria, on the quality of the local minima obtained: curriculum learning can be seen as a particular form of continuation method (a general strategy for global optimization of non-convex functions).
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, Jason Weston
ICML4
2009 Deep learning from temporal coherence in video
abstract
This work proposes a learning method for deep architectures that takes advantage of sequential data, in particular from the temporal coherence that naturally exists in unlabeled video recordings. That is, two successive frames are likely to contain the same object or objects. This coherence is used as a supervisory signal over the unlabeled data, and is used to improve the performance on a supervised task of interest. We demonstrate the effectiveness of this method on some pose invariant object and face recognition tasks. 1.
Hossein Mobahi, Ronan Collobert, Jason Weston
ICML3
2009 Polynomial Semantic Indexing
abstract
We present a class of nonlinear (polynomial) models that are discriminatively trained to directly map from the word content in a query-document or document-document pair to a ranking score. Dealing with polynomial models on word features is computationally challenging. We propose a low rank (but diagonal preserving) representation of our polynomial models to induce feasible memory and computation requirements. We provide an empirical study on retrieval tasks based on Wikipedia documents, where we obtain state-of-the-art performance while providing realistically scalable methods.
Jason Weston, David Grangier, Ronan Collobert, Kunihiko Sadamasa, Yanjun Qi, Corinna Cortes, Mehryar Mohri
NIPS2
2009 RANKPROP: a web server for protein remote homology detection
abstract
UNLABELLED: We present a large-scale implementation of the Rankprop protein homology ranking algorithm in the form of an openly accessible web server. We use the NRDB40 PSI-BLAST all-versus-all protein similarity network of 1.1 million proteins to construct the graph for the Rankprop algorithm, whereas previously, results were only reported for a database of 108 000 proteins. We also describe two algorithmic improvements to the original algorithm, including propagation from multiple homologs of the query and better normalization of ranking scores, that lead to higher accuracy and to scores with a probabilistic interpretation. AVAILABILITY: The Rankprop web server and source code are available at http://rankprop.gs.washington.edu
Iain Melvin, Jason Weston, Christina S. Leslie, William Stafford Noble
Bioinform.2
2008 A unified architecture for natural language processing: deep neural networks with multitask learning
abstract
We describe a single convolutional neural network architecture that, given a sentence, outputs a host of language processing predictions: part-of-speech tags, chunks, named entity tags, semantic roles, semantically similar words and the likelihood that the sentence makes sense (grammatically and semantically) using a language model. The entire network is trained jointly on all these tasks using weight-sharing, an instance of multitask learning. All the tasks use labeled data except the language model which is learnt from unlabeled text and represents a novel form of semi-supervised learning for the shared tasks. We show how both multitask learning and semi-supervised learning improve the generalization of the shared tasks, resulting in state-of-the-art-performance.
Ronan Collobert, Jason Weston
ICML2
2008 Large scale manifold transduction
abstract
We show how the regularizer of Transductive Support Vector Machines (TSVM) can be trained by stochastic gradient descent for linear models and multi-layer architectures. The resulting methods can be trained online, have vastly superior training and testing speed to existing TSVM algorithms, can encode prior knowledge in the network architecture, and obtain competitive error rates. We then go on to propose a natural generalization of the TSVM loss function that takes into account neighborhood and manifold information directly, unifying the two-stage Low Density Separation method into a single criterion, and leading to state-of-the-art results.
Michael Karlen, Jason Weston, Ayse Erkan, Ronan Collobert
ICML2
2008 Deep learning via semi-supervised embedding
abstract
Abstract. We show how nonlinear embedding algorithms popular for use with “shallow ” semi-supervised learning techniques such as kernel methods can be easily applied to deep multi-layer architectures, either as a regularizer at the output layer, or on each layer of the architecture. This trick provides a simple alternative to existing approaches to deep learning whilst yielding competitive error rates compared to those methods, and existing shallow semi-supervised techniques.
Jason Weston, Frédéric Ratle, Ronan Collobert
ICML1
2008 Large-Scale Clustering through Functional Embedding
Frédéric Ratle, Jason Weston, Matthew L. Miller
ECML/PKDD (2)2
2008 Combining classifiers for improved classification of proteins from sequence or structure
abstract
BACKGROUND: Predicting a protein's structural or functional class from its amino acid sequence or structure is a fundamental problem in computational biology. Recently, there has been considerable interest in using discriminative learning algorithms, in particular support vector machines (SVMs), for classification of proteins. However, because sufficiently many positive examples are required to train such classifiers, all SVM-based methods are hampered by limited coverage. RESULTS: In this study, we develop a hybrid machine learning approach for classifying proteins, and we apply the method to the problem of assigning proteins to structural categories based on their sequences or their 3D structures. The method combines a full-coverage but lower accuracy nearest neighbor method with higher accuracy but reduced coverage multiclass SVMs to produce a full coverage classifier with overall improved accuracy. The hybrid approach is based on the simple idea of "punting" from one method to another using a learned threshold. CONCLUSION: In cross-validated experiments on the SCOP hierarchy, the hybrid methods consistently outperform the individual component methods at all levels of coverage. Code and data sets are available at http://noble.gs.washington.edu/proj/sabretooth.
Iain Melvin, Jason Weston, Christina S. Leslie, William Stafford Noble
BMC Bioinform.2
2007 Fast Semantic Extraction Using a Novel Neural Network Architecture
Ronan Collobert, Jason Weston
ACL2
2007 Solving multiclass support vector machines with LaRank
abstract
Optimization algorithms for large margin multiclass recognizers are often too costly to handle ambitious problems with structured outputs and exponential numbers of classes. Optimization algorithms that rely on the full gradient are not effective because, unlike the solution, the gradient is not sparse and is very large. The LaRank algorithm sidesteps this difficulty by relying on a randomized exploration inspired by the perceptron algorithm. We show that this approach is competitive with gradient based optimizers on simple multiclass problems. Furthermore, a single LaRank pass over the training examples delivers test error rates that are nearly as good as those of the final solution.
Antoine Bordes, Léon Bottou, Patrick Gallinari, Jason Weston
ICML4
2007 SVM-Fold: a tool for discriminative multi-class protein fold and superfamily recognition
abstract
BACKGROUND: Predicting a protein's structural class from its amino acid sequence is a fundamental problem in computational biology. Much recent work has focused on developing new representations for protein sequences, called string kernels, for use with support vector machine (SVM) classifiers. However, while some of these approaches exhibit state-of-the-art performance at the binary protein classification problem, i.e. discriminating between a particular protein class and all other classes, few of these studies have addressed the real problem of multi-class superfamily or fold recognition. Moreover, there are only limited software tools and systems for SVM-based protein classification available to the bioinformatics community. RESULTS: We present a new multi-class SVM-based protein fold and superfamily recognition system and web server called SVM-Fold, which can be found at http://svm-fold.c2b2.columbia.edu. Our system uses an efficient implementation of a state-of-the-art string kernel for sequence profiles, called the profile kernel, where the underlying feature representation is a histogram of inexact matching k-mer frequencies. We also employ a novel machine learning approach to solve the difficult multi-class problem of classifying a sequence of amino acids into one of many known protein structural classes. Binary one-vs-the-rest SVM classifiers that are trained to recognize individual structural classes yield prediction scores that are not comparable, so that standard "one-vs-all" classification fails to perform well. Moreover, SVMs for classes at different levels of the protein structural hierarchy may make useful predictions, but one-vs-all does not try to combine these multiple predictions. To deal with these problems, our method learns relative weights between one-vs-the-rest classifiers and encodes information about the protein structural hierarchy for multi-class prediction. In large-scale benchmark results based on the SCOP database, our code weighting approach significantly improves on the standard one-vs-all method for both the superfamily and fold prediction in the remote homology setting and on the fold recognition problem. Moreover, our code weight learning algorithm strongly outperforms nearest-neighbor methods based on PSI-BLAST in terms of prediction accuracy on every structure classification problem we consider. CONCLUSION: By combining state-of-the-art SVM kernel methods with a novel multi-class algorithm, the SVM-Fold system delivers efficient and accurate protein fold and superfamily recognition.
Iain Melvin, Eugene Ie, Rui Kuang, Jason Weston, William Stafford Noble, Christina S. Leslie
BMC Bioinform.4
2007 Multi-class Protein Classification Using Adaptive Codes
Iain Melvin, Eugene Ie, Jason Weston, William Stafford Noble, Christina S. Leslie
J. Mach. Learn. Res.3
2007 The Need for Open Source Software in Machine Learning
Sören Sonnenburg, Mikio L. Braun, Cheng Soon Ong, Samy Bengio, Léon Bottou, Geoff Holmes 0001, Yann LeCun, Klaus-Robert Müller, Fernando Pereira 0003, Carl E. Rasmussen, Gunnar Rätsch, Bernhard Schölkopf, Alexander J. Smola, Pascal Vincent, Jason Weston, Robert C. Williamson
J. Mach. Learn. Res.15
2006 Trading convexity for scalability
abstract
Convex learning algorithms, such as Support Vector Machines (SVMs), are often seen as highly desirable because they offer strong practical properties and are amenable to theoretical analysis. However, in this work we show how non-convexity can provide scalability advantages over convexity. We show how concave-convex programming can be applied to produce (i) faster SVMs where training errors are no longer support vectors, and (ii) much faster Transductive SVMs.
Ronan Collobert, Fabian H. Sinz, Jason Weston, Léon Bottou
ICML3
2006 Inference with the Universum
abstract
In this paper we study a new framework introduced by Vapnik (1998) and Vapnik (2006) that is an alternative capacity concept to the large margin approach. In the particular case of binary classification, we are given a set of labeled examples, and a collection of "non-examples" that do not belong to either class of interest. This collection, called the Universum, allows one to encode prior knowledge by representing meaningful concepts in the same domain as the problem at hand. We describe an algorithm to leverage the Universum by maximizing the number of observed contradictions, and show experimentally that this approach delivers accuracy improvements over using labeled data alone.
Jason Weston, Ronan Collobert, Fabian H. Sinz, Léon Bottou, Vladimir Vapnik
ICML1
2006 Protein Ranking by Semi-Supervised Network Propagation
abstract
BACKGROUND: Biologists regularly search DNA or protein databases for sequences that share an evolutionary or functional relationship with a given query sequence. Traditional search methods, such as BLAST and PSI-BLAST, focus on detecting statistically significant pairwise sequence alignments and often miss more subtle sequence similarity. Recent work in the machine learning community has shown that exploiting the global structure of the network defined by these pairwise similarities can help detect more remote relationships than a purely local measure. METHODS: We review RankProp, a ranking algorithm that exploits the global network structure of similarity relationships among proteins in a database by performing a diffusion operation on a protein similarity network with weighted edges. The original RankProp algorithm is unsupervised. Here, we describe a semi-supervised version of the algorithm that uses labeled examples. Three possible ways of incorporating label information are considered: (i) as a validation set for model selection, (ii) to learn a new network, by choosing which transfer function to use for a given query, and (iii) to estimate edge weights, which measure the probability of inferring structural similarity. RESULTS: Benchmarked on a human-curated database of protein structures, the original RankProp algorithm provides significant improvement over local network search algorithms such as PSI-BLAST. Furthermore, we show here that labeled data can be used to learn a network without any need for estimating parameters of the transfer function, and that diffusion on this learned network produces better results than the original RankProp algorithm with a fixed network. CONCLUSION: In order to gain maximal information from a network, labeled and unlabeled data should be used to extract both local and global structure.
Jason Weston, Rui Kuang, Christina S. Leslie, William Stafford Noble
BMC Bioinform.1
2006 Large Scale Transductive SVMs
abstract
We show how the concave-convex procedure can be applied to transductive SVMs, which traditionally require solving a combinatorial search problem. This provides for the first time a highly scalable algorithm in the nonlinear case. Detailed experiments verify the utility of our approach. Software is available at http://www.kyb.tuebingen.mpg.de/bs/people/fabee/transduction.html.
Ronan Collobert, Fabian H. Sinz, Jason Weston, Léon Bottou
J. Mach. Learn. Res.3
2005 A general regression technique for learning transductions
abstract
The problem of learning a transduction, that is a string-to-string mapping, is a common problem arising in natural language processing and computational biology. Previous methods proposed for learning such mappings are based on classification techniques. This paper presents a new and general regression technique for learning transductions and reports the results of experiments showing its effectiveness. Our transduction learning consists of two phases: the estimation of a set of regression coefficients and the computation of the pre-image corresponding to this set of coefficients. A novel and conceptually cleaner formulation of kernel dependency estimation provides a simple framework for estimating the regression coefficients, and an efficient algorithm for computing the pre-image from the regression coefficients extends the applicability of kernel dependency estimation to output sequences. We report the results of a series of experiments illustrating the application of our regression technique for learning transductions.
Corinna Cortes, Mehryar Mohri, Jason Weston
ICML3
2005 Multi-class protein fold recognition using adaptive codes
abstract
We develop a novel multi-class classification method based on output codes for the problem of classifying a sequence of amino acids into one of many known protein structural classes, called folds. Our method learns relative weights between one-vs-all classifiers and encodes information about the protein structural hierarchy for multi-class prediction. Our code weighting approach significantly improves on the standard one-vs-all method for the fold recognition problem. In order to compare against widely used methods in protein sequence analysis, we also test nearest neighbor approaches based on the PSI-BLAST algorithm. Our code weight learning algorithm strongly outperforms these PSI-BLAST methods on every structure recognition problem we consider. 1.
Eugene Ie, Jason Weston, William Stafford Noble, Christina S. Leslie
ICML2
2005 Motif-based protein ranking by network propagation
abstract
MOTIVATION: Sequence similarity often suggests evolutionary relationships between protein sequences that can be important for inferring similarity of structure or function. The most widely-used pairwise sequence comparison algorithms for homology detection, such as BLAST and PSI-BLAST, often fail to detect less conserved remotely-related targets. RESULTS: In this paper, we propose a new general graph-based propagation algorithm called MotifProp to detect more subtle similarity relationships than pairwise comparison methods. MotifProp is based on a protein-motif network, in which edges connect proteins and the k-mer based motif features that they contain. We show that our new motif-based propagation algorithm can improve the ranking results over a base algorithm, such as PSI-BLAST, that is used to initialize the ranking. Despite the complex structure of the protein-motif network, MotifProp can be easily interpreted using the top-ranked motifs and motif-rich regions induced by the propagation, both of which are helpful for discovering conserved structural components in remote homologies.
Rui Kuang, Jason Weston, William Stafford Noble, Christina S. Leslie
Bioinform.2
2005 Semi-supervised protein classification using cluster kernels
abstract
MOTIVATION: Building an accurate protein classification system depends critically upon choosing a good representation of the input sequences of amino acids. Recent work using string kernels for protein data has achieved state-of-the-art classification performance. However, such representations are based only on labeled data--examples with known 3D structures, organized into structural classes--whereas in practice, unlabeled data are far more plentiful. RESULTS: In this work, we develop simple and scalable cluster kernel techniques for incorporating unlabeled data into the representation of protein sequences. We show that our methods greatly improve the classification performance of string kernels and outperform standard approaches for using unlabeled data, such as adding close homologs of the positive examples to the training data. We achieve equal or superior performance to previously presented cluster kernel methods and at the same time achieving far greater computational efficiency. AVAILABILITY: Source code is available at www.kyb.tuebingen.mpg.de/bs/people/weston/semiprot. The Spider matlab package is available at www.kyb.tuebingen.mpg.de/bs/people/spider. SUPPLEMENTARY INFORMATION: www.kyb.tuebingen.mpg.de/bs/people/weston/semiprot.
Jason Weston, Christina S. Leslie, Eugene Ie, Dengyong Zhou, André Elisseeff, William Stafford Noble
Bioinform.1
2005 Fast Kernel Classifiers with Online and Active Learning
abstract
Very high dimensional learning systems become theoretically possible when training examples are abundant. The computing cost then becomes the limiting factor. Any efficient learning algorithm should at least take a brief look at each example. But should all examples be given equal attention? This contribution proposes an empirical answer. We first present an online SVM algorithm based on this premise. LASVM yields competitive misclassification rates after a single pass over the training examples, outspeeding state-of-the-art SVM solvers. Then we show how active example selection can yield faster training, higher accuracies, and simpler models, using only a fraction of the training example labels.
Antoine Bordes, Seyda Ertekin, Jason Weston, Léon Bottou
J. Mach. Learn. Res.3
2004 Breaking SVM Complexity with Cross-Training
abstract
We propose to selectively remove examples from the training set using probabilistic estimates related to editing algorithms (Devijver and Kittler, 1982). This heuristic procedure aims at creating a separable distribution of training examples with minimal impact on the position of the decision boundary. It breaks the linear dependency between the number of SVs and the number of training examples, and sharply reduces the complexity of SVMs during both the training and prediction stages.
Gökhan H. Bakir, Léon Bottou, Jason Weston
NIPS3
2004 Mismatch string kernels for discriminative protein classification
abstract
MOTIVATION: Classification of proteins sequences into functional and structural families based on sequence homology is a central problem in computational biology. Discriminative supervised machine learning approaches provide good performance, but simplicity and computational efficiency of training and prediction are also important concerns. RESULTS: We introduce a class of string kernels, called mismatch kernels, for use with support vector machines (SVMs) in a discriminative approach to the problem of protein classification and remote homology detection. These kernels measure sequence similarity based on shared occurrences of fixed-length patterns in the data, allowing for mutations between patterns. Thus, the kernels provide a biologically well-motivated way to compare protein sequences without relying on family-based generative models such as hidden Markov models. We compute the kernels efficiently using a mismatch tree data structure, allowing us to calculate the contributions of all patterns occurring in the data in one pass while traversing the tree. When used with an SVM, the kernels enable fast prediction on test sequences. We report experiments on two benchmark SCOP datasets, where we show that the mismatch kernel used with an SVM classifier performs competitively with state-of-the-art methods for homology detection, particularly when very few training examples are available. Examination of the highest-weighted patterns learned by the SVM classifier recovers biologically important motifs in protein families and superfamilies.
Christina S. Leslie, Eleazar Eskin, Adiel Cohen, Jason Weston, William Stafford Noble
Bioinform.4
2003 Learning to Find Pre-Images
abstract
We consider the problem of reconstructing patterns from a feature map. Learning algorithms using kernels to operate in a reproducing kernel Hilbert space (RKHS) express their solutions in terms of input points mapped into the RKHS. We introduce a technique based on kernel princi- pal component analysis and regression to reconstruct corresponding pat- terns in the input space (aka pre-images) and review its performance in several applications requiring the construction of pre-images. The intro- duced technique avoids difficult and/or unstable numerical optimization, is easy to implement and, unlike previous methods, permits the compu- tation of pre-images in discrete input spaces.
Gökhan H. Bakir, Jason Weston, Bernhard Schölkopf
NIPS2
2003 Prediction on Spike Data Using Kernel Algorithms
abstract
We report and compare the performance of different learning algorithms based on data from cortical recordings. The task is to predict the orienta- tion of visual stimuli from the activity of a population of simultaneously recorded neurons. We compare several ways of improving the coding of the input (i.e., the spike data) as well as of the output (i.e., the orienta- tion), and report the results obtained using different kernel algorithms.
Jan Eichhorn, Andreas S. Tolias, Alexander Zien, Malte Kuss, Carl E. Rasmussen, Jason Weston, Nikos K. Logothetis, Bernhard Schölkopf
NIPS6
2003 Semi-supervised Protein Classification Using Cluster Kernels
abstract
A key issue in supervised protein classification is the representation of in- put sequences of amino acids. Recent work using string kernels for pro- tein data has achieved state-of-the-art classification performance. How- ever, such representations are based only on labeled data — examples with known 3D structures, organized into structural classes — while in practice, unlabeled data is far more plentiful. In this work, we de- velop simple and scalable cluster kernel techniques for incorporating un- labeled data into the representation of protein sequences. We show that our methods greatly improve the classification performance of string ker- nels and outperform standard approaches for using unlabeled data, such as adding close homologs of the positive examples to the training data. We achieve equal or superior performance to previously presented cluster kernel methods while achieving far greater computational efficiency.
Jason Weston, Christina S. Leslie, Dengyong Zhou, André Elisseeff, William Stafford Noble
NIPS1
2003 Learning with Local and Global Consistency
abstract
We consider the general problem of learning from labeled and unlabeled data, which is often called semi-supervised learning or transductive in- ference. A principled approach to semi-supervised learning is to design a classifying function which is suf(cid:2)ciently smooth with respect to the intrinsic structure collectively revealed by known labeled and unlabeled points. We present a simple algorithm to obtain such a smooth solution. Our method yields encouraging experimental results on a number of clas- si(cid:2)cation problems and demonstrates effective use of unlabeled data.
Dengyong Zhou, Olivier Bousquet, Thomas Navin Lal, Jason Weston, Bernhard Schölkopf
NIPS4
2003 Ranking on Data Manifolds
abstract
The Google search engine has enjoyed huge success with its web page ranking algorithm, which exploits global, rather than local, hyperlink structure of the web using random walks. Here we propose a simple universal ranking algorithm for data lying in the Euclidean space, such as text or image data. The core idea of our method is to rank the data with respect to the intrinsic manifold structure collectively revealed by a great amount of data. Encouraging experimental results from synthetic, image, and text data illustrate the validity of our method.
Dengyong Zhou, Jason Weston, Arthur Gretton, Olivier Bousquet, Bernhard Schölkopf
NIPS2
2003 Feature selection and transduction for prediction of molecular bioactivity for drug design
abstract
Abstract Motivation: In drug discovery a key task is to identify characteristics that separate active (binding) compounds from inactive (non-binding) ones. An automated prediction system can help reduce resources necessary to carry out this task. Results: Two methods for prediction of molecular bioactivity for drug design are introduced and shown to perform well in a data set previously studied as part of the KDD (Knowledge Discovery and Data Mining) Cup 2001. The data is characterized by very few positive examples, a very large number of features (describing three-dimensional properties of the molecules) and rather different distributions between training and test data. Two techniques are introduced specifically to tackle these problems: a feature selection method for unbalanced data and a classifier which adapts to the distribution of the the unlabeled test data (a so-called transductive method). We show both techniques improve identification performance and in conjunction provide an improvement over using only one of the techniques. Our results suggest the importance of taking into account the characteristics in this data which may also be relevant in other problems of a similar type. Availability: Matlab source code is available at http://www.kyb.tuebingen.mpg.de/bs/people/weston/kdd/kdd.html Contact: [email protected] Supplementary information: Supplementary material is available at http://www.kyb.tuebingen.mpg.de/bs/people/weston/kdd/kdd.html. * To whom correspondence should be addressed.
Jason Weston, Fernando Pérez-Cruz, Olivier Bousquet, Olivier Chapelle, André Elisseeff, Bernhard Schölkopf
Bioinform.1
2003 Use of the Zero-Norm with Linear Models and Kernel Methods
Jason Weston, André Elisseeff, Bernhard Schölkopf, Michael E. Tipping
J. Mach. Learn. Res.1
2003 Constructing Descriptive and Discriminative Nonlinear Features: Rayleigh Coefficients in Kernel Feature Spaces
abstract
We incorporate prior knowledge to construct nonlinear algorithms for invariant feature extraction and discrimination. Employing a unified framework in terms of a nonlinearized variant of the Rayleigh coefficient, we propose nonlinear generalizations of Fisher's discriminant and oriented PCA using support vector kernel functions. Extensive simulations show the utility of our approach.
Sebastian Mika, Gunnar Rätsch, Jason Weston, Bernhard Schölkopf, Alexander J. Smola, Klaus-Robert Müller
IEEE Trans. Pattern Anal. Mach. Intell.3
2002 A Kernel Approach for Learning from almost Orthogonal Patterns
Bernhard Schölkopf, Jason Weston, Eleazar Eskin, Christina S. Leslie, William Stafford Noble
ECML2
2002 Cluster Kernels for Semi-Supervised Learning
abstract
We propose a framework to incorporate unlabeled data in kernel classifier, based on the idea that two points in the same cluster are more likely to have the same label. This is achieved by modifying the eigenspectrum of the kernel matrix. Experimental results assess the validity of this approach.
Olivier Chapelle, Jason Weston, Bernhard Schölkopf
NIPS2
2002 Mismatch String Kernels for SVM Protein Classification
abstract
We introduce a class of string kernels, called mismatch kernels, for use with support vector machines (SVMs) in a discriminative approach to the protein classification problem. These kernels measure sequence sim- ilarity based on shared occurrences of  -length subsequences, counted with up to mismatches, and do not rely on any generative model for the positive training sequences. We compute the kernels efficiently using a mismatch tree data structure and report experiments on a benchmark SCOP dataset, where we show that the mismatch kernel used with an SVM classifier performs as well as the Fisher kernel, the most success- ful method for remote homology detection, while achieving considerable computational savings.
Christina S. Leslie, Eleazar Eskin, Jason Weston, William Stafford Noble
NIPS3
2002 Kernel Dependency Estimation
abstract
We consider the learning problem of finding a dependency between a general class of objects and another, possibly different, general class of objects. The objects can be for example: vectors, images, strings, trees or graphs. Such a task is made possible by employing similarity measures in both input and output spaces using ker(cid:173) nel functions, thus embedding the objects into vector spaces. We experimentally validate our approach on several tasks: mapping strings to strings, pattern recognition, and reconstruction from par(cid:173) tial images.
Jason Weston, Olivier Chapelle, André Elisseeff, Bernhard Schölkopf, Vladimir Vapnik
NIPS1
2002 A Kernel Approach for Learning from Almost Orthogonal Patterns
Bernhard Schölkopf, Jason Weston, Eleazar Eskin, Christina S. Leslie, William Stafford Noble
PKDD2
2002 Gene Selection for Cancer Classification using Support Vector Machines
Isabelle Guyon, Jason Weston, Stephen Barnhill, Vladimir Vapnik
Mach. Learn.2
2001 A kernel method for multi-labelled classification
abstract
This article presents a Support Vector Machine (SVM) like learning sys- tem to handle multi-label problems. Such problems are usually decom- posed into many two-class problems but the expressive power of such a system can be weak [5, 7]. We explore a new direct approach. It is based on a large margin ranking system that shares a lot of common proper- ties with SVMs. We tested it on a Yeast gene functional classification problem with positive results.
André Elisseeff, Jason Weston
NIPS2
2001 Gene functional classification from heterogeneous data
abstract
In our attempts to understand cellular function at the molecular level, we must be able to synthesize information from disparate types of genomic data. We consider the problem of inferring gene functional classifications from a heterogeneous data set consisting of DNA microarray expression measurements and phylogenetic profiles from whole-genome sequence comparisons. We demonstrate the application of the support vector machine (SVM) learning algorithm to this functional inference task. Our results suggest the importance of exploiting prior information about the heterogeneity of the data. In particular, we propose an SVM kernel function that is explicitly heterogeneous. We also show how to use knowledge about heterogeneity to aid in feature selection.
Paul Pavlidis, Jason Weston, Jinsong Cai, William Stafford Noble
RECOMB2
2000 Vicinal Risk Minimization
abstract
The Vicinal Risk Minimization principle establishes a bridge between generative models and methods derived from the Structural Risk Mini(cid:173) mization Principle such as Support Vector Machines or Statistical Reg(cid:173) ularization. We explain how VRM provides a framework which inte(cid:173) grates a number of existing algorithms, such as Parzen windows, Support Vector Machines, Ridge Regression, Constrained Logistic Classifiers and Tangent-Prop. We then show how the approach implies new algorithm(cid:173) s for solving problems usually associated with generative models. New algorithms are described for dealing with pattern recognition problems with very different pattern distributions and dealing with unlabeled data. Preliminary empirical results are presented.
Olivier Chapelle, Jason Weston, Léon Bottou, Vladimir Vapnik
NIPS2
2000 Feature Selection for SVMs
abstract
We introduce a method of feature selection for Support Vector Machines. The method is based upon finding those features which minimize bounds on the leave-one-out error. This search can be efficiently performed via gradient descent. The resulting algorithms are shown to be superior to some standard feature selection algorithms on both toy data and real-life problems of face recognition, pedestrian detection and analyzing DNA micro array data.
Jason Weston, Sayan Mukherjee 0001, Olivier Chapelle, Massimiliano Pontil, Tomaso A. Poggio, Vladimir Vapnik
NIPS1
1999 Support vector machines for multi-class pattern recognition
Jason Weston, Chris Watkins
ESANN1
1999 Leave-One-Out Support Vector Machines
Jason Weston
IJCAI1
1999 Transductive Inference for Estimating Values of Functions
Olivier Chapelle, Vladimir Vapnik, Jason Weston
NIPS3
1999 Invariant Feature Extraction and Classification in Kernel Spaces
Sebastian Mika, Gunnar Rätsch, Jason Weston, Bernhard Schölkopf, Alexander J. Smola, Klaus-Robert Müller
NIPS3