Richard Socher

dblp:79/128 · DBLP profile ↗
← Back
94ranked-venue papers
13as first author
7since 2021 · last 2021
0000-0002-3577-639XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 92 · 13 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
84 papers
Question answering and dialogue systems · 14% Deep learning architectures and training · 8% Language models and text generation · 8%
Databases, data mining, and information retrieval
5 papers
Data models and query languages · 47% Information retrieval · 33% Knowledge graphs · 20%

Topics — the 30 heaviest of 189, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
1.742020
A Simple Language Model for Task-Oriented Dialogue · NeurIPS 2020
Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference · EMNLP (1) 2020
TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented Dialogue · EMNLP (1) 2020
Natural language and speech › Question answering and dialogue systems › task-oriented dialogue
dialogue state tracking
1.542020
Non-Autoregressive Dialog State Tracking · ICLR 2020
Global-to-local Memory Pointer Networks for Task-Oriented Dialogue · ICLR (Poster) 2019
Transferable Multi-Domain State Generator for Task-Oriented Dialogue Systems · ACL (1) 2019
Machine learning › Trustworthy machine learning
interpretability
1.332021
BERTology Meets Biology: Interpreting Attention in Protein Language Models · ICLR 2021
ERASER: A Benchmark to Evaluate Rationalized NLP Models · ACL 2020
Interpretable Counting for Visual Question Answering · ICLR (Poster) 2018
Computer vision › Vision and language
visual question answering
1.242018
DCN+: Mixed Objective And Deep Residual Coattention for Question Answering · ICLR (Poster) 2018
Interpretable Counting for Visual Question Answering · ICLR (Poster) 2018
Dynamic Coattention Networks For Question Answering · ICLR (Poster) 2017
Natural language and speech › Language models and text generation
text summarization
1.132020
Evaluating the Factual Consistency of Abstractive Text Summarization · EMNLP (1) 2020
Neural Text Summarization: A Critical Evaluation · EMNLP/IJCNLP (1) 2019
Improving Abstraction in Text Summarization · EMNLP 2018
Machine learning › Trustworthy machine learning
robustness
1.032020
Learning From Noisy Anchors for One-Stage Object Detection · CVPR 2020
It's Morphin' Time! Combating Linguistic Discrimination with Inflectional Perturbations · ACL 2020
Efficient and Robust Question Answering from Minimal Context over Documents · ACL (1) 2018
Natural language and speech › Language models and text generation
language modeling
0.932018
Regularizing and Optimizing LSTM Language Models · ICLR (Poster) 2018
Pointer Sentinel Mixture Models · ICLR (Poster) 2017
Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling · ICLR (Poster) 2017
Natural language and speech › Information extraction and text analysis
semantic parsing
0.922021
GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing · ICLR 2021
SParC: Cross-Domain Semantic Parsing in Context · ACL (1) 2019
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.922020
DivideMix: Learning with Noisy Labels as Semi-supervised Learning · ICLR 2020
Learning From Noisy Anchors for One-Stage Object Detection · CVPR 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
explanation generation
0.822020
ESPRIT: Explaining Solutions to Physical Reasoning Tasks · ACL 2020
Explain Yourself! Leveraging Language Models for Commonsense Reasoning · ACL (1) 2019
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.822020
Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging · EMNLP (1) 2020
Augmented Cyclic Adversarial Learning for Low Resource Domain Adaptation · ICLR (Poster) 2019
Machine learning › Representation and self-supervised learning
pre-training
0.822021
GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing · ICLR 2021
Learned in Translation: Contextualized Word Vectors · NIPS 2017
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL
0.822019
Editing-Based SQL Query Generation for Cross-Domain Context-Dependent Questions · EMNLP/IJCNLP (1) 2019
CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases · EMNLP/IJCNLP (1) 2019
Computer vision › Vision and language › cross-modal attention
co-attention
0.622018
DCN+: Mixed Objective And Deep Residual Coattention for Question Answering · ICLR (Poster) 2018
Dynamic Coattention Networks For Question Answering · ICLR (Poster) 2017
Natural language and speech › Language models and text generation › text summarization
abstractive summarization
0.632020
A Deep Reinforced Model for Abstractive Summarization · ICLR (Poster) 2018
Evaluating the Factual Consistency of Abstractive Text Summarization · EMNLP (1) 2020
Neural Text Summarization: A Critical Evaluation · EMNLP/IJCNLP (1) 2019
Machine learning › Deep learning architectures and training
attention mechanism
0.522020
Tree-Structured Attention with Hierarchical Accumulation · ICLR 2020
Global-Locally Self-Attentive Encoder for Dialogue State Tracking · ACL (1) 2018
Natural language and speech › Information extraction and text analysis
textual entailment
0.522020
Universal Natural Language Processing with Limited Annotations: Try Few-shot Textual Entailment as a Start · EMNLP (1) 2020
A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks · EMNLP 2017
Natural language and speech › Language models and text generation
neural language model
0.522017
Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling · ICLR (Poster) 2017
Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks · ACL (1) 2015
Computer vision › Video understanding and tracking
action recognition
0.512021
WOAD: Weakly Supervised Online Action Detection in Untrimmed Videos · CVPR 2021
Machine learning › Trustworthy machine learning › interpretability
attention analysis
0.512021
BERTology Meets Biology: Interpreting Attention in Protein Language Models · ICLR 2021
Machine learning › Learning theory › classification › classification theory
bayes error rate
0.512021
Evaluating State-of-the-Art Classification Models Against Bayes Optimality · NeurIPS 2021
Machine learning › Learning theory › statistical learning theory › bayesian learning theory
bayes optimality
0.512021
Evaluating State-of-the-Art Classification Models Against Bayes Optimality · NeurIPS 2021
Machine learning › Learning theory
generalization
0.512021
Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts Generalization · ICML 2021
Machine learning › Optimization for machine learning
implicit regularization
0.512021
Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts Generalization · ICML 2021
Machine learning › Generative modeling › normalizing flow
invertible transformation
0.512021
Evaluating State-of-the-Art Classification Models Against Bayes Optimality · NeurIPS 2021
Machine learning › Deep learning architectures and training › loss landscape
loss landscape curvature
0.512021
Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts Generalization · ICML 2021
Machine learning › Generative modeling
normalizing flow
0.512021
Evaluating State-of-the-Art Classification Models Against Bayes Optimality · NeurIPS 2021
Computer vision › Video understanding and tracking › action detection
online action detection
0.512021
WOAD: Weakly Supervised Online Action Detection in Untrimmed Videos · CVPR 2021
Computer vision › Video understanding and tracking › action detection › temporal action localization
weakly supervised action detection
0.512021
WOAD: Weakly Supervised Online Action Detection in Untrimmed Videos · CVPR 2021
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model
0.512021
BERTology Meets Biology: Interpreting Attention in Protein Language Models · ICLR 2021

Methods — techniques the papers use, named apart from their topics

attention · 1.7LSTM · 1.2recursive neural network · 1.0neural network · 0.9weakly supervised learning · 0.8language model · 0.8data augmentation · 0.8attention mechanism · 0.8policy gradient · 0.8encoder-decoder · 0.7attention analysis · 0.5query-based extraction · 0.4pretrained multilingual initialization · 0.4editing-based query generation · 0.4cross-domain generalization · 0.4reward shaping · 0.3reinforcement learning · 0.3knowledge graph embedding · 0.3
YearPublicationVenuePosition
2021 WOAD: Weakly Supervised Online Action Detection in Untrimmed Videos
abstract
Online action detection in untrimmed videos aims to identify an action as it happens, which makes it very important for real-time applications. Previous methods rely on tedious annotations of temporal action boundaries for training, which hinders the scalability of online action detection systems. We propose WOAD, a weakly supervised framework that can be trained using only video-class labels. WOAD contains two jointly-trained modules, i.e., temporal proposal generator (TPG) and online action recognizer (OAR). Supervised by video-class labels, TPG works offline and targets at accurately mining pseudo frame-level labels for OAR. With the supervisory signals from TPG, OAR learns to conduct action detection in an online fashion. Experimental results on THUMOS’14, ActivityNet1.2 and ActivityNet1.3 show that our weakly-supervised method largely outperforms weakly-supervised baselines and achieves comparable performance to the previous strongly-supervised methods. Beyond that, WOAD is flexible to leverage strong supervision when it is available. When strongly supervised, our method obtains the state-of-the-art results in the tasks of both online per-frame action recognition and online detection of action start.
Mingfei Gao, Yingbo Zhou 0002, Ran Xu 0001, Richard Socher, Caiming Xiong
CVPR4
2021 GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing
Tao Yu 0009, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang 0002, Dragomir R. Radev, Richard Socher, Caiming Xiong
ICLR8
2021 BERTology Meets Biology: Interpreting Attention in Protein Language Models
Jesse Vig, Ali Madani, Lav R. Varshney, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani
ICLR5
2021 Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts Generalization
abstract
The early phase of training a deep neural network has a dramatic effect on the local curvature of the loss function. For instance, using a small learning rate does not guarantee stable optimization because the optimization trajectory has a tendency to steer towards regions of the loss surface with increasing local curvature. We ask whether this tendency is connected to the widely observed phenomenon that the choice of the learning rate strongly influences generalization. We first show that stochastic gradient descent (SGD) implicitly penalizes the trace of the Fisher Information Matrix (FIM), a measure of the local curvature, from the start of training. We argue it is an implicit regularizer in SGD by showing that explicitly penalizing the trace of the FIM can significantly improve generalization. We highlight that poor final generalization coincides with the trace of the FIM attaining a large value early in training, to which we refer as catastrophic Fisher explosion. Finally, to gain insight into the regularization effect of penalizing the trace of the FIM, we show that it limits memorization by reducing the learning speed of examples with noisy labels more than that of the examples with clean labels.
Stanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg, Huan Wang 0016, Caiming Xiong, Richard Socher, Kyunghyun Cho, Krzysztof J. Geras
ICML7
2021 DART: Open-Domain Structured Data Record to Text Generation
abstract
Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Linyong Nan, Dragomir R. Radev, Rui Zhang 0037, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma 0001, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta 0015, Tao Yu 0009, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, Nazneen Fatema Rajani
NAACL-HLT23
2021 Evaluating State-of-the-Art Classification Models Against Bayes Optimality
abstract
Evaluating the inherent difficulty of a given data-driven classification problem is important for establishing absolute benchmarks and evaluating progress in the field. To this end, a natural quantity to consider is the \emph{Bayes error}, which measures the optimal classification error theoretically achievable for a given data distribution. While generally an intractable quantity, we show that we can compute the exact Bayes error of generative models learned using normalizing flows. Our technique relies on a fundamental result, which states that the Bayes error is invariant under invertible transformation. Therefore, we can compute the exact Bayes error of the learned flow models by computing it for Gaussian base distributions, which can be done efficiently using Holmes-Diaconis-Ross integration. Moreover, we show that by varying the temperature of the learned flow models, we can generate synthetic datasets that closely resemble standard benchmark datasets, but with almost any desired Bayes error. We use our approach to conduct a thorough investigation of state-of-the-art classification models, and find that in some --- but not all --- cases, these models are capable of obtaining accuracy very near optimal. Finally, we use our method to evaluate the intrinsic "hardness" of standard benchmark datasets.
Ryan Theisen, Huan Wang 0016, Lav R. Varshney, Caiming Xiong, Richard Socher
NeurIPS5
2021 SummEval: Re-evaluating Summarization Evaluation
abstract
Abstract The scarcity of comprehensive up-to-date studies on evaluation metrics for text summarization and the lack of consensus regarding evaluation protocols continue to inhibit progress. We address the existing shortcomings of summarization evaluation methods along five dimensions: 1) we re-evaluate 14 automatic evaluation metrics in a comprehensive and consistent fashion using neural summarization model outputs along with expert and crowd-sourced human annotations; 2) we consistently benchmark 23 recent summarization models using the aforementioned automatic evaluation metrics; 3) we assemble the largest collection of summaries generated by models trained on the CNN/DailyMail news dataset and share it in a unified format; 4) we implement and share a toolkit that provides an extensible and unified API for evaluating summarization models across a broad range of automatic metrics; and 5) we assemble and share the largest and most diverse, in terms of model types, collection of human judgments of model-generated summaries on the CNN/Daily Mail dataset annotated by both expert judges and crowd-source workers. We hope that this work will help promote a more complete evaluation protocol for text summarization as well as advance research in developing evaluation metrics that better correlate with human judgments.
Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, Dragomir R. Radev
Trans. Assoc. Comput. Linguistics5
2020 ERASER: A Benchmark to Evaluate Rationalized NLP Models
abstract
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Jay DeYoung, Nazneen Fatema Rajani, Eric P. Lehman, Caiming Xiong, Richard Socher, Byron C. Wallace
ACL6
2020 Explicit Memory Tracker with Coarse-to-Fine Reasoning for Conversational Machine Reading
abstract
Yifan Gao, Chien-Sheng Wu, Shafiq Joty, Caiming Xiong, Richard Socher, Irwin King, Michael Lyu, Steven C.H. Hoi. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Yifan Gao 0001, Chien-Sheng Wu, Shafiq R. Joty, Caiming Xiong, Richard Socher, Irwin King, Michael R. Lyu, Steven C. H. Hoi
ACL5
2020 ESPRIT: Explaining Solutions to Physical Reasoning Tasks
abstract
Nazneen Fatema Rajani, Rui Zhang, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Nazneen Fatema Rajani, Rui Zhang 0037, Yi Chern Tan, Stephan Zheng, Jeremy Weiss, Aadit Vyas, Abhijit Gupta, Caiming Xiong, Richard Socher, Dragomir R. Radev
ACL9
2020 It's Morphin' Time! Combating Linguistic Discrimination with Inflectional Perturbations
abstract
Training on only perfect Standard English corpora predisposes pre-trained neural networks to discriminate against minorities from nonstandard linguistic backgrounds (e.g., African American Vernacular English, Colloquial Singapore English, etc.).We perturb the inflectional morphology of words to craft plausible and semantically similar adversarial examples that expose these biases in popular NLP models, e.g., BERT and Transformer, and show that adversarially fine-tuning them for a single epoch significantly improves robustness without sacrificing performance on clean data. 1
Samson Tan, Shafiq R. Joty, Min-Yen Kan, Richard Socher
ACL4
2020 Assessing Local Generalization Capability in Deep Models
abstract
While it has not yet been proven, empirical evidence suggests that model generalization is related to local properties of the optima, which can be described via the Hessian. We connect model generalization with the local property of a solution under the PAC-Bayes paradigm. In particular, we prove that model generalization ability is related to the Hessian, the higher-order “smoothness" terms characterized by the Lipschitz constant of the Hessian, and the scales of the parameters. Guided by the proof, we propose a metric to score the generalization capability of a model, as well as an algorithm that optimizes the perturbed model accordingly.
Huan Wang 0016, Nitish Shirish Keskar, Caiming Xiong, Richard Socher
AISTATS4
2020 Learning From Noisy Anchors for One-Stage Object Detection
abstract
State-of-the-art object detectors rely on regressing and classifying an extensive list of possible anchors, which are divided into positive and negative samples based on their intersection-over-union (IoU) with corresponding ground-truth objects. Such a harsh split conditioned on IoU results in binary labels that are potentially noisy and challenging for training. In this paper, we propose to mitigate noise incurred by imperfect label assignment such that the contributions of anchors are dynamically determined by a carefully constructed cleanliness score associated with each anchor. Exploring outputs from both regression and classification branches, the cleanliness scores, estimated without incurring any additional computational overhead, are used not only as soft labels to supervise the training of the classification branch but also sample re-weighting factors for improved localization and classification accuracy. We conduct extensive experiments on COCO, and demonstrate, among other things, the proposed approach steadily improves RetinaNet by ~2% with various backbones.
Hengduo Li, Zuxuan Wu, Chen Zhu 0001, Caiming Xiong, Richard Socher, Larry Davis 0001
CVPR5
2020 The Thieves on Sesame Street are Polyglots - Extracting Multilingual Models from Monolingual APIs
abstract
Pre-training in natural language processing makes it easier for an adversary with only query access to a victim model to reconstruct a local copy of the victim by training with gibberish input data paired with the victim's labels for that data.We discover that this extraction process extends to local copies initialized from a pre-trained, multilingual model while the victim remains monolingual.The extracted model learns the task from the monolingual victim, but it generalizes far better than the victim to several other languages.This is done without ever showing the multilingual, extracted model a well-formed input in any of the languages for the target task.We also demonstrate that a few real examples can greatly improve performance, and we analyze how these results shed light on how such extraction methods succeed.
Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, Richard Socher
EMNLP (1)4
2020 Evaluating the Factual Consistency of Abstractive Text Summarization
abstract
Currently used metrics for assessing summarization algorithms do not account for whether summaries are factually consistent with source documents.We propose a weakly-supervised, model-based approach for verifying factual consistency and identifying conflicts between source documents and a generated summary.Training data is generated by applying a series of rule-based transformations to the sentences of source documents.The factual consistency model is then trained jointly for three tasks: 1) identify whether sentences remain factually consistent after transformation, 2) extract a span in the source documents to support the consistency prediction, 3) extract a span in the summary sentence that is inconsistent if one exists.Transferring this model to summaries generated by several state-of-the art models reveals that this highly scalable approach substantially outperforms previous models, including those trained with strong supervision using standard datasets for natural language inference and fact checking.Additionally, human evaluation shows that the auxiliary span extraction tasks provide useful assistance in the process of verifying factual consistency.
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher
EMNLP (1)4
2020 TOD-BERT: Pre-trained Natural Language Understanding for Task-Oriented Dialogue
abstract
The underlying difference of linguistic patterns between general text and task-oriented dialogue makes existing pre-trained language models less useful in practice.In this work, we unify nine human-human and multi-turn task-oriented dialogue datasets for language modeling.To better model dialogue behavior during pre-training, we incorporate user and system tokens into the masked language modeling.We propose a contrastive objective function to simulate the response selection task.Our pre-trained task-oriented dialogue BERT (TOD-BERT) outperforms strong baselines like BERT on four downstream taskoriented dialogue applications, including intention recognition, dialogue state tracking, dialogue act prediction, and response selection.We also show that TOD-BERT has a stronger few-shot ability that can mitigate the data scarcity problem for task-oriented dialogue.
Chien-Sheng Wu, Steven C. H. Hoi, Richard Socher, Caiming Xiong
EMNLP (1)3
2020 Simple Data Augmentation with the Mask Token Improves Domain Adaptation for Dialog Act Tagging
abstract
The concept of Dialogue Act (DA) is universal across different task-oriented dialogue domains -the act of "request" carries the same speaker intention whether it is for restaurant reservation or flight booking.However, DA taggers trained on one domain do not generalize well to other domains, which leaves us with the expensive need for a large amount of annotated data in the target domain.In this work, we investigate how to better adapt DA taggers to desired target domains with only unlabeled data.We propose MASKAUGMENT, a controllable mechanism that augments text input by leveraging the pre-trained MASK token from BERT model.Inspired by consistency regularization, we use MASKAUGMENT to introduce an unsupervised teacher-student learning scheme to examine the domain adaptation of DA taggers.Our extensive experiments on the Simulated Dialogue (GSim) and Schema-Guided Dialogue (SGD) datasets show that MASKAUGMENT is useful in improving the cross-domain generalization for DA tagging.
Semih Yavuz, Kazuma Hashimoto, Wenhao Liu 0003, Nitish Shirish Keskar, Richard Socher, Caiming Xiong
EMNLP (1)5
2020 Universal Natural Language Processing with Limited Annotations: Try Few-shot Textual Entailment as a Start
abstract
A standard way to address different NLP problems is by first constructing a problem-specific dataset, then building a model to fit this dataset.To build the ultimate artificial intelligence, we desire a single machine that can handle diverse new problems, for which task-specific annotations are limited.We bring up textual entailment as a unified solver for such NLP problems.However, current research of textual entailment has not spilled much ink on the following questions: (i) How well does a pretrained textual entailment system generalize across domains with only a handful of domainspecific examples?and (ii) When is it worth transforming an NLP task into textual entailment?We argue that the transforming is unnecessary if we can obtain rich annotations for this task.Textual entailment really matters particularly when the target NLP task has insufficient annotations.Universal NLP 1 can be probably achieved through different routines.In this work, we introduce Universal Few-shot textual Entailment (UFO-ENTAIL).We demonstrate that this framework enables a pretrained entailment model to work well on new entailment domains in a few-shot setting, and show its effectiveness as a unified solver for several downstream NLP tasks such as question answering and coreference resolution when the end-task annotations are limited.
Wenpeng Yin 0001, Nazneen Fatema Rajani, Dragomir R. Radev, Richard Socher, Caiming Xiong
EMNLP (1)4
2020 Discriminative Nearest Neighbor Few-Shot Intent Detection by Transferring Natural Language Inference
abstract
Jianguo Zhang, Kazuma Hashimoto, Wenhao Liu, Chien-Sheng Wu, Yao Wan, Philip Yu, Richard Socher, Caiming Xiong. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jianguo Zhang 0005, Kazuma Hashimoto, Wenhao Liu 0003, Chien-Sheng Wu, Yao Wan 0001, Philip S. Yu, Richard Socher, Caiming Xiong
EMNLP (1)7
2020 Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question Answering
Akari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher, Caiming Xiong
ICLR4
2020 Non-Autoregressive Dialog State Tracking
Hung Le 0003, Richard Socher, Steven C. H. Hoi
ICLR2
2020 DivideMix: Learning with Noisy Labels as Semi-supervised Learning
Junnan Li 0001, Richard Socher, Steven C. H. Hoi
ICLR2
2020 Tree-Structured Attention with Hierarchical Accumulation
Xuan-Phi Nguyen, Shafiq R. Joty, Steven C. H. Hoi, Richard Socher
ICLR4
2020 Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
abstract
Acquiring abilities in the absence of a task-oriented reward function is at the frontier of reinforcement learning research. This problem has been studied through the lens of empowerment, which draws a connection between option discovery and information theory. Information-theoretic skill discovery methods have garnered much interest from the community, but little research has been conducted in understanding their limitations. Through theoretical analysis and empirical evidence, we show that existing algorithms suffer from a common limitation – they discover options that provide a poor coverage of the state space. In light of this, we propose Explore, Discover and Learn (EDL), an alternative approach to information-theoretic skill discovery. Crucially, EDL optimizes the same information-theoretic objective derived from the empowerment literature, but addresses the optimization problem using different machinery. We perform an extensive evaluation of skill discovery methods on controlled environments and show that EDL offers significant advantages, such as overcoming the coverage problem, reducing the dependence of learned skills on the initial state, and allowing the user to define a prior over which behaviors should be learned.
Victor Campos 0001, Alexander Trott, Caiming Xiong, Richard Socher, Xavier Giró-i-Nieto, Jordi Torres
ICML4
2020 An Investigation of Phone-Based Subword Units for End-to-End Speech Recognition
abstract
Phones and their context-dependent variants have been the standard modeling units for conventional speech recognition systems, while characters and subwords have demonstrated their effectiveness for end-to-end recognition systems.We investigate the use of phone-based subwords, in particular, byte pair encoder (BPE), as modeling units for end-to-end speech recognition.In addition, we also developed multi-level language model-based decoding algorithms based on a pronunciation dictionary.Besides the use of the lexicon, which is easily available, our system avoids the need of additional expert knowledge or processing steps from conventional systems.Experimental results show that phone-based BPEs tend to yield more accurate recognition systems than the character-based counterpart.In addition, further improvement can be obtained with a novel one-pass joint beam search decoder, which efficiently combines phone-and character-based BPE systems.For Switchboard, our phone-based BPE system achieves 6.8%/14.4% word error rate (WER) on the Switchboard/CallHome portion of the test set while joint decoding achieves 6.3%/13.3%WER.On Fisher + Switchboard, joint decoding leads to 4.9%/9.5% WER, setting new milestones for telephony speech recognition.
Guangsen Wang, Aadyot Bhatnagar, Yingbo Zhou 0002, Caiming Xiong, Richard Socher
INTERSPEECH6
2020 Towards Understanding Hierarchical Learning: Benefits of Neural Representations
abstract
Deep neural networks can empirically perform efficient hierarchical learning, in which the layers learn useful representations of the data. However, how they make use of the intermediate representations are not explained by recent theories that relate them to ``shallow learners'' such as kernels. In this work, we demonstrate that intermediate \emph{neural representations} add more flexibility to neural networks and can be advantageous over raw inputs. We consider a fixed, randomly initialized neural network as a representation function fed into another trainable network. When the trainable network is the quadratic Taylor model of a wide two-layer network, we show that neural representation can achieve improved sample complexities compared with the raw input: For learning a low-rank degree-$p$ polynomial ($p \geq 4$) in $d$ dimension, neural representation requires only $\widetilde{O}(d^{\ceil{p/2}})$ samples, while the best-known sample complexity upper bound for the raw input is $\widetilde{O}(d^{p-1})$. We contrast our result with a lower bound showing that neural representations do not improve over the raw input (in the infinite width limit), when the trainable network is instead a neural tangent kernel. Our results characterize when neural representations are beneficial, and may provide a new perspective on why depth is important in deep learning.
Minshuo Chen, Yu Bai 0017, Jason D. Lee, Tuo Zhao, Huan Wang 0016, Caiming Xiong, Richard Socher
NeurIPS7
2020 A Simple Language Model for Task-Oriented Dialogue
abstract
Task-oriented dialogue is often decomposed into three tasks: understanding user input, deciding actions, and generating a response. While such decomposition might suggest a dedicated model for each sub-task, we find a simple, unified approach leads to state-of-the-art performance on the MultiWOZ dataset. SimpleTOD is a simple approach to task-oriented dialogue that uses a single, causal language model trained on all sub-tasks recast as a single sequence prediction problem. This allows SimpleTOD to fully leverage transfer learning from pre-trained, open domain, causal language models such as GPT-2. SimpleTOD improves over the prior state-of-the-art in joint goal accuracy for dialogue state tracking, and our analysis reveals robustness to noisy annotations in this setting. SimpleTOD also improves the main metrics used to evaluate action decisions and response generation in an end-to-end setting: inform rate by 8.1 points, success rate by 9.7 points, and combined score by 7.2 points.
Ehsan Hosseini-Asl, Bryan McCann, Chien-Sheng Wu, Semih Yavuz, Richard Socher
NeurIPS5
2020 Online Structured Meta-learning
abstract
Learning quickly is of great importance for machine intelligence deployed in online platforms. With the capability of transferring knowledge from learned tasks, meta-learning has shown its effectiveness in online scenarios by continuously updating the model with the learned prior. However, current online meta-learning algorithms are limited to learn a globally-shared meta-learner, which may lead to sub-optimal results when the tasks contain heterogeneous information that are difficult to share. We overcome this limitation by proposing an online structured meta-learning (OSML) framework. Inspired by the knowledge organization of human and hierarchical feature representation, OSML explicitly disentangles the meta-learner as a meta-hierarchical graph with different knowledge blocks. When a new task is encountered, it constructs a meta-knowledge pathway by either utilizing the most relevant knowledge blocks or exploring new blocks. Through the meta-knowledge pathway, the model is able to quickly adapt to the new task. In addition, new knowledge is further incorporated into the selected blocks. Experiments on three datasets empirically demonstrate the effectiveness and interpretability of our proposed framework, not only under heterogeneous tasks but also under homogeneous settings.
Huaxiu Yao, Yingbo Zhou 0002, Mehrdad Mahdavi, Zhenhui Li, Richard Socher, Caiming Xiong
NeurIPS5
2020 Theory-Inspired Path-Regularized Differential Network Architecture Search
abstract
Despite its high search efficiency, differential architecture search (DARTS) often selects network architectures with dominated skip connections which lead to performance degradation. However, theoretical understandings on this issue remain absent, hindering the development of more advanced methods in a principled way. In this work, we solve this problem by theoretically analyzing the effects of various types of operations, e.g. convolution, skip connection and zero operation, to the network optimization. We prove that the architectures with more skip connections can converge faster than the other candidates, and thus are selected by DARTS. This result, for the first time, theoretically and explicitly reveals the impact of skip connections to fast network optimization and its competitive advantage over other types of operations in DARTS. Then we propose a theory-inspired path-regularized DARTS that consists of two key modules: (i) a differential group-structured sparse binary gate introduced for each operation to avoid unfair competition among operations, and (ii) a path-depth-wise regularization used to incite search exploration for deep architectures that often converge slower than shallow ones as shown in our theory and are not well explored during search. Experimental results on image classification tasks validate its advantages. Codes and models will be released.
Pan Zhou 0002, Caiming Xiong, Richard Socher, Steven C. H. Hoi
NeurIPS3
2019 Explain Yourself! Leveraging Language Models for Commonsense Reasoning
abstract
Deep learning models perform poorly on tasks that require commonsense reasoning, which often necessitates some form of worldknowledge or reasoning over information not immediately present in the input.We collect human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations in a new dataset called Common Sense Explanations (CoS-E).We use CoS-E to train language models to automatically generate explanations that can be used during training and inference in a novel Commonsense Auto-Generated Explanation (CAGE) framework.CAGE improves the state-of-the-art by 10% on the challenging CommonsenseQA task.We further study commonsense reasoning in DNNs using both human and auto-generated explanations including transfer to out-of-domain tasks.Empirical results indicate that we can effectively leverage language models for commonsense reasoning.
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, Richard Socher
ACL (1)4
2019 Transferable Multi-Domain State Generator for Task-Oriented Dialogue Systems
abstract
Over-dependence on domain ontology and lack of knowledge sharing across domains are two practical and yet less studied problems of dialogue state tracking.Existing approaches generally fall short in tracking unknown slot values during inference and often have difficulties in adapting to new domains.In this paper, we propose a TRAnsferable Dialogue statE generator (TRADE) that generates dialogue states from utterances using a copy mechanism, facilitating knowledge transfer when predicting (domain, slot, value) triplets not encountered during training.Our model is composed of an utterance encoder, a slot gate, and a state generator, which are shared across domains.Empirical results demonstrate that TRADE achieves state-of-the-art joint goal accuracy of 48.62% for the five domains of Mul-tiWOZ, a human-human dialogue dataset.In addition, we show its transferring ability by simulating zero-shot and few-shot dialogue state tracking for unseen domains.TRADE achieves 60.58% joint goal accuracy in one of the zero-shot domains, and is able to adapt to few-shot cases without forgetting already trained domains.* Work partially done while the first author was an intern at Salesforce Research.Usr: I am looking for a cheap restaurant in the centre of the
Chien-Sheng Wu, Andrea Madotto, Ehsan Hosseini-Asl, Caiming Xiong, Richard Socher, Pascale Fung
ACL (1)5
2019 SParC: Cross-Domain Semantic Parsing in Context
abstract
Tao Yu, Rui Zhang, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li, Heyang Er, Irene Li, Bo Pang, Tao Chen, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Vincent Zhang, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Tao Yu 0009, Rui Zhang 0037, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li 0002, Heyang Er, Irene Li, Bo Pang 0004, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Caiming Xiong, Richard Socher, Dragomir R. Radev
ACL (1)18
2019 AdaFrame: Adaptive Frame Selection for Fast Video Recognition
abstract
We present AdaFrame, a framework that adaptively selects relevant frames on a per-input basis for fast video recognition. AdaFrame contains a Long Short-Term Memory network augmented with a global memory that provides context information for searching which frames to use over time. Trained with policy gradient methods, AdaFrame generates a prediction, determines which frame to observe next, and computes the utility, i.e., expected future rewards, of seeing more frames at each time step. At testing time, AdaFrame exploits predicted utilities to achieve adaptive lookahead inference such that the overall computational costs are reduced without incurring a decrease in accuracy. Extensive experiments are conducted on two large-scale video benchmarks, FCVID and ActivityNet. AdaFrame matches the performance of using all frames with only 8.21 and 8.65 frames on FCVID and ActivityNet, respectively. We further qualitatively demonstrate learned frame usage can indicate the difficulty of making classification decisions; easier samples need fewer frames while harder ones require more, both at instance-level within the same class and at class-level among different categories.
Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, Larry Davis 0001
CVPR4
2019 WSLLN: Weakly Supervised Natural Language Localization Networks
abstract
Mingfei Gao, Larry Davis, Richard Socher, Caiming Xiong. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Mingfei Gao, Larry Davis 0001, Richard Socher, Caiming Xiong
EMNLP/IJCNLP (1)3
2019 Neural Text Summarization: A Critical Evaluation
abstract
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, Richard Socher. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, Richard Socher
EMNLP/IJCNLP (1)5
2019 CoSQL: A Conversational Text-to-SQL Challenge Towards Cross-Domain Natural Language Interfaces to Databases
abstract
Tao Yu, Rui Zhang, Heyang Er, Suyi Li, Eric Xue, Bo Pang, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Zihan Li, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Tao Chen, Alexander Fabbri, Zifan Li, Luyao Chen, Yuwen Zhang, Shreya Dixit, Vincent Zhang, Caiming Xiong, Richard Socher, Walter Lasecki, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Tao Yu 0009, Rui Zhang 0037, Heyang Er, Suyi Li 0002, Eric Xue 0001, Bo Pang 0004, Xi Victoria Lin, Yi Chern Tan, Tianze Shi, Youxuan Jiang, Michihiro Yasunaga, Sungrok Shim, Alexander R. Fabbri, Zifan Li, Shreya Dixit, Caiming Xiong, Richard Socher, Walter S. Lasecki, Dragomir R. Radev
EMNLP/IJCNLP (1)22
2019 Editing-Based SQL Query Generation for Cross-Domain Context-Dependent Questions
abstract
Rui Zhang, Tao Yu, Heyang Er, Sungrok Shim, Eric Xue, Xi Victoria Lin, Tianze Shi, Caiming Xiong, Richard Socher, Dragomir Radev. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Rui Zhang 0037, Tao Yu 0009, Heyang Er, Sungrok Shim, Eric Xue 0001, Xi Victoria Lin, Tianze Shi, Caiming Xiong, Richard Socher, Dragomir R. Radev
EMNLP/IJCNLP (1)9
2019 StartNet: Online Detection of Action Start in Untrimmed Videos
abstract
We propose StartNet to address Online Detection of Action Start (ODAS) where action starts and their associated categories are detected in untrimmed, streaming videos. Previous methods aim to localize action starts by learning feature representations that can directly separate the start point from its preceding background. It is challenging due to the subtle appearance difference near the action starts and the lack of training data. Instead, StartNet decomposes ODAS into two stages: action classification (using ClsNet) and start point localization (using LocNet). ClsNet focuses on per-frame labeling and predicts action score distributions online. Based on the predicted action scores of the past and current frames, LocNet conducts class-agnostic start detection by optimizing long-term localization rewards using policy gradient methods. The proposed framework is validated on two large-scale datasets, THUMOS'14 and ActivityNet. The experimental results show that StartNet significantly outperforms the state-of-the-art by 15%-30% p-mAP under the offset tolerance of 1-10 seconds on THUMOS'14, and achieves comparable performance on ActivityNet with 10 times smaller time offset.
Mingfei Gao, Larry Davis 0001, Richard Socher, Caiming Xiong
ICCV4
2019 A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, Richard Socher
ICLR (Poster)4
2019 Augmented Cyclic Adversarial Learning for Low Resource Domain Adaptation
Ehsan Hosseini-Asl, Yingbo Zhou 0002, Caiming Xiong, Richard Socher
ICLR (Poster)4
2019 Competitive experience replay
Alexander Trott, Richard Socher, Caiming Xiong
ICLR (Poster)3
2019 Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan Al-Regib, Zsolt Kira, Richard Socher, Caiming Xiong
ICLR (Poster)6
2019 Global-to-local Memory Pointer Networks for Task-Oriented Dialogue
Chien-Sheng Wu, Richard Socher, Caiming Xiong
ICLR (Poster)2
2019 Coarse-grain Fine-grain Coattention Network for Multi-evidence Question Answering
Victor Zhong, Caiming Xiong, Nitish Shirish Keskar, Richard Socher
ICLR (Poster)4
2019 Learn to Grow: A Continual Structure Learning Framework for Overcoming Catastrophic Forgetting
abstract
Addressing catastrophic forgetting is one of the key challenges in continual learning where machine learning systems are trained with sequential or streaming tasks. Despite recent remarkable progress in state-of-the-art deep learning, deep neural networks (DNNs) are still plagued with the catastrophic forgetting problem. This paper presents a conceptually simple yet general and effective framework for handling catastrophic forgetting in continual learning with DNNs. The proposed method consists of two components: a neural structure optimization component and a parameter learning and/or fine-tuning component. By separating the explicit neural structure learning and the parameter estimation, not only is the proposed method capable of evolving neural structures in an intuitively meaningful way, but also shows strong capabilities of alleviating catastrophic forgetting in experiments. Furthermore, the proposed method outperforms all other baselines on the permuted MNIST dataset, the split CIFAR100 dataset and the Visual Domain Decathlon dataset in continual learning setting.
Xilai Li, Yingbo Zhou 0002, Tianfu Wu 0001, Richard Socher, Caiming Xiong
ICML4
2019 Taming MAML: Efficient unbiased meta-reinforcement learning
abstract
While meta reinforcement learning (Meta-RL) methods have achieved remarkable success, obtaining correct and low variance estimates for policy gradients remains a significant challenge. In particular, estimating a large Hessian, poor sample efficiency and unstable training continue to make Meta-RL difficult. We propose a surrogate objective function named, Taming MAML (TMAML), that adds control variates into gradient estimation via automatic differentiation. TMAML improves the quality of gradient estimation by reducing variance without introducing bias. We further propose a version of our method that extends the meta-learning framework to learning the control variates themselves, enabling efficient and scalable learning from a distribution of MDPs. We empirically compare our approach with MAML and other variance-bias trade-off methods including DICE, LVC, and action-dependent control variates. Our approach is easy to implement and outperforms existing methods in terms of the variance and accuracy of gradient estimation, ultimately yielding higher performance across a variety of challenging Meta-RL environments.
Richard Socher, Caiming Xiong
ICML2
2019 On the Generalization Gap in Reparameterizable Reinforcement Learning
abstract
Understanding generalization in reinforcement learning (RL) is a significant challenge, as many common assumptions of traditional supervised learning theory do not apply. We focus on the special class of reparameterizable RL problems, where the trajectory distribution can be decomposed using the reparametrization trick. For this problem class, estimating the expected return is efficient and the trajectory can be computed deterministically given peripheral random variables, which enables us to study reparametrizable RL using supervised learning and transfer learning theory. Through these relationships, we derive guarantees on the gap between the expected and empirical return for both intrinsic and external errors, based on Rademacher complexity as well as the PAC-Bayes bound. Our bound suggests the generalization capability of reparameterizable RL is related to multiple factors including “smoothness” of the environment transition, reward and agent policy function class. We also empirically verify the relationship between the generalization gap and these factors through simulations.
Huan Wang 0016, Stephan Zheng, Caiming Xiong, Richard Socher
ICML4
2019 Keeping Your Distance: Solving Sparse Reward Tasks Using Self-Balancing Shaped Rewards
abstract
While using shaped rewards can be beneficial when solving sparse reward tasks, their successful application often requires careful engineering and is problem specific. For instance, in tasks where the agent must achieve some goal state, simple distance-to-goal reward shaping often fails, as it renders learning vulnerable to local optima. We introduce a simple and effective model-free method to learn from shaped distance-to-goal rewards on tasks where success depends on reaching a goal state. Our method introduces an auxiliary distance-based reward based on pairs of rollouts to encourage diverse exploration. This approach effectively prevents learning dynamics from stabilizing around local optima induced by the naive distance-to-goal reward shaping and enables policies to efficiently solve sparse reward tasks. Our augmented objective does not require any additional reward engineering or domain expertise to implement and converges to the original sparse objective as the agent learns to solve the task. We demonstrate that our method successfully solves a variety of hard-exploration tasks (including maze navigation and 3D construction in a Minecraft environment), where naive distance-based reward shaping otherwise fails, and intrinsic curiosity and reward relabeling strategies exhibit poor performance.
Alexander Trott, Stephan Zheng, Caiming Xiong, Richard Socher
NeurIPS4
2019 Genie: a generator of natural language semantic parsers for virtual assistant commands
abstract
To understand diverse natural language commands, virtual assistants today are trained with numerous labor-intensive, manually annotated sentences. This paper presents a methodology and the Genie toolkit that can handle new compound commands with significantly less manual effort. We advocate formalizing the capability of virtual assistants with a Virtual Assistant Programming Language (VAPL) and using a neural semantic parser to translate natural language into VAPL code. Genie needs only a small realistic set of input sentences for validating the neural model. Developers write templates to synthesize data; Genie uses crowdsourced paraphrases and data augmentation, along with the synthesized data, to train a semantic parser. We also propose design principles that make VAPL languages amenable to natural language translation. We apply these principles to revise ThingTalk, the language used by the Almond virtual assistant. We use Genie to build the first semantic parser that can support compound virtual assistants commands with unquoted free-form parameters. Genie achieves a 62% accuracy on realistic user inputs. We demonstrate Genie’s generality by showing a 19% and 31% improvement over the previous state of the art on a music skill, aggregate functions, and access control.
Giovanni Campagna, Silei Xu, Mehrad Moradshahi, Richard Socher, Monica S. Lam
PLDI4
2018 Global-Locally Self-Attentive Encoder for Dialogue State Tracking
abstract
Dialogue state tracking, which estimates user goals and requests given the dialogue context, is an essential part of taskoriented dialogue systems.In this paper, we propose the Global-Locally Self-Attentive Dialogue State Tracker (GLAD), which learns representations of the user utterance and previous system actions with global-local modules.Our model uses global modules to share parameters between estimators for different types (called slots) of dialogue states, and uses local modules to learn slot-specific features.We show that this significantly improves tracking of rare states and achieves stateof-the-art performance on the WoZ and DSTC2 state tracking tasks.GLAD obtains 88.1% joint goal accuracy and 97.1% request accuracy on WoZ, outperforming prior work by 3.7% and 5.5%.On DSTC2, our model obtains 74.5% joint goal accuracy and 97.5% request accuracy, outperforming prior work by 1.1% and 1.0%.
Victor Zhong, Caiming Xiong, Richard Socher
ACL (1)3
2018 Efficient and Robust Question Answering from Minimal Context over Documents
abstract
Neural models for question answering (QA) over documents have achieved significant performance improvements.Although effective, these models do not scale to large corpora due to their complex modeling of interactions between the document and the question.Moreover, recent work has shown that such models are sensitive to adversarial inputs.In this paper, we study the minimal context required to answer the question, and find that most questions in existing datasets can be answered with a small set of sentences.Inspired by this observation, we propose a simple sentence selector to select the minimal set of sentences to feed into the QA model.Our overall system achieves significant reductions in training (up to 15 times) and inference times (up to 13 times), with accuracy comparable to or better than the state-of-the-art on SQuAD, NewsQA, TriviaQA and SQuAD-Open.Furthermore, our experimental results and analyses show that our approach is more robust to adversarial inputs.
Sewon Min, Victor Zhong, Richard Socher, Caiming Xiong
ACL (1)3
2018 End-to-End Dense Video Captioning With Masked Transformer
abstract
Dense video captioning aims to generate text descriptions for all events in an untrimmed video. This involves both detecting and describing events. Therefore, all previous methods on dense video captioning tackle this problem by building two models, i.e. an event proposal and a captioning model, for these two sub-problems. The models are either trained separately or in alternation. This prevents direct influence of the language description to the event proposal, which is important for generating accurate descriptions. To address this problem, we propose an end-to-end transformer model for dense video captioning. The encoder encodes the video into appropriate representations. The proposal decoder decodes from the encoding with different anchors to form video event proposals. The captioning decoder employs a masking network to restrict its attention to the proposal event over the encoding feature. This masking network converts the event proposal to a differentiable mask, which ensures the consistency between the proposal and captioning during training. In addition, our model employs a self-attention mechanism, which enables the use of efficient non-recurrent structure during encoding and leads to performance improvements. We demonstrate the effectiveness of this end-to-end model on ActivityNet Captions and YouCookII datasets, where we achieved 10.12 and 6.58 METEOR score, respectively.
Luowei Zhou, Yingbo Zhou 0002, Jason J. Corso, Richard Socher, Caiming Xiong
CVPR4
2018 Improving Abstraction in Text Summarization
abstract
Abstractive text summarization aims to shorten long text documents into a human readable form that contains the most important facts from the original document.However, the level of actual abstraction as measured by novel phrases that do not appear in the source document remains low in existing approaches.We propose two techniques to improve the level of abstraction of generated summaries.First, we decompose the decoder into a contextual network that retrieves relevant parts of the source document, and a pretrained language model that incorporates prior knowledge about language generation.Second, we propose a novelty metric that is optimized directly through policy learning to encourage the generation of novel phrases.Our model achieves results comparable to state-of-the-art models, as determined by ROUGE scores and human evaluations, while achieving a significantly higher level of abstraction as measured by n-gram overlap with the source document.
Wojciech Kryscinski, Romain Paulus, Caiming Xiong, Richard Socher
EMNLP4
2018 Multi-Hop Knowledge Graph Reasoning with Reward Shaping
abstract
Multi-hop reasoning is an effective approach for query answering (QA) over incomplete knowledge graphs (KGs).The problem can be formulated in a reinforcement learning (RL) setup, where a policy-based agent sequentially extends its inference path until it reaches a target.However, in an incomplete KG environment, the agent receives low-quality rewards corrupted by false negatives in the training data, which harms generalization at test time.Furthermore, since no golden action sequence is used for training, the agent can be misled by spurious search trajectories that incidentally lead to the correct answer.We propose two modeling advances to address both issues: (1) we reduce the impact of false negative supervision by adopting a pretrained onehop embedding model to estimate the reward of unobserved facts; (2) we counter the sensitivity to spurious paths of on-policy RL by forcing the agent to explore a diverse set of paths using randomly generated edge masks.Our approach significantly improves over existing path-based KGQA models on several benchmark datasets and is comparable or better than embedding-based models.
Xi Victoria Lin, Richard Socher, Caiming Xiong
EMNLP2
2018 Improving End-to-End Speech Recognition with Policy Learning
abstract
Connectionist temporal classification (CTC) is widely used for maximum likelihood learning in end-to-end speech recognition models. However, there is usually a disparity between the negative maximum likelihood and the performance metric used in speech recognition, e.g., word error rate (WER). This results in a mismatch between the objective function and metric during training. We show that the above problem can be mitigated by jointly training with maximum likelihood and policy gradient. In particular, with policy learning we are able to directly optimize on the (otherwise non-differentiable) performance metric. We show that joint training improves relative performance by 4% to 13% for our end-to-end model as compared to the same model learned through maximum likelihood. The model achieves 5.53% WER on Wall Street Journal dataset, and 5.42% and 14.70% on Librispeech test-clean and test-other set, respectively.
Yingbo Zhou 0002, Caiming Xiong, Richard Socher
ICASSP3
2018 Non-Autoregressive Neural Machine Translation
Jiatao Gu, James Bradbury 0002, Caiming Xiong, Victor O. K. Li, Richard Socher
ICLR (Poster)5
2018 Regularizing and Optimizing LSTM Language Models
Stephen Merity, Nitish Shirish Keskar, Richard Socher
ICLR (Poster)3
2018 A Deep Reinforced Model for Abstractive Summarization
Romain Paulus, Caiming Xiong, Richard Socher
ICLR (Poster)3
2018 Hierarchical and Interpretable Skill Acquisition in Multi-task Reinforcement Learning
Tianmin Shu, Caiming Xiong, Richard Socher
ICLR (Poster)3
2018 Interpretable Counting for Visual Question Answering
Alexander Trott, Caiming Xiong, Richard Socher
ICLR (Poster)3
2018 DCN+: Mixed Objective And Deep Residual Coattention for Question Answering
Caiming Xiong, Victor Zhong, Richard Socher
ICLR (Poster)3
2018 A Multi-Discriminator CycleGAN for Unsupervised Non-Parallel Speech Domain Adaptation
abstract
Domain adaptation plays an important role for speech recognition models, in particular, for domains that have low resources. We propose a novel generative model based on cyclic-consistent generative adversarial network (CycleGAN) for unsupervised non-parallel speech domain adaptation. The proposed model employs multiple independent discriminators on the power spectrogram, each in charge of different frequency bands. As a result we have 1) better discriminators that focus on fine-grained details of the frequency features, and 2) a generator that is capable of generating more realistic domain-adapted spectrogram. We demonstrate the effectiveness of our method on speech recognition with gender adaptation, where the model only has access to supervised data from one gender during training, but is evaluated on the other at test time. Our model is able to achieve an average of $7.41\%$ on phoneme error rate, and $11.10\%$ word error rate relative performance improvement as compared to the baseline, on TIMIT and WSJ dataset, respectively. Qualitatively, our model also generates more natural sounding speech, when conditioned on data from the other domain.
Ehsan Hosseini-Asl, Yingbo Zhou 0002, Caiming Xiong, Richard Socher
INTERSPEECH4
2017 Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning
abstract
Attention-based neural encoder-decoder frameworks have been widely adopted for image captioning. Most methods force visual attention to be active for every generated word. However, the decoder likely requires little to no visual information from the image to predict non-visual words such as the and of. Other words that may seem visual can often be predicted reliably just from the language model e.g., sign after behind a red stop or phone following talking on a cell. In this paper, we propose a novel adaptive attention model with a visual sentinel. At each time step, our model decides whether to attend to the image (and if so, to which regions) or to the visual sentinel. The model decides whether to attend to the image and where, in order to extract meaningful information for sequential word generation. We test our method on the COCO image captioning 2015 challenge dataset and Flickr30K. Our approach sets the new state-of-the-art by a significant margin.
Jiasen Lu, Caiming Xiong, Devi Parikh, Richard Socher
CVPR4
2017 A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks
abstract
Transfer and multi-task learning have traditionally focused on either a single source-target pair or very few, similar tasks.Ideally, the linguistic levels of morphology, syntax and semantics would benefit each other by being trained in a single model.We introduce a joint many-task model together with a strategy for successively growing its depth to solve increasingly complex tasks.Higher layers include shortcut connections to lower-level task predictions to reflect linguistic hierarchies.We use a simple regularization term to allow for optimizing all model weights to improve one task's loss without exhibiting catastrophic interference of the other tasks.Our single end-to-end model obtains state-of-the-art or competitive results on five different tasks from tagging, parsing, relatedness, and entailment tasks.
Kazuma Hashimoto, Caiming Xiong, Yoshimasa Tsuruoka, Richard Socher
EMNLP4
2017 Quasi-Recurrent Neural Networks
James Bradbury 0002, Stephen Merity, Caiming Xiong, Richard Socher
ICLR (Poster)4
2017 Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
Hakan Inan, Khashayar Khosravi, Richard Socher
ICLR (Poster)3
2017 Pointer Sentinel Mixture Models
Stephen Merity, Caiming Xiong, James Bradbury 0002, Richard Socher
ICLR (Poster)4
2017 Dynamic Coattention Networks For Question Answering
Caiming Xiong, Victor Zhong, Richard Socher
ICLR (Poster)3
2017 Learned in Translation: Contextualized Word Vectors
abstract
Computer vision has benefited from initializing multiple deep layers with weights pretrained on large supervised training sets like ImageNet. Natural language processing (NLP) typically sees initialization of only the lowest layer of deep models with pretrained word vectors. In this paper, we use a deep LSTM encoder from an attentional sequence-to-sequence model trained for machine translation (MT) to contextualize word vectors. We show that adding these context vectors (CoVe) improves performance over using only unsupervised word and character vectors on a wide variety of common NLP tasks: sentiment analysis (SST, IMDb), question classification (TREC), entailment (SNLI), and question answering (SQuAD). For fine-grained sentiment analysis and entailment, CoVe improves performance of our baseline models to the state of the art.
Bryan McCann, James Bradbury 0002, Caiming Xiong, Richard Socher
NIPS4
2016 Ask Me Anything: Dynamic Memory Networks for Natural Language Processing
abstract
Most tasks in natural language processing can be cast into question answering (QA) problems over language input. We introduce the dynamic memory network (DMN), a neural network architecture which processes input sequences and questions, forms episodic memories, and generates relevant answers. Questions trigger an iterative attention process which allows the model to condition its attention on the inputs and the result of previous iterations. These results are then reasoned over in a hierarchical recurrent sequence model to generate answers. The DMN can be trained end-to-end and obtains state-of-the-art results on several types of tasks and datasets: question answering (Facebook’s bAbI dataset), text classification for sentiment analysis (Stanford Sentiment Treebank) and sequence modeling for part-of-speech tagging (WSJ-PTB). The training for these different tasks relies exclusively on trained word vector representations and input-question-answer triplets.
Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury 0002, Ishaan Gulrajani, Victor Zhong, Romain Paulus, Richard Socher
ICML9
2016 Dynamic Memory Networks for Visual and Textual Question Answering
abstract
Neural network architectures with memory and attention mechanisms exhibit certain reason- ing capabilities required for question answering. One such architecture, the dynamic memory net- work (DMN), obtained high accuracy on a variety of language tasks. However, it was not shown whether the architecture achieves strong results for question answering when supporting facts are not marked during training or whether it could be applied to other modalities such as images. Based on an analysis of the DMN, we propose several improvements to its memory and input modules. Together with these changes we introduce a novel input module for images in order to be able to answer visual questions. Our new DMN+ model improves the state of the art on both the Visual Question Answering dataset and the bAbI-10k text question-answering dataset without supporting fact supervision.
Caiming Xiong, Stephen Merity, Richard Socher
ICML3
2015 Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks
abstract
Kai Sheng Tai, Richard Socher, Christopher D. Manning. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Kai Sheng Tai, Richard Socher, Christopher D. Manning
ACL (1)2
2014 A Neural Network for Factoid Question Answering over Paragraphs
abstract
Text classification methods for tasks like factoid question answering typi-cally use manually defined string match-ing rules or bag of words representa-tions. These methods are ineffective when question text contains very few individual words (e.g., named entities) that are indicative of the answer. We introduce a recursive neural network (rnn) model that can reason over such input by modeling textual composition-ality. We apply our model, qanta, to a dataset of questions from a trivia competition called quiz bowl. Unlike previous rnn models, qanta learns word and phrase-level representations that combine across sentences to reason about entities. The model outperforms multiple baselines and, when combined with information retrieval methods, ri-vals the best human players. 1
Mohit Iyyer, Jordan L. Boyd-Graber, Leonardo Max Batista Claudino, Richard Socher, Hal Daumé III
EMNLP4
2014 Glove: Global Vectors for Word Representation
abstract
Recent methods for learning vector space representations of words have succeeded in capturing fine-grained semantic and syntactic regularities using vector arith-metic, but the origin of these regularities has remained opaque. We analyze and make explicit the model properties needed for such regularities to emerge in word vectors. The result is a new global log-bilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods. Our model efficiently leverages statistical information by training only on the nonzero elements in a word-word co-occurrence matrix, rather than on the en-tire sparse matrix or on individual context windows in a large corpus. The model pro-duces a vector space with meaningful sub-structure, as evidenced by its performance of 75 % on a recent word analogy task. It also outperforms related models on simi-larity tasks and named entity recognition. 1
Jeffrey Pennington, Richard Socher, Christopher D. Manning
EMNLP2
2014 Scaling short-answer grading by combining peer assessment with algorithmic scoring
abstract
Peer assessment helps students reflect and exposes them to different ideas. It scales assessment and allows large online classes to use open-ended assignments. However, it requires students to spend significant time grading. How can we lower this grading burden while maintaining quality? This paper integrates peer and machine grading to preserve the robustness of peer assessment and lower grading burden. In the identify-verify pattern, a grading algorithm first predicts a student grade and estimates confidence, which is used to estimate the number of peer raters required. Peers then identify key features of the answer using a rubric. Finally, other peers verify whether these feature labels were accurately applied. This pattern adjusts the number of peers that evaluate an answer based on algorithmic confidence and peer agreement. We evaluated this pattern with 1370 students in a large, online design class. With only 54% of the student grading time, the identify-verify pattern yields 80-90% of the accuracy obtained by taking the median of three peer scores, and provides more detailed feedback. A second experiment found that verification dramatically improves accuracy with more raters, with a 20% gain over the peer-median with four raters. However, verification also leads to lower initial trust in the grading system. The identify-verify pattern provides an example of how peer work and machine learning can combine to improve the learning experience.
Chinmay Kulkarni 0001, Richard Socher, Michael S. Bernstein, Scott R. Klemmer
L@S2
2014 Global Belief Recursive Neural Networks
Romain Paulus, Richard Socher, Christopher D. Manning
NIPS2
2014 Grounded Compositional Semantics for Finding and Describing Images with Sentences
abstract
Previous work on Recursive Neural Networks (RNNs) shows that these models can produce compositional feature vectors for accurately representing and classifying sentences or images. However, the sentence vectors of previous models cannot accurately represent visually grounded meaning. We introduce the DT-RNN model which uses dependency trees to embed sentences into a vector space in order to retrieve images that are described by those sentences. Unlike previous RNN-based models which use constituency trees, DT-RNNs naturally focus on the action and agents in a sentence. They are better able to abstract from the details of word order and syntactic expression. DT-RNNs outperform other recursive and recurrent neural networks, kernelized CCA and a bag-of-words baseline on the tasks of finding an image that fits a sentence description and vice versa. They also give more similar representations to sentences that describe the same image.
Richard Socher, Andrej Karpathy, Quoc V. Le, Christopher D. Manning, Andrew Y. Ng
Trans. Assoc. Comput. Linguistics1
2013 Parsing with Compositional Vector Grammars
Richard Socher, John Bauer, Christopher D. Manning, Andrew Y. Ng
ACL (1)1
2013 Better Word Representations with Recursive Neural Networks for Morphology
Thang Luong, Richard Socher, Christopher D. Manning
CoNLL2
2013 Recursive Deep Models for Semantic Compositionality Over a Sentiment Treebank
abstract
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, Christopher Potts. Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. 2013.
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, Christopher Potts
EMNLP1
2013 Bilingual Word Embeddings for Phrase-Based Machine Translation
abstract
We introduce bilingual word embeddings: semantic embeddings associated across two languages in the context of neural language models.We propose a method to learn bilingual embeddings from a large unlabeled corpus, while utilizing MT word alignments to constrain translational equivalence.The new embeddings significantly out-perform baselines in word semantic similarity.A single semantic similarity feature induced with bilingual embeddings adds near half a BLEU point to the results of NIST08 Chinese-English machine translation task.
Will Y. Zou, Richard Socher, Daniel M. Cer, Christopher D. Manning
EMNLP2
2013 Deep Learning for NLP (without Magic)
Richard Socher, Christopher D. Manning
HLT-NAACL1
2013 Reasoning With Neural Tensor Networks for Knowledge Base Completion
abstract
A common problem in knowledge representation and related fields is reasoning over a large joint knowledge graph, represented as triples of a relation between two entities. The goal of this paper is to develop a more powerful neural network model suitable for inference over these relationships. Previous models suffer from weak interaction between entities or simple linear projection of the vector space. We address these problems by introducing a neural tensor network (NTN) model which allow the entities and relations to interact multiplicatively. Additionally, we observe that such knowledge base models can be further improved by representing each entity as the average of vectors for the words in the entity name, giving an additional dimension of similarity by which entities can share statistical strength. We assess the model by considering the problem of predicting additional true relations between entities given a partial knowledge base. Our model outperforms previous models and can classify unseen relationships in WordNet and FreeBase with an accuracy of 86.2% and 90.0%, respectively.
Richard Socher, Danqi Chen 0001, Christopher D. Manning, Andrew Y. Ng
NIPS1
2013 Zero-Shot Learning Through Cross-Modal Transfer
abstract
This work introduces a model that can recognize objects in images even if no training data is available for the object class. The only necessary knowledge about unseen categories comes from unsupervised text corpora. Unlike previous zero-shot learning models, which can only differentiate between unseen classes, our model can operate on a mixture of objects, simultaneously obtaining state of the art performance on classes with thousands of training images and reasonable performance on unseen classes. This is achieved by seeing the distributions of words in texts as a semantic space for understanding what objects look like. Our deep learning model does not require any manually defined semantic or visual features for either words or images. Images are mapped to be close to semantic word vectors corresponding to their classes, and the resulting image embeddings can be used to distinguish whether an image is of a seen or unseen class. Then, a separate recognition model can be employed for each type. We demonstrate two strategies, the first gives high accuracy on unseen classes, while the second is conservative in its prediction of novelty and keeps the seen classes' accuracy high.
Richard Socher, Milind Ganjoo, Christopher D. Manning, Andrew Y. Ng
NIPS1
2012 Improving Word Representations via Global Context and Multiple Word Prototypes
Eric H. Huang, Richard Socher, Christopher D. Manning, Andrew Y. Ng
ACL (1)2
2012 Semantic Compositionality through Recursive Matrix-Vector Spaces
Richard Socher, Brody Huval, Christopher D. Manning, Andrew Y. Ng
EMNLP-CoNLL1
2012 Convolutional-Recursive Deep Learning for 3D Object Classification
abstract
Recent advances in 3D sensing technologies make it possible to easily record color and depth images which together can improve object recognition. Most current methods rely on very well-designed features for this new 3D modality. We in- troduce a model based on a combination of convolutional and recursive neural networks (CNN and RNN) for learning features and classifying RGB-D images. The CNN layer learns low-level translationally invariant features which are then given as inputs to multiple, fixed-tree RNNs in order to compose higher order fea- tures. RNNs can be seen as combining convolution and pooling into one efficient, hierarchical operation. Our main result is that even RNNs with random weights compose powerful features. Our model obtains state of the art performance on a standard RGB-D object dataset while being more accurate and faster during train- ing and testing than comparable architectures such as two-layer CNNs.
Richard Socher, Brody Huval, Bharath Putta Bath, Christopher D. Manning, Andrew Y. Ng
NIPS1
2011 Semi-Supervised Recursive Autoencoders for Predicting Sentiment Distributions
Richard Socher, Jeffrey Pennington, Eric H. Huang, Andrew Y. Ng, Christopher D. Manning
EMNLP1
2011 Parsing Natural Scenes and Natural Language with Recursive Neural Networks
Richard Socher, Cliff Chiung-Yu Lin, Andrew Y. Ng, Christopher D. Manning
ICML1
2011 Dynamic Pooling and Unfolding Recursive Autoencoders for Paraphrase Detection
abstract
Paraphrase detection is the task of examining two sentences and determining whether they have the same meaning. In order to obtain high accuracy on this task, thorough syntactic and semantic analysis of the two statements is needed. We introduce a method for paraphrase detection based on recursive autoencoders (RAE). Our unsupervised RAEs are based on a novel unfolding objective and learn feature vectors for phrases in syntactic trees. These features are used to measure the word- and phrase-wise similarity between two sentences. Since sentences may be of arbitrary length, the resulting matrix of similarity measures is of variable size. We introduce a novel dynamic pooling layer which computes a fixed-sized representation from the variable-sized matrices. The pooled representation is then used as input to a classifier. Our method outperforms other state-of-the-art approaches on the challenging MSRP paraphrase corpus.
Richard Socher, Eric H. Huang, Jeffrey Pennington, Andrew Y. Ng, Christopher D. Manning
NIPS1
2010 Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora
abstract
We propose a semi-supervised model which segments and annotates images using very few labeled images and a large unaligned text corpus to relate image regions to text labels. Given photos of a sports event, all that is necessary to provide a pixel-level labeling of objects and background is a set of newspaper articles about this sport and one to five labeled images. Our model is motivated by the observation that words in text corpora share certain context and feature similarities with visual objects. We describe images using visual words, a new region-based representation. The proposed model is based on kernelized canonical correlation analysis which finds a mapping between visual and textual words by projecting them into a latent meaning space. Kernels are derived from context and adjective features inside the respective visual and textual domains. We apply our method to a challenging dataset and rely on articles of the New York Times for textual features. Our model outperforms the state-of-the-art in annotation. In segmentation it compares favorably with other methods that use significantly more labeled training data.
Richard Socher, Li Fei-Fei 0001
CVPR1
2009 ImageNet: A large-scale hierarchical image database
abstract
The explosion of image data on the Internet has the potential to foster more sophisticated and robust models and algorithms to index, retrieve, organize and interact with images and multimedia data. But exactly how such data can be harnessed and organized remains a critical problem. We introduce here a new database called “ImageNet”, a large-scale ontology of images built upon the backbone of the WordNet structure. ImageNet aims to populate the majority of the 80,000 synsets of WordNet with an average of 500–1000 clean and full resolution images. This will result in tens of millions of annotated images organized by the semantic hierarchy of WordNet. This paper offers a detailed analysis of ImageNet in its current state: 12 subtrees with 5247 synsets and 3.2 million images in total. We show that ImageNet is much larger in scale and diversity and much more accurate than the current image datasets. Constructing such a large-scale database is a challenging task. We describe the data collection scheme with Amazon Mechanical Turk. Lastly, we illustrate the usefulness of ImageNet through three simple applications in object recognition, image classification and automatic object clustering. We hope that the scale, accuracy, diversity and hierarchical structure of ImageNet can offer unparalleled opportunities to researchers in the computer vision community and beyond.
Jia Deng 0001, Wei Dong 0003, Richard Socher, Li-Jia Li 0001, Kai Li 0001, Li Fei-Fei 0001
CVPR3
2009 Towards total scene understanding: Classification, annotation and segmentation in an automatic framework
abstract
Given an image, we propose a hierarchical generative model that classifies the overall scene, recognizes and segments each object component, as well as annotates the image with a list of tags. To our knowledge, this is the first model that performs all three tasks in one coherent framework. For instance, a scene of a `polo game' consists of several visual objects such as `human', `horse', `grass', etc. In addition, it can be further annotated with a list of more abstract (e.g. `dusk') or visually less salient (e.g. `saddle') tags. Our generative model jointly explains images through a visual model and a textual model. Visually relevant objects are represented by regions and patches, while visually irrelevant textual annotations are influenced directly by the overall scene class. We propose a fully automatic learning framework that is able to learn robust scene models from noisy Web data such as images and user tags from Flickr.com. We demonstrate the effectiveness of our framework by automatically classifying, annotating and segmenting images from eight classes depicting sport scenes. In all three tasks, our model significantly outperforms state-of-the-art algorithms.
Li-Jia Li 0001, Richard Socher, Li Fei-Fei 0001
CVPR2
2009 A Bayesian Analysis of Dynamics in Free Recall
abstract
We develop a probabilistic model of human memory performance in free recall experiments. In these experiments, a subject first studies a list of words and then tries to recall them. To model these data, we draw on both previous psychological research and statistical topic models of text documents. We assume that memories are formed by assimilating the semantic meaning of studied words (represented as a distribution over topics) into a slowly changing latent context (represented in the same space). During recall, this context is reinstated and used as a cue for retrieving studied words. By conceptualizing memory retrieval as a dynamic latent variable model, we are able to use Bayesian inference to represent uncertainty and reason about the cognitive processes underlying memory. We present a particle filter algorithm for performing approximate posterior inference, and evaluate our model on the prediction of recalled words in experimental data. By specifying the model hierarchically, we are also able to capture inter-subject variability.
Richard Socher, Samuel Gershman, Adler J. Perotte, Per B. Sederberg, David M. Blei, Kenneth A. Norman
NIPS1