EDBT 2026 Demo / reviewers in the wild / expert
Jacob Andreas
dblp:97/8154
· DBLP profile ↗
97ranked-venue papers
15as first author
65since 2021 · last 2026
0000-0002-3141-5845ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 95 · 15 first-author · 64 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Modeling Student Learning with 3.8 Million Program Traces
Alexis Ross, Megha Srivastava, Jeremiah J. Blanchard, Jacob Andreas |
AIED | 4 |
| 2025 | Language and Experience: A Computational Model of Social Learning in Complex Novel Tasks
Cédric Colas, Tracey Mills, Ben Prystawski, Michael Henry Tessler, Noah D. Goodman, Jacob Andreas, Josh Tenenbaum |
CogSci | 6 |
| 2025 | Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
Lionel Wong, Katie Collins, Lance Ying, Cedegao E. Zhang, Adrian Weller, Tobias Gerstenberg, Timothy J. O'Donnell, Alexander K. Lew, Jacob Andreas, Tyler Brooke-Wilson, Josh Tenenbaum |
CogSci | 9 |
| 2025 | Learning How Hard to Think: Input-Adaptive Allocation of LM ComputationabstractComputationally intensive decoding procedures---including search, reranking, and self-critique---can improve the quality of language model (LM) outputs in problems spanning code generation, numerical reasoning, and dialog.
Existing work typically applies the same decoding procedure for every input to an LM. But not all inputs require the same amount of computation to process. Can we allocate decoding computation adaptively, using more resources to answer questions whose answers will be harder to compute? We present an approach that predicts the distribution of rewards given an input and computation budget, then allocates additional computation to inputs for which it is predicted to be most useful. We apply this approach in two decoding procedures: first, an adaptive best-of-$k$ procedure that dynamically selects the number of samples to generate as input to a reranker; second, a routing procedure that dynamically responds to a query using a decoding procedure that is expensive but accurate, or one that is cheaper but less capable. Across a suite of programming, mathematics, and dialog tasks, we show that accurate computation-allocation procedures can be learned, and reduce computation by up to 50% at no cost to quality. Mehul Damani, Idan Shenfeld, Andi Peng, Andreea Bobu, Jacob Andreas |
ICLR | 5 |
| 2025 | Eliciting Human Preferences with Language ModelsabstractLanguage models (LMs) can be directed to perform user- and context-dependent
tasks by using labeled examples or natural language prompts.
But selecting examples or writing prompts can be challenging---especially in tasks that require users to precisely articulate nebulous preferences or reason about complex edge cases. For such tasks, we introduce **Generative Active Task Elicitation (GATE)**, a method for using *LMs themselves* to guide the task specification process. GATE is a learning framework in which models elicit and infer human preferences through free-form, language-based interaction with users.
We identify prototypical challenges that users face when specifying preferences, and design three preference modeling tasks to study these challenges:
content recommendation, moral reasoning, and email validation.
In preregistered experiments, we show that LMs that learn to perform these tasks using GATE (by interactively querying users with open-ended questions) obtain preference specifications that are more informative than user-written prompts or examples. GATE matches existing task specification methods in the moral reasoning task, and significantly outperforms them in the content recommendation and email validation tasks. Users additionally report that interactive task elicitation requires less effort than prompting or example labeling and surfaces considerations that they did not anticipate on their own. Our findings suggest that LM-driven elicitation can be a powerful tool for aligning models to complex human preferences and values. Belinda Z. Li, Alex Tamkin, Noah D. Goodman, Jacob Andreas |
ICLR | 4 |
| 2025 | The Surprising Effectiveness of Test-Time Training for Few-Shot LearningabstractLanguage models (LMs) have shown impressive performance on tasks within their training distribution, but often struggle with structurally novel tasks even when given a small number of in-context task examples. We investigate the effectiveness of test-time training (TTT)—temporarily updating model parameters during inference using a loss derived from input data—as a mechanism for improving LMs’ reasoning and few-shot learning capabilities. On the Abstraction and Reasoning Corpus (ARC), performing TTT with in-context examples yields up to $6\times$ higher accuracy compared to fine-tuned baselines—reaching $53.0%$ on the public validation set with an 8B-parameter LM and $61.9%$ when ensembled with program-synthesis methods, matching average human performance. On BIG-Bench Hard (BBH), TTT on in-context examples surpasses standard few-shot prompting in the $10$-shot setting by $7.3$ percentage points ($50.5%$ to $57.8%$). Our findings highlight the limitations of in-context learning for novel tasks and demonstrate the potential of test-time training to enhance language model adaptability. Ekin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu, Jyothish Pari, Jacob Andreas |
ICML | 8 |
| 2025 | A Hitchhiker's Guide to Scaling Law EstimationabstractScaling laws predict the loss of a target machine learning model by extrapolating from easier-to-train models with fewer parameters or smaller training sets. This provides an efficient way for practitioners and researchers alike to compare pretraining decisions involving optimizers, datasets, and model architectures. Despite the widespread use of scaling laws to model the dynamics of language model training, there has been little work on understanding how to best estimate and interpret them. We collect (and release) a large-scale dataset containing losses and downstream evaluations for 485 previously published pretrained models. We use these to estimate more than 1000 scaling laws, then derive a set of best practices for estimating scaling laws in new model families. We find that fitting scaling laws to intermediate checkpoints of training runs (and not just their final losses) substantially improves accuracy, and that—all else equal—estimates of performance are generally most accurate when derived from other models of similar sizes. However, because there is a significant degree of variability across model seeds, training multiple small models is sometimes more useful than training a single large one. Moreover, while different model families differ in scaling behavior, they are often similar enough that a target model’s behavior can be predicted from a single model with the same architecture, along with scaling parameter estimates derived from other model families. Leshem Choshen, Yang Zhang 0001, Jacob Andreas |
ICML | 3 |
| 2025 | (How) Do Language Models Track State?abstractTransformer language models (LMs) exhibit behaviors—from storytelling to code generation—that seem to require tracking the unobserved state of an evolving world. How do they do this? We study state tracking in LMs trained or fine-tuned to compose permutations (i.e., to compute the order of a set of objects after a sequence of swaps). Despite the simple algebraic structure of this problem, many other tasks (e.g., simulation of finite automata and evaluation of boolean expressions) can be reduced to permutation composition, making it a natural model for state tracking in general. We show that LMs consistently learn one of two state tracking mechanisms for this task. The first closely resembles the “associative scan” construction used in recent theoretical work by Liu et al. (2023) and Merrill et al. (2024). The second uses an easy-to-compute feature (permutation parity) to partially prune the space of outputs, and then refines this with an associative scan. LMs that learn the former algorithm tend to generalize better and converge faster, and we show how to steer LMs toward one or the other with intermediate training tasks that encourage or suppress the heuristics. Our results demonstrate that transformer LMs, whether pre-trained or fine-tuned, can learn to implement efficient and interpretable state-tracking mechanisms, and the emergence of these mechanisms can be predicted and controlled. Code and data are available at https://github.com/belindal/state-tracking Belinda Z. Li, Zifan Carl Guo, Jacob Andreas |
ICML | 3 |
| 2025 | A Probabilistic Framework for LLM Hallucination Detection via Belief Tree PropagationabstractBairu Hou, Yang Zhang, Jacob Andreas, Shiyu Chang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Bairu Hou, Yang Zhang 0001, Jacob Andreas, Shiyu Chang |
NAACL (Long Papers) | 3 |
| 2025 | Automated Detection of Visual Attribute Reliance with a Self-Reflective AgentabstractWhen a vision model performs image recognition, which visual attributes drive its predictions? Detecting unintended reliance on specific visual features is critical for ensuring model robustness, preventing overfitting, and avoiding spurious correlations. We introduce an automated framework for detecting such dependencies in trained vision models. At the core of our method is a self-reflective agent that systematically generates and tests hypotheses about visual attributes that a model may rely on. This process is iterative: the agent refines its hypotheses based on experimental outcomes and uses a self-evaluation protocol to assess whether its findings accurately explain model behavior. When inconsistencies arise, the agent self-reflects over its findings and triggers a new cycle of experimentation. We evaluate our approach on a novel benchmark of 130 models designed to exhibit diverse visual attribute dependencies across 18 categories. Our results show that the agent's performance consistently improves with self-reflection, with a significant performance increase over non-reflective baselines. We further demonstrate that the agent identifies real-world visual attribute dependencies in state-of-the-art models, including CLIP's vision encoder and the YOLOv8 object detector. Christy Li, Josep López Camuñas, Jake Thomas Touchet, Jacob Andreas, Àgata Lapedriza, Antonio Torralba 0001, Tamar Rott Shaham |
NeurIPS | 4 |
| 2025 | LoRA vs Full Fine-tuning: An Illusion of EquivalenceabstractFine-tuning is a crucial paradigm for adapting pre-trained large language models to downstream tasks. Recently, methods like Low-Rank Adaptation (LoRA) have been shown to effectively fine-tune LLMs with an extreme reduction in trainable parameters. But, \emph{are their learned solutions really equivalent?} We study how LoRA and full-finetuning change pre-trained models by analyzing the model's weight matrices through the lens of their spectral properties. We find that LoRA and full fine-tuning yield weight matrices whose singular value decompositions exhibit very different structure: weight matrices trained with LoRA have new, high-ranking singular vectors, which we call \emph{intruder dimensions}, while those trained with full fine-tuning do not. Further, we extend the finding that LoRA forgets less than full fine-tuning and find its forgetting is vastly localized to the intruder dimension -- by causally intervening on the intruder dimensions by changing their associated singular values post-fine-tuning, we show that they cause forgetting. Moreover, scaling them down significantly improves modeling of the pre-training distribution with a minimal drop in downstream task performance. Given this, we should expect accumulating intruder dimensions to be harmful and lead to more forgetting. This will be amplified during continual learning because of sequentially fine-tuning, and we show that LoRA models do accumulate intruder dimensions here tend to perform worse in this setting, emphasizing the practicality of our findings. Reece Shuttleworth, Jacob Andreas, Antonio Torralba 0001, Pratyusha Sharma |
NeurIPS | 2 |
| 2025 | Learning Linear Attention in Polynomial TimeabstractPrevious research has explored the expressivity of Transformer models in simulating Boolean circuits or Turing machines. However, the efficient learnability of Transformers from data has remained an open question. Our study addresses this gap by providing the first polynomial-time learnability results (specifically strong, agnostic PAC learning) for single-layer Transformers with linear attention. We show that learning the optimal multi head linear attention can be recast as finding the optimal kernel predictor in a suitably defined RKHS. Moving to generalization, we construct an algorithm that, given a dataset, checks in polynomial time whether the set of best fit multi head linear attention networks on this data all perform an identical computation--a powerful notion for out of distribution generalization. We empirically validate our theoretical findings on several canonical tasks: learning random linear attention networks, key--value associations, and learning to execute finite automata. Our findings bridge a critical gap between theoretical expressivity and learnability of Transformer models. Morris Yau, Ekin Akyürek, Jiayuan Mao, Josh Tenenbaum, Stefanie Jegelka, Jacob Andreas |
NeurIPS | 6 |
| 2025 | Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models
Anna A. Ivanova, Aalok Sathe, Benjamin Lipkin, Unnathi U. Kumar, Setayesh Radkani, Thomas Hikaru Clark, Carina Kauf, Jennifer Hu 0001, R. T. Pramod, Gabriel Grand, Vivian C. Paulun, Maria Ryskina, Ekin Akyürek, Ethan Wilcox, Nafisa Rashid, Leshem Choshen, Roger Levy, Evelina Fedorenko, Josh Tenenbaum, Jacob Andreas |
Trans. Assoc. Comput. Linguistics | 20 |
| 2024 | Toward In-Context Teaching: Adapting Examples to Students' MisconceptionsabstractWhen a teacher provides examples for a student to study, these examples must be informative, enabling a student to progress from their current state toward a target concept or skill.Good teachers must therefore simultaneously infer what students already know and adapt their teaching to students' changing state of knowledge.There is increasing interest in using computational models, particularly large language models, as pedagogical tools.As students, language models in particular have shown a remarkable ability to adapt to new tasks given small numbers of examples.But how effectively can these models adapt as teachers to students of different types?To study this question, we introduce a suite of models and evaluation methods we call ADAPT.ADAPT has two components: (1) a collection of simulated Bayesian student models that can be used for evaluation of automated teaching methods; (2) a platform for evaluation with human students, to characterize the real-world effectiveness of these methods.We additionally introduce (3) ATOM, a new probabilistic method for adaptive teaching that jointly infers students' past beliefs and optimizes for the correctness of future beliefs.In evaluations of simulated students across three learning domains (fraction arithmetic, English morphology, function learning), ATOM systematically outperforms LLM-based and standard Bayesian teaching methods.In human experiments, both ATOM and LLMs outperform non-adaptive random example selection.Our results highlight both the difficulty of the adaptive teaching task and the potential of learned adaptive methods for solving it. Alexis Ross, Jacob Andreas |
ACL (1) | 2 |
| 2024 | Naturalistic Transmission of Causal Knowledge between Machines and Humans
Cédric Colas, Tracey Mills, Ben Prystawski, Michael Henry Tessler, Noah D. Goodman, Jacob Andreas, Josh Tenenbaum |
CogSci | 6 |
| 2024 | Loose LIPS Sink Ships: Asking Questions in Battleship with Language-Informed Program Sampling
Gabriel Grand, Valerio Pepe, Jacob Andreas, Josh Tenenbaum |
CogSci | 3 |
| 2024 | Language-to-Code Translation with a Single Labeled ExampleabstractKaj Bostrom, Harsh Jhamtani, Hao Fang, Sam Thomson, Richard Shin, Patrick Xia, Benjamin Van Durme, Jason Eisner, Jacob Andreas. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Kaj Bostrom, Harsh Jhamtani, Hao Fang 0002, Sam Thomson, Richard Shin, Patrick Xia 0002, Benjamin Van Durme, Jason Eisner, Jacob Andreas |
EMNLP | 9 |
| 2024 | MisinfoEval: Generative AI in the Era of "Alternative Facts"abstractThe spread of misinformation on social media platforms threatens democratic processes, contributes to massive economic losses, and endangers public health.Many efforts to address misinformation focus on a knowledge deficit model and propose interventions for improving users' critical thinking through access to facts.Such efforts are often hampered by challenges with scalability, and by platform users' personal biases.The emergence of generative AI presents promising opportunities for countering misinformation at scale across ideological barriers.In this paper, we introduce a framework (Mis-infoEval) for generating and comprehensively evaluating large language model (LLM) based misinformation interventions.We present (1) an experiment with a simulated social media environment to measure effectiveness of misinformation interventions, and (2) a second experiment with personalized explanations tailored to the demographics and beliefs of users with the goal of countering misinformation by appealing to their pre-existing values.Our findings confirm that LLM-based interventions are highly effective at correcting user behavior (improving overall user accuracy at reliability labeling by up to 41.72%).Furthermore, we find that users favor more personalized interventions when making decisions about news reliability and users shown personalized interventions have significantly higher accuracy at identifying misinformation. Saadia Gabriel, Liang Lyu 0001, James Siderius, Marzyeh Ghassemi, Jacob Andreas, Asuman E. Ozdaglar |
EMNLP | 5 |
| 2024 | LILO: Learning Interpretable Libraries by Compressing and Documenting CodeabstractWhile large language models (LLMs) now excel at code generation, a key aspect of software development is the art of refactoring: consolidating code into libraries of reusable and readable programs. In this paper, we introduce LILO, a neurosymbolic framework that iteratively synthesizes, compresses, and documents code to build libraries tailored to particular problem domains. LILO combines LLM-guided program synthesis with recent algorithmic advances in automated refactoring from Stitch: a symbolic compression system that efficiently identifies optimal lambda abstractions across large code corpora. To make these abstractions interpretable, we introduce an auto-documentation (AutoDoc) procedure that infers natural language names and docstrings based on contextual examples of usage. In addition to improving human readability, we find that AutoDoc boosts performance by helping LILO's synthesizer to interpret and deploy learned abstractions. We evaluate LILO on three inductive program synthesis benchmarks for string editing, scene reasoning, and graphics composition. Compared to existing neural and symbolic methods—including the state-of-the-art library learning algorithm DreamCoder—LILO solves more complex tasks and learns richer libraries that are grounded in linguistic knowledge. Gabriel Grand, Lionel Wong, Matthew Bowers, Theo X. Olausson, Muxin Liu, Josh Tenenbaum, Jacob Andreas |
ICLR | 7 |
| 2024 | Linearity of Relation Decoding in Transformer Language ModelsabstractMuch of the knowledge encoded in transformer language models (LMs) may be expressed in terms of relations: relations between words and their synonyms, entities and their attributes, etc. We show that, for a subset of relations, this computation is well-approximated by a single linear transformation on the subject representation. Linear relation representations may be obtained by constructing a first-order approximation to the LM from a single prompt, and they exist for a variety of factual, commonsense, and linguistic relations. However, we also identify many cases in which LM predictions capture relational knowledge accurately, but this knowledge is not linearly encoded in their representations. Our results thus reveal a simple, interpretable, but heterogeneously deployed knowledge representation strategy in transformer LMs. Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, David Bau |
ICLR | 6 |
| 2024 | Modeling Boundedly Rational Agents with Latent Inference BudgetsabstractWe study the problem of modeling a population of agents pursuing unknown goals subject to unknown computational constraints. In standard models of bounded rationality, sub-optimal decision-making is simulated by adding homoscedastic noise to optimal decisions rather than actually simulating constrained inference. In this work, we introduce a latent inference budget model (L-IBM) that models these constraints explicitly, via a latent variable (inferred jointly with a model of agents’ goals) that controls the runtime of an iterative inference algorithm. L-IBMs make it possible to learn agent models using data from diverse populations of suboptimal actors. In three modeling tasks—inferring navigation goals from routes, inferring communicative intents from human utterances, and predicting next moves in human chess games—we show that L-IBMs match or outperforms Boltzmann models of decision-making under uncertainty. Moreover, the inferred inference budgets are themselves meaningful, efficient to compute, and correlated with measures of player skill, partner skill and task difficulty. Athul Paul Jacob, Abhishek Gupta 0004, Jacob Andreas |
ICLR | 3 |
| 2024 | The Consensus Game: Language Model Generation via Equilibrium SearchabstractWhen applied to question answering and other text generation tasks, language models (LMs) may be queried generatively (by sampling answers from their output distribution) or discriminatively (by using them to score or rank a set of candidate answers). These procedures sometimes yield very different predictions. How do we reconcile mutually incompatible scoring procedures to obtain coherent LM predictions? We introduce a new, a training-free, game-theoretic procedure for language model decoding. Our approach casts language model decoding as a regularized imperfect-information sequential signaling game—which we term the concensus game—in which a generator seeks to communicate an abstract correctness parameter using natural language sentences to a discriminator. We develop computational procedures for finding approximate equilibria of this game, resulting in a decoding algorithm we call equilibrium-ranking. Applied to a large number of tasks (including reading comprehension, commonsense reasoning, mathematical problem-solving, and assistive dialog), equilibrium-ranking consistently improves performance over existing LM decoding procedures. These improvements are sometimes substantial—on multiple benchmarks, we observe that applying equilibrium-ranking to LLaMA-7B outperforms the much larger LLaMA-65B and PaLM-540B models. Athul Paul Jacob, Yikang Shen, Gabriele Farina, Jacob Andreas |
ICLR | 4 |
| 2024 | Learning with Language-Guided State AbstractionsabstractWe describe a framework for using natural language to design state abstractions for imitation learning.
Generalizable policy learning in high-dimensional observation spaces is facilitated by well-designed state representations, which can surface important features of an environment and hide irrelevant ones.
These state representations are typically manually specified, or derived from other labor-intensive labeling procedures.
Our method, LGA (\textit{language-guided abstraction}), uses a combination of natural language supervision and background knowledge from language models (LMs) to automatically build state representations tailored to unseen tasks.
In LGA, a user first provides a (possibly incomplete) description of a target task in natural language; next, a pre-trained LM translates this task description into a state abstraction function that masks out irrelevant features; finally, an imitation policy is trained using a small number of demonstrations and LGA-generated abstract states.
Experiments on simulated robotic tasks show that LGA yields state abstractions similar to those designed by humans, but in a fraction of the time, and that these abstractions improve generalization and robustness in the presence of spurious correlations and ambiguous specifications.
We illustrate the utility of the learned abstractions on mobile manipulation tasks with a Spot robot. Andi Peng, Ilia Sucholutsky, Belinda Z. Li, Theodore R. Sumers, Thomas L. Griffiths 0001, Jacob Andreas, Julie A. Shah |
ICLR | 6 |
| 2024 | Learning Grounded Action Abstractions from LanguageabstractEffective planning in the real world requires not only world knowledge, but the ability to leverage that knowledge to build the right representation of the task at hand. Decades of hierarchical planning techniques have used domain-specific temporal action abstractions to support efficient and accurate planning, almost always relying on human priors and domain knowledge to decompose hard tasks into smaller subproblems appropriate for a goal or set of goals. This paper describes Ada (Action Domain Acquisition), a framework for automatically constructing task-specific planning representations using task-general background knowledge from language models (LMs). Starting with a general-purpose hierarchical planner and a low-level goal-conditioned policy, Ada interactively learns a library of planner-compatible high-level action abstractions and low-level controllers adapted to a particular domain of planning tasks. On two language-guided interactive planning benchmarks (Mini Minecraft and ALFRED Household Tasks), Ada strongly outperforms other approaches that use LMs for sequential decision-making, offering more accurate plans and better generalization to complex tasks. Lionel Wong, Jiayuan Mao, Pratyusha Sharma, Zachary S. Siegel, Jiahai Feng, Noa Korneev, Josh Tenenbaum, Jacob Andreas |
ICLR | 8 |
| 2024 | In-Context Language Learning: Architectures and AlgorithmsabstractSome neural language models (LMs) exhibit a remarkable capacity for in-context learning (ICL): they can fit predictors to datasets provided as input. While the mechanisms underlying ICL are well-studied in the context of synthetic problems like in-context linear regression, there is still some divergence between these model problems and the “real” ICL exhibited by LMs trained on large text corpora. In this paper, we study ICL through the lens of a new family of model problems we term in context language learning (ICLL). In ICLL, LMs are presented with a set of strings from a formal language, and must generate additional strings from the same language. We focus on in- context learning of regular languages generated by random finite automata. We evaluate a diverse set of neural sequence models on regular ICLL tasks. We first show that Transformers significantly outperform neural sequence models with recurrent or convolutional representations on ICLL tasks. Next, we provide evidence that they do so by computing in-context n-gram statistics using specialized attention heads. Finally, we show that hard-wiring these heads into neural models improves performance not just on synthetic ICLL, but natural language modeling, reducing the perplexity of 340M-parameter Transformers by up to 1.14 points (6.7%) on the SlimPajama dataset. Our results highlight the usefulness of in-context formal language learning as a tool for understanding ICL in models of natural text. Ekin Akyürek, Bailin Wang, Jacob Andreas |
ICML | 4 |
| 2024 | Decomposing Uncertainty for Large Language Models through Input Clarification EnsemblingabstractUncertainty decomposition refers to the task of decomposing the total uncertainty of a predictive model into aleatoric (data) uncertainty, resulting from inherent randomness in the data-generating process, and epistemic (model) uncertainty, resulting from missing information in the model's training data. In large language models (LLMs) specifically, identifying sources of uncertainty is an important step toward improving reliability, trustworthiness, and interpretability, but remains an important open research question. In this paper, we introduce an uncertainty decomposition framework for LLMs, called input clarification ensembling, which can be applied to any pre-trained LLM. Our approach generates a set of clarifications for the input, feeds them into an LLM, and ensembles the corresponding predictions. We show that, when aleatoric uncertainty arises from ambiguity or under-specification in LLM inputs, this approach makes it possible to factor an (un-clarified) LLM's predictions into separate aleatoric and epistemic terms, using a decomposition similar to the one employed by Bayesian neural networks. Empirical evaluations demonstrate that input clarification ensembling provides accurate and reliable uncertainty quantification on several language processing tasks. Code and data are available at https://github.com/UCSB-NLP-Chang/llm_uncertainty. Bairu Hou, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang, Yang Zhang 0001 |
ICML | 4 |
| 2024 | A Multimodal Automated Interpretability AgentabstractThis paper describes MAIA, a Multimodal Automated Interpretability Agent. MAIA is a system that uses neural models to automate neural model understanding tasks like feature interpretation and failure mode discovery. It equips a pre-trained vision-language model with a set of tools that support iterative experimentation on subcomponents of other models to explain their behavior. These include tools commonly used by human interpretability researchers: for synthesizing and editing inputs, computing maximally activating exemplars from real-world datasets, and summarizing and describing experimental results. Interpretability experiments proposed by MAIA compose these tools to describe and explain system behavior. We evaluate applications of MAIA to computer vision models. We first characterize MAIA’s ability to describe (neuron-level) features in learned representations of images. Across several trained models and a novel dataset of synthetic vision neurons with paired ground-truth descriptions, MAIA produces descriptions comparable to those generated by expert human experimenters. We then show that MAIA can aid in two additional interpretability tasks: reducing sensitivity to spurious features, and automatically identifying inputs likely to be mis-classified. Tamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram, Evan Hernandez, Jacob Andreas, Antonio Torralba 0001 |
ICML | 6 |
| 2024 | Natural Language Decomposition and Interpretation of Complex Utterances
Harsh Jhamtani, Hao Fang 0002, Patrick Xia 0002, Eran Levy, Jacob Andreas, Benjamin Van Durme |
IJCAI | 5 |
| 2024 | Regularized Conventions: Equilibrium Computation as a Model of Pragmatic ReasoningabstractAthul Jacob, Gabriele Farina, Jacob Andreas. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Athul Paul Jacob, Gabriele Farina, Jacob Andreas |
NAACL-HLT | 3 |
| 2024 | Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual TasksabstractZhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas, Yoon Kim. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen 0003, Bailin Wang, Najoung Kim, Jacob Andreas |
NAACL-HLT | 8 |
| 2024 | Visual Grounding Helps Learn Word Meanings in Low-Data RegimesabstractChengxu Zhuang, Evelina Fedorenko, Jacob Andreas. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chengxu Zhuang, Evelina Fedorenko, Jacob Andreas |
NAACL-HLT | 3 |
| 2024 | Algorithmic Capabilities of Random TransformersabstractTrained transformer models have been found to implement interpretable procedures for tasks like arithmetic and associative recall, but little is understood about how the circuits that implement these procedures originate during training. To what extent do they depend on the supervisory signal provided to models, and to what extent are they attributable to behavior already present in models at the beginning of training? To investigate these questions, we investigate what functions can be learned by randomly initialized transformers in which only the embedding layers are optimized, so that the only input--output mappings learnable from data are those already implemented (up to a choice of encoding scheme) by the randomly initialized model. We find that these random transformers can perform a wide range of meaningful algorithmic tasks, including modular arithmetic, in-weights and in-context associative recall, decimal addition, parenthesis balancing, and even some aspects of natural language text generation. Our results indicate that some algorithmic capabilities are present in transformers (and accessible via appropriately structured inputs) even before these models are trained. Ziqian Zhong, Jacob Andreas |
NeurIPS | 2 |
| 2024 | Decision-Oriented Dialogue for Human-AI CollaborationabstractAbstract We describe a class of tasks called decision-oriented dialogues, in which AI assistants such as large language models (LMs) must collaborate with one or more humans via natural language to help them make complex decisions. We formalize three domains in which users face everyday decisions: (1) choosing an assignment of reviewers to conference papers, (2) planning a multi-step itinerary in a city, and (3) negotiating travel plans for a group of friends. In each of these settings, AI assistants and users have disparate abilities that they must combine to arrive at the best decision: Assistants can access and process large amounts of information, while users have preferences and constraints external to the system. For each task, we build a dialogue environment where agents receive a reward based on the quality of the final decision they reach. We evaluate LMs in self-play and in collaboration with humans and find that they fall short compared to human assistants, achieving much lower rewards despite engaging in longer dialogues. We highlight a number of challenges models face in decision-oriented dialogues, ranging from goal-directed behavior to reasoning and optimization, and release our environments as a testbed for future work. Jessy Lin, Nicholas Tomlin, Jacob Andreas, Jason Eisner |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | LexSym: Compositionality as Lexical SymmetryabstractIn tasks like semantic parsing, instruction following, and question answering, standard deep networks fail to generalize compositionally from small datasets.Many existing approaches overcome this limitation with model architectures that enforce a compositional process of sentence interpretation.In this paper, we present a domain-general and model-agnostic formulation of compositionality as a constraint on symmetries of data distributions rather than models.Informally, we prove that whenever a task can be solved by a compositional model, there is a corresponding data augmentation scheme-a procedure for transforming examples into other well-formed examples-that imparts compositional inductive bias on any model trained to solve the same task.We describe a procedure called LEXSYM that discovers these transformations automatically, then applies them to training data for ordinary neural sequence models.Unlike existing compositional data augmentation procedures, LEXSYM can be deployed agnostically across text, structured data, and even images.It matches or surpasses state-of-the-art, task-specific models on COGS semantic parsing, SCAN and ALCHEMY instruction following, and CLEVR-COGENT visual question answering datasets. Ekin Akyürek, Jacob Andreas |
ACL (1) | 2 |
| 2023 | Alignment via Mutual InformationabstractMany language learning tasks require learners to infer correspondences between data in two modalities.Often, these alignments are manyto-many and context-sensitive.For example, translating into morphologically rich languages requires learning not just how words, but morphemes, should be translated; words and morphemes may have different meanings (or groundings) depending on the context in which they are used.We describe an informationtheoretic approach to context-sensitive, manyto-many alignment.Our approach first trains a masked sequence model to place distributions over missing spans in (source, target) sequences.Next, it uses this model to compute pointwise mutual information between source and target spans conditional on context.Finally, it aligns spans with high mutual information.We apply this approach to two learning problems: character-based word translation (using alignments for joint morphological segmentation and lexicon learning) and visually grounded reference resolution (using alignments to jointly localize referents and learn word meanings).In both cases, our proposed approach outperforms both structured and neural baselines, showing that conditional mutual information offers an effective framework for formalizing alignment problems in general domains. Shinjini Ghosh, Ramón Fernandez Astudillo, Tahira Naseem, Jacob Andreas |
CoNLL | 5 |
| 2023 | Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?abstractNeural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal representations of truthfulness.Past work has found that these two procedures sometimes disagree, and that probes tend to be more accurate than LM outputs.This has led some researchers to conclude that LMs "lie" or otherwise encode non-cooperative communicative intents.Is this an accurate description of today's LMs, or can query-probe disagreement arise in other ways?We identify three different classes of disagreement, which we term confabulation, deception, and heterogeneity.In many cases, the superiority of probes is simply attributable to better calibration on uncertain answers rather than a greater fraction of correct, high-confidence answers.In some cases, queries and probes perform better on different subsets of inputs, and accuracy can further be improved by ensembling the two. 1 Kevin Liu, Stephen Casper, Dylan Hadfield-Menell, Jacob Andreas |
EMNLP | 4 |
| 2023 | Pushdown Layers: Encoding Recursive Structure in Transformer Language ModelsabstractRecursion is a prominent feature of human language, and fundamentally challenging for self-attention due to the lack of an explicit recursive-state tracking mechanism.Consequently, Transformer language models poorly capture long-tail recursive structure and exhibit sample-inefficient syntactic generalization.This work introduces Pushdown Layers, a new self-attention layer that models recursive state via a stack tape that tracks estimated depths of every token in an incremental parse of the observed prefix.Transformer LMs with Pushdown Layers are syntactic language models that autoregressively and synchronously update this stack tape as they predict new tokens, in turn using the stack tape to softly modulate attention over tokens-for instance, learning to "skip" over closed constituents.When trained on a corpus of strings annotated with silver constituency parses, Transformers equipped with Pushdown Layers achieve dramatically better and 3-5x more sample-efficient syntactic generalization, while maintaining similar perplexities.Pushdown Layers are a drop-in replacement for standard self-attention.We illustrate this by finetuning GPT2-medium with Pushdown Layers on an automatically parsed WikiText-103, leading to improvements on several GLUE text classification tasks. Shikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. Manning |
EMNLP | 3 |
| 2023 | What learning algorithm is in-context learning? Investigations with linear models
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma 0001, Denny Zhou |
ICLR | 3 |
| 2023 | Characterizing intrinsic compositionality in transformers with Tree Projections
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, Christopher D. Manning |
ICLR | 3 |
| 2023 | Guiding Pretraining in Reinforcement Learning with Large Language ModelsabstractReinforcement learning algorithms typically struggle in the absence of a dense, well-shaped reward function. Intrinsically motivated exploration methods address this limitation by rewarding agents for visiting novel states or transitions, but these methods offer limited benefits in large environments where most discovered novelty is irrelevant for downstream tasks. We describe a method that uses background knowledge from text corpora to shape exploration. This method, called ELLM (Exploring with LLMs) rewards an agent for achieving goals suggested by a language model prompted with a description of the agent's current state. By leveraging large-scale language model pretraining, ELLM guides agents toward human-meaningful and plausibly useful behaviors without requiring a human in the loop. We evaluate ELLM in the Crafter game environment and the Housekeep robotic simulator, showing that ELLM-trained agents have better coverage of common-sense behaviors during pretraining and usually match or improve performance on a range of downstream tasks. Olivia Watkins, Cédric Colas, Trevor Darrell, Pieter Abbeel, Abhishek Gupta 0004, Jacob Andreas |
ICML | 8 |
| 2023 | PromptBoosting: Black-Box Text Classification with Ten Forward PassesabstractWe describe PromptBoosting, a query-efficient procedure for building a text classifier from a neural language model (LM) without access to the LM’s parameters, gradients, or hidden representations. This form of "black-box" classifier training has become increasingly important as the cost of training and inference in large-scale LMs has grown. But existing black-box LM classifier learning approaches are themselves computationally inefficient, typically specializing LMs to the target task by searching in a large space of (discrete or continuous) prompts using zeroth-order optimization methods. Instead of directly optimizing in prompt space, PromptBoosting obtains a small pool of prompts via a gradient-free approach and then constructs a large pool of weak learners by pairing these prompts with different elements of the LM’s output distribution. These weak learners are then ensembled using the AdaBoost algorithm. The entire learning process requires only a small number of forward passes and no backward pass. Experiments show that PromptBoosting achieves state-of-the-art performance in multiple black-box few-shot classification tasks, and matches or outperforms full fine-tuning in both few-shot and standard learning paradigms, while training 10x faster than existing black-box methods. Bairu Hou, Joe O'Connor, Jacob Andreas, Shiyu Chang, Yang Zhang 0001 |
ICML | 3 |
| 2023 | FIND: A Function Description Benchmark for Evaluating Interpretability MethodsabstractLabeling neural network submodules with human-legible descriptions is useful for many downstream tasks: such descriptions can surface failures, guide interventions, and perhaps even explain important model behaviors. To date, most mechanistic descriptions of trained networks have involved small models, narrowly delimited phenomena, and large amounts of human labor. Labeling all human-interpretable sub-computations in models of increasing size and complexity will almost certainly require tools that can generate and validate descriptions automatically. Recently, techniques that use learned models in-the-loop for labeling have begun to gain traction, but methods for evaluating their efficacy are limited and ad-hoc. How should we validate and compare open-ended labeling tools? This paper introduces FIND (Function INterpretation and Description), a benchmark suite for evaluating the building blocks of automated interpretability methods. FIND contains functions that resemble components of trained neural networks, and accompanying descriptions of the kind we seek to generate. The functions are procedurally constructed across textual and numeric domains, and involve a range of real-world complexities, including noise, composition, approximation, and bias. We evaluate methods that use pretrained language models (LMs) to produce code-based and natural language descriptions of function behavior. Additionally, we introduce a new interactive method in which an Automated Interpretability Agent (AIA) generates function descriptions. We find that an AIA, built with an off-the-shelf LM augmented with black-box access to functions, can sometimes infer function structure—acting as a scientist by forming hypotheses, proposing experiments, and updating descriptions in light of new data. However, FIND also reveals that LM-based descriptions capture global function behavior while missing local details. These results suggest that FIND will be useful for characterizing the performance of more sophisticated interpretability methods before they are applied to real-world models. Sarah Schwettmann, Tamar Rott Shaham, Joanna Materzynska, Neil Chowdhury, Shuang Li 0013, Jacob Andreas, David Bau, Antonio Torralba 0001 |
NeurIPS | 6 |
| 2023 | The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural NetworksabstractDo neural networks, trained on well-understood algorithmic tasks, reliably rediscover known algorithms? Several recent studies, on tasks ranging from group operations to in-context linear regression, have suggested that the answer is yes. Using modular addition as a prototypical problem, we show that algorithm discovery in neural networks is sometimes more complex: small changes to model hyperparameters and initializations can induce discovery of qualitatively different algorithms from a fixed training set, and even learning of multiple different solutions in parallel. In modular addition, we specifically show that models learn a known *Clock* algorithm, a previously undescribed, less intuitive, but comprehensible procedure we term the *Pizza* algorithm, and a variety of even more complex procedures. Our results show that even simple learning problems can admit a surprising diversity of solutions, motivating the development of new tools for mechanistically characterizing the behavior of neural networks across the algorithmic phase space. Ziqian Zhong, Ziming Liu 0001, Max Tegmark, Jacob Andreas |
NeurIPS | 4 |
| 2022 | Skill Induction and Planning with Latent LanguageabstractWe present a framework for learning hierarchical policies from demonstrations, using sparse natural language annotations to guide the discovery of reusable skills for autonomous decision-making.We formulate a generative model of action sequences in which goals generate sequences of high-level subtask descriptions, and these descriptions generate sequences of low-level actions.We describe how to train this model using primarily unannotated demonstrations by parsing demonstrations into sequences of named high-level subtasks, using only a small number of seed annotations to ground language in action.In trained models, natural language commands index a combinatorial library of skills; agents can use these skills to plan by generating high-level instruction sequences tailored to novel goals.We evaluate this approach in the ALFRED household simulation environment, providing natural language annotations for only 10% of demonstrations.It achieves task completion rates comparable to state-of-the-art models (outperforming several recent methods with access to ground-truth plans during training and evaluation) while providing structured and human-readable high-level plans. 1 Pratyusha Sharma, Antonio Torralba 0001, Jacob Andreas |
ACL (1) | 3 |
| 2022 | Identifying concept libraries from language about object structure
Catherine Wong, William P. McCarthy, Gabriel Grand, Yoni Friedman, Josh Tenenbaum, Jacob Andreas, Robert D. Hawkins, Judith E. Fan |
CogSci | 6 |
| 2022 | Hierarchical Phrase-Based Sequence-to-Sequence LearningabstractWe describe a neural transducer that maintains the flexibility of standard sequence-to-sequence (seq2seq) models while incorporating hierarchical phrases as a source of inductive bias during training and as explicit constraints during inference.Our approach trains two models: a discriminative parser based on a bracketing transduction grammar whose derivation tree hierarchically aligns source and target phrases, and a neural seq2seq model that learns to translate the aligned phrases one-by-one.We use the same seq2seq model to translate at all phrase scales, which results in two inference modes: one mode in which the parser is discarded and only the seq2seq component is used at the sequence-level, and another in which the parser is combined with the seq2seq model.Decoding in the latter mode is done with the cube-pruned CKY algorithm, which is more involved but can make use of new translation rules during inference.We formalize our model as a sourceconditioned synchronous grammar and develop an efficient variational inference algorithm for training.When applied on top of both randomly initialized and pretrained seq2seq models, we find that both inference modes performs well compared to baselines on small scale machine translation benchmarks. Bailin Wang, Ivan Titov 0001, Jacob Andreas |
EMNLP | 3 |
| 2022 | Subspace Regularizers for Few-Shot Class Incremental Learning
Afra Feyza Akyürek, Ekin Akyürek, Derry Wijaya, Jacob Andreas |
ICLR | 4 |
| 2022 | Natural Language Descriptions of Deep Visual Features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba 0001, Jacob Andreas |
ICLR | 6 |
| 2022 | Modeling Strong and Human-Like Gameplay with KL-Regularized SearchabstractWe consider the task of accurately modeling strong human policies in multi-agent decision-making problems, given examples of human behavior. Imitation learning is effective at predicting human actions but may not match the strength of expert humans (e.g., by sometimes committing blunders), while self-play learning and search techniques such as AlphaZero lead to strong performance but may produce policies that differ markedly from human behavior. In chess and Go, we show that regularized search algorithms that penalize KL divergence from an imitation-learned policy yield higher prediction accuracy of strong humans and better performance than imitation learning alone. We then introduce a novel regret minimization algorithm that is regularized based on the KL divergence from an imitation-learned policy, and show that using this algorithm for search in no-press Diplomacy yields a policy that matches the human prediction accuracy of imitation learning while being substantially stronger. Athul Paul Jacob, David J. Wu 0002, Gabriele Farina, Adam Lerer, Hengyuan Hu, Anton Bakhtin, Jacob Andreas, Noam Brown |
ICML | 7 |
| 2022 | Quantifying Adaptability in Pre-trained Language Models with 500 TasksabstractBelinda Li, Jane Yu, Madian Khabsa, Luke Zettlemoyer, Alon Halevy, Jacob Andreas. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Belinda Z. Li, Jane Dwivedi-Yu, Madian Khabsa, Luke Zettlemoyer, Alon Y. Halevy, Jacob Andreas |
NAACL-HLT | 6 |
| 2022 | Pre-Trained Language Models for Interactive Decision-MakingabstractLanguage model (LM) pre-training is useful in many language processing tasks. But can pre-trained LMs be further leveraged for more general machine learning problems? We propose an approach for using LMs to scaffold learning and generalization in general sequential decision-making problems. In this approach, goals and observations are represented as a sequence of embeddings, and a policy network initialized with a pre-trained LM predicts the next action. We demonstrate that this framework enables effective combinatorial generalization across different environments and supervisory modalities. We begin by assuming access to a set of expert demonstrations, and show that initializing policies with LMs and fine-tuning them via behavior cloning improves task completion rates by 43.6% in the VirtualHome environment. Next, we integrate an active data gathering procedure in which agents iteratively interact with the environment, relabel past "failed" experiences with new goals, and update their policies in a self-supervised loop. Active data gathering further improves combinatorial generalization, outperforming the best baseline by 25.1%. Finally, we explain these results by investigating three possible factors underlying the effectiveness of the LM-based policy. We find that sequential input representations (vs. fixed-dimensional feature vectors) and LM-based weight initialization are both important for generalization. Surprisingly, however, the format of the policy inputs encoding (e.g. as a natural language string vs. an arbitrary sequential encoding) has little influence. Together, these results suggest that language modeling induces representations that are useful for modeling not just language, but also goals and plans; these representations can aid learning and generalization even outside of language processing. Shuang Li 0013, Xavier Puig, Chris Paxton 0001, Yilun Du, Clinton Wang, Linxi Fan, Tao Chen 0046, De-An Huang, Ekin Akyürek, Anima Anandkumar, Jacob Andreas, Igor Mordatch, Antonio Torralba 0001, Yuke Zhu |
NeurIPS | 11 |
| 2021 | Lexicon Learning for Few Shot Sequence ModelingabstractEkin Akyurek, Jacob Andreas. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ekin Akyürek, Jacob Andreas |
ACL/IJCNLP (1) | 2 |
| 2021 | Implicit Representations of Meaning in Neural Language ModelsabstractBelinda Z. Li, Maxwell Nye, Jacob Andreas. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Belinda Z. Li, Maxwell I. Nye, Jacob Andreas |
ACL/IJCNLP (1) | 3 |
| 2021 | What Context Features Can Transformer Language Models Use?abstractJoe O’Connor, Jacob Andreas. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Joe O'Connor, Jacob Andreas |
ACL/IJCNLP (1) | 2 |
| 2021 | Value-Agnostic Conversational Semantic ParsingabstractEmmanouil Antonios Platanios, Adam Pauls, Subhro Roy, Yuchen Zhang, Alexander Kyte, Alan Guo, Sam Thomson, Jayant Krishnamurthy, Jason Wolfe, Jacob Andreas, Dan Klein. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Emmanouil A. Platanios, Adam Pauls, Subhro Roy, Yuchen Zhang 0002, Alexander Kyte, Alan Guo, Sam Thomson, Jayant Krishnamurthy, Jason Andrew Wolfe, Jacob Andreas, Daniel Klein 0001 |
ACL/IJCNLP (1) | 10 |
| 2021 | Language as a bootstrap for compositional visual reasoning
Catherine Wong, Yoni Friedman, Jacob Andreas, Josh Tenenbaum |
CogSci | 3 |
| 2021 | The Low-Dimensional Linear Geometry of Contextualized Word RepresentationsabstractBlack-box probing models can reliably extract linguistic features like tense, number, and syntactic role from pretrained word representations.However, the manner in which these features are encoded in representations remains poorly understood.We present a systematic study of the linear geometry of contextualized word representations in ELMO and BERT.We show that a variety of linguistic features (including structured dependency relationships) are encoded in low-dimensional subspaces.We then refine this geometric picture, showing that there are hierarchical relations between the subspaces encoding general linguistic categories and more specific ones, and that lowdimensional feature encodings are distributed rather than aligned to individual neurons.Finally, we demonstrate that these linear subspaces are causally related to model behavior, and can be used to perform fine-grained manipulation of BERT's output distribution. Evan Hernandez, Jacob Andreas |
CoNLL | 2 |
| 2021 | How Do Neural Sequence Models Generalize? Local and Global Cues for Out-of-Distribution PredictionabstractAfter a neural sequence model encounters an unexpected token, can its behavior be predicted?We show that RNN and transformer language models exhibit structured, consistent generalization in out-of-distribution contexts.We begin by introducing two idealized models of generalization in next-word prediction: a local context model in which generalization is consistent with the last word observed, and a global context model in which generalization is consistent with the global structure of the input.In experiments in English, Finnish, Mandarin, and random regular languages, we demonstrate that neural language models interpolate between these two forms of generalization: their predictions are well-approximated by a log-linear combination of local and global predictive distributions.We then show that, in some languages, noise mediates the two forms of generalization: noise applied to input tokens encourages global generalization, while noise in history representations encourages local generalization.Finally, we offer a preliminary theoretical explanation of these results by proving that the observed interpolation behavior is expected in log-linear models with a particular feature correlation structure.These results help explain the effectiveness of two popular regularization schemes and show that aspects of sequence model generalization can be understood and controlled. D. Anthony Bau, Jacob Andreas |
EMNLP (1) | 2 |
| 2021 | Toward a Visual Concept Vocabulary for GAN Latent SpaceabstractA large body of recent work has identified transformations in the latent spaces of generative adversarial networks (GANs) that consistently and interpretably transform generated images. But existing techniques for identifying these transformations rely on either a fixed vocabulary of prespecified visual concepts, or on unsupervised disentanglement techniques whose alignment with human judgments about perceptual salience is unknown. This paper introduces a new method for building open-ended vocabularies of primitive visual concepts represented in a GAN’s latent space. Our approach is built from three components: (1) automatic identification of perceptually salient directions based on their layer selectivity; (2) human annotation of these directions with free-form, compositional natural language descriptions; and (3) decomposition of these annotations into a visual concept vocabulary, consisting of distilled directions labeled with single words. Experiments show that concepts learned with our approach are reliable and composable—generalizing across classes, contexts, and observers, and enabling fine-grained manipulation of image style and content. Sarah Schwettmann, Evan Hernandez, David Bau, Samuel Klein, Jacob Andreas, Antonio Torralba 0001 |
ICCV | 5 |
| 2021 | Learning to Recombine and Resample Data For Compositional Generalization
Ekin Akyürek, Afra Feyza Akyürek, Jacob Andreas |
ICLR | 3 |
| 2021 | Representing Partial Programs with Blended Abstract Semantics
Maxwell I. Nye, Yewen Pu, Matthew Bowers, Jacob Andreas, Josh Tenenbaum, Armando Solar-Lezama |
ICLR | 4 |
| 2021 | Leveraging Language to Learn Program Abstractions and Search HeuristicsabstractInductive program synthesis, or inferring programs from examples of desired behavior, offers a general paradigm for building interpretable, robust, andgeneralizable machine learning systems. Effective program synthesis depends on two key ingredients: a strong library of functions from which to build programs, and an efficient search strategy for finding programs that solve a given task. We introduce LAPS (Language for Abstraction and Program Search), a technique for using natural language annotations to guide joint learning of libraries and neurally-guided search models for synthesis. When integrated into a state-of-the-art library learning system (DreamCoder), LAPS produces higher-quality libraries and improves search efficiency and generalization on three domains {–} string editing, image composition, and abstract reasoning about scenes {–} even when no natural language hints are available at test time. Catherine Wong, Kevin Ellis, Josh Tenenbaum, Jacob Andreas |
ICML | 4 |
| 2021 | Multitasking Inhibits Semantic DriftabstractWhen intelligent agents communicate to accomplish shared goals, how do these goals shape the agents' language?We study the dynamics of learning in latent language policies (LLPs), in which instructor agents generate natural-language subgoal descriptions and executor agents map these descriptions to lowlevel actions.LLPs can solve challenging long-horizon reinforcement learning problems and provide a rich model for studying taskoriented language use.But previous work has found that LLP training is prone to semantic drift (use of messages in ways inconsistent with their original natural language meanings).Here, we demonstrate theoretically and empirically that multitask training is an effective counter to this problem: we prove that multitask training eliminates semantic drift in a well-studied family of signaling games, and show that multitask training of neural LLPs in a complex strategy game reduces drift and while improving sample efficiency. Athul Paul Jacob, Mike Lewis, Jacob Andreas |
NAACL-HLT | 3 |
| 2021 | Compositional Generalization for Neural Semantic Parsing via Span-level Supervised AttentionabstractPengcheng Yin, Hao Fang, Graham Neubig, Adam Pauls, Emmanouil Antonios Platanios, Yu Su, Sam Thomson, Jacob Andreas. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Hao Fang 0002, Graham Neubig, Adam Pauls, Emmanouil A. Platanios, Yu Su 0001, Sam Thomson, Jacob Andreas |
NAACL-HLT | 8 |
| 2021 | Teachable Reinforcement Learning via Advice DistillationabstractTraining automated agents to complete complex tasks in interactive environments is challenging: reinforcement learning requires careful hand-engineering of reward functions, imitation learning requires specialized infrastructure and access to a human expert, and learning from intermediate forms of supervision (like binary preferences) is time-consuming and extracts little information from each human intervention. Can we overcome these challenges by building agents that learn from rich, interactive feedback instead? We propose a new supervision paradigm for interactive learning based on "teachable" decision-making systems that learn from structured advice provided by an external teacher. We begin by formalizing a class of human-in-the-loop decision making problems in which multiple forms of teacher-provided advice are available to a learner. We then describe a simple learning algorithm for these problems that first learns to interpret advice, then learns from advice to complete tasks even in the absence of human supervision. In puzzle-solving, navigation, and locomotion domains, we show that agents that learn from advice can acquire new skills with significantly less human supervision than standard reinforcement learning algorithms and often less than imitation learning. Olivia Watkins, Abhishek Gupta 0004, Trevor Darrell, Pieter Abbeel, Jacob Andreas |
NeurIPS | 5 |
| 2020 | Good-Enough Compositional Data AugmentationabstractWe propose a simple data augmentation protocol aimed at providing a compositional inductive bias in conditional and unconditional sequence models.Under this protocol, synthetic training examples are constructed by taking real training examples and replacing (possibly discontinuous) fragments with other fragments that appear in at least one similar environment.The protocol is model-agnostic and useful for a variety of tasks.Applied to neural sequence-to-sequence models, it reduces error rate by as much as 87% on diagnostic tasks from the SCAN dataset and 16% on a semantic parsing task.Applied to n-gram language models, it reduces perplexity by roughly 1% on small corpora in several languages. Jacob Andreas |
ACL | 1 |
| 2020 | Experience Grounds LanguageabstractYonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, Joseph Turian. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Y. Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, Nicolas Pinto, Joseph P. Turian |
EMNLP (1) | 4 |
| 2020 | Joint Modeling of Chest Radiographs and Radiology Reports for Pulmonary Edema Assessment
Geeticka Chauhan, Ruizhi Liao 0001, William M. Wells III, Jacob Andreas, Seth J. Berkowitz, Steven Horng, Peter Szolovits, Polina Golland |
MICCAI (2) | 4 |
| 2020 | Compositional Explanations of NeuronsabstractWe describe a procedure for explaining neurons in deep representations by identifying compositional logical concepts that closely approximate neuron behavior. Compared to prior work that uses atomic labels as explanations, analyzing neurons compositionally allows us to more precisely and expressively characterize their behavior. We use this procedure to answer several questions on interpretability in models for vision and natural language processing. First, we examine the kinds of abstractions learned by neurons. In image classification, we find that many neurons learn highly abstract but semantically coherent visual concepts, while other polysemantic neurons detect multiple unrelated features; in natural language inference (NLI), neurons learn shallow lexical heuristics from dataset biases. Second, we see whether compositional explanations give us insight into model performance: vision neurons that detect human-interpretable concepts are positively correlated with task performance, while NLI neurons that fire for shallow heuristics are negatively correlated with task performance. Finally, we show how compositional explanations provide an accessible way for end users to produce simple "copy-paste" adversarial examples that change model behavior in predictable ways. Jesse Mu, Jacob Andreas |
NeurIPS | 2 |
| 2020 | A Benchmark for Systematic Generalization in Grounded Language UnderstandingabstractHumans easily interpret expressions that describe unfamiliar situations composed from familiar parts ("greet the pink brontosaurus by the ferris wheel"). Modern neural networks, by contrast, struggle to interpret novel compositions. In this paper, we introduce a new benchmark, gSCAN, for evaluating compositional generalization in situated language understanding. Going beyond a related benchmark that focused on syntactic aspects of generalization, gSCAN defines a language grounded in the states of a grid world, facilitating novel evaluations of acquiring linguistically motivated rules. For example, agents must understand how adjectives such as 'small' are interpreted relative to the current world state or how adverbs such as 'cautiously' combine with new verbs. We test a strong multi-modal baseline model and a state-of-the-art compositional method finding that, in most cases, they fail dramatically when generalization requires systematic compositional rules. Laura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt, Brenden M. Lake |
NeurIPS | 2 |
| 2020 | Task-Oriented Dialogue as Dataflow SynthesisabstractWe describe an approach to task-oriented dialogue in which dialogue state is represented as a dataflow graph. A dialogue agent maps each user utterance to a program that extends this graph. Programs include metacomputation operators for reference and revision that reuse dataflow fragments from previous turns. Our graph-based state enables the expression and manipulation of complex user intents, and explicit metacomputation makes these intents easier for learned models to predict. We introduce a new dataset, SMCalFlow, featuring complex dialogues about events, weather, places, and people. Experiments show that dataflow graphs and metacomputation substantially improve representability and predictability in these natural dialogues. Additional experiments on the MultiWOZ dataset show that our dataflow representation enables an otherwise off-the-shelf sequence-to-sequence model to match the best existing task-specific state tracking model. The SMCalFlow dataset, code for replicating experiments, and a public leaderboard are available at https://www.microsoft.com/en-us/research/project/dataflow-based-dialogue-semantic-machines . Jacob Andreas, John Bufe, David Burkett, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, Hao Fang 0002, Alan Guo, David Hall 0006, Kristin Hayes, Kellie Hill, Diana Ho, Wendy Iwaszuk, Smriti Jha, Daniel Klein 0001, Jayant Krishnamurthy, Theo Lanman, Percy Liang, Christopher H. Lin, Ilya Lintsbakh, Andy McGovern, Aleksandr Nisnevich, Adam Pauls, Dmitrij Petters, Brent Read, Dan Roth 0001, Subhro Roy, Jesse Rusak, Beth Short, Div Slomin, Ben Snyder, Stephon Striplin, Yu Su 0001, Zachary Tellman, Sam Thomson, Andrei Vorobev, Izabela Witoszko, Jason Andrew Wolfe, Abby Wray, Yuchen Zhang 0002, Alexander Zotov |
Trans. Assoc. Comput. Linguistics | 1 |
| 2019 | Measuring Compositionality in Representation Learning
Jacob Andreas |
ICLR (Poster) | 1 |
| 2019 | Guiding Policies with Language via Meta-Learning
John D. Co-Reyes, Abhishek Gupta 0004, Suvansh Sanjeev, Nick Altieri 0001, Jacob Andreas, John DeNero, Pieter Abbeel, Sergey Levine |
ICLR (Poster) | 5 |
| 2019 | A Survey of Reinforcement Learning Informed by Natural LanguageabstractTo be successful in real-world tasks, Reinforcement Learning (RL) needs to exploit the compositional, relational, and hierarchical structure of the world, and learn to transfer it to the task at hand. Recent advances in representation learning for language make it possible to build models that acquire world knowledge from text corpora and integrate this knowledge into downstream decision making problems. We thus argue that the time is right to investigate a tight integration of natural language understanding into RL in particular. We survey the state of the field, including work on instruction following, text games, and learning from textual domain knowledge. Finally, we call for the development of new environments as well as further investigation into the potential uses of recent Natural Language Processing (NLP) techniques for such tasks. Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob N. Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, Tim Rocktäschel |
IJCAI | 5 |
| 2018 | Explainable Neural Computation via Stack Neural Module Networks
Ronghang Hu, Jacob Andreas, Trevor Darrell, Kate Saenko |
ECCV (7) | 2 |
| 2018 | Can Deep Reinforcement Learning Solve Erdos-Selfridge-Spencer Games?abstractDeep reinforcement learning has achieved many recent successes, but our understanding of its strengths and limitations is hampered by the lack of rich environments in which we can fully characterize optimal behavior, and correspondingly diagnose individual actions against such a characterization. Here we consider a family of combinatorial games, arising from work of Erdos, Selfridge, and Spencer, and we propose their use as environments for evaluating and comparing different approaches to reinforcement learning. These games have a number of appealing features: they are challenging for current learning approaches, but they form (i) a low-dimensional, simply parametrized environment where (ii) there is a linear closed form solution for optimal behavior from any state, and (iii) the difficulty of the game can be tuned by changing environment parameters in an interpretable way. We use these Erdos-Selfridge-Spencer games not only to compare different algorithms, but test for generalization, make comparisons to supervised learning, analyse multiagent play, and even develop a self play algorithm. Maithra Raghu, Alex Irpan, Jacob Andreas, Robert D. Kleinberg, Quoc V. Le, Jon M. Kleinberg |
ICML | 3 |
| 2018 | Learning with Latent LanguageabstractJacob Andreas, Dan Klein, Sergey Levine. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Jacob Andreas, Daniel Klein 0001, Sergey Levine |
NAACL-HLT | 1 |
| 2018 | Unified Pragmatic Models for Generating and Following InstructionsabstractDaniel Fried, Jacob Andreas, Dan Klein. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Daniel Fried, Jacob Andreas, Daniel Klein 0001 |
NAACL-HLT | 2 |
| 2018 | Speaker-Follower Models for Vision-and-Language NavigationabstractNavigation guided by natural language instructions presents a challenging reasoning problem for instruction followers. Natural language instructions typically identify only a few high-level decisions and landmarks rather than complete low-level motor behaviors; much of the missing information must be inferred based on perceptual context. In machine learning settings, this is doubly challenging: it is difficult to collect enough annotated data to enable learning of this reasoning process from scratch, and also difficult to implement the reasoning process using generic sequence models. Here we describe an approach to vision-and-language navigation that addresses both these issues with an embedded speaker model. We use this speaker model to (1) synthesize new instructions for data augmentation and to (2) implement pragmatic reasoning, which evaluates how well candidate action sequences explain an instruction. Both steps are supported by a panoramic action space that reflects the granularity of human-generated instructions. Experiments show that all three components of this approach---speaker-driven data augmentation, pragmatic reasoning and panoramic action space---dramatically improve the performance of a baseline instruction follower, more than doubling the success rate over the best existing approach on a standard benchmark. Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Daniel Klein 0001, Trevor Darrell |
NeurIPS | 5 |
| 2017 | Translating NeuraleseabstractSeveral approaches have recently been proposed for learning decentralized deep multiagent policies that coordinate via a differentiable communication channel.While these policies are effective for many tasks, interpretation of their induced communication strategies has remained a challenge.Here we propose to interpret agents' messages by translating them.Unlike in typical machine translation problems, we have no parallel data to learn from.Instead we develop a translation model based on the insight that agent messages and natural language strings mean the same thing if they induce the same belief about the world in a listener.We present theoretical guarantees and empirical evidence that our approach preserves both the semantics and pragmatics of messages by ensuring that players communicating through a translation layer do not suffer a substantial loss in reward relative to players with a common language.1 Jacob Andreas, Anca D. Dragan, Daniel Klein 0001 |
ACL (1) | 1 |
| 2017 | A Minimal Span-Based Neural Constituency ParserabstractIn this work, we present a minimal neural model for constituency parsing based on independent scoring of labels and spans.We show that this model is not only compatible with classical dynamic programming techniques, but also admits a novel greedy top-down inference algorithm based on recursive partitioning of the input.We demonstrate empirically that both prediction schemes are competitive with recent work, and when combined with basic extensions to the scoring model are capable of achieving state-of-the-art single-model performance on the Penn Treebank (91.79 F1) and strong performance on the French Treebank (82.23 F1). Mitchell Stern, Jacob Andreas, Daniel Klein 0001 |
ACL (1) | 2 |
| 2017 | Modeling Relationships in Referential Expressions with Compositional Modular NetworksabstractPeople often refer to entities in an image in terms of their relationships with other entities. For example, the black cat sitting under the table refers to both a black cat entity and its relationship with another table entity. Understanding these relationships is essential for interpreting and grounding such natural language expressions. Most prior work focuses on either grounding entire referential expressions holistically to one region, or localizing relationships based on a fixed set of categories. In this paper we instead present a modular deep architecture capable of analyzing referential expressions into their component parts, identifying entities and relationships mentioned in the input expression and grounding them all in the scene. We call this approach Compositional Modular Networks (CMNs): a novel architecture that learns linguistic analysis and visual inference end-to-end. Our approach is built around two types of neural modules that inspect local regions and pairwise interactions between regions. We evaluate CMNs on multiple referential expression datasets, outperforming state-of-the-art approaches on all tasks. Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, Kate Saenko |
CVPR | 3 |
| 2017 | Analogs of Linguistic Structure in Deep RepresentationsabstractWe investigate the compositional structure of message vectors computed by a deep network trained on a communication game.By comparing truth-conditional representations of encoder-produced message vectors to human-produced referring expressions, we are able to identify aligned (vector, utterance) pairs with the same meaning.We then search for structured relationships among these aligned pairs to discover simple vector space transformations corresponding to negation, conjunction, and disjunction.Our results suggest that neural representations are capable of spontaneously developing a "syntax" with functional analogues to qualitative properties of natural language.1 Jacob Andreas, Daniel Klein 0001 |
EMNLP | 1 |
| 2017 | Learning to Reason: End-to-End Module Networks for Visual Question AnsweringabstractNatural language questions are inherently compositional, and many are most easily answered by reasoning about their decomposition into modular sub-problems. For example, to answer “is there an equal number of balls and boxes?” we can look for balls, look for boxes, count them, and compare the results. The recently proposed Neural Module Network (NMN) architecture [3, 2] implements this approach to question answering by parsing questions into linguistic substructures and assembling question-specific deep networks from smaller modules that each solve one subtask. However, existing NMN implementations rely on brittle off-the-shelf parsers, and are restricted to the module configurations proposed by these parsers rather than learning them from data. In this paper, we propose End-to-End Module Networks (N2NMNs), which learn to reason by directly predicting instance-specific network layouts without the aid of a parser. Our model learns to generate network structures (by imitating expert demonstrations) while simultaneously learning network parameters (using the downstream task loss). Experimental results on the new CLEVR dataset targeted at compositional question answering show that N2NMNs achieve an error reduction of nearly 50% relative to state-of-the-art attentional approaches, while discovering interpretable network architectures specialized for each question. Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Kate Saenko |
ICCV | 2 |
| 2017 | Modular Multitask Reinforcement Learning with Policy SketchesabstractWe describe a framework for multitask deep reinforcement learning guided by policy sketches. Sketches annotate tasks with sequences of named subtasks, providing information about high-level structural relationships among tasks but not how to implement them—specifically not providing the detailed guidance used by much previous work on learning policy abstractions for RL (e.g. intermediate rewards, subtask completion signals, or intrinsic motivations). To learn from sketches, we present a model that associates every subtask with a modular subpolicy, and jointly maximizes reward over full task-specific policies by tying parameters across shared subpolicies. Optimization is accomplished via a decoupled actor–critic training objective that facilitates learning common behaviors from multiple dissimilar reward functions. We evaluate the effectiveness of our approach in three environments featuring both discrete and continuous control, and with sparse rewards that can be obtained only after completing a number of high-level subgoals. Experiments show that using our approach to learn policies guided by sketches gives better performance than existing techniques for learning task-specific or shared policies, while naturally inducing a library of interpretable primitive behaviors that can be recombined to rapidly adapt to new tasks. Jacob Andreas, Daniel Klein 0001, Sergey Levine |
ICML | 1 |
| 2016 | Neural Module NetworksabstractVisual question answering is fundamentally compositional in nature-a question like where is the dog? shares substructure with questions like what color is the dog? and where is the cat? This paper seeks to simultaneously exploit the representational capacity of deep networks and the compositional linguistic structure of questions. We describe a procedure for constructing and learning neural module networks, which compose collections of jointly-trained neural "modules" into deep networks for question answering. Our approach decomposes questions into their linguistic substructures, and uses these structures to dynamically instantiate modular networks (with reusable components for recognizing dogs, classifying colors, etc.). The resulting compound networks are jointly trained. We evaluate our approach on two challenging datasets for visual question answering, achieving state-of-the-art results on both the VQA natural image dataset and a new dataset of complex questions about abstract shapes. Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Daniel Klein 0001 |
CVPR | 1 |
| 2016 | Reasoning about Pragmatics with Neural Listeners and SpeakersabstractWe present a model for contrastively describing scenes, in which context-specific behavior results from a combination of inferencedriven pragmatics and learned semantics.Like previous learned approaches to language generation, our model uses a simple featuredriven architecture (here a pair of neural "listener" and "speaker" models) to ground language in the world.Like inference-driven approaches to pragmatics, our model actively reasons about listener behavior when selecting utterances.For training, our approach requires only ordinary captions, annotated without demonstration of the pragmatic behavior the model ultimately exhibits.In human evaluations on a referring expression game, our approach succeeds 81% of the time, compared to 69% using existing techniques. Jacob Andreas, Daniel Klein 0001 |
EMNLP | 1 |
| 2016 | Learning to Compose Neural Networks for Question AnsweringabstractJacob Andreas, Marcus Rohrbach, Trevor Darrell, Dan Klein. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Daniel Klein 0001 |
HLT-NAACL | 1 |
| 2015 | Alignment-Based Compositional Semantics for Instruction FollowingabstractThis paper describes an alignment-based model for interpreting natural language instructions in context.We approach instruction following as a search over plans, scoring sequences of actions conditioned on structured observations of text and the environment.By explicitly modeling both the low-level compositional structure of individual actions and the high-level structure of full plans, we are able to learn both grounded representations of sentence meaning and pragmatic constraints on interpretation.To demonstrate the model's flexibility, we apply it to a diverse set of benchmark tasks.On every task, we outperform strong task-specific baselines, and achieve several new state-of-the-art results. Jacob Andreas, Daniel Klein 0001 |
EMNLP | 1 |
| 2015 | When and why are log-linear models self-normalizing?abstractSeveral techniques have recently been proposed for training "self-normalized" discriminative models.These attempt to find parameter settings for which unnormalized model scores approximate the true label probability.However, the theoretical properties of such techniques (and of self-normalization generally) have not been investigated.This paper examines the conditions under which we can expect self-normalization to work.We characterize a general class of distributions that admit self-normalization, and prove generalization bounds for procedures that minimize empirical normalizer variance.Motivated by these results, we describe a novel variant of an established procedure for training self-normalized models.The new procedure avoids computing normalizers for most training examples, and decreases training time by as much as factor of ten while preserving model quality. Jacob Andreas, Daniel Klein 0001 |
HLT-NAACL | 1 |
| 2015 | On the Accuracy of Self-Normalized Log-Linear ModelsabstractCalculation of the log-normalizer is a major computational obstacle in applications of log-linear models with large output spaces. The problem of fast normalizer computation has therefore attracted significant attention in the theoretical and applied machine learning literature. In this paper, we analyze a recently proposed technique known as ``self-normalization'', which introduces a regularization term in training to penalize log normalizers for deviating from zero. This makes it possible to use unnormalized model scores as approximate probabilities. Empirical evidence suggests that self-normalization is extremely effective, but a theoretical understanding of why it should work, and how generally it can be applied, is largely lacking.We prove upper bounds on the loss in accuracy due to self-normalization, describe classes of input distributionsthat self-normalize easily, and construct explicit examples of high-variance input distributions. Our theoretical results make predictions about the difficulty of fitting self-normalized models to several classes of distributions, and we conclude with empirical validation of these predictions on both real and synthetic datasets. Jacob Andreas, Maxim Rabinovich, Michael I. Jordan, Daniel Klein 0001 |
NIPS | 1 |
| 2014 | Grounding Language with Points and Paths in Continuous SpacesabstractWe present a model for generating pathvalued interpretations of natural language text.Our model encodes a map from natural language descriptions to paths, mediated by segmentation variables which break the language into a discrete set of events, and alignment variables which reorder those events.Within an event, lexical weights capture the contribution of each word to the aligned path segment.We demonstrate the applicability of our model on three diverse tasks: a new color description task, a new financial news task and an established direction-following task.On all three, the model outperforms strong baselines, and on a hard variant of the direction-following task it achieves results close to the state-of-the-art system described in Vogel and Jurafsky (2010). Jacob Andreas, Daniel Klein 0001 |
CoNLL | 1 |
| 2014 | Unsupervised Transcription of Piano Music
Taylor Berg-Kirkpatrick, Jacob Andreas, Daniel Klein 0001 |
NIPS | 2 |
| 2013 | Parsing Graphs with Hyperedge Replacement Grammars
David Chiang 0001, Jacob Andreas, Daniel Bauer 0002, Karl Moritz Hermann, Bevan K. Jones, Kevin Knight |
ACL (1) | 2 |
| 2012 | Semantics-Based Machine Translation with Hyperedge Replacement Grammars
Bevan K. Jones, Jacob Andreas, Daniel Bauer 0002, Karl Moritz Hermann, Kevin Knight |
COLING | 2 |
| 2012 | Annotating Agreement and Disagreement in Threaded Discussion
Jacob Andreas, Sara Rosenthal, Kathy McKeown |
LREC | 1 |
| 2010 | Towards Semi-Automated Annotation for Prepositional Phrase Attachment
Sara Rosenthal, William Lipovsky, Kathy McKeown, Kapil Thadani, Jacob Andreas |
LREC | 5 |