Jonathan D. Cohen 0003

dblp:31/5509-3 · also Jonathan Cohen 0003 · DBLP profile ↗
← Back
72ranked-venue papers
0as first author
27since 2021 · last 2025
0000-0003-2316-0763ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 64 · 24 since 2021Applied, interdisciplinary, general and emerging computing · 42 · 15 since 2021Systems, architecture and hardware · 3 · 1 since 2021Theory of computation · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 A Quantum Model of Arousal and Yerkes Dodson Law
Jonathan D. Cohen 0003, Jerome R. Busemeyer
CogSci2
2025 Understanding Task Representations in Neural Networks via Bayesian Ablation
Andrew Nam, Declan Campbell, Thomas L. Griffiths 0001, Jonathan D. Cohen 0003, Sarah-Jane Leslie
CogSci4
2025 AI-enhanced semantic feature norms for 786 concepts
Siddharth Suresh, Kushin Mukherjee, Tyler Giallanza, Mia Patil, Xizheng Yu, Jonathan D. Cohen 0003, Timothy T. Rogers
CogSci6
2025 Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models
abstract
Many recent studies have found evidence for emergent reasoning capabilities in large language models (LLMs), but debate persists concerning the robustness of these capabilities, and the extent to which they depend on structured reasoning mechanisms. To shed light on these issues, we study the internal mechanisms that support abstract reasoning in LLMs. We identify an emergent symbolic architecture that implements abstract reasoning via a series of three computations. In early layers, symbol abstraction heads convert input tokens to abstract variables based on the relations between those tokens. In intermediate layers, symbolic induction heads perform sequence induction over these abstract variables. Finally, in later layers, retrieval heads predict the next token by retrieving the value associated with the predicted abstract variable. These results point toward a resolution of the longstanding debate between symbolic and neural network approaches, suggesting that emergent reasoning in neural networks depends on the emergence of symbolic mechanisms.
Yukang Yang, Declan Campbell, Kaixuan Huang, Mengdi Wang 0001, Jonathan D. Cohen 0003, Taylor W. Webb
ICML5
2025 Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers
abstract
We present causal head gating (CHG), a scalable method for interpreting the functional roles of attention heads in transformer models. CHG learns soft gates over heads and assigns them a causal taxonomy—facilitating, interfering, or irrelevant—based on their impact on task performance. Unlike prior approaches in mechanistic interpretability, which are hypothesis-driven and require prompt templates or target labels, CHG applies directly to any dataset using standard next-token prediction. We evaluate CHG across multiple large language models (LLMs) in the Llama 3 model family and diverse tasks, including syntax, commonsense, and mathematical reasoning, and show that CHG scores yield causal, not merely correlational, insight validated via ablation and causal mediation analyses. We also introduce contrastive CHG, a variant that isolates sub-circuits for specific task components. Our findings reveal that LLMs contain multiple sparse task-sufficient sub-circuits, that individual head roles depend on interactions with others (low modularity), and that instruction following and in-context learning rely on separable mechanisms.
Andrew Nam, Henry Conklin, Yukang Yang, Thomas L. Griffiths 0001, Jonathan D. Cohen 0003, Sarah-Jane Leslie
NeurIPS5
2024 Anxiety symptoms of major depression associated with increased willingness to exert cognitive, but not physical effort
Laura Bustamante, Deanna M. Barch, Johanne Solis, Temitope Oshinowo, Ivan Grahek, Anna B. Konova, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci8
2024 A Relational Inductive Bias for Dimensional Abstraction in Neural Networks
Declan Campbell, Jonathan D. Cohen 0003
CogSci2
2024 Human-Like Geometric Abstraction in Large Pre-trained Neural Networks
Declan Campbell, Sreejan Kumar, Tyler Giallanza, Jonathan D. Cohen 0003, Thomas L. Griffiths 0001
CogSci4
2024 Learning expectations shape initial cognitive control allocation
Javier Alejandro Masís, Sebastian Musslick, Jonathan D. Cohen 0003
CogSci3
2024 Abstractors and relational cross-attention: An inductive bias for explicit relational reasoning in Transformers
abstract
An extension of Transformers is proposed that enables explicit relational reasoning through a novel module called the *Abstractor*. At the core of the Abstractor is a variant of attention called *relational cross-attention*. The approach is motivated by an architectural inductive bias for relational learning that disentangles relational information from object-level features. This enables explicit relational reasoning, supporting abstraction and generalization from limited data. The Abstractor is first evaluated on simple discriminative relational tasks and compared to existing relational architectures. Next, the Abstractor is evaluated on purely relational sequence-to-sequence tasks, where dramatic improvements are seen in sample efficiency compared to standard Transformers. Finally, Abstractors are evaluated on a collection of tasks based on mathematical problem solving, where consistent improvements in performance and sample efficiency are observed.
Awni Altabaa, Taylor W. Webb, Jonathan D. Cohen 0003, John D. Lafferty
ICLR3
2024 Slot Abstractors: Toward Scalable Abstract Visual Reasoning
abstract
Abstract visual reasoning is a characteristically human ability, allowing the identification of relational patterns that are abstracted away from object features, and the systematic generalization of those patterns to unseen problems. Recent work has demonstrated strong systematic generalization in visual reasoning tasks involving multi-object inputs, through the integration of slot-based methods used for extracting object-centric representations coupled with strong inductive biases for relational abstraction. However, this approach was limited to problems containing a single rule, and was not scalable to visual reasoning problems containing a large number of objects. Other recent work proposed Abstractors, an extension of Transformers that incorporates strong relational inductive biases, thereby inheriting the Transformer’s scalability and multi-head architecture, but it has yet to be demonstrated how this approach might be applied to multi-object visual inputs. Here we combine the strengths of the above approaches and propose Slot Abstractors, an approach to abstract visual reasoning that can be scaled to problems involving a large number of objects and multiple relations among them. The approach displays state-of-the-art performance across four abstract visual reasoning tasks, as well as an abstract reasoning task involving real-world images.
Shanka Subhra Mondal, Jonathan D. Cohen 0003, Taylor W. Webb
ICML2
2024 Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem
abstract
Recent work has documented striking heterogeneity in the performance of state-of-the-art vision language models (VLMs), including both multimodal language models and text-to-image models. These models are able to describe and generate a diverse array of complex, naturalistic images, yet they exhibit surprising failures on basic multi-object reasoning tasks -- such as counting, localization, and simple forms of visual analogy -- that humans perform with near perfect accuracy. To better understand this puzzling pattern of successes and failures, we turn to theoretical accounts of the binding problem in cognitive science and neuroscience, a fundamental problem that arises when a shared set of representational resources must be used to represent distinct entities (e.g., to represent multiple objects in an image), necessitating the use of serial processing to avoid interference. We find that many of the puzzling failures of state-of-the-art VLMs can be explained as arising due to the binding problem, and that these failure modes are strikingly similar to the limitations exhibited by rapid, feedforward processing in the human brain.
Declan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata, Kia Ghods, Amogh Joshi 0004, Alexander Ku, Steven Frankland, Thomas L. Griffiths 0001, Jonathan D. Cohen 0003, Taylor W. Webb
NeurIPS10
2024 Erratum: Multitasking Capacity: Hardness Results and Improved Constructions
abstract
Abstract. We correct an error in the appendix of [N. Alon et al., SIAM J. Discrete Math., 34 (2020), pp. 885–903] and prove that it is NP-hard to approximate the size of a maximum induced matching of a bipartite graph within any constant factor.
Noga Alon, Jonathan D. Cohen 0003, Thomas L. Griffiths 0001, Pasin Manurangsi, Daniel Reichman 0001, Igor Shinkar, Tal Wagner
SIAM J. Discret. Math.2
2023 When to choose: Information seeking in the speed-accuracy tradeoff
Javier Alejandro Masís, David Melnikoff, Lisa Feldman Barrett, Jonathan D. Cohen 0003
CogSci4
2023 Advances in the Study of Visual and Multisensory Objects
Aleksandra Mroczko-Wasowicz, Casey O'Callaghan, Jonathan D. Cohen 0003, Brian J. Scholl, Philip J. Kellman
CogSci3
2023 Continuous and Discrete Transitions during Task-Switching
Harrison Ritz, William Wolf, Jonathan D. Cohen 0003
CogSci3
2023 Beyond Transformers for Function Learning
Simon N. Segert, Jonathan D. Cohen 0003
CogSci2
2023 Learning to reason over visual objects
Shanka Subhra Mondal, Taylor W. Webb, Jonathan D. Cohen 0003
ICLR3
2023 Systematic Visual Reasoning through Object-Centric Relational Abstraction
abstract
Human visual reasoning is characterized by an ability to identify abstract patterns from only a small number of examples, and to systematically generalize those patterns to novel inputs. This capacity depends in large part on our ability to represent complex visual inputs in terms of both objects and relations. Recent work in computer vision has introduced models with the capacity to extract object-centric representations, leading to the ability to process multi-object visual inputs, but falling short of the systematic generalization displayed by human reasoning. Other recent models have employed inductive biases for relational abstraction to achieve systematic generalization of learned abstract rules, but have generally assumed the presence of object-focused inputs. Here, we combine these two approaches, introducing Object-Centric Relational Abstraction (OCRA), a model that extracts explicit representations of both objects and abstract relations, and achieves strong systematic generalization in tasks (including a novel dataset, CLEVR-ART, with greater visual complexity) involving complex visual displays.
Taylor W. Webb, Shanka Subhra Mondal, Jonathan D. Cohen 0003
NeurIPS3
2023 Disentangling Abstraction from Statistical Pattern Matching in Human and Machine Learning
abstract
The ability to acquire abstract knowledge is a hallmark of human intelligence and is believed by many to be one of the core differences between humans and neural network models. Agents can be endowed with an inductive bias towards abstraction through meta-learning, where they are trained on a distribution of tasks that share some abstract structure that can be learned and applied. However, because neural networks are hard to interpret, it can be difficult to tell whether agents have learned the underlying abstraction, or alternatively statistical patterns that are characteristic of that abstraction. In this work, we compare the performance of humans and agents in a meta-reinforcement learning paradigm in which tasks are generated from abstract rules. We define a novel methodology for building "task metamers" that closely match the statistics of the abstract tasks but use a different underlying generative process, and evaluate performance on both abstract and metamer tasks. We find that humans perform better at abstract tasks than metamer tasks whereas common neural network architectures typically perform worse on the abstract tasks than the matched metamers. This work provides a foundation for characterizing differences between humans and machine learning that can be used in future work towards developing machines with more human-like behavior.
Sreejan Kumar, Ishita Dasgupta 0001, Nathaniel D. Daw, Jonathan D. Cohen 0003, Thomas L. Griffiths 0001
PLoS Comput. Biol.4
2022 Distill: Domain-Specific Compilation for Cognitive Models
abstract
Computational models of cognition enable a better understanding of the human brain and behavior, psychiatric and neurological illnesses, clinical interventions to treat illnesses, and also offer a path towards human-like artificial intelligence. Cognitive models are also, however, laborious to develop, requiring composition of many types of computational tasks, and suffer from poor performance as they are generally designed using high-level languages like Python. In this work, we present Distill, a domain-specific compilation tool to accelerate cognitive models while continuing to offer cognitive scientists the ability to develop their models in flexible high-level languages. Distill uses domain-specific knowledge to compile Python-based cognitive models into LLVM IR, carefully stripping away features like dynamic typing and memory management that add performance overheads without being necessary for the underlying computation of the models. The net effect is an average of 27 × performance improvement in model execution over state-of-the-art techniques using Pyston and PyPy. Distill also repurposes classical compiler data flow analyses to reveal properties about data flow in cognitive models that are useful to cognitive scientists. Distill is publicly available, integrated in the PsyNeuLink cognitive modeling environment, and is already being used by researchers in the brain sciences.
Ján Veselý, Raghavendra Pradyumna Pothukuchi, Ketaki Joshi, Samyak Gupta, Jonathan D. Cohen 0003, Abhishek Bhattacharjee
CGO5
2022 Maximum Entropy Function Learning
Simon N. Segert, Jonathan D. Cohen 0003
CogSci2
2022 Using natural language and program abstractions to instill human inductive biases in machines
abstract
Strong inductive biases give humans the ability to quickly learn to perform a variety of tasks. Although meta-learning is a method to endow neural networks with useful inductive biases, agents trained by meta-learning may sometimes acquire very different strategies from humans. We show that co-training these agents on predicting representations from natural language task descriptions and programs induced to generate such tasks guides them toward more human-like inductive biases. Human-generated language descriptions and program induction models that add new learned primitives both contain abstract concepts that can compress description length. Co-training on these representations result in more human-like behavior in downstream meta-reinforcement learning agents than less abstract controls (synthetic language descriptions, program induction without learned primitives), suggesting that the abstraction supported by these representations is key.
Sreejan Kumar, Carlos G. Correa, Ishita Dasgupta 0001, Raja Marjieh, Michael Y. Hu, Robert D. Hawkins, Jonathan D. Cohen 0003, Nathaniel D. Daw, Karthik Narasimhan, Thomas L. Griffiths 0001
NeurIPS7
2021 Modelling the development of counting with memory-augmented neural networks
Zachary Dulberg, Taylor W. Webb, Jonathan D. Cohen 0003
CogSci3
2021 Regression, encoding, control: an integrated approach to shared representations with distributed coding
Gregory Henselman-Petrusek, Tyler Giallanza, Sebastian Musslick, Jonathan D. Cohen 0003
CogSci4
2021 Meta-Learning of Structured Task Distributions in Humans and Machines
Sreejan Kumar, Ishita Dasgupta 0001, Jonathan D. Cohen 0003, Nathaniel D. Daw, Thomas L. Griffiths 0001
ICLR3
2021 Emergent Symbols through Binding in External Memory
Taylor W. Webb, Ishan Sinha, Jonathan D. Cohen 0003
ICLR3
2020 People Do Not Just Plan, They Plan to Plan
Mark K. Ho, David Abel, Jonathan D. Cohen 0003, Michael L. Littman, Thomas L. Griffiths 0001
AAAI3
2020 Determinantal Point Processes for Memory and Structured Inference
Steven Frankland, Jonathan D. Cohen 0003
CogSci2
2020 Mental effort: One construct, many faces?
Sebastian Musslick, Maria Wirzberger, Ivan Grahek, Laura Bustamante, Amitai Shenhav, Jonathan D. Cohen 0003
CogSci6
2020 A Novel Quantum Approach to the Dynamics of Decision Making
Lena Rosendahl, Anastasia S. Bizyaeva, Jonathan D. Cohen 0003
CogSci3
2020 A memory-augmented neural network model of abstract sequential reasoning
Ishan Sinha, Jonathan D. Cohen 0003, Taylor W. Webb
CogSci2
2020 Learning Representations that Support Extrapolation
abstract
Extrapolation – the ability to make inferences that go beyond the scope of one’s experiences – is a hallmark of human intelligence. By contrast, the generalization exhibited by contemporary neural network algorithms is largely limited to interpolation between data points in their training corpora. In this paper, we consider the challenge of learning representations that support extrapolation. We introduce a novel visual analogy benchmark that allows the graded evaluation of extrapolation as a function of distance from the convex domain defined by the training data. We also introduce a simple technique, temporal context normalization, that encourages representations that emphasize the relations between objects. We find that this technique enables a significant improvement in the ability to extrapolate, considerably outperforming a number of competitive techniques.
Taylor W. Webb, Zachary Dulberg, Steven Frankland, Alexander A. Petrov, Randall C. O'Reilly, Jonathan D. Cohen 0003
ICML6
2020 Multitasking Capacity: Hardness Results and Improved Constructions
abstract
We consider the problem of determining the maximal $\alpha \in (0,1]$ such that every matching $M$ of size $k$ (or at most $k$) in a bipartite graph $G$ contains an induced matching of size at least $\alpha |M|$. This measure was recently introduced in [N. Alon et al., Adv. Neural Inf. Process. Syst., 2017, pp. 2097--2106] and is motivated by computational models in cognitive neuroscience as well as by modeling interference in radio and communication networks. We prove various hardness results for computing $\alpha$ either exactly or approximately. En route to our results, we also consider the maximum connected matching problem: determining the largest matching $N$ in a graph $G$ such that every two edges in $N$ are connected by an edge. We prove a nearly optimal $n^{1-\epsilon}$ hardness of approximation result (under randomized reductions) for connected matching in bipartite graphs (with both sides of cardinality $n$). Toward this end we define bipartite half-covers: a new combinatorial object that may be of independent interest. To our knowledge, the best previous hardness result for the maximum connected matching problem was that it is hard to approximate within some constant $\beta>1$. Finally, we demonstrate the existence of bipartite graphs with $n$ vertices on each side of average degree $d$, achieving $\alpha=1/2-\epsilon$ for matchings of size sufficiently smaller than $n/d$. This nearly matches the trivial upper bound of $1/2$ on $\alpha$ which holds for any graph containing a path of length 3.
Noga Alon, Jonathan D. Cohen 0003, Thomas L. Griffiths 0001, Pasin Manurangsi, Daniel Reichman 0001, Igor Shinkar, Tal Wagner, Alexander Y. Ku
SIAM J. Discret. Math.2
2019 Extracting and Utilizing Abstract, Structured Representations for Analogy
Steven Frankland, Taylor W. Webb, Alexander A. Petrov, Randall C. O'Reilly, Jonathan D. Cohen 0003
CogSci5
2019 Stability-Flexibility Dilemma in Cognitive Control: A Dynamical System Perspective
Sebastian Musslick, Anastasia S. Bizyaeva, Shamay Agaron, Naomi Ehrich Leonard, Jonathan D. Cohen 0003
CogSci5
2019 A Mechanistic Account of Constraints on Control-Dependent Processing: Shared Representation, Conflict and Persistence
Sebastian Musslick, Jonathan D. Cohen 0003
CogSci2
2019 Decomposing Individual Differences in Cognitive Control: A Model-Based Approach
Sebastian Musslick, Jonathan D. Cohen 0003, Amitai Shenhav
CogSci2
2019 Understanding interactions amongst cognitive control, learning and representation
Sebastian Musslick, Abigail Novick Hoskin, Taylor W. Webb, Steven Frankland, Jonathan D. Cohen 0003, Rebecca L. Jackson, Matthew A. Lambon Ralph, Lang Chen, Timothy T. Rogers, Randall C. O'Reilly, Alexander A. Petrov
CogSci5
2019 Asymmetric Switch Costs as a Function of Task Strength
Markus Spitzer 0002, Sebastian Musslick, Michael Shvartsman, Amitai Shenhav, Jonathan D. Cohen 0003
CogSci5
2019 A tradeoff between generalization and perceptual capacity in recurrent neural networks
Taylor W. Webb, Steven Frankland, Simon N. Segert, Alexander A. Petrov, Randall C. O'Reilly, Jonathan D. Cohen 0003
CogSci6
2018 Matrix-normal models for fMRI analysis
abstract
Multivariate analysis of fMRI data has bene- fited substantially from advances in machine learning. Most recently, a range of prob- abilistic latent variable models applied to fMRI data have been successful in a variety of tasks, including identifying similarity pat- terns in neural data, combining multi-subject datasets, and mapping between brain and be- havior. Although these methods share some underpinnings, they have been developed as distinct methods, with distinct algorithms and software tools. We show how the matrix- variate normal (MN) formalism can unify some of these methods into a single frame- work. In doing so, we gain the ability to reuse noise modeling assumptions, algorithms, and code across models. Our primary theoretical contribution shows how some of these meth- ods can be written as instantiations of the same model, allowing us to generalize them to flexibly modeling structured noise covari- ances. Our formalism permits novel model variants and improved estimation strategies for SRM and RSA using substantially fewer parameters. We empirically demonstrate ad- vantages of our two new methods: for MN-RSA, we show up to 10x improvement in run- time, up to 6x improvement in RMSE, and more conservative behavior under the null. For MN-SRM, our method grants a modest improvement to out-of-sample reconstruction while relaxing the orthonormality constraint of SRM. We also provide a software prototyp- ing tool for MN models that can flexibly reuse noise covariance assumptions and algorithms across models.
Michael Shvartsman, Narayanan Sundaram, Mikio C. Aoi, Adam Charles, Theodore L. Willke, Jonathan D. Cohen 0003
AISTATS6
2018 Novel methods for measuring the cost of cognitive control in a patch foraging task and a demand selection task with Stroop
Laura Bustamante, Augustus Baker, Allison Burton, Amitai Shenhav, Chloe Hoeber, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci7
2018 Feature Ratings and Empirical Dimension-Specific Similarity Explain Distinct Aspects of Semantic Similarity Judgments
Marius Catalin Iordan, Cameron T. Ellis, Michael Lesnick, Daniel N. Osherson, Jonathan D. Cohen 0003
CogSci5
2018 Estimating the costs of cognitive control from task performance: theoretical validation and potential pitfalls
Sebastian Musslick, Jonathan D. Cohen 0003, Amitai Shenhav
CogSci2
2018 Constraints associated with cognitive control and the stability-flexibility dilemma
Sebastian Musslick, Seong Jun Jang, Michael Shvartsman, Amitai Shenhav, Jonathan D. Cohen 0003
CogSci5
2018 Efficiency of learning vs. processing: Towards a normative theory of multitasking
Yotam Sagiv, Sebastian Musslick, Yael Niv, Jonathan D. Cohen 0003
CogSci4
2017 Mechanisms of overharvesting in patch foraging
Gary Kane, Aaron M. Bornstein, Amitai Shenhav, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci6
2017 Adaptive response priors in context-dependent decision-making
Olga Lositsky, Michael Shvartsman, Robert C. Wilson, Jonathan D. Cohen 0003
CogSci4
2017 Multitasking Capability Versus Learning Efficiency in Neural Network Architectures
Sebastian Musslick, Andrew Saxe, Kayhan Özcimder, Biswadip Dey, Greg Henselman, Jonathan D. Cohen 0003
CogSci6
2017 A Formal Approach to Modeling the Cost of Cognitive Control
Kayhan Özcimder, Biswadip Dey, Sebastian Musslick, Giovanni Petri, Nesreen K. Ahmed, Theodore L. Willke, Jonathan D. Cohen 0003
CogSci7
2017 A graph-theoretic approach to multitasking
abstract
A key feature of neural network architectures is their ability to support the simultaneous interaction among large numbers of units in the learning and processing of representations. However, how the richness of such interactions trades off against the ability of a network to simultaneously carry out multiple independent processes -- a salient limitation in many domains of human cognition -- remains largely unexplored. In this paper we use a graph-theoretic analysis of network architecture to address this question, where tasks are represented as edges in a bipartite graph $G=(A \cup B, E)$. We define a new measure of multitasking capacity of such networks, based on the assumptions that tasks that \emph{need} to be multitasked rely on independent resources, i.e., form a matching, and that tasks \emph{can} be performed without interference if they form an induced matching. Our main result is an inherent tradeoff between the multitasking capacity and the average degree of the network that holds \emph{regardless of the network architecture}. These results are also extended to networks of depth greater than $2$. On the positive side, we demonstrate that networks that are random-like (e.g., locally sparse) can have desirable multitasking properties. Our results shed light into the parallel-processing limitations of neural systems and provide insights that may be useful for the analysis and design of parallel architectures.
Noga Alon, Daniel Reichman 0001, Igor Shinkar, Tal Wagner, Sebastian Musslick, Jonathan D. Cohen 0003, Thomas L. Griffiths 0001, Biswadip Dey, Kayhan Özcimder
NIPS6
2017 Noise correlations in the human brain and their impact on pattern classification
abstract
Multivariate decoding methods, such as multivoxel pattern analysis (MVPA), are highly effective at extracting information from brain imaging data. Yet, the precise nature of the information that MVPA draws upon remains controversial. Most current theories emphasize the enhanced sensitivity imparted by aggregating across voxels that have mixed and weak selectivity. However, beyond the selectivity of individual voxels, neural variability is correlated across voxels, and such noise correlations may contribute importantly to accurate decoding. Indeed, a recent computational theory proposed that noise correlations enhance multivariate decoding from heterogeneous neural populations. Here we extend this theory from the scale of neurons to functional magnetic resonance imaging (fMRI) and show that noise correlations between heterogeneous populations of voxels (i.e., voxels selective for different stimulus variables) contribute to the success of MVPA. Specifically, decoding performance is enhanced when voxels with high vs. low noise correlations (measured during rest or in the background of the task) are selected during classifier training. Conversely, voxels that are strongly selective for one class in a GLM or that receive high classification weights in MVPA tend to exhibit high noise correlations with voxels selective for the other class being discriminated against. Furthermore, we use simulations to show that this is a general property of fMRI data and that selectivity and noise correlations can have distinguishable influences on decoding. Taken together, our findings demonstrate that if there is signal in the data, the resulting above-chance classification accuracy is modulated by the magnitude of noise correlations.
Vikranth R. Bejjanki, Rava Azeredo da Silveira, Jonathan D. Cohen 0003, Nicholas B. Turk-Browne
PLoS Comput. Biol.3
2016 Real-time full correlation matrix analysis of fMRI data
abstract
Real-time functional magnetic resonance imaging (rtfMRI) is an emerging approach for studying the functioning of the human brain. Computational challenges combined with high data velocity have to this point restricted rtfMRI analyses to studying regions of the brain independently. However, given that neural processing is accomplished via functional interactions among brain regions, neuroscience could stand to benefit from rtfMRI analyses of full-brain interactions. In this paper, we extend such an offline analysis method, full correlation matrix analysis (FCMA), to enable its use in rtfMRI studies. Specifically, we introduce algorithms capable of processing real-time data for all stages of the FCMA machine learning workflow: incremental feature selection, model updating, and real-time classification. We also present an actor-model based distributed system designed to support FCMA and other rtfMRI analysis methods. Experiments show that our system successfully analyzes a stream of brain volumes and returns neurofeedback with less than 180 ms of lag. Our real-time FCMA implementation provides the same accuracy as an optimized offline FCMA toolbox while running 3.6–6.2x faster.
Yida Wang 0003, Bryn Keller, Mihai Capota, Michael J. Anderson, Narayanan Sundaram, Jonathan D. Cohen 0003, Kai Li 0001, Nicholas B. Turk-Browne, Theodore L. Willke
IEEE BigData6
2016 Boredom, Information-Seeking and Exploration
Andra Geana, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci4
2016 Information-Seeking, Learning and the Marginal Value Theorem: A Normative Approach to Adaptive Exploration
Andra Geana, Nathaniel D. Daw, Jonathan D. Cohen 0003
CogSci4
2016 Controlled vs. Automatic Processing: A Graph-Theoretic Approach to the Analysis of Serial vs. Parallel Processing in Neural Network Architectures
Sebastian Musslick, Biswadip Dey, Kayhan Özcimder, Md. Mostofa Ali Patwary, Theodore L. Willke, Jonathan D. Cohen 0003
CogSci6
2015 A Theory of Decision Making Under Dynamic Context
abstract
The dynamics of simple decisions are well understood and modeled as a class of random walk models (e.g. Laming, 1968; Ratcliff, 1978; Busemeyer and Townsend, 1993; Usher and McClelland, 2001; Bogacz et al., 2006). However, most real-life decisions include a rich and dynamically-changing influence of additional information we call context. In this work, we describe a computational theory of decision making under dynamically shifting context. We show how the model generalizes the dominant existing model of fixed-context decision making (Ratcliff, 1978) and can be built up from a weighted combination of fixed-context decisions evolving simultaneously. We also show how the model generalizes re- cent work on the control of attention in the Flanker task (Yu et al., 2009). Finally, we show how the model recovers qualitative data patterns in another task of longstanding psychological interest, the AX Continuous Performance Test (Servan-Schreiber et al., 1996), using the same model parameters.
Michael Shvartsman, Vaibhav Srivastava, Jonathan D. Cohen 0003
NIPS3
2015 Full correlation matrix analysis of fMRI data on Intel® Xeon Phi™ coprocessors
abstract
Full correlation matrix analysis (FCMA) is an unbiased approach for exhaustively studying interactions among brain regions in functional magnetic resonance imaging (fMRI) data from human participants. In order to answer neuroscientific questions efficiently, we are developing a closed-loop analysis system with FCMA on a cluster of nodes with Intel® Xeon Phi™ coprocessors. Here we propose several ideas for data-driven algorithmic modification to improve the performance on the coprocessor. Our experiments with real datasets show that the optimized single-node code runs 5x-16x faster than the baseline implementation using the well-known Intel® MKL and LibSVM libraries, and that the cluster implementation achieves near linear speedup on 5760 cores.
Yida Wang 0003, Michael J. Anderson, Jonathan D. Cohen 0003, Alexander Heinecke, Kai Li 0001, Nadathur Satish, Narayanan Sundaram, Nicholas B. Turk-Browne, Theodore L. Willke
SC3
2012 A Decision Task in a Social Context: Human Experiments, Models, and Analyses of Behavioral Data
abstract
To investigate the influence of information about fellow group members in a constrained decision-making context, we develop four two-armed bandit tasks in which subjects freely select one of two options ($A$or$B$) and are informed of the resulting reward following each choice. Rewards are determined by the fraction$x$of past$A$choices by two functions$f_{A}(x),f_{B}(x)$(unknown to the subject) which intersect at a matching point$\bar{x}$that does not generally represent globally optimal behavior. Playing individually, subjects typically remain close to the matching point, although some discover the optimum. Each task is designed to probe a different type of behavior, and subjects work in parallel in groups of five with feedback of other group members' choices, of their rewards, of both, or with no knowledge of others' behavior. We employ a soft-max choice model that emerges from a drift-diffusion process, commonly used to model perceptual decision making with noisy stimuli. Here the stimuli are replaced by estimates of expected rewards produced by a temporal-difference reinforcement-learning algorithm, augmented to include appropriate feedback terms. Models are fitted for each task and feedback condition, and we compare choice allocations averaged across subjects and individual choice sequences to highlight differences between tasks and intersubject differences. The most complex model, involving both choice and reward feedback, contains only four parameters, but nonetheless reveals significant differences in individual strategies. Strikingly, we find that rewards feedback can be either detrimental or advantageous to performance, depending upon the task.
Andrea Nedic, Damon Tomlin, Philip Holmes, Deborah A. Prentice, Jonathan D. Cohen 0003
Proc. IEEE5
2011 Finding neural correlates of drift diffusion processes in EEG oscillations
Marieke K. van Vugt, Patrick Simen, Jonathan D. Cohen 0003
CogSci3
2009 Sequential Effects in Two-Choice Reaction Time Tasks: Decomposition and Synthesis of Mechanisms
abstract
Performance on serial tasks is influenced by first- and higher-order sequential effects, respectively, due to the immediately previous and earlier trials. As response-to-stimulus interval (RSI) increases, the pattern of reaction times transits from a benefit-only mode, traditionally ascribed to automatic facilitation (AF), to a cost-benefit mode, due to strategic expectancy (SE). To illuminate the sources of such effects, we develop a connectionist network of two mutually inhibiting neural decision units subject to feedback from previous trials. A study of separate biasing mechanisms shows that residual decision unit activity can lead to only first-order AF, but higher-order AF can result from strategic priming mediated by conflict monitoring, which we instantiate in two distinct versions. A further mechanism mediates expectation-related biases that grow during RSI toward saturation levels determined by weighted repetition (or alternation) sequence lengths. Equipped with these mechanisms, the network, consistent with known neurophysiology, accounts for several sets of behavioral data over a wide range of RSIs. The results also suggest that practice speeds up all the mechanisms rather than adjusting their relative strengths.
Juan Gao, KongFatt Wong-Lin, Philip Holmes, Patrick Simen, Jonathan D. Cohen 0003
Neural Comput.5
2008 Learning to Use Working Memory in Partially Observable Environments through Dopaminergic Reinforcement
abstract
Working memory is a central topic of cognitive neuroscience because it is critical for solving real world problems in which information from multiple temporally distant sources must be combined to generate appropriate behavior. However, an often neglected fact is that learning to use working memory effectively is itself a difficult problem. The Gating" framework is a collection of psychological models that show how dopamine can train the basal ganglia and prefrontal cortex to form useful working memory representations in certain types of problems. We bring together gating with ideas from machine learning about using finite memory systems in more general problems. Thus we present a normative Gating model that learns, by online temporal difference methods, to use working memory to maximize discounted future rewards in general partially observable settings. The model successfully solves a benchmark working memory problem, and exhibits limitations similar to those observed in human experiments. Moreover, the model introduces a concise, normative definition of high level cognitive concepts such as working memory and cognitive control in terms of maximizing discounted future rewards."
Michael T. Todd, Yael Niv, Jonathan D. Cohen 0003
NIPS3
2008 Sequential effects: Superstition or rational behavior?
abstract
In a variety of behavioral tasks, subjects exhibit an automatic and apparently sub-optimal sequential effect: they respond more rapidly and accurately to a stimulus if it reinforces a local pattern in stimulus history, such as a string of repetitions or alternations, compared to when it violates such a pattern. This is often the case even if the local trends arise by chance in the context of a randomized design, such that stimulus history has no predictive power. In this work, we use a normative Bayesian framework to examine the hypothesis that such idiosyncrasies may reflect the inadvertent engagement of fundamental mechanisms critical for adapting to changing statistics in the natural environment. We show that prior belief in non-stationarity can induce experimentally observed sequential effects in an otherwise Bayes-optimal algorithm. The Bayesian algorithm is shown to be well approximated by linear-exponential filtering of past observations, a feature also apparent in the behavioral data. We derive an explicit relationship between the parameters and computations of the exact Bayesian algorithm and those of the approximate linear-exponential filter. Since the latter is equivalent to a leaky-integration process, a commonly used model of neuronal dynamics underlying perceptual decision-making and trial-to-trial dependencies, our model provides a principled account of why such dynamics are useful. We also show that near-optimal tuning of the leaky-integration process is possible, using stochastic gradient descent based only on the noisy binary inputs. This is a proof of concept that not only can neurons implement near-optimal prediction based on standard neuronal dynamics, but that they can also learn to tune the processing parameters without explicitly representing probabilities.
Angela J. Yu, Jonathan D. Cohen 0003
NIPS2
2008 A Neural Network Model of the Eriksen Task: Reduction, Analysis, and Data Fitting
abstract
We analyze a neural network model of the Eriksen task: a two-alternative forced-choice task in which subjects must correctly identify a central stimulus and disregard flankers that may or may not be compatible with it. We linearize and decouple the model, deriving a reduced drift-diffusion process with variable drift rate that describes the accumulation of net evidence in favor of either alternative, and we use this to analytically describe how accuracy and response time data depend on model parameters. Such analyses both assist parameter tuning in network models and suggest explanations of changing drift rates in terms of attention. We compare our results with numerical simulations of the full nonlinear model and with empirical data and show good fits to both with fewer parameters.
Yuan Sophie Liu, Philip Holmes, Jonathan D. Cohen 0003
Neural Comput.3
2008 Optimization of Decision Making in Multilayer Networks: The Role of Locus Coeruleus
abstract
Previous theoretical work has shown that a single-layer neural network can implement the optimal decision process for simple, two-alternative forced-choice (2AFC) tasks. However, it is likely that the mammalian brain comprises multilayer networks, raising the question of whether and how optimal performance can be approximated in such an architecture. Here, we present theoretical work suggesting that the noradrenergic nucleus locus coeruleus (LC) may help optimize 2AFC decision making in the brain. This is based on the observations that neurons of the LC selectively fire following the presentation of salient stimuli in decision tasks and that the corresponding release of norepinephrine can transiently increase the responsivity, or gain, of cortical processing units. We describe computational simulations that investigate the role of such gain changes in optimizing performance of 2AFC decision making. In the tasks we model, no prior cueing or knowledge of stimulus onset time is assumed. Performance is assessed in terms of the rate of correct responses over time (the reward rate). We first present the results of a single-layer model that accumulates (integrates) sensory input and implements the decision process as a threshold crossing. Gain transients, representing the modulatory effect of the LC, are driven by separate threshold crossings in this layer. We optimize over all free parameters to determine the maximum reward rate achievable by this model and compare it to the maximum reward rate when gain is held fixed. We find that the dynamic gain mechanism yields no improvement in reward for this single-layer model. We then examine a two-layer model, in which competing sensory accumulators in the first layer (capable of implementing the task relevant decision) pass activity to response accumulators in a second layer. Again, we compare a version in which threshold crossing in the first (decision) layer elicits an LC response (and a concomitant increase in gain) with a fixed-gain version of the model. Here, we find that gain transients modeling the LC phasic response yield an improvement in reward rate of 12% to 24%. Furthermore, we show that the timing characteristics of these gain transients agree with observations concerning LC firing patterns reported in recent experimental studies. This provides converging evidence for the hypothesis that the LC optimizes processes underlying 2AFC decision making in multilayer networks.
Eric Shea-Brown, Mark S. Gilzenrat, Jonathan D. Cohen 0003
Neural Comput.3
2006 Rapid decision threshold modulation by reward rate in a neural network
Patrick Simen, Jonathan D. Cohen 0003, Philip Holmes
Neural Networks2
2005 An exploration-exploitation model based on norepinepherine and dopamine activity
abstract
We propose a model by which dopamine (DA) and norepinepherine (NE) combine to alternate behavior between relatively exploratory and exploitative modes. The model is developed for a target detection task for which there is extant single neuron recording data available from locus coeruleus (LC) NE neurons. An exploration-exploitation trade-off is elicited by regularly switching which of the two stimuli are rewarded. DA functions within the model to change synaptic weights according to a reinforcement learning algorithm. Exploration is mediated by the state of LC firing, with higher tonic and lower phasic activity producing greater response variability. The opposite state of LC function, with lower baseline firing rate and greater phasic responses, favors exploitative behavior. Changes in LC firing mode result from combined measures of response conflict and reward rate, where response conflict is monitored using models of anterior cingulate cortex (ACC). Increased long-term response conflict and decreased reward rate, which occurs following reward contingency switch, favors the higher tonic state of LC function and NE release. This increases exploration, and facilitates discovery of the new target.
Samuel M. McClure, Mark S. Gilzenrat, Jonathan D. Cohen 0003
NIPS3
2002 Simplified dynamics in a model of noradrenergic modulation of cognitive performance
Mark S. Gilzenrat, Benjamin D. Holmes, Janusz Rajkowski, Gary Aston-Jones, Jonathan D. Cohen 0003
Neural Networks5
1997 Online Analysis of Functional MRI Datasets on Parallel Platforms
Nigel H. Goddard, Greg Hood, Jonathan D. Cohen 0003, William F. Eddy, Christopher R. Genovese, Douglas C. Noll, Leigh E. Nystrom
J. Supercomput.3
1994 A Computational Model of Prefrontal Cortex Function
abstract
Accumulating data from neurophysiology and neuropsychology have suggested two information processing roles for prefrontal cor(cid:173) tex (PFC): 1) short-term active memory; and 2) inhibition. We present a new behavioral task and a computational model which were developed in parallel. The task was developed to probe both of these prefrontal functions simultaneously, and produces a rich set of behavioral data that act as constraints on the model. The model is implemented in continuous-time, thus providing a natural framework in which to study the temporal dynamics of processing in the task. We show how the model can be used to examine the be(cid:173) havioral consequences of neuromodulation in PFC . Specifically, we use the model to make novel and testable predictions regarding the behavioral performance of schizophrenics, who are hypothesized to suffer from reduced dopaminergic tone in this brain area.
Todd S. Braver, Jonathan D. Cohen 0003, David Servan-Schreiber
NIPS2
1989 The Effect of Catecholamines on Performance: From Unit to System Behavior
David Servan-Schreiber, Harry Printz, Jonathan D. Cohen 0003
NIPS3