James L. McClelland

dblp:49/5831 · also Jay McClelland · DBLP profile ↗
← Back
41ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-8217-405XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021
YearPublicationVenuePosition
2025 Humans learn proactively in ways that language models don't
Simon Jerome Han, James L. McClelland
CogSci2
2025 Extending a Mathematical Theory of the Emergence of Knowledge from the Experience to Capture Learning Dynamics in Transformers
Sabrina Jones, James L. McClelland
CogSci2
2024 Symbolic Variables in Distributed Networks that Count
Satchel Grant, Zhengxuan Wu, James L. McClelland, Noah D. Goodman
CogSci3
2024 SODA: Bottleneck Diffusion Models for Representation Learning
abstract
We introduce SODA, a self-supervised diffusion model, designed for representation learning. The model incorpo-rates an image encoder, which distills a source view into a compact representation, that, in turn, guides the generation of related novel views. We show that by imposing a tight bottleneck between the encoder and a denoising decoder, and leveraging novel view synthesis as a self-supervised ob-jective, we can turn diffusion models into strong represen-tation learners, capable of capturing visual semantics in an unsupervised manner. To the best of our knowledge, SODA is the first diffusion model to succeed at ImageNet linear-probe classification, and, at the same time, it accomplishes reconstruction, editing and synthesis tasks across a wide range of datasets. Further investigation reveals the disentangled nature of its emergent latent space, that serves as an effective interface to control and manipulate the produced images. All in all, we aim to shed light on the exciting and promising potential of diffusion models, not only for image generation, but also for learning rich and robust represen-tations. See our website at soda-diffusion.github.io.
Drew A. Hudson, Daniel Zoran, Mateusz Malinowski, Andrew K. Lampinen, Andrew Jaegle, James L. McClelland, Loïc Matthey, Felix Hill, Alexander Lerchner
CVPR6
2022 Tell me why! Explanations support learning relational and causal structure
abstract
Inferring the abstract relational and causal structure of the world is a major challenge for reinforcement-learning (RL) agents. For humans, language{—}particularly in the form of explanations{—}plays a considerable role in overcoming this challenge. Here, we show that language can play a similar role for deep RL agents in complex environments. While agents typically struggle to acquire relational and causal knowledge, augmenting their experience by training them to predict language descriptions and explanations can overcome these limitations. We show that language can help agents learn challenging relational tasks, and examine which aspects of language contribute to its benefits. We then show that explanations can help agents to infer not only relational but also causal structure. Language can shape the way that agents to generalize out-of-distribution from ambiguous, causally-confounded training, and explanations even allow agents to learn to perform experimental interventions to identify causal relationships. Our results suggest that language description and explanation may be powerful tools for improving agent learning and generalization.
Andrew K. Lampinen, Nicholas A. Roy, Ishita Dasgupta 0001, Stephanie C. Y. Chan, Allison C. Tam, James L. McClelland, Adam Santoro, Neil C. Rabinowitz, Jane X. Wang, Felix Hill
ICML6
2022 Data Distributional Properties Drive Emergent In-Context Learning in Transformers
abstract
Large transformer-based models are able to perform in-context few-shot learning, without being explicitly trained for it. This observation raises the question: what aspects of the training regime lead to this emergent behavior? Here, we show that this behavior is driven by the distributions of the training data itself. In-context learning emerges when the training data exhibits particular distributional properties such as burstiness (items appear in clusters rather than being uniformly distributed over time) and having a large number of rarely occurring classes. In-context learning also emerges more strongly when item meanings or interpretations are dynamic rather than fixed. These properties are exemplified by natural language, but are also inherent to naturalistic data in a wide range of other domains. They also depart significantly from the uniform, i.i.d. training distributions typically used for standard supervised learning. In our initial experiments, we found that in-context learning traded off against more conventional weight-based learning, and models were unable to achieve both simultaneously. However, our later experiments uncovered that the two modes of learning could co-exist in a single model when it was trained on data following a skewed Zipfian distribution -- another common property of naturalistic data, including language. In further experiments, we found that naturalistic data distributions were only able to elicit in-context learning in transformers, and not in recurrent models. Our findings indicate how the transformer architecture works together with particular properties of the training data to drive the intriguing emergent in-context learning behaviour of large language models, and indicate how future work might encourage both in-context and in-weights learning in domains beyond language.
Stephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang, Aaditya K. Singh, Pierre H. Richemond, James L. McClelland, Felix Hill
NeurIPS7
2022 A weighted constraint satisfaction approach to human goal-directed decision making
James L. McClelland
PLoS Comput. Biol.2
2021 Are people still smarter than machines? If so, why?
James L. McClelland
CogSci1
2020 Generative Continual Concept Learning
abstract
After learning a concept, humans are also able to continually generalize their learned concepts to new domains by observing only a few labeled instances without any interference with the past learned knowledge. In contrast, learning concepts efficiently in a continual learning setting remains an open challenge for current Artificial Intelligence algorithms as persistent model retraining is necessary. Inspired by the Parallel Distributed Processing learning and the Complementary Learning Systems theories, we develop a computational model that is able to expand its previously learned concepts efficiently to new domains using a few labeled samples. We couple the new form of a concept to its past learned forms in an embedding space for effective continual learning. Doing so, a generative distribution is learned such that it is shared across the tasks in the embedding space and models the abstract concepts. This procedure enables the model to generate pseudo-data points to replay the past experience to tackle catastrophic forgetting.
Soheil Kolouri, Praveen K. Pilly, James L. McClelland
AAAI4
2020 Cognitive consequences of structured education in a connectionist model of analogical reasoning
David G. T. Barrett, Felix Hill, Adam Santoro, James L. McClelland
CogSci4
2020 A computational model of learning to count in a multimodal, interactive environment
Silvester Sabathiel, James L. McClelland, Trygve Solstad
CogSci2
2020 Human-like learning Framework for frequency-skewed multi-level classification
Amarjot Singh, James L. McClelland
CogSci2
2020 Environmental drivers of systematicity and generalization in a situated agent
Felix Hill, Andrew K. Lampinen, Rosália G. Schneider, Stephen Clark, Matt M. Botvinick, James L. McClelland, Adam Santoro
ICLR6
2019 Symposium in Memory of Jeff Elman: Language Learning, Prediction, and Temporal Dynamics
James L. McClelland, Ken McRae
CogSci1
2019 Modeling Number Sense Acquisition in A Number Board Game by Coordinating Verbal, Visual, and Grounded Action Components
Arianna Yuan, James L. McClelland
CogSci2
2018 Can Generic Neural Networks Estimate Numerosity Like Humans?
Sharon Chen, Zhenglong Zhou, Mengting Fang, James L. McClelland
CogSci4
2018 Can a Recurrent Neural Network Learn to Count Things?
Mengting Fang, Zhenglong Zhou, Sharon Chen, James L. McClelland
CogSci4
2017 Geometric Concept Acquisition in a Dueling Deep Q-Network
Alex Kuefler, Mykel J. Kochenderfer, James L. McClelland
CogSci3
2017 Analogies Emerge from Learning Dyamics in Neural Networks
Andrew K. Lampinen, Shaw Hsu, James L. McClelland
CogSci3
2017 Neural responses decrease while performance increases with practice: A neural network model
Milena Rabovsky, Steven Hansen 0001, James L. McClelland
CogSci3
2016 Tutorial Workshop on Contemporary Deep Neural Network Models
James L. McClelland, Steven Hansen 0001, Andrew M. Saxe
CogSci1
2016 N400 amplitudes reflect change in a probabilistic representation of meaning: Evidence from a connectionist model
Milena Rabovsky, Steven Hansen 0001, James L. McClelland
CogSci3
2016 Emergence of Euclidean geometrical intuitions in hierarchical generative models
Arianna Yuan, Te-Lin Wu, James L. McClelland
CogSci3
2015 Connecting learning, memory, and representation in math education
Martha W. Alibali, Chuck Kalish, Timothy T. Rogers, Christine M. Massey, Philip J. Kellman, Vladimir M. Sloutsky, James L. McClelland, Kevin W. Mickey
CogSci7
2014 Two Plus Three Is Five: Discovering Efficient Addition Strategies without Metacognition
Steven Hansen 0001, Cameron R. L. McKenzie, James L. McClelland
CogSci3
2014 A neural network model of learning mathematical equivalence
Kevin W. Mickey, James L. McClelland
CogSci2
2013 From symbols to analog magnitudes: A process model of fraction comparison, with fits to experimental data
Cameron R. L. McKenzie, James L. McClelland
CogSci2
2013 Running circles around symbol manipulation in trigonometry
Kevin W. Mickey, James L. McClelland
CogSci2
2013 Learning hierarchical categories in deep neural networks
Andrew M. Saxe, James L. McClelland, Surya Ganguli
CogSci2
2013 Progressive Development of the Number Sense in a Deep Neural Network
Will Y. Zou, James L. McClelland
CogSci2
2011 Estimating the strength of unlabeled information during semi-supervised learning
Brenden M. Lake, James L. McClelland
CogSci2
2011 A PDP model of the simultaneous perception of multiple objects
abstract
Illusory conjunctions in normal and simultanagnosic subjects are two instances where the visual features of multiple objects are incorrectly ‘bound’ together. A connectionist model explores how multiple objects could be perceived correctly in normal subjects given sufficient time, but could give rise to illusory conjunctions with damage or time pressure. In this model, perception of two objects benefits from lateral connections between hidden layers modelling aspects of the ventral and dorsal visual pathways. As with simultanagnosia, simulations of dorsal lesions impair multi-object recognition. In contrast, a large ventral lesion has minimal effect on dorsal functioning, akin to dissociations between simple object manipulation (retained in visual form agnosia and semantic dementia) and object discrimination (impaired in these disorders) [Hodges, J.R., Bozeat, S., Lambon Ralph, M.A., Patterson, K., and Spatt, J. (2000), ‘The Role of Conceptual Knowledge: Evidence from Semantic Dementia’, Brain, 123, 1913–1925; Milner, A.D., and Goodale, M.A. (2006), The Visual Brain in Action (2nd ed.), New York: Oxford]. It is hoped that the functioning of this model might suggest potential processes underlying dorsal and ventral contributions to the correct perception of multiple objects.
Cynthia M. Henderson, James L. McClelland
Connect. Sci.2
2007 Guest Editorial: Convergent Approaches to the Understanding of Autonomous Mental Development
abstract
The eight articles in this special issue focus on convergent approaches to the understanding of autonomous mental development. The goal is to promote the effort to build the necessary bridges between research areas to help foster the investigation of the emergence of intelligent, autonomous, and organized behaviour in children and robotic systems.
James L. McClelland, Kim Plunkett, Juyang Weng
IEEE Trans. Evol. Comput.1
2000 Normal and impaired processing in quasi-regular domains of language: the case of English past-tense verbs
Karalyn Patterson, Matthew A. Lambon Ralph, Helen Bird, John R. Hodges, James L. McClelland
INTERSPEECH5
1999 Information Factorization in Connectionist Models of Perception
Javier R. Movellan, James L. McClelland
NIPS2
1997 A Hippocampal Model of Recognition Memory
Randall C. O'Reilly, Kenneth A. Norman, James L. McClelland
NIPS3
1991 Graded State Machines: The Representation of Temporal Contingencies in Simple Recurrent Networks
David Servan-Schreiber, Axel Cleeremans, James L. McClelland
Mach. Learn.3
1990 Learning and Applying Contextual Constraints in Sentence Comprehension
Mark F. St. John, James L. McClelland
Artif. Intell.2
1989 Finite State Automata and Simple Recurrent Networks
abstract
We explore a network architecture introduced by Elman (1988) for predicting successive elements of a sequence. The network uses the pattern of activation over a set of hidden units from time-step t−1, together with element t, to predict element t + 1. When the network is trained with strings from a particular finite-state grammar, it can learn to be a perfect finite-state recognizer for the grammar. When the network has a minimal number of hidden units, patterns on the hidden units come to correspond to the nodes of the grammar, although this correspondence is not necessary for the network to act as a perfect finite-state recognizer. We explore the conditions under which the network can carry information about distant sequential contingencies across intervening elements. Such information is maintained with relative ease if it is relevant at each intermediate step; it tends to be lost when intervening elements do not depend on it. At first glance this may suggest that such networks are not relevant to natural language, in which dependencies may span indefinite distances. However, embeddings in natural language are not completely independent of earlier information. The final simulation shows that long distance sequential contingencies can be encoded by the network even if only subtle statistical properties of embedded strings depend on the early information.
Axel Cleeremans, David Servan-Schreiber, James L. McClelland
Neural Comput.3
1988 Learning Subsequential Structure in Simple Recurrent Networks
David Servan-Schreiber, Axel Cleeremans, James L. McClelland
NIPS3
1987 Learning Representations by Recirculation
Geoffrey E. Hinton, James L. McClelland
NIPS2