Robert A. Jacobs

dblp:32/638 · DBLP profile ↗
← Back
31ranked-venue papers
10as first author
1since 2021 · last 2023
0000-0001-6607-908XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 9 first-authorApplied, interdisciplinary, general and emerging computing · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Probabilistic and Bayesian machine learning · 22% Reinforcement learning · 17% Motion planning and robot control · 16%
Theoretical computer science
1 paper
Computational geometry · 100%

Topics — the 13 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control › robot learning
movement primitives
0.112008
Learning to Combine Motor Primitives Via Greedy Additive Regression · J. Mach. Learn. Res. 2008
Machine learning › Representation and self-supervised learning › computational neuroscience › neural coding
neural population coding
0.012011
Probabilistic Modeling of Dependencies Among Visual Short-Term Memory Representations · NIPS 2011
Computer vision › 3D vision
motion estimation
0.012002
Visual Development Aids the Acquisition of Motion Velocity Sensitivities · NIPS 2002
Computer vision › 3D vision › stereo vision
stereo matching
0.012001
Visual Development and the Acquisition of Binocular Disparity Sensitivities · ICML 2001
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.012002
Visual Development Aids the Acquisition of Motion Velocity Sensitivities · NIPS 2002
Machine learning › Efficient and distributed learning
divide-and-conquer learning
0.011993
Supervised Learning and Divide-and-Conquer: A Statistical Approach · ICML 1993
Machine learning › Learning theory
statistical learning theory
0.011993
Supervised Learning and Divide-and-Conquer: A Statistical Approach · ICML 1993
Machine learning › Learning paradigms
supervised learning
0.011993
Supervised Learning and Divide-and-Conquer: A Statistical Approach · ICML 1993
Bioinformatics and computational biology
computational neuroscience
0.012001
Visual Development and the Acquisition of Binocular Disparity Sensitivities · ICML 2001
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.011991
Hierarchies of Adaptive Experts · NIPS 1991
Machine learning › Deep learning architectures and training
mixture of experts
0.011991
Hierarchies of Adaptive Experts · NIPS 1991
Robotics › Motion planning and robot control › robot modeling
forward models
0.011989
Learning to Control an Unstable System with Forward Modeling · NIPS 1989
Machine learning › Reinforcement learning
model-based reinforcement learning
0.011989
Learning to Control an Unstable System with Forward Modeling · NIPS 1989

Methods — techniques the papers use, named apart from their topics

behavioral shaping · 0.1multivariate gaussian modeling · 0.1gaussian process kernel · 0.1motor primitives · 0.1greedy additive regression · 0.1multi-scale feature learning · 0.0coarse-to-fine training · 0.0statistical analysis · 0.0divide-and-conquer · 0.0competitive learning · 0.0
YearPublicationVenuePosition
2023 Rapid runtime learning by curating small datasets of high-quality items obtained from memory
abstract
We propose the "runtime learning" hypothesis which states that people quickly learn to perform unfamiliar tasks as the tasks arise by using task-relevant instances of concepts stored in memory during mental training. To make learning rapid, the hypothesis claims that only a few class instances are used, but these instances are especially valuable for training. The paper motivates the hypothesis by describing related ideas from the cognitive science and machine learning literatures. Using computer simulation, we show that deep neural networks (DNNs) can learn effectively from small, curated training sets, and that valuable training items tend to lie toward the centers of data item clusters in an abstract feature space. In a series of three behavioral experiments, we show that people can also learn effectively from small, curated training sets. Critically, we find that participant reaction times and fitted drift rates are best accounted for by the confidences of DNNs trained on small datasets of highly valuable items. We conclude that the runtime learning hypothesis is a novel conjecture about the relationship between learning and memory with the potential for explaining a wide variety of cognitive phenomena.
Joseph German, Guofeng Cui, Chenliang Xu, Robert A. Jacobs
PLoS Comput. Biol.4
2020 Modeling Human Cognitive Flexibility with Extemporaneous Networks
Joseph German, Robert A. Jacobs
CogSci2
2019 Human Visual Object Similarity Judgments are Viewpoint-Invariant and Part-Based as Revealed via Metric Learning
Joseph German, Robert A. Jacobs
CogSci2
2016 A 3D shape inference model matches human visual object similarity judgments better than deep convolutional neural networks
Goker Erdogan, Robert A. Jacobs
CogSci2
2015 From Sensory Signals to Modality-Independent Conceptual Representations: A Probabilistic Language of Thought Approach
abstract
People learn modality-independent, conceptual representations from modality-specific sensory signals. Here, we hypothesize that any system that accomplishes this feat will include three components: a representational language for characterizing modality-independent representations, a set of sensory-specific forward models for mapping from modality-independent representations to sensory signals, and an inference algorithm for inverting forward models-that is, an algorithm for using sensory signals to infer modality-independent representations. To evaluate this hypothesis, we instantiate it in the form of a computational model that learns object shape representations from visual and/or haptic signals. The model uses a probabilistic grammar to characterize modality-independent representations of object shape, uses a computer graphics toolkit and a human hand simulator to map from object representations to visual and haptic features, respectively, and uses a Bayesian inference algorithm to infer modality-independent object representations from visual and/or haptic signals. Simulation results show that the model infers identical object representations when an object is viewed, grasped, or both. That is, the model's percepts are modality invariant. We also report the results of an experiment in which different subjects rated the similarity of pairs of objects in different sensory conditions, and show that the model provides a very accurate account of subjects' ratings. Conceptually, this research significantly contributes to our understanding of modality invariance, an important type of perceptual constancy, by demonstrating how modality-independent representations can be acquired and used. Methodologically, it provides an important contribution to cognitive modeling, particularly an emerging probabilistic language-of-thought approach, by showing how symbolic and statistical approaches can be combined in order to understand aspects of human perception.
Goker Erdogan, Ilker Yildirim, Robert A. Jacobs
PLoS Comput. Biol.3
2014 Transfer of object shape knowledge across visual and haptic modalities
Goker Erdogan, Ilker Yildirim, Robert A. Jacobs
CogSci3
2011 A Nonparametric Bayesian Model of Visual Short-Term Memory
A. Emin Orhan, Robert A. Jacobs
CogSci2
2011 An Ideal Observer Model of Visual Short-Term Memory Predicts Human Capacity - Precision Tradeoffs
Chris R. Sims, Robert A. Jacobs, David C. Knill
CogSci2
2011 Probabilistic Modeling of Dependencies Among Visual Short-Term Memory Representations
abstract
Extensive evidence suggests that items are not encoded independently in visual short-term memory (VSTM). However, previous research has not quantitatively considered how the encoding of an item influences the encoding of other items. Here, we model the dependencies among VSTM representations using a multivariate Gaussian distribution with a stimulus-dependent mean and covariance matrix. We report the results of an experiment designed to determine the specific form of the stimulus-dependence of the mean and the covariance matrix. We find that the magnitude of the covariance between the representations of two items is a monotonically decreasing function of the difference between the items' feature values, similar to a Gaussian process with a distance-dependent, stationary kernel function. We further show that this type of covariance function can be explained as a natural consequence of encoding multiple stimuli in a population of neurons with correlated responses.
A. Emin Orhan, Robert A. Jacobs
NIPS2
2008 Learning to Combine Motor Primitives Via Greedy Additive Regression
Manu Chhabra, Robert A. Jacobs
J. Mach. Learn. Res.2
2007 Behavioral Shaping for Geometric Concepts
Manu Chhabra, Robert A. Jacobs, Daniel Stefankovic
J. Mach. Learn. Res.2
2006 Properties of Synergies Arising from a Theory of Optimal Motor Behavior
abstract
We consider the properties of motor components, also known as synergies, arising from a computational theory (in the sense of Marr, 1982) of optimal motor behavior. An actor's goals were formalized as cost functions, and the optimal control signals minimizing the cost functions were calculated. Optimal synergies were derived from these optimal control signals using a variant of nonnegative matrix factorization. This was done using two different simulated two--joint arms--an arm controlled directly by torques applied at the joints and an arm in which forces were applied by muscles--and two types of motor tasks-reaching tasks and via-point tasks. Studies of the motor synergies reveal several interesting findings. First, optimal motor actions can be generated by summing a small number of scaled and time-shifted motor synergies, indicating that optimal movements can be planned in a low-dimensional space by using optimal motor synergies as motor primitives or building blocks. Second, some optimal synergies are task independent--they arise regardless of the task context-whereas other synergies are task dependent--they arise in the context of one task but not in the contexts of other tasks. Biological organisms use a combination of task--independent and task--dependent synergies. Our work suggests that this may be an efficient combination for generating optimal motor actions from motor primitives. Third, optimal motor actions can be rapidly acquired by learning new linear combinations of optimal motor synergies. This result provides further evidence that optimal motor synergies are useful motor primitives. Fourth, synergies with similar properties arise regardless if one uses an arm controlled by torques applied at the joints or an arm controlled by muscles, suggesting that synergies, when considered in "movement space," are more a reflection of task goals and constraints than of fine details of the underlying hardware.
Manu Chhabra, Robert A. Jacobs
Neural Comput.2
2006 The Costs of Ignoring High-Order Correlations in Populations of Model Neurons
abstract
Investigators debate the extent to which neural populations use pairwise and higher-order statistical dependencies among neural responses to represent information about a visual stimulus. To study this issue, three statistical decoders were used to extract the information in the responses of model neurons about the binocular disparities present in simulated pairs of left-eye and right-eye images: (1) the full joint probability decoder considered all possible statistical relations among neural responses as potentially important; (2) the dependence tree decoder also considered all possible relations as potentially important, but it approximated high-order statistical correlations using a computationally tractable procedure; and (3) the independent response decoder, which assumed that neural responses are statistically independent, meaning that all correlations should be zero and thus can be ignored. Simulation results indicate that high-order correlations among model neuron responses contain significant information about binocular disparities and that the amount of this high-order information increases rapidly as a function of neural population size. Furthermore, the results highlight the potential importance of the dependence tree decoder to neuroscientists as a powerful but still practical way of approximating high-order correlations among neural responses.
Melchi M. Michel, Robert A. Jacobs
Neural Comput.2
2003 Developmental Constraints Aid the Acquisition of Binocular Disparity Sensitivities
abstract
This article considers the hypothesis that systems learning aspects of visual perception may benefit from the use of suitably designed developmental progressions during training. We report the results of simulations in which four models were trained to detect binocular disparities in pairs of visual images. Three of the models were developmental models in the sense that the nature of their visual input changed during the course of training. These models received a relatively impoverished visual input early in training, and the quality of this input improved as training progressed. One model used a coarse-scale-to-multiscale developmental progression, another used a fine-scale-to-multiscale progression, and the third used a random progression. The final model was nondevelopmental in the sense that the nature of its input remained the same throughout the training period. The simulation results show that the two developmental models whose progressions were organized by spatial frequency content consistently outperformed the nondevelopmental and random developmental models. We speculate that the superior performance of these two models is due to two important features of their developmental progressions: (1) these models were exposed to visual inputs at a single scale early in training, and (2) the spatial scale of their inputs progressed in an orderly fashion from one scale to a neighboring scale during training. Simulation results consistent with these speculations are presented. We conclude that suitably designed developmental sequences can be useful to systems learning to detect binocular disparities. The idea that visual development can aid visual learning is a viable hypothesis in need of study.
Melissa Dominguez, Robert A. Jacobs
Neural Comput.2
2003 A Developmental Approach Aids Motor Learning
abstract
Bernstein (1967) suggested that people attempting to learn to perform a difficult motor task try to ameliorate the degrees-of-freedom problem through the use of a developmental progression. Early in training, people maintain a subset of their control parameters (e.g., joint positions) at constant settings and attempt to learn to perform the task by varying the values of the remaining parameters. With practice, people refine and improve this early-learned control strategy by also varying those parameters that were initially held constant. We evaluated Bernstein's proposed developmental progression using six neural network systems and found that a network whose training included developmental progressions of both its trajectory and its feedback gains outperformed all other systems. These progressions, however, yielded performance benefits only on motor tasks that were relatively difficult to learn. We conclude that development can indeed aid motor learning.
Volodymyr Ivanchenko, Robert A. Jacobs
Neural Comput.2
2003 Visual Development and the Acquisition of Motion Velocity Sensitivities
abstract
We consider the hypothesis that systems learning aspects of visual perception may benefit from the use of suitably designed developmental progressions during training. Four models were trained to estimate motion velocities in sequences of visual images. Three of the models were developmental models in the sense that the nature of their visual input changed during the course of training. These models received a relatively impoverished visual input early in training, and the quality of this input improved as training progressed. One model used a coarse-to-multiscale developmental progression (it received coarse-scale motion features early in training and finer-scale features were added to its input as training progressed), another model used a fine-to-multiscale progression, and the third model used a random progression. The final model was nondevelopmental in the sense that the nature of its input remained the same throughout the training period. The simulation results show that the coarse-to-multiscale model performed best. Hypotheses are offered to account for this model's superior performance, and simulation results evaluating these hypotheses are reported. We conclude that suitably designed developmental sequences can be useful to systems learning to estimate motion velocities. The idea that visual development can aid visual learning is a viable hypothesis in need of further study.
Robert A. Jacobs, Melissa Dominguez
Neural Comput.1
2002 Visual Development Aids the Acquisition of Motion Velocity Sensitivities
abstract
We consider the hypothesis that systems learning aspects of visual per- ception may benefit from the use of suitably designed developmental pro- gressions during training. Four models were trained to estimate motion velocities in sequences of visual images. Three of the models were “de- velopmental models” in the sense that the nature of their input changed during the course of training. They received a relatively impoverished visual input early in training, and the quality of this input improved as training progressed. One model used a coarse-to-multiscale develop- mental progression (i.e. it received coarse-scale motion features early in training and finer-scale features were added to its input as training progressed), another model used a fine-to-multiscale progression, and the third model used a random progression. The final model was non- developmental in the sense that the nature of its input remained the same throughout the training period. The simulation results show that the coarse-to-multiscale model performed best. Hypotheses are offered to account for this model’s superior performance. We conclude that suit- ably designed developmental sequences can be useful to systems learn- ing to estimate motion velocities. The idea that visual development can aid visual learning is a viable hypothesis in need of further study.
Robert A. Jacobs, Melissa Dominguez
NIPS1
2002 Factorial Hidden Markov Models and the Generalized Backfitting Algorithm
abstract
Previous researchers developed new learning architectures for sequential data by extending conventional hidden Markov models through the use of distributed state representations. Although exact inference and parameter estimation in these architectures is computationally intractable, Ghahramani and Jordan (1997) showed that approximate inference and parameter estimation in one such architecture, factorial hidden Markov models (FHMMs), is feasible in certain circumstances. However, the learning algorithm proposed by these investigators, based on variational techniques, is difficult to understand and implement and is limited to the study of real-valued data sets. This chapter proposes an alternative method for approximate inference and parameter estimation in FHMMs based on the perspective that FHMMs are a generalization of a well-known class of statistical models known as generalized additive models (GAMs; Hastie & Tibshirani, 1990). Using existing statistical techniques for GAMs as a guide, we have developed the generalized backfitting algorithm. This algorithm computes customized error signals for each hidden Markov chain of an FHMM and then trains each chain one at a time using conventional techniques from the hidden Markov models literature. Relative to previous perspectives on FHMMs, we believe that the viewpoint taken here has a number of advantages. First, it places FHMMs on firm statistical foundations by relating them to a class of models that are well studied in the statistics community, yet it generalizes this class of models in an interesting way. Second, it leads to an understanding of how FHMMs can be applied to many different types of time-series data, including Bernoulli and multinomial data, not just data that are real valued. Finally, it leads to an effective learning procedure for FHMMs that is easier to understand and easier to implement than existing learning procedures. Simulation results suggest that FHMMs trained with the generalized backfitting algorithm are a practical and powerful tool for analyzing sequential data.
Robert A. Jacobs, Wenxin Jiang 0003, Martin A. Tanner
Neural Comput.1
2001 Visual Development and the Acquisition of Binocular Disparity Sensitivities
Melissa Dominguez, Robert A. Jacobs
ICML2
1999 Modeling the Combination of Motion, Stereo, and Vergence Angle Cues to Visual Depth
abstract
Three models of visual cue combination were simulated: a weak fusion model, a modified weak model, and a strong model. Their relative strengths and weaknesses are evaluated on the basis of their performances on the tasks of judging the depth and shape of an ellipse. The models differ in the amount of interaction that they permit among the cues of stereo, motion, and vergence angle. Results suggest that the constrained nonlinear interaction of the modified weak model allows better performance than either the linear interaction of the weak model or the unconstrained nonlinear interaction of the strong model. Further examination of the modified weak model revealed that its weighting of motion and stereo cues was dependent on the task, the viewing distance, and, to a lesser degree, the noise model. Although the dependencies were sensible from a computational viewpoint, they were sometimes inconsistent with psychophysical experimental data. In a second set of experiments, the modified weak model was given contradictory motion and stereo information. One cue was informative in the sense that it indicated an ellipse, while the other cue indicated a flat surface. The modified weak model rapidly reweighted its use of stereo and motion cues as a function of each cue's informativeness. Overall, the simulation results suggest that relative to the weak and strong models, the modified weak fusion model is a good candidate model of the combination of motion, stereo, and vergence angle cues, although the results also highlight areas in which this model needs modification or further elaboration.
Ione Fine, Robert A. Jacobs
Neural Comput.2
1997 Bias/Variance Analyses of Mixtures-of-Experts Architectures
abstract
This article investigates the bias and variance of mixtures-of-experts (ME) architectures. The variance of an ME architecture can be expressed as the sum of two terms: the first term is related to the variances of the expert networks that comprise the architecture and the second term is related to the expert networks' covariances. One goal of this article is to study and quantify a number of properties of ME architectures via the metrics of bias and variance. A second goal is to clarify the relationships between this class of systems and other systems that have recently been proposed. It is shown that in contrast to systems that produce unbiased experts whose estimation errors are uncorrelated, ME architectures produce biased experts whose estimates are negatively correlated.
Robert A. Jacobs
Neural Comput.1
1997 A Bayesian Approach to Model Selection in Hierarchical Mixtures-of-Experts Architectures
abstract
There does not exist a statistical model that shows good performance on all tasks. Consequently, the model selection problem is unavoidable; investigators must decide which model is best at summarizing the data for each task of interest. This article presents an approach to the model selection problem in hierarchical mixtures-of-experts architectures. These architectures combine aspects of generalized linear models with those of finite mixture models in order to perform tasks via a recursive "divide-and-conquer" strategy. Markov chain Monte Carlo methodology is used to estimate the distribution of the architectures' parameters. One part of our approach to model selection attempts to estimate the worth of each component of an architecture so that relatively unused components can be pruned from the architecture's structure. A second part of this approach uses a Bayesian hypothesis testing procedure in order to differentiate inputs that carry useful information from nuisance inputs. Simulation results suggest that the approach presented here adheres to the dictum of Occam's razor; simple architectures that are adequate for summarizing the data are favored over more complex structures. Copyright 1997 Elsevier Science Ltd. All Rights Reserved.
Robert A. Jacobs, Fengchun Peng, Martin A. Tanner
Neural Networks1
1995 Methods for combining experts' probability assessments
abstract
This article reviews statistical techniques for combining multiple probability distributions. The framework is that of a decision maker who consults several experts regarding some events. The experts express their opinions in the form of probability distributions. The decision maker must aggregate the experts' distributions into a single distribution that can be used for decision making. Two classes of aggregation methods are reviewed. When using a supra Bayesian procedure, the decision maker treats the expert opinions as data that may be combined with its own prior distribution via Bayes' rule. When using a linear opinion pool, the decision maker forms a linear combination of the expert opinions. The major feature that makes the aggregation of expert opinions difficult is the high correlation or dependence that typically occurs among these opinions. A theme of this paper is the need for training procedures that result in experts with relatively independent opinions or for aggregation methods that implicitly or explicitly model the dependence among the experts. Analyses are presented that show that m dependent experts are worth the same as k independent experts where k < or = m. In some cases, an exact value for k can be given; in other cases, lower and upper bounds can be placed on k.
Robert A. Jacobs
Neural Comput.1
1994 Hierarchical Mixtures of Experts and the EM Algorithm
abstract
We present a tree-structured architecture for supervised learning. The statistical model underlying the architecture is a hierarchical mixture model in which both the mixture coefficients and the mixture components are generalized linear models (GLIM's). Learning is treated as a maximum likelihood problem; in particular, we present an Expectation-Maximization (EM) algorithm for adjusting the parameters of the architecture. We also develop an on-line learning algorithm in which the parameters are updated incrementally. Comparative simulation results are presented in the robot dynamics domain.
Michael I. Jordan, Robert A. Jacobs
Neural Comput.2
1993 Supervised Learning and Divide-and-Conquer: A Statistical Approach
Michael I. Jordan, Robert A. Jacobs
ICML2
1993 Learning piecewise control strategies in a modular neural network architecture
abstract
The authors describe a multinetwork, or modular, neural network architecture that learns to perform control tasks using a piecewise control strategy. The architecture's networks compete to learn the training patterns. As a result, a plant's parameter space is adaptively partitioned into a number of regions, and a different network learns a control law in each region. This learning process is described in a probabilistic framework and learning algorithms that perform gradient ascent in a log-likelihood function are discussed. Simulations show that the modular architecture's performance is superior to that of a single network on a multipayload robot motion control task.>
Robert A. Jacobs, Michael I. Jordan
IEEE Trans. Syst. Man Cybern.1
1991 Hierarchies of Adaptive Experts
Michael I. Jordan, Robert A. Jacobs
NIPS2
1991 Adaptive Mixtures of Local Experts
abstract
We present a new supervised learning procedure for systems composed of many separate networks, each of which learns to handle a subset of the complete set of training cases. The new procedure can be viewed either as a modular version of a multilayer supervised network, or as an associative version of competitive learning. It therefore provides a new link between these two apparently different approaches. We demonstrate that the learning procedure divides up a vowel discrimination task into appropriate subtasks, each of which can be solved by a very simple expert network.
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, Geoffrey E. Hinton
Neural Comput.1
1990 A Competitive Modular Connectionist Architecture
Robert A. Jacobs, Michael I. Jordan
NIPS1
1989 Learning to Control an Unstable System with Forward Modeling
Michael I. Jordan, Robert A. Jacobs
NIPS2
1988 Increased rates of convergence through learning rate adaptation
Robert A. Jacobs
Neural Networks1