VLDB 2026 Research / reviewers in the wild / expert
Rajesh P. N. Rao
dblp:r/RajeshPNRao · also Raj P. N. Rao
· DBLP profile ↗
80ranked-venue papers
21as first author
11since 2021 · last 2026
0000-0003-0682-8952ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 66 · 18 first-author · 7 since 2021Systems, architecture and hardware · 13Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorTheory of computation · 3 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixed-Density Diffuser: Efficient Planning with Non-Uniform Temporal ResolutionabstractRecent studies demonstrate that diffusion planners benefit from sparsestep planning over single-step planning.Training models to skip steps in their trajectories helps capture long-term dependencies without additional memory or computational cost.However, predicting excessively sparse plans degrades performance.We hypothesize this temporal density threshold is non-uniform across a planning horizon and that certain parts of a predicted trajectory should be more densely generated.We propose Mixed-Density Diffuser (MDD), a diffusion planner where the densities throughout the horizon are tunable hyperparameters.We show that MDD surpasses the SOTA Diffusion Veteran (DV) framework across the Maze2D, Franka Kitchen, and Antmaze Datasets for Deep Data-Driven Reinforcement Learning (D4RL) task domains, achieving a new SOTA on the D4RL benchmark. Crimson Stambaugh, Rajesh P. N. Rao |
ESANN | 2 |
| 2026 | Predictive coding with spiking neural networks: A survey
Antony W. N'Dri, William Gebhardt, Céline Teulière, Fleur Zeldenrust, Rajesh P. N. Rao, Jochen Triesch, Alexander Ororbia |
Neural Networks | 5 |
| 2026 | A survey on neuro-mimetic deep learning via predictive codingabstractArtificial intelligence (AI) is rapidly becoming one of the key technologies of this century. The majority of results in AI thus far have been achieved using deep neural networks trained with a learning algorithm called error backpropagation, always considered biologically implausible. To this end, recent works have studied learning algorithms for deep neural networks inspired by the neurosciences. One such theory, called predictive coding (PC), has shown promising properties that make it potentially valuable for the machine learning community: it can model information processing in different areas of the brain, can be used in control and robotics, has a solid mathematical foundation in variational inference, and performs its computations asynchronously. Inspired by such properties, works that propose novel PC-like algorithms are starting to be present in multiple sub-fields of machine learning and AI at large. Here, we survey such efforts by first providing a broad overview of the history of PC to provide common ground for the understanding of the recent developments, then by describing current efforts and results, and concluding with a large discussion of possible implications and ways forward. Tommaso Salvatori, Ankur Mali, Christopher L. Buckley, Thomas Lukasiewicz, Rajesh P. N. Rao, Karl J. Friston, Alexander Ororbia |
Neural Networks | 5 |
| 2024 | Active Predictive Coding: A Unifying Neural Model for Active Perception, Compositional Learning, and Hierarchical PlanningabstractThere is growing interest in predictive coding as a model of how the brain learns through predictions and prediction errors. Predictive coding models have traditionally focused on sensory coding and perception. Here we introduce active predictive coding (APC) as a unifying model for perception, action, and cognition. The APC model addresses important open problems in cognitive science and AI, including (1) how we learn compositional representations (e.g., part-whole hierarchies for equivariant vision) and (2) how we solve large-scale planning problems, which are hard for traditional reinforcement learning, by composing complex state dynamics and abstract actions from simpler dynamics and primitive actions. By using hypernetworks, self-supervised learning, and reinforcement learning, APC learns hierarchical world models by combining task-invariant state transition networks and task-dependent policy networks at multiple abstraction levels. We illustrate the applicability of the APC model to active visual perception and hierarchical planning. Our results represent, to our knowledge, the first proof-of-concept demonstration of a unified approach to addressing the part-whole learning problem in vision, the nested reference frames learning problem in cognition, and the integrated state-action hierarchy learning problem in reinforcement learning. Rajesh P. N. Rao, Dimitrios C. Gklezakos, Vishwas Sathish |
Neural Comput. | 1 |
| 2024 | Dynamic predictive coding: A model of hierarchical sequence learning and prediction in the neocortexabstractWe introduce dynamic predictive coding, a hierarchical model of spatiotemporal prediction and sequence learning in the neocortex. The model assumes that higher cortical levels modulate the temporal dynamics of lower levels, correcting their predictions of dynamics using prediction errors. As a result, lower levels form representations that encode sequences at shorter timescales (e.g., a single step) while higher levels form representations that encode sequences at longer timescales (e.g., an entire sequence). We tested this model using a two-level neural network, where the top-down modulation creates low-dimensional combinations of a set of learned temporal dynamics to explain input sequences. When trained on natural videos, the lower-level model neurons developed space-time receptive fields similar to those of simple cells in the primary visual cortex while the higher-level responses spanned longer timescales, mimicking temporal response hierarchies in the cortex. Additionally, the network's hierarchical sequence representation exhibited both predictive and postdictive effects resembling those observed in visual motion processing in humans (e.g., in the flash-lag illusion). When coupled with an associative memory emulating the role of the hippocampus, the model allowed episodic memories to be stored and retrieved, supporting cue-triggered recall of an input sequence similar to activity recall in the visual cortex. When extended to three hierarchical levels, the model learned progressively more abstract temporal representations along the hierarchy. Taken together, our results suggest that cortical processing and learning of sequences can be interpreted as dynamic predictive coding based on a hierarchical spatiotemporal generative model of the visual world. Linxing Preston Jiang, Rajesh P. N. Rao |
PLoS Comput. Biol. | 2 |
| 2023 | Dynamic Predictive Coding Explains Both Prediction and Postdiction in Visual Motion Perception
Linxing Preston Jiang, Rajesh P. N. Rao |
CogSci | 2 |
| 2023 | Active Predictive Coding: A Unified Neural Framework for Learning Hierarchical World Models for Perception and Planning
Rajesh P. N. Rao, Dimitrios C. Gklezakos, Vishwas Sathish |
CogSci | 1 |
| 2023 | Expressive probabilistic sampling in recurrent neural networksabstractIn sampling-based Bayesian models of brain function, neural activities are assumed to be samples from probability distributions that the brain uses for probabilistic computation. However, a comprehensive understanding of how mechanistic models of neural dynamics can sample from arbitrary distributions is still lacking. We use tools from functional analysis and stochastic differential equations to explore the minimum architectural requirements for $\textit{recurrent}$ neural circuits to sample from complex distributions. We first consider the traditional sampling model consisting of a network of neurons whose outputs directly represent the samples ($\textit{sampler-only}$ network). We argue that synaptic current and firing-rate dynamics in the traditional model have limited capacity to sample from a complex probability distribution. We show that the firing rate dynamics of a recurrent neural circuit with a separate set of output units can sample from an arbitrary probability distribution. We call such circuits $\textit{reservoir-sampler networks}$ (RSNs). We propose an efficient training procedure based on denoising score matching that finds recurrent and output weights such that the RSN implements Langevin sampling. We empirically demonstrate our model's ability to sample from several complex data distributions using the proposed neural dynamics and discuss its applicability to developing the next generation of sampling-based Bayesian brain models. Shirui Chen, Linxing Jiang, Rajesh P. N. Rao, Eric Shea-Brown |
NeurIPS | 3 |
| 2022 | Repairing Brain-Computer Interfaces with Fault-Based Data AcquisitionabstractBrain-computer interfaces (BCIs) decode recorded neural signals from the brain and/or stimulate the brain with encoded neural signals. BCIs span both hardware and software and have a wide range of applications in restorative medicine, from restoring movement through prostheses and robotic limbs to restoring sensation and communication through spellers. BCIs also have applications in diagnostic medicine, e.g., providing clinicians with data for detecting seizures, sleep patterns, or emotions. Cailin Winston, Caleb Winston, Chloe N. Winston, Claris Winston, Cleah Winston, Rajesh P. N. Rao, René Just |
ICSE | 6 |
| 2022 | Touching the Void: Intracranial Stimulation for NeuroHaptic Feedback in Virtual RealityabstractDirect cortical stimulation of the somatosensory cortex (SI-DCS) has been shown to evoke distinct and localizable percepts, exploitable as neurohaptic feedback. In this study, we leveraged a novel virtual reality (VR) experimental platform to evaluate SI-DCS neurohaptic feedback during naturalistic object interaction. Two human subjects implanted with intracranial electrodes for seizure localization were asked to discriminate between visually identical virtual objects based on their distinct SI-DCS neurohaptic profiles. In a binary discrimination task, neurohaptic feedback was either present or absent while grasping a virtual object. In the ternary discrimination task, neurohaptic feedback was either present in one of two distinct neurohaptic sequences or absent. Both subjects performed significantly above chance in binary and ternary discrimination, demonstrating the efficacy of S1-DCS as neurohaptic feedback. Successful ternary discrimination also demonstrated that different sequences of amplitude-modulated SI-DCS at a single pair of electrodes can evoke discriminable neurohaptic percepts. Moreover, amplitude-modulated SI-DCS sequences were shown to elicit sensorimimetic percepts described as “bumpy” and “smooth” in Subject 1, and as a sensation of movement in the paralyzed hand of Subject 2. Our study demonstrates the reliability and discriminability of both simple and complex SI-DCS for neurohaptic feedback during immersive VR object interaction and supports the use of immersive VR for neurohaptic design towards the development of functional brain computer interface. Courtnie Paschall, Jason S. Hauptman, Rajesh P. N. Rao, Jeffrey G. Ojemann, Jeffrey Herron |
SMC | 3 |
| 2022 | Human intracortical responses to varying electrical stimulation conditions are separable in low-dimensional subspacesabstractElectrical stimulation is a powerful tool for targeted neurorehabilitation, and recent work in adaptive stimulation where stimulation can be adjusted in real-time has shown promise in improving stimulation outcomes and reducing stimulation-induced side effects. Mapping the relationship between electrical stimulation input and neural activity response can help reveal their interactions and can give us tools to iterate and improve on our stimulation protocols. Here, we introduce methods for identifying low-dimensional subspaces of human intracortical responses to electrical stimulation in invasive electroencephalography. In epilepsy patients (n=4) undergoing clinical monitoring, we applied a stimulation protocol of varying amplitude and frequency in 5-second intervals to capture a range of responses to different stimulation conditions. We characterized these responses using time-frequency spectral power, applied baseline subtraction and outlier removal procedures, and performed principal component analysis across frequencies. We identified that intracortical responses to different stimulation conditions can be represented in a 3-dimensional subspace, accounting for more than 95% of the variance. Using support vector machine classification, we demonstrated separability of intracortical responses in different stimulation conditions across subjects, where this separability was contingent on applying baseline subtraction and outlier removal. Our results represent a first step towards building a predictive model of neural response from stimulation input, an important prerequisite for adaptive closed-loop stimulation for targeted neurorehabilitation. Samantha Sun, Lila Levinson, Courtnie Paschall, Jeffrey Herron, Kurt E. Weaver, Jason S. Hauptman, Jeffrey G. Ojemann, Rajesh P. N. Rao |
SMC | 9 |
| 2018 | AJILE Movement Prediction: Multimodal Deep Learning for Natural Human Neural Recordings and VideoabstractDeveloping useful interfaces between brains and machines is a grand challenge of neuroengineering. An effective interface has the capacity to not only interpret neural signals, but predict the intentions of the human to perform an action in the near future; prediction is made even more challenging outside well-controlled laboratory experiments. This paper describes our approach to detect and to predict natural human arm movements in the future, a key challenge in brain computer interfacing that has never before been attempted. We introduce the novel Annotated Joints in Long-term ECoG (AJILE) dataset; AJILE includes automatically annotated poses of 7 upper body joints for four human subjects over 670 total hours (more than 72 million frames), along with the corresponding simultaneously acquired intracranial neural recordings. The size and scope of AJILE greatly exceeds all previous datasets with movements and electrocorticography (ECoG), making it possible to take a deep learning approach to movement prediction. We propose a multimodal model that combines deep convolutional neural networks (CNN) with long short-term memory (LSTM) blocks, leveraging both ECoG and video modalities. We demonstrate that our models are able to detect movements and predict future movements up to 800 msec before movement initiation. Further, our multimodal movement prediction models exhibit resilience to simulated ablation of input neural signals. We believe a multimodal approach to natural neural decoding that takes context into account is critical in advancing bioelectronic technologies and human neuroscience. Nancy Xin Ru Wang, Ali Farhadi, Rajesh P. N. Rao, Bingni W. Brunton |
AAAI | 3 |
| 2018 | Learning Graph-Structured Sum-Product Networks for Probabilistic Semantic MapsabstractWe introduce Graph-Structured Sum-Product Networks (GraphSPNs), a probabilistic approach to structured prediction for problems where dependencies between latent variables are expressed in terms of arbitrary, dynamic graphs. While many approaches to structured prediction place strict constraints on the interactions between inferred variables, many real-world problems can be only characterized using complex graph structures of varying size, often contaminated with noise when obtained from real data. Here, we focus on one such problem in the domain of robotics. We demonstrate how GraphSPNs can be used to bolster inference about semantic, conceptual place descriptions using noisy topological relations discovered by a robot exploring large-scale office spaces. Through experiments, we show that GraphSPNs consistently outperform the traditional approach based on undirected graphical models, successfully disambiguating information in global semantic maps built from uncertain, noisy local evidence. We further exploit the probabilistic nature of the model to infer marginal distributions over semantic descriptions of as yet unexplored places and detect spatial environment configurations that are novel and incongruent with the known evidence. Kaiyu Zheng, Andrzej Pronobis, Rajesh P. N. Rao |
AAAI | 3 |
| 2018 | Electrocorticographic Dynamics Predict Sustained Grasping and Upper-Limb Kinetic OutputabstractDetermining hand grasping forces and kinetic direction and magnitude are important for the development of dexterous neural prosthetics. While many earlier decoding methods have successfully predicted upper-limb kinematic output from cortical signals in the sensorimotor parietal and premotor regions, the full extent of the regions that characterize kinetic behavior is unknown. In this study, we found that neural dynamics based on electrocorticography (ECoG) recorded from the human brain surface can successfully encode structured and unstructured grasping and arm kinetic output. We found a time-averaged linear relationship between gamma band spectral ECoG power and sustained grasping force output with visual feedback. In the kinetic grasping task, we obtained classification accuracy of 47% (25% = chance) using quadratic discriminant analysis. Additionally, we also found a similar linear relationship between spectral power and cued isometric force generation, without concurrent visual feedback, in arm force application; this feature could also classify arm force output categories with an accuracy of 41% (33% = chance). In addition, we applied quadratic discriminant analysis with top 12 principal components to attain approximately 26% accuracy (chance is 16%) in determining arm kinetic force direction. We found that the gamma band spectral power from both experiments in posterior parietal cortex, as well as projections of high gamma variability along the top principal components, can be successfully used to predict sustained kinetic outputs and to explain both structured and unstructured force output variability. Our findings contribute to a deeper understanding of neural dynamics for fine kinetic behavior. Katie Ly, Jing Wu 0015, Lila Levinson, Benjamin R. Shuman, Katherine Muterspaugh Steele, Jeffrey G. Ojemann, Rajesh P. N. Rao |
SMC | 7 |
| 2017 | Learning deep generative spatial models for mobile robotsabstractWe propose a new probabilistic framework that allows mobile robots to autonomously learn deep, generative models of their environments that span multiple levels of abstraction. Unlike traditional approaches that combine engineered models for low-level features, geometry, and semantics, our approach leverages recent advances in Sum-Product Networks (SPNs) and deep learning to learn a single, universal model of the robot's spatial environment. Our model is fully probabilistic and generative, and represents a joint distribution over spatial information ranging from low-level geometry to semantic interpretations. Once learned, it is capable of solving a wide range of tasks: from semantic classification of places, uncertainty estimation, and novelty detection, to generation of place appearances based on semantic information and prediction of missing data in partial observations. Experiments on laser-range data from a mobile robot show that the proposed universal model obtains performance superior to state-of-the-art models fine-tuned to one specific task, such as Generative Adversarial Networks (GANs) or SVMs. Andrzej Pronobis, Rajesh P. N. Rao |
IROS | 2 |
| 2016 | Initiative in Robot Assistance during Collaborative Task ExecutionabstractCollaborative robots are quickly gaining momentum in real-world settings. This has motivated many new research questions in human-robot collaboration. In this paper, we address the questions of whether and when a robot should take initiative during joint human-robot task execution. We develop a system capable of autonomously tracking and performing table-top object manipulation tasks with humans and we implement three different initiative models to trigger robot actions. Human-initiated help gives control of robot action timing to the user; robot-initiated reactive help triggers robot assistance when it detects that the user needs help; and robot-initiated proactive help makes the robot help whenever it can. We performed a user study (N=18) to compare these trigger mechanisms in terms of task performance, usage characteristics, and subjective preference. We found that people collaborate best with a proactive robot, yielding better team fluency and high subjective ratings. However, they prefer having control of when the robot should help, rather than working with a reactive robot that only helps when it is needed. Jimmy Baraglia, Maya Cakmak, Yukie Nagai, Rajesh P. N. Rao, Minoru Asada |
HRI | 4 |
| 2016 | Autonomous question answering with mobile robots in human-populated environmentsabstractAutonomous mobile robots will soon become ubiquitous in human-populated environments. Besides their typical applications in fetching, delivery, or escorting, such robots present the opportunity to assist human users in their daily tasks by gathering and reporting up-to-date knowledge about the environment. In this paper, we explore this use case and present an end-to-end framework that enables a mobile robot to answer natural language questions about the state of a large-scale, dynamic environment asked by the inhabitants of that environment. The system parses the question and estimates an initial viewpoint that is likely to contain information for answering the question based on prior environment knowledge. Then, it autonomously navigates towards the viewpoint while dynamically adapting to changes and new information. The output of the system is an image of the most relevant part of the environment that allows the user to obtain an answer to their question. We additionally demonstrate the benefits of a continuously operating information gathering robot by showing how the system can answer retrospective questions about the past state of the world using incidentally recorded sensory data. We evaluate our approach with a custom mobile robot deployed in a university building, with questions collected from occupants of the building. We demonstrate our system's ability to respond to these questions in different environmental conditions. Mike Chung 0001, Andrzej Pronobis, Maya Cakmak, Dieter Fox, Rajesh P. N. Rao |
IROS | 5 |
| 2016 | Bayesian Inference and Online Learning in Poisson Neuronal NetworksabstractMotivated by the growing evidence for Bayesian computation in the brain, we show how a two-layer recurrent network of Poisson neurons can perform both approximate Bayesian inference and learning for any hidden Markov model. The lower-layer sensory neurons receive noisy measurements of hidden world states. The higher-layer neurons infer a posterior distribution over world states via Bayesian inference from inputs generated by sensory neurons. We demonstrate how such a neuronal network with synaptic plasticity can implement a form of Bayesian inference similar to Monte Carlo methods such as particle filtering. Each spike in a higher-layer neuron represents a sample of a particular hidden world state. The spiking activity across the neural population approximates the posterior distribution over hidden states. In this model, variability in spiking is regarded not as a nuisance but as an integral feature that provides the variability necessary for sampling during inference. We demonstrate how the network can learn the likelihood model, as well as the transition probabilities underlying the dynamics, using a Hebbian learning rule. We present results illustrating the ability of the network to perform inference and learning for arbitrary hidden Markov models. Yanping Huang, Rajesh P. N. Rao |
Neural Comput. | 2 |
| 2016 | Spontaneous Decoding of the Timing and Content of Human Object Perception from Cortical Surface Recordings Reveals Complementary Information in the Event-Related Potential and Broadband Spectral ChangeabstractThe link between object perception and neural activity in visual cortical areas is a problem of fundamental importance in neuroscience. Here we show that electrical potentials from the ventral temporal cortical surface in humans contain sufficient information for spontaneous and near-instantaneous identification of a subject's perceptual state. Electrocorticographic (ECoG) arrays were placed on the subtemporal cortical surface of seven epilepsy patients. Grayscale images of faces and houses were displayed rapidly in random sequence. We developed a template projection approach to decode the continuous ECoG data stream spontaneously, predicting the occurrence, timing and type of visual stimulus. In this setting, we evaluated the independent and joint use of two well-studied features of brain signals, broadband changes in the frequency power spectrum of the potential and deflections in the raw potential trace (event-related potential; ERP). Our ability to predict both the timing of stimulus onset and the type of image was best when we used a combination of both the broadband response and ERP, suggesting that they capture different and complementary aspects of the subject's perceptual state. Specifically, we were able to predict the timing and type of 96% of all stimuli, with less than 5% false positive rate and a ~20ms error in timing. Kai J. Miller, Gerwin Schalk, Dora Hermes, Jeffrey G. Ojemann, Rajesh P. N. Rao |
PLoS Comput. Biol. | 5 |
| 2016 | Cortico-Cortical Interactions during Acquisition and Use of a Neuroprosthetic SkillabstractA motor cortex-based brain-computer interface (BCI) creates a novel real world output directly from cortical activity. Use of a BCI has been demonstrated to be a learned skill that involves recruitment of neural populations that are directly linked to BCI control as well as those that are not. The nature of interactions between these populations, however, remains largely unknown. Here, we employed a data-driven approach to assess the interaction between both local and remote cortical areas during the use of an electrocorticographic BCI, a method which allows direct sampling of cortical surface potentials. Comparing the area controlling the BCI with remote areas, we evaluated relationships between the amplitude envelopes of band limited powers as well as non-linear phase-phase interactions. We found amplitude-amplitude interactions in the high gamma (HG, 70-150 Hz) range that were primarily located in the posterior portion of the frontal lobe, near the controlling site, and non-linear phase-phase interactions involving multiple frequencies (cross-frequency coupling between 8-11 Hz and 70-90 Hz) taking place over larger cortical distances. Further, strength of the amplitude-amplitude interactions decreased with time, whereas the phase-phase interactions did not. These findings suggest multiple modes of cortical communication taking place during BCI use that are specialized for function and depend on interaction distance. Jeremiah D. Wander, Devapratim Sarma, Lise A. Johnson, Eberhard E. Fetz, Rajesh P. N. Rao, Jeffrey G. Ojemann, Felix Darvas |
PLoS Comput. Biol. | 5 |
| 2015 | Robot Programming by Demonstration with situated spatial language understandingabstractRobot Programming by Demonstration (PbD) allows users to program a robot by demonstrating the desired behavior. Providing these demonstrations typically involves moving the robot through a sequence of states, often by physically manipulating it. This requires users to be co-located with the robot and have the physical ability to manipulate it. In this paper, we present a natural language based interface for PbD that removes these requirements and enables hands-free programming. We focus on programming object manipulation actions-our key insight is that such actions can be decomposed into known types of manipulator movements that are naturally described using spatial language; e.g., object reference expressions and prepositions. Our method takes a natural language command and the current world state to infer the intended movement command and its parametrization. We implement this method on a two-armed mobile manipulator and demonstrate the different types of manipulation actions that can be programmed with it. We compare it to a kinesthetic PbD interface and we demonstrate our method's ability to deal with incomplete language. Maxwell Forbes, Rajesh P. N. Rao, Luke Zettlemoyer, Maya Cakmak |
ICRA | 2 |
| 2015 | Designing information gathering robots for human-populated environmentsabstractAdvances in mobile robotics have enabled robots that can autonomously operate in human-populated environments. Although primary tasks for such robots might be fetching, delivery, or escorting, they present an untapped potential as information gathering agents that can answer questions for the community of co-inhabitants. In this paper, we seek to better understand requirements for such information gathering robots (InfoBots) from the perspective of the user requesting the information. We present findings from two studies: (i) a user survey conducted in two office buildings and (ii) a 4-day long deployment in one of the buildings, during which inhabitants of the building could ask questions to an InfoBot through a web-based interface. These studies allow us to characterize the types of information that InfoBots can provide for their users. Mike Chung 0001, Andrzej Pronobis, Maya Cakmak, Dieter Fox, Rajesh P. N. Rao |
IROS | 5 |
| 2014 | Non-intrusive tongue machine interfaceabstractThere has been recent interest in designing systems that use the tongue as an input interface. Prior work however either require surgical procedures or in-mouth sensor placements. In this paper, we introduce TongueSee, a non-intrusive tongue machine interface that can recognize a rich set of tongue gestures using electromyography (EMG) signals from the surface of the skin. We demonstrate the feasibility and robustness of TongueSee with experimental studies to classify six tongue gestures across eight participants. TongueSee achieves a classification accuracy of 94.17% and a false positive probability of 0.000358 per second using three-protrusion preamble design. Qiao Zhang 0001, Shyamnath Gollakota, Ben Taskar, Rajesh P. N. Rao |
CHI | 4 |
| 2014 | Robot Programming by Demonstration with Crowdsourced Action FixesabstractProgramming by Demonstration (PbD) can allow end-users to teach robots new actions simply by demonstrating them. However, learning generalizable actions requires a large number of demonstrations that is unreasonable to expect from end-users. In this paper, we explore the idea of using crowdsourcing to collect action demonstrations from the crowd. We propose a PbD framework in which the end-user provides an initial seed demonstration, and then the robot searches for scenarios in which the action will not work and requests the crowd to fix the action for these scenarios. We use instance-based learning with a simple yet powerful action representation that allows an intuitive visualization of the action. Crowd workers directly interact with these visualizations to fix them. We demonstrate the utility of our approach with a user study involving local crowd workers (N=31) and analyze the collected data and the impact of alternative design parameters so as to inform a real-world deployment of our system. Maxwell Forbes, Mike Chung 0001, Maya Cakmak, Rajesh P. N. Rao |
HCOMP | 4 |
| 2014 | Accelerating imitation learning through crowdsourcingabstractAlthough imitation learning is a powerful technique for robot learning and knowledge acquisition from näıve human users, it often suffers from the need for expensive human demonstrations. In some cases the robot has an insufficient number of useful demonstrations, while in others its learning ability is limited by the number of users it directly interacts with. We propose an approach that overcomes these shortcomings by using crowdsourcing to collect a wider variety of examples from a large pool of human demonstrators online. We present a new goal-based imitation learning framework which utilizes crowdsourcing as a major source of human demonstration data. We demonstrate the effectiveness of our approach experimentally on a scenario where the robot learns to build 2D object models on a table from basic building blocks using knowledge gained from locals and online crowd workers. In addition, we show how the robot can use this knowledge to support human-robot collaboration tasks such as goal inference through object-part classification and missing-part prediction. We report results from a user study involving fourteen local demonstrators and hundreds of crowd workers on 16 different model building tasks. Mike Chung 0001, Maxwell Forbes, Maya Cakmak, Rajesh P. N. Rao |
ICRA | 4 |
| 2014 | Neurons as Monte Carlo Samplers: Bayesian Inference and Learning in Spiking Networks
Yanping Huang, Rajesh P. N. Rao |
NIPS | 2 |
| 2012 | Automatic extraction of command hierarchies for adaptive brain-robot interfacingabstractRecent advances in neuroscience and robotics have allowed initial demonstrations of brain-computer interfaces (BCIs) for controlling wheeled and humanoid robots. However, further advances have proved challenging due to the low throughput of the interfaces and the high degrees-of-freedom (DOF) of the robots. In this paper, we build on our previous work on Hierarchical BCIs (HBCIs) which seek to mitigate this problem. We extend HBCIs to allow training of arbitrarily complex tasks, with training no longer restricted to a particular robot state space (such as Cartesian space for a navigation task). We present two algorithms for learning command hierarchies by automatically extracting patterns from a user's command history. The first algorithm builds an arbitrary-level hierarchical structure (a “control grammar”) whose elements can represent skills, whole tasks, collections of tasks, etc. The user “executes” single symbols from this grammar, which produce sequences of lower-level commands. The second algorithm, which is probabilistic, also learns sequences which can be executed as high-level commands, but does not build an explicit hierarchical structure. Both algorithms provide a de facto form of dictionary compression, which enhances the effective throughput of the BCI. We present results from two human subjects who successfully used the hierarchical BCI to control a simulated PR2 robot using brain signals recorded non-invasively through electroencephalography (EEG). Matthew J. Bryan, Griffin Nicoll, Vibinash Thomas, Mike Chung 0001, Joshua R. Smith 0001, Rajesh P. N. Rao |
ICRA | 6 |
| 2012 | How Prior Probability Influences Decision Making: A Unifying Probabilistic ModelabstractHow does the brain combine prior knowledge with sensory evidence when making decisions under uncertainty? Two competing descriptive models have been proposed based on experimental data. The first posits an additive offset to a decision variable, implying a static effect of the prior. However, this model is inconsistent with recent data from a motion discrimination task involving temporal integration of uncertain sensory evidence. To explain this data, a second model has been proposed which assumes a time-varying influence of the prior. Here we present a normative model of decision making that incorporates prior knowledge in a principled way. We show that the additive offset model and the time-varying prior model emerge naturally when decision making is viewed within the framework of partially observable Markov decision processes (POMDPs). Decision making in the model reduces to (1) computing beliefs given observations and prior information in a Bayesian manner, and (2) selecting actions based on these beliefs to maximize the expected sum of future rewards. We show that the model can explain both data previously explained using the additive offset model as well as more recent data on the time-varying influence of prior knowledge on decision making. Yanping Huang, Abram L. Friesen, Timothy D. Hanks, Michael N. Shadlen, Rajesh P. N. Rao |
NIPS | 5 |
| 2012 | Fast Structured Prediction Using Large Margin Sigmoid Belief Networks
Xu Miao, Rajesh P. N. Rao |
Int. J. Comput. Vis. | 2 |
| 2012 | Complementary Kernel Density Estimation
Xu Miao, Rajesh P. N. Rao |
Pattern Recognit. Lett. | 3 |
| 2011 | Gaze Following as Goal Inference: A Bayesian Model
Abram L. Friesen, Rajesh P. N. Rao |
CogSci | 2 |
| 2011 | A Hierarchical Architecture for Adaptive Brain-Computer Interfacing
Mike Chung 0001, Willy Cheung, Reinhold Scherer, Rajesh P. N. Rao |
IJCAI | 4 |
| 2010 | A rational decision making framework for inhibitory controlabstractIntelligent agents are often faced with the need to choose actions with uncertain consequences, and to modify those actions according to ongoing sensory processing and changing task demands. The requisite ability to dynamically modify or cancel planned actions is known as inhibitory control in psychology. We formalize inhibitory control as a rational decision-making problem, and apply to it to the classical stop-signal task. Using Bayesian inference and stochastic control tools, we show that the optimal policy systematically depends on various parameters of the problem, such as the relative costs of different action choices, the noise level of sensory inputs, and the dynamics of changing environmental demands. Our normative model accounts for a range of behavioral data in humans and animals in the stop-signal task, suggesting that the brain implements statistically optimal, dynamically adaptive, and reward-sensitive decision-making in the context of inhibitory control problems. Pradeep Shenoy, Rajesh P. N. Rao, Angela J. Yu |
NIPS | 2 |
| 2010 | Entropy, the Indus Script, and Language: A Reply to R. SproatabstractIn a recent LastWords column (Sproat 2010), Richard Sproat laments the reviewing practices of “general science journals” after dismissing our work and that of Lee, Jonathan, and Ziman (2010) as “useless” and “trivially and demonstrably wrong.” Although we expect such categorical statements to have already raised some red flags in the minds of readers, we take this opportunity to present a more accurate description of our work, point out the straw man argument used in Sproat (2010), and provide a more complete characterization of the Indus script debate. A separate response by Lee and colleagues in this issue provides clarification of issues not covered here. Rajesh P. N. Rao, Nisha Yadav, Mayank N. Vahia, Hrishikesh Joglekar, Ronojoy Adhikari, Iravatham Mahadevan |
Comput. Linguistics | 1 |
| 2010 | "Social" robots are psychological agents for infants: A test of gaze following
Andrew N. Meltzoff, Rechele Brooks, Aaron P. Shon, Rajesh P. N. Rao |
Neural Networks | 4 |
| 2009 | Large Margin Boltzmann Machines
Xu Miao, Rajesh P. N. Rao |
IJCAI | 2 |
| 2009 | Using eigenposes for lossless periodic human motion imitationabstractProgramming a humanoid robot to perform an action that takes the robot's complex dynamics into account is a challenging problem. Traditional approaches typically require highly accurate prior knowledge of the robot's dynamics and environment in order to devise complex control algorithms for generating a stable dynamic motion. Training using human motion capture is an intuitive and flexible approach to programming a robot but directly applying motion capture data to a robot usually results in dynamically unstable motion. Optimization using high-dimensional motion capture data in the humanoid full-body joint-space is also typically intractable. In previous work, we proposed an approach that uses dimensionality reduction to achieve tractable imitation-based learning in humanoids without the need for a physics-based dynamics model. This work was based on a 3D ¿eigenpose¿ representation. However, for some motion patterns, using only three dimensions for eigenposes is insufficient. In this paper, we propose a new method for motion optimization based on high-dimensional eigenpose data. A one-dimensional computationally efficient motion-phase optimization method is implemented along with a newly developed cylindrical coordinate transformation technique for hyperdimensional subspaces. This results in a fast learning algorithm and very accurate motion imitation. We demonstrate the new algorithm on a Fujitsu HOAP-2 humanoid robot model in a dynamic simulator and show that a dynamically stable sidestep motion can be successfully learned by imitating a human demonstrator. Rawichote Chalodhorn, Rajesh P. N. Rao |
IROS | 2 |
| 2008 | Feasibility and pragmatics of classifying working memory load with an electroencephalographabstractA reliable and unobtrusive measurement of working memory load could be used to evaluate the efficacy of interfaces and to provide real-time user-state information to adaptive systems. In this paper, we describe an experiment we con-ducted to explore some of the issues around using an elec-troencephalograph (EEG) for classifying working memory load. Within this experiment, we present our classification methodology, including a novel feature selection scheme that seems to alleviate the need for complex drift modeling and artifact rejection. We demonstrate classification accuracies of up to 99% for 2 memory load levels and up to 88% for 4 levels. We also present results suggesting that we can do this with shorter windows, much less training data, and a smaller number of EEG channels, than reported previously. Finally, we show results suggesting that the models we construct transfer across variants of the task, implying some level of generality. We believe these findings extend prior work and bring us a step closer to the use of such technologies in HCI research. David B. Grimes, Desney S. Tan, Scott E. Hudson, Pradeep Shenoy, Rajesh P. N. Rao |
CHI | 5 |
| 2008 | Learning nonparametric policies by imitationabstractA long cherished goal in artificial intelligence has been the ability to endow a robot with the capacity to learn and generalize skills from watching a human teacher. Such an ability to learn by imitation has remained hard to achieve due to a number of factors, including the problem of learning in high-dimensional spaces and the problem of uncertainty. In this paper, we propose a new probabilistic approach to the problem of teaching a high degree-of-freedom robot (in particular, a humanoid robot) flexible and generalizable skills via imitation of a human teacher. The robot uses inference in a graphical model to learn sensor-based dynamics and infer a stable plan from a teacherpsilas demonstration of an action. The novel contribution of this work is a method for learning a nonparametric policy which generalizes a fixed action plan to operate over a continuous space of task variation. A notable feature of the approach is that it does not require any knowledge of the physics of the robot or the environment. By leveraging advances in probabilistic inference and Gaussian process regression, the method produces a nonparametric policy for sensor-based feedback control in continuous state and action spaces. We present experimental and simulation results using a Fujitsu HOAP-2 humanoid robot demonstrating imitation-based learning of a task involving lifting objects of different weights from a single human demonstration. David B. Grimes, Rajesh P. N. Rao |
IROS | 2 |
| 2007 | Active Imitation Learning
Aaron P. Shon, Rajesh P. N. Rao |
AAAI | 3 |
| 2007 | Imitation Learning Using Graphical Models
Rajesh P. N. Rao |
ECML | 2 |
| 2007 | Towards a Real-Time Bayesian Imitation System for a Humanoid RobotabstractImitation learning, or programming by demonstration (PbD), holds the promise of allowing robots to acquire skills from humans with domain-specific knowledge, who nonetheless are inexperienced at programming robots. We have prototyped a real-time, closed-loop system for teaching a humanoid robot to interact with objects in its environment. The system uses nonparametric Bayesian inference to determine an optimal action given a configuration of objects in the world and a desired future configuration. We describe our prototype implementation, show imitation of simple motor acts on a humanoid robot, and discuss extensions to the system Aaron P. Shon, Joshua J. Storz, Rajesh P. N. Rao |
ICRA | 3 |
| 2007 | Learning to Walk through Imitation
Rawichote Chalodhorn, David B. Grimes, Keith Grochow, Rajesh P. N. Rao |
IJCAI | 4 |
| 2007 | Learning full-body motions from monocular vision: dynamic imitation in a humanoid robotabstractIn an effort to ease the burden of programming motor commands for humanoid robots, a computer vision technique is developed for converting a monocular video sequence of human poses into stabilized robot motor commands for a humanoid robot. The human teacher wears a multi-colored body suit while performing a desired set of actions. Leveraging the colors of the body suit, the system detects the most probable locations of the different body parts and joints in the image. Then, by exploiting the known dimensions of the body suit, a user specified number of candidate 3D poses are generated for each frame. Using human to robot joint correspondences, the estimated 3D poses for each frame are then mapped to corresponding robot motor commands. An initial set of kinematically valid motor commands is generated using an approximate best path search through the pose candidates for each frame. Finally a learning-based probabilistic dynamic balance model obtains a dynamically stable imitative sequence of motor commands. We demonstrate the viability of the approach by presenting results showing full-body imitation of human actions by a Fujitsu HOAP-2 humanoid robot. Jeffrey B. Cole, David B. Grimes, Rajesh P. N. Rao |
IROS | 3 |
| 2007 | Learning the Lie Groups of Visual InvarianceabstractA fundamental problem in biological and machine vision is visual invariance: How are objects perceived to be the same despite transformations such as translations, rotations, and scaling? In this letter, we describe a new, unsupervised approach to learning invariances based on Lie group theory. Unlike traditional approaches that sacrifice information about transformations to achieve invariance, the Lie group approach explicitly models the effects of transformations in images. As a result, estimates of transformations are available for other purposes, such as pose estimation and visuomotor control. Previous approaches based on first-order Taylor series expansions of images can be regarded as special cases of the Lie group approach, which utilizes a matrix-exponential-based generative model of images and can handle arbitrarily large transformations. We present an unsupervised expectation-maximization algorithm for learning Lie transformation operators directly from image data containing examples of transformations. Our experimental results show that the Lie operators learned by the algorithm from an artificial data set containing six types of affine transformations closely match the analytically predicted affine operators. We then demonstrate that the algorithm can also recover novel transformation operators from natural image sequences. We conclude by showing that the learned operators can be used to both generate and estimate transformations in images, thereby providing a basis for achieving visual invariance. Xu Miao, Rajesh P. N. Rao |
Neural Comput. | 2 |
| 2006 | Learning Humanoid Motion Dynamics through Sensory-motor Mapping in Reduced Dimensional SpacesabstractOptimization of robot dynamics for a given human motion is an intuitive way to approach the problem of learning complex human behavior by imitation. In this paper, we propose a methodology based on a learning approach that performs optimization of humanoid dynamics in a low-dimensional subspace. We compactly represent the kinematic information of humanoid motion in a low dimensional subspace. Motor commands in the low dimensional subspace are mapped to the expected sensory feedback. We select optimal motor commands based on sensory-motor mapping that also satisfy our kinematic constraints. Finally, we obtain a set of novel postures that result in superior motion dynamics compared to the initial motion. We demonstrate results of the optimized motion on both a dynamics simulator and a real humanoid robot Rawichote Chalodhorn, David B. Grimes, Gabriel Y. Maganis, Rajesh P. N. Rao, Minoru Asada |
ICRA | 4 |
| 2006 | Planning and Acting in Uncertain Environments using Probabilistic InferenceabstractAn important problem in robotics is planning and selecting actions for goal-directed behavior in noisy uncertain environments. The problem is typically addressed within the framework of partially observable Markov decision processes (POMDPs). Although efficient algorithms exist for learning policies for MDPs, these algorithms do not generalize easily to POMDPs. In this paper, we propose a framework for planning and action selection based on probabilistic inference in graphical models. Unlike previous approaches based on MAP inference, our approach utilizes the most probable explanation (MPE) of variables in a graphical model, allowing tractable and efficient inference of actions. It generalizes easily to complex partially observable environments. Furthermore, it allows rewards and costs to be incorporated in a straightforward manner as part of the inference process. We investigate the application of our approach to the problem of robot navigation by testing it on a suite of well-known POMDP benchmarks. Our results demonstrate that the proposed method can beat or match the performance of recently proposed specialized POMDP solvers. Rajesh P. N. Rao |
IROS | 2 |
| 2006 | Learning Nonparametric Models for Probabilistic ImitationabstractLearning by imitation represents an important mechanism for rapid acquisition of new behaviors in humans and robots. A critical requirement for learning by imitation is the ability to handle uncertainty arising from the observation process as well as the imitator's own dynamics and interactions with the environment. In this paper, we present a new probabilistic method for inferring imitative actions that takes into account both the observations of the teacher as well as the imitator's dynamics. Our key contribution is a nonparametric learning method which generalizes to systems with very different dynamics. Rather than relying on a known forward model of the dynamics, our approach learns a nonparametric forward model via exploration. Leveraging advances in approximate inference in graphical models, we show how the learned forward model can be directly used to plan an imitating sequence. We provide experimental results for two systems: a biomechanical model of the human arm and a 25-degrees-of-freedom humanoid robot. We demonstrate that the proposed method can be used to learn appropriate motor inputs to the model arm which imitates the desired movements. A second set of results demonstrates dynamically stable full-body imitation of a human teacher by the humanoid robot. David B. Grimes, Daniel R. Rashid, Rajesh P. N. Rao |
NIPS | 3 |
| 2006 | A probabilistic model of gaze imitation and shared attention
Matt Hoffman 0001, David B. Grimes, Aaron P. Shon, Rajesh P. N. Rao |
Neural Networks | 4 |
| 2005 | Real-Time Classification of Electromyographic Signals for Robotic Control
Beau Crawford, Kai J. Miller, Pradeep Shenoy, Rajesh P. N. Rao |
AAAI | 4 |
| 2005 | Probabilistic Gaze Imitation and Saliency Learning in a Robotic HeadabstractImitation is a powerful mechanism for transferring knowledge from an instructor to a naïve observer, one that is deeply contingent on a state of shared attention between these two agents. In this paper we present Bayesian algorithms that implement the core of an imitation learning framework. We use gaze imitation, coupled with task-dependent saliency learning, to build a state of shared attention between the instructor and observer. We demonstrate the performance of our algorithms in a gaze following and saliency learning task implemented on an active vision robotic head. Our results suggest that the ability to follow gaze and learn instructor-and task-specific saliency models could play a crucial role in building systems capable of complex forms of human-robot interaction. Aaron P. Shon, David B. Grimes, Chris L. Baker, Matt Hoffman 0001, Rajesh P. N. Rao |
ICRA | 6 |
| 2005 | Learning Shared Latent Structure for Image Synthesis and Robotic ImitationabstractWe propose an algorithm that uses Gaussian process regression to learn common hidden structure shared between corresponding sets of heterogenous observations. The observation spaces are linked via a single, reduced-dimensionality latent variable space. We present results from two datasets demonstrating the algorithms's ability to synthesize novel data from learned correspondences. We first show that the method can learn the nonlinear mapping between corresponding views of objects, filling in missing data as needed to synthesize novel views. We then show that the method can learn a mapping between human degrees of freedom and robotic degrees of freedom for a humanoid robot, allowing robotic imitation of human poses from motion capture data. Aaron P. Shon, Keith Grochow, Aaron Hertzmann, Rajesh P. N. Rao |
NIPS | 4 |
| 2005 | Goal-Based Imitation as Probabilistic Inference over Graphical ModelsabstractHumans are extremely adept at learning new skills by imitating the actions of others. A progression of imitative abilities has been observed in children, ranging from imitation of simple body movements to goalbased imitation based on inferring intent. In this paper, we show that the problem of goal-based imitation can be formulated as one of inferring goals and selecting actions using a learned probabilistic graphical model of the environment. We first describe algorithms for planning actions to achieve a goal state using probabilistic inference. We then describe how planning can be used to bootstrap the learning of goal-dependent policies by utilizing feedback from the environment. The resulting graphical model is then shown to be powerful enough to allow goal-based imitation. Using a simple maze navigation task, we illustrate how an agent can infer the goals of an observed teacher and imitate the teacher even when the goals are uncertain and the demonstration is incomplete. Rajesh P. N. Rao |
NIPS | 2 |
| 2005 | Learning temporal clusters with synaptic facilitation and lateral inhibition
Chris L. Baker, Aaron P. Shon, Rajesh P. N. Rao |
Neurocomputing | 3 |
| 2005 | Implementing belief propagation in neural circuits
Aaron P. Shon, Rajesh P. N. Rao |
Neurocomputing | 2 |
| 2005 | Bilinear Sparse Coding for Invariant VisionabstractRecent algorithms for sparse coding and independent component analysis (ICA) have demonstrated how localized features can be learned from natural images. However, these approaches do not take image transformations into account. We describe an unsupervised algorithm for learning both localized features and their transformations directly from images using a sparse bilinear generative model. We show that from an arbitrary set of natural images, the algorithm produces oriented basis filters that can simultaneously represent features in an image and their transformations. The learned generative model can be used to translate features to different locations, thereby reducing the need to learn the same feature at multiple locations, a limitation of previous approaches to sparse coding and ICA. Our results suggest that by explicitly modeling the interaction between local image features and their transformations, the sparse bilinear approach can provide a basis for achieving transformation-invariant vision. David B. Grimes, Rajesh P. N. Rao |
Neural Comput. | 2 |
| 2004 | Hierarchical Bayesian Inference in Networks of Spiking NeuronsabstractThere is growing evidence from psychophysical and neurophysiological studies that the brain utilizes Bayesian principles for inference and de- cision making. An important open question is how Bayesian inference for arbitrary graphical models can be implemented in networks of spik- ing neurons. In this paper, we show that recurrent networks of noisy integrate-and-fire neurons can perform approximate Bayesian inference for dynamic and hierarchical graphical models. The membrane potential dynamics of neurons is used to implement belief propagation in the log domain. The spiking probability of a neuron is shown to approximate the posterior probability of the preferred state encoded by the neuron, given past inputs. We illustrate the model using two examples: (1) a motion de- tection network in which the spiking probability of a direction-selective neuron becomes proportional to the posterior probability of motion in a preferred direction, and (2) a two-level hierarchical network that pro- duces attentional effects similar to those observed in visual cortical areas V2 and V4. The hierarchical model offers a new Bayesian interpretation of attentional modulation in V2 and V4. Rajesh P. N. Rao |
NIPS | 1 |
| 2004 | Dynamic Bayesian Networks for Brain-Computer InterfacesabstractWe describe an approach to building brain-computer interfaces (BCI) based on graphical models for probabilistic inference and learning. We show how a dynamic Bayesian network (DBN) can be used to infer probability distributions over brain- and body-states during planning and execution of actions. The DBN is learned directly from observed data and allows measured signals such as EEG and EMG to be interpreted in terms of internal states such as intent to move, preparatory activity, and movement execution. Unlike traditional classification-based approaches to BCI, the proposed approach (1) allows continuous tracking and predic- tion of internal states over time, and (2) generates control signals based on an entire probability distribution over states rather than binary yes/no decisions. We present preliminary results of brain- and body-state es- timation using simultaneous EEG and EMG signals recorded during a self-paced left/right hand movement task. Pradeep Shenoy, Rajesh P. N. Rao |
NIPS | 2 |
| 2004 | Bayesian Computation in Recurrent Neural CircuitsabstractA large number of human psychophysical results have been successfully explained in recent years using Bayesian models. However, the neural implementation of such models remains largely unclear. In this article, we show that a network architecture commonly used to model the cerebral cortex can implement Bayesian inference for an arbitrary hidden Markov model. We illustrate the approach using an orientation discrimination task and a visual motion detection task. In the case of orientation discrimination, we show that the model network can infer the posterior distribution over orientations and correctly estimate stimulus orientation in the presence of significant noise. In the case of motion detection, we show that the resulting model network exhibits direction selectivity and correctly computes the posterior probabilities over motion direction and position. When used to solve the well-known random dots motion discrimination task, the model generates responses that mimic the activities of evidence-accumulating neurons in cortical areas LIP and FEF. The framework we introduce posits a new interpretation of cortical activities in terms of log posterior probabilities of stimuli occurring in the natural world. Rajesh P. N. Rao |
Neural Comput. | 1 |
| 2003 | Probabilistic Bilinear Models for Appearance-Based VisionabstractWe present a probabilistic approach to learning object representations based on the "content and style" bilinear generative model of Tenenbaum and Freeman. In contrast to their earlier SVD-based approach, our approach models images using particle filters. We maintain separate particle filters to represent the content and style spaces, allowing us to define arbitrary weighting functions over the particles to help estimate the content/style densities. We combine this approach with a new EM-based method for learning basis vectors that describe content-style mixing. Using a particle-based representation permits good reconstruction despite reduced dimensionality, and increases storage capacity and computational efficiency. We describe how learning the distributions using particle filters allows us to efficiently compute a probabilistic "novelty" term. Our example application considers a dataset of faces under different lighting conditions. The system classifies faces of people it has seen before, and can identify previously unseen faces as new content. Using a probabilistic definition of novelty in conjunction with learning content-style separability provides a crucial building block for designing real-world, real-time object recognition systems. David B. Grimes, Aaron P. Shon, Rajesh P. N. Rao |
ICCV | 3 |
| 2003 | Learning temporal patterns by redistribution of synaptic efficacy
Aaron P. Shon, Rajesh P. N. Rao |
Neurocomputing | 2 |
| 2002 | A Bilinear Model for Sparse CodingabstractRecent algorithms for sparse coding and independent component analy- sis (ICA) have demonstrated how localized features can be learned from natural images. However, these approaches do not take image transfor- mations into account. As a result, they produce image codes that are redundant because the same feature is learned at multiple locations. We describe an algorithm for sparse coding based on a bilinear generative model of images. By explicitly modeling the interaction between im- age features and their transformations, the bilinear approach helps reduce redundancy in the image code and provides a basis for transformation- invariant vision. We present results demonstrating bilinear sparse coding of natural images. We also explore an extension of the model that can capture spatial relationships between the independent features of an ob- ject, thereby providing a new framework for parts-based object recogni- tion. David B. Grimes, Rajesh P. N. Rao |
NIPS | 2 |
| 2001 | Optimal Smoothing in Visual Motion PerceptionabstractWhen a flash is aligned with a moving object, subjects perceive the flash to lag behind the moving object. Two different models have been proposed to explain this "flash-lag" effect. In the motion extrapolation model, the visual system extrapolates the location of the moving object to counteract neural propagation delays, whereas in the latency difference model, it is hypothesized that moving objects are processed and perceived more quickly than flashed objects. However, recent psychophysical experiments suggest that neither of these interpretations is feasible (Eagleman & Sejnowski, 2000a, 2000b, 2000c), hypothesizing instead that the visual system uses data from the future of an event before committing to an interpretation. We formalize this idea in terms of the statistical framework of optimal smoothing and show that a model based on smoothing accounts for the shape of psychometric curves from a flash-lag experiment involving random reversals of motion direction. The smoothing model demonstrates how the visual system may enhance perceptual accuracy by relying not only on data from the past but also on data collected from the immediate future of an event. Rajesh P. N. Rao, David M. Eagleman, Terrence J. Sejnowski |
Neural Comput. | 1 |
| 2001 | Spike-Timing-Dependent Hebbian Plasticity as Temporal Difference LearningabstractA spike-timing-dependent Hebbian mechanism governs the plasticity of recurrent excitatory synapses in the neocortex: synapses that are activated a few milliseconds before a postsynaptic spike are potentiated, while those that are activated a few milliseconds after are depressed. We show that such a mechanism can implement a form of temporal difference learning for prediction of input sequences. Using a biophysical model of a cortical neuron, we show that a temporal difference rule used in conjunction with dendritic backpropagating action potentials reproduces the temporally asymmetric window of Hebbian plasticity observed physio-logically. Furthermore, the size and shape of the window vary with the distance of the synapse from the soma. Using a simple example, we show how a spike-timing-based temporal difference learning rule can allow a network of neocortical neurons to predict an input a few milliseconds before the input's expected arrival. Rajesh P. N. Rao, Terrence J. Sejnowski |
Neural Comput. | 1 |
| 2000 | A single-spike model of predictive codingabstractThe standard cortical model assumes that the meaning of a neuron's signal is contained in its firing rate. While that model has been used to interpret a voluminous amount of experimental data, it does not address the question of timing, or how recipient neurons can decode this signal in time to predict behavioral results. We propose a model based on coincident firing of large groups of neurons. We show using an example of predictive coding, how the cortex can support vast amounts of non-interfering parallel computation. Dana H. Ballard, Rajesh P. N. Rao, Zuohua Zhang |
Neurocomputing | 2 |
| 2000 | Corrigendum to "Upward separation for FewP and related classes"
Rajesh P. N. Rao, Jörg Rothe, Osamu Watanabe 0001 |
Inf. Process. Lett. | 1 |
| 1999 | Predictive Sequence Learning in Recurrent Neocortical Circuits
Rajesh P. N. Rao, Terrence J. Sejnowski |
NIPS | 1 |
| 1998 | Learning Lie Groups for Invariant Visual Perception
Rajesh P. N. Rao, Daniel L. Ruderman |
NIPS | 1 |
| 1998 | Hierarchical Learning of Navigational Behaviors in an Autonomous Robot using a Predictive Sparse Distributed Memory
Rajesh P. N. Rao, Olac Fuentes |
Mach. Learn. | 1 |
| 1997 | Dynamic Appearance-Based RecognitionabstractWe describe a hierarchical appearance-based method for learning, recognizing, and predicting arbitrary spatiotemporal sequences of images. The method, which implements a robust hierarchical form of the Kalman filter derived from the Minimum Description Length (MDL) principle, includes as a special case several well-known object encoding techniques including eigenspace methods for static recognition. Successive levels of the hierarchical filter implement dynamic models operating over successively larger spatial and temporal scales. Each hierarchical level predicts the recognition state at a lower level and modifies its own recognition state using the residual error between the prediction and the actual lower-level state. Simultaneously, on a longer time scale, the filter learns an internal model of input dynamics by adapting its generative and state transition matrices at each level to minimize prediction errors. The resulting prediction/learning scheme thereby implements an on-line form of the well-known Expectation-Maximization (EM) algorithm from statistics. We present experimental results demonstrating the method's efficacy in mediating robust spatiotemporal recognition in a variety of scenarios containing varying degrees of occlusions and clutter. Rajesh P. N. Rao |
CVPR | 1 |
| 1997 | Correlates of Attention in a Model of Dynamic Visual Recognition
Rajesh P. N. Rao |
NIPS | 1 |
| 1997 | Dynamic Model of Visual Recognition Predicts Neural Response Properties in the Visual CortexabstractThe responses of visual cortical neurons during fixation tasks can be significantly modulated by stimuli from beyond the classical receptive field. Modulatory effects in neural responses have also been recently reported in a task where a monkey freely views a natural scene. In this article, we describe a hierarchical network model of visual recognition that explains these experimental observations by using a form of the extended Kalman filter as given by the minimum description length (MDL) principle. The model dynamically combines input-driven bottom-up signals with expectation-driven top-down signals to predict current recognition state. Synaptic weights in the model are adapted in a Hebbian manner according to a learning rule also derived from the MDL principle. The resulting prediction-learning scheme can be viewed as implementing a form of expectation-maximization (EM) algorithm. The architecture of the model posits an active computational role of the reciprocal connections between adjoining visual cortical areas in determining neural response properties. In particular, the model demonstrates the possible role of feedback from higher cortical areas in mediating neurophysiological effects due to stimuli from beyond the classical receptive field. Simulations of the model are provided that help explain the experimental observations regarding neural responses in both free viewing and fixation conditions. Rajesh P. N. Rao, Dana H. Ballard |
Neural Comput. | 1 |
| 1995 | Object Indexing Using an Iconic Sparse Distributed MemoryabstractA general-purpose object indexing technique is described that combines the virtues of principal component analysis with the favorable matching properties of high-dimensional spaces to achieve high-precision recognition. An object is represented by a set of high-dimensional iconic feature vectors comprised of the responses of derivatives of Gaussian filters at a range of orientations and scales. Since these filters can be shown to form the eigenvectors of arbitrary images containing both natural and man-made structures, they are well-suited for indexing in disparate domains. The indexing algorithm uses an active vision system in conjunction with a modified form of Kanerva's (1988, 1993) sparse distributed memory which facilitates interpolation between views and provides a convenient platform for learning the association between an object's appearance and its identity. The robustness of the indexing method was experimentally confirmed by subjecting the method to a range of viewing conditions and the accuracy was verified using a well-known model database containing a number of complex 3D objects under varying pose.> Rajesh P. N. Rao, Dana H. Ballard |
ICCV | 1 |
| 1995 | Natural Basis Functions and Topographic Memory for Face Recognition
Rajesh P. N. Rao, Dana H. Ballard |
IJCAI | 1 |
| 1995 | Modeling Saccadic Targeting in Visual Search
Rajesh P. N. Rao, Gregory J. Zelinsky, Mary M. Hayhoe, Dana H. Ballard |
NIPS | 1 |
| 1995 | An Active Vision Architecture Based on Iconic RepresentationsabstractActive vision systems have the capability of continuously interacting with the environment. The rapidly changing environment of such systems means that it is attractive to replace static representations with visual routines that compute information on demand. Such routines place a premium on image data structures that are easily computed and used. The purpose of this paper is to propose a general active vision architecture based on efficiently computable iconic representations. This architecture employs two primary visual routines, one for identifying the visual image near the fovea (object identification), and another for locating a stored prototype on the retina (object location). This design allows complex visual behaviors to be obtained by composing these two routines with different parameters. The iconic representations are comprised of high-dimensional feature vectors obtained from the responses of an ensemble of Gaussian derivative spatial filters at a number of orientations and scales. These representations are stored in two separate memories. One memory is indexed by image coordinates while the other is indexed by object coordinates. Object location matches a localized set of model features with image features at all possible retinal locations. Object identification matches a foveal set of image features with all possible model features. We present experimental results for a near real-time implementation of these routines on a pipeline image processor and suggest relatively simple strategies for tackling the problems of occlusions and scale variations. We also discuss two additional visual routines, one for top-down foveal targeting using log-polar sensors and another for looming detection, which are facilitated by the proposed architecture. Rajesh P. N. Rao, Dana H. Ballard |
Artif. Intell. | 1 |
| 1995 | A Note on P-Selective Sets and ClosenessabstractWe investigate the class of sets that form sparse symmetric differences with P-selective sets. Intuitively, this class (denoted by PSEL-close) comprises of sets that can, in a certain sense, be approximated by P-selective sets. A primary motivation behind the introduction of this new class is to unify the separate approaches that have been undertaken for sparse and P-selective sets. In order to establish PSEL-close as a distinct class, we first prove a theorem separating it from both the encompassing class of P/poly and the subclasses of P-selective and sparse sets. We then prove that no ⩽mp-hard set for E can be in PSEL-close. The proof of this theorem relies on techniques from the work of Berman and Hartmanis (1977) and Schöning (1986), and generalizes their results in a straightforward manner. Rajesh P. N. Rao |
Inf. Process. Lett. | 1 |
| 1994 | Seeing Behind Occlusions
Dana H. Ballard, Rajesh P. N. Rao |
ECCV (1) | 2 |
| 1994 | Learning Saccadic Eye Movements Using Multiscale Spatial FiltersabstractWe describe a framework for learning saccadic eye movements using a photometric representation of target points in natural scenes. The rep(cid:173) resentation takes the form of a high-dimensional vector comprised of the responses of spatial filters at different orientations and scales. We first demonstrate the use of this response vector in the task of locating pre(cid:173) viously foveated points in a scene and subsequently use this property in a multisaccade strategy to derive an adaptive motor map for delivering accurate saccades. Rajesh P. N. Rao, Dana H. Ballard |
NIPS | 1 |
| 1994 | Upward Separation for FewP and Related Classes
Rajesh P. N. Rao, Jörg Rothe, Osamu Watanabe 0001 |
Inf. Process. Lett. | 1 |