EDBT 2026 Demo / reviewers in the wild / expert
Dana H. Ballard
dblp:68/6429
· DBLP profile ↗
91ranked-venue papers
24as first author
1since 2021 · last 2021
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 76 · 18 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 13 first-authorApplied, interdisciplinary, general and emerging computing · 7 · 2 first-authorSystems, architecture and hardware · 6 · 1 first-authorHuman-computer interaction and ubiquitous computing · 4Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
41 papers |
Reinforcement learning · 60% Trustworthy machine learning · 17% Motion planning and robot control · 6% | |
| Human-computer interaction and pervasive computing
7 papers |
Wearable and physiological sensing · 62% Human-AI interaction · 35% Human-robot interaction · 2% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% Medical and health informatics · 0% |
Topics — the 30 heaviest of 98, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
imitation learning |
1.1 | 3 | 2020 | Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset · AAAI 2020 Leveraging Human Guidance for Deep Reinforcement Learning Tasks · IJCAI 2019 Learning Attention Model From Human for Visuomotor Tasks · AAAI 2018 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.8 | 2 | 2021 | Machine versus Human Attention in Deep Reinforcement Learning Tasks · NeurIPS 2021 Learning Attention Model From Human for Visuomotor Tasks · AAAI 2018 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | Machine versus Human Attention in Deep Reinforcement Learning Tasks · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › interpretability › visual explanation
saliency analysis |
0.5 | 1 | 2021 | Machine versus Human Attention in Deep Reinforcement Learning Tasks · NeurIPS 2021 |
Wearable and physiological sensing › eye tracking › gaze-based interaction
gaze-assisted interaction |
0.4 | 1 | 2020 | Human Gaze Assisted Artificial Intelligence: A Review · IJCAI 2020 |
Wearable and physiological sensing › eye tracking
gaze behavior modeling |
0.4 | 1 | 2020 | Human Gaze Assisted Artificial Intelligence: A Review · IJCAI 2020 |
Human-AI interaction
human-centered AI |
0.4 | 1 | 2020 | Human Gaze Assisted Artificial Intelligence: A Review · IJCAI 2020 |
Robotics › Motion planning and robot control › robot learning
visuomotor learning |
0.3 | 1 | 2018 | AGIL: Learning Attention from Human for Visuomotor Tasks · ECCV (11) 2018 |
Knowledge, reasoning and agents › Multi-agent systems
heterogeneous multi-agent systems |
0.2 | 1 | 2016 | Decision-Making Policies for Heterogeneous Autonomous Multi-Agent Systems with Safety Constraints · IJCAI 2016 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.2 | 1 | 2016 | Decision-Making Policies for Heterogeneous Autonomous Multi-Agent Systems with Safety Constraints · IJCAI 2016 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
safe multi-agent reinforcement learning |
0.2 | 1 | 2016 | Decision-Making Policies for Heterogeneous Autonomous Multi-Agent Systems with Safety Constraints · IJCAI 2016 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
modular reinforcement learning |
0.2 | 1 | 2015 | Global Policy Construction in Modular Reinforcement Learning · AAAI 2015 |
Machine learning › Reinforcement learning
temporal difference learning |
0.2 | 1 | 2015 | Global Policy Construction in Modular Reinforcement Learning · AAAI 2015 |
Computer vision › Video understanding and tracking
motion representation |
0.2 | 1 | 2014 | Efficient Codes for Inverse Dynamics During Walking · AAAI 2014 |
Bioinformatics and computational biology
computational neuroscience |
0.2 | 1 | 2014 | Efficient Codes for Inverse Dynamics During Walking · AAAI 2014 |
Bioinformatics and computational biology › computational neuroscience › neural coding
efficient coding |
0.2 | 1 | 2014 | Efficient Codes for Inverse Dynamics During Walking · AAAI 2014 |
Wearable and physiological sensing
eye tracking |
0.1 | 1 | 2020 | Atari-HEAD: Atari Human Eye-Tracking and Demonstration Dataset · AAAI 2020 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.1 | 1 | 2018 | AGIL: Learning Attention from Human for Visuomotor Tasks · ECCV (11) 2018 |
Computer vision › Image recognition and object detection
saliency prediction |
0.1 | 1 | 2018 | Learning Attention Model From Human for Visuomotor Tasks · AAAI 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.0 | 1 | 2004 | On the Integration of Grounding Language and Learning Objects · AAAI 2004 |
Machine learning › Reinforcement learning › goal-conditioned reinforcement learning
multi-goal reinforcement learning |
0.0 | 1 | 2003 | Multiple-Goal Reinforcement Learning with Modular Sarsa(0) · IJCAI 2003 |
Machine learning › Reinforcement learning
oculomotor control |
0.0 | 1 | 2003 | Eye Movements for Reward Maximization · NIPS 2003 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › decision making under uncertainty
value of information |
0.0 | 1 | 2003 | Eye Movements for Reward Maximization · NIPS 2003 |
Immersive interaction
virtual reality |
0.0 | 1 | 1999 | Recognizing Evoked Potentials in a Virtual Environment · NIPS 1999 |
Information retrieval
indexing |
0.0 | 1 | 1998 | Phonetic Set Indexing for Fast Lexical Access · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Information retrieval
search engines |
0.0 | 1 | 1998 | Phonetic Set Indexing for Fast Lexical Access · IEEE Trans. Pattern Anal. Mach. Intell. 1998 |
Computer vision › Image recognition and object detection
object recognition |
0.0 | 2 | 1995 | Object Indexing Using an Iconic Sparse Distributed Memory · ICCV 1995 Viewer Independent Shape Recognition · IEEE Trans. Pattern Anal. Mach. Intell. 1983 |
Robotics › Robot manipulation
autonomous manipulation |
0.0 | 2 | 1995 | Teleassistance: Contextual Guidance for Autonomous Manipulation · AAAI 1994 Remote Teleassistance · ICRA 1995 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › subspace learning
basis learning |
0.0 | 1 | 1995 | Natural Basis Functions and Topographic Memory for Face Recognition · IJCAI 1995 |
Computer vision › Face, body and person analysis
face recognition |
0.0 | 1 | 1995 | Natural Basis Functions and Topographic Memory for Face Recognition · IJCAI 1995 |
Methods — techniques the papers use, named apart from their topics
visual attention model · 1.0saliency map · 1.0gaze prediction · 0.9survey · 0.4imitation learning · 0.3gaze data collection · 0.3deep neural network · 0.3attention learning · 0.3reinforcement learning · 0.3policy learning · 0.2supervised regression · 0.2sparse coding · 0.2motion capture · 0.2inverse kinematics · 0.2uncertainty estimation · 0.0eye tracking · 0.0sparse distributed memory · 0.0principal component analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Machine versus Human Attention in Deep Reinforcement Learning TasksabstractDeep reinforcement learning (RL) algorithms are powerful tools for solving visuomotor decision tasks. However, the trained models are often difficult to interpret, because they are represented as end-to-end deep neural networks. In this paper, we shed light on the inner workings of such trained models by analyzing the pixels that they attend to during task execution, and comparing them with the pixels attended to by humans executing the same tasks. To this end, we investigate the following two questions that, to the best of our knowledge, have not been previously studied. 1) How similar are the visual representations learned by RL agents and humans when performing the same task? and, 2) How do similarities and differences in these learned representations explain RL agents' performance on these tasks? Specifically, we compare the saliency maps of RL agents against visual attention models of human experts when learning to play Atari games. Further, we analyze how hyperparameters of the deep RL algorithm affect the learned representations and saliency maps of the trained agents. The insights provided have the potential to inform novel algorithms for closing the performance gap between human experts and RL agents. Sihang Guo, Bo Liu 0042, Dana H. Ballard, Mary M. Hayhoe, Peter Stone 0001 |
NeurIPS | 5 |
| 2020 | Atari-HEAD: Atari Human Eye-Tracking and Demonstration DatasetabstractLarge-scale public datasets have been shown to benefit research in multiple areas of modern artificial intelligence. For decision-making research that requires human data, high-quality datasets serve as important benchmarks to facilitate the development of new methods by providing a common reproducible standard. Many human decision-making tasks require visual attention to obtain high levels of performance. Therefore, measuring eye movements can provide a rich source of information about the strategies that humans use to solve decision-making tasks. Here, we provide a large-scale, high-quality dataset of human actions with simultaneously recorded eye movements while humans play Atari video games. The dataset consists of 117 hours of gameplay data from a diverse set of 20 games, with 8 million action demonstrations and 328 million gaze samples. We introduce a novel form of gameplay, in which the human plays in a semi-frame-by-frame manner. This leads to near-optimal game decisions and game scores that are comparable or better than known human records. We demonstrate the usefulness of the dataset through two simple applications: predicting human gaze and imitating human demonstrated actions. The quality of the data leads to promising results in both tasks. Moreover, using a learned human gaze model to inform imitation learning leads to an 115% increase in game performance. We interpret these results as highlighting the importance of incorporating human visual attention in models of decision making and demonstrating the value of the current dataset to the research community. We hope that the scale and quality of this dataset can provide more opportunities to researchers in the areas of visual attention, imitation learning, and reinforcement learning. Calen Walshe, Zhuode Liu, Lin Guan 0003, Karl S. Muller, Jake Alden Whritner, Luxin Zhang, Mary M. Hayhoe, Dana H. Ballard |
AAAI | 9 |
| 2020 | Human Gaze Assisted Artificial Intelligence: A ReviewabstractHuman gaze reveals a wealth of information about internal cognitive state. Thus, gaze-related research has significantly increased in computer vision, natural language processing, decision learning, and robotics in recent years. We provide a high-level overview of the research efforts in these fields, including collecting human gaze data sets, modeling gaze behaviors, and utilizing gaze information in various applications, with the goal of enhancing communication between these research areas. We discuss future challenges and potential applications that work towards a common goal of human-centered artificial intelligence. Akanksha Saran, Bo Liu 0042, Sihang Guo, Scott Niekum, Dana H. Ballard, Mary M. Hayhoe |
IJCAI | 7 |
| 2020 | Parallel Neural Multiprocessing with Gamma Frequency LatenciesabstractThe Poisson variability in cortical neural responses has been typically modeled using spike averaging techniques, such as trial averaging and rate coding, since such methods can produce reliable correlates of behavior. However, mechanisms that rely on counting spikes could be slow and inefficient and thus might not be useful in the brain for computations at timescales in the 10 millisecond range. This issue has motivated a search for alternative spike codes that take advantage of spike timing and has resulted in many studies that use synchronized neural networks for communication. Here we focus on recent studies that suggest that the gamma frequency may provide a reference that allows local spike phase representations that could result in much faster information transmission. We have developed a unified model (gamma spike multiplexing) that takes advantage of a single cycle of a cell's somatic gamma frequency to modulate the generation of its action potentials. An important consequence of this coding mechanism is that it allows multiple independent neural processes to run in parallel, thereby greatly increasing the processing capability of the cortex. System-level simulations and preliminary analysis of mouse cortical cell data are presented as support for the proposed theoretical model. Dana H. Ballard |
Neural Comput. | 2 |
| 2019 | Leveraging Human Guidance for Deep Reinforcement Learning TasksabstractReinforcement learning agents can learn to solve sequential decision tasks by interacting with the environment. Human knowledge of how to solve these tasks can be incorporated using imitation learning, where the agent learns to imitate human demonstrated decisions. However, human guidance is not limited to the demonstrations. Other types of guidance could be more suitable for certain tasks and require less human effort. This survey provides a high-level overview of five recent learning frameworks that primarily rely on human guidance other than conventional, step-by-step action demonstrations. We review the motivation, assumption, and implementation of each framework. We then discuss possible future research directions. Faraz Torabi, Lin Guan 0003, Dana H. Ballard, Peter Stone 0001 |
IJCAI | 4 |
| 2018 | Learning Attention Model From Human for Visuomotor TasksabstractA wealth of information regarding intelligent decision making is conveyed by human gaze and visual attention, hence, modeling and exploiting such information might be a promising way to strengthen algorithms like deep reinforcement learning. We collect high-quality human action and gaze data while playing Atari games. Using these data, we train a deep neural network that can predict human gaze positions and visual attention with high accuracy. Luxin Zhang, Zhuode Liu, Mary M. Hayhoe, Dana H. Ballard |
AAAI | 5 |
| 2018 | AGIL: Learning Attention from Human for Visuomotor Tasks
Zhuode Liu, Luxin Zhang, Jake Alden Whritner, Karl S. Muller, Mary M. Hayhoe, Dana H. Ballard |
ECCV (11) | 7 |
| 2018 | Modeling sensory-motor decisions in natural behaviorabstractAlthough a standard reinforcement learning model can capture many aspects of reward-seeking behaviors, it may not be practical for modeling human natural behaviors because of the richness of dynamic environments and limitations in cognitive resources. We propose a modular reinforcement learning model that addresses these factors. Based on this model, a modular inverse reinforcement learning algorithm is developed to estimate both the rewards and discount factors from human behavioral data, which allows predictions of human navigation behaviors in virtual reality with high accuracy across different subjects and with different tasks. Complex human navigation trajectories in novel environments can be reproduced by an artificial agent that is based on the modular model. This model provides a strategy for estimating the subjective value of actions and how they influence sensory-motor decisions in natural behavior. Matthew H. Tong, Yuchen Cui, Constantin A. Rothkopf, Dana H. Ballard, Mary M. Hayhoe |
PLoS Comput. Biol. | 6 |
| 2016 | Decision-Making Policies for Heterogeneous Autonomous Multi-Agent Systems with Safety Constraints
Yue Yu 0004, Mahmoud El Chamie, Behçet Açikmese, Dana H. Ballard |
IJCAI | 5 |
| 2015 | Global Policy Construction in Modular Reinforcement LearningabstractWe propose a modular reinforcement learning algorithm which decomposes a Markov decision process into independent modules. Each module is trained using Sarsa(lambda). We introduce three algorithms for forming global policy from modules policies, and demonstrate our results using a 2D grid world. Zhao Song 0002, Dana H. Ballard |
AAAI | 3 |
| 2014 | Efficient Codes for Inverse Dynamics During WalkingabstractEfficient codes have been used effectively in both computer science and neuroscience to better understand the information processing in visual and auditory encoding and discrimination tasks. In this paper, we explore the use of efficient codes for representing information relevant to human movements during locomotion. Specifically, we apply motion capture data to a physical model of the human skeleton to compute joint angles (inverse kinematics) and joint torques (inverse dynamics); then, by treating the resulting paired dataset as a supervised regression problem, we investigate the effect of sparsity in mapping from angles to torques. The results of our investigation suggest that sparse codes can indeed represent salient features of both the kinematic and dynamic views of human locomotion movements. However, sparsity appears to be only one parameter in building a model of inverse dynamics; we also show that the "encoding" process benefits significantly by integrating with the "regression" process for this task. In addition, we show that, for this task, simple coding and decoding methods are not sufficient to model the extremely complex inverse dynamics mapping. Finally, we use our results to argue that representations of movement are critical to modeling and understanding these movements. Leif Johnson, Dana H. Ballard |
AAAI | 2 |
| 2014 | Classifying movements using efficient kinematic codes
Leif Johnson, Dana H. Ballard |
CogSci | 2 |
| 2013 | Efficient codes for multi-modal pose regression
Leif Johnson, Joseph L. Cooper, Dana H. Ballard |
CogSci | 3 |
| 2013 | A soft barrier model for predicting human visuomotor behavior in a driving task
Leif Johnson, Brian T. Sullivan, Mary M. Hayhoe, Dana H. Ballard |
CogSci | 4 |
| 2013 | Unified losses for multi-modal pose coding and regressionabstractSparsity and redundancy reduction have been shown to be useful in machine learning, but empirical evaluation has been performed primarily on classification tasks using datasets of natural images and sounds. Similarly, the performance of unsupervised feature learning followed by supervised fine-tuning has primarily focused on classification tasks. In comparison, relatively little work has investigated the use of sparse codes for representing human movements and poses, or for using these codes in regression tasks with movement data. This paper defines a basic coding and regression architecture for evaluating the impact of sparsity when coding human pose information, and tests the performance of several coding methods within this framework for the task of mapping from a kinematic (joint angle) modality to a dynamic (joint torque) one. In addition, we evaluate the performance of unified loss functions defined on the same class of models. We show that, while sparse codes are useful for effective mappings between modalities, their primary benefit for this task seems to be in admitting overcomplete codebooks. We make use of the proposed architecture to examine in detail the sources of error for each stage in the model under various coding strategies. Furthermore, we show that using a unified loss function that passes gradient information between stages of the coding and regression architecture provides substantial reductions in overall error. Leif Johnson, Joseph L. Cooper, Dana H. Ballard |
IJCNN | 3 |
| 2012 | Realtime, Physics-Based Marker Following
Joseph L. Cooper, Dana H. Ballard |
MIG | 2 |
| 2011 | Novelty detection using Growing Neural Gas for visuo-spatial memoryabstractDetecting visual changes in environments is an important computation with many applications in robotics and computer vision. Security cameras, remotely operated vehicles, and sentry robots could all benefit from robust change detection capability. We conjecture that if one has a mobile camera system the number of visual scenes that are experienced is limited (compared to the space of all possible scenes) and that the scenes do not frequently undergo major changes between observations. These assumptions can be exploited to ease the task of change detection and reduce the computational complexity of processing visual information by utilizing memory to store previous computations. We demonstrate a method to learn the distribution of visual features in an environment via a self-organizing map. Additionally, the spatial distribution of these features can be learned if a positional signal is available. Our method uses a low dimensional representation of visual features to rapidly detect changes in current visual inputs. The model encodes spatially-distributed color histograms of real world visual scenes captured by a camera moved through an environment. The distribution of the color histograms is learned using a self-organizing map with location (when available) and color data. We present tests of the model on detecting changes in an indoor environment. Dmitry Kit, Brian T. Sullivan, Dana H. Ballard |
IROS | 3 |
| 2009 | Predictive Feedback Can Account for Biphasic Responses in the Lateral Geniculate NucleusabstractBiphasic neural response properties, where the optimal stimulus for driving a neural response changes from one stimulus pattern to the opposite stimulus pattern over short periods of time, have been described in several visual areas, including lateral geniculate nucleus (LGN), primary visual cortex (V1), and middle temporal area (MT). We describe a hierarchical model of predictive coding and simulations that capture these temporal variations in neuronal response properties. We focus on the LGN-V1 circuit and find that after training on natural images the model exhibits the brain's LGN-V1 connectivity structure, in which the structure of V1 receptive fields is linked to the spatial alignment and properties of center-surround cells in the LGN. In addition, the spatio-temporal response profile of LGN model neurons is biphasic in structure, resembling the biphasic response structure of neurons in cat LGN. Moreover, the model displays a specific pattern of influence of feedback, where LGN receptive fields that are aligned over a simple cell receptive field zone of the same polarity decrease their responses while neurons of opposite polarity increase their responses with feedback. This phase-reversed pattern of influence was recently observed in neurophysiology. These results corroborate the idea that predictive feedback is a general coding strategy in the brain. Janneke F. M. Jehee, Dana H. Ballard |
PLoS Comput. Biol. | 2 |
| 2007 | A unified model of early word learning: Integrating statistical and social cues
Chen Yu 0001, Dana H. Ballard |
Neurocomputing | 2 |
| 2007 | Modeling embodied visual behaviorsabstractTo make progess in understanding human visuomotor behavior, we will need to understand its basic components at an abstract level. One way to achieve such an understanding would be to create a model of a human that has a sufficient amount of complexity so as to be capable of generating such behaviors. Recent technological advances have been made that allow progress to be made in this direction. Graphics models that simulate extensive human capabilities can be used as platforms from which to develop synthetic models of visuomotor behavior. Currently, such models can capture only a small portion of a full behavioral repertoire, but for the behaviors that they do model, they can describe complete visuomotor subsystems at a useful level of detail. The value in doing so is that the body's elaborate visuomotor structures greatly simplify the specification of the abstract behaviors that guide them. The net result is that, essentially, one is faced with proposing an embodied “operating system” model for picking the right set of abstract behaviors at each instant. This paper outlines one such model. A centerpiece of the model uses vision to aid the behavior that has the most to gain from taking environmental measurements. Preliminary tests of the model against human performance in realistic VR environments show that main features of the model show up in human behavior. Nathan Sprague, Dana H. Ballard, Al Robinson |
ACM Trans. Appl. Percept. | 2 |
| 2006 | Behavior Recognition in Human Object Interactions with a Task ModelabstractBehavior recognition can be greatly levitated by a model of the task with which the behavior is associated. We present a Markovian task model that captures the temporal relations between subtasks, and provides prior information that helps behavior recognition from simple and noisy sensory data. A Dynamic Bayesian Network is used to integrate multi-model evidence and infer underlying task states. Experiments demonstrate that this system can recognize human behavior in everyday tasks such as making sandwiches. Weilie Yi, Dana H. Ballard |
AVSS | 2 |
| 2006 | Motor Synergies for Coordinated Movements in HumanoidsabstractSynthesizing automatons whole body movements is difficult, especially in humanoids with as high Degrees of Freedoms (DOFs) as humans. Based on a biologically inspired control model we proposed, three different ways of motor synergies over multiple motor routines are discussed to compose complex movements in a 33 DOF humanoid. Motor routine is a movement unit which implements a functional task and only involves those active joints participating in the task. We demonstrate the humanoid doing coordinated walking, sitting and rising, reaching, object manipulation etc. Xue Gu, Dana H. Ballard |
IROS | 2 |
| 2006 | Modeling the Brain's Operating System Using Virtual HumanoidsabstractMuch of the allocation of human resources to tasks is studied under the rubric of "attention". However this is a very low-dimensional characterization of a system that has many degrees of freedom. To make progess in understanding human brain resource allocations, we will need to understand its basic functions at an abstract level. One way of accomplishing such an integration is to create a model of a human that has a useful amount of complexity. Essentially, one is faced with proposing an embodied "operating system" model that can be tested against human performance. Recently, technological advances have been made that allow progress in this direction. Graphic models that simulate extensive human capabilities can be used as platforms to develop synthetic models of visuo-motor behavior. Currently, such models can capture only a small portion of a full behavioral repertoire, but for the behaviors that they do model, they can describe complete visuo-motor subsystems at a level of detail that can be tested against human performance in realistic environments. This paper outlines one such model and shows both that it can produce interesting new hypotheses as to the role of vision and also that it can greatly enhance our understanding of a more multifacted characterization attention in visuo-motor tasks. Dana H. Ballard, Nathan Sprague |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2004 | On the Integration of Grounding Language and Learning Objects
Chen Yu 0001, Dana H. Ballard |
AAAI | 2 |
| 2004 | A single spike model of predictive coding
Zuohua Zhang, Dana H. Ballard |
Neurocomputing | 2 |
| 2004 | A multimodal learning interface for grounding spoken language in sensory perceptionsabstractWe present a multimodal interface that learns words from natural interactions with users. In light of studies of human language development, the learning system is trained in an unsupervised mode in which users perform everyday tasks while providing natural language descriptions of their behaviors. The system collects acoustic signals in concert with user-centric multisensory information from nonspeech modalities, such as user's perspective video, gaze positions, head directions, and hand movements. A multimodal learning algorithm uses this data to first spot words from continuous speech and then associate action verbs and object names with their perceptually grounded meanings. The central ideas are to make use of nonspeech contextual information to facilitate word spotting, and utilize body movements as deictic references to associate temporally cooccurring data from different modalities and build hypothesized lexical items. From those items, an EM-based method is developed to select correct word--meaning pairs. Successful learning is demonstrated in the experiments of three natural tasks: "unscrewing a jar," "stapling a letter," and "pouring water." Chen Yu 0001, Dana H. Ballard |
ACM Trans. Appl. Percept. | 2 |
| 2003 | A multimodal learning interface for word acquisitionabstractWe present a multimodal interface that learns words from natural interactions with users. The system can be trained in an unsupervised mode in which users perform everyday tasks while providing natural language descriptions of their behavior. We collect acoustic signals in concert with user-centric multisensory information from non-speech modalities, such as user's perspective video, gaze positions, head directions and hand movements. A multimodal learning algorithm is developed that firstly spots words from continuous speech and then associates action verbs and object names with their grounded meanings. The central idea is to make use of non-speech contextual information to facilitate word spotting, and utilize temporal correlations of data from different modalities to build hypothesized lexical items. From those items, an EM-based method selects correct word-meaning pairs. Successful learning has been demonstrated in the experiment of the natural task of "stapling papers". Dana H. Ballard, Chen Yu 0001 |
ICASSP (5) | 1 |
| 2003 | A multimodal learning interface for grounding spoken language in sensory perceptionsabstractMost speech interfaces are based on natural language processing techniques that use pre-defined symbolic representations of word meanings and process only linguistic information. To understand and use language like their human counterparts in multimodal human-computer interaction, computers need to acquire spoken language and map it to other sensory perceptions. This paper presents a multimodal interface that learns to associate spoken language with perceptual features by being situated in users' everyday environments and sharing user-centric multisensory information. The learning interface is trained in unsupervised mode in which users perform everyday tasks while providing natural language descriptions of their behaviors. We collect acoustic signals in concert with multisensory information from non-speech modalities, such as user's perspective video, gaze positions, head directions and hand movements. The system firstly estimates users' focus of attention from eye and head cues. Attention, as represented by gaze fixation, is used for spotting the target object of user interest. Attention switches are calculated and used to segment an action sequence into action units which are then categorized by mixture hidden Markov models. A multimodal learning algorithm is developed to spot words from continuous speech and then associate them with perceptually grounded meanings extracted from visual perception and action. Successful learning has been demonstrated in the experiments of three natural tasks: "unscrewing a jar", "stapling a letter" and "pouring water". Chen Yu 0001, Dana H. Ballard |
ICMI | 2 |
| 2003 | Multiple-Goal Reinforcement Learning with Modular Sarsa(0)
Nathan Sprague, Dana H. Ballard |
IJCAI | 2 |
| 2003 | Eye Movements for Reward MaximizationabstractUniversity of Rochester Rochester, NY 14627 [email protected] Recent eye tracking studies in natural tasks suggest that there is a tight link between eye movements and goal directed motor actions. However, most existing models of human eye movements provide a bottom up ac- count that relates visual attention to attributes of the visual scene. The purpose of this paper is to introduce a new model of human eye move- ments that directly ties eye movements to the ongoing demands of be- havior. The basic idea is that eye movements serve to reduce uncertainty about environmental variables that are task relevant. A value is assigned to an eye movement by estimating the expected cost of the uncertainty that will result if the movement is not made. If there are several candidate eye movements, the one with the highest expected value is chosen. The model is illustrated using a humanoid graphic figure that navigates on a sidewalk in a virtual urban environment. Simulations show our protocol is superior to a simple round robin scheduling mechanism. Nathan Sprague, Dana H. Ballard |
NIPS | 2 |
| 2002 | Vision in natural and virtual environmentsabstractOur knowledge of the way that the visual system operates in everyday behavior has, until recently, been very limited. This information is critical not only for understanding visual function, but also for understanding the consequences of various kinds of visual impairment, and for the development of interfaces between human and artificial systems. The development of eye trackers that can be mounted on the head now allows monitoring of gaze without restricting the observer's movements. Observations of natural behavior have demonstrated the highly task-specific and directed nature of fixation patterns, and reveal considerable regularity between observers. Eye, head, and hand coordination also reveals much greater flexibility and task-specificity than previously supposed. Experimental examination of the issues raised by observations of natural behavior requires the development of complex virtual environments that can be manipulated by the experimenter at critical points during task performance. Experiments where we monitored gaze in a simulated driving environment demonstrate that visibility of task relevant information depends critically on active search initiated by the observer according to an internally generated schedule, and this schedule depends on learnt regularities in the environment. In another virtual environment where observers copied toy models we showed that regularities in the spatial structure are used by observers to control eye movement targeting. Other experiments in a virtual environment with haptic feedback show that even simple visual properties like size are not continuously available or processed automatically by the visual system, but are dynamically acquired and discarded according to the momentary task demands. Mary M. Hayhoe, Dana H. Ballard, Jochen Triesch, Hiroyuki Shinoda 0002, Pilar Aivar, Brian T. Sullivan |
ETRA | 2 |
| 2002 | Saccade contingent updating in virtual realityabstractWe are interested in saccade contingent scene updates where the visual information presented in a display is altered while a saccadic eye movement of an unconstrained, freely moving observer is in progress. Since saccades typically last only several tens of milliseconds depending on their size, this poses dif cult constraints on the latency of detection. We have integrated two complementary eye trackers in a virtual reality helmet to simultaneously 1) detect saccade onsets with very low latency and 2) track the gaze with high precision albeit higher latency. In a series of experiments we demonstrate the system s capability of detecting saccade onsets with suf ciently low latency to make scene changes while a saccade is still progressing. While the method was developed to facilitate studies of human visual perception and attention, it may nd interesting applications in human-computer interaction and computer graphics. Jochen Triesch, Brian T. Sullivan, Mary M. Hayhoe, Dana H. Ballard |
ETRA | 4 |
| 2002 | Attentional Object Spotting by Integrating Multimodal InputabstractAn intelligent human-computer interface is expected to allow computers to work with users in a cooperative manner. To achieve this goal, computers need to be aware of user attention and provide assistance without explicit user requests. Cognitive studies of eye movements suggest that in accomplishing well-learned tasks, the performer's focus of attention is locked onto ongoing work and more than 90% of eye movements are closely related to the objects being manipulated in the tasks. In light of this, we have developed an attentional object spotting system that integrates multimodal data consisting of eye position, head position and video from the "first-person" perspective. To detect the user's focus of attention, we modeled eye gaze and head movements using a hidden Markov model (HMM) representation. For each attentional point in time, the object of user interest is automatically extracted and recognized. We report the results of experiments on finding attentional objects in the natural task of "making a peanut-butter sandwich". Chen Yu 0001, Dana H. Ballard, Shenghuo Zhu |
ICMI | 2 |
| 2002 | Distributed synchrony
Zuohua Zhang, Dana H. Ballard |
Neurocomputing | 2 |
| 2000 | A single-spike model of predictive codingabstractThe standard cortical model assumes that the meaning of a neuron's signal is contained in its firing rate. While that model has been used to interpret a voluminous amount of experimental data, it does not address the question of timing, or how recipient neurons can decode this signal in time to predict behavioral results. We propose a model based on coincident firing of large groups of neurons. We show using an example of predictive coding, how the cortex can support vast amounts of non-interfering parallel computation. Dana H. Ballard, Rajesh P. N. Rao, Zuohua Zhang |
Neurocomputing | 1 |
| 2000 | Single trial P3 epoch recognition in a virtual environmentabstractMost visual evoked potential (EP) experiments require a subject to sit still and stare at a computer screen. Virtual reality (VR) can provide more natural-like experimental environments while maintaining environmental control, however subjects are more likely to move and create other artifacts while immersed in a VR environment. We show that a P3 evoked potential (EP) is obtained when subjects stop/go at stoplights in a virtual town and that this potential can be reliably detected on a single trial basis in spite of the extra artifacts generated by subjects in VR. We describe the use of a robust Kalman filter for recognition and note that an average recognition rate of 84.5% is obtained using this filter. The benefits and drawbacks of working in an immersive VR environment are described and future directions suggested. Jessica D. Bayliss, Dana H. Ballard |
Neurocomputing | 2 |
| 1999 | Recognizing Evoked Potentials in a Virtual Environment
Jessica D. Bayliss, Dana H. Ballard |
NIPS | 2 |
| 1998 | Visual Routines for Autonomous DrivingabstractThe paper describes visual routines based on models of color and shape, as well as crucial issues involving the scheduling of such routines. The visual routines are developed in a unique platform. The view from a car driving in a simulated world is feed into a Datacube pipeline video processor. The use of this simulation provides a flexible environment from which to set crucial image processing parameters of the individual routines. In addition to the simulations, the routines are also tested in similar images generated by driving in the real world, to assure the generalizability of the simulation. Garbis Salgian, Dana H. Ballard |
ICCV | 2 |
| 1998 | Category Learning Through Multi-Modality SensingabstractHumans and other animals learn to form complex categories without receiving a target output, or teaching signal, with each input pattern. In contrast, most computer algorithms that emulate such performance assume the brain is provided with the correct output at the neuronal level or require grossly unphysiological methods of information propagation. Natural environments do not contain explicit labeling signals, but they do contain important information in the form of temporal correlations between sensations to different sensory modalities, and humans are affected by this correlational structure (Howells, 1944; McGurk & MacDonald, 1976; MacDonald & McGurk, 1978; Zellner & Kautz, 1990; Durgin & Proffitt, 1996). In this article we describe a simple, unsupervised neural network algorithm that also uses this natural structure. Using only the co-occurring patterns of lip motion and sound signals from a human speaker, the network learns separate visual and auditory speech classifiers that perform comparably to supervised networks. Virginia R. de Sa, Dana H. Ballard |
Neural Comput. | 2 |
| 1998 | Phonetic Set Indexing for Fast Lexical AccessabstractA novel nonsequential indexing mechanism (termed phonetic set indexing) has been evaluated for the purpose of fast word pre-selection. Our approach to handling the lexical access problem stems from the primary observation that the set of phones which are present in the transcription of a word is sparsely distributed across the vocabulary, and is thus suitable as an indexing key for retrieving a short-list of word possibilities. Ramesh R. Sarukkai, Dana H. Ballard |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1997 | Dynamic Model of Visual Recognition Predicts Neural Response Properties in the Visual CortexabstractThe responses of visual cortical neurons during fixation tasks can be significantly modulated by stimuli from beyond the classical receptive field. Modulatory effects in neural responses have also been recently reported in a task where a monkey freely views a natural scene. In this article, we describe a hierarchical network model of visual recognition that explains these experimental observations by using a form of the extended Kalman filter as given by the minimum description length (MDL) principle. The model dynamically combines input-driven bottom-up signals with expectation-driven top-down signals to predict current recognition state. Synaptic weights in the model are adapted in a Hebbian manner according to a learning rule also derived from the MDL principle. The resulting prediction-learning scheme can be viewed as implementing a form of expectation-maximization (EM) algorithm. The architecture of the model posits an active computational role of the reciprocal connections between adjoining visual cortical areas in determining neural response properties. In particular, the model demonstrates the possible role of feedback from higher cortical areas in mediating neurophysiological effects due to stimuli from beyond the classical receptive field. Simulations of the model are provided that help explain the experimental observations regarding neural responses in both free viewing and fixation conditions. Rajesh P. N. Rao, Dana H. Ballard |
Neural Comput. | 2 |
| 1997 | Word set probability boosting for improved spontaneous dialog recognitionabstractBased on the observation that the unpredictable nature of conversational speech makes it almost impossible to reliably model sequential word constraints, the notion of word set error criteria is proposed for improved recognition of spontaneous dialogs. The single-pass adaptive boosting (AB) algorithm enables the language model weights to be tuned using the word set error criteria. In the two-pass version of the algorithm, the basic idea is to predict a set of words based on some a priori information, and perform a rescoring pass wherein the probabilities of the words in the predicted word set are amplified or boosted in some manner. An adaptive gradient descent procedure for tuning the word boosting factor is formulated, which enables the boost factors to be incrementally adjusted to maximize the accuracy of the speech recognition system outputs on held-out training data using the word set error criteria. Two novel models which predict the required word sets are presented: (i) utterance triggers, which capture within-utterance long-distance word interdependencies, and (ii) dialog triggers, which capture local temporal dialog-oriented word relations. The proposed trigger and adaptive boosting (TAB) algorithm, and the single-pass adaptive boosting (AB) algorithm are experimentally tested on a subset of the TRAINS-93 spontaneous dialogs and the TRAINS-95 semispontaneous corpus, and the results summarized. Ramesh R. Sarukkai, Dana H. Ballard |
IEEE Trans. Speech Audio Process. | 2 |
| 1996 | A novel word pre-selection method based on phonetic set indexingabstractThe possibility of pre-fetching words using the phoneme sequence output of automatic speech recognition systems has been explored. A novel non-sequential indexing mechanism (termed phonetic set indexing) has been evaluated for the purpose of fast word pre-selection. Our approach to handling the lexical access problem stems from the primacy observation that the set of phonemes which are present in a transcription of a word is sparsely distributed across the vocabulary, and is thus suitable as an indexing key for retrieving a short list of word possibilities. The fixed dimensionality of the phonetic set representation also makes it suitable for implementation with compact bit strings, thus enabling fast lexical access using low-level bit operations. Ramesh R. Sarukkai, Dana H. Ballard |
ICASSP | 2 |
| 1996 | Improved spontaneous dialogue recognition using dialogue and utterance triggers by adaptive probability boosting
Ramesh R. Sarukkai, Dana H. Ballard |
ICSLP | 2 |
| 1995 | Object Indexing Using an Iconic Sparse Distributed MemoryabstractA general-purpose object indexing technique is described that combines the virtues of principal component analysis with the favorable matching properties of high-dimensional spaces to achieve high-precision recognition. An object is represented by a set of high-dimensional iconic feature vectors comprised of the responses of derivatives of Gaussian filters at a range of orientations and scales. Since these filters can be shown to form the eigenvectors of arbitrary images containing both natural and man-made structures, they are well-suited for indexing in disparate domains. The indexing algorithm uses an active vision system in conjunction with a modified form of Kanerva's (1988, 1993) sparse distributed memory which facilitates interpolation between views and provides a convenient platform for learning the association between an object's appearance and its identity. The robustness of the indexing method was experimentally confirmed by subjecting the method to a range of viewing conditions and the accuracy was verified using a well-known model database containing a number of complex 3D objects under varying pose.> Rajesh P. N. Rao, Dana H. Ballard |
ICCV | 2 |
| 1995 | Remote TeleassistanceabstractWe contrast the effects of communication latency on our human/robot control technique, called teleassistance, versus traditional teleoperation. In teleassistance, a human operator uses hand signs to guide an otherwise autonomous robot manipulator through a given task. Each sign signals a context switch and provides task-centered reference frames for the robot's autonomous servo-motor routines. The signs are natural, such as pointing to an object to indicate the desire to reach toward it as well as the axis along which to reach. The robot is a Utah/MIT hand mounted on a Puma 760 arm. For both teleassistance and teleoperation, the operator wears an EXOS hand master, a polhemus arm position sensor and a Virtual Research helmet that is coupled to binocular cameras mounted on a second Puma 760. The use of a video helmet allows for remote control of the robot and the simulation of communication lag-time. Experimental results suggest that teleassistance scales well with latency delays, unlike teleoperation. Polly K. Pook, Dana H. Ballard |
ICRA | 2 |
| 1995 | Natural Basis Functions and Topographic Memory for Face Recognition
Rajesh P. N. Rao, Dana H. Ballard |
IJCAI | 2 |
| 1995 | The distance set representation of speech segments
Ramesh R. Sarukkai, Dana H. Ballard |
EUROSPEECH | 2 |
| 1995 | Modeling Saccadic Targeting in Visual Search
Rajesh P. N. Rao, Gregory J. Zelinsky, Mary M. Hayhoe, Dana H. Ballard |
NIPS | 4 |
| 1995 | An Active Vision Architecture Based on Iconic RepresentationsabstractActive vision systems have the capability of continuously interacting with the environment. The rapidly changing environment of such systems means that it is attractive to replace static representations with visual routines that compute information on demand. Such routines place a premium on image data structures that are easily computed and used. The purpose of this paper is to propose a general active vision architecture based on efficiently computable iconic representations. This architecture employs two primary visual routines, one for identifying the visual image near the fovea (object identification), and another for locating a stored prototype on the retina (object location). This design allows complex visual behaviors to be obtained by composing these two routines with different parameters. The iconic representations are comprised of high-dimensional feature vectors obtained from the responses of an ensemble of Gaussian derivative spatial filters at a number of orientations and scales. These representations are stored in two separate memories. One memory is indexed by image coordinates while the other is indexed by object coordinates. Object location matches a localized set of model features with image features at all possible retinal locations. Object identification matches a foveal set of image features with all possible model features. We present experimental results for a near real-time implementation of these routines on a pipeline image processor and suggest relatively simple strategies for tackling the problems of occlusions and scale variations. We also discuss two additional visual routines, one for top-down foveal targeting using log-polar sensors and another for looming detection, which are facilitated by the proposed architecture. Rajesh P. N. Rao, Dana H. Ballard |
Artif. Intell. | 2 |
| 1994 | Teleassistance: Contextual Guidance for Autonomous Manipulation
Polly K. Pook, Dana H. Ballard |
AAAI | 2 |
| 1994 | Seeing Behind Occlusions
Dana H. Ballard, Rajesh P. N. Rao |
ECCV (1) | 1 |
| 1994 | Hierarchical Self-Organization in Genetic programming
Justinian P. Rosca, Dana H. Ballard |
ICML | 2 |
| 1994 | Cross-coding networks for speech classificationabstractWhat kind of internal representations develop with networks that transform speech of one speaker to that of another? This question is addressed in this paper by a novel supervised coding scheme: cross-coding. Instead of performing auto-association, we train networks to map speech of many speakers to speech of a particular speaker, with intermediate bottlenecks. The internal representations developed are then input to another network trained to label the corresponding sounds. Interestingly, the cross-codings seem to have captured speaker invariant properties in the different sounds. Experiments with multispeaker syllable recognition task show that the proposed scheme outperforms the corresponding multilayered net. Ramesh R. Sarukkai, Dana H. Ballard |
ICPR (2) | 2 |
| 1994 | Deictic teleassistanceabstractWe present a simple sign language for teleassistance inspired by the work of the Bernstein (1967) and by psychophysical evidence in hand-eye coordination. In our schema, a teleoperator uses hand signs to guide an otherwise autonomous robot manipulator through a given task. Each sign signals a context switch and provides a hand-centered reference frame for the robot's servomotor routines. The signs are natural, such as pointing to an object to indicate the desire to reach toward it as well as the axis along which to reach. These signs are called deictic from the Greek word for pointing to stress their indicative and relative nature. The task example is opening a door using a Utah/MIT hand mounted on a Puma 760 arm. The teleoperator wears an EXOS hand master and polhemus sensor. Three variations of nearest neighbor pattern classification are tested for online recognition of the sign language. The simplest, in which the operator signs each pose once before starting, is the best for this task. The dual-control strategy of teleassistance combines teleoperation and autonomous servo control to their advantage. The use of a symbolic sign language helps to alleviate many problems inherent to literal master/slave teleoperation. Conversely, the integration of global operator guidance and hand-centered coordinate frames permits the servo routines to position the robot in relative coordinates and interpret feedback within a constrained context, significantly simplifying the computation and reducing the need for detailed task models.> Polly K. Pook, Dana H. Ballard |
IROS | 2 |
| 1994 | Learning Saccadic Eye Movements Using Multiscale Spatial FiltersabstractWe describe a framework for learning saccadic eye movements using a photometric representation of target points in natural scenes. The rep(cid:173) resentation takes the form of a high-dimensional vector comprised of the responses of spatial filters at different orientations and scales. We first demonstrate the use of this response vector in the task of locating pre(cid:173) viously foveated points in a scene and subsequently use this property in a multisaccade strategy to derive an adaptive motor map for delivering accurate saccades. Rajesh P. N. Rao, Dana H. Ballard |
NIPS | 2 |
| 1994 | Using intermediate objects to improve the efficiency of visual search
Lambert E. Wixson, Dana H. Ballard |
Int. J. Comput. Vis. | 2 |
| 1993 | Head-centered orientation strategies in animate visionabstractThe authors consider orienting, that is, establishing and maintaining a spatial relation between a motorized pair of cameras (the eye-head system) and a static or a moving object tracked over time. Motivated by physiological evidence, they propose a simple set of vision-based strategies aimed to perform head, eye, and body movements in a complex environment. Fixation is shown to be an essential feature in visual servoing, and is used to decouple control on head rotational degrees of freedom, making possible a metric-less approach to the orientation problem. An implementation of these strategies, using a binocular camera system mounted on a PUMA 700 robotic system, demontrated the effectiveness of the approach.> Enrico Grosso, Dana H. Ballard |
ICCV | 2 |
| 1993 | Report on Workshop on High Performance Computing and Communications for Grand Challenge Applications: Computer Vision, Speech and Natural Language Processing, and Artificial IntelligenceabstractThe findings of a workshop, the goals of which were to identify applications, research problems, and designs of high performance computing and communications (HPCC) systems for supporting applications are discussed. In computer vision, the main scientific issues are machine learning, surface reconstruction, inverse optics and integration, model acquisition, and perception and action. In speech and natural language processing (SNLP), issues were identified statistical analysis in corpus-based speech and language understanding, search strategies for language analysis, auditory and vocal-tract modeling, integration of multiple levels of speech and language analyses, and connectionist systems. In AI, important issues that need immediate attention include the development of efficient machine learning and heuristic search methods that can adapt to different architectural configurations, and the design and construction of scalable and verifiable knowledge bases, active memories, and artificial neural networks.> Benjamin W. Wah, Thomas S. Huang, Aravind K. Joshi, Dan I. Moldovan, Yiannis Aloimonos, Ruzena Bajcsy, Dana H. Ballard, Doug DeGroot, Kenneth A. De Jong, Charles R. Dyer, Scott E. Fahlman, Ralph Grishman, Lynette Hirschman, Richard E. Korf, Stephen E. Levinson, Daniel P. Miranker, N. H. Morgan, Sergei Nirenburg, Tomaso A. Poggio, Edward M. Riseman, Craig Stanfil, Salvatore J. Stolfo, Steven L. Tanimoto, Charles C. Weems |
IEEE Trans. Knowl. Data Eng. | 7 |
| 1992 | Low resolution cues for guiding saccadic eye movementsabstractThe high-resolution field of view of the human eye only covers a tiny fraction of the total field of view, which allows for great economy in computational resources but forces the visual system to solve other problems that would not exist with uniformly high resolution. One of these is how to determine where to redirect the fovea, given only the low-resolution information available in the periphery. The advent of spatially variant receptor arrays for cameras has made it imperative that computational solutions to this problem be found. Color has been traditionally associated with foveal vision, but it is shown that color cues are well preserved under low resolution, and an algorithm for locating objects based on color histograms that is both effective under low resolution and computationally efficient is illustrated.> Michael J. Swain, Roger E. Kahn, Dana H. Ballard |
CVPR | 3 |
| 1992 | A Note on Learning Vector Quantization
Virginia R. de Sa, Dana H. Ballard |
NIPS | 2 |
| 1992 | Principles of animate vision
Dana H. Ballard |
CVGIP Image Underst. | 1 |
| 1991 | Animate Vision
Dana H. Ballard |
Artif. Intell. | 1 |
| 1991 | Egomotion perception using visual trackingabstractThe ability of a biological organism to visually track a perceptually significant feature in its environment has been argued to be an important feedback mechanism guiding locomotion. This paper analyzes the constraints available from the visual motion stimuli in the context of tracking. Our aim is to show that the act of tracking simplifies the decoding of egomotion parameters from motion stimuli. The constraints obtainable under tracking are utilized to analyze a possible egomotion decoding strategy for a binocular robot eye system, modeled after the human ocular tracking (smooth pursuit) mechanism. The main result of the paper is in the derivation of a closed‐form solution of the egomotion parameters using feedback information concerning the movement of the tracking motors over time. The theoretical results are verified by experiments. We believe that the active tracking approach presented here is a more simple, practical, and manageable technique in a robot navigation setting, compared to passive methods. Amit Bandopadhay, Dana H. Ballard |
Comput. Intell. | 2 |
| 1991 | Color indexing
Michael J. Swain, Dana H. Ballard |
Int. J. Comput. Vis. | 2 |
| 1991 | Learning to Perceive and Act by Trial and Error
Steven D. Whitehead, Dana H. Ballard |
Mach. Learn. | 2 |
| 1990 | Indexing via color histogramsabstractThis paper shows color histograms to be stable object representations over change in view, and demonstrates that they can differentiate among a large number of objects. The authors introduce a technique called histogram intersection for efficiently matching model and image histograms. Color can also be used to search for the location of an object. An algorithm called histogram backprojection performs this task efficiently in crowded scenes.> Michael J. Swain, Dana H. Ballard |
ICCV | 2 |
| 1990 | Active Perception and Reinforcement Learning
Steven D. Whitehead, Dana H. Ballard |
ML | 2 |
| 1990 | Active Perception and Reinforcement LearningabstractThis paper considers adaptive control architectures that integrate active sensorimotor systems with decision systems based on reinforcement learning. One unavoidable consequence of active perception is that the agent's internal representation often confounds external world states. We call this phenomenon perceptual aliasing and show that it destabilizes existing reinforcement learning algorithms with respect to the optimal decision policy. A new decision system that overcomes these difficulties is described. The system incorporates a perceptual subcycle within the overall decision cycle and uses a modified learning algorithm to suppress the effects of perceptual aliasing. The result is a control architecture that learns not only how to solve a task but also where to focus its attention in order to collect necessary sensory information. Steven D. Whitehead, Dana H. Ballard |
Neural Comput. | 2 |
| 1989 | A Role for Anticipation in Reactive Systems that Learn
Steven D. Whitehead, Dana H. Ballard |
ML | 2 |
| 1989 | Reference Frames for Animate Vision
Dana H. Ballard |
IJCAI | 1 |
| 1989 | Behavioural constraints on animate vision
Dana H. Ballard |
Image Vis. Comput. | 1 |
| 1988 | Eye Fixation And Early Vision: Kinetic DepthabstractOne aspect of primate intelligence is the ability to coordinate eye movements in the process of solving complex tasks. Primate eye movements have been studied in several disciplines but little work has been directed toward a computational theory that shows how the eye nouements can confer specific advantages in problem-solving behaviors, This paper outlines some of the elements of such a theory, emphasizing one point: the advantages of an active system in choosing an external frame of reference for the computations of early vision. These advantages are illustrated by the real-time computation of depth map with a monocular, fixating vision system. Dana H. Ballard, Altan Ozcandarli |
ICCV | 1 |
| 1988 | Fixed Point Analysis for Recurrent Networks
Patrice Y. Simard, Mary B. Ottaway, Dana H. Ballard |
NIPS | 3 |
| 1987 | Modular Learning in Neural Networks
Dana H. Ballard |
AAAI | 1 |
| 1986 | Parallel Logical Inference and Energy Minimization
Dana H. Ballard |
AAAI | 1 |
| 1986 | Task frames: Primitives for sensory-motor coordination
Dana H. Ballard, Leo Hartman |
Comput. Vis. Graph. Image Process. | 1 |
| 1985 | Self-calibration in robot manipulatorsabstractThe development of fast recursive methods for computing manipulator inverse dynamics has made possible open loop control strategies. However, for these strategies to work, an accurate plant model is required. Two key components of the plant are frictional terms and link/load inertias. This paper shows how these components may be computed using force and moment sensing. Amitabha Mukerjee, Dana H. Ballard |
ICRA | 2 |
| 1985 | Transformational Form Perception in 3D: Constraints, Algorithms, Implementation
Dana H. Ballard, Hiromi Tanaka |
IJCAI | 1 |
| 1984 | Task Frames in Robot Manipulation
Dana H. Ballard |
AAAI | 1 |
| 1984 | Parameter Nets
Dana H. Ballard |
Artif. Intell. | 1 |
| 1983 | Boundary Conditions in Multiple Intrinsic Images
Bernhard H. Stuth, Dana H. Ballard |
IJCAI | 2 |
| 1983 | Rigid body motion from depth and optical flow
Dana H. Ballard, O. A. Kimball |
Comput. Vis. Graph. Image Process. | 1 |
| 1983 | Viewer Independent Shape RecognitionabstractAn important problem in vision is to detect the presence of a known rigid 3-D object. The general 3-D object recognition task can be thought of as building a description of the object that must have at least two parts: 1) the internal description of the object itself (with respect to an object-centered frame); and 2) the transformation of the object-centered frame to the viewer-centered (image) frame. The reason for this decomposition is parsimony: different views of the object should have minimal impact on its description. This is achieved by factoring the object's description into two sets of parameters, one which is view-independent (the object-centered component) and one which is view-varying (the viewing transformation). Often a description of the object is known beforehand and the task reduces to finding the objectframe to viewer-frame transformation. This paper describes a method for handling this case: a known object is detected by finding changes in orientation, translation, and scale of the object from its canonical description. The method is a Hough technique and has the characteristic insensitivity to occlusion and noise. Dana H. Ballard, Daniel Sabbah |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1981 | Parameter Networks: Towards a Theory of Low-Level Vision
Dana H. Ballard |
IJCAI | 1 |
| 1981 | On Shapes
Dana H. Ballard, Daniel Sabbah |
IJCAI | 1 |
| 1981 | Generalizing the Hough transform to detect arbitrary shapes
Dana H. Ballard |
Pattern Recognit. | 1 |
| 1979 | Anatomical models for medical imagesabstractWe are working towards a general model of the anatomy seen in medical images that can be used to find the boundaries of organs. We argue that such a model must have four components: - multiple resolution copies of the image (pyrarn ids) ; - geometric shape models for organs; - network-like models of inter-organ relationships; and - a top-down control structure that plans ahead. Dana H. Ballard, Uri Shani, Robert B. Schudy |
COMPSAC | 1 |
| 1977 | An Approach to Knowledge-Directed Image Analysis
Dana H. Ballard, Jay M. Feldman |
IJCAI | 1 |
| 1976 | A Ladder-Structured Decision Tree for Recognizing Tumors in Chest RadiographsabstractWe describe a hierarchic computer procedure for the detection of nodular tumors in a chest radiograph. The radiograph is scanned and consolidated into several resolutions which are enhanced and analyzed by a hierarchic tumor recognition process. The hierarchic structure of the tumor recognition process has the form of a ladder-like decision tree. The major steps in the decision tree are: 1) find the lung regions within the chest radiograph, 2) find candidate nodule sites (potential tumor locations) within the lung regions, 3) find boundaries for most of these sites, 4) find nodules from among the candidate nodule boundaries, and 5) find tumors from among the nodules. The first three steps locate potential nodules in the radiograph. The last two steps classify the potential nodules into nonnodules, nodules which are not tumors, and nodules which are tumors. Dana H. Ballard, Jack Sklansky |
IEEE Trans. Computers | 1 |
| 1974 | An Algorithm for the Solution of Constrained Generalised Polynomial Programming ProblemsabstractAn algorithm is presented for the solution of a class of constrained, nonlinear programming problems. The problems considered may be formulated as generalised polynomials. This class of problems, which encompasses linear, quadratic and geometric programming problems, can be extended to include functions which are the ratios of generalised polynomials. Computational experience with some typical examples is also reviewed. Dana H. Ballard, C. O. Jelinek, Roland Schinzinger |
Comput. J. | 1 |