James J. Little

dblp:18/2287 · also James Little 0001, Jim Little 0001 · DBLP profile ↗
← Back
114ranked-venue papers
13as first author
7since 2021 · last 2026
0000-0002-4411-2544ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 87 · 9 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 7 first-author · 6 since 2021Systems, architecture and hardware · 27 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4Theory of computation · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Test-Time Consistency in Vision Language Models
abstract
Vision-Language Models (VLMs) have achieved impressive performance across a wide range of multimodal tasks, yet they often exhibit inconsistent behavior when faced with semantically equivalent inputs—undermining their reliability and robustness. Recent benchmarks, such as MM-R3, highlight that even state-of-the-art VLMs can produce divergent response across semantically equivalent inputs, despite maintaining high average accuracy. Prior work addresses this issue by modifying model architectures or conducting large-scale fine-tuning on curated datasets. In contrast, we propose a simple and effective test-time consistency framework that enhances semantic consistency without supervised re-training. Our method is entirely post-hoc, model-agnostic, and applicable to any VLM with access to its weights. Given a single test point, we enforce consistent predictions via two complementary objectives: (i) a Cross-Entropy Agreement Loss that aligns predictive distributions across semantically equivalent inputs, and (ii) a Pseudo-Label Consistency Loss that draws outputs toward a self-averaged consensus. Our method is plug-and-play, and leverages information from a single test-input itself to improve consistency. Experiments on the MM-R3benchmark show that our framework yields substantial gains in consistency across state-of-the-art models, establishing a new direction for inference-time adaptation in VLMs.1
Shih-Han Chou, Shivam Chandhok, James J. Little, Leonid Sigal
WACV3
2024 Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
abstract
The emergence of attention-based transformer models has led to their extensive use in various tasks, due to their superior generalization and transfer properties. Recent research has demonstrated that such models, when prompted appropriately, are excellent for few-shot inference. However, such techniques are under-explored for dense prediction tasks like semantic segmentation. In this work, we examine the effectiveness of prompting a transformer-decoder with learned visual prompts for the generalized few-shot segmentation (GFSS) task. Our goal is to achieve strong performance not only on novel categories with limited examples, but also to retain performance on base cat-egories. We propose an approach to learn visual prompts with limited examples. These learned visual prompts are used to prompt a multiscale transformer decoder to facilitate accurate dense predictions. Additionally, we introduce a unidirectional causal attention mechanism between the novel prompts, learned with limited examples, and the base prompts, learned with abundant data. This mecha-nism enriches the novel prompts without deteriorating the base class performance. Overall, this form of prompting helps us achieve state-of-the-art performance for GFSS on two different benchmark datasets: COCO-20i and Pascal-Si, without the need for test-time optimization (or transduction). Furthermore, test-time optimization leveraging unlabelled test data can be used to improve the prompts, which we refer to as transductive prompt tuning.
Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal, James J. Little
CVPR4
2024 Framework-agnostic Semantically-aware Global Reasoning for Segmentation
abstract
Recent advances in pixel-level tasks (e.g. segmentation) illustrate the benefit of of long-range interactions between aggregated region-based representations that can enhance local features. However, such aggregated representations, often in the form of attention, fail to model the underlying semantics of the scene (e.g. individual objects and, by extension, their interactions). In this work, we address the issue by proposing a component that learns to project image features into latent representations and reason between them using a transformer encoder to generate contextualized and scene-consistent representations which are fused with original image features. Our design encourages the latent regions to represent semantic concepts by ensuring that the activated regions are spatially disjoint and the union of such regions corresponds to a connected object segment. The proposed semantic global reasoning (SGR) component is end-to-end trainable and can be easily added to a wide variety of backbones (CNN or transformer-based) and segmentation heads (per-pixel or mask classification) to consistently improve the segmentation results on different datasets. In addition, our latent tokens are semantically interpretable and diverse and provide a rich set of features that can be transferred to downstream tasks like object detection and segmentation, with improved performance. Furthermore, we also proposed metrics to quantify the semantics of latent tokens at both class & instance level.
Mir Rayat Imtiaz Hossain, Leonid Sigal, James J. Little
WACV3
2024 Implicit and explicit commonsense for multi-sentence video captioning
abstract
Existing dense or paragraph video captioning approaches rely on holistic representations of videos, possibly coupled with learned object/action representations, to condition hierarchical language decoders. However, they fundamentally lack the commonsense knowledge of the world required to reason about progression of events, causality, and even the function of certain objects within a scene. To address this limitation we propose a novel video captioning Transformer-based model, that takes into account both implicit (visuo-lingual and purely linguistic) and explicit (knowledge-base) commonsense knowledge. We show that these forms of knowledge, in isolation and in combination, enhance the quality of produced captions. Further, inspired by imitation learning, we propose a new task of instruction generation, where the goal is to produce a set of linguistic instructions from a video demonstration of its performance. We formalize the task using the ALFRED dataset generated using an AI2-THOR environment. While instruction generation is conceptually similar to paragraph captioning, it differs in the fact that it exhibits stronger object persistence, as well as spatially-aware and causal sentence structure. We show that our commonsense knowledge enhanced approach produces significant improvements on this task (up to 57% in METEOR and 8.5% in CIDEr), as well as the state-of-the-art result on more traditional video captioning in the ActivityNet Captions dataset.
Shih-Han Chou, James J. Little, Leonid Sigal
Comput. Vis. Image Underst.2
2022 Bootstrapping Human Optical Flow and Pose
Aritro Roy Arko, James J. Little, Kwang Moo Yi
BMVC2
2022 ElePose: Unsupervised 3D Human Pose Estimation by Predicting Camera Elevation and Learning Normalizing Flows on 2D Poses
abstract
Human pose estimation from single images is a challenging problem that is typically solved by supervised learning. Unfortunately, labeled training data does not yet exist for many human activities since 3D annotation requires dedicated motion capture systems. Therefore, we propose an unsupervised approach that learns to predict a 3D human pose from a single image while only being trained with 2D pose data, which can be crowd-sourced and is already widely available. To this end, we estimate the 3D pose that is most likely over random projections, with the likelihood estimated using normalizing flows on 2D poses. While previous work requires strong priors on camera rotations in the training data set, we learn the distribution of camera angles which significantly improves the performance. Another part of our contribution is to stabilize training with normalizing flows on high-dimensional 3D pose data by first projecting the 2D poses to a linear subspace. We outperform the state-of-the-art unsupervised human pose estimation methods on the benchmark datasets Human3.6M and MPI-INF-3DHP in many metrics.
Bastian Wandt, James J. Little, Helge Rhodin
CVPR2
2021 AutoRetouch: Automatic Professional Face Retouching
abstract
Face retouching is one of the most time-consuming steps in professional photography pipelines. The existing auto-mated approaches blindly apply smoothing on the skin, destroying the delicate texture of the face. We present the first automatic face retouching approach that produces high-quality professional-grade results in less than two seconds. Unlike previous work, we show that our method preserves textures and distinctive features while retouching the skin. We demonstrate that our trained models generalize across datasets and are suitable for low-resolution cellphone images. Finally, we release the first large-scale, professionally retouched dataset with our baseline to encourage further work on the presented problem.
Alireza Shafaei, James J. Little, Mark Schmidt 0001
WACV2
2019 Pan-tilt-zoom SLAM for Sports Videos
Jikai Lu, James J. Little
BMVC3
2019 A Less Biased Evaluation of Out-of-distribution Sample Detectors
Alireza Shafaei, Mark Schmidt 0001, James J. Little
BMVC3
2019 Spatio-temporal Relational Reasoning for Video Question Answering
Gursimran Singh, Leonid Sigal, James J. Little
BMVC3
2019 Learning Sports Camera Selection From Internet Videos
abstract
This work addresses camera selection, the task of predicting which camera should be on air from multiple candidate cameras for soccer broadcast. The task is challenging because of the scarcity of learning data with all candidate views. Meanwhile, broadcast videos are freely available on the Internet (e.g. Youtube). However, these videos only record the selected camera views, omitting the other candidate views. To overcome this problem, we first introduce a random survival forest (RSF) method to impute the incomplete data effectively. Then, we propose a spatial-appearance heatmap to describe foreground objects (e.g. players and balls) in an image. To evaluate the performance of our system, we collect the largest-ever dataset for soccer broadcasting camera selection. It has one main game which has all candidate views and twelve auxiliary games which only have the broadcast view. Our method significantly outperforms state-of-the-art methods on this challenging dataset. Further analysis suggests that the improvement in performance is indeed from the extra information from auxiliary games.
Keyu Lu, Sijia Tian, James J. Little
WACV4
2018 Exploiting Temporal Information for 3D Human Pose Estimation
Mir Rayat Imtiaz Hossain, James J. Little
ECCV (10)2
2018 LSQ++: Lower Running Time and Higher Recall in Multi-codebook Quantization
Julieta Martinez 0001, Shobhit Zakhmi, Holger H. Hoos, James J. Little
ECCV (16)4
2018 Exploiting Points and Lines in Regression Forests for RGB-D Camera Relocalization
abstract
Camera relocalization plays a vital role in many robotics and computer vision applications, such as self-driving cars and virtual reality. Recent random forests based methods exploit randomly sampled pixel comparison features to predict 3D world locations for 2D image locations to guide the camera pose optimization. However, these point features are only sampled randomly in images, without considering geometric information such as lines, leading to large errors with the existence of poorly textured areas or in motion blur. Line segments are more robust in these environments. In this work, we propose to jointly exploit points and lines within the framework of uncertainty driven regression forests. The proposed approach is thoroughly evaluated on three publicly available datasets against several strong state-of-the-art baselines in terms of several different error metrics. Experimental results prove the efficacy of our method, showing superior or on-par state-of-the-art performance.
Lili Meng, Frederick Tung, James J. Little, Julien Valentin, Clarence W. de Silva
IROS3
2018 Video-based Human Fall Detection in Smart Homes Using Deep Learning
abstract
Automatic human fall detection is a challenging task of healthcare in smart homes, and video cameras have been proved to be efficient in addressing this problem. Although existing methods perform relatively well, they are all built upon "hand-crafted" features, thus constraining the performance of the model to some presumed conditions and scenarios, and making it vulnerable to any deviation from the assumed settings. In this paper, we propose a deep-learning-based approach for human fall detection, using long short-term memory neural network. Our model is not restricted to any specific circumstances, and performance evaluations show that it outperforms all the existing methods.
Anahita Shojaei-Hashemi, Panos Nasiopoulos, James J. Little, Mahsa T. Pourazad
ISCAS3
2018 Camera Selection for Broadcasting Soccer Games
abstract
When broadcasting events such as soccer games, human operators constantly select the camera with the best viewpoint to cover the whole event. Modeling the prediction of which camera should be on air will assist automatic sports broadcasts and influence millions of viewers. In this paper, we propose a proof-of-concept method to automatically select cameras for broadcasting soccer games. First, a random forest based regressor smoothly predicts the visual importance of short video clips using deep convolutional features. Then, the predictions from multiple candidate cameras are regularized by a novel camera duration cumulative distribution function (CDF), naturally guiding the camera selection. We apply our approach to real soccer broadcasts with a professional human operator's result as a reference. The quantitative experiments demonstrate that our method outperforms two alternatives in terms of prediction accuracy. Moreover, the video generated by our method is preferred in the user study experiment, exhibiting its practicality.
Lili Meng, James J. Little
WACV3
2018 A Two-Point Method for PTZ Camera Calibration in Sports
abstract
Calibrating narrow field of view soccer cameras is challenging because there are very few field markings in the image. Unlike previous solutions, we propose a two-point method, which requires only two point correspondences given the prior knowledge of base location and orientation of a pan-tilt-zoom (PTZ) camera. We deploy this new calibration method to annotate pan-tilt-zoom data from soccer videos. The collected data are used as references for new images. We also propose a fast random forest method to predict pan-tilt angles without image-to-image feature matching, leading to an efficient calibration method for new images. We demonstrate our system on synthetic data and two real soccer datasets. Our two-point approach achieves superior performance over the state-of-the-art method.
Fangrui Zhu, James J. Little
WACV3
2018 Lightweight convolutional neural networks for player detection and classification
Keyu Lu, James J. Little, Hangen He
Comput. Vis. Image Underst.3
2018 Resolving Occlusion in Active Visual Target Search of High-Dimensional Robotic Systems
abstract
We propose an algorithm for handling visual occlusions that disrupt visual tracking of high-dimensional eye-in-hand systems. Our algorithm allows a robot to look behind an occluder during active visual target search and reacquire its target in an online manner. A particle filter continuously estimates the target location and an enhanced observation model updates the target belief state. Meanwhile, we build a simple but efficient map of the occluder boundaries to compute potential occlusion-clearing motions. Our mixed-initiative cost function balances the goal of gaining more information about the target and occluder boundary while minimizing the sensor action cost. A data-driven planner uses informed samples to strike a balance between target search and information gain to avoid exhaustive mapping of the three-dimensional occluder into Configuration space. We demonstrate the capabilities of our algorithm in simulation and a real-world experiment. We also show that our proposed solvers outperform a common approach in the literature. Our results indicate that our algorithm can quickly obtain clear views of the target when occlusion is persistent and significant camera motion is required.
Sina Radmard, David Meger, James J. Little, Elizabeth A. Croft
IEEE Trans. Robotics3
2017 Light Cascaded Convolutional Neural Networks for Accurate Player Detection
Keyu Lu, James J. Little, Hangen He
BMVC3
2017 A Simple Yet Effective Baseline for 3d Human Pose Estimation
abstract
Following the success of deep convolutional networks, state-of-the-art methods for 3d human pose estimation have focused on deep end-to-end systems that predict 3d joint locations given raw image pixels. Despite their excellent performance, it is often not easy to understand whether their remaining error stems from a limited 2dpose (visual) understanding, or from a failure to map 2d poses into 3dimensional positions. With the goal of understanding these sources of error, we set out to build a system that given 2d joint locations predicts 3d positions. Much to our surprise, we have found that, with current technology, “lifting” ground truth 2djoint locations to 3d space is a task that can be solved with a remarkably low error rate: a relatively simple deep feedforward network outperforms the best reported result by about 30% on Human3.6M, the largest publicly available 3d pose estimation benchmark. Furthermore, training our system on the output of an off-the-shelf state-of-the-art 2d detector (i.e., using images as input) yields state of the art results - this includes an array of systems that have been trained end-to-end specifically for this task. Our results indicate that a large portion of the error of modern deep 3d pose estimation systems stems from their visual analysis, and suggests directions to further advance the state of the art in 3d human pose estimation.
Julieta Martinez 0001, Mir Rayat Imtiaz Hossain, Javier Romero 0002, James J. Little
ICCV4
2017 MF3D: Model-free 3D semantic scene parsing
abstract
We present a novel model-free method for online 3D semantic scene parsing from video sequences. MF3D (Model-Free 3D) is different from conventional methods for 3D scene parsing in that voxel labelling is approached via search-based label transfer instead of discriminative classification. This non-parametric approach makes MF3D easy to scale with an online growth in the database, as no model re-training is required with the addition of new examples or categories. Experimental results on the KITTI benchmark demonstrate that our model-free approach enables accurate online 3D scene parsing while retaining extensibility to new categories. In addition, we show that unsupervised binary encoding (hashing) techniques can be easily incorporated into our framework for scalability to larger databases.
Frederick Tung, James J. Little
ICRA2
2017 Backtracking regression forests for accurate camera relocalization
abstract
Camera relocalization plays a vital role in many robotics and computer vision tasks, such as global localization, recovery from tracking failure, and loop closure detection. Recent random forests based methods directly predict 3D world locations for 2D image locations to guide the camera pose optimization. During training, each tree greedily splits the samples to minimize the spatial variance. However, these greedy splits often produce uneven sub-trees in training or incorrect 2D-3D correspondences in testing. To address these problems, we propose a sample-balanced objective to encourage equal numbers of samples in the left and right sub-trees, and a novel backtracking scheme to remedy the incorrect 2D-3D correspondence predictions. Furthermore, we extend the regression forests based methods to use local features in both training and testing stages for outdoor RGB-only applications. Experimental results on publicly available indoor and outdoor datasets demonstrate the efficacy of our approach, which shows superior or on-par accuracy with several state-of-the-art methods.
Lili Meng, Frederick Tung, James J. Little, Julien Valentin, Clarence W. de Silva
IROS4
2017 Where should cameras look at soccer games: Improving smoothness using the overlapped hidden Markov model
James J. Little
Comput. Vis. Image Underst.2
2016 SSP: Supervised Sparse Projections for Large-Scale Retrieval in High Dimensions
Frederick Tung, James J. Little
ACCV (1)2
2016 Exploiting Random RGB and Sparse Features for Camera Pose Estimation
Lili Meng, Frederick Tung, James J. Little, Clarence W. de Silva
BMVC4
2016 Play and Learn: Using Video Games to Train Computer Vision Models
Alireza Shafaei, James J. Little, Mark Schmidt 0001
BMVC2
2016 Factorized Binary Codes for Large-Scale Nearest Neighbor Search
Frederick Tung, James J. Little
BMVC2
2016 Learning Online Smooth Predictors for Realtime Camera Planning Using Recurrent Decision Trees
abstract
We study the problem of online prediction for realtime camera planning, where the goal is to predict smooth trajectories that correctly track and frame objects of interest (e.g., players in a basketball game). The conventional approach for training predictors does not directly consider temporal consistency, and often produces undesirable jitter. Although post-hoc smoothing (e.g., via a Kalman filter) can mitigate this issue to some degree, it is not ideal due to overly stringent modeling assumptions (e.g., Gaussian noise). We propose a recurrent decision tree framework that can directly incorporate temporal consistency into a data-driven predictor, as well as a learning algorithm that can efficiently learn such temporally smooth models. Our approach does not require any post-processing, making online smooth predictions much easier to generate when the noise model is unknown. We apply our approach to sports broadcasting: given noisy player detections, we learn where the camera should look based on human demonstrations. Our experiments exhibit significant improvements over conventional baselines and showcase the practicality of our approach.
Hoang Minh Le 0002, Peter Carr 0001, Yisong Yue, James J. Little
CVPR5
2016 Revisiting Additive Quantization
Julieta Martinez 0001, Joris Clement, Holger H. Hoos, James J. Little
ECCV (2)4
2016 Efficient video-based retrieval of human motion with flexible alignment
abstract
We present a novel and scalable approach for retrieval and flexible alignment of 3d human motion examples given a video query. Our method efficiently searches a large set of motion capture (mocap) files accounting for speed variations in motion. To align a short video clip with a part of a longer mocap sequence, we experiment with different feature representations comparable across the two modalities. We also evaluate two different Dynamic Time Warping (DTW) approaches that allow sub-sequence matching and suggest additional local constraints for a smooth alignment. Finally, to quantify video-based mocap retrieval, we introduce a benchmark providing a novel set of per-frame action labels for 2 000 files of the CMU-mocap dataset, as well as a collection of realistic video queries taken from YouTube. Our experiments show that temporal flexibility is not only required for the correct alignment of pose and motion, but it also improves the retrieval accuracy.
Ankur Gupta 0004, John He, Julieta Martinez 0001, James J. Little, Robert J. Woodham
WACV4
2016 Scene parsing by nonparametric label transfer of content-adaptive windows
Frederick Tung, James J. Little
Comput. Vis. Image Underst.2
2015 Bank of Quantization Models: A Data-Specific Approach to Learning Binary Codes for Large-Scale Retrieval Applications
abstract
We explore a novel paradigm in learning binary codes for large-scale image retrieval applications. Instead of learning a single globally optimal quantization model as in previous approaches, we encode the database points in a data-specific manner using a bank of quantization models. Each individual database point selects the quantization model that minimizes its individual quantization error. We apply the idea of a bank of quantization models to data independent and data-driven hashing methods for learning binary codes, obtaining state-of-the-art performance on three benchmark datasets.
Frederick Tung, Julieta Martinez 0001, Holger H. Hoos, James J. Little
WACV4
2015 Improving scene attribute recognition using web-scale object detectors
Frederick Tung, James J. Little
Comput. Vis. Image Underst.2
2014 Unlabelled 3D Motion Examples Improve Cross-View Action Recognition
Ankur Gupta 0004, Alireza Shafaei, James J. Little, Robert J. Woodham
BMVC3
2014 3D Pose from Motion for Cross-View Action Recognition via Non-linear Circulant Temporal Encoding
abstract
We describe a new approach to transfer knowledge across views for action recognition by using examples from a large collection of unlabelled mocap data. We achieve this by directly matching purely motion based features from videos to mocap. Our approach recovers 3D pose sequences without performing any body part tracking. We use these matches to generate multiple motion projections and thus add view invariance to our action recognition model. We also introduce a closed form solution for approximate non-linear Circulant Temporal Encoding (nCTE), which allows us to efficiently perform the matches in the frequency domain. We test our approach on the challenging unsupervised modality of the IXMAS dataset, and use publicly available motion capture data for matching. Without any additional annotation effort, we are able to significantly outperform the current state of the art.
Ankur Gupta 0004, Julieta Martinez 0001, James J. Little, Robert J. Woodham
CVPR3
2014 CollageParsing: Nonparametric Scene Parsing by Adaptive Overlapping Windows
Frederick Tung, James J. Little
ECCV (6)2
2014 Ensuring safety in human-robot dialog - A cost-directed approach
abstract
We present an approach for detecting potentially unsafe commands in human-robot dialog, where a robotic system evaluates task cost in input commands to ask input-specific, directed questions to ensure safe task execution. The goal is to reduce risk, both to the robot and the environment, by asking context-appropriate questions. Given an input program, (i.e., a sequence of commands) the system evaluates a set of likely alternate programs along with their likelihood and cost, and these are given as input to a Decision Function to decide whether to execute the task or confirm the plan from the human partner. A process called token-risk grounding identifies the costly commands in the programs, and specifically asks the human user to clarify those commands. We evaluate our system in two simulated robot tasks, and also on-board the Willow Garage PR2 and TurtleBot robots in an indoor task setting. In both sets of evaluations, the results show that the system is able to identify specific commands that contribute to high task cost, and present users the option to either confirm or modify those commands. In addition to ensuring task safety, this results in an overall reduction in robot reprogramming time.
Junaed Sattar, James J. Little
ICRA2
2014 Bayesian Optimization with an Empirical Hardness Model for approximate Nearest Neighbour Search
abstract
Nearest Neighbour Search in high-dimensional spaces is a common problem in Computer Vision. Although no algorithm better than linear search is known, approximate algorithms are commonly used to tackle this problem. The drawback of using such algorithms is that their performance depends highly on parameter tuning. While this process can be automated using standard empirical optimization techniques, tuning is still time-consuming. In this paper, we propose to use Empirical Hardness Models to reduce the number of parameter configurations that Bayesian Optimization has to try, speeding up the optimization process. Evaluation on standard benchmarks of SIFT and GIST descriptors shows the viability of our approach.
Julieta Martinez 0001, James J. Little, Nando de Freitas
WACV2
2014 Task-based control of articulated human pose detection for OpenVL
abstract
Human pose detection is the foundation for many applications, particularly those using gestures as part of a natural user interface. We introduce a novel task-based control method for human pose detection, encoding specialist knowledge in a descriptive abstraction for application by non-experts, such as developers, artists and students. The abstraction hides the details of a set of algorithms which specialise either in different estimations of pose (e.g. articulated, body part) or under different conditions (e.g. occlusion, clutter). Users describe the conditions of their problem, which is used to select the most suitable algorithm (and automatically set up the parameters). The task-based control is evaluated with images described using the abstraction. Expected outcomes are compared to results and demonstrate that describing the conditions is sufficient to allow the abstraction to produce the required result.
Georgii Oleinikov, Gregor Miller, James J. Little, Sidney S. Fels
WACV3
2013 Modeling nonconvex workspace constraints from diverse demonstration sets for Constrained Manipulator Visual Servoing
abstract
This paper presents a novel framework for solving the Constrained Manipulator Visual Servoing (CMVS) problem. Classical eye-in-hand visual servoing relies on a reference image to capture the end-effector positioning task, but non-convex workspace constraints (such as whole-arm collision and camera occlusion constraints) are not represented. An explicit CAD model of the workspace is typically required for collision avoidance and visibility planning algorithms. In our novel CMVS framework, during the reference image capture process, we leverage the user's kinesthetic and visual capabilities to obtain a set of qualitatively-diverse demonstrations that provide information about the robot's work environment. We investigate methods for identifying the topology of the feasible regions represented directly in the control space of the robot (i.e., image-space and joint-space). We use a combination of stochastic modeling and graphical methods to describe the feasible space, capturing both the inter-group and intra-group variations. Specifically, our method uses the inter-groups variations to build a map that describes the global connectivity of the space, while exploiting the intra-group variations to automatically derive the appropriate gains in the control law. For a given target object, we apply online Gaussian Mixture Regression to the relevant feasible space regions to provide an idealized trajectory for tracking in image-space and in joint-space. We illustrate the key advantages of our approach through a set of visual servoing experiments on a Barrett WAM 7-DOF manipulator with a Sony XC-HR70 camera.
Ambrose Chan, Elizabeth A. Croft, James J. Little
ICRA3
2013 Overcoming unknown occlusions in eye-in-hand visual search
abstract
We propose a method for handling persistent visual occlusions that disrupt visual tracking for eye-in-hand systems. Our approach allows a robot to “look behind” an occluder and re-acquire its target. To allow efficient planning, we avoid exhaustive mapping of the 3D occluder into configuration space, and instead use informed samples to strike a balance between target search and information gain. A particle filter continuously estimates the target location when it is not visible. Meanwhile, we build a simple but effective map of the occluder's extents to compute potential occlusion-clearing motions using very few calls to efficient approximations of inverse kinematics. Our mixed-initiative cost function balances the goal of directly locating the target with the goal of gaining information through mapping the occluder. Monte-Carlo optimization with efficient data-driven proposals allows us to approximate one-step solutions efficiently. Experimental evaluation performed on a realistic simulator shows that our method can quickly obtain clear views of the target, even when occlusions are persistent and significant camera motion is required.
Sina Radmard, David Meger, Elizabeth A. Croft, James J. Little
ICRA4
2013 3D spatial relationships for improving object detection
abstract
This work demonstrates how 3D qualitative spatial relationships can be used to improve object detection by differentiating between true and false positive detections. Our method identifies the most likely subset of 3D detections using seven types of 3D relationships and adjusts detection confidence scores to improve the average precision. A model is learned using a structured support vector machine [1] from examples of 3D layouts of objects in offices and kitchens. We test our method on synthetic detections to determine how factors such as localization accuracy, number of detections and detection scores change the effectiveness of 3D spatial relationships for improving object detection rates. Finally, we describe a technique for generating 3D detections from 2D image-based object detections and demonstrate how our method improves the average precision of these 3D detections.
Tristram Southey, James J. Little
ICRA2
2013 Recognition of Action in Broadcast Basketball Videos on the Basis of Global and Local Pairwise Representation
abstract
A new feature-representation method for recognizing actions in broadcast videos, which focuses on the relationship between human actions and camera motions, is proposed. With this method, key point trajectories are extracted as motion features in spatio-temporal sub-regions called "spatio-temporal multiscale bags" (STMBs). Global representations and local representations from one sub-region in the STMBs are then combined to create a "glocal pair wise representation" (GPR). The GPR considers the co-occurrence of camera motions and human actions. Finally, two-stage SVM classifiers are trained with STMB-based GPRs, and specified human actions in video sequences are identified. It was experimentally confirmed that the proposed method can robustly detect specific human actions in broadcast basketball videos.
Masahide Naemura, Mahito Fujii, James J. Little
ISM4
2013 Learning to Track and Identify Players from Broadcast Sports Videos
abstract
Tracking and identifying players in sports videos filmed with a single pan-tilt-zoom camera has many applications, but it is also a challenging problem. This paper introduces a system that tackles this difficult task. The system possesses the ability to detect and track multiple players, estimates the homography between video frames and the court, and identifies the players. The identification system combines three weak visual cues, and exploits both temporal and mutual exclusion constraints in a Conditional Random Field (CRF). In addition, we propose a novel Linear Programming (LP) Relaxation algorithm for predicting the best player identification in a video clip. In order to reduce the number of labeled training data required to learn the identification system, we make use of weakly supervised learning with the assistance of play-by-play texts. Experiments show promising results in tracking, homography estimation, and identification. Moreover, weakly supervised learning with play-by-play texts greatly reduces the number of labeled training examples required. The identification system can achieve similar accuracies by using merely 200 labels in weakly supervised learning, while a strongly supervised approach needs a least 20,000 labels.
Wei-Lwun Lu, Jo-Anne Ting, James J. Little, Kevin Murphy 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Incremental Learning for Video-Based Gait Recognition With LBP Flow
abstract
Gait analysis provides a feasible approach for identification in intelligent video surveillance. However, the effectiveness of the dominant silhouette-based approaches is overly dependent upon background subtraction. In this paper, we propose a novel incremental framework based on optical flow, including dynamics learning, pattern retrieval, and recognition. It can greatly improve the usability of gait traits in video surveillance applications. Local binary pattern (LBP) is employed to describe the texture information of optical flow. This representation is called LBP flow, which performs well as a static representation of gait movement. Dynamics within and among gait stances becomes the key consideration for multiframe detection and tracking, which is quite different from existing approaches. To simulate the natural way of knowledge acquisition, an individual hidden Markov model (HMM) representing the gait dynamics of a single subject incrementally evolves from a population model that reflects the average motion process of human gait. It is beneficial for both tracking and recognition and makes the training process of the HMM more robust to noise. Extensive experiments on widely adopted databases have been carried out to show that our proposed approach achieves excellent performance.
Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001, James J. Little
IEEE Trans. Cybern.5
2013 View-Invariant Discriminative Projection for Multi-View Gait-Based Human Identification
abstract
Existing methods for multi-view gait-based identification mainly focus on transforming the features of one view to the features of another view, which is technically sound but has limited practical utility. In this paper, we propose a view-invariant discriminative projection (ViDP) method, to improve the discriminative ability of multi-view gait features by a unitary linear projection. It is implemented by iteratively learning the low dimensional geometry and finding the optimal projection according to the geometry. By virtue of ViDP, the multi-view gait features can be directly matched without knowing or estimating the viewing angles. The ViDP feature projected from gait energy image achieves promising performance in the experiments of multi-view gait-based identification. We suggest that it is possible to construct a gait-based identification system for arbitrary probe views, by incorporating the information of gallery data with sufficient viewing angles. In addition, ViDP performs even better than the state-of-the-art view transformation methods, which are trained for the combination of gallery and probe viewing angles in every evaluation.
Maodi Hu, Yunhong Wang 0001, Zhaoxiang Zhang 0001, James J. Little, Di Huang 0001
IEEE Trans. Inf. Forensics Secur.4
2012 Combinational Subsequence Matching for Human Identification from General Actions
Maodi Hu, Yunhong Wang 0001, James J. Little
ACCV (3)3
2012 Fine-Grained Categorization for 3D Scene Understanding
abstract
Fine-grained categorization of object classes is receiving increased attention, since it promises to automate classification tasks that are difficult even for humans, such as the distinction between different animal species. In this paper, we consider fine-grained categorization for a different reason: following the intuition that fine-grained categories encode metric information, we aim to generate metric constraints from fine-grained cate-gory predictions, for the benefit of 3D scene-understanding. To that end, we propose two novel methods for fine-grained classification, both based on part information, as well as a new fine-grained category data set of car types. We demonstrate superior performance of our methods to state-of-the-art classifiers, and show first promising results for estimating the depth of objects from fine-grained category predictions from a monocular camera. 1
Michael Stark 0003, Jonathan Krause, Bojan Pepik, David Meger, James J. Little, Bernt Schiele, Daphne Koller
BMVC5
2011 Navigation and obstacle avoidance help (NOAH) for older adults with cognitive impairment: a pilot study
abstract
Many older adults with cognitive impairment are excluded from powered wheelchair use because of safety concerns. This leads to reduced mobility, and in turn, higher dependence on caregivers. In this paper, we describe an intelligent wheelchair that uses computer vision and machine learning methods to provide adaptive navigation assistance to users with cognitive impairment. We demonstrate the performance of the system in a user study with the target population. We show that the collision avoidance module of the system successfully decreases the number of collisions for all participants. We also show that the wayfinding module assists users with memory and vision impairments. We share feedback from the users on various aspects of the intelligent wheelchair system. In addition, we provide our own observations and insights on the target population and their use of intelligent wheelchairs. Finally, we suggest directions for future work.
Pooja Viswanathan, James J. Little, Alan K. Mackworth, Alex Mihailidis
ASSETS2
2011 Explicit Occlusion Reasoning for 3D Object Detection
abstract
Consider the problem of recognizing an object that is partially occluded in an image. The visible portions are likely to match learned appearance models for the object, but hidden portions will not. The (hypothetical) ideal system would consider only the visible object information, correctly ignoring all occluded regions. In purely 2D recognition, this requires inferring the occlusion present, which is a significant challenge since the number of possible occlusion masks is, in principle, exponential. We simplify the problem, considering only a small subset of the most likely occlusions (top, bottom, left, and right halves) and noting that some mismatch is tolerable. We train partial-object detectors tailored exactly to each of these few cases. In addition, we reason about objects in 3D and incorporate sensed geometry, as from an RGB-depth camera, along with visual imagery. This allows explicit occlusion masks to be constructed for each object hypothesis. The masks specify how much to trust each partial template, based on their overlap with visible object regions. Only the visible evidence contributes to our object reasoning.
David Meger, Christian Wojek, James J. Little, Bernt Schiele
BMVC3
2011 Identifying players in broadcast sports videos using conditional random fields
abstract
We are interested in the problem of automatic tracking and identification of players in broadcast sport videos shot with a moving camera from a medium distance. While there are many good tracking systems, there are fewer methods that can identify the tracked players. Player identification is challenging in such videos due to blurry facial features (due to fast camera motion and low-resolution) and rarely visible jersey numbers (which, when visible, are deformed due to player movements). We introduce a new system consisting of three components: a robust tracking system, a robust person identification system, and a conditional random field (CRF) model that can perform joint probabilistic inference about the player identities. The resulting system is able to achieve a player recognition accuracy up to 85% on unlabeled NBA basketball clips.
Wei-Lwun Lu, Jo-Anne Ting, Kevin Murphy 0002, James J. Little
CVPR4
2011 Constrained manipulator visual servoing (CMVS): Rapid robot programming in cluttered workspaces
abstract
This paper presents a model-free optimization framework for the visual servoing of eye-in-hand manipulators in cluttered environments. Visual feedback is used to solve for a set of feasible trajectories that bring the robot end-effector to a target object at a previously untaught location under a number of challenging constraints (i.e., whole-arm collisions, object occlusions, robot's joint limits, camera's sensing limits). A novel controller is proposed, which exploits the natural by-products of the teach-by-showing process, to help the robot navigate this non-convex space. Examining the user-demonstrated trajectories that lead up to the reference image, we use a combination of stochastic optimization techniques and classical optimization techniques to extract the relevant cost functions and constraints for servoing. We hypothesize that we can leverage the user's sensory capabilities and knowledge of the workspace to alleviate the burden of modeling system constraints explicitly. We verify this hypothesis via realistic experiments on a Barrett WAM 7-DOF manipulator equipped with a Sony XC-HR70 camera to show the comparative efficacy of this approach.
Ambrose Chan, Elizabeth A. Croft, James J. Little
IROS3
2011 Mobile 3D object detection in clutter
abstract
This paper presents a method for multi-view 3D robotic object recognition targeted for cluttered indoor scenes. We explicitly model occlusions that cause failures in visual detectors by learning a generative appearance-occlusion model from a training set containing annotated 3D objects, images and point clouds. A Bayesian 3D object likelihood incorporates visual information from many views as well as geometric priors for object size and position. An iterative, sampling-based inference technique determines object locations based on the model. We also contribute a novel robot-collected data set with images and point clouds from multiple views of 60 scenes, with over 600 manually annotated 3D objects accounting for over ten thousand bounding boxes. This data has been released to the community. Our results show that our system is able to robustly recognize objects in realistic scenes, significantly improving recognition performance in clutter.
David Meger, James J. Little
IROS2
2010 Multiple Viewpoint Recognition and Localization
Scott Helmer, David Meger, Marius Muja, James J. Little, David G. Lowe
ACCV (1)4
2010 Viewpoint detection models for sequential embodied object category recognition
abstract
This paper proposes a method for learning viewpoint detection models for object categories that facilitate sequential object category recognition and viewpoint planning. We have examined such models for several state-of-the-art object detection methods. Our learning procedure has been evaluated using an exhaustive multiview category database recently collected for multiview category recognition research. Our approach has been evaluated on a simulator that is based on real images that have previously been collected. Simulation results verify that our viewpoint planning approach requires fewer viewpoints for confident recognition. Finally, we illustrate the applicability of our method as a component of a completely autonomous visual recognition platform that has previously been demonstrated in an object category recognition competition.
David Meger, Ankur Gupta 0004, James J. Little
ICRA3
2010 Path Planning for Improved Visibility Using a Probabilistic Road Map
abstract
This paper focuses on the challenges of vision-based motion planning for industrial manipulators. Our approach is aimed at planning paths that are within the sensing and actuation limits of industrial hardware and software. Building on recent advances in path planning, our planner augments probabilistic road maps with vision-based constraints. The resulting planner finds collision-free paths that simultaneously avoid occlusions of an image target and keep the target within the field of view of the camera. The planner can be applied to eye-in-hand visual-target-tracking tasks for manipulators that use point-to-point commands with interpolated joint motion.
Matthew A. Baumann, Simon Léonard, Elizabeth A. Croft, James J. Little
IEEE Trans. Robotics4
2009 Planning collision-free and occlusion-free paths for industrial manipulators with eye-to-hand configuration
abstract
This paper presents a motion planning algorithm for industrial manipulators with the simultaneous constraints of avoiding collisions and avoiding the occlusion of specified pixellated regions of an eye-to-hand camera. The system uses a probabilistic roadmap to satisfy the constraints imposed by the command interface of typical industrial manipulators and uses dynamic collision checking to ensure collision-free motion. In the context of a task monitored by a camera, we enhance a probabilistic roadmap with a dynamic occlusion checking algorithm that is able to determine which pixels of the camera are occluded by the robot during each motion segment. The occlusion algorithm is formulated as collision algorithm where the field of view of the camera is represented as a quadtree of frustums. The proposed algorithm is demonstrated in industrial bin picking simulations where the gripper must not occlude the targeted object throughout the task.
Simon Léonard, Elizabeth A. Croft, James J. Little
IROS3
2009 Tracking and recognizing actions of multiple hockey players using the boosted particle filter
Wei-Lwun Lu, Kenji Okuma, James J. Little
Image Vis. Comput.3
2009 Autonomous vision-based robotic exploration and mapping using hybrid maps and particle filters
Robert Sim, James J. Little
Image Vis. Comput.2
2009 A Hybrid Conditional Random Field for Estimating the Underlying Ground Surface From Airborne LiDAR Data
abstract
Recent advances in airborne light detection and ranging (LiDAR) technology allow rapid and inexpensive generation of digital surface models (DSMs), 3-D point clouds of buildings, vegetations, cars, and natural terrain features over large regions. However, in many applications, such as flood modeling and landslide prediction, digital terrain models (DTMs), the topography of the bare-Earth surface, are needed. This paper introduces a novel machine learning approach to automatically extract DTMs from their corresponding DSMs. We first classify each point as being either ground or nonground, using supervised learning techniques applied to a variety of features. For the points which are classified as ground, we use the LiDAR measurements as an estimate of the surface height, but, for the nonground points, we have to interpolate between nearby values, which we do using a Gaussian random field. Since our model contains both discrete and continuous latent variables, and is a discriminative (rather than generative) probabilistic model, we call it ahybridconditionalrandomfield. We show that a MaximumaPosterioriestimate of the surface height can be efficiently estimated by using a variant of the Expectation Maximization algorithm. Experiments demonstrate that the accuracy of this learning-based approach outperforms the previous best systems, based on manually tuned heuristics.
Wei-Lwun Lu, Kevin Murphy 0002, James J. Little, Alla Sheffer, Hongbo Fu 0001
IEEE Trans. Geosci. Remote. Sens.3
2008 Trajectory specification via sparse waypoints for eye-in-hand robots requiring continuous target visibility
abstract
This paper presents several methods of managing field of view constraints of an eye-in-hand system for vision- based pose control with limited controller input. Herein, the possible inverse kinematic solutions for a desired relative camera pose are evaluated to determine whether the interpolated trajectories satisfy field of view constraints for the target of interest. If no immediately feasible trajectory exists, additional waypoints are specified to guide the robot towards its goal while maintaining visibility. The insertion of an additional visible and feasible waypoint divides the problem into two sub-problems of the same form, but of lesser difficulty by reducing the robot's interpolation distance. Virtual image-based visual servoing (IBVS) is used to generate an ideal image trajectory to guide the selection of waypoints. A damped least- squares inverse kinematics solution is implemented to handle robot singularities. The methods are simulated for a CRS-A465 robot with a Sony XC-HR70 camera.
Ambrose Chan, Elizabeth A. Croft, James J. Little
ICRA3
2008 Informed visual search: Combining attention and object recognition
abstract
This paper studies the sequential object recognition problem faced by a mobile robot searching for specific objects within a cluttered environment. In contrast to current state-of-the-art object recognition solutions which are evaluated on databases of static images, the system described in this paper employs an active strategy based on identifying potential objects using an attention mechanism and planning to obtain images of these objects from numerous viewpoints. We demonstrate the use of a bag-of-features technique for ranking potential objects, and show that this measure outperforms geometric matching for invariance across viewpoints. Our system implements informed visual search by prioritising map locations and re-examining promising locations first. Experimental results demonstrate that our system is a highly competent object recognition system that is capable of locating numerous challenging objects amongst distractors.
Per-Erik Forssén, David Meger, Scott Helmer, James J. Little, David G. Lowe
ICRA5
2008 Dynamic visibility checking for vision-based motion planning
abstract
An important problem in position-based visual servoing (PBVS) is to guarantee that a target will remain within the field of view for the duration of the task. In this paper, we propose a dynamic visibility checking algorithm that, given a parametrized trajectory of the camera, determines if an arbitrary 3D target will remain within the field of view. We reformulate this problem as the problem of determining if the 3D coordinates of the target collide with the frustum formed by the camera field of view during the camera trajectory. To solve this problem, our algorithm computes and compares the shortest distance between the target and the frustum with the length of the trajectory described by the target in the camera's coordinate frame. Furthermore, we demonstrate that our algorithm can be combined with path planning algorithms and, in particular, probabilistic roadmaps (PRM). Results suggest that our algorithm is computationally efficient even when the target moves in the vicinity of image borders. In simulations, we use our dynamic visibility checking algorithm in conjunction with a PRM to plan collision free paths while providing the guarantee that a specific target will not leave the field of view.
Simon Léonard, Elizabeth A. Croft, James J. Little
ICRA3
2008 Occlusion-free path planning with a probabilistic roadmap
abstract
We present a novel algorithm for path planning that avoids occlusions of a visual target for an ldquoeye-in-handrdquo sensor on an articulated robot arm. We compute paths using a probabilistic roadmap to avoid collisions between the robot and obstacles, while penalizing trajectories that do not maintain line-of-sight. The system determines the space from which line-of-sight is unimpeded to the target (the visible region). We assign penalties to trajectories within the roadmap proportional to the distance the camera travels while outside the visible region. Using Dijkstrapsilas algorithm, we compute paths of minimal occlusion (maximal visibility) through the roadmap. In our experiments, we compare a shortest-distance path to the minimal-occlusion path and discuss the impact of the improved visibility.
Matthew A. Baumann, Donna C. Dupuis, Simon Léonard, Elizabeth A. Croft, James J. Little
IROS5
2008 Optimizing Multiple Object Tracking and Best View Video Synthesis
abstract
We study schemes to tackle problems of optimizing multiple object tracking and best-view video synthesis. A novel linear relaxation method is proposed for the class of multiple object tracking problems where the inter-object interaction metric is convex and the intra-object term quantifying object state continuity may use any metric. This scheme models object tracking as multi-path searching. It explicitly models track interaction, such as object spatial layout consistency or mutual occlusion, and optimizes multiple object tracks simultaneously. The proposed scheme does not rely on track initialization and complex heuristics. It has much less average complexity than previous efficient exhaustive search methods such as extended dynamic programming and can find the global optimum with high probability. Given the tracking data from our method, optimizing best-view video synthesis using multiple-view videos is further studied, which is formulated as a recursive decision problem and optimized by a dynamic programming approach. The proposed object tracking and best-view synthesis methods have found successful applications in MyView - a system to enhance media content presentation of multiple-view video.
Hao Jiang 0007, Sidney S. Fels, James J. Little
IEEE Trans. Multim.3
2007 The UBC Semantic Robot Vision System
Scott Helmer, David Meger, Per-Erik Forssén, Tristram Southey, Sancho McCann, Pooyan Fazli, James J. Little, David G. Lowe
AAAI7
2007 A Linear Programming Approach for Multiple Object Tracking
abstract
We propose a linear programming relaxation scheme for the class of multiple object tracking problems where the inter-object interaction metric is convex and the intra-object term quantifying object state continuity may use any metric. The proposed scheme models object tracking as a multi-path searching problem. It explicitly models track interaction, such as object spatial layout consistency or mutual occlusion, and optimizes multiple object tracks simultaneously. The proposed scheme does not rely on track initialization and complex heuristics. It has much less average complexity than previous efficient exhaustive search methods such as extended dynamic programming and is found to be able to find the global optimum with high probability. We have successfully applied the proposed method to multiple object tracking in video streams.
Hao Jiang 0007, Sidney S. Fels, James J. Little
CVPR3
2007 OpenVL: Towards A Novel Software Architecture for Computer Vision
abstract
This paper presents our progress on OpenVL -a novel software architecture to address efficiency through facilitating hardware acceleration, reusability and scalability for computer vision. A logical image understanding pipeline is introduced to allow parallel processing. As well, we discuss our middleware -VLUT that enables applications to operate transparently over a heterogeneous collection of hardware implementations. OpenVL works as a state machine, with an event-driven mechanism to provide users with application-level interaction. Various explicit or implicit synchronization and communication methods are supported among distributed processes in the logical pipelines. The intent of OpenVL is to allow users to quickly and easily recover useful information from multiple scenes across various software environments and hardware platforms. We implement two different human tracking systems to validate the critical underlying concepts of OpenVL.
Changsong Shen, Sidney S. Fels, James J. Little
CVPR3
2007 Decision theoretic task coordination for a visually-guided interactive mobile robot
abstract
In this paper, we present a visually-guided mobile robot that is capable of executing a task requiring complex human-robot interaction (HRI). The robot delivers verbal messages among the inhabitants of an office-like environment. Essential to the robot's robust performance is our behavior-based robot control architecture enhanced with a state of the art decision theoretic planner that takes into account the temporal characteristics of the robot's actions. The decision theoretic layer is based on the partially observable Markov decision process (POMDP) framework allowing us to achieve principled coordination of complex subtasks implemented as robot behaviors/skills. We compute approximate POMDP policies using the randomized point-based value iteration algorithm and we present heuristics for improving its computational efficiency.
Pantelis Elinas, James J. Little
IROS2
2007 A Study of the Rao-Blackwellised Particle Filter for Efficient and Accurate Vision-Based SLAM
Robert Sim, Pantelis Elinas, James J. Little
Int. J. Comput. Vis.3
2007 Value-Directed Human Behavior Analysis from Video Using Partially Observable Markov Decision Processes
abstract
This paper presents a method for learning decision theoretic models of human behaviors from video data. Our system learns relationships between the movements of a person, the context in which they are acting, and a utility function. This learning makes explicit that the meaning of a behavior to an observer is contained in its relationship to actions and outcomes. An agent wishing to capitalize on these relationships must learn to distinguish the behaviors according to how they help the agent to maximize utility. The model we use is a partially observable Markov decision process, or POMDP. The video observations are integrated into the POMDP using a dynamic Bayesian network that creates spatial and temporal abstractions amenable to decision making at the high level. The parameters of the model are learned from training data using an a posteriori constrained optimization technique based on the expectation-maximization algorithm. The system automatically discovers classes of behaviors and determines which are important for choosing actions that optimize over the utility of possible outcomes. This type of learning obviates the need for labeled data from expert knowledge about which behaviors are significant and removes bias about what behaviors may be useful to recognize in a particular situation. We show results in three interactions: a single player imitation game, a gestural robotic control problem, and a card game played by two people.
Jesse Hoey, James J. Little
IEEE Trans. Pattern Anal. Mach. Intell.2
2006 Robust Visual Tracking for Multiple Targets
Yizheng Cai, Nando de Freitas, James J. Little
ECCV (4)3
2006 σSLAM: Stereo Vision SLAM using the Rao-Blackwellised Particle Filter and a Novel Mixture Proposal Distribution
abstract
We consider the problem of simultaneous localization and mapping (SLAM) using the Rao-Blackwellised particle filter (RBPF) for the class of indoor mobile robots equipped only with stereo vision. Our goal is to construct dense metric maps of natural 3D point landmarks for large cyclic environments in the absence of accurate landmark position measurements and motion estimates. Our work differs from other approaches because landmark estimates are derived from stereo vision and motion estimates are based on sparse optical flow. We distinguish between landmarks using the scale invariant feature transform (SIFT). This is in contrast to current popular approaches that rely on reliable motion models derived from odometric hardware and accurate landmark measurements obtained with laser sensors. Since our approach depends on a particle filter whose main component is the proposal distribution, we develop and evaluate a novel mixture proposal distribution that allows us to robustly close large loops. We validate our approach experimentally for long camera trajectories processing thousands of images at reasonable frame rates
Pantelis Elinas, Robert Sim, James J. Little
ICRA3
2006 Autonomous vision-based exploration and mapping using hybrid maps and Rao-Blackwellised particle filters
abstract
This paper addresses the problem of exploring and mapping an unknown environment using a robot equipped with a stereo vision sensor. The main contribution of our work is a fully automatic mapping system that operates without the use of active ranger sensors (such as laser or sonic transducers), can operate in real-time and can consistently produce accurate maps of large-scale environments. Our approach implements a Rao-Blackwellised particle filter (RBPF) to solve the simultaneous localization and mapping problem and uses efficient data structures for real-time data association, mapping, and spatial reasoning. We employ a hybrid map representation that infers 3D point landmarks from image features to achieve precise localization, coupled with occupancy grids for safe navigation. This paper describes our framework and implementation, and presents our exploration method, and experimental results illustrating the functionality of the system
Robert Sim, James J. Little
IROS2
2005 Vision-based global localization and mapping for mobile robots
abstract
We have previously developed a mobile robot system which uses scale-invariant visual landmarks to localize and simultaneously build three-dimensional (3-D) maps of unmodified environments. In this paper, we examine global localization, where the robot localizes itself globally, without any prior location estimate. This is achieved by matching distinctive visual landmarks in the current frame to a database map. A Hough transform approach and a RANSAC approach for global localization are compared, showing that RANSAC is much more efficient for matching specific features, but much worse for matching nonspecific features. Moreover, robust global localization can be achieved by matching a small submap of the local region built from multiple frames. This submap alignment algorithm for global localization can be applied to map building, which can be regarded as alignment of multiple 3-D submaps. A global minimization procedure is carried out using the loop closure constraint to avoid the effects of slippage and drift accumulation. Landmark uncertainty is taken into account in the submap alignment and the global minimization process. Experiments show that global localization can be achieved accurately using the scale-invariant landmarks. Our approach of pairwise submap alignment with backward correction in a consistent manner produces a better global 3-D map.
Stephen Se, David G. Lowe, James J. Little
IEEE Trans. Robotics3
2004 Value Directed Learning of Gestures and Facial Displays
Jesse Hoey, James J. Little
CVPR (2)2
2004 Decision Theoretic Modeling of Human Facial Displays
Jesse Hoey, James J. Little
ECCV (3)2
2004 A Boosted Particle Filter: Multitarget Detection and Tracking
Kenji Okuma, Ali Taleghani, Nando de Freitas, James J. Little, David G. Lowe
ECCV (1)4
2004 Environment modeling with stereo vision
abstract
We consider the problem of creating compact surface-based environment models from stereo vision images taken from a stereo-camera equipped mobile robot. The stereo images can be quite complex and correlation stereo suffers from considerable noise at ranges over a few metres. We construct the environment models by segmenting the scene viewed from a stereo camera into rectangular planar surfaces through the use of the patchlets surface element data structure. Patchlets are the projection of the stereo pixels onto detected surfaces in the scene. They have position, orientation, size and sensor-based confidence measures. The confidence measures allow proper weighting of patchlet parameters when aggregating patchlets into larger surfaces.
Don Ray Murray, James J. Little
IROS2
2003 Bayesian Clustering of Optical Flow Fields
abstract
We present a method for unsupervised learning of classes of motions in video. We project optical flow fields to a complete, orthogonal, a-priori set of basis functions in a probabilistic fashion, which improves the estimation of the projections by incorporating uncertainties in the flows. We then cluster the projections using a mixture of feature-weighted Gaussians over optical flow fields. The resulting model extracts a concise probabilistic description of the major classes of optical flow present. The method is demonstrated on a video of a person's facial expressions.
Jesse Hoey, James J. Little
ICCV2
2003 Ordering Points for Incremental TIN Construction from DEMs
James J. Little
GeoInformatica1
2002 Waiting with José, a Vision-Based Mobile Robot
abstract
Jose is a visually guided autonomous robotic waiter. He circulates around a room populated by groups of people, politely serving appetizers to humans. The serving task combines elements of robotics with human computer interaction, challenging control architecture with multiple task integration. This paper describes our purely vision-based approach to this task. Methods for mapping, localization and navigation are presented and discussed, including issues of safety for both robots and humans. Our work on human-robot interaction is covered, as well as our solutions to various tasks specific to serving food. We present results of our methods from sample experiments in our laboratory. We further discuss our experiences at the 2001 AAAI mobile robot "Hors D'oeuvres Anyone?" competition, at which Jose took first prize.
Pantelis Elinas, Jesse Hoey, Darrell Lahey, Jefferson D. Montgomery, Don Ray Murray, Stephen Se, James J. Little
ICRA7
2002 Vision-based mapping with backward correction
abstract
We consider the problem of creating a consistent alignment of multiple 3D submaps containing distinctive visual landmarks in an unmodified environment. An efficient map alignment algorithm based on landmark specificity is proposed to align submaps. This is followed by a global minimization using the close-the-loop constraint. Landmark uncertainty is taken into account in the pairwise alignment and the global minimization process. Experiments show that the pairwise alignment of submaps with backward correction produces a consistent global 3D map. Our vision-based mapping approach using sparse 3D data is different from other existing approaches which use dense 2D range data from laser or sonar rangefinders.
Stephen Se, David G. Lowe, James J. Little
IROS3
2002 Global localization using distinctive visual features
abstract
We have previously developed a mobile robot system which uses scale invariant visual landmarks to localize and simultaneously build a 3D map of the environment In this paper, we look at global localization, also known as the kidnapped robot problem, where the robot localizes itself globally, without any prior location estimate. This is achieved by matching distinctive landmarks in the current frame to a database map. A Hough transform approach and a random sample consensus (RANSAC) approach for global localization are compared, showing that RANSAC is much more efficient. Moreover, robust global localization can be achieved by matching a small sub-map of the local region built from multiple frames.
Stephen Se, David G. Lowe, James J. Little
IROS3
2001 Vision-based Mobile Robot Localization And Mapping using Scale-Invariant Features
abstract
A key component of a mobile robot system is the ability to localize itself accurately and build a map of the environment simultaneously. In this paper, a vision-based mobile robot localization and mapping algorithm is described which uses scale-invariant image features as landmarks in unmodified dynamic environments. These 3D landmarks are localized and robot ego-motion is estimated by matching them, taking into account the feature viewpoint variation. With our Triclops stereo vision system, experiments show that these features are robustly matched between views, 3D landmarks are tracked, robot pose is estimated and a 3D map is built.
Stephen Se, David G. Lowe, James J. Little
ICRA3
2001 Local and global localization for mobile robots using visual landmarks
abstract
Our mobile robot system uses scale-invariant visual landmarks to localize itself and build a 3D map of the environment simultaneously. As image features are not noise-free, we carry out error analysis and use Kalman filters to track the 3D landmarks, resulting in a database map with landmark positional uncertainty. By matching a set of landmarks as a whole, our robot can localize itself globally based on the database containing landmarks of sufficient distinctiveness. Experiments show that recognition of position within a map without any prior estimate can be achieved using the scale-invariant landmarks.
Stephen Se, David G. Lowe, James J. Little
IROS3
2001 Structural Lines, TINs, and DEMs
James J. Little
Algorithmica1
2000 Representation and Recognition of Complex Human Motion
abstract
The quest for a vision system capable of representing and recognizing arbitrary motions benefits from a low dimensional, non-specific representation of flow fields, to be used in high level classification tasks. We present Zernike polynomials as an ideal candidate for such a representation. The basis of Zernike polynomials is complete and orthogonal and can be used for describing many types of motion at many scales. Starting from image sequences, locally smooth image velocities are derived using a robust estimation procedure, from which are computed compact representations of the flow using the Zernike basis. Continuous density hidden Markov models are trained using the temporal sequences of vectors thus obtained, and are used for subsequent classification. We present results of our method applied to image sequences of facial expressions both with and without significant rigid head motion and to sequences of lip motion from a known database. We demonstrate that the Zernike representation yields results competitive with those obtained using principal components, while not committing to specific types of motion. It is therefore ideal as a fundamental building block for a vision system capable of classifying arbitrary motion types.
Jesse Hoey, James J. Little
CVPR2
2000 Deforming Surface Features Lines in Intrinsic Coordinates
abstract
Significant local structure of terrain surfaces can be described by structural lines which are connected sets of points where the surface is approximately cylindrical, i.e., the ratio of the principal curvatures is large. At each point curvature is maximal in the curvature direction associated with the curvature with larger absolute value. These lines form the skeleton of the surface for constructing triangulated approximations. Significant structures are best identified at coarse scales but need to be deformed to fine scale before use. Standard snake algorithms using proximity in the image plane do not suffice. Earlier work used the maximal curvature field as an intrinsic (curvature) coordinate system. This leads to shrinkage of the open curves under internal forces. A better solution is to restrict the movement of the lines to the principal curvature directions forming an intrinsic coordinate system over the surface. We evaluate the results of deforming lines by utilizing the deformed lines as structural lines for triangulation. Points at coarse scale move along curvature lines to the proper location at fine scale. The resulting fine scale lines are better positioned for terrain representation than those derived from proximity alone and produce more compact triangulations.
James J. Little
ICPR1
1999 Cooperative Robot Localization with Vision-Based Mapping
abstract
Two stereo vision-based mobile robots navigate and autonomously explore their environment safely while building occupancy grid maps of the environment. A novel landmark recognition system allows one robot to automatically find suitable landmarks in the environment. The second robot uses these landmarks to localize itself relative to the first robot's reference frame, even when the current state of the map is incomplete. The robots have a common local reference frame so that they can collaborate on tasks, without having a prior map of the environment. Stereo vision processing and map updates are done at 5 Hz and the robots move at 200 cm/s. Using occupancy grids the robots can robustly explore unstructured and dynamic environments. The map is used for path planning and landmark detection. Landmark detection uses the map's corner features and least-squares optimization to find the transformation between the robots' coordinate frames. The results provide very accurate relative localization without requiring highly accurate sensors. Accuracy of better than 2 cm was achieved in experiments.
Cullen Jennings, Don Ray Murray, James J. Little
ICRA3
1999 Reflectance and Shape from Images Using a Collinear Light Source
Jiping Lu, James J. Little
Int. J. Comput. Vis.2
1998 Selecting stable image features for robot localization using stereo
abstract
To navigate and recognize where it is, a mobile robot must be able to identify its current location. In an unknown initial position, a robot needs to refer to its environment to determine its location in an external coordinate system. Even with a known initial position, drift in odometry causes the estimated position to deviate from the correct position, requiring correction. We show how to find landmarks without models. We use dense stereo data from our mobile robot's trinocular system to discover image regions that will be stable over widely differing viewpoints. We find image brightness "corners" in images and select those that do not straddle depth discontinuities in the stereo depth data. Selecting corners only in regions of nearly planar stereo data results in landmarks that can be seen in images taken from different viewpoints.
James J. Little, Jiping Lu, Don Ray Murray
IROS1
1998 Structural lines for triangulations of terrain
abstract
To build triangulated approximations to terrain surfaces from dense elevation models, we find structural lines for the initial skeleton of the triangulation. We describe the surfaces as regions; the spines of these regions are the structural lines on the surface. We use these lines as a skeleton of the surface, points and edges that initialize the triangulation. Simple local curvature analysis of the terrain results in many lines whose significance is only local. We show how the curvature analysis can be extended to find local regions around the structural lines. We use various properties of the regions to assign a significance to the lines, to rank them for inclusion in the skeleton. We then use the lines to approximate the surface with a compact triangulation.
James J. Little
WACV1
1996 Geometric and Photometric Constraints for Surface Recovery
abstract
In this paper we present a novel approach to surface recovery from an image sequence of a rotating object. In this approach, the object is illuminated under a collinear light source (where the light source lies on or near the optical axis) and rotated on a controlled turntable. A wire-frame of 3D curves on the object surface is extracted by using shading and occluding contours in the image sequence. Then the whole object surface is recovered by interpolating the surface between curves on the wire-frame. The interpolation can be done by using geometric or photometric constraints. The photometric method uses shading information and is more powerful than geometric methods. The experimental results on real image sequence of matte and specular surfaces show that the technique is feasible and promising.
Jiping Lu, James J. Little
CVPR2
1995 Reflectance Function Estimation and Shape Recovery from Image Sequence of a Rotating Object
abstract
We describe a technique for surface recovery of a rotating object illuminated under a collinear light source (where the light source lies on or near the optical axis). We show that the surface reflectance function can be directly estimated from the image sequence without any assumption on the reflectance property of the object surface. From the image sequence, the 3D locations of some singular surface points are calculated and their brightness values are extracted for the estimation of the reflectance function. We also show that the surface can be recovered by using shading information in two images of the rotating object. Iteratively using the first-order Taylor series approximation and the estimated reflectance function, the depth and orientation of the surface can be recovered simultaneously. The experimental results on real image sequences of both matte and specular surfaces demonstrate that the technique is feasible and robust.>
Jiping Lu, James J. Little
ICCV2
1994 Complementary data fusion for limited-angle tomography
abstract
Ambiguity in the solution of inverse problems arises when data are insufficient to define a unique solution (i.e., the problem is ill-posed). Data fusion has the potential to reduce this ambiguity by using other sensory data that complement the original data. This paper examines the application of data fusion to limited-angle computed tomography (CT) to resolve ambiguity. While CT in its conventional form is ill-posed with a small null space, limited-angle CT has a much larger null space. Structures that lie primarily in the null space of the limited-angle Radon transform are particularly prone to ambiguity. We describe a novel constraint-based data fusion system that fuses spatial support and ultrasound measurements with x-ray data. The ensuing problem is less ambiguous, has a reduced null space, and permits accurate reconstruction of a sandwich structure where otherwise impossible.>
Jeffrey E. Boyd, James J. Little
CVPR2
1994 Cooperative Analysis of Multiple Frames by Visual Echoes
abstract
Many computational vision tasks, such as trinocular stereo disparity calculations, multiple baseline stereo measurements, multi-frame motion analysis, motion-stereo (binocular or trinocular) techniques, vergence control, and structure from uniform three-dimensional acceleration, involve determination of disparities among multiple frames. The paper introduces a uniform approach to multiframe analysis where equal disparities between multiple frames reinforce one another in time and space. Consequently, in the presence of constant displacements, the increased number of frames leads to an increase in the detection and accuracy of disparity estimation. In the case of unequal separations-for instance due to acceleration-disparities between every two frames are calculated, providing a rate of change in disparities.>
Esfandiar Bandari, James J. Little
ICIP (3)2
1994 Vision servers and their clients
abstract
Robotic applications impose hard real-time demands on their vision components. To accommodate the realtime constraints, the visual component of robotic systems are often simplified by narrowing the scope of the vision system for a particular task. Another option is to build a generalized vision (sensor) processor and provides multiple interfaces, of differing scales and content, to other modules in the robot. Both options can be implemented in many ways, depending on computational resources. The tradeoffs among these alternatives become clear when we study the vision process as a server whose clients request information about the world. We model the interface on client-server relations in user interfaces and operating systems. We examine the relation of this model to robot and vision sensor architecture and explore its application to a variety of vision sensor implementations.
James J. Little
ICPR (3)1
1994 A unified recognition and stereo vision system for size assessment of fish
abstract
This paper presents a unified recognition and stereo vision system which locates objects and determines their distances and sizes given stereo video input. Unlike other such systems, the recognition stage precedes and provides input to stereo processing. Model-based recognition is accomplished in two stages. The first stage seeks feature matches by comparing the absolute orientation, relative orientation and relative length of each image and model segments to find matching chains of segments. The second stage verifies candidate matches by comparing the relative locations of matched image features and corresponding model features. Models are generated semi-automatically from images of the desired objects. In addition to providing distance estimates, feature-based stereo information is used to disambiguate multiple or questionable matches. Although quite general, the system is described in the context of its motivating task of assessing the size of sea-cage salmon non-invasively.>
Andrew Naiberg, James J. Little
WACV2
1994 Parallel Solutions to Geometric Problems in the Scan Model of Computation
Guy E. Blelloch, James J. Little
J. Comput. Syst. Sci.2
1993 Visual echo analysis
abstract
The term visual echoes is introduced as a common framework for the analysis of multi-frame optical flow, binocular and trinocular stereo, stationary texture and boundary symmetries. The authors examined cepstral filtering, a powerful nonlinear adaptive technique for the retrieval of echoes, as a common methodology to address these visual routines. They consider the application of cepstral analysis to computational vision, review improvements to traditional methods, and provide a comparison with other routines presently used. A general multievidential correlation approach is introduced which lends itself to several computational techniques. CepsCorr, it is called, is a simple general technique that can accept different matching routines as its measurement kernel. The evidence provided by each iteration of cepsCorr can then be combined to provide a more accurate estimate of motion or binocular disparity.>
Esfandiar Bandari, James J. Little
ICCV2
1993 Silt: A distributed bit-parallel architecture for early vision
Michael Bolotski, Rod Barman, James J. Little, Daniel Camporese
Int. J. Comput. Vis.3
1992 Spatial-quefrency approach to optical echo analysis
abstract
A methodology for optical flow analysis based on cepstral filtering is introduced. The power cepstrum is extended to multiframe analysis. A correlative cepstral technique, cepsCorr, is developed. It significantly increases the signal-to-noise ratio, reduces ambiguities, and it provides a predictive or multievidence approach to visual motion analysis.>
Esfandiar Bandari, James J. Little
CVPR2
1990 Direct Evidence for Occlusion in Stereo and Motion
James J. Little, Walter E. Gillett
ECCV1
1990 Silt: the bit-parallel approach
abstract
A particular form of parallelism, called bit-parallelism, is introduced. A bit-parallel organization distributes each bit of a data item to a different processor. Bit-parallelism allows computation that is sublinear with word size for such operations as integer addition, arithmetic shifts, and data moves. The implications of bit-parallelism for system architecture are analyzed. An implementation of a bit-parallel architecture based on a mesh with a bypass network is presented. Using a conservative estimate for cycle time, a Silt processor performs 64-b integer additions more than 10 times faster than the Connection Machine-2. Using current CMOS technology, a 16 M processor Silt system would be capable of nearly 500 billion 32-b adds per second. The application of the architecture to low-level vision algorithms is discussed.>
Rod Barman, Michael Bolotski, Daniel Camporese, James J. Little
ICPR (2)4
1990 Direct evidence for occlusion in stereo and motion
James J. Little, Walter E. Gillett
Image Vis. Comput.1
1989 Algorithmic Techniques for Computer Vision on a Fine-Grained Parallel Machine
abstract
The authors describe several fundamentally useful primitive operations and routines and illustrate their usefulness in a wide range of familiar version processes. These operations are described in terms of a vector machine model of parallel computation. They use a parallel vector model because vector models can be mapped onto a wide range of architectures. They also describe implementing these primitives on a particular fine-grained machine, the connection machine. It is found that these primitives are applicable in a variety of vision tasks. Grid permutations are useful in many early vision algorithms, such as Gaussian convolution, edge detection, motion, and stereo computation. Scan primitives facilitate simple, efficient solutions of many problems in middle- and high-level vision. Pointer jumping, using permutation operations, permits construction of extended image structures in logarithmic time. Methods such as outer products, which rely on a variety of primitives, play an important role of many high-level algorithms.>
James J. Little, Guy E. Blelloch, Todd A. Cass
IEEE Trans. Pattern Anal. Mach. Intell.1
1988 Parallel Optical Flow Using Local Voting
abstract
We describe a parallel algorithm for computing optical flow from short-range motion. Regularizing optical flow computation leads to a forruulation which minimizes matching error and, at the same time, maximises smoothness of the optical flow. We develop an approximation to the full regularization computation in which corresponding points are found by comparing local patches of the images. Selection aniong competing matches is performed using a winner-take-all scheme. The algorithm accommodates many different image transformations uniformly, with siniilar results, from brightness to edges. The optical flow computed froni different image transformations, such as edge detection and direct brightness computation, can be simply combined. The algorithm is easily implemented using local operations on a finegrained computer, and has been implemented on a Connection Machine. Experiments with natural images show that the scheme is effective and robust against noise. The algorithm leads to dense optical flow fields ; in addition, inforniation from matching facilitates segmentation.
James J. Little, Heinrich H. Bülthoff, Tomaso A. Poggio
ICCV1
1988 Neural mapping and parallel optical flow computation for autonomous navigation
Heinrich H. Bülthoff, James J. Little, Hanspeter A. Mallot
Neural Networks2
1985 Extended Gaussian images, mixed volumes, shape reconstruction
abstract
The Extended Gaussian Image (EGI) of an object records the variation of surface area with surface orientation. The EGI is a unique representation for convex objects. For a polyhedron, each face is represented by its normal and its area. The inversion problem (from an EGI to a description in terms of vertices and faces) is solved for convex polyhedra, by providing an algorithm giving an iterative solution by a minimization[Little,1983]. The algorithm employs a geometric construction, the mixed volume, which was used in Minkowski's proof [1897] of the existence and uniqueness of an inverse. The mixed volume measures similarity of shape for convex objects.
James J. Little
SCG1
1985 Determining Object Attitude from Extended Gaussian Images
James J. Little
IJCAI1
1983 An Iterative Method for Reconstructing Convex Polyhedra From External Guassian Images
James J. Little
AAAI1
1979 Automatic extraction of Irregular Network digital terrain models
abstract
For representation of terrain, an efficient alternative to dense grids is the Triangulated Irregular Network (TIN), which represents a surface as a set of non-overlapping contiguous triangular facets, of irregular size and shape. The source of digital terrain data is increasingly dense raster models produced by automated orthophoto machines or by direct sensors such as synthetic aperture radar. A method is described for automatically extracting a TIN model from dense raster data. An initial approximation is constructed by automatically triangulating a set of feature points derived from the raster model. The method works by local incremental refinement of this model by the addition of new points until a uniform approximation of specified tolerance is obtained. Empirical results show that substantial savings in storage can be obtained.
Robert J. Fowler, James J. Little
SIGGRAPH2