EDBT 2026 Demo / reviewers in the wild / expert
James L. Crowley
dblp:50/3289
· DBLP profile ↗
93ranked-venue papers
31as first author
6since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 69 · 24 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 12 first-author · 2 since 2021Systems, architecture and hardware · 17 · 7 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 12 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
32 papers |
Segmentation and scene understanding · 45% Image recognition and object detection · 28% 3D vision · 14% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-robot interaction · 64% Ubiquitous computing and smart environments · 36% |
Topics — the 30 heaviest of 71, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › object discovery
unsupervised object discovery |
1.2 | 2 | 2023 | TokenCut: Segmenting Objects in Images and Videos With Self-Supervised Transformer and Normalized Cut · IEEE Trans. Pattern Anal. Mach. Intell. 2023 Self-Supervised Transformers for Unsupervised Object Discovery using Normalized Cut · CVPR 2022 |
Computer vision › Segmentation and scene understanding
object segmentation |
0.7 | 1 | 2023 | TokenCut: Segmenting Objects in Images and Videos With Self-Supervised Transformer and Normalized Cut · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection |
0.7 | 1 | 2023 | TokenCut: Segmenting Objects in Images and Videos With Self-Supervised Transformer and Normalized Cut · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › Segmentation and scene understanding › image segmentation › graph-based segmentation
normalized cuts |
0.6 | 1 | 2022 | Self-Supervised Transformers for Unsupervised Object Discovery using Normalized Cut · CVPR 2022 |
Computer vision › Segmentation and scene understanding › saliency detection
unsupervised saliency detection |
0.6 | 1 | 2022 | Self-Supervised Transformers for Unsupervised Object Discovery using Normalized Cut · CVPR 2022 |
Computer vision › Video understanding and tracking
video object segmentation |
0.2 | 1 | 2023 | TokenCut: Segmenting Objects in Images and Videos With Self-Supervised Transformer and Normalized Cut · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › 3D vision
scene flow estimation |
0.2 | 1 | 2014 | Dense Semi-rigid Scene Flow Estimation from RGBD Images · ECCV (7) 2014 |
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection |
0.2 | 1 | 2022 | Self-Supervised Transformers for Unsupervised Object Discovery using Normalized Cut · CVPR 2022 |
Computer vision › Image recognition and object detection
object recognition |
0.1 | 4 | 2000 | Recognition without Correspondence using Multidimensional Receptive Field Histograms · Int. J. Comput. Vis. 2000 Object Recognition Using Coloured Receptive Fields · ECCV (1) 2000 Transinformation for Active Object Recognition · ICCV 1998 |
Human-robot interaction
social robot |
0.1 | 1 | 2008 | Learning polite behavior with situation models · HRI 2008 |
Computer vision › 3D vision › motion estimation
displacement estimation |
0.1 | 1 | 2014 | Dense Semi-rigid Scene Flow Estimation from RGBD Images · ECCV (7) 2014 |
Computer vision › Video understanding and tracking
activity recognition |
0.1 | 2 | 2000 | A Probabilistic Sensor for the Perception and Recognition of Activities · ECCV (1) 2000 Probabilistic Recognition of Activity using Local Appearance · CVPR 1999 |
Robotics › Robot navigation and mapping › localization
position estimation |
0.0 | 3 | 1998 | Position Estimation Using Principal Components of Range Data · ICRA 1998 A Comparison of Position Estimation Techniques Using Occupancy Grids · ICRA 1994 World modeling and position estimation for a mobile robot using ultrasonic ranging · ICRA 1989 |
Ubiquitous computing and smart environments
context-aware computing |
0.0 | 1 | 2002 | Perceptual Components for Context Aware Computing · UbiComp 2002 |
Robotics › Robot navigation and mapping › localization › robot localization
mobile robot localization |
0.0 | 3 | 1998 | Position Estimation Using Principal Components of Range Data · ICRA 1998 Position estimation for a mobile robot using vision and odometry · ICRA 1992 World modeling and position estimation for a mobile robot using ultrasonic ranging · ICRA 1989 |
Robotics › Robot navigation and mapping › robot mapping
environment modeling |
0.0 | 3 | 1998 | Position Estimation Using Principal Components of Range Data · ICRA 1998 World modeling and position estimation for a mobile robot using ultrasonic ranging · ICRA 1989 Dynamic world modeling for an intelligent mobile robot using a rotating ultra-sonic ranging device · ICRA 1985 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.0 | 1 | 2000 | Object Recognition Using Coloured Receptive Fields · ECCV (1) 2000 |
Computer vision › 3D vision
local feature descriptor |
0.0 | 1 | 2000 | Local Scale Selection for Gaussian Based Description Techniques · ECCV (1) 2000 |
Machine learning › Deep learning architectures and training › convolutional neural network
receptive field |
0.0 | 1 | 2000 | Object Recognition Using Coloured Receptive Fields · ECCV (1) 2000 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic classifier
bayesian recognition |
0.0 | 1 | 1999 | Probabilistic Recognition of Activity using Local Appearance · CVPR 1999 |
Robotics › Robot navigation and mapping
active vision |
0.0 | 2 | 1994 | Integration and Control of Reactive Visual Processes · ECCV (2) 1994 Gaze Control for a Binocular Camera Head · ECCV 1992 |
Robotics › Robot navigation and mapping › active perception
active object recognition |
0.0 | 1 | 1998 | Transinformation for Active Object Recognition · ICCV 1998 |
Computer vision › 3D vision
camera calibration |
0.0 | 2 | 1993 | Dynamic calibration of an active stereo head · ICCV 1993 Maintaining stereo calibration by tracking image points · CVPR 1993 |
Computer vision › 3D vision › point cloud analysis › point cloud learning › point cloud representation learning
range view representation |
0.0 | 1 | 1998 | Position Estimation Using Principal Components of Range Data · ICRA 1998 |
Computational photography and imaging
color constancy |
0.0 | 1 | 1998 | Comprehensive Colour Image Normalization · ECCV (1) 1998 |
Computer vision › 3D vision
3d reconstruction |
0.0 | 3 | 1993 | Dynamic calibration of an active stereo head · ICCV 1993 Measurement and Integration of 3-D Structures By Tracking Edge Lines · ECCV 1990 Dynamic World Modeling Using Vertical Line Stereo · ECCV 1990 |
Computer vision › Face, body and person analysis
face detection |
0.0 | 1 | 1997 | Multi-Modal Tracking of Faces for Video Communications · CVPR 1997 |
Computer vision › Face, body and person analysis
face tracking |
0.0 | 1 | 1997 | Multi-Modal Tracking of Faces for Video Communications · CVPR 1997 |
Software testing
automated testing |
0.0 | 1 | 1996 | Issues in the Full Scale Use of Formal Methods for Automated Testing · ISSTA 1996 |
Software testing
specification-based testing |
0.0 | 1 | 1996 | Issues in the Full Scale Use of Formal Methods for Automated Testing · ISSTA 1996 |
Methods — techniques the papers use, named apart from their topics
self-supervised transformer · 1.2graph cuts · 1.2normalized cut · 0.7spectral clustering · 0.6rigid object pose · 0.3semi-rigid scene flow · 0.2RGB-D · 0.2q-learning · 0.2credit assignment · 0.2analogy-based learning · 0.2spatio-temporal filter · 0.0joint statistics · 0.0bayes rule · 0.0kalman filter · 0.0cross-correlation · 0.0color histogram matching · 0.0blink detection · 0.0PD controller · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Accommodating Missing Modalities in Time-Continuous Multimodal Emotion RecognitionabstractDecades of research indicate that emotion recognition is more effective when drawing information from multiple modalities. But what if some modalities are sometimes missing? To address this problem, we propose a novel Transformer-based architecture for recognizing valence and arousal in a time-continuous manner even with missing input modalities. We use a coupling of cross-attention and self-attention mechanisms to emphasize relationships between modalities during time and enhance the learning process on weak salient inputs. Experimental results on the Ulm-TSST dataset show that our model exhibits an improvement of the concordance correlation coefficient evaluation of 37% when predicting arousal values and 30% when predicting valence values, compared to a late-fusion baseline approach. Juan Vazquez-Rodriguez, Grégoire Lefebvre, Julien Cumin, James L. Crowley |
ACII | 4 |
| 2023 | TokenCut: Segmenting Objects in Images and Videos With Self-Supervised Transformer and Normalized CutabstractIn this paper, we describe a graph-based algorithm that uses the features obtained by a self-supervised transformer to detect and segment salient objects in images and videos. With this approach, the image patches that compose an image or video are organised into a fully connected graph, in which the edge between each pair of patches is labeled with a similarity score based on the features learned by the transformer. Detection and segmentation of salient objects can then be formulated as a graph-cut problem and solved using the classical Normalized Cut algorithm. Despite the simplicity of this approach, it achieves state-of-the-art results on several common image and video detection and segmentation tasks. For unsupervised object discovery, this approach outperforms the competing approaches by a margin of 6.1%, 5.7%, and 2.6% when tested with the VOC07, VOC12, and COCO20 K datasets. For the unsupervised saliency detection task in images, this method improves the score for Intersection over Union (IoU) by 4.4%, 5.6% and 5.2%. When tested with the ECSSD, DUTS, and DUT-OMRON datasets. This method also achieves competitive results for unsupervised video object segmentation tasks with the DAVIS, SegTV2, and FBMS datasets. Yangtao Wang, Xi Shen 0001, Yuan Yuan 0002, Yuming Du, Maomao Li, Shell Xu Hu, James L. Crowley, Dominique Vaufreydaz |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Emotion Recognition with Pre-Trained Transformers Using Multimodal SignalsabstractIn this paper, we address the problem of multimodal emotion recognition from multiple physiological signals. We demonstrate that a Transformer-based approach is suitable for this task. In addition, we present how such models may be pre-trained in a multimodal scenario to improve emotion recognition performances. We evaluate the benefits of using multimodal inputs and pre-training with our approach on a state-of-the-art dataset. Juan Vazquez-Rodriguez, Grégoire Lefebvre, Julien Cumin, James L. Crowley |
ACII | 4 |
| 2022 | Self-Supervised Transformers for Unsupervised Object Discovery using Normalized CutabstractTransformers trained with self-supervision using selfdistillation loss (DINO) have been shown to produce attention maps that highlight salient foreground objects. In this paper, we show a graph-based method that uses the selfsupervised transformer features to discover an object from an image. Visual tokens are viewed as nodes in a weighted graph with edges representing a connectivity score based on the similarity of tokens. Foreground objects can then be segmented using a normalized graph-cut to group self-similar regions. We solve the graph-cut problem using spectral clustering with generalized eigen-decomposition and show that the second smallest eigenvector provides a cutting solution since its absolute value indicates the likelihood that a token belongs to a foreground object. Despite its simplicity, this approach significantly boosts the performance of unsupervised object discovery: we improve over the recent state-of-the-art LOST by a margin of 6.9%, 8.1%, and 8.1% respectively on the VOC07, VOC12, and COCO20K. The performance can be further improved by adding a second stage class-agnostic detector (CAD). Our proposed method can be easily extended to unsupervised saliency detection and weakly supervised object detection. For unsupervised saliency detection, we improve IoU for 4.9%, 5.2%, 12.9% on ECSSD, DUTS, DUT-OMRON respectively compared to state-of-the-art. For weakly supervised object detection, we achieve competitive performance on CUB and ImageNet. Our code is available at: https://www.m-psi.fr/Papers/TokenCut2022/ Yangtao Wang, Xi Shen 0001, Shell Xu Hu, Yuan Yuan 0002, James L. Crowley, Dominique Vaufreydaz |
CVPR | 5 |
| 2022 | Transformer-Based Self-Supervised Learning for Emotion RecognitionabstractIn order to exploit representations of time-series signals, such as physiological signals, it is essential that these representations capture relevant information from the whole signal. In this work, we propose to use a Transformer-based model to process electrocardiograms (ECG) for emotion recognition. Attention mechanisms of the Transformer can be used to build contextualized representations for a signal, giving more importance to relevant parts. These representations may then be processed with a fully-connected network to predict emotions.To overcome the relatively small size of datasets with emotional labels, we employ self-supervised learning. We gathered several ECG datasets with no labels of emotion to pre-train our model, which we then fine-tuned for emotion recognition on the AMIGOS dataset. We show that our approach reaches state-of-the-art performances for emotion recognition using ECG signals on AMIGOS. More generally, our experiments show that transformers and pre-training are promising strategies for emotion recognition with physiological signals. Juan Vazquez-Rodriguez, Grégoire Lefebvre, Julien Cumin, James L. Crowley |
ICPR | 4 |
| 2021 | PSINES: Activity and Availability Prediction for Adaptive Ambient IntelligenceabstractAutonomy and adaptability are essential components of ambient intelligence. For example, in smart homes, proactive acting and occupants advising, adapted to current and future contexts of living, are essential to go beyond limitations of previous domotic services. To reach such autonomy and adaptability, ambient systems need to automatically grasp their users’ ambient context. In particular, users’ activities and availabilities for communication are valuable pieces of contextual information that can help such systems to adapt to user needs and behaviours. While significant research work exists on activity recognition in homes, less attention has been given to prediction of future activities, as well as to availability recognition and prediction in general. In this article, we investigate several Dynamic Bayesian Network (DBN) architectures for activity and availability prediction of occupants in homes, including our novel model, called Past SItuations to predict the NExt Situation (PSINES). This predictive architecture utilizes context information, sensor event aggregations, and latent user cognitive states to accurately predict future home situations based on previous situations. We experimentally evaluate PSINES, as well as intermediate DBN architectures, on multiple state-of-the-art datasets, with prediction accuracies of up to 89.52% for activity and 82.08% for availability on the Orange4Home dataset. Julien Cumin, Grégoire Lefebvre, Fano Ramparany, James L. Crowley |
ACM Trans. Auton. Adapt. Syst. | 4 |
| 2019 | Deep learning investigation for chess player attention prediction using eye-tracking and game dataabstractThis article reports on an investigation of the use of convolutional neural networks to predict the visual attention of chess players. The visual attention model described in this article has been created to generate saliency maps that capture hierarchical and spatial features of chessboard, in order to predict the probability fixation for individual pixels Using a skip-layer architecture of an autoencoder, with a unified decoder, we are able to use multiscale features to predict saliency of part of the board at different scales, showing multiple relations between pieces. We have used scan path and fixation data from players engaged in solving chess problems, to compute 6600 saliency maps associated to the corresponding chess piece configurations. This corpus is completed with synthetically generated data from actual games gathered from an online chess platform. Experiments realized using both scan-paths from chess players and the CAT2000 saliency dataset of natural images, highlights several results. Deep features, pretrained on natural images, were found to be helpful in training visual attention prediction for chess. The proposed neural network architecture is able to generate meaningful saliency maps on unseen chess configurations with good scores on standard metrics. This work provides a baseline for future work on visual attention prediction in similar contexts. Justin Le Louedec, Thomas Guntz, James L. Crowley, Dominique Vaufreydaz |
ETRA | 3 |
| 2018 | Put That There: 20 Years of Research on Multimodal InteractionabstractHumans interact with the world using five major senses: sight, hearing, touch, smell, and taste. Almost all interaction with the environment is naturally multimodal, as audio, tactile or paralinguistic cues provide confirmation for physical actions and spoken language interaction. Multimodal interaction seeks to fully exploit these parallel channels for perception and action to provide robust, natural interaction. Richard Bolt's "Put That There" (1980) provided an early paradigm that demonstrated the power of multimodality and helped attract researchers from a variety of disciplines to study a new approach for post-WIMP computing that moves beyond desktop graphical user interfaces (GUI). In this talk, I will look back to the origins of the scientific community of multimodal interaction, and review some of the more salient results that have emerged over the last 20 years, including results in machine perception, system architectures, visualization, and computer to human communications. Recently, a number of game-changing technologies such as deep learning, cloud computing, and planetary scale data collection have emerged to provide robust solutions to historically hard problems. As a result, scientific understanding of multimodal interaction has taken on new relevance as construction of practical systems has become feasible. I will discuss the impact of these new technologies and the opportunities and challenges that they raise. I will conclude with a discussion of the importance of convergence with cognitive science and cognitive systems to provide foundations for intelligent, human-centered interactive systems that learn and fully understand humans and human-to-human social interaction, in order to provide services that surpass the abilities of the most intelligent human servants. James L. Crowley |
ICMI | 1 |
| 2018 | Defining the Pose of Any 3D Rigid Object and an Associated Distance
Romain Brégier, Frederic Devernay, Laetitia Leyrit, James L. Crowley |
Int. J. Comput. Vis. | 4 |
| 2014 | Dense Semi-rigid Scene Flow Estimation from RGBD Images
Julian Quiroga, Thomas Brox, Frederic Devernay, James L. Crowley |
ECCV (7) | 4 |
| 2014 | Local Binary Patterns Calculated over Gaussian Derivative ImagesabstractIn this paper we present a new static descriptor for facial image analysis. We combine Gaussian derivatives with Local Binary Patterns to provide a robust and powerful descriptor especially suited to extracting texture from facial images. Gaussian features in the form of image derivatives form the input to the Linear Binary Pattern(LBP) operator instead of the original image. The proposed descriptor is tested for face recognition and smile detection. For face recognition we use the CMU-PIE and the YaleB+extended YaleB database. Smile detection is performed on the benchmark GENKI 4k database. With minimal machine learning our descriptor outperforms the state of the art at smile detection and compares favourably with the state of the art at face recognition. Varun Jain, James L. Crowley, Augustin Lux |
ICPR | 2 |
| 2014 | Scale Normalized Radial Fourier Transform as a Robust Image DescriptorabstractWe present a new visual descriptor that combines a multi-scale Laplacian Profile with a Radial Discrete Fourier Transform. This descriptor exists at every position and scale in an image and provides a local feature vector that is both discriminant and robust to changes in orientation and scale. It has a variable description length, and thus can be easily adapted for a variety of applications, ranging from simple detection tasks on low power computing platforms to complex tasks requiring highly discriminant detectors. To demonstrate the discriminant power of this descriptor we employ it in its most compact form to construct a cascade of linear classifiers for detecting people in images. We compare this detector to cascades classifiers constructed using Haar wavelets, Gaussian derivatives and variable size block HOG descriptors. Our experiments show that a cascade with this descriptor performs well against the other three detectors when tested using a common publicly available data set. We examine the stability of the descriptor to changes in image rotation and scaling for different description lengths. Evanthia Mavridou, Manh-Dung Hoang, James L. Crowley, Augustin Lux |
ICPR | 3 |
| 2014 | Local scene flow by tracking in intensity and depth
Julian Quiroga, Frederic Devernay, James L. Crowley |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Local/global scene flow estimationabstractThe scene flow describes the 3D motion of every point in a scene between two time steps. We present a novel method to estimate a dense scene flow using intensity and depth data. It is well known that local methods are more robust under noise while global techniques yield dense motion estimation. We combine local and global constraints to solve for the scene flow in a variational framework. An adaptive TV (Total Variation) regularization is used to preserve motion discontinuities. Besides, we constrain the motion using a set of 3D correspondences to deal with large displacements. In the experimentation our approach outperforms previous scene flow from intensity and depth methods in terms of accuracy. Julian Quiroga, Frederic Devernay, James L. Crowley |
ICIP | 3 |
| 2012 | Extracting planar structures efficiently with revisited BetaSAC
Marion Decrouez, Romain Dupont, François Gaspard, James L. Crowley |
ICPR | 4 |
| 2012 | Planning with Inaccurate Temporal RulesabstractWe use a temporal pattern model called Temporal Interval Tree Associative Rules (Tita rules). This pattern model has been introduced in a previous work. The model can express uncertainty, temporal inaccuracy, the usual time point operators, synchronicity, incomplete orders, chaining, disjunctive time constraints and temporal negation. This pattern model is initially designed to be used for temporal learning. In this paper, we use Tita rules as world description models for a Planning and Scheduling task. We present an efficient temporal planning algorithm able to deal with uncertainty, temporal inaccuracy, discontinuous (or disjunctive) time constraints and predictable but imprecisely time located exogenous events. We evaluate our technique by joining a learning algorithm and our planning algorithm into a simple reactive cognitive architecture that we apply on with virtual robot. Mathieu Guillame-Bert, James L. Crowley |
ICTAI | 2 |
| 2011 | New Approach on Temporal Data Mining for Symbolic Time Sequences: Temporal Tree Associate RulesabstractWe introduce a temporal pattern model called Temporal Tree Associative Rule (TTA rule). This pattern model can be used to express both uncertainty and temporal inaccuracy of temporal events expressed as Symbolic Time Sequences. Among other things, TTA rules can express the usual time point operators, synchronicity, order, chaining, as well as temporal negation. TTA rule is designed to allows predictions with optimum temporal precision. Using this representation, we present an algorithm that can be used to extract Temporal Tree Associative rules from large data sets of symbolic time sequences. This algorithm is a mining heuristic based on entropy maximisation and statistical independence analysis. We discuss the evaluation of probabilistic temporal rules, evaluate our technique with an experiment and discuss the results. Mathieu Guillame-Bert, James L. Crowley |
ICTAI | 2 |
| 2010 | BetaSAC: A New Conditional Sampling For RANSACabstractWe present a new strategy for RANSAC sampling named BetaSAC, in reference to the beta distribution. Our proposed sampler builds a hypothesis set incrementally, selecting data points conditional on the previous data selected for the set. Such a sampling is shown to provide more suitable samples in terms of inlier ratio but also of consistency and potential to lead to an accurate parameters estimation. The algorithm is presented as a general framework, easily implemented and able to exploit any kind of prior information on the potential of a sample. As with PROSAC, BetaSAC converges towards RANSAC in the worst case. The benefits of the method are demonstrated on the homography estimation problem. Antoine Méler, Marion Decrouez, James L. Crowley |
BMVC | 3 |
| 2010 | "How old are you?" : Age Estimation with Tensors of Binary Gaussian Receptive MapsabstractIn this paper we describe experiments with a method to automatically estimate human age from facial images. This system extends recent results with the use of a tensorial representation from Gaussian receptive field responses for face recognition to the problem of estimating age. Among other results, we show that inclusion of fourth order Gaussian receptive fields can improve recognition. We describe an optimal tensorial configuration and compare the use of two different configurations with Multilinear Principal Component Analysis to reduce tensor order and Relevance Vector Machines as regressor. Experimental results are demonstrated for the FG-NET and the MORPH aging datasets. John A. Ruiz-Hernandez, James L. Crowley, Augustin Lux |
BMVC | 2 |
| 2009 | Face Recognition using Tensors of Census Transform Histograms from Gaussian Features MapsabstractInternational audience John A. Ruiz-Hernandez, James L. Crowley, Antoine Méler, Augustin Lux |
BMVC | 2 |
| 2009 | A comparison of three methods for measure of Time to ContactabstractTime to contact (TTC) is a biologically inspired method for obstacle detection and reactive control of motion that does not require scene reconstruction or 3D depth estimation. Estimating TTC is difficult because it requires a stable and reliable estimate of the rate of change of distance between image features. In this paper we propose a new method to measure time to contact, active contour affine scale (ACAS). We experimentally and analytically compare ACAS with two other recently proposed methods: scale invariant ridge segments (SIRS), and image brightness derivatives (IBD). Our results show that ACAS provides a more accurate estimation of TTC when the image flow may be approximated by an affine transformation, while SIRS provides an estimate that is generally valid, but may not always be as accurate as ACAS, and IBD systematically over-estimate time to contact. Guillem Alenyà, Amaury Nègre, James L. Crowley |
IROS | 3 |
| 2009 | Detecting small group activities from multimodal observations
Oliver Brdiczka, Jérôme Maisonnasse, Patrick Reignier, James L. Crowley |
Appl. Intell. | 4 |
| 2009 | Detecting Human Behavior Models From Multimodal Observation in a Smart HomeabstractThis paper addresses learning and recognition of human behavior models from multimodal observation in a smart home environment. The proposed approach is part of a framework for acquiring a high-level contextual model for human behavior in an augmented environment. A 3-D video tracking system creates and tracks entities (persons) in the scene. Further, a speech activity detector analyzes audio streams coming from head set microphones and determines for each entity, whether the entity speaks or not. An ambient sound detector detects noises in the environment. An individual role detector derives basic activity like ldquowalkingrdquo or ldquointeracting with tablerdquo from the extracted entity properties of the 3-D tracker. From the derived multimodal observations, different situations like ldquoaperitifrdquo or ldquopresentationrdquo are learned and detected using statistical models (HMMs). The objective of the proposed general framework is two-fold: the automatic offline analysis of human behavior recordings and the online detection of learned human behavior models. To evaluate the proposed approach, several multimodal recordings showing different situations have been conducted. The obtained results, in particular for offline analysis, are very good, showing that multimodality as well as multiperson observation generation are beneficial for situation recognition. Oliver Brdiczka, Matthieu Langet, Jérôme Maisonnasse, James L. Crowley |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2009 | Learning Situation Models in a Smart HomeabstractThis paper addresses the problem of learning situation models for providing context-aware services. Context for modeling human behavior in a smart environment is represented by a situation model describing environment, users, and their activities. A framework for acquiring and evolving different layers of a situation model in a smart environment is proposed. Different learning methods are presented as part of this framework: role detection per entity, unsupervised extraction of situations from multimodal data, supervised learning of situation representations, and evolution of a predefined situation model with feedback. The situation model serves as frame and support for the different methods, permitting to stay in an intuitive declarative framework. The proposed methods have been integrated into a whole system for smart home environment. The implementation is detailed, and two evaluations are conducted in the smart home environment. The obtained results validate the proposed approach. Oliver Brdiczka, James L. Crowley, Patrick Reignier |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2008 | Face detection by cascade of Gaussian derivates classifiers calculated with a half-octave pyramidabstractThis paper presents a method for object detection based on a cascade of scale and orientation normalized Gaussian derivative classifiers learnt with Adaboost. Normalized Gaussian derivatives provide a small but powerful feature set for rapid learning using Adaboost. Real time detection is made possible by use of a fast integer coefficient algorithm that computes a half-octave Gaussian pyramid with linear algorithmic complexity using a cascade of binomial kernel filters. The method is demonstrated by training a boosted classifier for frontal face detection using standard data sets. Experiments demonstrate that this approach can provide detection rates that are comparable or superior to those obtained with integral images while dramatically reducing the required training effort. John A. Ruiz-Hernandez, Augustin Lux, James L. Crowley |
FG | 3 |
| 2008 | Learning polite behavior with situation modelsabstractIn this paper, we describe experiments with methods for learning the appropriateness of behaviors based on a model of the current social situation. We first review different approaches for social robotics, and present a new approach based on situation modeling. We then review algorithms for social learning and propose three modifications to the classical Q-Learning algorithm. We describe five experiments with progressively complex algorithms for learning the appropriateness of behaviors. The first three experiments illustrate how social factors can be used to improve learning by controlling learning rate. In the fourth experiment we demonstrate that proper credit assignment improves the effectiveness of reinforcement learning for social interaction. In our fifth experiment we show that analogy can be used to accelerate learning rates in contexts composed of many situations. Rémi Barraquand, James L. Crowley |
HRI | 2 |
| 2007 | Agent based middleware infrastructure for autonomous context-aware ubiquitous computing services
John Soldatos 0001, Ippokratis Pandis, Kostas Stamatis, Lazaros Polymenakos, James L. Crowley |
Comput. Commun. | 5 |
| 2007 | Context-aware environments: from specification to implementationabstractAbstract: This paper deals with the problem of implementing a context model for a smart environment. The problem has already been addressed several times using many different data‐ or problem‐driven methods. In order to separate the modelling phase from implementation, we first represent the context model by a network of situations. Then, different implementations can be automatically generated from this context model depending on user needs and underlying perceptual components. Two different implementations are proposed in this paper: a deterministic one based on Petri nets and a probabilistic one based on hidden Markov models. Both implementations are illustrated and applied to real‐world problems. Patrick Reignier, Oliver Brdiczka, Dominique Vaufreydaz, James L. Crowley, Jérôme Maisonnasse |
Expert Syst. J. Knowl. Eng. | 4 |
| 2006 | User-Centric Design of a Vision System for Interactive ApplicationsabstractDespite great promise of vision-based user interfaces, commercial employment of such systems remains marginal. Most vision-based interactive systems are one-time, "proof of concept" prototypes that demonstrate the interest of a particular image treatment applied to interaction. In general, vision systems require parameter tuning, both during setup and at runtime, and are thus difficult to handle by nonexperts in computer vision. In this paper, we present a pragmatic, developer-centric, service-oriented framework for the construction of vision-based interactive systems. Our framework is designed to allow developers unfamiliar with vision to use computer vision as an interaction modality. To achieve this goal, we address specific developer- and interaction-centric requirements during the design of our system. We validate our approach with an implementation of standard GUI widgets (buttons and sliders) based on computer vision. Stanislaw Borkowski, James L. Crowley, Julien Letessier, François Bérard |
ICVS | 2 |
| 2006 | Real-time stereo and optical flow data fusionabstractIn this paper, we propose a real-time method to detect obstacles using theoretical models of the ground plane, first in a 3D point cloud given by a stereo camera, and then in an optical flow field given by one of the stereo pair's camera. The idea of our method is to combine two partial occupancy grids from both sensor modalities with an occupancy grid framework. The two methods do not have the same range, precision and resolution. For example, the stereo method is precise for close objects but cannot see further than 7 m (with our lenses), while the optical flow method can see considerably further but has lower accuracy. Experiments that have been carried on the CyCab mobile robot and on a tractor demonstrate that we can combine the advantages of both algorithms to build local occupancy grids from incomplete data (optical flow from a monocular camera cannot give depth information without time integration) Christophe Braillon, Kane Usher, Cédric Pradalier, James L. Crowley, Christian Laugier |
IROS | 4 |
| 2006 | Extracting Activities from Multimodal Observation
Oliver Brdiczka, Jérôme Maisonnasse, Patrick Reignier, James L. Crowley |
KES (2) | 4 |
| 2004 | Introduction to the special issue: International Conference on Vision Systems
James L. Crowley, Justus H. Piater |
Mach. Vis. Appl. | 1 |
| 2004 | Brand identification using Gaussian derivative histograms
Daniela Hall, Fabien Pelisson, Olivier Riff, James L. Crowley |
Mach. Vis. Appl. | 4 |
| 2003 | Brand Identification Using Gaussian Derivative Histograms
Fabien Pelisson, Daniela Hall, Olivier Riff, James L. Crowley |
ICVS | 4 |
| 2002 | Perceptual Components for Context Aware Computing
James L. Crowley, Joëlle Coutaz, Gaëtan Rey, Patrick Reignier |
UbiComp | 1 |
| 2002 | Context aware observation of human activitiesabstractInteractive environments combine perception, action and communication to extend human-computer interaction. We believe that a fundamental challenge for interactive environments is developing models and methods for "context awareness". We present an ontology for context awareness for interactive environments. We show how the elements of this ontology correspond to the elements of a software architecture for observing situation and context. Within this framework, context predicts the evolution of the situation, and provides "meaning" for objects and events. Context also provides a specification for assembling federations of processes to measure properties, determine relations and detect events. James L. Crowley |
ICME (1) | 1 |
| 2001 | Continuity properties of the appearance manifold for mobile robot position estimation
James L. Crowley, F. Pourraz |
Image Vis. Comput. | 1 |
| 2000 | A Probabilistic Sensor for the Perception and Recognition of Activities
Olivier Chomat, Jérôme Martin, James L. Crowley |
ECCV (1) | 3 |
| 2000 | Local Scale Selection for Gaussian Based Description Techniques
Olivier Chomat, Vincent Colin de Verdière, Daniela Hall, James L. Crowley |
ECCV (1) | 4 |
| 2000 | Object Recognition Using Coloured Receptive Fields
Daniela Hall, Vincent Colin de Verdière, James L. Crowley |
ECCV (1) | 3 |
| 2000 | A Probabilistic Sensor for the Perception of ActivitiesabstractThis paper presents a new technique for the perception of activities using a statistical description of spatio-temporal properties. With this approach, the probability of an activity in a spatio-temporal image sequence is computed by applying a Bayes rule to the joint statistics of the responses of motion energy receptive fields. A set of motion energy receptive fields is designed in order to sample the power spectrum of a moving texture. Their structure relates to the spatio-temporal energy models of Adelson and Bergen where measures of local visual motion information are extracted comparing the outputs of triad of Gabor energy filters. Then the probability density function required for the Bayes rule is estimated for each class of activity by computing multi-dimensional histograms from the outputs from the set of receptive fields. The perception of activities is achieved according to the Bayes rule. The result at a given time is the map of the conditional probabilities that each pixel belongs to an activity of the training set. The approach is validated with experiments in the perception of activities of walking persons in a visual surveillance scenario. Results are robust to changes in illumination conditions, to occlusions and to changes in texture. Olivier Chomat, James L. Crowley |
FG | 2 |
| 2000 | Robust Face Tracking Using ColorabstractWe discuss a new robust tracking technique applied to histograms of intensity-normalized color. This technique supports a video codec based on orthonormal basis coding. Orthonormal basis coding can be very efficient when the images to be coded have been normalized in size and position. However an imprecise tracking procedure can have a negative impact on the efficiency and the quality of reconstruction of this technique, since it may increase the size of the required basis space. The face tracking procedure described in this paper has certain advantages, such as greater stability, higher precision, and less jitter, over conventional tracking techniques using color histograms. In addition to those advantages, the features of the tracked object such as mean and variance are mathematically describable. Karl Schwerdt, James L. Crowley |
FG | 2 |
| 2000 | A Sound MagicBoard
Christophe Le Gal, Ali Erdem Özcan, Karl Schwerdt, James L. Crowley |
ICMI | 4 |
| 2000 | Estimating the Pose of Phicons for Human Computer Interaction
Daniela Hall, James L. Crowley |
ICMI | 2 |
| 2000 | Visual Recognition of Emotional States
Karl Schwerdt, Daniela Hall, James L. Crowley |
ICMI | 3 |
| 2000 | Object Detection Using ColorabstractPresents a method to detect objects in images using colour. The system learns by example how characteristic each discrete colour is of a class of objects using two sets of training images; one set is of images known to contain at least one target object, while the other set is of images known to be free of the target objects. Using the training images a lookup table is constructed, in which each discrete colour is assigned a value which approximates the probability that pixels with this discrete colour are part of a target object. This table is called the colour profile. To detect the presence of target objects in a new image the colour profile values are back projected onto the image and the image is thresholded. A suitable threshold is calculated automatically and the system is able to recognize if the target objects are not able to be distinguished by the colour profile method. Applying the colour profiling approach to a difficult inspection task has shown it to be a robust, accurate and fast solution. Nicola Duffy, James L. Crowley, Gerard Lacey |
ICPR | 2 |
| 2000 | Recognition without Correspondence using Multidimensional Receptive Field Histograms
Bernt Schiele, James L. Crowley |
Int. J. Comput. Vis. | 2 |
| 1999 | Probabilistic Recognition of Activity using Local AppearanceabstractThis paper addresses the problem of probabilistic recognition of activities from local spatio-temporal appearance. Joint statistics of space-time filters are employed to define histograms which characterize the activities to be recognized. These histograms provide the joint probability density functions required for recognition using Bayes rule. The result is a technique for recognition of activities which is robust to partial occlusions as well as changes in illumination. In this paper the framework and background for this approach is first described. Then the family of spatio-temporal receptive fields used for characterizing activities is presented. This is followed by a review of probabilistic recognition of patterns from joint statistics of receptive field responses. The approach is validated with the results of experiments in the discrimination of persons walking in different directions, and the recognition of a simple set of hand gestures in an augmented reality scenario. Olivier Chomat, James L. Crowley |
CVPR | 2 |
| 1999 | Face-Tracking and Coding for Video Compression
William E. Vieux, Karl Schwerdt, James L. Crowley |
ICVS | 3 |
| 1998 | Comprehensive Colour Image Normalization
Graham D. Finlayson, Bernt Schiele, James L. Crowley |
ECCV (1) | 3 |
| 1998 | Visual Recognition Using Local Appearance
Vincent Colin de Verdière, James L. Crowley |
ECCV (1) | 2 |
| 1998 | Active Hand Tracking
Jérôme Martin, Vincent E. Devin, James L. Crowley |
FG | 3 |
| 1998 | Transinformation for Active Object RecognitionabstractThis article develops an analogy between object recognition and the transmission of information through a channel based on the statistical representation of the appearances of 3D objects. This analogy provides a means to quantitatively evaluate the contribution of individual receptive field vectors, and to predict the performance of the object recognition process. Transinformation also provides a quantitative measure of the discrimination provided by each viewpoint, thus permitting the determination of the most discriminant viewpoints. As an application, the article develops an active object recognition algorithm which is able to resolve ambiguities inherent in a single-view recognition algorithm. Bernt Schiele, James L. Crowley |
ICCV | 2 |
| 1998 | Position Estimation Using Principal Components of Range DataabstractDescribes an approach to mobile robot position estimation based on principal component analysis of laser range data. An eigenspace is constructed from the principal components of a large number of range data sets. The structure of an environment, as seen by a range sensor, is represented as a family of surfaces in this space. Subsequent range data sets from the environment project as a point in this space. Associating this point to the family of surfaces gives a set of candidate positions and orientations (poses) for the sensor. These candidate poses correspond to positions and orientations in the environment which have similar range profiles. A Kalman filter can used to select the most likely candidate pose based on coherence with small movements. The first part of this paper describes how a relatively small number of depth profiles of an environment can be used to generate a complete eigenspace. This space is used to build a representation of the range scan profiles obtained from a regular grid of positions and orientations (poses). This representation has the form of a family of surfaces (a manifold). This representation converts the problem of associating a range profile to possible positions and orientations into a table lookup. As a side benefit, the method provides a simple means to detect obstacles in a range profile. The final section of the paper reviews the use of estimation theory to determine the correct pose hypothesis by tracking. James L. Crowley, Frank Wallner, Bernt Schiele |
ICRA | 1 |
| 1997 | Multi-Modal Tracking of Faces for Video CommunicationsabstractVisual processes to detect and track faces for video compression and transmission. The system is based on an architecture in which a supervisor selects and activates visual processes in cyclic manner. Control of visual processes is made possible by a confidence factor which accompanies each observation. Fusion of results into a unified estimation for tracking is made possible by estimating a covariance matrix with each observation. Visual processes for face tracking are described using blink detection, normalised color histogram matching, and cross correlation (SSD and NCC). Ensembles of visual processes are organised into processing states so as to provide robust tracking. Transition between states is determined by events detected by processes. The result of face detection is fed into recursive estimator (Kalman filter). The output from the estimator drives a PD controller for a pan/tilt/zoom camera. The resulting system provides robust and precise tracking which operates continuously at approximately 20 images per second on a 150 megahertz computer workstation. James L. Crowley, François Bérard |
CVPR | 1 |
| 1997 | Robust Computer Vision for Computer Mediated Communication
François Bérard, Joëlle Coutaz, James L. Crowley |
INTERACT | 3 |
| 1997 | Eigen-Space Coding as a Means to Support Privacy in Computer Mediated Communication
Joëlle Coutaz, James L. Crowley, François Bérard |
INTERACT | 2 |
| 1997 | Appearance based process for visual navigationabstractDescribes the use of appearance based vision for defining visual processes for navigation. A visual processes which transform images to commands and events. A family of visual processes are defined by associating the appearance of a scene from a given viewpoint with the simple trajectories. Appearance is captured as a set of low-resolution images. Energy normalised cross correlation is used to maintain heading, to estimate confidence and to servo control a robot vehicle while following a path. Experimental results are presented which compare results with a single camera, a pair of parallel cameras and a pair of divergent cameras. The most accurate (and robust) navigation is found with a pair of cameras which are slightly divergent. Stephen D. Jones, Claus Andresen, James L. Crowley |
IROS | 3 |
| 1996 | Uncalibrated Visual Tasks via Linear Interaction
Carlo Colombo, James L. Crowley |
ECCV (2) | 2 |
| 1996 | Object Recognition Using Multidimensional Receptive Field Histograms
Bernt Schiele, James L. Crowley |
ECCV (1) | 2 |
| 1996 | Coordination of Perceptual Processes for Computer Mediated CommunicationabstractIn Computer Mediated Communication such as desktop video conferencing, static video cameras provide a restricted field of view of remote sites. The effective field of view can be enlarged, while maintaining the user's freedom of movement, by slaving a remote controlled camera to movements of the user's head. This paper concerns techniques for tracking of faces. We demonstrate that robustness and reliability can be increased by combining multiple perceptual processes such as eye blink detection, skin color histogram and cross correlation, that adapt to a variety of operating conditions. We illustrate our technique with CoMedi, a media-space currently under development. Joëlle Coutaz, François Bérard, James L. Crowley |
FG | 3 |
| 1996 | Experimental performance characterization of adaptive filtersabstractAdaptive filters and image enhancement techniques have repeatedly been suggested to make feature extraction more robust. Few comparative analysis exists between competing techniques and even the existing ones evaluate only the effect of the filter on the image but not its effect on a feature extraction process. We present an experimental approach for the evaluation of low level vision system components in a system framework. Our method is applied to adaptive filters. By analyzing the system architecture we chose an evaluation level and derive the evaluation criteria and parameter control strategies. The results show that none of the tested techniques performs better than linear filters or the Canny edge detector. Guido Appenzeller, James L. Crowley |
ICPR | 2 |
| 1996 | Probabilistic object recognition using multidimensional receptive field histogramsabstractThis paper describes a probabilistic object recognition technique which does not require correspondence matching of images. This technique is an extension of our earlier work (1996) on object recognition using matching of multi-dimensional receptive field histograms. In the earlier paper we have shown that multi-dimensional receptive field histograms can be matched to provide object recognition which is robust in the face of changes in viewing position and independent of image plane rotation and scale. In this paper we extend this method to compute the probability of the presence of an object in an image. The paper begins with a review of the method and previously presented experimental results. We then extend the method for histogram matching to obtain a genuine probability of the presence of an object. We present experimental results on a database of 100 objects showing that the approach is capable recognizing all objects correctly by using only a small portion of the image. Our results show that receptive field histograms provide a technique for object recognition which is robust, has low computational cost and a computational complexity which is linear with the number of pixels. Bernt Schiele, James L. Crowley |
ICPR | 2 |
| 1996 | Where to look next and what to look forabstractThe authors (1996) introduced the use of multidimensional receptive field histograms for probabilistic object recognition. In this paper we reverse the object recognition problem by asking the question "where should we look?", when we want to verify the presence of an object, to track an object or to actively explore a scene. This paper describes the statistical framework from which we obtain a network of salient points for an object. This network of salient points may be used for fixation control in the context of active object recognition. Bernt Schiele, James L. Crowley |
IROS | 2 |
| 1996 | Issues in the Full Scale Use of Formal Methods for Automated TestingabstractExperience from a full scale effort to apply formal methods to automated testing in the open systems software arena is described. The formal method applied in this work is based upon the Clemson Automated Testing System (CATS) which includes a formal specification language, a set of guidelines describing how to use the method effectively, and tool support capable of translating formal specifications into executable tests. This method is currently being used to develop a full scale test suite for IEEE's Ada Language Binding to POSIX. Following an overview of CATS, an experience report consisting of results, lessons learned and future directions is presented. James L. Crowley, James F. Leathrum Jr., K. A. Liburdy |
ISSTA | 1 |
| 1995 | Comparison of kinematic and visual servoing for fixationabstractThis paper describes experiments in estimation and control of the fixation point with a 10 degree of freedom head-neck-body system. A system has been constructed which uses standard kinematic techniques to estimate the fixation point. Two approaches have been investigated to control the fixation point: 3D kinematic estimation and 2D visual servoing. Kinematics is shown to be efficient provided that there is a sufficiently precise model of the kinematic chain. In particular, large errors in the kinematic model can cause the system to oscillate. In contrast, visual servoing is extremely robust with respect to errors in the kinematic model, but require a much larger (factor of 10) number of cycles to converge. The two approaches are found to be complementary. Kinematics can be used to perform a ballistic saccade to the fixation point, while visual servoing serves to correct small errors with a "micro-saccade". This approach mimics saccadic control found in the human visual system. James L. Crowley, Mouafak Mesrabi, François Chaumette |
IROS (1) | 1 |
| 1994 | Integration and Control of Reactive Visual Processes
James L. Crowley, Jean Marc Bedrune, Morten Bekker |
ECCV (2) | 1 |
| 1994 | Self-Supervised Neural System for Reactive NavigationabstractThis paper deals with an artificial neural system for a mobile robot reactive navigation in an unknown, cluttered environment. A task of a presented system is to provide a steering angle signal letting a robot reach a goal while avoiding collisions with obstacles. Basic reactive navigation methods are briefly characterized, a special attention is paid to neural approaches. Then a qualitative description of a presented system is given. The main parts of the system are: the Fuzzy-ART classifier performing a perceptual space partitioning, and the neural associative memory, storing system's experience and superposing influences of different behaviours. Preliminary tests show that the learning by trial-and-error is efficient, as well in a case of beginning from scratch, as after some disturbances of either system's or environmental characteristics.> Artur Dubrawski, James L. Crowley |
ICRA | 2 |
| 1994 | A Comparison of Position Estimation Techniques Using Occupancy GridsabstractA mobile robot requires perception of its local environment for both sensor based locomotion and for position estimation. Occupancy grids, based on ultrasonic range data, provide a robust description of the local environment for locomotion. Unfortunately, current techniques for position estimation based on occupancy grids are both unreliable and computationally expensive. This paper reports on experiments with four techniques for position estimation using occupancy grids. A world modeling technique based on combining global and local occupancy grids is described. Techniques are described for extracting line segments from an occupancy grid based on a Hough transform. The use of an extended Kalman filter for position estimation is then adapted to this framework. Four matching techniques are presented for obtaining the innovation vector required by the Kalman filter equations. Experimental results show that matching of segments extracted from the both the local and global occupancy grids gives results which are superior to a direct matching of grids, or to a mixed matching of segments to grids.> Bernt Schiele, James L. Crowley |
ICRA | 2 |
| 1994 | Navigation with constraints for an autonomous mobile robotabstractThis paper describes a navigation planning algorithm for a mobile robot capable of autonomous navigation in a quasi-structured, partially known and dynamic environment. This algorithm uses a description of the environment containing both operational and geometric information. An interactive program permits an operator to create the description of the environment composed of structural elements such as walls and furniture. This description is completed by the addition of operational information in the form of a network of places, routes, landmarks and recharge stations. Places and routes include constraints on the actions of the robot as well as other information needed for planning and plan execution. A separate interactive program permits an operator to compose a surveillance mission as parallel sequences of navigation tasks and surveillance tasks, subject to constraints on energy, time, risk and position uncertainty. The planning algorithm described in this paper computes an optimal path for each navigation task according to the optimisation criterion and constraints. The authors introduce the notion of efficient path applied to a new best first search algorithm solving a multiple constraints problem, The paths determination relies on a state representation adapted to deal with environment constraints. The authors demonstrate that the complexity characteristics of their algorithm are similar to those of the A* algorithm. The planning system described in this paper has been implemented on a laboratory workstation and tested using radio-modem communication with a mobile platform equipped with ultrasonic range sensors and an active stereo vision system. This system has been developed for the MITHRA family of autonomous surveillance robots as part of project EUREKA EU 110.> Olivier Causse, James L. Crowley |
IROS | 2 |
| 1994 | Response to comments by Professor Dickmanns
James L. Crowley |
Signal Process. | 1 |
| 1993 | Maintaining stereo calibration by tracking image pointsabstractAn important problem in active 3-D vision is updating the camera calibration matrix as the focus, aperture, zoom or vergence angle of the cameras changes dynamically. Techniques are presented to compute the projection matrix from five-and-a-half points in a scene without matrix inversion, and to correct the projective transformation matrix by tracking reference points. The authors' experiments show that a change of focus can be corrected by an affine transform obtained by tracking three points. For a change in camera vergence, a projective correction, based on tracking four image points, is slightly more precise than an affine correction matrix. It is shown how stereo reconstruction makes it possible to 'hop' a reference frame from one object to another. Any set of four non-coplanar points in the scene may define such a reference frame. It is shown how to keep the reference frame locked onto a set of four points as a stereo head is translated or rotated. These techniques make it possible to reconstruct the shape of an object in in its intrinsic coordinates without having to match new observations to a partially reconstructed description.> James L. Crowley, Philippe Bobet, Cordelia Schmid |
CVPR | 1 |
| 1993 | Dynamic calibration of an active stereo headabstractThe authors present a method for using objects in a scene to define the reference frame for 3-D reconstruction. They first present a simple technique to calibrate an orthographic projection from four non-coplanar reference points. It is then shown that the observation of two additional known scene points can provide the complete perspective projection. When used with a known object, this technique permits a calibration of the full projective transformation matrix. For an arbitrary non-coplanar set of four points, this calibration provides an affine basis for the reconstruction of local scene structure. When the four points define three orthogonal vectors, the basis is orthogonal, with a metric defined by the lengths of the three vectors. This technique is demonstrated for the case of a cube. Results are presented in which five and a half points on the cube are sufficient to compute the projective transformation for an orthogonal basis by direct observation without matrix inversion. Experiments are outlined for reducing the imprecision due to pixel quantization and noise.> James L. Crowley, Philippe Bobet, Cordelia Schmid |
ICCV | 1 |
| 1993 | A man machine interface for a mobile robotabstractA man machine interface (MMI) designed for human interaction with an autonomous mobile robot is presented. The authors define a high level language for mission specification. This language allows an operator to describe surveillance missions. The model chosen provided an efficient interaction and allowed to implement fundamentals properties for a good man/machine interaction. Several examples were detailed to prove the significance of a simultaneous design phase of the MMI and the functional core of the application, in this case the supervisor. Olivier Causse, James L. Crowley |
IROS | 2 |
| 1993 | Layered Control of a Binocular Camera HeadabstractThis paper describes a layered control system for a binocular stereo head. It begins by a discussion of the principles of layered control. It then describes the mechanical configuration for a binocular camera head with six degrees of freedom. A device level controller is presented which permits an active vision system to command the position of a binocular gaze point in the scene. The final section describes the design of perceptual actions which exploit this device level controller. James L. Crowley, Philippe Bobet, Mouafak Mesrabi |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1993 | Auto-calibration by direct observation of objects
James L. Crowley, Philippe Bobet, Cordelia Schmid |
Image Vis. Comput. | 1 |
| 1993 | Principles and techniques for sensor data fusion
James L. Crowley, Yves Demazeau |
Signal Process. | 1 |
| 1992 | Gaze Control for a Binocular Camera Head
James L. Crowley, Philippe Bobet, Mouafak Mesrabi |
ECCV | 1 |
| 1992 | Position estimation for a mobile robot using vision and odometryabstractThe authors describe a method for locating a mobile robot moving in a known environment. This technique combines position estimation from odometry with observations of the environment from a mobile camera. Fixed objects in the world provide landmarks which are listed in a database. The system calculates the angle to each landmark and then orients the camera. An extended Kalman filter is used to correct the error between the observed and estimated angle to each landmark. Results from experiments in a real environment are presented.> Frédéric Chenavier, James L. Crowley |
ICRA | 2 |
| 1992 | Measurement and integration of 3-D structures by tracking edge lines
James L. Crowley, Patrick Stelmaszyk, Thomas Skordas, Pierre Puget |
Int. J. Comput. Vis. | 1 |
| 1990 | Dynamic World Modeling Using Vertical Line Stereo
James L. Crowley, Philippe Bobet, Karen B. Sarachik |
ECCV | 1 |
| 1990 | Measurement and Integration of 3-D Structures By Tracking Edge Lines
James L. Crowley, Patrick Stelmaszyk |
ECCV | 1 |
| 1989 | World modeling and position estimation for a mobile robot using ultrasonic rangingabstractThe author describes a system for dynamically maintaining a description of the limits to free space for a mobile robot using a belt of ultrasonic range sensors. A model is presented for the uncertainty inherent in such sensors, and the projection of range measurements into external Cartesian coordinates is described. Line segments are then expressed by a set of parameters represented by an estimate and a precision. A process is presented for extracting line segments from adjacent collinear range measurements, and a fast algorithm is presented for matching these line segments to a model of the limits to free space of the robot. A side effect of matching observations to a local model is a correction to the estimated position of the robot at the time that the observation was made. A Kalman filter update equation is developed to permit the correspondence of a line segment to the model to be applied as a correction to estimated position. Examples of segment extraction, position correction and modeling are presented using real ultrasonic data.> James L. Crowley |
ICRA | 1 |
| 1989 | Asynchronous control of orientation and displacement in a robot vehicleabstractThe design of a general-purpose vehicle controller is presented. The controller organization is presented as a three-layer structure. The top layer is an interpreter which assures a control protocol based on asynchronous commands and independent control of orientation and forward displacement. The middle layer is control loop which maintains an estimate of the vehicle's position and orientation, as well as their uncertainties. This control loop generates commands for vehicle displacements in terms of a 'virtual vehicle'. The bottom layer is a translator between the 'virtual vehicle' and the physical vehicle on which the controller is implemented.> James L. Crowley |
ICRA | 1 |
| 1988 | Measuring Image Flow By Tracking Edge-linesabstractThis paper describes a technique for measuring the movement of edge-lines in a sequence of images by maintalning an image plane model. Edge-lines are expressed as a set of parameter vectors representing the center-point, orientation and length of a segment. Each parameter vector is composed of an estimate, a temporal derivative, and their covariance matrix. Line segment parameters in the flow model are updated using a Kalman filter. The eorrespondance of observed edge-lines segments to segments predicted from the flow model is determined by a linear complexity algorithm using distance normalized by covariance. The existence of segments in the flow model is controlled using a confidence factor. This technique is in everyday use as part of a larger system for building 3-D scene descriptions using a camera mounted on a robot arm. A near video-rate hardware implementation is currently under development James L. Crowley, Patrick Stelmaszyk, Christophe Discours |
ICCV | 1 |
| 1987 | Using the composite surface model for perceptual tasksabstractA Composite Surface Model is a structure in which streams of information from diverse sensory sources are integrated into a unified model of the immediate environment. The composite surface model then serves as the basis for planning and executing actions, for learning about objects, and for interpreting the world in terms of known objects. This paper reviews current progress in developing a composite surface model using geometric information. Principles for composite modeling are described, and the role of the composite model in a task oriented robotic system is presented. A set of geometric primitives for surfaces patches, contours and vertices are then defined. A family of interface functions are presented which permit the composite surface model to be used by other processes within a task-oriented robotic system. The use of these interface functions is illustrated by a procedure for the task of finding an object. James L. Crowley |
ICRA | 1 |
| 1987 | Coordination of Action and Perception in a Surveillance Robot
James L. Crowley |
IJCAI | 1 |
| 1987 | Multiple Resolution Representation and Probabilistic Matching of 2-D Gray-Scale ShapeabstractOne approach to pattern classification is to match a structural description of a pattern to models which describe the structural properties of pattern classes. The central problem in structural pattern matching is to determine the correspondence between the symbols which comprise a model and symbols which describe a pattern. The difficulty of determining this correspondence depends critically on the representation that is used to describe patterns. This correspondence presents a probabilistic representation for structural models of pattern classes. Both pattern descriptions and models for pattern classes are based on symbols which represent grayscale information at multiple resolutions. A pattern description is given by a tree of symbols with attribute values. Structural models are represented by a tree of symbols with probabilistic attributes. The position and scale (resolution) of the symbols, as well as other ``features,'' are represented by these attributes. An algorithm is presented for determining the correspondence between symbols in a description of a pattern and symbols in a model of a pattern class. This algorithm uses the connectivity between symbols at different scales to constrain the search for correspondence. An interactive training program for learning models of pattern classes is described, and some conclusions from the work are presented. James L. Crowley, Arthur C. Sanderson |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1986 | Representation and maintenance of a composite surface modelabstractThis paper describes the current state of an on-going project to develop techniques for maintaining a coherent model of the surfaces in a scene from a variety of sources of surface information. The introduction describes the motivation for viewing abstract descriptions of surface data as hypotheses with an explicit representation of the uncertainty of spatial attributes and the uncertainty of existence. Primitive elements are then defined for representing surface patches and contours, and the internal constraints within these primitives are described. Finally, an algorithm is presented for updating a composite surface model based on 3-D information expressed with these primitives. James L. Crowley |
ICRA | 1 |
| 1985 | Dynamic world modeling for an intelligent mobile robot using a rotating ultra-sonic ranging deviceabstractA system which performs task-oriented navigation for an intelligent mobile robot is described in this paper. This navigation system is based on a dynamically maintained model of the local environment, called the "Composite Local Model." The Composite Local Model integrates information from a rotating sonar sensor, the robot's touch sensor and a pre-learned Global Model as the robot moves through its environment. Techniques are described for constructing a line segment description of the most recent sensor scan (the Sensor Model), and for integrating such descriptions to build up a model of the immediate environment (the Composite Local Model). Model integration is based on a process of reinforcing the confidence in consistent information while decaying the confidence in inconsistent information. The estimated position of the robot is corrected by the difference in position between observed sensor signals and the corresponding symbols in the Composite Local Model. This system is useful for navigation in a finite, pre-learned domain such as a house, office, or factory. James L. Crowley |
ICRA | 1 |
| 1985 | Navigation for an intelligent mobile robotabstractA navigation system is described for a mobile robot equipped with a rotating ultrasonic range sensor. This navigation system is based on a dynamically maintained model of the local environment, called the composite local model. The composite local model integrates information from the rotating range sensor, the robot's touch sensor, and a pre-learned global model as the robot moves through its environment. Techniques are described for constructing a line segment description of the most recent sensor scan (the sensor model), and for integrating such descriptions to build up a model of the immediate environment (the composite local model). The estimated position of the robot is corrected by the difference in position between observed sensor signals and the corresponding symbols in the composite local model. A learning technique is described in which the robot develops a global model and a network of places. The network of places is used in global path planning, while the segments are recalled from the global model to assist in local path execution. This system is useful for navigation in a finite, pre-learned domain such as a house, office, or factory. James L. Crowley |
IEEE J. Robotics Autom. | 1 |
| 1984 | A Representation for Shape Based on Peaks and Ridges in the Difference of Low-Pass TransformabstractThis paper defines a multiple resolution representation for the two-dimensional gray-scale shapes in an image. This representation is constructed by detecting peaks and ridges in the difference of lowpass (DOLP) transform. Descriptions of shapes which are encoded in this representation may be matched efficiently despite changes in size, orientation, or position. Motivations for a multiple resolution representation are presented first, followed by the definition of the DOLP transform. Techniques are then presented for encoding a symbolic structural description of forms from the DOLP transform. This process involves detecting local peaks and ridges in each bandpass image and in the entire three-dimensional space defined by the DOLP transform. Linking adjacent peaks in different bandpass images gives a multiple resolution tree which describes shape. Peaks which are local maxima in this tree provide landmarks for aligning, manipulating, and matching shapes. Detecting and linking the ridges in each DOLP bandpass image provides a graph which links peaks within a shape in a bandpass image and describes the positions of the boundaries of the shape at multiple resolutions. Detecting and linking the ridges in the DOLP three-space describes elongated forms and links the largest peaks in the tree. The principles for determining the correspondence between symbols in pairs of such descriptions are then described. Such correspondence matching is shown to be simplified by using the correspondence at lower resolutions to constrain the possible correspondence at higher resolutions. James L. Crowley, Alice C. Parker |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1984 | Fast Computation of the Difference of Low-Pass TransformabstractThis paper defines the difference of low-pass (DOLP) transform and describes a fast algorithm for its computation. The DOLP is a reversible transform which converts an image into a set of bandpass images. A DOLP transform is shown to require O(N2) multiplies and produce O(N log(N)) samples from an N sample image. When Gaussian low-pass filters are used, the result is a set of images which have been convolved with difference of Gaussian (DOG) filters from an exponential set of sizes. A fast computation technique based on ``resampling'' is described and shown to reduce the DOLP transform complexity to O(N log(N)) multiplies and O(N) storage locations. A second technique, ``cascaded convolution with expansion,'' is then defined and also shown to reduce the computational cost to O(N log(N)) multiplies. Combining these two techniques yields an algorithm for a DOLP transform that requires O(N) storage cells and requires O(N) multiplies. James L. Crowley, Richard M. Stern |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |