Chris McCool

dblp:92/3022 · also Christopher McCool · DBLP profile ↗
← Back
34ranked-venue papers
4as first author
7since 2021 · last 2025
0000-0002-0577-1299ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 3 first-author · 7 since 2021Systems, architecture and hardware · 15 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-authorSecurity and privacy · 3Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 A Dataset and Benchmark for Shape Completion of Fruits for Agricultural Robotics
abstract
As the world population is expected to reach 10 billion by 2050, our agricultural production system needs to double its productivity despite a decline of human workforce in the agricultural sector. Autonomous robotic systems are one promising pathway to increase productivity by taking over labor-intensive manual tasks like fruit picking. To be effective, such systems need to monitor and interact with plants and fruits precisely, which is challenging due to the cluttered nature of agricultural environments causing, for example, strong occlusions. Thus, being able to estimate the complete 3D shapes of objects in presence of occlusions is crucial for automating operations such as fruit harvesting. In this paper, we propose the first publicly available 3D shape completion dataset for agricultural vision systems. We provide an RGB-D dataset for estimating the 3D shape of fruits. Specifically, our dataset contains RGB-D frames of single sweet peppers in lab conditions but also in a commercial greenhouse. For each fruit, we additionally collected high-precision point clouds that we use as ground truth. For acquiring the ground truth shape, we developed a measuring process that allows us to record data of real sweet pepper plants, both in the lab and in the greenhouse with high precision, and determine the shape of the sensed fruits. We release our dataset, consisting of almost 7,000 RGB-D frames belonging to more than 100 different fruits. We provide segmented RGB-D frames, with camera intrinsics to easily obtain colored point clouds, together with the corresponding high-precision, occlusion-free point clouds obtained with a high-precision laser scanner. We additionally enable evaluation of shape completion approaches on a hidden test set through a public challenge on a benchmark server.
Federico Magistri, Thomas Läbe, Elias Marks, Sumanth Nagulavancha, Yue Pan 0009, Claus Smitt, Lasse Klingbeil, Michael Halstead, Heiner Kuhlmann, Chris McCool, Jens Behley, Cyrill Stachniss
ICRA10
2023 Knowledge Distillation for Efficient Panoptic Semantic Segmentation: Applied to Agriculture
abstract
Panoptic segmentation provides both holistic and detailed image parsing information at both the pixel and the instance level. However, the computational burdens restrict its applications in real-time scenarios. A potential approach to learn more efficient models is to employ knowledge distillation. However, previous knowledge distillation schemes have focused mainly on classification with limited attention given to rearession-related tasks which is key for panoptic segmentation. In this paper, we establish a logits-based, a hints-based, and a combination-based scheme for panoptic knowledge distillation by using logits from the final layers and features in the middle layers. Then we explore different combinations of balancing weights for optimal solutions according to different network structures and datasets. To validate our proposed approach, various experiments on different datasets have been conducted and efficient networks with higher performance have been obtained. We show that knowledge distillation can be applied to develop accurate ResNet-34 networks improving their panoptic quality on things by an absolute amount of 4.1 points for sweet pepper (glasshouse environment) and 2.2 points for sugar beet (arable farming environment). These student ResNet-34 networks are able to run inference at faster than a framerate of 53Hz on computing infrastructure similar to PATHoBot (a glasshouse robot). To the best of our knowledge, this is the first work to propose knowledge distillation schemes for panoptic semantic segmentation.
Maohui Li, Michael Halstead, Chris McCool
IROS3
2023 Panoptic Mapping with Fruit Completion and Pose Estimation for Horticultural Robots
abstract
Monitoring plants and fruits at high resolution play a key role in the future of agriculture. Accurate 3D information can pave the way to a diverse number of robotic applications in agriculture ranging from autonomous harvesting to precise yield estimation. Obtaining such 3D information is non-trivial as agricultural environments are often repetitive and cluttered, and one has to account for the partial observability of fruit and plants. In this paper, we address the problem of jointly estimating complete 3D shapes of fruit and their pose in a 3D multi-resolution map built by a mobile robot. To this end, we propose an online multi-resolution panoptic mapping system where regions of interest are represented with a higher resolution. We exploit data to learn a general fruit shape representation that we use at inference time together with an occlusion-aware differentiable rendering pipeline to complete partial fruit observations and estimate the 7 DoF pose of each fruit in the map. The experiments presented in this paper, evaluated both in the controlled environment and in a commercial greenhouse, show that our novel algorithm yields higher completion and pose estimation accuracy than existing methods, with an improvement of 41 % in completion accuracy and 52 % in pose estimation accuracy while keeping a low inference time of 0.6 s in average.
Yue Pan 0009, Federico Magistri, Thomas Läbe, Elias Marks, Claus Smitt, Chris McCool, Jens Behley, Cyrill Stachniss
IROS6
2022 Towards Autonomous Visual Navigation in Arable Fields
abstract
Autonomous navigation of a robot in agricultural fields is essential for every task from crop monitoring to weed management and fertilizer application. Many current approaches rely on accurate GPS, however, such technology is expensive and can be impacted by lack of coverage. As such, autonomous navigation through sensors that can interpret their environment (such as cameras) is important to achieve the goal of autonomy in agriculture. In this paper, we introduce a purely vision-based navigation scheme that is able to reliably guide the robot through row-crop fields using computer vision and signal processing techniques without manual intervention. Independent of any global localization or mapping, this approach is able to accurately follow the crop-rows and switch between the rows, only using onboard cameras. The proposed navigation scheme can be deployed in a wide range of fields with different canopy shapes in various growth stages, creating a crop agnostic navigation approach. This was completed under various illumination conditions using simulated and real fields where we achieve an average navigation accuracy of 3.82cm with minimal human intervention (hyper-parameter tuning) on BonnBot-I.
Michael Halstead, Chris McCool
IROS3
2022 BonnBot-I: A Precise Weed Management and Crop Monitoring Platform
abstract
Cultivation and weeding are two of the primary tasks performed by farmers today. A recent challenge for weeding is the desire to reduce herbicide and pesticide treatments while maintaining crop quality and quantity. In this paper we introduce BonnBot-I a precise weed management platform which can also performs field monitoring. Driven by crop monitoring approaches which can accurately locate and classify plants (weed and crop) we further improve their performance by fusing the platform available GNSS and wheel odometry. This improves tracking accuracy of our crop monitoring approach from a normalized average error of 8.3% to 3.5%, evaluated on a new publicly available corn dataset. We also present a novel arrangement of weeding tools mounted on linear actuators evaluated in simulated environments. We replicate weed distributions from a real field, using the results from our monitoring approach, and show the validity of our work-space division techniques which require significantly less movement (a 50% reduction) to achieve similar results. Overall, BonnBot-I is a significant step forward in precise weed management with a novel method of selectively spraying and controlling weeds in an arable field.
Michael Halstead, Chris McCool
IROS3
2021 PATHoBot: A Robot for Glasshouse Crop Phenotyping and Intervention
abstract
We present PATHoBot an autonomous crop surveying and intervention robot for glasshouse environments. The aim of this platform is to autonomously gather high quality data and also estimate key phenotypic parameters. To achieve this we retro-fit an off-the-shelf pipe-rail trolley with an array of multi-modal cameras, navigation sensors and a robotic arm for close surveying tasks and intervention. In this paper we describe PATHoBot design choices made to ensure proper operation in a commercial glasshouse environment. As a surveying platform we collect a number of datasets which include both sweet pepper and tomatoes. We show how PATHoBot enables novel surveillance approaches by first improving our previous work on fruit counting by incorporating wheel odometry and depth information. We find that by introducing re-projection and depth information we are able to achieve an absolute improvement of 20 points over the baseline technique in an "in the wild" situation. Finally, we present a 3D mapping case study, further showcasing PATHoBot’s crop surveying capabilities.
Claus Smitt, Michael Halstead, Tobias Zaenker, Maren Bennewitz, Chris McCool
ICRA5
2021 Viewpoint Planning for Fruit Size and Position Estimation
abstract
Modern agricultural applications require knowledge about the position and size of fruits on plants. However, occlusions from leaves typically make obtaining this information difficult. We present a novel viewpoint planning approach that builds up an octree of plants with labeled regions of interest (ROIs), i.e., fruits. Our method uses this octree to sample viewpoint candidates that increase the information around the fruit regions and evaluates them using a heuristic utility function that takes into account the expected information gain. Our system automatically switches between ROI targeted sampling and exploration sampling, which considers general frontier voxels, depending on the estimated utility. When the plants have been sufficiently covered with the RGB-D sensor, our system clusters the ROI voxels and estimates the position and size of the detected fruits. We evaluated our approach in simulated scenarios and compared the resulting fruit estimations with the ground truth. The results demonstrate that our combined approach outperforms a sampling method that does not explicitly consider the ROIs to generate viewpoints in terms of the number of discovered ROI cells. Furthermore, we show the real-world applicability by testing our framework on a robotic arm equipped with an RGB-D camera installed on an automated pipe-rail trolley in a capsicum glasshouse.
Tobias Zaenker, Claus Smitt, Chris McCool, Maren Bennewitz
IROS3
2020 LiDAR Panoptic Segmentation for Autonomous Driving
abstract
Truly autonomous driving without the need for human intervention can only be attained when self-driving cars fully understand their surroundings. Most of these vehicles rely on a suite of active and passive sensors. LiDAR sensors are a cornerstone in most of these hardware stacks, and leveraging them as a complement to other passive sensors such as RGB cameras is an enticing goal. Understanding the semantic class of each point in a LiDAR sweep is important, as well as knowing to which instance of that class it belongs to. To this end, we present a novel, single-stage, and real-time capable panoptic segmentation approach using a shared encoder with a semantic and instance decoder. We leverage the geometric information of the LiDAR scan to perform a novel, distance- aware tri-linear upsampling, which allows our approach to use larger output strides than using transpose convolutions leading to substantial savings in computation time. Our experimental evaluation and ablation studies for each module show that combining our geometric and semantic embeddings with our learned, variable instance thresholds, a category-specific loss, and the novel trilinear upsampling module leads to higher panoptic quality. We will release the code of our approach in our LiDAR processing library LiDAR-Bonnetal [27].
Andres Milioto, Jens Behley, Chris McCool, Cyrill Stachniss
IROS3
2019 Improving Underwater Obstacle Detection using Semantic Image Segmentation
abstract
This paper presents two novel approaches for improving image-based underwater obstacle detection by combining sparse stereo point clouds with monocular semantic image segmentation. Generating accurate image-based obstacle maps in cluttered underwater environments, such as coral reefs, are essential for robust robotic path planning and navigation. However, these maps can be challenged by factors including visibility, lighting and dynamic objects (e.g. fish) that may lead to falsely identified free space or dynamic objects which trajectory planners may react to undesirably. We propose combining feature-based stereo matching with learning-based segmentation to produce a more robust obstacle map. This approach considers direct binary learning of the presence or absence of underwater obstacles, as well as a multiclass learning approach to classify their distance (near, mid and far) in the scene. An enhancement to the binary map is also shown by including depth information from sparse stereo matching to produce 3D obstacle maps of the scene. The performance is evaluated using field data collected in cluttered, and at times, visually degraded coral reef environments. The results show improved image-wide obstacle detection, rejection of transient objects (such as fish), and range estimation compared to feature-based sparse and dense stereo point clouds alone.
Bilal Arain, Chris McCool, Paul Rigby, Daniel Cagara, Matthew Dunbabin
ICRA2
2019 3D Move to See: Multi-perspective visual servoing towards the next best view within unstructured and occluded environments
abstract
In this paper we present a novel approach termed 3D Move to See (3DMTS) which is based on the principle of finding the next best view using a 3D camera array and a robotic manipulator to obtain multiple samples of the scene from different perspectives. Distinct from traditional visual servoing and next best view approaches, the proposed method uses simultaneously-captured multiple views, scene segmentation and an objective function applied to each perspective to estimate a gradient representing the direction of the next best view in a “single shot”. The method is demonstrated within simulation and on a real robot containing a custom 3D camera array for the challenging scenario of robotic harvesting in a highly occluded and unstructured environment. We show, on a real robotic platform, that by moving the eye-in-hand camera using the gradient of an objective function leads to a locally optimal view of the object of interest, even amongst occlusions. The overall performance of the 3DMTS approach obtains a mean increase in target size of 29.3% compared to a baseline method using a single RGB-D camera, which obtained 9.17%. The results demonstrate qualitatively and quantitatively that the 3DMTS method performed better in most scenarios, and yielded three times the target size compared to the baseline method. Increasing the target size in the image given occlusions can improve robotic systems detecting key object features for further manipulation tasks, such as grasping and harvesting.
Chris Lehnert, Dorian Tsai, Anders P. Eriksson, Chris McCool
IROS4
2017 Towards unsupervised weed scouting for agricultural robotics
abstract
Weed scouting is an important part of modern integrated weed management but can be time consuming and sparse when performed manually. Automated weed scouting and weed destruction has typically been performed using classification systems able to classify a set group of species known a priori. This greatly limits deployability as classification systems must be retrained for any field with a different set of weed species present within them. In order to overcome this limitation, this paper works towards developing a clustering approach to weed scouting which can be utilized in any field without the need for prior species knowledge. We demonstrate our system using challenging data collected in the field from an agricultural robotics platform. We show that considerable improvements can be made by (i) learning low-dimensional (bottleneck) features using a deep convolutional neural network to represent plants in general and (ii) tying views of the same area (plant) together. Deploying this algorithm on in-field data collected by AgBotII, we are able to successfully cluster cotton plants from grasses without prior knowledge or training for the specific plants in the field.
David Hall 0003, Feras Dayoub, Jason Kulk, Chris McCool
ICRA4
2017 The ACRV picking benchmark: A robotic shelf picking benchmark to foster reproducible research
abstract
Robotic challenges like the Amazon Picking Challenge (APC) or the DARPA Challenges are an established and important way to drive scientific progress. They make research comparable on a well-defined benchmark with equal test conditions for all participants. However, such challenge events occur only occasionally, are limited to a small number of contestants, and the test conditions are very difficult to replicate after the main event. We present a new physical benchmark challenge for robotic picking: the ACRV Picking Benchmark. Designed to be reproducible, it consists of a set of 42 common objects, a widely available shelf, and exact guidelines for object arrangement using stencils. A well-defined evaluation protocol enables the comparison of complete robotic systems - including perception and manipulation - instead of sub-systems only. Our paper also describes and reports results achieved by an open baseline system based on a Baxter robot.
Jürgen Leitner, Adam W. Tow, Niko Sünderhauf, Jake E. Dean, Joseph W. Durham, Matthew Cooper 0005, Markus Eich, Chris Lehnert, Ruben Mangels, Chris McCool, Peter Kujala, Lachlan Nicholson, Trung Pham, James Sergeant, Liao Wu, Fangyi Zhang, Ben Upcroft, Peter I. Corke
ICRA10
2017 A transplantable system for weed classification by agricultural robotics
abstract
This work presents a rapidly deployable system for automated precision weeding with minimal human labeling time. This overcomes a limiting factor in robotic precision weeding related to the use of vision-based classification systems trained for species that may not be relevant to specific farms. We present a novel approach to overcome this problem by employing unsupervised weed scouting, weed-group labeling, and finally, weed classification that is trained on the labeled scouting data. This work demonstrates a novel labeling approach designed to maximize labeling accuracy whilst needing to label as few images as possible. The labeling approach is able to provide the best classification results of any of the examined exemplar-based labeling approaches whilst needing to label over seven times fewer images than full data labeling.
David Hall 0003, Feras Dayoub, Tristan Perez, Chris McCool
IROS4
2016 Sweet pepper pose detection and grasping for automated crop harvesting
abstract
This paper presents a method for estimating the 6DOF pose of sweet-pepper (capsicum) crops for autonomous harvesting via a robotic manipulator. The method uses the Kinect Fusion algorithm to robustly fuse RGB-D data from an eye-in-hand camera combined with a colour segmentation and clustering step to extract an accurate representation of the crop. The 6DOF pose of the sweet peppers is then estimated via a nonlinear least squares optimisation by fitting a superellipsoid to the segmented sweet pepper. The performance of the method is demonstrated on a real 6DOF manipulator with a custom gripper. The method is shown to estimate the 6DOF pose successfully enabling the manipulator to grasp sweet peppers for a range of different orientations. The results obtained improve largely on the performance of grasping when compared to a naive approach, which does not estimate the orientation of the crop.
Chris Lehnert, Inkyu Sa, Chris McCool, Ben Upcroft, Tristan Perez
ICRA3
2016 Visual detection of occluded crop: For automated harvesting
abstract
This paper presents a novel crop detection system applied to the challenging task of field sweet pepper (capsicum) detection. The field-grown sweet pepper crop presents several challenges for robotic systems such as the high degree of occlusion and the fact that the crop can have a similar colour to the background (green on green). To overcome these issues, we propose a two-stage system that performs per-pixel segmentation followed by region detection. The output of the segmentation is used to search for highly probable regions and declares these to be sweet pepper. We propose the novel use of the local binary pattern (LBP) to perform crop segmentation. This feature improves the accuracy of crop segmentation from an AUC of 0.10, for previously proposed features, to 0.56. Using the LBP feature as the basis for our two-stage algorithm, we are able to detect 69.2% of field grown sweet peppers in three sites. This is an impressive result given that the average detection accuracy of people viewing the same colour imagery is 66.8%.
Chris McCool, Inkyu Sa, Feras Dayoub, Chris Lehnert, Tristan Perez, Ben Upcroft
ICRA1
2016 Fine-grained classification via mixture of deep convolutional neural networks
abstract
We present a novel deep convolutional neural network (DCNN) system for fine-grained image classification, called a mixture of DCNNs (MixDCNN). The fine-grained image classification problem is characterised by large intra-class variations and small inter-class variations. To overcome these problems our proposed MixDCNN system partitions images into K subsets of similar images and learns an expert DCNN for each subset. The output from each of the K DCNNs is combined to form a single classification decision. In contrast to previous techniques, we provide a formulation to perform joint end-to-end training of the K DCNNs simultaneously. Extensive experiments, on three datasets using two network structures (AlexNet and GoogLeNet), show that the proposed MixDCNN system consistently outperforms other methods. It provides a relative improvement of 12.7% and achieves state-of-the-art results on two datasets.
ZongYuan Ge, Alex Bewley, Chris McCool, Peter I. Corke, Ben Upcroft, Conrad Sanderson
WACV3
2015 Fine-grained bird species recognition via hierarchical subset learning
abstract
We propose a novel method to improve fine-grained bird species classification based on hierarchical subset learning. We first form a similarity tree where classes with strong visual correlations are grouped into subsets. An expert local classifier with strong discriminative power to distinguish visually similar classes is then learnt for each subset. On the challenging Caltech200-2011 bird dataset we show that using the hierarchical approach with features derived from a deep convolutional neural network leads to the average accuracy improving from 64.5% to 72.7%, a relative improvement of 12.7%.
ZongYuan Ge, Chris McCool, Conrad Sanderson, Alex Bewley, Zetao Chen, Peter I. Corke
ICIP2
2015 Modelling local deep convolutional neural network features to improve fine-grained image classification
abstract
We propose a local modelling approach using deep convolutional neural networks (CNNs) for fine-grained image classification. Recently, deep CNNs trained from large datasets have considerably improved the performance of object recognition. However, to date there has been limited work using these deep CNNs as local feature extractors. This partly stems from CNNs having internal representations which are high dimensional, thereby making such representations difficult to model using stochastic models. To overcome this issue, we propose to reduce the dimensionality of one of the internal fully connected layers, in conjunction with layer-restricted retraining to avoid retraining the entire network. The distribution of low-dimensional features obtained from the modified layer is then modelled using a Gaussian mixture model. Comparative experiments show that considerable performance improvements can be achieved on the challenging Fish and UEC FOOD-100 datasets.
ZongYuan Ge, Chris McCool, Conrad Sanderson, Peter I. Corke
ICIP2
2015 Evaluation of Features for Leaf Classification in Challenging Conditions
abstract
Fine-grained leaf classification has concentrated on the use of traditional shape and statistical features to classify ideal images. In this paper we evaluate the effectiveness of traditional hand-crafted features and propose the use of deep convolutional neural network (Conv Net) features. We introduce a range of condition variations to explore the robustness of these features, including: translation, scaling, rotation, shading and occlusion. Evaluations on the Flavia dataset demonstrate that in ideal imaging conditions, combining traditional and Conv Net features yields state-of-the art performance with an average accuracy of 97.3%±0:6% compared to traditional features which obtain an average accuracy of 91.2%±1:6%. Further experiments show that this combined classification approach consistently outperforms the best set of traditional features by an average of 5.7% for all of the evaluated condition variations.
David Hall 0003, Chris McCool, Feras Dayoub, Niko Sünderhauf, Ben Upcroft
WACV2
2014 Local inter-session variability modelling for object classification
abstract
Object classification is plagued by the issue of session variation. Session variation describes any variation that makes one instance of an object look different to another, for instance due to pose or illumination variation. Recent work in the challenging task of face verification has shown that session variability modelling provides a mechanism to overcome some of these limitations. However, for computer vision purposes, it has only been applied in the limited setting of face verification. In this paper we propose a local region based intersession variability (ISV) modelling approach, and apply it to challenging real-world data. We propose a region based session variability modelling approach so that local session variations can be modelled, termed Local ISV. We then demonstrate the efficacy of this technique on a challenging real-world fish image database which includes images taken underwater, providing significant real-world session variations. This Local ISV approach provides a relative performance improvement of, on average, 23% on the challenging MOBIO, Multi-PIE and SCface face databases. It also provides a relative performance improvement of 35% on our challenging fish image dataset.
Kaneswaran Anantharajah, ZongYuan Ge, Chris McCool, Simon Denman, Clinton Fookes, Peter I. Corke, Dian Tjondronegoro, Sridha Sridharan
WACV3
2014 Summarisation of short-term and long-term videos using texture and colour
abstract
We present a novel approach to video summarisation that makes use of a Bag-of-visual-Textures (BoT) approach. Two systems are proposed, one based solely on the BoT approach and another which exploits both colour information and BoT features. On 50 short-term videos from the Open Video Project we show that our BoT and fusion systems both achieve state-of-the-art performance, obtaining an average F-measure of 0.83 and 0.86 respectively, a relative improvement of 9% and 13% when compared to the previous state-of-the-art. When applied to a new underwater surveillance dataset containing 33 long-term videos, the proposed system reduces the amount of footage by a factor of 27, with only minor degradation in the information content. This order of magnitude reduction in video data represents significant savings in terms of time and potential labour cost when manually reviewing such footage.
Johanna Carvajal, Chris McCool, Conrad Sanderson
WACV2
2014 Bi-modal biometric authentication on mobile phones in challenging conditions
Elie Khoury 0001, Laurent El Shafey, Chris McCool, Manuel Günther, Sébastien Marcel
Image Vis. Comput.3
2013 A Scalable Formulation of Probabilistic Linear Discriminant Analysis: Applied to Face Recognition
abstract
In this paper, we present a scalable and exact solution for probabilistic linear discriminant analysis (PLDA). PLDA is a probabilistic model that has been shown to provide state-of-the-art performance for both face and speaker recognition. However, it has one major drawback: At training time estimating the latent variables requires the inversion and storage of a matrix whose size grows quadratically with the number of samples for the identity (class). To date, two approaches have been taken to deal with this problem, to 1) use an exact solution that calculates this large matrix and is obviously not scalable with the number of samples or 2) derive a variational approximation to the problem. We present a scalable derivation which is theoretically equivalent to the previous nonscalable solution and thus obviates the need for a variational approximation. Experimentally, we demonstrate the efficacy of our approach in two ways. First, on labeled faces in the wild, we illustrate the equivalence of our scalable implementation with previously published work. Second, on the large Multi-PIE database, we illustrate the gain in performance when using more training samples per identity (class), which is made possible by the proposed scalable formulation of PLDA.
Laurent El Shafey, Chris McCool, Roy Wallace, Sébastien Marcel
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Bi-modal authentication in mobile environments using session variability modelling
Petr Motlícek, Laurent El Shafey, Roy Wallace, Chris McCool, Sébastien Marcel
ICPR4
2012 Bob: a free signal processing and machine learning toolbox for researchers
abstract
Bob is a free signal processing and machine learning toolbox originally developed by the Biometrics group at Idiap Research Institute, Switzerland. The toolbox is designed to meet the needs of researchers by reducing development time and efficiently processing data. Firstly, Bob provides a researcher-friendly Python environment for rapid development. Secondly, efficient processing of large amounts of multimedia data is provided by fast C++ implementations of identified bottlenecks. The Python environment is integrated seamlessly with the C++ library, which ensures the library is easy to use and extensible. Thirdly, Bob supports reproducible research through its integrated experimental protocols for several databases. Finally, a strong emphasis is placed on code clarity, documentation, and thorough unit testing. Bob is thus an attractive resource for researchers due to this unique combination of ease of use, efficiency, extensibility and transparency. Bob is an open-source library and an ongoing community effort.
André Anjos, Laurent El Shafey, Roy Wallace, Manuel Günther, Chris McCool, Sébastien Marcel
ACM Multimedia5
2012 Cross-Pollination of Normalization Techniques From Speaker to Face Authentication Using Gaussian Mixture Models
abstract
This paper applies score and feature normalization techniques to parts-based Gaussian mixture model (GMM) face authentication. In particular, we propose to utilize techniques that are well established in state-of-the-art speaker authentication, and apply them to the face authentication task. For score normalization, T-, Z- and ZT-norm techniques are evaluated. For feature normalization, we propose a generalization of feature warping to 2D images, which is applied to discrete cosine transform (DCT) features prior to modeling. Evaluation is performed on a range of challenging databases relevant to forensics and security, including surveillance and access control scenarios. The normalization techniques are shown to generalize well to the face authentication task, resulting in relative improvements in half total error rate (HTER) of between 17% and 62%.
Roy Wallace, Mitchell McLaren, Chris McCool, Sébastien Marcel
IEEE Trans. Inf. Forensics Secur.3
2011 Inter-session variability modelling and joint factor analysis for face authentication
abstract
This paper applies inter-session variability modelling and joint factor analysis to face authentication using Gaussian mixture models. These techniques, originally developed for speaker authentication, aim to explicitly model and remove detrimental within-client (inter-session) variation from client models. We apply the techniques to face authentication on the publicly-available BANCA, SCface and MO- BIO databases. We propose a face authentication protocol for the challenging SCface database, and provide the first results on the MOBIO still face protocol. The techniques provide relative reductions in error rate of up to 44%, using only limited training data. On the BANCA database, our results represent a 31% reduction in error rate when benchmarked against previous work.
Roy Wallace, Mitchell McLaren, Chris McCool, Sébastien Marcel
IJCB3
2010 A principled approach to remove false alarms by modelling the context of a face detector
abstract
Face detection [1, 6] is the task of classifying a sub-window as being a face or not. There are many ways to obtain sub-windows from an image, with the sliding window approach being the most well known. This can result in multiple detections and false alarms. A merging and pruning heuristic algorithm is then typically used to output the final detections [3]. Recent work has been done to overcome the limitations of the sliding window approach by using a branch-and-bound technique to evaluate all possible sub-windows in an efficient way [2]. A different approach was recently proposed in [4] and [5] where they show that the score distribution is significantly different around a true object location than around a false alarm location. We propose a model to enhance a given face classifier, by discriminating false detections (sub-windows) from true detections using the contextual information. Our approach follows the work of [4, 5], but we propose a more discriminative approach and we extract a larger variety of features. We investigate the detection distribution around some sub-window (which we call the context) from which we compute features from every possible axis combination (location and scale). The main advantages of our method is that it can be initialized with any sub-window collection and it poses no restriction regarding the object classifier to run on top of. To build the context of a target sub-window Tsw = (x,y,s), we sample in the 3D space of location (x,y) and scale (s) to collect detections. Then the context of Tsw consists of collection of 4D points C(Tsw) = {(xi,yi,si,msi)i=1,..}, where ms is the classifier score. We propose two strategies for context sampling: full and axis. The full strategy consists of sampling by varying the location and scale at the same time, while the axis strategy the sampling is done just along one axis at a time. The feature vectors are defined by their attribute and the axis combination (x, y and s) used to obtain the attribute. We use 5 attributes that capture the global information (counts), the geometry of the detection distribution (hits) and the detection confidence (score) obtained from the face classifier.
Cosmin Atanasoaei, Chris McCool, Sébastien Marcel
BMVC2
2010 On the vulnerability of face verification systems to hill-climbing attacks
Javier Galbally, Chris McCool, Julian Fierrez, Sébastien Marcel, Javier Ortega-Garcia
Pattern Recognit.2
2010 Feature distribution modelling techniques for 3D face verification
Chris McCool, Jordi Sanchez-Riera, Sébastien Marcel
Pattern Recognit. Lett.1
2010 An Evaluation of Video-to-Video Face Verification
abstract
Person recognition using facial features, e.g., mug-shot images, has long been used in identity documents. However, due to the widespread use of web-cams and mobile devices embedded with a camera, it is now possible to realize facial video recognition, rather than resorting to just still images. In fact, facial video recognition offers many advantages over still image recognition; these include the potential of boosting the system accuracy and deterring spoof attacks. This paper presents an evaluation of person identity verification using facial video data, organized in conjunction with the International Conference on Biometrics (ICB 2009). It involves 18 systems submitted by seven academic institutes. These systems provide for a diverse set of assumptions, including feature representation and preprocessing variations, allowing us to assess the effect of adverse conditions, usage of quality information, query selection, and template construction for video-to-video face authentication.
Norman Poh, Chi-Ho Chan, Josef Kittler, Sébastien Marcel, Chris McCool, Enrique Argones-Rúa, José Luis Alba-Castro, Mauricio Villegas, Roberto Paredes, Vitomir Struc, Nikola Pavesic, Albert Ali Salah, Hui Fang 0003, Nicholas Costen
IEEE Trans. Inf. Forensics Secur.5
2008 3D face verification using a free-parts approach
Chris McCool, Vinod Chandran, Sridha Sridharan, Clinton Fookes
Pattern Recognit. Lett.1
2006 Combined 2D/3D Face Recognition Using Log-Gabor Templates
abstract
The addition of Three Dimensional (3D) data has the potential to greatly improve the accuracy of Face Recognition Technologies by providing complementary information. In this paper a new method combining intensity and range images and providing insensitivity to expression variation based on Log-Gabor Templates is presented. By breaking a single image into 75 semi-independent observations the reliance of the algorithm upon any particular part of the face is relaxed allowing robustness in the presence of occulusions, distortions and facial expressions. Also presented is a new distance measure based on the Mahalanobis Cosine metric which has desirable discriminatory characteristics in both the 2D and 3D domains. Using the 3D database collected by University of Notre Dame for the Face Recognition Grand Challenge (FRGC), benchmarking results are presented demonstrating the performance of the proposed methods.
Jamie Cook, Chris McCool, Vinod Chandran, Sridha Sridharan
AVSS2
2006 Feature Modelling of PCA Difference Vectors for 2D and 3D Face Recognition
abstract
This paper examines the the effectiveness of feature modelling to conduct 2D and 3D face recognition. In particular, PCA difference vectors are modelled using Gaussian Mixture Models (GMMs) which describe Intra-Personal (IP) and Extra-Personal (EP) variations. Two classifiers, an IP and IPEP classifier, are formed using these GMMs and their performance is compared to that of the Mahalanobis cosine metric (MahCosine). The best results for the 2D and 3D face modalities are obtained with the IP and IPEP classifiers respectively. The multi-modal fusion of these two systems provided consistent performance improvement across the FRGC database v2.0.
Chris McCool, Jamie Cook, Vinod Chandran, Sridha Sridharan
AVSS1