Nicolas Pugeault

dblp:35/1348 · DBLP profile ↗
← Back
33ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-3455-6280ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 2 since 2021Systems, architecture and hardware · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Robot navigation and mapping · 22% 3D vision · 16% Video understanding and tracking · 14%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping › localization
global localization
0.822020
SeDAR: Reading Floorplans Like a Human - Using Deep Learning to Enable Human-Inspired Localisation · Int. J. Comput. Vis. 2020
SeDAR - Semantic Detection and Ranging: Humans can Localise without LiDAR, can Robots? · ICRA 2018
Computer vision › Video understanding and tracking
sign language recognition
0.532014
Sign Spotting Using Hierarchical Sequential Patterns with Temporal Intervals · CVPR 2014
Sign language recognition using sub-units · J. Mach. Learn. Res. 2012
Sign Language Recognition using Sequential Pattern Trees · CVPR 2012
Computer vision › Vision and language › image captioning › neural image captioning
attention-based image captioning
0.412019
Human Attention in Image Captioning: Dataset and Analysis · ICCV 2019
Computer vision › Vision and language
image captioning
0.412019
Human Attention in Image Captioning: Dataset and Analysis · ICCV 2019
Machine learning › Trustworthy machine learning
interpretability
0.412019
Understanding and Visualizing Deep Visual Saliency Models · CVPR 2019
Computer vision › Image recognition and object detection
saliency prediction
0.412019
Understanding and Visualizing Deep Visual Saliency Models · CVPR 2019
Robotics › Robot navigation and mapping
localization
0.312018
SeDAR - Semantic Detection and Ranging: Humans can Localise without LiDAR, can Robots? · ICRA 2018
Computer vision › 3D vision › visual localization
semantic localization
0.312018
SeDAR - Semantic Detection and Ranging: Humans can Localise without LiDAR, can Robots? · ICRA 2018
Robotics › Robot navigation and mapping › active perception
active reconstruction
0.312017
Taking the Scenic Route to 3D: Optimising Reconstruction from Moving Cameras · ICCV 2017
Robotics › Motion planning and robot control
path planning
0.312017
Taking the Scenic Route to 3D: Optimising Reconstruction from Moving Cameras · ICCV 2017
Computer vision › 3D vision
structure from motion
0.312017
Taking the Scenic Route to 3D: Optimising Reconstruction from Moving Cameras · ICCV 2017
Computer vision › Video understanding and tracking › sign language recognition
sign spotting
0.212014
Sign Spotting Using Hierarchical Sequential Patterns with Temporal Intervals · CVPR 2014
Computer vision › 3D vision › 3d object recognition
multi-view object recognition
0.212013
Multi-view object recognition using view-point invariant shape relations and appearance information · ICRA 2013
Computer vision › Image recognition and object detection
object recognition
0.212013
Multi-view object recognition using view-point invariant shape relations and appearance information · ICRA 2013
Computer vision › Video understanding and tracking
gesture recognition
0.112012
Sign language recognition using sub-units · J. Mach. Learn. Res. 2012
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
sequential classification
0.112012
Sign Language Recognition using Sequential Pattern Trees · CVPR 2012
Robotics › Robot navigation and mapping › localization › map-based localization
floor-plan-based localization
0.112020
SeDAR: Reading Floorplans Like a Human - Using Deep Learning to Enable Human-Inspired Localisation · Int. J. Comput. Vis. 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
0.112020
SeDAR: Reading Floorplans Like a Human - Using Deep Learning to Enable Human-Inspired Localisation · Int. J. Comput. Vis. 2020
Visualization and visual analytics
eye tracking analysis
0.112019
Human Attention in Image Captioning: Dataset and Analysis · ICCV 2019
Visualization and visual analytics
visual analytics
0.112019
Human Attention in Image Captioning: Dataset and Analysis · ICCV 2019
Robotics › Autonomous driving
driving model learning
0.112010
Learning Pre-attentive Driving Behaviour from Holistic Visual Features · ECCV (6) 2010
Computer vision › Segmentation and scene understanding
scene understanding
0.112018
SeDAR - Semantic Detection and Ranging: Humans can Localise without LiDAR, can Robots? · ICRA 2018
Natural language and speech › Information extraction and text analysis › data annotation
semantic annotation
0.112018
SeDAR - Semantic Detection and Ranging: Humans can Localise without LiDAR, can Robots? · ICRA 2018
Computer vision › 3D vision
object pose estimation
0.112009
A Probabilistic Framework for 3D Visual Object Representation · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Computer vision › 3D vision
object representation
0.112009
A Probabilistic Framework for 3D Visual Object Representation · IEEE Trans. Pattern Anal. Mach. Intell. 2009
Computer vision › 3D vision
3d scene understanding
0.112007
A Scene Representation Based on Multi-Modal 2D and 3D Features · ICCV 2007
Computer vision › 3D vision
depth estimation
0.112007
A Scene Representation Based on Multi-Modal 2D and 3D Features · ICCV 2007
Robotics › Robot manipulation › assembly › assembly task
assembly manipulation
0.012013
Multi-view object recognition using view-point invariant shape relations and appearance information · ICRA 2013
Robotics › Motion planning and robot control › robot learning
object learning
0.012007
A Scene Representation Based on Multi-Modal 2D and 3D Features · ICCV 2007
Robotics › Robot manipulation › grasping
vision-based grasping
0.012007
A Scene Representation Based on Multi-Modal 2D and 3D Features · ICCV 2007

Methods — techniques the papers use, named apart from their topics

soft attention · 0.8saliency · 0.8lyapunov exponent regularization · 0.8dreamer v3 · 0.8convolutional neural network · 0.8semantic labeling · 0.4deep learning · 0.4neuron analysis · 0.4fine-tuning · 0.4feature extraction · 0.4
YearPublicationVenuePosition
2026 Splat-Portrait: Generalizing Talking Heads with Gaussian Splatting
Melonie de Almeida, Daniela Ivanova, Nicolas Pugeault, Paul Henderson
MMM (1)4
2026 Guest Editorial: Special Issue for the British Machine Vision Conference (BMVC), 2024 (Glasgow, Scotland, UK)
Carlos Francisco Moreno-García, Gerardo Aragon-Camarasa, Edmond S. L. Ho, Paul Henderson, Nicolas Pugeault, Jungong Han, Sergio Escalera
Int. J. Comput. Vis.5
2025 Beyond Reconstruction: A Physics Based Neural Deferred Shader for Photo-Realistic Rendering
Zhuo He, Paul Henderson, Nicolas Pugeault
ICANN (4)3
2024 Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
Xuri Ge, Joemon M. Jose, Nicolas Pugeault, Paul Henderson
ICPR (21)4
2024 The Bad Batches: Enhancing Self-Supervised Learning in Image Classification Through Representative Batch Curation
abstract
The pursuit of learning robust representations without human supervision is a longstanding challenge. The recent advancements in self-supervised contrastive learning approaches have demonstrated high performance across various representation learning challenges. However, current methods depend on the random transformation of training examples, resulting in some cases of unrepresentative positive pairs that can have a large impact on learning. This limitation not only impedes the convergence of the learning process but the robustness of the learnt representation as well as requiring larger batch sizes to improve robustness to such bad batches. This paper attempts to alleviate the influence of false positive and false negative pairs by employing pairwise similarity calculations through the Fréchet ResNet Distance (FRD), thereby obtaining robust representations from unlabelled data. The effectiveness of the proposed method is substantiated by empirical results, where a linear classifier trained on self-supervised contrastive representations achieved an impressive 87.74% top-1 accuracy on STL10 and 99.31% on the Flower102 dataset. These results emphasize the potential of the proposed approach in pushing the boundaries of the state-of-the-art in self-supervised contrastive learning, particularly for image classification tasks.
Özgü Göksu, Nicolas Pugeault
IJCNN2
2024 Enhancing Robustness in Deep Reinforcement Learning: A Lyapunov Exponent Approach
abstract
Deep reinforcement learning agents achieve state-of-the-art performance in a wide range of simulated control tasks. However, successful applications to real-world problems remain limited. One reason for this dichotomy is because the learnt policies are not robust to observation noise or adversarial attacks. In this paper, we investigate the robustness of deep RL policies to a single small state perturbation in deterministic continuous control tasks. We demonstrate that RL policies can be deterministically chaotic, as small perturbations to the system state have a large impact on subsequent state and reward trajectories. This unstable non-linear behaviour has two consequences: first, inaccuracies in sensor readings, or adversarial attacks, can cause significant performance degradation; second, even policies that show robust performance in terms of rewards may have unpredictable behaviour in practice. These two facets of chaos in RL policies drastically restrict the application of deep RL to real-world problems. To address this issue, we propose an improvement on the successful Dreamer V3 architecture, implementing Maximal Lyapunov Exponent regularisation. This new approach reduces the chaotic state dynamics, rendering the learnt policies more resilient to sensor noise or adversarial attacks and thereby improving the suitability of deep reinforcement learning for real-world applications.
Rory Young, Nicolas Pugeault
NeurIPS2
2023 The role of noise in denoising models for anomaly detection in medical images
Antanas Kascenas, Pedro Sanchez, Patrick Schrempf, William Clackett, Shadia Mikhael, Jeremy Voisey, Keith A. Goatman, Alexander J. Weir, Nicolas Pugeault, Sotirios A. Tsaftaris, Alison O'Neil
Medical Image Anal.10
2020 Image Captioning Through Image Transformer
Sen He 0001, Wentong Liao, Hamed Rezazadegan Tavakoli, Michael Ying Yang, Bodo Rosenhahn, Nicolas Pugeault
ACCV (4)6
2020 Real-time Facial Expression Recognition "In The Wild" by Disentangling 3D Expression from Identity
abstract
Human emotions analysis has been the focus of many studies, especially in the field of Affective Computing, and is important for many applications, e.g. human-computer intelligent interaction, stress analysis, interactive games, animations, etc. Solutions for automatic emotion analysis have also benefited from the development of deep learning approaches and the availability of vast amount of visual facial data on the internet. This paper proposes a novel method for human emotion recognition from a single RGB image. We construct a largescale dataset of facial videos (FaceVid), rich in facial dynamics, identities, expressions, appearance and 3D pose variations. We use this dataset to train a deep Convolutional Neural Network for estimating expression parameters of a 3D Morphable Model and combine it with an effective back-end emotion classifier. Our proposed framework runs at 50 frames per second and is capable of robustly estimating parameters of 3D expression variation and accurately recognizing facial expressions from in the-wild images. We present extensive experimental evaluation that shows that the proposed method outperforms the compared techniques in estimating the 3D expression parameters and achieves state-of-the-art performance in recognising the basic emotions from facial images, as well as recognising stress from facial videos.
Mohammad Rami Koujan, Luma Alharbawee, Giorgos A. Giannakakis, Nicolas Pugeault, Anastasios Roussos
FG4
2020 Transformer-Encoder Detector Module: Using Context to Improve Robustness to Adversarial Attacks on Object Detection
abstract
Deep neural network approaches have demonstrated high performance in object recognition (CNN) and detection (Faster-RCNN) tasks, but experiments have shown that such architectures are vulnerable to adversarial attacks (FFF, UAP): low amplitude perturbations, barely perceptible by the human eye, can lead to a drastic reduction in labelling performance. This article proposes a new context module, called Transformer-Encoder Detector Module, that can be applied to an object detector to (i) improve the labelling of object instances; and (ii) improve the detector's robustness to adversarial attacks. The proposed model achieves higher mAP, F1 scores and AUC average score of up to 13% compared to the baseline Faster-RCNN detector, and an mAP score 8 points higher on images subjected to FFF or UAP attacks due to the inclusion of both contextual and visual features extracted from scene and encoded into the model. The result demonstrates that a simple ad-hoc context module can improve the reliability of object detectors significantly.
Faisal Alamri, Sinan Kalkan, Nicolas Pugeault
ICPR3
2020 SeDAR: Reading Floorplans Like a Human - Using Deep Learning to Enable Human-Inspired Localisation
abstract
Abstract The use of human-level semantic information to aid robotic tasks has recently become an important area for both Computer Vision and Robotics. This has been enabled by advances in Deep Learning that allow consistent and robust semantic understanding. Leveraging this semantic vision of the world has allowed human-level understanding to naturally emerge from many different approaches. Particularly, the use of semantic information to aid in localisation and reconstruction has been at the forefront of both fields. Like robots, humans also require the ability to localise within a structure. To aid this, humans have designed high-level semantic maps of our structures called floorplans. We are extremely good at localising in them, even with limited access to the depth information used by robots. This is because we focus on the distribution of semantic elements, rather than geometric ones. Evidence of this is that humans are normally able to localise in a floorplan that has not been scaled properly. In order to grant this ability to robots, it is necessary to use localisation approaches that leverage the same semantic information humans use. In this paper, we present a novel method for semantically enabled global localisation. Our approach relies on the semantic labels present in the floorplan. Deep Learning is leveraged to extract semantic labels from RGB images, which are compared to the floorplan for localisation. While our approach is able to use range measurements if available, we demonstrate that they are unnecessary as we can achieve results comparable to state-of-the-art without them.
Oscar Mendez 0001, Simon Hadfield, Nicolas Pugeault, Richard Bowden
Int. J. Comput. Vis.3
2019 Understanding and Visualizing Deep Visual Saliency Models
abstract
Recently, data-driven deep saliency models have achieved high performance and have outperformed classical saliency models, as demonstrated by results on datasets such as the MIT300 and SALICON. Yet, there remains a large gap between the performance of these models and the inter-human baseline. Some outstanding questions include what have these models learned, how and where they fail, and how they can be improved. This article attempts to answer these questions by analyzing the representations learned by individual neurons located at the intermediate layers of deep saliency models. To this end, we follow the steps of existing deep saliency models, that is borrowing a pre-trained model of object recognition to encode the visual features and learning a decoder to infer the saliency. We consider two cases when the encoder is used as a fixed feature extractor and when it is fine-tuned, and compare the inner representations of the network. To study how the learned representations depend on the task, we fine-tune the same network using the same image set but for two different tasks: saliency prediction versus scene classification. Our analyses reveal that: 1) some visual regions (e.g. head, text, symbol, vehicle) are already encoded within various layers of the network pre-trained for object recognition, 2) using modern datasets, we find that fine-tuning pre-trained models for saliency prediction makes them favor some categories (e.g. head) over some others (e.g. text), 3) although deep models of saliency outperform classical models on natural images, the converse is true for synthetic stimuli (e.g. pop-out search arrays), an evidence of significant difference between human and data-driven saliency models, and 4) we confirm that, after-fine tuning, the change in inner-representations is mostly due to the task and not the domain shift in the data.
Sen He 0001, Hamed Rezazadegan Tavakoli, Ali Borji, Yang Mi, Nicolas Pugeault
CVPR5
2019 Human Attention in Image Captioning: Dataset and Analysis
abstract
In this work, we present a novel dataset consisting of eye movements and verbal descriptions recorded synchronously over images. Using this data, we study the differences in human attention during free-viewing and image captioning tasks. We look into the relationship between human atten- tion and language constructs during perception and sen- tence articulation. We also analyse attention deployment mechanisms in the top-down soft attention approach that is argued to mimic human attention in captioning tasks, and investigate whether visual saliency can help image caption- ing. Our study reveals that (1) human attention behaviour differs in free-viewing and image description tasks. Hu- mans tend to fixate on a greater variety of regions under the latter task, (2) there is a strong relationship between de- scribed objects and attended objects (97% of the described objects are being attended), (3) a convolutional neural net- work as feature encoder accounts for human-attended re- gions during image captioning to a great extent (around 78%), (4) soft-attention mechanism differs from human at- tention, both spatially and temporally, and there is low correlation between caption scores and attention consis- tency scores. These indicate a large gap between humans and machines in regards to top-down attention, and (5) by integrating the soft attention model with image saliency, we can significantly improve the model's performance on Flickr30k and MSCOCO benchmarks. The dataset can be found at: https://github.com/SenHe/ Human-Attention-in-Image-Captioning.
Sen He 0001, Hamed Rezazadegan Tavakoli, Ali Borji, Nicolas Pugeault
ICCV4
2018 Aggregated Sparse Attention for Steering Angle Prediction
abstract
In this paper, we apply the attention mechanism to autonomous driving for steering angle prediction. We propose the first model, applying the recently introduced sparse attention mechanism to visual domain, as well as the aggregated extension for this model. We show the improvement of the proposed method, comparing to no attention as well as to different types of attention.
Sen He 0001, Dmitry Kangin, Yang Mi, Nicolas Pugeault
ICPR4
2018 SeDAR - Semantic Detection and Ranging: Humans can Localise without LiDAR, can Robots?
abstract
How does a person work out their location using a floorplan? It is probably safe to say that we do not explicitly measure depths to every visible surface and try to match them against different pose estimates in the floorplan. And yet, this is exactly how most robotic scan-matching algorithms operate. Similarly, we do not extrude the 2D geometry present in the floorplan into 3D and try to align it to the real-world. And yet, this is how most vision-based approaches localise. Humans do the exact opposite. Instead of depth, we use high level semantic cues. Instead of extruding the floorplan up into the third dimension, we collapse the 3D world into a 2D representation. Evidence of this is that many of the floorplans we use in everyday life are not accurate, opting instead for high levels of discriminative landmarks. In this work, we use this insight to present a global localisation approach that relies solely on the semantic labels present in the floorplan and extracted from RGB images. While our approach is able to use range measurements if available, we demonstrate that they are unnecessary as we can achieve results comparable to state-of-the-art without them.
Oscar Mendez 0001, Simon Hadfield, Nicolas Pugeault, Richard Bowden
ICRA3
2018 Continuous Control With a Combination of Supervised and Reinforcement Learning
abstract
Reinforcement learning methods have recently achieved impressive results on a wide range of control problems. However, especially with complex inputs, they still require an extensive amount of training data in order to converge to a meaningful solution. This limits their applicability to complex input spaces such as video signals, and makes them impractical for use in complex real world problems, including many of those for video based control. Supervised learning, on the contrary, is capable of learning on a relatively limited number of samples, but relies on arbitrary hand-labelling of data rather than task-derived reward functions, and hence do not yield independent control policies. In this article we propose a novel, model-free approach, which uses a combination of reinforcement and supervised learning for autonomous control and paves the way towards policy based control in real world environments. We use SpeedDreams/TORCS video game to demonstrate that our approach requires much less samples (hundreds of thousands against millions or tens of millions) comparing to the state-of-the-art reinforcement learning techniques on similar data, and at the same time overcomes both supervised and reinforcement learning approaches in terms of quality. Additionally, we demonstrate applicability of the method to MuJoCo control problems.
Dmitry Kangin, Nicolas Pugeault
IJCNN2
2017 Taking the Scenic Route to 3D: Optimising Reconstruction from Moving Cameras
abstract
Reconstruction of 3D environments is a problem that has been widely addressed in the literature. While many approaches exist to perform reconstruction, few of them take an active role in deciding where the next observations should come from. Furthermore, the problem of travelling from the camera's current position to the next, known as pathplanning, usually focuses on minimising path length. This approach is ill-suited for reconstruction applications, where learning about the environment is more valuable than speed of traversal. We present a novel Scenic Route Planner that selects paths which maximise information gain, both in terms of total map coverage and reconstruction accuracy. We also introduce a new type of collaborative behaviour into the planning stage called opportunistic collaboration, which allows sensors to switch between acting as independent Structure from Motion (SfM) agents or as a variable baseline stereo pair. We show that Scenic Planning enables similar performance to state-of-the-art batch approaches using less than 0.00027% of the possible stereo pairs (3% of the views). Comparison against length-based pathplanning approaches show that our approach produces more complete and more accurate maps with fewer frames. Finally, we demonstrate the Scenic Pathplanner's ability to generalise to live scenarios by mounting cameras on autonomous ground-based sensor platforms and exploring an environment.
Oscar Mendez 0001, Simon Hadfield, Nicolas Pugeault, Richard Bowden
ICCV3
2016 Next-Best Stereo: Extending Next-Best View Optimisation For Collaborative Sensors
Oscar Mendez 0001, Simon Hadfield, Nicolas Pugeault, Richard Bowden
BMVC3
2015 Using surfaces and surface relations in an Early Cognitive Vision system
Dirk Kraft, Wail Mustafa, Mila Popovic, Jeppe Barsøe Jessen, Anders Glent Buch, Thiusius Rajeeth Savarimuthu, Nicolas Pugeault, Norbert Krüger
Mach. Vis. Appl.7
2014 Sign Spotting Using Hierarchical Sequential Patterns with Temporal Intervals
abstract
This paper tackles the problem of spotting a set of signs occuring in videos with sequences of signs. To achieve this, we propose to model the spatio-temporal signatures of a sign using an extension of sequential patterns that contain temporal intervals called Sequential Interval Patterns (SIP). We then propose a novel multi-class classifier that organises different sequential interval patterns in a hierarchical tree structure called a Hierarchical SIP Tree (HSP-Tree). This allows one to exploit any subsequence sharing that exists between different SIPs of different classes. Multiple trees are then combined together into a forest of HSP-Trees resulting in a strong classifier that can be used to spot signs. We then show how the HSP-Forest can be used to spot sequences of signs that occur in an input video. We have evaluated the method on both concatenated sequences of isolated signs and continuous sign sequences. We also show that the proposed method is superior in robustness and accuracy to a state of the art sign recogniser when applied to spotting a sequence of signs.
Eng-Jon Ong, Nicolas Pugeault, Richard Bowden
CVPR2
2013 Multi-view object recognition using view-point invariant shape relations and appearance information
abstract
We present an object recognition system coding shape by view-point invariant geometric relations and appearance. In our intelligent work-cell, the system can observe the work space of the robot by 3 pairs of Kinect and stereo cameras allowing for reliable and complete object information. We show that in such a set-up we can achieve high performance already with a low number of training samples. We show this by training the system to classify 56 objects using Random Forest algorithm. This indicates that our approach can be used in contexts such as assembly manipulation which require high reliability of object recognition.
Wail Mustafa, Nicolas Pugeault, Norbert Krüger
ICRA2
2012 Sign Language Recognition using Sequential Pattern Trees
abstract
This paper presents a novel, discriminative, multi-class classifier based on Sequential Pattern Trees. It is efficient to learn, compared to other Sequential Pattern methods, and scalable for use with large classifier banks. For these reasons it is well suited to Sign Language Recognition. Using deterministic robust features based on hand trajectories, sign level classifiers are built from sub-units. Results are presented both on a large lexicon single signer data set and a multi-signer Kinect™ data set. In both cases it is shown to out perform the non-discriminative Markov model approach and be equivalent to previous, more costly, Sequential Pattern (SP) techniques.
Eng-Jon Ong, Helen Cooper, Nicolas Pugeault, Richard Bowden
CVPR3
2012 Sign language recognition using sub-units
Helen Cooper, Eng-Jon Ong, Nicolas Pugeault, Richard Bowden
J. Mach. Learn. Res.3
2011 Temporal accumulation of oriented visual features
Nicolas Pugeault, Norbert Krüger
J. Vis. Commun. Image Represent.1
2010 Learning Pre-attentive Driving Behaviour from Holistic Visual Features
Nicolas Pugeault, Richard Bowden
ECCV (6)1
2010 A compact harmonic code for early vision based on anisotropic frequency channels
Silvio P. Sabatini, Giulia Gastaldi, Fabio Solari, Karl Pauwels, Marc M. Van Hulle, Javier Díaz 0001, Eduardo Ros Vidal, Nicolas Pugeault, Norbert Krüger
Comput. Vis. Image Underst.8
2010 Using multi-modal 3D contours and their relations for vision and robotics
Emre Baseski, Nicolas Pugeault, Sinan Kalkan, Leon Bodenhagen, Justus H. Piater, Norbert Krüger
J. Vis. Commun. Image Represent.2
2009 Learning Objects and Grasp Affordances through Autonomous Exploration
Dirk Kraft, Renaud Detry, Nicolas Pugeault, Emre Baseski, Justus H. Piater, Norbert Krüger
ICVS3
2009 A Probabilistic Framework for 3D Visual Object Representation
abstract
We present an object representation framework that encodes probabilistic spatial relations between 3D features and organizes these features in a hierarchy. Features at the bottom of the hierarchy are bound to local 3D descriptors. Higher level features recursively encode probabilistic spatial configurations of more elementary features. The hierarchy is implemented in a Markov network. Detection is carried out by a belief propagation algorithm, which infers the pose of high-level features from local evidence and reinforces local evidence from globally consistent knowledge, effectively producing a likelihood for the pose of the object in the detection scene. We also present a simple learning algorithm that autonomously builds hierarchies from local object descriptors. We explain how to use our framework to estimate the pose of a known object in an unknown scene. Experiments demonstrate the robustness of hierarchies to input noise, viewpoint changes, and occlusions.
Renaud Detry, Nicolas Pugeault, Justus H. Piater
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Accumulated Visual Representation for Cognitive Vision
abstract
In this paper we present a scheme for accumulating local visual information in 3D, under known motion. Information about the object’s 3D shape is provided by reconstructing local contour descriptors. This shape information is accumulated over time in three ways: 1) disambiguation: erroneous stereo correspondences that are unsuccessfully tracked are discarded. We make use of aspect cues to increase the data association selectivity. 2) correction: the full pose of the reconstructed features is corrected over time using an Kalman Filter approach. 3) completeness: multiple 2 1/2D representations become merged, constructing a full 3D representation of the object. The described system is evaluated quantitatively on three different scenarios.
Nicolas Pugeault, Florentin Wörgötter, Norbert Krüger
BMVC1
2007 A Scene Representation Based on Multi-Modal 2D and 3D Features
abstract
Visually extracted 2D and 3D information have their own advantages and disadvantages that complement each other. Therefore, it is important to be able to switch between the different dimensions according to the requirements of the problem and use them together to combine the reliability of 2D information with the richness of 3D information. In this article, we use 2D and 3D information in a feature-based vision system and demonstrate their complementary properties on different applications (namely: depth prediction, scene interpretation, grasping from vision and object learning).
Emre Baseski, Nicolas Pugeault, Sinan Kalkan, Dirk Kraft, Florentin Wörgötter, Norbert Krüger
ICCV2
2004 Early Cognitive Vision: Using Gestalt-Laws for Task-Dependent, Active Image-Processing
Florentin Wörgötter, Norbert Krüger, Nicolas Pugeault, Dirk Calow, Markus Lappe, Karl Pauwels, Marc M. Van Hulle, Sovira Tan, Alan Johnston
Nat. Comput.3
2003 Multi-Modal Matching Applied to Stereo
abstract
We introduce a compact coding of image information which explicitly sep- arates visual information into geometric information (orientation) and struc- tural information (phase and colour) and temporal information (optic flow). We investigate the importance of these visual attributes for stereo match- ing on a large data set. From these investigation we can conclude that it is the combination of different attributes that gives the best results. Concrete weights for the relative importance of different visual attributes are statisti- cally determined.
Nicolas Pugeault, Norbert Krüger
BMVC1