George Vogiatzis

dblp:36/2989 · DBLP profile ↗
← Back
40ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0002-3226-0603ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 9 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
3D vision · 95% Trustworthy machine learning · 4% Generative modeling · 1%
Computer graphics and multimedia
4 papers
Computational photography and imaging · 62% Geometric modeling and processing · 31% Rendering · 7%

Topics — the 28 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.712023
Self-supervised Monocular Depth Estimation: Let's Talk About The Weather · ICCV 2023
Computer vision › 3D vision › depth estimation
self-supervised depth estimation
0.712023
Self-supervised Monocular Depth Estimation: Let's Talk About The Weather · ICCV 2023
Computer vision › 3D vision › pose estimation › learning-based pose estimation
self-supervised pose estimation
0.712023
Self-supervised Monocular Depth Estimation: Let's Talk About The Weather · ICCV 2023
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.662014
Large Scale Multi-view Stereopsis Evaluation · CVPR 2014
Multiview Photometric Stereo · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Using Multiple Hypotheses to Improve Depth-Maps for Multi-View Stereo · ECCV (1) 2008
Computer vision › 3D vision
photometric stereo
0.542012
Self-calibrated, Multi-spectral Photometric Stereo for 3D Face Capture · Int. J. Comput. Vis. 2012
Overcoming Shadows in 3-Source Photometric Stereo · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Video Normals from Colored Lights · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › 3D vision
3d reconstruction
0.342016
Multiview Photometric Stereo · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Large-Scale Data for Multiple-View Stereopsis · Int. J. Comput. Vis. 2016
Reconstruction in the Round Using Photometric Normals and Silhouettes · CVPR (2) 2006
Computer vision › 3D vision › depth estimation
depth map fusion
0.222012
A Generative Model for Online Depth Fusion · ECCV (5) 2012
Probabilistic visibility for multi-view stereo · CVPR 2007
Machine learning › Trustworthy machine learning › robustness
data augmentation for robustness
0.212023
Self-supervised Monocular Depth Estimation: Let's Talk About The Weather · ICCV 2023
Computational photography and imaging
photometric stereo
0.222008
Shadows in Three-Source Photometric Stereo · ECCV (1) 2008
Non-rigid Photometric Stereo with Colored Lights · ICCV 2007
Computational photography and imaging › spectral imaging
multispectral imaging
0.112012
Self-calibrated, Multi-spectral Photometric Stereo for 3D Face Capture · Int. J. Comput. Vis. 2012
Computer vision › 3D vision
3d shape acquisition
0.112011
Video Normals from Colored Lights · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › 3D vision › 3d reconstruction › non-rigid reconstruction
deformable surface reconstruction
0.112011
Video Normals from Colored Lights · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › 3D vision › photometric stereo
multispectral photometric stereo
0.112011
Video Normals from Colored Lights · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › 3D vision › 3d reconstruction
surface reconstruction
0.112011
Video Normals from Colored Lights · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › 3D vision › 3d reconstruction › shape from silhouette
visual hull
0.122008
Multiview Photometric Stereo · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Reconstruction in the Round Using Photometric Normals and Silhouettes · CVPR (2) 2006
Computer vision › 3D vision › depth estimation
depth map refinement
0.112008
Using Multiple Hypotheses to Improve Depth-Maps for Multi-View Stereo · ECCV (1) 2008
Computational photography and imaging › shape and reflectance estimation
shape from shading
0.112008
Shadows in Three-Source Photometric Stereo · ECCV (1) 2008
Geometric modeling and processing › surface reconstruction › shape reconstruction
shape from shadow
0.112008
Shadows in Three-Source Photometric Stereo · ECCV (1) 2008
Geometric modeling and processing
3d reconstruction
0.112007
Non-rigid Photometric Stereo with Colored Lights · ICCV 2007
Geometric modeling and processing › 3d reconstruction › non-rigid 3d reconstruction
deformable object reconstruction
0.112007
Non-rigid Photometric Stereo with Colored Lights · ICCV 2007
Computational photography and imaging › photometric stereo
multispectral photometric stereo
0.112007
Non-rigid Photometric Stereo with Colored Lights · ICCV 2007
Computer vision › 3D vision › photometric stereo
multi-view photometric stereo
0.112006
Reconstruction in the Round Using Photometric Normals and Silhouettes · CVPR (2) 2006
Computer vision › 3D vision
3d shape reconstruction
0.112005
Using Frontier Points to Recover Shape, Reflectance and Illumunation · ICCV 2005
Rendering › inverse rendering
reflectance and illumination estimation
0.112005
Using Frontier Points to Recover Shape, Reflectance and Illumunation · ICCV 2005
Computer vision › 3D vision › geometric estimation › 3d registration
surface registration
0.012011
Video Normals from Colored Lights · IEEE Trans. Pattern Anal. Mach. Intell. 2011
Computer vision › 3D vision
camera pose estimation
0.012008
Multiview Photometric Stereo · IEEE Trans. Pattern Anal. Mach. Intell. 2008
Computer vision › 3D vision › 3d reconstruction
shape from silhouette
0.012006
Reconstruction in the Round Using Photometric Normals and Silhouettes · CVPR (2) 2006
Computer vision › 3D vision › 3d scene modeling › scene representation
volumetric scene representation
0.012005
Multi-View Stereo via Volumetric Graph-Cuts · CVPR (2) 2005

Methods — techniques the papers use, named apart from their topics

generative model · 0.8pseudo supervised loss · 0.7computer graphics augmentation · 0.7photometric stereo · 0.5self-calibration · 0.3graph cuts · 0.2structured light scanning · 0.2evaluation protocol · 0.2silhouette-based reconstruction · 0.12d tracking · 0.1zero-depth-silhouette initialization · 0.1optical flow · 0.1frontier points · 0.1
YearPublicationVenuePosition
2025 LKSVC: A Novel VANET Caching Method by Integrating Location-Based K-Means Clustering into Spiking Neural Network
Yuanchen Li, George Vogiatzis
AINA (1)4
2025 PSNVC: A Novel Physical-Based Vanet Caching Framework by Integrating SDN and NDN Paradigm
abstract
With the growing demand for on-board services, Vehicular Ad-hoc Network (VANET) caching technologies have attracted significant attention in both academia and industry. In VANET, vehicles request data by broadcasting to network nodes and caching the received data locally. However, the increasing variety of services, escalating data volumes, and the dynamic nature of VANET topologies pose challenges for traditional caching methods in providing high Quality of Service (QoS). This paper proposes a novel physicalbased VANET caching framework, PSNVC, combining Software-Defined Network (SDN) and Named Data Network (NDN). PSNVC uses SDN controllers to distribute globally popular data, embeds an NDN-based request rebroadcast waiting time scheme in vehicle nodes, and adopts physical-based packets broadcasting strategies which is a new VANET caching framework to replace the conventional model. The PSNVC approach addresses the challenges posed by highly dynamic topologies and limited signal transmission ranges. Empirical simulations demonstrate that this method improves VANET caching performance by reducing link load by at least 33%, increasing the local satisfaction ratio by more than 10%, and maintaining the one-hop hit ratio.
Yuanchen Li, George Vogiatzis
WCNC4
2025 A dual-aligned knowledge self-distillation framework for visible-infrared cross-modal person re-identification
abstract
• Dual alignment knowledge self-distillation to better capture modality-invariant/specific features for VI-ReID • Temperature-modulated alignment and confidence-based selective masking to enhance model reliability. • CutSwap augmentation to improve model robustness against intra-class variations and modality discrepancies. • State-of-the-art performance on SYSU-MM01 and RegDB benchmarks. Visible-infrared person re-identification (VI-ReID) significantly enhances identity retrieval across different illumination conditions by matching visible and infrared modalities. However, existing contrastive-learning-based approaches predominantly focus on cross-modal feature alignment, thus undermining model reliability in complex scenarios. To address this challenge, we introduce a Dual Alignment Knowledge Distillation (DAKD) framework that leverages comprehensive self-distillation at both instance and class levels. Our framework incorporates a temperature-modulated alignment strategy, capturing rich modality-invariant generalities as well as modality-specific discriminative details. Additionally, we propose a confidence-based selective masking mechanism that guides the distillation towards confident and informative teacher predictions. To further enhance robustness against modality discrepancies and intra-class variations, we develop a dedicated augmentation technique, CutSwap, which exchanges image channels to simulate realistic cross-modality variations. Extensive experiments on the benchmark SYSU-MM01 and RegDB datasets demonstrate superior performance compared to other state-of-the-art methods, achieving rank-1 accuracies of 76.31% and 94.83%, respectively and validating the efficacy of DAKD in maintaining robust cross-modal alignment while preserving essential identity-specific discriminative information.
Siyuan Deng, Kunhao Yuan, Gerald Schaefer, Shihua Zhou, George Vogiatzis, Yifan Wang 0008, Hui Fang 0003
Knowl. Based Syst.5
2024 BaseBoostDepth: Exploiting Larger Baselines For Self-supervised Monocular Depth Estimation
Kieran Saunders, Luis Manso, George Vogiatzis
BMVC3
2024 Multi-person 3D pose estimation from unlabelled data
abstract
Abstract Its numerous applications make multi-human 3D pose estimation a remarkably impactful area of research. Nevertheless, it presents several challenges, especially when approached using multiple views and regular RGB cameras as the only input. First, each person must be uniquely identified in the different views. Secondly, it must be robust to noise, partial occlusions, and views where a person may not be detected. Thirdly, many pose estimation approaches rely on environment-specific annotated datasets that are frequently prohibitively expensive and/or require specialised hardware. Specifically, this is the first multi-camera, multi-person data-driven approach that does not require an annotated dataset. In this work, we address these three challenges with the help of self-supervised learning. In particular, we present a three-staged pipeline and a rigorous evaluation providing evidence that our approach performs faster than other state-of-the-art algorithms, with comparable accuracy, and most importantly, does not require annotated datasets. The pipeline is composed of a 2D skeleton detection step, followed by a Graph Neural Network to estimate cross-view correspondences of the people in the scenario, and a Multi-Layer Perceptron that transforms the 2D information into 3D pose estimations. Our proposal comprises the last two steps, and it is compatible with any 2D skeleton detector as input. These two models are trained in a self-supervised manner, thus avoiding the need for datasets annotated with 3D ground-truth poses.
Daniel Rodriguez-Criado, Pilar Bachiller-Burgos, George Vogiatzis, Luis Manso
Mach. Vis. Appl.3
2023 Self-supervised Monocular Depth Estimation: Let's Talk About The Weather
abstract
Current, self-supervised depth estimation architectures rely on clear and sunny weather scenes to train deep neural networks. However, in many locations, this assumption is too strong. For example in the UK (2021), 149 days consisted of rain. For these architectures to be effective in real-world applications, we must create models that can generalise to all weather conditions, times of the day and image qualities. Using a combination of computer graphics and generative models, one can augment existing sunny-weather data in a variety of ways that simulate adverse weather effects. While it is tempting to use such data augmentations for self-supervised depth, in the past this was shown to degrade performance instead of improving it. In this paper, we put forward a method that uses augmentations to remedy this problem. By exploiting the correspondence between unaugmented and augmented data we introduce a pseudo-supervised loss for both depth and pose estimation. This brings back some of the benefits of supervised learning while still not requiring any labels. We also make a series of practical recommendations which collectively offer a reliable, efficient framework for weather-related augmentation of self-supervised depth from monocular video. We present extensive testing to show that our method, Robust-Depth, achieves SotA performance on the KITTI dataset while significantly surpassing SotA on challenging, adverse condition data such as DrivingStereo, Foggy CityScape and NuScenes-Night. The project website can be found at https://kieran514.github.io/Robust-Depth-Project/.
Kieran Saunders, George Vogiatzis, Luis Manso
ICCV2
2022 Learning an augmentation strategy for sparse datasets
Renato Barros Arantes, George Vogiatzis, Diego R. Faria
Image Vis. Comput.2
2021 An Efficient Industrial System for Vehicle Tyre (Tire) Detection and Text Recognition Using Deep Learning
abstract
This paper addresses the challenge of reading low contrast text on tyre sidewall images of vehicles in motion. It presents first of its kind, a full scale industrial system which can read tyre codes when installed along driveways such as at gas stations or parking lots with vehicles driving under 10 mph. Tyre circularity is first detected using a circular Hough transform with dynamic radius detection. The detected tyre arches are then unwarped into rectangular patches. A cascade of convolutional neural network (CNN) classifiers is then applied for text recognition. Firstly, a novel proposal generator for the code localization is introduced by integrating convolutional layers producing HOG-like (Histogram of Oriented Gradients) features into a CNN. The proposals are then filtered using a deep network. After the code is localized, character detection and recognition are carried out using two separate deep CNNs. The results (accuracy, repeatability and efficiency) are impressive and show promise for the intended application.
Wajahat Kazmi, Ian T. Nabney, George Vogiatzis, Peter Rose, Alexander Codd
IEEE Trans. Intell. Transp. Syst.3
2020 Variational Recurrent Sequence-to-Sequence Retrieval for Stepwise Illustration
Vishwash Batra, Aparajita Haldar, Yulan He 0001, Hakan Ferhatosmanoglu, George Vogiatzis, Tanaya Guha
ECIR (1)5
2020 Look and Listen: A Multi-modality Late Fusion Approach to Scene Classification for Autonomous Machines
abstract
The novelty of this study consists in a multi-modality approach to scene classification, where image and audio complement each other in a process of deep late fusion. The approach is demonstrated on a difficult classification problem, consisting of two synchronised and balanced datasets of 16,000 data objects, encompassing 4.4 hours of video of 8 environments with varying degrees of similarity. We first extract video frames and accompanying audio at one second intervals. The image and the audio datasets are first classified independently, using a fine-tuned VGG16 and an evolutionary optimised deep neural network, with accuracies of 89.27% and 93.72%, respectively. This is followed by late fusion of the two neural networks to enable a higher order function, leading to accuracy of 96.81% in this multi-modality classifier with synchronised video frames and audio clips. The tertiary neural network implemented for late fusion outperforms classical state-of-the-art classifiers by around 3% when the two primary networks are considered as feature generators. We show that situations where a single-modality may be confused by anomalous data points are now corrected through an emerging higher order integration. Prominent examples include a water feature in a city misclassified as a river by the audio classifier alone and a densely crowded street misclassified as a forest by the image classifier alone. Both are examples which are correctly classified by our multi-modality approach.
Jordan J. Bird, Diego R. Faria, Cristiano Premebida, Anikó Ekárt, George Vogiatzis
IROS5
2020 Goal Density-based Hindsight Experience Prioritization for Multi-Goal Robot Manipulation Reinforcement Learning
abstract
Reinforcement learning for multi-goal robot manipulation tasks is usually challenging, especially when sparse rewards are provided. It often requires millions of data collected before a stable strategy is learned. Recent algorithms like Hindsight Experience Replay (HER) have accelerated the learning process greatly by replacing the original desired goal with one of the achieved points (substitute goals) alongside the same trajectory. However, the selection of previous experience to learn is naively sampled in HER, in which the trajectory selection and the substitute goal sampling is completely random. In this paper, we discuss an experience prioritization strategy for HER that improves the learning efficiency. We propose the Goal Density-based hindsight experience Prioritization (GDP) method that focuses on utilizing the density distribution of the achieved points and prioritizes achieved points which are rarely seen in the replay buffer. These points are used as substitute goals for HER. In addition, we propose an Prioritization Switching with Ensembling Strategy (PSES) method to switch different experience prioritization algorithms during learning, which allows to select the best performance during each learning stage. We evaluate our method with several OpenAI Gym robotic manipulation tasks. The results show that GDP accelerates the learning process in most tasks and can be improved when combining with other prioritization methods using PSES.
Yingyi Kuang, Abraham Itzhak Weinberg, George Vogiatzis, Diego R. Faria
RO-MAN3
2020 Multi-camera Torso Pose Estimation using Graph Neural Networks
abstract
Estimating the location and orientation of humans is an essential skill for service and assistive robots. To achieve a reliable estimation in a wide area such as an apartment, multiple RGBD cameras are frequently used. Firstly, these setups are relatively expensive. Secondly, they seldom perform an effective data fusion using the multiple camera sources at an early stage of the processing pipeline. Occlusions and partial views make this second point very relevant in these scenarios. The proposal presented in this paper makes use of graph neural networks to merge the information acquired from multiple camera sources, achieving a mean absolute error below 125mm for the location and 10° for the orientation using low-resolution RGB images. The experiments, conducted in an apartment with three cameras, benchmarked two different graph neural network implementations and a third architecture based on fully connected layers. The software used has been released as open-source in a public repository1.
Daniel Rodriguez-Criado, Pilar Bachiller-Burgos, Pablo Bustos, George Vogiatzis, Luis Manso
RO-MAN4
2019 QuiltGAN: An Adversarially Trained, Procedural Algorithm for Texture Generation
Renato Barros Arantes, George Vogiatzis, Diego R. Faria
ICVS2
2019 Learning non-metric visual similarity for image retrieval
Noa Garcia, George Vogiatzis
Image Vis. Comput.2
2018 Asymmetric Spatio-Temporal Embeddings for Large-Scale Image-to-Video Retrieval
Noa Garcia, George Vogiatzis
BMVC2
2018 Neural Caption Generation for News Images
Vishwash Batra, Yulan He 0001, George Vogiatzis
LREC3
2017 A deep learning pipeline for semantic facade segmentation
Radwa Fathalla, George Vogiatzis
BMVC2
2017 Optimization of Facade Segmentation Based on Layout Priors
Radwa Fathalla, George Vogiatzis
CAIP (1)2
2016 Large-Scale Data for Multiple-View Stereopsis
Henrik Aanæs, Rasmus R. Jensen, George Vogiatzis, Engin Tola, Anders Bjorholm Dahl
Int. J. Comput. Vis.3
2014 Detection of multiple meaningful primitive geometric models
Radwa Fathalla, George Vogiatzis
BMVC2
2014 Large Scale Multi-view Stereopsis Evaluation
abstract
The seminal multiple view stereo benchmark evaluations from Middlebury and by Strecha et al. have played a major role in propelling the development of multi-view stereopsis methodology. Although seminal, these benchmark datasets are limited in scope with few reference scenes. Here, we try to take these works a step further by proposing a new multi-view stereo dataset, which is an order of magnitude larger in number of scenes and with a significant increase in diversity. Specifically, we propose a dataset containing 80 scenes of large variability. Each scene consists of 49 or 64 accurate camera positions and reference structured light scans, all acquired by a 6-axis industrial robot. To apply this dataset we propose an extension of the evaluation protocol from the Middlebury evaluation, reflecting the more complex geometry of some of our scenes. The proposed dataset is used to evaluate the state of the art multiview stereo algorithms of Tola et al., Campbell et al. and Furukawa et al. Hereby we demonstrate the usability of the dataset as well as gain insight into the workings and challenges of multi-view stereopsis. Through these experiments we empirically validate some of the central hypotheses of multi-view stereopsis, as well as determining and reaffirming some of the central challenges.
Rasmus R. Jensen, Anders Lindbjerg Dahl, George Vogiatzis, Engin Tola, Henrik Aanæs
CVPR3
2012 A Generative Model for Online Depth Fusion
Oliver J. Woodford, George Vogiatzis
ECCV (5)2
2012 Self-calibrated, Multi-spectral Photometric Stereo for 3D Face Capture
George Vogiatzis, Carlos Hernández 0002
Int. J. Comput. Vis.1
2011 Video-based, real-time multi-view stereo
George Vogiatzis, Carlos Hernández 0002
Image Vis. Comput.1
2011 Video Normals from Colored Lights
abstract
We present an algorithm and the associated single-view capture methodology to acquire the detailed 3D shape, bends, and wrinkles of deforming surfaces. Moving 3D data has been difficult to obtain by methods that rely on known surface features, structured light, or silhouettes. Multispectral photometric stereo is an attractive alternative because it can recover a dense normal field from an untextured surface. We show how to capture such data, which in turn allows us to demonstrate the strengths and limitations of our simple frame-to-frame registration over time. Experiments were performed on monocular video sequences of untextured cloth and faces with and without white makeup. Subjects were filmed under spatially separated red, green, and blue lights. Our first finding is that the color photometric stereo setup is able to produce smoothly varying per-frame reconstructions with high detail. Second, when these 3D reconstructions are augmented with 2D tracking results, one can register both the surfaces and relax the homogenous-color restriction of the single-hue subject. Quantitative and qualitative experiments explore both the practicality and limitations of this simple multispectral capture system.
Gabriel J. Brostow, Carlos Hernández 0002, George Vogiatzis, Björn Stenger, Roberto Cipolla
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Overcoming Shadows in 3-Source Photometric Stereo
abstract
Light occlusions are one of the most significant difficulties of photometric stereo methods. When three or more images are available without occlusion, the local surface orientation is overdetermined so that shape can be computed and the shadowed pixels can be discarded. In this paper, we look at the challenging case when only two images are available without occlusion, leading to a one degree of freedom ambiguity per pixel in the local orientation. We show that, in the presence of noise, integrability alone cannot resolve this ambiguity and reconstruct the geometry in the shadowed regions. As the problem is ill-posed in the presence of noise, we describe two regularization schemes that improve the numerical performance of the algorithm while preserving the data. Finally, the paper describes how this theory applies in the framework of color photometric stereo where one is restricted to only three images and light occlusions are common. Experiments on synthetic and real image sequences are presented.
Carlos Hernández 0002, George Vogiatzis, Roberto Cipolla
IEEE Trans. Pattern Anal. Mach. Intell.2
2010 Automatic 3D object segmentation in multiple views using volumetric graph-cuts
Neill D. F. Campbell, George Vogiatzis, Carlos Hernández 0002, Roberto Cipolla
Image Vis. Comput.2
2008 Using Multiple Hypotheses to Improve Depth-Maps for Multi-View Stereo
Neill D. F. Campbell, George Vogiatzis, Carlos Hernández 0002, Roberto Cipolla
ECCV (1)2
2008 Shadows in Three-Source Photometric Stereo
Carlos Hernández 0002, George Vogiatzis, Roberto Cipolla
ECCV (1)2
2008 Reconstructing relief surfaces
George Vogiatzis, Philip Torr 0001, Steven M. Seitz, Roberto Cipolla
Image Vis. Comput.1
2008 Multiview Photometric Stereo
abstract
This paper addresses the problem of obtaining complete, detailed reconstructions of textureless shiny objects. We present an algorithm which uses silhouettes of the object, as well as images obtained under changing illumination conditions. In contrast with previous photometric stereo techniques, ours is not limited to a single viewpoint but produces accurate reconstructions in full 3D. A number of images of the object are obtained from multiple viewpoints, under varying lighting conditions. Starting from the silhouettes, the algorithm recovers camera motion and constructs the object's visual hull. This is then used to recover the illumination and initialise a multi-view photometric stereo scheme to obtain a closed surface reconstruction. There are two main contributions in this paper: Firstly we describe a robust technique to estimate light directions and intensities and secondly, we introduce a novel formulation of photometric stereo which combines multiple viewpoints and hence allows closed surface reconstructions. The algorithm has been implemented as a practical model acquisition system. Here, a quantitative evaluation of the algorithm on synthetic data is presented together with complete reconstructions of challenging real objects. Finally, we show experimentally how even in the case of highly textured objects, this technique can greatly improve on correspondence-based multi-view stereo results.
Carlos Hernández Esteban, George Vogiatzis, Roberto Cipolla
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Automatic 3D Object Segmentation in Multiple Views using Volumetric Graph-Cuts
abstract
We propose an algorithm for automatically obtaining a segmentation of a rigid object in a sequence of images that are calibrated for camera pose and intrinsic parameters. Until recently, the best segmentation results have been obtained by interactive methods that require manual labelling of image regions. Our method requires no user input but instead relies on the camera fixating on the object of interest during the sequence. We begin by learning a model of the object is colour, from the image pixels around the fixation points. We then extract image edges and combine these with the object colour information in a volumetric binary MRF model. The globally optimal segmentation of 3D space is obtained by a graph-cut optimisation. From this segmentation an improved colour model is extracted and the whole process is iterated until convergence. Our first finding is that the fixation constraint, which requires that the object of interest is more or less central in the image, is enough to determine what to segment and initialise an automatic segmentation process. Second, we find that by performing a single segmentation in 3D, we implicitly exploit a 3D rigidity constraint, expressed as silhouette coherency, which significantly improves silhouette quality over independent 2D segmentations. We demonstrate the validity of our approach by providing segmentation results on real sequences.
Neill D. F. Campbell, George Vogiatzis, Carlos Hernández Esteban, Roberto Cipolla
BMVC2
2007 Probabilistic visibility for multi-view stereo
abstract
We present a new formulation to multi-view stereo that treats the problem as probabilistic 3D segmentation. Previous work has used the stereo photo-consistency criterion as a detector of the boundary between the 3D scene and the surrounding empty space. Here we show how the same criterion can also provide a foreground/background model that can predict if a 3D location is inside or outside the scene. This model replaces the commonly used naive foreground model based on ballooning which is known to perform poorly in concavities. We demonstrate how the probabilistic visibility is linked to previous work on depth-map fusion and we present a multi-resolution graph-cut implementation using the new ballooning term that is very efficient both in terms of computation time and memory requirements.
Carlos Hernández 0002, George Vogiatzis, Roberto Cipolla
CVPR2
2007 Non-rigid Photometric Stereo with Colored Lights
abstract
We present an algorithm and the associated capture methodology to acquire and track the detailed 3D shape, bends, and wrinkles of deforming surfaces. Moving 3D data has been difficult to obtain by methods that rely on known surface features, structured light, or silhouettes. Multispec- tral photometric stereo is an attractive alternative because it can recover a dense normal field from an un-textured surface. We show how to capture such data and register it over time to generate a single deforming surface. Experiments were performed on video sequences of un- textured cloth, filmed under spatially separated red, green, and blue light sources. Our first finding is that using zero- depth-silhouettes as the initial boundary condition already produces rather smoothly varying per-frame reconstructions with high detail. Second, when these 3D reconstructions are augmented with 2D optical flow, one can register the first frame's reconstruction to every subsequent frame.
Carlos Hernández 0002, George Vogiatzis, Gabriel J. Brostow, Björn Stenger, Roberto Cipolla
ICCV2
2007 Multiview Stereo via Volumetric Graph-Cuts and Occlusion Robust Photo-Consistency
abstract
This paper presents a volumetric formulation for the multi-view stereo problem which is amenable to a computationally tractable global optimisation using Graph-cuts. Our approach is to seek the optimal partitioning of 3D space into two regions labelled as "object" and "empty" under a cost functional consisting of the following two terms: (1) A term that forces the boundary between the two regions to pass through photo-consistent locations and (2) a ballooning term that inflates the "object" region. To take account of the effect of occlusion on the first term we use an occlusion robust photo-consistency metric based on Normalised Cross Correlation, which does not assume any geometric knowledge about the reconstructed object. The globally optimal 3D partitioning can be obtained as the minimum cut solution of a weighted graph.
George Vogiatzis, Carlos Hernández Esteban, Philip Torr 0001, Roberto Cipolla
IEEE Trans. Pattern Anal. Mach. Intell.1
2006 Reconstruction in the Round Using Photometric Normals and Silhouettes
abstract
This paper addresses the problem of obtaining complete, detailed reconstructions of shiny textureless objects. We present an algorithm which uses silhouettes of the object, as well as images obtained under varying illumination conditions. In contrast with previous photometric stereo techniques, ours is not limited to a single viewpoint and produces accurate reconstructions in full 3D. A number of images of the object are obtained from multiple viewpoints, under varying lighting conditions. Starting from the silhouettes, the algorithm recovers camera motion and constructs the object’s visual hull. This is then used to recover the illumination and initialise a multi-view photometric stereo scheme to obtain a closed surface reconstruction. The contributions of the paper are twofold: Firstly we describe a robust technique to estimate light directions and intensities and secondly, we introduce a novel formulation of photometric stereo which combines multiple viewpoints and hence allows closed surface reconstructions. The algorithm has been implemented as a practical model acquisition system. Here, a quantitative evaluation of the algorithm on synthetic data is presented together with a complete reconstruction of a challenging real object.
George Vogiatzis, Carlos Hernández 0002, Roberto Cipolla
CVPR (2)1
2005 Multi-View Stereo via Volumetric Graph-Cuts
abstract
This paper presents a novel formulation for the multi-view scene reconstruction problem. While this formulation benefits from a volumetric scene representation, it is amenable to a computationally tractable global optimisation using Graph-cuts. The algorithm proposed uses the visual hull of the scene to infer occlusions and as a constraint on the topology of the scene. A photo consistency-based surface cost functional is defined and discretised with a weighted graph. The optimal surface under this discretised functional is obtained as the minimum cut solution of the weighted graph. Our method provides a viewpoint independent surface regularisation, approximate handling of occlusions and a tractable optimisation scheme. Promising experimental results on real scenes as well as a quantitative evaluation on a synthetic scene are presented.
George Vogiatzis, Philip Torr 0001, Roberto Cipolla
CVPR (2)1
2005 Using Frontier Points to Recover Shape, Reflectance and Illumunation
abstract
We describe a method to recover the surface reflectance and the 3D shape of a non-Lambertian object as well as illumination, from a collection of images. It is based on the so-called frontier points, which are extracted from the outlines of an object. Frontier points provide 3D locations on the object surface where the surface normal is known. This information is exploited to infer the surface reflectance of the object and the light distribution of the scene both under varying illumination and fixed vantage point, and under varying vantage point and fixed illumination. We also show how to apply frontier points for shape recovery in photometric stereo. The effectiveness of frontier points for recovering reflectance, illumination and shape is confirmed by a number of experiments on both real and synthetic data.
George Vogiatzis, Paolo Favaro, Roberto Cipolla
ICCV1
2004 Reconstructing Relief Surfaces
abstract
This paper generalizes Markov Random Field (MRF) stereo methods to the generation of surface relief (height) fields rather than disparity or depth maps. This generalization enables the reconstruction of complete object models using the same algorithms that have been previously used to compute depth maps in binocular stereo. In contrast to traditional dense stereo where the parametrization is image based, here we advocate a parametrization by a height field over any base surface. In practice, the base surface is a coarse approximation to the true geometry, e.g., a bounding box, visual hull or triangulation of sparse correspondences, and is assigned or computed using other means. A dense set of sample points is defined on the base surface, each with a fixed normal direction and unknown height value. The estimation of heights for the sample points is achieved by a belief propagation technique. Our method provides a viewpoint independent smoothness constraint, a more compact parametrization and explicit handling of occlusions. We present experimental results on real scenes as well as a quantitative evaluation on an artificial scene.
George Vogiatzis, Philip Torr 0001, Steven M. Seitz, Roberto Cipolla
BMVC1
2003 Bayesian Stochastic Mesh Optimization for 3D reconstruction
abstract
We describe a mesh based approach to the problem of structure from motion. The input to the algorithm is a small set of images, sparse noisy feature correspondences (such as those provided by a Harris corner detector and cross correlation) and the camera geometry plus calibration. The output is a 3D mesh, that when projected onto each view, is visually consistent with the images. There are two contributions in this paper. The first is a Bayesian formulation in which simplicity and smoothness assumptions are encoded in the prior distribution. The resulting posterior is optimized by simulated annealing. The second and more important contribution is a way to make this optimization scheme more efficient. Generic simulated annealing has been long studied in computer vision and is thought to be highly inefficient. This is often because the proposal distribution searches regions of space which are far from the modes. In order to improve the performance of simulated annealing it has long been acknowledged that choice of the correct proposal distribution is of paramount importance to convergence. Taking inspiration from RANSAC andimportance sampling we craft a proposal distribution that is tailored to the problem of structure from motion. This makes our approach particularly robust to noise and ambiguity. We show results for an artificial object and an architectural scene.
George Vogiatzis, Philip Torr 0001, Roberto Cipolla
BMVC1