EDBT 2026 Demo / reviewers in the wild / expert
Jean-Yves Guillemaut
dblp:98/339
· DBLP profile ↗
45ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0001-8223-5505ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 21 · 6 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MVL-Net: Pairwise Learning for Multi-View Multiple People LabellingabstractIn the multi-view domain, it is challenging to correctly label multiple people across viewpoints because of occlusions, visual ambiguities, appearance variation, etc. Deep learning, although having witnessed remarkable success in computer vision tasks, still remains underexplored for the multi-view labelling task, due to the lack of labelled multi-view datasets. In this paper, we propose a novel end-to-end deep neural network named Multi-View Labelling network (MVL-net) that addresses this issue. To overcome the dataset shortage, a large-scale multi-view dataset is generated by combining 3D human models and panoramic backgrounds, along with human poses and realistic rendering. In the proposed MVL-net, we first incorporate Transformer blocks to capture the non-local information for multi-view feature extraction. A matching net is then introduced to achieve multiple people labelling, by predicting matching confidence scores for pairwise instances from two views, thus addressing the problem of the unknown number of people when labelling across views. An additional geometry feature obtained from the epipolar geometry is integrated to leverage multi-view cues during training. To the best of our knowledge, the MVL-net is the first work using deep learning to train a multi-view labelling network. Comprehensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed method, which outperforms the existing state-of-the-art approaches. Yue Zhang 0082, Akin Caliskan, Mai Xu, Adrian Hilton 0001, Jean-Yves Guillemaut |
IEEE Trans. Multim. | 5 |
| 2024 | Learning Self-Shadowing for Clothed Human Bodies
Farshad Einabadi, Jean-Yves Guillemaut, Adrian Hilton 0001 |
EGSR (ST) | 2 |
| 2023 | A toolkit of approaches for digital mapping and correction of visual distortionabstractVisual distortion, known as metamorphopsia, is a serious visual deficit with no effective clinical treatment and which cannot be corrected by traditional optical glasses. In this paper, we introduce a toolkit of approaches for digitally mapping and correcting visual distortion, that might eventually be incorporated in a low vision aid VR headset. We describe three different approaches spanning data-driven and generative designs and leveraging either uniocular or binocular cues. We present our proposed demonstrator, our evaluation roadmap, and challenges for the field. Initial tests with simulated data demonstrate the effectiveness of the approach. Once clinically validated, we hope these approaches will enable accurate mapping of visual distortion and eventually lead to the development of ‘digital glasses’ capable of correcting the effects of metamorphopsia and restoring healthy vision. Ye Ling, David M. Frohlich, Tom H. Williamson, Jean-Yves Guillemaut |
ASSETS | 4 |
| 2023 | Learning Projective Shadow Textures for Neural Rendering of Human Cast Shadows from Silhouettes
Farshad Einabadi, Jean-Yves Guillemaut, Adrian Hilton 0001 |
EGSR (ST) | 2 |
| 2023 | A Family of Approaches for Full 3D Reconstruction of Objects with Complex Surface ReflectanceabstractAbstract 3D reconstruction of general scenes remains an open challenge with current techniques often reliant on assumptions on the scene’s surface reflectance, which restrict the range of objects that can be modelled. Helmholtz Stereopsis offers an appealing framework to make the modelling process agnostic to surface reflectance. However, previous formulations have been almost exclusively limited to 2.5D modelling. To address this gap, this paper introduces a family of reconstruction approaches that exploit Helmholtz reciprocity to produce complete 3D models of objects with arbitrary unknown reflectance. This includes an approach based on the fusion of (orthographic or perspective) view-dependent reconstructions, a volumetric approach optimising surface location within a voxel grid, and a mesh-based formulation optimising vertices positions of a given mesh topology. The contributed approaches are evaluated on synthetic and real datasets, including novel full 3D datasets publicly released with this paper, with experimental comparison against a wide range of competing methods. Results demonstrate the benefits of the different approaches and their abilities to achieve high quality full 3D reconstructions of complex objects. Gianmarco Addari, Jean-Yves Guillemaut |
Int. J. Comput. Vis. | 2 |
| 2022 | There and Back Again: 3D Sign Language Generation from Text Using Back-TranslationabstractWe introduce the first method to automatically generate 3D mesh sequences from text, inspired by the challenging problem of Sign Language Production (SLP). The approach only requires simple 2D annotations for training, which can be automatically extracted from video. Rather than incorporating high-definition or motion capture data, we propose back-translation as a powerful paradigm for supervision: By first addressing the arguably simpler problem of translating 2D pose sequences to text, we can leverage this to drive a transformer-based architecture to translate text to 2D poses. These are then used to drive a 3D mesh generator. Our mesh generator Pose2Mesh uses temporal information, to enforce temporal coherence and significantly reduce processing time. The approach is evaluated by generating 2D pose, and 3D mesh sequences in DGS (German Sign Language) from German language sentences. An extensive analysis of the approach and its sub-networks is conducted, reporting BLEU and ROUGE scores, as well as Mean 2D Joint Distance. Our proposed Text2Pose model outperforms the current state-of-the-art in SLP, and we establish the first benchmark for the complex task of text-to-3D-mesh-sequence generation with our Text2Mesh model. Stephanie Stoll, Armin Mustafa, Jean-Yves Guillemaut |
3DV | 3 |
| 2022 | Finite Aperture StereoabstractAbstract Multi-view stereo remains a popular choice when recovering 3D geometry, despite performance varying dramatically according to the scene content. Moreover, typical pinhole camera assumptions fail in the presence of shallow depth of field inherent to macro-scale scenes; limiting application to larger scenes with diffuse reflectance. However, the presence of defocus blur can itself be considered a useful reconstruction cue, particularly in the presence of view-dependent materials. With this in mind, we explore the complimentary nature of stereo and defocus cues in the context of multi-view 3D reconstruction; and propose a complete pipeline for scene modelling from a finite aperature camera that encompasses image formation, camera calibration and reconstruction stages. As part of our evaluation, an ablation study reveals how each cue contributes to the higher performance observed over a range of complex materials and geometries. Though of lesser concern with large apertures, the effects of image noise are also considered. By introducing pre-trained deep feature extraction into our cost function, we show a step improvement over per-pixel comparisons; as well as verify the cross-domain applicability of networks using largely in-focus training data applied to defocused images. Finally, we compare to a number of modern multi-view stereo methods, and demonstrate how the use of both cues leads to a significant increase in performance across several synthetic and real datasets. Matthew Bailey 0003, Adrian Hilton 0001, Jean-Yves Guillemaut |
Int. J. Comput. Vis. | 3 |
| 2021 | A Novel Multi-View Labelling Network Based on Pairwise LearningabstractCorrect labelling of multiple people from different viewpoints in complex scenes is a challenging task due to occlusions, visual ambiguities, as well as variations in appearance and illumination. In recent years, deep learning approaches have proved very successful at improving the performance of a wide range of recognition and labelling tasks such as person re-identification and video tracking. However, to date, applications to multi-view tasks have proved more challenging due to the lack of suitably labelled multi-view datasets, which are difficult to collect and annotate. The contributions of this paper are two-fold. First, a synthetic dataset is generated by combining 3D human models and panoramas along with human poses and appearance detail rendering to overcome the shortage of real dataset for multi-view labelling. Second, a novel framework named Multi-View Labelling network (MVL-net) is introduced to leverage the new dataset and unify the multi-view multiple people detection, segmentation and labelling tasks in complex scenes. To the best of our knowledge, this is the first work using deep learning to train a multi-view labelling network. Experiments conducted on both synthetic and real datasets demonstrate that the proposed method outperforms the existing state-of-the-art approaches. Yue Zhang 0082, Akin Caliskan, Adrian Hilton 0001, Jean-Yves Guillemaut |
ICIP | 4 |
| 2021 | Deep Neural Models for Illumination Estimation and Relighting: A SurveyabstractAbstract Scene relighting and estimating illumination of a real scene for insertion of virtual objects in a mixed‐reality scenario are well‐studied challenges in the computer vision and graphics fields. Classical inverse rendering approaches aim to decompose a scene into its orthogonal constituting elements, namely scene geometry, illumination and surface materials, which can later be used for augmented reality or to render new images under novel lighting or viewpoints. Recently, the application of deep neural computing to illumination estimation, relighting and inverse rendering has shown promising results. This contribution aims to bring together in a coherent manner current advances in this conjunction. We examine in detail the attributes of the proposed approaches, presented in three categories: scene illumination estimation, relighting with reflectance‐aware scene‐specific representations and finally relighting as image‐to‐image transformations. Each category is concluded with a discussion on the main characteristics of the current methods and possible future trends. We also provide an overview of current publicly available datasets for neural lighting applications. Farshad Einabadi, Jean-Yves Guillemaut, Adrian Hilton 0001 |
Comput. Graph. Forum | 2 |
| 2021 | Temporally Coherent General Dynamic Scene ReconstructionabstractAbstract Existing techniques for dynamic scene reconstruction from multiple wide-baseline cameras primarily focus on reconstruction in controlled environments, with fixed calibrated cameras and strong prior constraints. This paper introduces a general approach to obtain a 4D representation of complex dynamic scenes from multi-view wide-baseline static or moving cameras without prior knowledge of the scene structure, appearance, or illumination. Contributions of the work are: an automatic method for initial coarse reconstruction to initialize joint estimation; sparse-to-dense temporal correspondence integrated with joint multi-view segmentation and reconstruction to introduce temporal coherence; and a general robust approach for joint segmentation refinement and dense reconstruction of dynamic scenes by introducing shape constraint. Comparison with state-of-the-art approaches on a variety of complex indoor and outdoor scenes, demonstrates improved accuracy in both multi-view segmentation and dense reconstruction. This paper demonstrates unsupervised reconstruction of complete temporally coherent 4D scene models with improved non-rigid object segmentation and shape reconstruction and its application to various applications such as free-view rendering and virtual reality. Armin Mustafa, Marco Volino, Hansung Kim 0001, Jean-Yves Guillemaut, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 4 |
| 2021 | Full-Reference Stereoscopic Video Quality Assessment Using a Motion Sensitive HVS ModelabstractStereoscopic video quality assessment has become a major research topic in recent years. Existing stereoscopic video quality metrics are predominantly based on stereoscopic image quality metrics extended to the time domain via for example temporal pooling. These approaches do not explicitly consider the motion sensitivity of the Human Visual System (HVS). To address this limitation, this paper introduces a novel HVS model inspired by physiological findings characterising the motion sensitive response of complex cells in the primary visual cortex (V1 area). The proposed HVS model generalises previous HVS models, which characterised the behaviour of simple and complex cells but ignored motion sensitivity, by estimating optical flow to measure scene velocity at different scales and orientations. The local motion characteristics (direction and amplitude) are used to modulate the output of complex cells. The model is applied to develop a new type of full-reference stereoscopic video quality metrics which uniquely combine non-motion sensitive and motion sensitive energy terms to mimic the response of the HVS. A tailored two-stage multi-variate stepwise regression algorithm is introduced to determine the optimal contribution of each energy term. The two proposed stereoscopic video quality metrics are evaluated on three stereoscopic video datasets. Results indicate that they achieve average correlations with subjective scores of 0.9257 (PLCC), 0.9338 and 0.9120 (SRCC), 0.8622 and 0.8306 (KRCC), and outperform previous stereoscopic video quality metrics including other recent HVS-based metrics. Chathura Galkandage, Janko Calic, Safak Dogan, Jean-Yves Guillemaut |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | A Novel Depth from Defocus Framework Based on a Thick Lens Camera ModelabstractReconstruction approaches based on monocular defocus analysis such as Depth from Defocus (DFD) often utilise the thin lens camera model. Despite this widespread adoption, there are inherent limitations associated with it. Coupled with invalid parameterisation commonplace in literature, the overly-simplified image formation it describes leads to inaccurate defocus modelling; especially in macro-scale scenes. As a result, DFD reconstructions based around this model are not geometrically consistent, and are typically restricted to single-view applications. Subsequently, the handful of existing approaches which attempt to include additional viewpoints have had only limited success.In this work, we address these issues by instead utilising a thick lens camera model, and propose a novel calibration procedure to accurately parameterise it. The effectiveness of our model and calibration is demonstrated with a novel DFD reconstruction framework. We achieve highly detailed, geometrically accurate and complete 3D models of real-world scenes from multi-view focal stacks. To our knowledge, this is the first time DFD has been successfully applied to complete scene modelling in this way. Matthew Bailey 0003, Jean-Yves Guillemaut |
3DV | 2 |
| 2019 | Light Field Compression using Eigen TexturesabstractLight fields are becoming an increasingly popular method of digital content production for visual effects and virtual/augmented reality as they capture a view dependent representation enabling photo realistic rendering over a range of viewpoints. Light field video is generally captured using arrays of cameras resulting in tens to hundreds of images of a scene at each time instance. An open problem is how to efficiently represent the data preserving the view-dependent detail of the surface in such a way that is compact to store and efficient to render. In this paper we show that constructing an Eigen texture basis representation from the light field using an approximate 3D surface reconstruction as a geometric proxy provides a compact representation that maintains view-dependent realism. We demonstrate that the proposed method is able to reduce storage requirements by > 95% while maintaining the visual quality of the captured data. An efficient view-dependent rendering technique is also proposed which is performed in eigen space allowing smooth continuous viewpoint interpolation through the light field. Marco Volino, Armin Mustafa, Jean-Yves Guillemaut, Adrian Hilton 0001 |
3DV | 3 |
| 2019 | A family of globally optimal branch-and-bound algorithms for 2D-3D correspondence-free registrationabstractWe present a family of methods for 2D–3D registration spanning both deterministic and non-deterministic branch-and-bound approaches. Critically, the methods exhibit invariance to the underlying scene primitives, enabling e.g. points and lines to be treated on an equivalent basis, potentially enabling a broader range of problems to be tackled while maximising available scene information, all scene primitives being simultaneously considered. Being a branch-and-bound based approach, the method furthermore enjoys intrinsic guarantees of global optimality; while branch-and-bound approaches have been employed in a number of computer vision contexts, the proposed method represents the first time that this strategy has been applied to the 2D–3D correspondence-free registration problem from points and lines. Within the proposed procedure, deterministic and probabilistic procedures serve to speed up the nested branch-and-bound search while maintaining optimality. Experimental evaluation with synthetic and real data indicates that the proposed approach significantly increases both accuracy and robustness compared to the state of the art. David Windridge, Jean-Yves Guillemaut |
Pattern Recognit. | 3 |
| 2019 | Hybrid Modeling of Non-Rigid Scenes From RGBD CamerasabstractRecent advances in sensor technology have introduced low-cost RGB video plus depth sensors, such as the Kinect, which enable simultaneous acquisition of color and depth images at video rates. This paper introduces a framework for representation of general dynamic scenes from video plus depth acquisition. A hybrid representation is proposed which combines the advantages of prior surfel graph surface segmentation and modeling work with the higher resolution surface reconstruction capability of volumetric fusion techniques. The contributions are: 1) extension of a prior piecewise surfel graph modeling approach for improved accuracy and completeness; 2) combination of this surfel graph modeling with a truncated signed distance function surface fusion to generate dense geometry; and 3) proposal of means for validation of the reconstructed a 4D scene model against the input data and efficient storage of any unmodeled regions via residual depth maps. The approach allows arbitrary dynamic scenes to be efficiently represented with a temporally consistent structure and enhanced levels of detail and completeness where possible, but gracefully falls back to raw measurements where no structure can be inferred. The representation is shown to facilitate creative manipulation of real scene data which would previously require more complex capture set-ups or manual processing. Charles Malleson, Jean-Yves Guillemaut, Adrian Hilton 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Bayesian Helmholtz Stereopsis with Integrability PriorabstractHelmholtz Stereopsis is a 3D reconstruction method uniquely independent of surface reflectance. Yet, its sub-optimal maximum likelihood formulation with drift-prone normal integration limits performance. Via three contributions this paper presents a complete novel pipeline for Helmholtz Stereopsis. First, we propose a Bayesian formulation replacing the maximum likelihood problem by a maximum a posteriori one. Second, a tailored prior enforcing consistency between depth and normal estimates via a novel metric related to optimal surface integrability is proposed. Third, explicit surface integration is eliminated by taking advantage of the accuracy of prior and high resolution of the coarse-to-fine approach. The pipeline is validated quantitatively and qualitatively against alternative formulations, reaching sub-millimetre accuracy and coping with complex geometry and reflectance. Nadejda Roubtsova, Jean-Yves Guillemaut |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | 4D Temporally Coherent Light-Field VideoabstractLight-field video has recently been used in virtual and augmented reality applications to increase realism and immersion. However, existing light-field methods are generally limited to static scenes due to the requirement to acquire a dense scene representation. The large amount of data and the absence of methods to infer temporal coherence pose major challenges in storage, compression and editing compared to conventional video. In this paper, we propose the first method to extract a spatio-temporally coherent light-field video representation. A novel method to obtain Epipolar Plane Images (EPIs) from a spare lightfield camera array is proposed. EPIs are used to constrain scene flow estimation to obtain 4D temporally coherent representations of dynamic light-fields. Temporal coherence is achieved on a variety of light-field datasets. Evaluation of the proposed light-field scene flow against existing multi-view dense correspondence approaches demonstrates a significant improvement in accuracy of temporal coherence. Armin Mustafa, Marco Volino, Jean-Yves Guillemaut, Adrian Hilton 0001 |
3DV | 3 |
| 2017 | A generalised framework for saliency-based point feature detectionabstractHere we present a novel, histogram-based salient point feature detector that may naturally be applied to both images and 3D data. Existing point feature detectors are often modality specific, with 2D and 3D feature detectors typically constructed in separate ways. As such, their applicability in a 2D-3D context is very limited, particularly where the 3D data is obtained by a LiDAR scanner. By contrast, our histogram-based approach is highly generalisable and as such, may be meaningfully applied between 2D and 3D data. Using the generalised approach, we propose salient point detectors for images, and both untextured and textured 3D data. The approach naturally allows for the detection of salient 3D points based jointly on both the geometry and texture of the scene, allowing for broader applicability. The repeatability of the feature detectors is evaluated using a range of datasets including image and LiDAR input from indoor and outdoor scenes. Experimental results demonstrate a significant improvement in terms of 2D-2D and 2D-3D repeatability compared to existing multi-modal feature detectors. David Windridge, Jean-Yves Guillemaut |
Comput. Vis. Image Underst. | 3 |
| 2017 | Colour Helmholtz Stereopsis for Reconstruction of Dynamic Scenes with Arbitrary Unknown ReflectanceabstractHelmholtz Stereopsis is a powerful technique for reconstruction of scenes with arbitrary reflectance properties. However, previous formulations have been limited to static objects due to the requirement to sequentially capture reciprocal image pairs (i.e. two images with the camera and light source positions mutually interchanged). In this paper, we propose colour Helmholtz Stereopsis-a novel framework for Helmholtz Stereopsis based on wavelength multiplexing. To address the new set of challenges introduced by multispectral data acquisition, the proposed Colour Helmholtz Stereopsis pipeline uniquely combines a tailored photometric calibration for multiple camera/light source pairs, a novel procedure for spatio-temporal surface chromaticity calibration and a state-of-the-art Bayesian formulation necessary for accurate reconstruction from a minimal number of reciprocal pairs. In this framework, reflectance is spatially unconstrained both in terms of its chromaticity and the directional component dependent on the illumination incidence and viewing angles. The proposed approach for the first time enables modelling of dynamic scenes with arbitrary unknown and spatially varying reflectance using a practical acquisition set-up consisting of a small number of cameras and light sources. Experimental results demonstrate the accuracy and flexibility of the technique on a variety of static and dynamic scenes with arbitrary unknown BRDF and chromaticity ranging from uniform to arbitrary and spatially varying. Nadejda Roubtsova, Jean-Yves Guillemaut |
Int. J. Comput. Vis. | 2 |
| 2016 | Temporally Coherent 4D Reconstruction of Complex Dynamic ScenesabstractThis paper presents an approach for reconstruction of 4D temporally coherent models of complex dynamic scenes. No prior knowledge is required of scene structure or camera calibration allowing reconstruction from multiple moving cameras. Sparse-to-dense temporal correspondence is integrated with joint multi-view segmentation and reconstruction to obtain a complete 4D representation of static and dynamic objects. Temporal coherence is exploited to overcome visual ambiguities resulting in improved reconstruction of complex scenes. Robust joint segmentation and reconstruction of dynamic objects is achieved by introducing a geodesic star convexity constraint. Comparative evaluation is performed on a variety of unstructured indoor and outdoor dynamic scenes with hand-held cameras and multiple people. This demonstrates reconstruction of complete temporally coherent 4D scene models with improved nonrigid object segmentation and shape reconstruction. Armin Mustafa, Hansung Kim 0001, Jean-Yves Guillemaut, Adrian Hilton 0001 |
CVPR | 3 |
| 2015 | Globally Optimal 2D-3D Registration from Points or Lines without CorrespondencesabstractWe present a novel approach to 2D-3D registration from points or lines without correspondences. While there exist established solutions in the case where correspondences are known, there are many situations where it is not possible to reliably extract such correspondences across modalities, thus requiring the use of a correspondence-free registration algorithm. Existing correspondence-free methods rely on local search strategies and consequently have no guarantee of finding the optimal solution. In contrast, we present the first globally optimal approach to 2D-3D registration without correspondences, achieved by a Branch-and-Bound algorithm. Furthermore, a deterministic annealing procedure is proposed to speed up the nested branch-and-bound algorithm used. The theoretical and practical advantages this brings are demonstrated on a range of synthetic and real data where it is observed that the proposed approach is significantly more robust to high proportions of outliers compared to existing approaches. David Windridge, Jean-Yves Guillemaut |
ICCV | 3 |
| 2015 | General Dynamic Scene Reconstruction from Multiple View VideoabstractThis paper introduces a general approach to dynamic scene reconstruction from multiple moving cameras without prior knowledge or limiting constraints on the scene structure, appearance, or illumination. Existing techniques or dynamic scene reconstruction from multiple wide-baseline camera views primarily focus on accurate reconstruction in controlled environments, where the cameras are fixed and calibrated and background is known. These approaches are not robust for general dynamic scenes captured with sparse moving cameras. Previous approaches for outdoor dynamic scene reconstruction assume prior knowledge of the static background appearance and structure. The primary contributions of this paper are twofold: an automatic method for initial coarse dynamic scene segmentation and reconstruction without prior knowledge of background appearance or structure, and a general robust approach for joint segmentation refinement and dense reconstruction of dynamic scenes from multiple wide-baseline static or moving cameras. Evaluation is performed on a variety of indoor and outdoor scenes with cluttered backgrounds and multiple dynamic non-rigid objects such as people. Comparison with state-of-the-art approaches demonstrates improved accuracy in both multiple view segmentation and dense reconstruction. The proposed approach also eliminates the requirement for prior knowledge of scene structure and appearance. Armin Mustafa, Hansung Kim 0001, Jean-Yves Guillemaut, Adrian Hilton 0001 |
ICCV | 3 |
| 2015 | A generalisable framework for saliency-based line segment detectionabstractHere we present a novel, information-theoretic salient line segment detector. Existing line detectors typically only use the image gradient to search for potential lines. Consequently, many lines are found, particularly in repetitive scenes. In contrast, our approach detects lines that define regions of significant divergence between pixel intensity or colour statistics. This results in a novel detector that naturally avoids the repetitive parts of a scene while detecting the strong, discriminative lines present. We furthermore use our approach as a saliency filter on existing line detectors to more efficiently detect salient line segments. The approach is highly generalisable, depending only on image statistics rather than image gradient; and this is demonstrated by an extension to depth imagery. Our work is evaluated against a number of other line detectors and a quantitative evaluation demonstrates a significant improvement over existing line detectors for a range of image transformations. David Windridge, Jean-Yves Guillemaut |
Pattern Recognit. | 3 |
| 2014 | Structured Representation of Non-Rigid Surfaces from Single View 3D Point TracksabstractThis work considers the problem of structured representation of dynamic surfaces from incomplete 3D point tracks from a single viewpoint. The surface is segmented into a set of connected regions each of which can be represented by a fixed intrinsic shape and a parametrised rigid/non-rigid motion trajectory. Neither the model parameters nor the point-to-model assignments are known upfront. Motion and geometric shape parameters are estimated in alternation with a graph-cuts based point-to-model assignment. This modelling process facilitates in-filling of missing data as well as de-noising of measurements by temporal integration while adding meaningful structure to the geometry and reducing storage cost by an order of magnitude. Experiments are presented for real and synthetic sequences to validate the approach and show how a single tuning parameter can be used to trade modelling error with extrapolation level and storage cost. Charles Malleson, Martin Klaudiny, Jean-Yves Guillemaut, Adrian Hilton 0001 |
3DV | 3 |
| 2014 | Colour Helmholtz Stereopsis for Reconstruction of Complex Dynamic ScenesabstractHelmholtz Stereopsis (HS) is a powerful technique for reconstruction of scenes with arbitrary reflectance properties. However, previous formulations have been limited to static objects due to the requirement to sequentially capture reciprocal image pairs (i.e. Two images with the camera and light source positions mutually interchanged). In this paper, we propose colour HS - a novel variant of the technique based on wavelength multiplexing. To address the new set of challenges introduced by multispectral data acquisition, the proposed novel pipeline for colour HS uniquely combines a tailored photometric calibration for multiple camera/light source pairs, a novel procedure for surface chromaticity calibration and the state-of-the-art Bayesian HS suitable for reconstruction from a minimal number of reciprocal pairs. Experimental results including quantitative and qualitative evaluation demonstrate that the method is suitable for flexible (single-shot) reconstruction of static scenes and reconstruction of dynamic scenes with complex surface reflectance properties. Nadejda Roubtsova, Jean-Yves Guillemaut |
3DV | 2 |
| 2014 | Intrinsic Textures for Relightable Free-Viewpoint Video
James Imber, Jean-Yves Guillemaut, Adrian Hilton 0001 |
ECCV (2) | 2 |
| 2013 | Interactive Animation of 4D Performance CaptureabstractA 4D parametric motion graph representation is presented for interactive animation from actor performance capture in a multiple camera studio. The representation is based on a 4D model database of temporally aligned mesh sequence reconstructions for multiple motions. High-level movement controls such as speed and direction are achieved by blending multiple mesh sequences of related motions. A real-time mesh sequence blending approach is introduced, which combines the realistic deformation of previous nonlinear solutions with efficient online computation. Transitions between different parametric motion spaces are evaluated in real time based on surface shape and motion similarity. Four-dimensional parametric motion graphs allow real-time interactive character animation while preserving the natural dynamics of the captured performance. Dan Casas, Margara Tejera, Jean-Yves Guillemaut, Adrian Hilton 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2012 | 4D parametric motion graphs for interactive animationabstractA 4D parametric motion graph representation is presented for interactive animation from actor performance capture in a multiple camera studio. The representation is based on a 4D model database of temporally aligned mesh sequence reconstructions for multiple motions. High-level movement controls such as speed and direction are achieved by blending multiple mesh sequences of related motions. A real-time mesh sequence blending approach is introduced which combines the realistic deformation of previous non-linear solutions with efficient online computation. Transitions between different parametric motion spaces are evaluated in real-time based on surface shape and motion similarity. 4D parametric motion graphs allow real-time interactive character animation while preserving the natural dynamics of the captured performance. Dan Casas, Margara Tejera, Jean-Yves Guillemaut, Adrian Hilton 0001 |
I3D | 3 |
| 2012 | Parametric animation of performance-captured mesh sequencesabstractABSTRACT In this paper, we introduce an approach to high‐level parameterisation of captured mesh sequences of actor performance for real‐time interactive animation control. High‐level parametric control is achieved by non‐linear blending between multiple mesh sequences exhibiting variation in a particular movement. For example, walking speed is parameterised by blending fast and slow walk sequences. A hybrid non‐linear mesh sequence blending approach is introduced to approximate the natural deformation of non‐linear interpolation techniques whilst maintaining the real‐time performance of linear mesh blending. Quantitative results show that the hybrid approach gives an accurate real‐time approximation of offline non‐linear deformation. An evaluation of the approach shows good performance not only for entire meshes but also with specific mesh areas. Results are presented for single and multi‐dimensional parametric control of walking (speed/direction), jumping (height/distance) and reaching (height) from captured mesh sequences. This approach allows continuous real‐time control of high‐level parameters such as speed and direction whilst maintaining the natural surface dynamics of captured movement. Copyright © 2012 John Wiley & Sons, Ltd. Dan Casas, Margara Tejera, Jean-Yves Guillemaut, Adrian Hilton 0001 |
Comput. Animat. Virtual Worlds | 3 |
| 2012 | Outdoor Dynamic 3-D Scene ReconstructionabstractExisting systems for 3-D reconstruction from multiple view video use controlled indoor environments with uniform illumination and backgrounds to allow accurate segmentation of dynamic foreground objects. In this paper, we present a portable system for 3-D reconstruction of dynamic outdoor scenes that require relatively large capture volumes with complex backgrounds and nonuniform illumination. This is motivated by the demand for 3-D reconstruction of natural outdoor scenes to support film and broadcast production. Limitations of existing multiple view 3-D reconstruction techniques for use in outdoor scenes are identified. Outdoor 3-D scene reconstruction is performed in three stages: 1) 3-D background scene modeling using spherical stereo image capture; 2) multiple view segmentation of dynamic foreground objects by simultaneous video matting across multiple views; and 3) robust 3-D foreground reconstruction and multiple view segmentation refinement in the presence of segmentation and calibration errors. Evaluation is performed on several outdoor productions with complex dynamic scenes including people and animals. Results demonstrate that the proposed approach overcomes limitations of previous indoor multiple view reconstruction approaches enabling high-quality free-viewpoint rendering and 3-D reference models for production. Hansung Kim 0001, Jean-Yves Guillemaut, Takeshi Takai, Muhammad Sarim, Adrian Hilton 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Temporal trimap propagation for video matting using inferential statisticsabstractThis paper introduces a statistical inference framework to temporally propagate trimap labels from sparsely defined key frames to estimate trimaps for the entire video sequence. Trimap is a fundamental requirement for digital image and video matting approaches. Statistical inference is coupled with Bayesian statistics to allow robust trimap labelling in the presence of shadows, illumination variation and overlap between the foreground and background appearance. Results demonstrate that trimaps are sufficiently accurate to allow high quality video matting using existing natural image matting algorithms. Quantitative evaluation against ground-truth demonstrates that the approach achieves accurate matte estimation with less amount of user interaction compared to the state-of-the-art techniques. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut |
ICIP | 3 |
| 2011 | Parametric Control of Captured Mesh Sequences for Real-Time Animation
Dan Casas, Margara Tejera, Jean-Yves Guillemaut, Adrian Hilton 0001 |
MIG | 3 |
| 2011 | Joint Multi-Layer Segmentation and Reconstruction for Free-Viewpoint Video Applications
Jean-Yves Guillemaut, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 1 |
| 2010 | Moving Camera Registration for Multiple Camera Setups in Dynamic ScenesabstractThis paper describes a method to register a moving (principal) camera, given a set of fully calibrated static cameras (witnesses) viewing a dynamic scene, a common scenario in broadcasting and film production. Our ultimate aim is to equip the existing free-viewpoint video algorithms with the ability to exploit any available moving cameras in generic dynamic scenes, and to facilitate 3D content production by augmented reality and stereoscopic rendering. Evren Imre, Jean-Yves Guillemaut, Adrian Hilton 0001 |
BMVC | 2 |
| 2010 | Stereoscopic content production of complex dynamic scenes using a wide-baseline monoscopic camera set-upabstractConventional stereoscopic video content production requires use of dedicated stereo camera rigs which is both costly and lacking video editing flexibility. In this paper, we propose a novel approach which only requires a small number of standard cameras sparsely located around a scene to automatically convert the monocular inputs into stereoscopic streams. The approach combines a probabilistic spatio-temporal segmentation framework with a state-of-the-art multi-view graph-cut reconstruction algorithm, thus providing full control of the stereoscopic settings at render time. Results with studio sequences of complex human motion demonstrate the suitability of the method for high quality stereoscopic content generation with minimum user interaction. Jean-Yves Guillemaut, Muhammad Sarim, Adrian Hilton 0001 |
ICIP | 1 |
| 2010 | Natural image matting for multiple wide-baseline viewsabstractIn this paper we present a novel approach to estimate the alpha mattes of a foreground object captured by a wide-baseline circular camera rig provided a single key frame trimap. Bayesian inference coupled with camera calibration information are used to propagate high confidence trimaps labels across the views. Recent techniques have been developed to estimate an alpha matte of an image using multiple views but they are limited to narrow baseline views with low foreground variation. The proposed wide-baseline trimap propagation is robust to inter-view foreground appearance changes, shadows and similarity in foreground/background appearance for cameras with opposing views enabling high quality alpha matte extraction using any state-of-the-art image matting algorithm. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut, Takeshi Takai, Hansung Kim 0001 |
ICIP | 3 |
| 2010 | Multi-label propagation for coherent video segmentation and artistic stylizationabstractWe present a new algorithm for segmenting video frames into temporally stable colored regions, applying our technique to create artistic stylizations (e.g. cartoons and paintings) from real video sequences. Our approach is based on a multi-label graph cut applied to successive frames, in which the color data term and label priors are incrementally updated and propagated over time. We demonstrate coherent segmentation and stylization over a variety of home videos. Tinghuai Wang, Jean-Yves Guillemaut, John P. Collomosse |
ICIP | 2 |
| 2009 | Non-parametric Patch based Video MattingabstractIn computer vision, matting is the process of accurate foreground estimation in images and videos. In this paper we presents a novel patch based approach to video matting relying on non-parametric statistics to represent image variations in appearance. This overcomes the limitation of parametric algorithms which only rely on strong colour correlation between the nearby pixels. Initially we construct a clean background by utilising the foreground object’s movement across the background. For a given frame, a trimap is constructed using the background and the last frame’s trimap. A patch-based approach is used to estimate the foreground colour for every unknown pixel and finally the alpha matte is extracted. Quantitative evaluation shows that the technique performs better, in terms of the accuracy and the required user interaction, than the current state-of-the-art parametric approaches. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut |
BMVC | 3 |
| 2009 | Robust graph-cut scene segmentation and reconstruction for free-viewpoint video of complex dynamic scenesabstractCurrent state-of-the-art image-based scene reconstruction techniques are capable of generating high-fidelity 3D models when used under controlled capture conditions. However, they are often inadequate when used in more challenging outdoor environments with moving cameras. In this case, algorithms must be able to cope with relatively large calibration and segmentation errors as well as input images separated by a wide-baseline and possibly captured at different resolutions. In this paper, we propose a technique which, under these challenging conditions, is able to efficiently compute a high-quality scene representation via graph-cut optimisation of an energy function combining multiple image cues with strong priors. Robustness is achieved by jointly optimising scene segmentation and multiple view reconstruction in a view-dependent manner with respect to each input camera. Joint optimisation prevents propagation of errors from segmentation to reconstruction as is often the case with sequential approaches. View-dependent processing increases tolerance to errors in on-the-fly calibration compared to global approaches. We evaluate our technique in the case of challenging outdoor sports scenes captured with manually operated broadcast cameras and demonstrate its suitability for high-quality free-viewpoint video. Jean-Yves Guillemaut, Joe Kilner, Adrian Hilton 0001 |
ICCV | 1 |
| 2009 | Non-parametric natural image mattingabstractNatural image matting is an extremely challenging image processing problem due to its ill-posed nature. It often requires skilled user interaction to aid definition of foreground and background regions. Current algorithms use these predefined regions to build local foreground and background colour models. In this paper we propose a novel approach which uses non-parametric statistics to model image appearance variations. This technique overcomes the limitations of previous parametric approaches which are purely colour-based and thereby unable to model natural image structure. The proposed technique consists of three successive stages: (i) background colour estimation, (ii) foreground colour estimation, (iii) alpha estimation. Colour estimation uses patch-based matching techniques to efficiently recover the optimum colour by comparison against patches from the known regions. Quantitative evaluation against ground truth demonstrates that the technique produces better results and successfully recovers fine details such as hair where many other algorithms fail. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut, Hansung Kim 0001 |
ICIP | 3 |
| 2009 | Objective quality assessment in free-viewpoint video production
Joe Kilner, Jonathan Starck, Jean-Yves Guillemaut, Adrian Hilton 0001 |
Signal Process. Image Commun. | 3 |
| 2008 | The normalised image of the absolute conic and its application for zooming camera calibration
Jean-Yves Guillemaut, John Illingworth |
Pattern Recognit. | 1 |
| 2006 | General Pose Face Recognition Using Frontal Face Model
Jean-Yves Guillemaut, Josef Kittler, Mohammad Sadeghi 0001, William J. Christmas |
CIARP | 1 |
| 2005 | Using Points at Infinity for Parameter Decoupling in Camera CalibrationabstractThe majority of camera calibration methods, including the Gold Standard algorithm, use point-based information and simultaneously estimate all calibration parameters. In contrast, we propose a novel calibration method that exploits line orientation information and decouples the problem into two simpler stages. We formulate the problem as minimization of the lateral displacement between single projected image lines and their vanishing points. Unlike previous vanishing point methods, parallel line pairs are not required. Additionally, the invariance properties of vanishing points mean that multiple images related by pure translation can be used to increase the calibration data set size without increasing the number of estimated parameters. We compare this method with vanishing point methods and the Gold Standard algorithm and demonstrate that it has comparable performance. Jean-Yves Guillemaut, Alberto S. Aguado, John Illingworth |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Using Points at Infinity for Parameter Decoupling in Camera CalibrationabstractWe consider the problem of decoupling translation and rotation for a collec-tion of 3D data related to 2D images by a projection. The main contribution is to show that equations describing image formation can be decoupled to form a component independent of translation. The decomposition is based on the invariance properties of the projection of points at infinity; if image formation is expressed in terms of points at infinity, general motions can be reduced to pure rotations. Contrary to other methods based on vanishing points, our approach does not require parallel directions to be present in the scene. We use the invariance property to simplify camera calibration equa-tions. We consider three cases for a known calibration object: full calibration, pose estimation and internal calibration by pure translation. Experiments on synthetic and real data show that the decomposition can obtain similar re-sults to a full parameter search. For pure translation, the decomposition can be effectively used to obtain more accurate parameters, without increasing the dimensionality of the problem. 1 Jean-Yves Guillemaut, Alberto S. Aguado, John Illingworth |
BMVC | 1 |