EDBT 2026 Demo / reviewers in the wild / expert
Adrian Hilton 0001
dblp:63/34
· DBLP profile ↗
161ranked-venue papers
16as first author
26since 2021 · last 2026
0000-0003-4223-238XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 121 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 88 · 11 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reciprocal Teaching: Dynamic Multi-Model Teacher-Student Learning for Multiple Noisy AnnotationsabstractAs datasets grow, expert-based annotation becomes impractical, making crowdsourcing a scalable alternative. In crowdsourcing, samples are typically annotated by multiple workers and aggregated via majority voting, which ignores annotator-specific biases and introduces noisy labels that impair downstream models. Traditional multi-rater methods attempt to model annotator biases but often overfit with many classes or few, noisy annotators. Learning with Noisy Labels (LNL) methods offer more robust strategies for handling noisy labels, but their assumption of a single noisy label per sample makes extending them to multi-annotator settings non-trivial. To bridge this gap, we propose the Reciprocal Teacher-student Learning from Multi-rater Noisy Annotation (RETINA), which trains annotator-specific models and employs a dynamic teacher–student process to separate clean from noisy samples. Progress in multi-rater learning has also been limited by benchmarks with few classes, fixed noise rates, and no control over annotators. To address this, we introduce the Synthetic MRL (SynMRL) benchmark that contains many classes and controllable noise and annotator settings for systematic evaluation. Experiments on synthetic and real-world data show that RETINA outperforms existing multi-rater methods, particularly in high-noise, low-annotator, many-class settings. Wenjie Ai, Cuong Nguyen 0006, Adrian Hilton 0001, Gustavo Carneiro 0001 |
WACV | 3 |
| 2026 | An Effective-Efficient Approach for Dense Multi-Label Action DetectionabstractAbstract Unlike the sparse label action detection task, where a single action occurs in each timestamp of a video, in a dense multi-label scenario, actions can overlap temporally. To address this challenging task, it is necessary to simultaneously learn (i) co-occurrence action relationships and (ii) temporal dependencies. Current methods model co-occurrence action relationships by explicitly embedding class relations into the transformer network architecture. However, these approaches are not computationally efficient, as the network needs to compute all possible pair action class relations. In this paper, we overcome this by introducing a novel framework trained through a novel learning paradigm that allows the network to benefit from explicitly modelling temporal co-occurrence action dependencies during training without imposing their computational overhead during inference. Furthermore, to model temporal information, recent approaches extract multi-scale temporal features through hierarchical transformer-based networks. However, the self-attention mechanism in transformers inherently loses temporal positional information. We argue that combining this with multiple sub-sampling processes in hierarchical designs can lead to further loss of positional information. Preserving this information is essential for accurate action detection. In this paper, we address this issue by proposing a novel transformer network that (a) employs a non-hierarchical structure when modelling different ranges of temporal dependencies and (b) embeds relative positional encoding in its transformer layers. We evaluate the performance of our proposed approach on two challenging dense multi-label benchmark datasets and show that our method improves the current state-of-the-art results by 1.1% and 0.6% per-frame mAP on the Charades and MultiTHUMOS datasets, respectively, achieving new state-of-the-art per-frame mAP results at 26.5% and 44.6%, respectively. We also performed extensive ablation studies to examine the impact of the different components of our proposed approach. Our code will be released upon paper publication Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | Improving Gaussian Splatting with Localized Points ManagementabstractPoint management is critical for optimizing 3D Gaussian Splatting models, as point initiation (e.g., via structure from motion) is often distributionally inappropriate. Typically, Adaptive Density Control (ADC) algorithm is adopted, leveraging view-averaged gradient magnitude thresholding for point densification, opacity thresholding for pruning, and regular all-points opacity reset. We reveal that this strategy is limited in tackling intricate/special image regions (e.g., transparent) due to inability of identifying all 3D zones requiring point densification, and lacking an appropriate mechanism to handle ill-conditioned points with negative impacts (e.g., occlusion due to false high opacity). To address these limitations, we propose a Localized Point Management (LPM) strategy, capable of identifying those error-contributing zones in greatest need for both point addition and geometry calibration. Zone identification is achieved by leveraging the underlying multiview geometry constraints, subject to image rendering errors. We apply point densification in the identified zones and then reset the opacity of the points in front of these regions, creating a new opportunity to correct poorly conditioned points. Serving as a versatile plugin, LPM can be seamlessly integrated into existing static 3D and dynamic 4D Gaussian Splatting models with minimal additional cost. Experimental evaluations validate the efficacy of our LPM in boosting a variety of existing 3D/4D models both quantitatively and qualitatively. Notably, LPM improves both static 3DGS and dynamic SpaceTimeGS to achieve state-of-the-art rendering quality while retaining real-time speeds, excelling on challenging datasets such as Tanks & Temples and the Neural 3D Video dataset. Haosen Yang 0003, Chenhao Zhang 0006, Wenqing Wang 0002, Marco Volino, Adrian Hilton 0001, Li Zhang 0040, Xiatian Zhu |
CVPR | 5 |
| 2025 | NarrativeBridge: Enhancing Video Captioning with Causal-Temporal NarrativeabstractExisting video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by characters or agents. This lack of narrative restricts models’ ability to generate text descriptions that capture the causal and temporal dynamics inherent in video content. To address this gap, we propose NarrativeBridge, an approach comprising of: (1) a novel Causal-Temporal Narrative (CTN) captions benchmark generated using a large language model and few-shot prompting, explicitly encoding cause-effect temporal relationships in video descriptions; and (2) a Cause-Effect Network (CEN) with separate encoders for capturing cause and effect dynamics, enabling effective learning and generation of captions with causal-temporal narrative. Extensive experiments demonstrate that CEN significantly outperforms state-of-the-art models in articulating the causal and temporal aspects of video content: 17.88 and 17.44 CIDEr on the MSVD-CTN and MSRVTT-CTN datasets, respectively. Cross-dataset evaluations further showcase CEN’s strong generalization capabilities. The proposed framework understands and generates nuanced text descriptions with intricate causal-temporal narrative structures present in videos, addressing a critical limitation in video captioning. For project details, visit https://narrativebridge.github.io/. Asmar Nadeem, Faegheh Sardari, Robert Dawes, Syed Sameed Husain, Adrian Hilton 0001, Armin Mustafa |
ICLR | 5 |
| 2025 | Risk Estimation of Knee Osteoarthritis Progression via Predictive Multi-task Modelling from Efficient Diffusion Model Using X-Ray Images
Adrian Hilton 0001, Gustavo Carneiro 0001 |
MICCAI (14) | 2 |
| 2025 | MVL-Net: Pairwise Learning for Multi-View Multiple People LabellingabstractIn the multi-view domain, it is challenging to correctly label multiple people across viewpoints because of occlusions, visual ambiguities, appearance variation, etc. Deep learning, although having witnessed remarkable success in computer vision tasks, still remains underexplored for the multi-view labelling task, due to the lack of labelled multi-view datasets. In this paper, we propose a novel end-to-end deep neural network named Multi-View Labelling network (MVL-net) that addresses this issue. To overcome the dataset shortage, a large-scale multi-view dataset is generated by combining 3D human models and panoramic backgrounds, along with human poses and realistic rendering. In the proposed MVL-net, we first incorporate Transformer blocks to capture the non-local information for multi-view feature extraction. A matching net is then introduced to achieve multiple people labelling, by predicting matching confidence scores for pairwise instances from two views, thus addressing the problem of the unknown number of people when labelling across views. An additional geometry feature obtained from the epipolar geometry is integrated to leverage multi-view cues during training. To the best of our knowledge, the MVL-net is the first work using deep learning to train a multi-view labelling network. Comprehensive experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed method, which outperforms the existing state-of-the-art approaches. Yue Zhang 0082, Akin Caliskan, Mai Xu, Adrian Hilton 0001, Jean-Yves Guillemaut |
IEEE Trans. Multim. | 4 |
| 2024 | ANIM: Accurate Neural Implicit Model for Human Reconstruction from a Single RGB-D ImageabstractRecent progress in human shape learning, shows that neural implicit models are effective in generating 3D hu-man surfaces from limited number of views, and even from a single RGB image. However, existing monocular approaches still struggle to recover fine geometric details such as face, hands or cloth wrinkles. They are also easily prone to depth ambiguities that result in distorted geome-tries along the camera optical axis. In this paper, we ex-plore the benefits of incorporating depth observations in the reconstruction process by introducing ANIM, a novel method that reconstructs arbitrary 3D human shapes from single-view RGB-D images with an unprecedented level of accuracy. Our model learns geometric details from both multi-resolution pixel-aligned and voxel-aligned features to leverage depth information and enable spatial relation-ships, mitigating depth ambiguities. We further enhance the quality of the reconstructed shape by introducing a depth-supervision strategy, which improves the accuracy of the signed distance field estimation of points that lie on the re-constructed surface. Experiments demonstrate that ANIM outperforms state-of-the-art works that use RGB, surface normals, point cloud or RGB-D data as input. In addition, we introduce ANIM-Real, a new multi-modal dataset comprising highquality scans paired with consumer-grade RGB-D camera, and our protocol to fine-tune ANIM, enabling highquality reconstruction from real-world human capture. https://marcopesavento.github.io/Anim/ Marco Pesavento, Yuanlu Xu, Nikolaos Sarafianos, Robert Maier 0001, Chun-Han Yao, Marco Volino, Edmond Boyer, Adrian Hilton 0001, Tony Tung |
CVPR | 9 |
| 2024 | COSMU: Complete 3D Human Shape from Monocular Unconstrained Images
Marco Pesavento, Marco Volino, Adrian Hilton 0001 |
ECCV (39) | 3 |
| 2024 | CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing
Faegheh Sardari, Armin Mustafa, Philip J. B. Jackson, Adrian Hilton 0001 |
ECCV (11) | 4 |
| 2024 | ForecasterFlexOBM: A Multi-View Audio-Visual Dataset for Flexible Object-Based Media ProductionabstractLeveraging machine learning techniques, in the context of object-based media production, could enable provision of personalized media experiences to diverse audiences. To fine-tune and evaluate techniques for personalization applications, as well as more broadly, datasets which bridge the gap between research and production are needed. We introduce and release such a dataset, themed around a UK weather forecast and shot against a blue-screen background, of three professional actors/presenters – one male and one female (English) and one female (British Sign Language). Scenes include both production and research-oriented examples, with a range of dialogues and actions. Capture techniques consisted of a synchronized 4K resolution 16-camera array, production-typical microphones plus professional audio mix, a 16-channel microphone array with collocated Grasshopper3 camera, and a photogrammetry array. We demonstrate applications relevant to virtual production and creation of personalized media including neural radiance fields, shadow casting, action/event detection, speaker source tracking and video captioning. Davide Berghi, Craig Cieciura, Farshad Einabadi, Maxine Glancy, Oliver C. Camilleri, Philip Foster, Asmar Nadeem, Faegheh Sardari, Jinzheng Zhao, Marco Volino, Armin Mustafa, Philip J. B. Jackson, Adrian Hilton 0001 |
ICME | 13 |
| 2024 | Learning Self-Shadowing for Clothed Human Bodies
Farshad Einabadi, Jean-Yves Guillemaut, Adrian Hilton 0001 |
EGSR (ST) | 3 |
| 2024 | CAD - Contextual Multi-modal Alignment for Dynamic AVQAabstractIn the context of Audio Visual Question Answering (AVQA) tasks, the audio and visual modalities could be learnt on three levels: 1) Spatial, 2) Temporal, and 3) Semantic. Existing AVQA methods suffer from two major shortcomings; the audio-visual (AV) information passing through the network isn’t aligned on Spatial and Temporal levels; and, inter-modal (audio and visual) Semantic information is often not balanced within a context; this results in poor performance. In this paper, we propose a novel end-to-end Contextual Multi-modal Alignment (CAD) network that addresses the challenges in AVQA methods by i) introducing a parameter-free stochastic Contextual block that ensures robust audio and visual alignment on the Spatial level; ii) proposing a pre-training technique for dynamic audio and visual alignment on Temporal level in a self-supervised setting, and iii) introducing a cross-attention mechanism to balance audio and visual information on Semantic level. The proposed novel CAD network improves the overall performance over the state-of-the-art methods on average by 9.4% on the MUSIC-AVQA dataset. We also demonstrate that our proposed contributions to AVQA can be added to the existing methods to improve their performance without additional complexity requirements. Asmar Nadeem, Adrian Hilton 0001, Robert Dawes, Graham A. Thomas, Armin Mustafa |
WACV | 2 |
| 2024 | SyDog-Video: A Synthetic Dog Video Dataset for Temporal Pose EstimationabstractAbstract We aim to estimate the pose of dogs from videos using a temporal deep learning model as this can result in more accurate pose predictions when temporary occlusions or substantial movements occur. Generally, deep learning models require a lot of data to perform well. To our knowledge, public pose datasets containing videos of dogs are non existent. To solve this problem, and avoid manually labelling videos as it can take a lot of time, we generated a synthetic dataset containing 500 videos of dogs performing different actions using Unity3D. Diversity is achieved by randomising parameters such as lighting, backgrounds, camera parameters and the dog’s appearance and pose. We evaluate the quality of our synthetic dataset by assessing the model’s capacity to generalise to real data. Usually, networks trained on synthetic data perform poorly when evaluated on real data, this is due to the domain gap. As there was still a domain gap after improving the quality of the synthetic dataset and inserting diversity, we bridged the domain gap by applying 2 different methods: fine-tuning and using a mixed dataset to train the network. Additionally, we compare the model pre-trained on synthetic data with models pre-trained on a real-world animal pose datasets. We demonstrate that using the synthetic dataset is beneficial for training models with (small) real-world datasets. Furthermore, we show that pre-training the model with the synthetic dataset is the go to choice rather than pre-training on real-world datasets for solving the pose estimation task from videos of dogs. Moira Shooter, Charles Malleson, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 3 |
| 2023 | Learning Projective Shadow Textures for Neural Rendering of Human Cast Shadows from Silhouettes
Farshad Einabadi, Jean-Yves Guillemaut, Adrian Hilton 0001 |
EGSR (ST) | 3 |
| 2022 | Super-Resolution 3D Human Shape from a Single Low-Resolution Image
Marco Pesavento, Marco Volino, Adrian Hilton 0001 |
ECCV (2) | 3 |
| 2022 | Finite Aperture StereoabstractAbstract Multi-view stereo remains a popular choice when recovering 3D geometry, despite performance varying dramatically according to the scene content. Moreover, typical pinhole camera assumptions fail in the presence of shallow depth of field inherent to macro-scale scenes; limiting application to larger scenes with diffuse reflectance. However, the presence of defocus blur can itself be considered a useful reconstruction cue, particularly in the presence of view-dependent materials. With this in mind, we explore the complimentary nature of stereo and defocus cues in the context of multi-view 3D reconstruction; and propose a complete pipeline for scene modelling from a finite aperature camera that encompasses image formation, camera calibration and reconstruction stages. As part of our evaluation, an ablation study reveals how each cue contributes to the higher performance observed over a range of complex materials and geometries. Though of lesser concern with large apertures, the effects of image noise are also considered. By introducing pre-trained deep feature extraction into our cost function, we show a step improvement over per-pixel comparisons; as well as verify the cross-domain applicability of networks using largely in-focus training data applied to defocused images. Finally, we compare to a number of modern multi-view stereo methods, and demonstrate how the use of both cues leads to a significant increase in performance across several synthetic and real datasets. Matthew Bailey 0003, Adrian Hilton 0001, Jean-Yves Guillemaut |
Int. J. Comput. Vis. | 2 |
| 2022 | 4D Temporally Coherent Multi-Person Semantic Reconstruction and SegmentationabstractAbstract We introduce the first approach to solve the challenging problem of automatic 4D visual scene understanding for complex dynamic scenes with multiple interacting people from multi-view video. Our approach simultaneously estimates a detailed model that includes a per-pixel semantically and temporally coherent reconstruction, together with instance-level segmentation exploiting photo-consistency, semantic and motion information. We further leverage recent advances in 3D pose estimation to constrain the joint semantic instance segmentation and 4D temporally coherent reconstruction. This enables per person semantic instance segmentation of multiple interacting people in complex dynamic scenes. Extensive evaluation of the joint visual scene understanding framework against state-of-the-art methods on challenging indoor and outdoor sequences demonstrates a significant ( $$\approx 40\%$$ ≈ 40 % ) improvement in semantic segmentation, reconstruction and scene flow accuracy. In addition to the evaluation on several indoor and outdoor scenes, the proposed joint 4D scene understanding framework is applied to challenging outdoor sports scenes in the wild captured with manually operated wide-baseline broadcast cameras. Armin Mustafa, Chris Russell 0001, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 3 |
| 2021 | Multi-Person Implicit Reconstruction From a Single ImageabstractWe present a new end-to-end learning framework to obtain detailed and spatially coherent reconstructions of multiple people from a single image. Existing multi-person methods suffer from two main drawbacks: they are often model-based and therefore cannot capture accurate 3D models of people with loose clothing and hair; or they require manual intervention to resolve occlusions or interactions. Our method addresses both limitations by introducing the first end-to-end learning approach to perform model-free implicit reconstruction for realistic 3D capture of multiple clothed people in arbitrary poses (with occlusions) from a single image. Our network simultaneously estimates the 3D geometry of each person and their 6DOF spatial locations, to obtain a coherent multi-human reconstruction. In addition, we introduce a new synthetic dataset that depicts images with a varying number of inter-occluded humans and a variety of clothing and hair styles. We demonstrate robust, high-resolution reconstructions on images of multiple humans with complex occlusions, loose clothing and a large variety of poses and scenes. Our quantitative evaluation on both synthetic and real world datasets demonstrates state-of-the-art performance with significant improvements in the accuracy and completeness of the reconstructions over competing approaches. Armin Mustafa, Akin Caliskan, Lourdes Agapito, Adrian Hilton 0001 |
CVPR | 4 |
| 2021 | Attention-based Multi-Reference Learning for Image Super-ResolutionabstractThis paper proposes a novel Attention-based Multi-Reference Super-resolution network (AMRSR) that, given a low-resolution image, learns to adaptively transfer the most similar texture from multiple reference images to the super-resolution output whilst maintaining spatial coherence. The use of multiple reference images together with attention-based sampling is demonstrated to achieve significantly improved performance over state-of-the-art reference super-resolution approaches on multiple benchmark datasets. Reference super-resolution approaches have recently been proposed to overcome the ill-posed problem of image super-resolution by providing additional information from a high-resolution reference image. Multi-reference super-resolution extends this approach by providing a more diverse pool of image features to overcome the inherent information deficit whilst maintaining memory efficiency. A novel hierarchical attention-based sampling approach is introduced to learn the similarity between low-resolution image features and multiple reference images based on a perceptual loss. Ablation demonstrates the contribution of both multi-reference and hierarchical attention-based sampling to overall performance. Perceptual and quantitative ground-truth evaluation demonstrates significant improvement in performance even when the reference images deviate significantly from the target image. The project website can be found at https://marcopesavento.github.io/AMRSR/ Marco Pesavento, Marco Volino, Adrian Hilton 0001 |
ICCV | 3 |
| 2021 | A Novel Multi-View Labelling Network Based on Pairwise LearningabstractCorrect labelling of multiple people from different viewpoints in complex scenes is a challenging task due to occlusions, visual ambiguities, as well as variations in appearance and illumination. In recent years, deep learning approaches have proved very successful at improving the performance of a wide range of recognition and labelling tasks such as person re-identification and video tracking. However, to date, applications to multi-view tasks have proved more challenging due to the lack of suitably labelled multi-view datasets, which are difficult to collect and annotate. The contributions of this paper are two-fold. First, a synthetic dataset is generated by combining 3D human models and panoramas along with human poses and appearance detail rendering to overcome the shortage of real dataset for multi-view labelling. Second, a novel framework named Multi-View Labelling network (MVL-net) is introduced to leverage the new dataset and unify the multi-view multiple people detection, segmentation and labelling tasks in complex scenes. To the best of our knowledge, this is the first work using deep learning to train a multi-view labelling network. Experiments conducted on both synthetic and real datasets demonstrate that the proposed method outperforms the existing state-of-the-art approaches. Yue Zhang 0082, Akin Caliskan, Adrian Hilton 0001, Jean-Yves Guillemaut |
ICIP | 3 |
| 2021 | Visually Supervised Speaker Detection and Localization via Microphone ArrayabstractActive speaker detection (ASD) is a multi-modal task that aims to identify who, if anyone, is speaking from a set of candidates. Current audio-visual approaches for ASD typically rely on visually pre-extracted face tracks (sequences of consecutive face crops) and the respective monaural audio. However, their recall rate is often low as only the visible faces are included in the set of candidates. Monaural audio may successfully detect the presence of speech activity but fails in localizing the speaker due to the lack of spatial cues. Our solution extends the audio front-end using a microphone array. We train an audio convolutional neural network (CNN) in combination with beamforming techniques to regress the speaker’s horizontal position directly in the video frames. We propose to generate weak labels using a pre-trained active speaker detector on pre-extracted face tracks. Our pipeline embraces the "student-teacher" paradigm, where a trained "teacher" network is used to produce pseudo-labels visually. The "student" network is an audio network trained to generate the same results. At inference, the student network can independently localize the speaker in the visual frames directly from the audio input. Experimental results on newly collected data prove that our approach significantly outperforms a variety of other baselines as well as the teacher network itself. It results in an excellent speech activity detector too. Davide Berghi, Adrian Hilton 0001, Philip J. B. Jackson |
MMSP | 2 |
| 2021 | Temporally Consistent 3D Human Pose Estimation Using Dual 360° CamerasabstractThis paper presents a 3D human pose estimation system that uses a stereo pair of 360° sensors to capture the complete scene from a single location. The approach combines the advantages of omnidirectional capture, the accuracy of multiple view 3D pose estimation and the portability of monocular acquisition. Joint monocular belief maps for joint locations are estimated from 360° images and are used to fit a 3D skeleton to each frame. Temporal data association and smoothing is performed to produce accurate 3D pose estimates throughout the sequence. We evaluate our system on the Panoptic Studio dataset, as well as real 360° video for tracking multiple people, demonstrating an average Mean Per Joint Position Error of 12.47cm with 30cm baseline cameras. We also demonstrate improved capabilities over perspective and 360° multi-view systems when presented with limited camera views of the subject. Matthew Shere, Hansung Kim 0001, Adrian Hilton 0001 |
WACV | 3 |
| 2021 | Deep Neural Models for Illumination Estimation and Relighting: A SurveyabstractAbstract Scene relighting and estimating illumination of a real scene for insertion of virtual objects in a mixed‐reality scenario are well‐studied challenges in the computer vision and graphics fields. Classical inverse rendering approaches aim to decompose a scene into its orthogonal constituting elements, namely scene geometry, illumination and surface materials, which can later be used for augmented reality or to render new images under novel lighting or viewpoints. Recently, the application of deep neural computing to illumination estimation, relighting and inverse rendering has shown promising results. This contribution aims to bring together in a coherent manner current advances in this conjunction. We examine in detail the attributes of the proposed approaches, presented in three categories: scene illumination estimation, relighting with reflectance‐aware scene‐specific representations and finally relighting as image‐to‐image transformations. Each category is concluded with a discussion on the main characteristics of the current methods and possible future trends. We also provide an overview of current publicly available datasets for neural lighting applications. Farshad Einabadi, Jean-Yves Guillemaut, Adrian Hilton 0001 |
Comput. Graph. Forum | 3 |
| 2021 | Temporally Coherent General Dynamic Scene ReconstructionabstractAbstract Existing techniques for dynamic scene reconstruction from multiple wide-baseline cameras primarily focus on reconstruction in controlled environments, with fixed calibrated cameras and strong prior constraints. This paper introduces a general approach to obtain a 4D representation of complex dynamic scenes from multi-view wide-baseline static or moving cameras without prior knowledge of the scene structure, appearance, or illumination. Contributions of the work are: an automatic method for initial coarse reconstruction to initialize joint estimation; sparse-to-dense temporal correspondence integrated with joint multi-view segmentation and reconstruction to introduce temporal coherence; and a general robust approach for joint segmentation refinement and dense reconstruction of dynamic scenes by introducing shape constraint. Comparison with state-of-the-art approaches on a variety of complex indoor and outdoor scenes, demonstrates improved accuracy in both multi-view segmentation and dense reconstruction. This paper demonstrates unsupervised reconstruction of complete temporally coherent 4D scene models with improved non-rigid object segmentation and shape reconstruction and its application to various applications such as free-view rendering and virtual reality. Armin Mustafa, Marco Volino, Hansung Kim 0001, Jean-Yves Guillemaut, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 5 |
| 2021 | Channel and spatial attention based deep object co-segmentation
Jia Chen 0026, Yasong Chen, Guoqin Ning, Mingwen Tong, Adrian Hilton 0001 |
Knowl. Based Syst. | 6 |
| 2021 | Acoustic Room Modelling Using 360 Stereo CamerasabstractIn this paper we propose a pipeline for estimating acoustic 3D room structure with geometry and attribute prediction using spherical 360$^{\circ }$cameras. Instead of setting microphone arrays with loudspeakers to measure acoustic parameters for specific rooms, a simple and practical single-shot capture of the scene using a stereo pair of 360 cameras can be used to simulate those acoustic parameters. We assume that the room and objects can be represented as cuboids aligned to the main axes of the room coordinate (Manhattan world). The scene is captured as a stereo pair using off-the-shelf consumer spherical 360 cameras. A cuboid-based 3D room geometry model is estimated by correspondence matching between captured images and semantic labelling using a convolutional neural network (SegNet). The estimated geometry is used to produce frequency-dependent acoustic predictions of the scene. This is, to our knowledge, the first attempt in the literature to use visual geometry estimation and object classification algorithms to predict acoustic properties. Results are compared to measurements through calculated reverberant spatial audio object parameters used for reverberation reproduction customized to the given loudspeaker set up. Hansung Kim 0001, Luca Remaggi, Sam Fowler, Philip J. B. Jackson, Adrian Hilton 0001 |
IEEE Trans. Multim. | 5 |
| 2020 | Message from the 3DV 2020 Program Chairs
Adrian Hilton 0001, Zuzana Kukelova, Stephen Lin 0001, Jun Sato |
3DV | 1 |
| 2020 | Multi-view Consistency Loss for Improved Single-Image 3D Reconstruction of Clothed People
Akin Caliskan, Armin Mustafa, Evren Imre, Adrian Hilton 0001 |
ACCV (1) | 4 |
| 2020 | Semantic Estimation of 3D Body Shape and Pose using Minimal Cameras
Andrew Gilbert, Matthew Trumble, Adrian Hilton 0001, John P. Collomosse |
BMVC | 3 |
| 2020 | 3D Multi Person Tracking With Dual 360° CamerasabstractPerson tracking is an often studied facet of computer vision, with applications in security, automated driving and entertainment. However, despite the advantages they offer, few current solutions work for 360° cameras, due to projection distortion. This paper presents a simple yet robust method for 3D tracking of multiple people in a scene from a pair of 360° cameras. By using 2D pose information, rather than potentially unreliable 3D position or repeated colour information, we create a tracker that is both appearance independent as well as capable of operating at narrow baseline. Our results demonstrate state of the art performance on 360° scenes, as well as the capability to handle vertical axis rotation. Matthew Shere, Hansung Kim 0001, Adrian Hilton 0001 |
ICIP | 3 |
| 2020 | EdgeNet: Semantic Scene Completion from a Single RGB- D ImageabstractSemantic scene completion is the task of predicting a complete 3D representation of volumetric occupancy with corresponding semantic labels for a scene from a single point of view. In this paper, we present EdgeNet, a new end-to-end neural network architecture that fuses information from depth and RGB, explicitly representing RGB edges in 3D space. Previous works on this task used either depth-only or depth with colour by projecting 2D semantic labels generated by a 2D segmentation network into the 3D volume, requiring a two step training process. Our EdgeNet representation encodes colour information in 3D space using edge detection and flipped truncated signed distance, which improves semantic completion scores especially in hard to detect classes. We achieved state-of-the-art scores on both synthetic and real datasets with a simpler and a more computationally efficient training pipeline than competing approaches. Aloisio Dourado, Teófilo Emídio de Campos, Hansung Kim 0001, Adrian Hilton 0001 |
ICPR | 4 |
| 2020 | Real-Time Multi-person Motion Capture from Multi-view Video and IMUsabstractAbstract A real-time motion capture system is presented which uses input from multiple standard video cameras and inertial measurement units (IMUs). The system is able to track multiple people simultaneously and requires no optical markers, specialized infra-red cameras or foreground/background segmentation, making it applicable to general indoor and outdoor scenarios with dynamic backgrounds and lighting. To overcome limitations of prior video or IMU-only approaches, we propose to use flexible combinations of multiple-view, calibrated video and IMU input along with a pose prior in an online optimization-based framework, which allows the full 6-DoF motion to be recovered including axial rotation of limbs and drift-free global position. A method for sorting and assigning raw input 2D keypoint detections into corresponding subjects is presented which facilitates multi-person tracking and rejection of any bystanders in the scene. The approach is evaluated on data from several indoor and outdoor capture environments with one or more subjects and the trade-off between input sparsity and tracking performance is discussed. State-of-the-art pose estimation performance is obtained on the Total Capture (mutli-view video and IMU) and Human 3.6M (multi-view video) datasets. Finally, a live demonstrator for the approach is presented showing real-time capture, solving and character animation using a light-weight, commodity hardware setup. Charles Malleson, John P. Collomosse, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | Semantically Coherent 4D Scene Flow of Dynamic ScenesabstractAbstract Simultaneous semantically coherent object-based long-term 4D scene flow estimation, co-segmentation and reconstruction is proposed exploiting the coherence in semantic class labels both spatially, between views at a single time instant, and temporally, between widely spaced time instants of dynamic objects with similar shape and appearance. In this paper we propose a framework for spatially and temporally coherent semantic 4D scene flow of general dynamic scenes from multiple view videos captured with a network of static or moving cameras. Semantic coherence results in improved 4D scene flow estimation, segmentation and reconstruction for complex dynamic scenes. Semantic tracklets are introduced to robustly initialize the scene flow in the joint estimation and enforce temporal coherence in 4D flow, semantic labelling and reconstruction between widely spaced instances of dynamic objects. Tracklets of dynamic objects enable unsupervised learning of long-term flow, appearance and shape priors that are exploited in semantically coherent 4D scene flow estimation, co-segmentation and reconstruction. Comprehensive performance evaluation against state-of-the-art techniques on challenging indoor and outdoor sequences with hand-held moving cameras shows improved accuracy in 4D scene flow, segmentation, temporally coherent semantic labelling, and reconstruction of dynamic scenes. Armin Mustafa, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Inpainting of Wide-Baseline Multiple Viewpoint VideoabstractWe describe a non-parametric algorithm for multiple-viewpoint video inpainting. Uniquely, our algorithm addresses the domain of wide baseline multiple-viewpoint video (MVV) with no temporal look-ahead in near real time speed. A Dictionary of Patches (DoP) is built using multi-resolution texture patches reprojected from geometric proxies available in the alternate views. We dynamically update the DoP over time, and a Markov Random Field optimisation over depth and appearance is used to resolve and align a selection of multiple candidates for a given patch, this ensures the inpainting of large regions in a plausible manner conserving both spatial and temporal coherence. We demonstrate the removal of large objects (e.g., people) on challenging indoor and outdoor MVV exhibiting cluttered, dynamic backgrounds and moving cameras. Andrew Gilbert, Matthew Trumble, Adrian Hilton 0001, John P. Collomosse |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2019 | Dynamic Surface Animation using Generative NetworksabstractThis paper presents techniques to animate realistic human-like motion using a compressed learnt model from 4D volumetric performance capture data. Sequences of 4D dynamic geometry representing a human performing an arbitrary motion are encoded through a generative network into a compact space representation, whilst maintaining the original properties, such as, surface dynamics. An animation framework is proposed which computes an optimal motion graph using the novel capabilities of compression and generative synthesis properties of the network. This approach significantly reduces the memory space requirements, improves quality of animation, and facilitates the interpolation between motions. The framework optimises the number of transitions in the graph with respect to the shape and motion of the dynamic content. This generates a compact graph structure with low edge connectivity, and maintains realism when transitioning between motions. Finally, it demonstrates that generative networks facilitate the computation of novel poses, and provides a compact motion graph representation of captured dynamic shape enabling real-time interactive animation and interpolation of novel poses to smoothly transition between motions. João Regateiro, Adrian Hilton 0001, Marco Volino |
3DV | 2 |
| 2019 | Light Field Compression using Eigen TexturesabstractLight fields are becoming an increasingly popular method of digital content production for visual effects and virtual/augmented reality as they capture a view dependent representation enabling photo realistic rendering over a range of viewpoints. Light field video is generally captured using arrays of cameras resulting in tens to hundreds of images of a scene at each time instance. An open problem is how to efficiently represent the data preserving the view-dependent detail of the surface in such a way that is compact to store and efficient to render. In this paper we show that constructing an Eigen texture basis representation from the light field using an approximate 3D surface reconstruction as a geometric proxy provides a compact representation that maintains view-dependent realism. We demonstrate that the proposed method is able to reduce storage requirements by > 95% while maintaining the visual quality of the captured data. An efficient view-dependent rendering technique is also proposed which is performed in eigen space allowing smooth continuous viewpoint interpolation through the light field. Marco Volino, Armin Mustafa, Jean-Yves Guillemaut, Adrian Hilton 0001 |
3DV | 4 |
| 2019 | U4D: Unsupervised 4D Dynamic Scene UnderstandingabstractWe introduce the first approach to solve the challenging problem of unsupervised 4D visual scene understanding for complex dynamic scenes with multiple interacting people from multi-view video. Our approach simultaneously estimates a detailed model that includes a per-pixel semantically and temporally coherent reconstruction, together with instance-level segmentation exploiting photo-consistency, semantic and motion information. We further leverage recent advances in 3D pose estimation to constrain the joint semantic instance segmentation and 4D temporally coherent reconstruction. This enables per person semantic instance segmentation of multiple interacting people in complex dynamic scenes. Extensive evaluation of the joint visual scene understanding framework against state-of-the-art methods on challenging indoor and outdoor sequences demonstrates a significant (approx 40%) improvement in semantic segmentation, reconstruction and scene flow accuracy. Armin Mustafa, Chris Russell 0001, Adrian Hilton 0001 |
ICCV | 3 |
| 2019 | Spectral Analysis Network for Deep Representation Learning and Image ClusteringabstractDeep representation learning is a crucial procedure in multimedia analysis and attracts increasing attention. Most of the popular techniques rely on convolutional neural network and require a large amount of labeled data in the training procedure. However, it is time consuming or even impossible to obtain the label information in some tasks due to cost limitation. Thus, it is necessary to develop unsupervised deep representation learning techniques. This paper proposes a new network structure for unsupervised deep representation learning based on spectral analysis, which is a popular technique with solid theory foundations. Compared with the existing spectral analysis methods, the proposed network structure has at least three advantages. Firstly, it can identify the local similarities among images in patch level and thus more robust against occlusion. Secondly, through multiple consecutive spectral analysis procedures, the proposed network can learn more clustering-friendly representations and is capable to reveal the deep correlations among data samples. Thirdly, it can elegantly integrate different spectral analysis procedures, so that each spectral analysis procedure can have their individual strengths in dealing with different data sample distributions. Extensive experimental results show the effectiveness of the proposed methods on various image clustering tasks. Adrian Hilton 0001, Jianmin Jiang |
ICME | 2 |
| 2019 | Immersive Spatial Audio Reproduction for VR/AR Using Room Acoustic Modelling from 360° ImagesabstractRecent progresses in Virtual Reality (VR) and Augmented Reality (AR) allow us to experience various VR/AR applications in our daily life. In order to maximise the immersiveness of user in VR/AR environments, a plausible spatial audio reproduction synchronised with visual information is essential. In this paper, we propose a simple and efficient system to estimate room acoustic for plausible reproducton of spatial audio using 360° cameras for VR/AR applications. A pair of 360° images is used for room geometry and acoustic property estimation. A simplified 3D geometric model of the scene is estimated by depth estimation from captured images and semantic labelling using a convolutional neural network (CNN). The real environment acoustics are characterised by frequency-dependent acoustic predictions of the scene. Spatially synchronised audio is reproduced based on the estimated geometric and acoustic properties in the scene. The reconstructed scenes are rendered with synthesised spatial audio as VR/AR content. The results of estimated room geometry and simulated spatial audio are evaluated against the actual measurements and audio calculated from ground-truth Room Impulse Responses (RIRs) recorded in the rooms. Hansung Kim 0001, Luca Hernaggi, Philip J. B. Jackson, Adrian Hilton 0001 |
VR | 4 |
| 2019 | OCEAN: Object-centric arranging network for self-supervised visual representations learning
Changjae Oh, Bumsub Ham, Hansung Kim 0001, Adrian Hilton 0001, Kwanghoon Sohn |
Expert Syst. Appl. | 4 |
| 2019 | Fusing Visual and Inertial Sensors with Semantics for 3D Human Pose EstimationabstractWe propose an approach to accurately estimate 3D human pose by fusing multi-viewpoint video (MVV) with inertial measurement unit (IMU) sensor data, without optical markers, a complex hardware setup or a full body model. Uniquely we use a multi-channel 3D convolutional neural network to learn a pose embedding from visual occupancy and semantic 2D pose estimates from the MVV in a discretised volumetric probabilistic visual hull. The learnt pose stream is concurrently processed with a forward kinematic solve of the IMU data and a temporal model (LSTM) exploits the rich spatial and temporal long range dependencies among the solved joints, the two streams are then fused in a final fully connected layer. The two complementary data sources allow for ambiguities to be resolved within each sensor modality, yielding improved accuracy over prior methods. Extensive evaluation is performed with state of the art performance reported on the popular Human 3.6M dataset (Ionescu et al. in Intell IEEE Trans Pattern Anal Mach 36(7):1325–1339, 2014), the newly released TotalCapture dataset and a challenging set of outdoor videos TotalCaptureOutdoor. We release the new hybrid MVV dataset (TotalCapture) comprising of multi-viewpoint video, IMU and accurate 3D skeletal joint ground truth derived from a commercial motion capture system. The dataset is available online at http://cvssp.org/data/totalcapture/ . Andrew Gilbert, Matthew Trumble, Charles Malleson, Adrian Hilton 0001, John P. Collomosse |
Int. J. Comput. Vis. | 4 |
| 2019 | Hybrid Modeling of Non-Rigid Scenes From RGBD CamerasabstractRecent advances in sensor technology have introduced low-cost RGB video plus depth sensors, such as the Kinect, which enable simultaneous acquisition of color and depth images at video rates. This paper introduces a framework for representation of general dynamic scenes from video plus depth acquisition. A hybrid representation is proposed which combines the advantages of prior surfel graph surface segmentation and modeling work with the higher resolution surface reconstruction capability of volumetric fusion techniques. The contributions are: 1) extension of a prior piecewise surfel graph modeling approach for improved accuracy and completeness; 2) combination of this surfel graph modeling with a truncated signed distance function surface fusion to generate dense geometry; and 3) proposal of means for validation of the reconstructed a 4D scene model against the input data and efficient storage of any unmodeled regions via residual depth maps. The approach allows arbitrary dynamic scenes to be efficiently represented with a temporally consistent structure and enhanced levels of detail and completeness where possible, but gracefully falls back to raw measurements where no structure can be inferred. The representation is shown to facilitate creative manipulation of real scene data which would previously require more complex capture set-ups or manual processing. Charles Malleson, Jean-Yves Guillemaut, Adrian Hilton 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | MSFD: Multi-Scale Segmentation-Based Feature Detection for Wide-Baseline Scene ReconstructionabstractA common problem in wide-baseline matching is the sparse and non-uniform distribution of correspondences when using conventional detectors, such as SIFT, SURF, FAST, A-KAZE, and MSER. In this paper, we introduce a novel segmentation-based feature detector (SFD) that produces an increased number of accurate features for wide-baseline matching. A multi-scale SFD is proposed using bilateral image decomposition to produce a large number of scale-invariant features for wide-baseline reconstruction. All input images are over-segmented into regions using any existing segmentation technique, such as Watershed, Mean-shift, and simple linear iterative clustering. Feature points are then detected at the intersection of the boundaries of three or more regions. The detected feature points are local maxima of the image function. The key advantage of feature detection based on segmentation is that it does not require global threshold setting and can, therefore, detect features throughout the image. A comprehensive evaluation demonstrates that SFD gives an increased number of features that are accurately localized and matched between wide-baseline camera views; the number of features for a given matching error increases by a factor of 3-5 compared with SIFT; feature detection and matching performance are maintained with increasing baseline between views; multi-scale SFD improves matching performance at varying scales. Application of SFD to sparse multi-view wide-baseline reconstruction demonstrates a factor of 10 increases in the number of reconstructed points with improved scene coverage compared with SIFT/MSER/A-KAZE. Evaluation against ground-truth shows that SFD produces an increased number of wide-baseline matches with a reduced error. Armin Mustafa, Hansung Kim 0001, Adrian Hilton 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Human-Centric Scene Understanding from Single View 360 VideoabstractIn this paper, we propose an approach to indoor scene understanding from observation of people in single view spherical video. As input, our approach takes a centrally located spherical video capture of an indoor scene, estimating the 3D localisation of human actions performed throughout the long term capture. The central contribution of this work is a deep convolutional encoder-decoder network trained on a synthetic dataset to reconstruct regions of affordance from captured human activity. The predicted affordance segmentation is then applied to compose a reconstruction of the complete 3D scene, integrating the affordance segmentation into 3D space. The mapping learnt between human activity and affordance segmentation demonstrates that omnidirectional observation of human activity can be applied to scene understanding tasks such as 3D reconstruction. We show that our approach using only observation of people performs well against previous approaches, allowing reconstruction of occluded regions and labelling of scene affordances. Sam Fowler, Hansung Kim 0001, Adrian Hilton 0001 |
3DV | 3 |
| 2018 | Hybrid Skeleton Driven Surface Registration for Temporally Consistent Volumetric VideoabstractThis paper presents a hybrid skeleton-driven surface registration (HSDSR) approach to generate temporally consistent meshes from multiple view video of human subjects. 2D pose detections from multiple view video are used to estimate 3D skeletal pose on a per-frame basis. The 3D pose is embedded into a 3D surface reconstruction allowing any frame to be reposed into the shape from any other frame in the captured sequence. Skeletal motion transfer is performed by selecting a reference frame from the surface reconstruction data and reposing it to match the pose estimation of other frames in a sequence. This allows an initial coarse alignment to be performed prior to refinement by a patch-based non-rigid mesh deformation. The proposed approach overcomes limitations of previous work by reposing a reference mesh to match the pose of a target mesh reconstruction, providing a closer starting point for further non-rigid mesh deformation. It is shown that the proposed approach is able to achieve comparable results to existing model-based and model-free approaches. Finally, it is demonstrated that this framework provides an intuitive way for artists and animators to edit volumetric video. João Regateiro, Marco Volino, Adrian Hilton 0001 |
3DV | 3 |
| 2018 | Volumetric Performance Capture from Minimal Camera Viewpoints
Andrew Gilbert, Marco Volino, John P. Collomosse, Adrian Hilton 0001 |
ECCV (11) | 4 |
| 2018 | Deep Autoencoder for Combined Human Pose Estimation and Body Model Upscaling
Matthew Trumble, Andrew Gilbert, Adrian Hilton 0001, John P. Collomosse |
ECCV (10) | 3 |
| 2018 | Non-Zero Diffusion Particle Flow SMC-PHD Filter for Audio-Visual Multi-Speaker TrackingabstractThe sequential Monte Carlo probability hypothesis density (SMC-PHD) filter has been shown to be promising for audio-visual multi-speaker tracking. Recently, the zero diffusion particle flow (ZPF) has been used to mitigate the weight degeneracy problem in the SMC-PHD filter. However, this leads to a substantial increase in the computational cost due to the migration of particles from prior to posterior distribution with a partial differential equation. This paper proposes an alternative method based on the non-zero diffusion particle flow (NPF) to adjust the particle states by fitting the particle distribution with the posterior probability density using the non-zero diffusion. This property allows efficient computation of the migration of particles. Results from the AV16.3 dataset demonstrate that we can significantly mitigate the weight degeneracy problem with a smaller computational cost as compared with the ZPF based SMC-PHD filter. Yang Liu 0175, Adrian Hilton 0001, Jonathon A. Chambers, Yuxin Zhao 0001, Wenwu Wang 0001 |
ICASSP | 2 |
| 2018 | Acoustic Reflector Localization and ClassificationabstractThe process of understanding acoustic properties of environments is important for several applications, such as spatial audio, augmented reality and source separation. In this paper, multichannel room impulse responses are recorded and transformed into their direction of arrival (DOA)-time domain, by employing a superdirective beamformer. This domain can be represented as a 2D image. Hence, a novel image processing method is proposed to analyze the DOA-time domain, and estimate the reflection times of arrival and DOAs. The main acoustically reflective objects are then localized. Recent studies in acoustic reflector localization usually assume the room to be free from furniture. Here, by analyzing the scattered reflections, an algorithm is also proposed to binary classify reflectors into room boundaries and interior furniture. Experiments were conducted in four rooms. The classification algorithm showed high quality performance, also improving the localization accuracy, for non-static listener scenarios. Luca Remaggi, Hansung Kim 0001, Philip J. B. Jackson, Filippo Maria Fazi, Adrian Hilton 0001 |
ICASSP | 5 |
| 2018 | AVSU: Workshop on Audio-Visual Scene Understanding for Immersive MultimediaabstractThis workshop aims to provide a forum to exchange ideas in scene understanding techniques researched in audio and visual communities, and to ultimately unlock the creative potential of joint audio-visual signal processing to deliver a step change in various multimedia applications. Papers and talks presented in this workshop will contribute to the emerging technology for audio and visual information that can improve traditional approaches for multimedia content production and reproduction. The goals of this workshop are to (1) present and discuss the latest trends in audio and computer vision fields for the common research goals, (2) understand state-of-the-art techniques and bottlenecks in the other's discipline for the common topics, (3) investigate research opportunities of joint audio-visual scene understandings in multimedia content production. This workshop will be a good opportunity to bring together leading experts in audio processing and computer vision, and will bridge the gap between two research fields in multimedia content production and reproduction. Adrian Hilton 0001, Hong-Goo Kang, Hansung Kim 0001, Kwanghoon Sohn |
ACM Multimedia | 1 |
| 2018 | Scalable object instance recognition based on keygraph matching
Estephan Dazzi, Teófilo Emídio de Campos, Adrian Hilton 0001, Roberto Marcondes Cesar Junior |
Pattern Recognit. Lett. | 3 |
| 2018 | Multimodal Visual Data Registration for Web-Based Visualization in Media ProductionabstractRecent developments of video and sensing technology have led to large volumes of digital media data. Current media production relies on videos from the principal camera together with a wide variety of heterogeneous source of supporting data [photos, light detection and ranging point clouds, witness video camera, high dynamic range imaging, and depth imagery]. Registration of visual data acquired from various 2D and 3D sensing modalities is challenging because current matching and registration methods are not appropriate due to differences in structure, format, and noise characteristics for multimodal data. A combined 2D/3D visualization of this registered data allows an integrated overview of the entire data set. For such a visualization, a Web-based context presents several advantages. In this paper, we propose a unified framework for registration and visualization of this type of visual media data. A new feature description and matching method is proposed, adaptively considering local geometry, semiglobal geometry, and color information in the scene for more robust registration. The resulting registered 2D/3D multimodal visual data are too large to be downloaded and viewed directly via the Web browser, while maintaining an acceptable user experience. Thus, we employ hierarchical techniques for compression and restructuring to enable efficient transmission and visualization over the Web, leading to interactive visualization as registered point clouds, 2D images, and videos in the browser, improving on the current state-of-the-art techniques for Web-based visualization of big media data. This is the first unified 3D Web-based visualization of multimodal visual media production data sets. The proposed pipeline is tested on big multimodal data set typical of film and broadcast production, which are made publicly available. The proposed feature description method shows two times higher precision of feature matching and more stable registration performance than existing 3D feature descriptors. Hansung Kim 0001, Alun Evans, Josep Blat, Adrian Hilton 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | An Audio-Visual System for Object-Based Audio: From Recording to ListeningabstractObject-based audio is an emerging representation for audio content, where content is represented in a reproduction-format-agnostic way and, thus, produced once for consumption on many different kinds of devices. This affords new opportunities for immersive, personalized, and interactive listening experiences. This paper introduces an end-to-end object-based spatial audio pipeline, from sound recording to listening. A high-level system architecture is proposed, which includes novel audio-visual interfaces to support object-based capture and listener-tracked rendering, and incorporates a proposed component for objectification, that is, recording content directly into an object-based form. Text-based and extensible metadata enable communication between the system components. An open architecture for object rendering is also proposed. The system's capabilities are evaluated in two parts. First, listener-tracked reproduction of metadata automatically estimated from two moving talkers is evaluated using an objective binaural localization model. Second, object-based scene capture with audio extracted using blind source separation (to remix between two talkers) and beamforming (to remix a recording of a jazz group) is evaluated with perceptually motivated objective and subjective experiments. These experiments demonstrate that the novel components of the system add capabilities beyond the state of the art. Finally, we discuss challenges and future perspectives for object-based audio workflows. Philip Coleman, Andreas Franck, Jon Francombe, Qingju Liu, Teófilo Emídio de Campos, Richard J. Hughes 0001, Dylan Menzies, Marcos F. Simón Gálvez, James Woodcock, Philip J. B. Jackson, Frank Melchior, Chris Pike, Filippo Maria Fazi, Trevor J. Cox, Adrian Hilton 0001 |
IEEE Trans. Multim. | 16 |
| 2018 | Multiple Speaker Tracking in Spatial Audio via PHD Filtering and Depth-Audio FusionabstractIn the object-based spatial audio system, positions of the audio objects (e.g., speakers/talkers or voices) presented in the sound scene are required as important metadata attributes for object acquisition and reproduction. Binaural microphones are often used as a physical device to mimic human hearing and to monitor and analyze the scene, including localization and tracking of multiple speakers. The binaural audio tracker, however, is usually prone to the errors caused by room reverberation and background noise. To address this limitation, we present a multimodal tracking method by fusing the binaural audio with depth information (from a depth sensor, e.g., Kinect). More specifically, the probability hypothesis density (PHD) filtering framework is first applied to the depth stream, and a novel clutter intensity model is proposed to improve the robustness of the PHD filter when an object is occluded either by other objects or due to the limited field of view of the depth sensor. To compensate misdetections in the depth stream, a novel gap filling technique is presented to map audio azimuths obtained from the binaural audio tracker to 3D positions, using speaker-dependent spatial constraints learned from the depth stream. With our proposed method, both the errors in the binaural tracker and the misdetections in the depth tracker can be significantly reduced. Real-room recordings are used to show the improved performance of the proposed method in removing outliers and reducing misdetections. Qingju Liu, Wenwu Wang 0001, Teófilo Emídio de Campos, Philip J. B. Jackson, Adrian Hilton 0001 |
IEEE Trans. Multim. | 5 |
| 2017 | 3D Room Geometry Reconstruction Using Audio-Visual SensorsabstractIn this paper we propose a cuboid-based air-tight indoor room geometry estimation method using combination of audio-visual sensors. Existing vision-based 3D reconstruction methods are not applicable for scenes with transparent or reflective objects such as windows and mirrors. In this work we fuse multi-modal sensory information to overcome the limitations of purely visual reconstruction for reconstruction of complex scenes including transparent and mirror surfaces. A full scene is captured by 360$^{\circ}$ cameras and acoustic room impulse responses (RIRs) recorded by a loudspeaker and compact microphone array. Depth information of the scene is recovered by stereo matching from the captured images and estimation of major acoustic reflector locations from the sound. The coordinate systems for audio-visual sensors are aligned into a unified reference frame and plane elements are reconstructed from audio-visual data. Finally cuboid proxies are fitted to the planes to generate a complete room model. Experimental results show that the proposed system generates complete representations of the room structures regardless of transparent windows, featureless walls and shiny surfaces. Hansung Kim 0001, Luca Remaggi, Philip J. B. Jackson, Filippo Maria Fazi, Adrian Hilton 0001 |
3DV | 5 |
| 2017 | Real-Time Full-Body Motion Capture from Video and IMUsabstractA real-time full-body motion capture system is presented which uses input from a sparse set of inertial measurement units (IMUs) along with images from two or more standard video cameras and requires no optical markers or specialized infra-red cameras. A real-time optimization-based framework is proposed which incorporates constraints from the IMUs, cameras and a prior pose model. The combination of video and IMU data allows the full 6-DOF motion to be recovered including axial rotation of limbs and drift-free global position. The approach was tested using both indoor and outdoor captured data. The results demonstrate the effectiveness of the approach for tracking a wide range of human motion in real time in unconstrained indoor/outdoor scenes. Charles Malleson, Andrew Gilbert, Matthew Trumble, John P. Collomosse, Adrian Hilton 0001, Marco Volino |
3DV | 5 |
| 2017 | 4D Temporally Coherent Light-Field VideoabstractLight-field video has recently been used in virtual and augmented reality applications to increase realism and immersion. However, existing light-field methods are generally limited to static scenes due to the requirement to acquire a dense scene representation. The large amount of data and the absence of methods to infer temporal coherence pose major challenges in storage, compression and editing compared to conventional video. In this paper, we propose the first method to extract a spatio-temporally coherent light-field video representation. A novel method to obtain Epipolar Plane Images (EPIs) from a spare lightfield camera array is proposed. EPIs are used to constrain scene flow estimation to obtain 4D temporally coherent representations of dynamic light-fields. Temporal coherence is achieved on a variety of light-field datasets. Evaluation of the proposed light-field scene flow against existing multi-view dense correspondence approaches demonstrates a significant improvement in accuracy of temporal coherence. Armin Mustafa, Marco Volino, Jean-Yves Guillemaut, Adrian Hilton 0001 |
3DV | 4 |
| 2017 | Towards Complete Scene Reconstruction from Single-View Depth and Human Motion
Sam Fowler, Hansung Kim 0001, Adrian Hilton 0001 |
BMVC | 3 |
| 2017 | Total Capture: 3D Human Pose Estimation Fusing Video and Inertial Sensors
Matthew Trumble, Andrew Gilbert, Charles Malleson, Adrian Hilton 0001, John P. Collomosse |
BMVC | 4 |
| 2017 | Semantically Coherent Co-Segmentation and Reconstruction of Dynamic ScenesabstractIn this paper we propose a framework for spatially and temporally coherent semantic co-segmentation and reconstruction of complex dynamic scenes from multiple static or moving cameras. Semantic co-segmentation exploits the coherence in semantic class labels both spatially, between views at a single time instant, and temporally, between widely spaced time instants of dynamic objects with similar shape and appearance. We demonstrate that semantic coherence results in improved segmentation and reconstruction for complex scenes. A joint formulation is proposed for semantically coherent object-based co-segmentation and reconstruction of scenes by enforcing consistent semantic labelling between views and over time. Semantic tracklets are introduced to enforce temporal coherence in semantic labelling and reconstruction between widely spaced instances of dynamic objects. Tracklets of dynamic objects enable unsupervised learning of appearance and shape priors that are exploited in joint segmentation and reconstruction. Evaluation on challenging indoor and outdoor sequences with hand-held moving cameras shows improved accuracy in segmentation, temporally coherent semantic labelling and 3D reconstruction of dynamic scenes. Armin Mustafa, Adrian Hilton 0001 |
CVPR | 2 |
| 2017 | Computer Vision in Sports
Thomas B. Moeslund, Graham A. Thomas, Adrian Hilton 0001, Peter Carr 0001, Irfan A. Essa |
Comput. Vis. Image Underst. | 3 |
| 2017 | Computer vision for sports: Current applications and research topics
Graham A. Thomas, Rikke Gade, Thomas B. Moeslund, Peter Carr 0001, Adrian Hilton 0001 |
Comput. Vis. Image Underst. | 5 |
| 2016 | Room Layout Estimation with Object and Material Attributes Information Using a Spherical CameraabstractIn this paper we propose a pipeline for estimating 3D room layout with object and material attribute prediction using a spherical stereo image pair. We assume that the room and objects can be represented as cuboids aligned to the main axes of the room coordinate (Manhattan world). A spherical stereo alignment algorithm is proposed to align two spherical images to the global world coordinate system. Depth information of the scene is estimated by stereo matching between images. Cubic projection images of the spherical RGB and estimated depth are used for object and material attribute detection. A single Convolutional Neural Network is designed to assign object and attribute labels to geometrical elements built from the spherical image. Finally simplified room layout is reconstructed by cuboid fitting. The reconstructed cuboid-based model shows the structure of the scene with object information and material attributes. Hansung Kim 0001, Teófilo Emídio de Campos, Adrian Hilton 0001 |
3DV | 3 |
| 2016 | Temporally Coherent 4D Reconstruction of Complex Dynamic ScenesabstractThis paper presents an approach for reconstruction of 4D temporally coherent models of complex dynamic scenes. No prior knowledge is required of scene structure or camera calibration allowing reconstruction from multiple moving cameras. Sparse-to-dense temporal correspondence is integrated with joint multi-view segmentation and reconstruction to obtain a complete 4D representation of static and dynamic objects. Temporal coherence is exploited to overcome visual ambiguities resulting in improved reconstruction of complex scenes. Robust joint segmentation and reconstruction of dynamic objects is achieved by introducing a geodesic star convexity constraint. Comparative evaluation is performed on a variety of unstructured indoor and outdoor dynamic scenes with hand-held cameras and multiple people. This demonstrates reconstruction of complete temporally coherent 4D scene models with improved nonrigid object segmentation and shape reconstruction. Armin Mustafa, Hansung Kim 0001, Jean-Yves Guillemaut, Adrian Hilton 0001 |
CVPR | 4 |
| 2016 | 4D Match Trees for Non-rigid Surface Alignment
Armin Mustafa, Hansung Kim 0001, Adrian Hilton 0001 |
ECCV (1) | 3 |
| 2016 | Identity association using PHD filters in multiple head tracking with depth sensorsabstractThe work on 3D human pose estimation has been through a significant amount of progress in recent years, particularly due to the widespread availability of commodity depth sensors. However, most pose estimation methods follow a tracking-as-detection approach which does not explicitly handle occlusions, thus introducing outliers and identity association issues when multiple targets are involved. To address these issues, we propose a new method based on Probability Hypothesis Density (PHD) filter. In this method, the PHD filter with a novel clutter intensity model is used to remove outliers in the 3D head detection results, followed by an identity association scheme with occlusion detection for the targets. Experimental results show that our proposed method greatly mitigates the outliers, and correctly associates identities to individual detections with low computational cost. Qingju Liu, Teófilo Emídio de Campos, Wenwu Wang 0001, Adrian Hilton 0001 |
ICASSP | 4 |
| 2016 | Big Data Analysis for Media ProductionabstractA typical high-end film production generates several terabytes of data per day, either as footage from multiple cameras or as background information regarding the set (laser scans, spherical captures, etc). This paper presents solutions to improve the integration of the multiple data sources, and understand their quality and content, which are useful both to support creative decisions on-set (or near it) and enhance the postproduction process. The main cinema specific contributions, tested on a multisource production dataset made publicly available for research purposes, are the monitoring and quality assurance of multicamera set-ups, multisource registration and acceleration of 3-D reconstruction, anthropocentric visual analysis techniques for semantic content annotation, and integrated 2-D–3-D web visualization tools. We discuss as well improvements carried out in basic techniques for acceleration, clustering and visualization, which were necessary to deal with the very large multisource data, and can be applied to other big data problems in diverse application fields. Josep Blat, Alun Evans, Hansung Kim 0001, Evren Imre, Lukás Polok, Viorela Ila, Nikos Nikolaidis 0001, Pavel Zemcík, Anastasios Tefas, Pavel Smrz, Adrian Hilton 0001, Ioannis Pitas |
Proc. IEEE | 11 |
| 2016 | Mean-Shift and Sparse Sampling-Based SMC-PHD Filtering for Audio Informed Visual Speaker TrackingabstractThe probability hypothesis density (PHD) filter based on sequential Monte Carlo (SMC) approximation (also known as SMC-PHD filter) has proven to be a promising algorithm for multispeaker tracking. However, it has a heavy computational cost as surviving, spawned, and born particles need to be distributed in each frame to model the state of the speakers and to estimate jointly the variable number of speakers with their states. In particular, the computational cost is mostly caused by the born particles as they need to be propagated over the entire image in every frame to detect the new speaker presence in the view of the visual tracker. In this paper, we propose to use the audio data to improve the visual SMC-PHD (V-SMC-PHD) filter by using the direction of arrival angles of the audio sources to determine when to propagate the born particles and reallocate the surviving and spawned particles. The tracking accuracy of the audio-visual SMC-PHD (AV-SMC-PHD) algorithm is further improved by using a modified mean-shift algorithm to search and climb density gradients iteratively to find the peak of the probability distribution, and the extra computational complexity introduced by mean-shift is controlled with a sparse sampling technique. These improved algorithms, named as AVMS-SMC-PHD and sparse-AVMS-SMC-PHD, respectively, are compared systematically with AV-SMC-PHD and V-SMC-PHD based on the AV16.3, AMI, and CLEAR datasets. Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Adrian Hilton 0001, Josef Kittler |
IEEE Trans. Multim. | 4 |
| 2015 | Segmentation Based Features for Wide-Baseline Multi-view ReconstructionabstractA common problem in wide-baseline stereo is the sparse and non-uniform distribution of correspondences when using conventional detectors such as SIFT, SURF, FAST and MSER. In this paper we introduce a novel segmentation based feature detector SFD that produces an increased number of 'good' features for accurate wide-baseline reconstruction. Each image is segmented into regions by over-segmentation and feature points are detected at the intersection of the boundaries for three or more regions. Segmentation-based feature detection locates features at local maxima giving a relatively large number of feature points which are consistently detected across wide-baseline views and accurately localised. A comprehensive comparative performance evaluation with previous feature detection approaches demonstrates that: SFD produces a large number of features with increased scene coverage, detected features are consistent across wide-baseline views for images of a variety of indoor and outdoor scenes, and the number of wide-baseline matches is increased by an order of magnitude compared to alternative detector-descriptor combinations. Sparse scene reconstruction from multiple wide-baseline stereo views using the SFD feature detector demonstrates at least a factor six increase in the number of reconstructed points with reduced error distribution compared to SIFT when evaluated against ground-truth and similar computational cost to SURF/FAST. Armin Mustafa, Hansung Kim 0001, Evren Imre, Adrian Hilton 0001 |
3DV | 4 |
| 2015 | FaceDirector: Continuous Control of Facial Performance in VideoabstractWe present a method to continuously blend between multiple facial performances of an actor, which can contain different facial expressions or emotional states. As an example, given sad and angry video takes of a scene, our method empowers the movie director to specify arbitrary weighted combinations and smooth transitions between the two takes in post-production. Our contributions include (1) a robust nonlinear audio-visual synchronization technique that exploits complementary properties of audio and visual cues to automatically determine robust, dense spatiotemporal correspondences between takes, and (2) a seamless facial blending approach that provides the director full control to interpolate timing, facial expression, and local appearance, in order to generate novel performances after filming. In contrast to most previous works, our approach operates entirely in image space, avoiding the need of 3D facial reconstruction. We demonstrate that our method can synthesize visually believable performances with applications in emotion transition, performance correction, and timing control. Charles Malleson, Jean-Charles Bazin, Oliver Wang, Derek Bradley, Thabo Beeler, Adrian Hilton 0001, Alexander Sorkine-Hornung |
ICCV | 6 |
| 2015 | General Dynamic Scene Reconstruction from Multiple View VideoabstractThis paper introduces a general approach to dynamic scene reconstruction from multiple moving cameras without prior knowledge or limiting constraints on the scene structure, appearance, or illumination. Existing techniques or dynamic scene reconstruction from multiple wide-baseline camera views primarily focus on accurate reconstruction in controlled environments, where the cameras are fixed and calibrated and background is known. These approaches are not robust for general dynamic scenes captured with sparse moving cameras. Previous approaches for outdoor dynamic scene reconstruction assume prior knowledge of the static background appearance and structure. The primary contributions of this paper are twofold: an automatic method for initial coarse dynamic scene segmentation and reconstruction without prior knowledge of background appearance or structure, and a general robust approach for joint segmentation refinement and dense reconstruction of dynamic scenes from multiple wide-baseline static or moving cameras. Evaluation is performed on a variety of indoor and outdoor scenes with cluttered backgrounds and multiple dynamic non-rigid objects such as people. Comparison with state-of-the-art approaches demonstrates improved accuracy in both multiple view segmentation and dense reconstruction. The proposed approach also eliminates the requirement for prior knowledge of scene structure and appearance. Armin Mustafa, Hansung Kim 0001, Jean-Yves Guillemaut, Adrian Hilton 0001 |
ICCV | 4 |
| 2015 | Coverage evaluation of camera networks for facilitating big-data management in film productionabstractFilm production inherently generates large amounts of data-at a rate of 27TB/hour for a conventional multicamera setup [1]. In this paper, we propose a video coverage monitoring framework for such setups, which enables the identification of problematic sensor configurations, and therefore significantly reduces the data volume by eliminating unusable material before it is generated. Our approach involves analysing the projection of a set of 3D volume elements on the cameras, to verify whether they satisfy a number of constraints predicting the success of specified tasks. We demonstrate the utility of the proposed framework on three use cases, and conclude that our approach facilitates the development of tools with considerable practical value. Evren Imre, Adrian Hilton 0001 |
ICIP | 2 |
| 2015 | Multi-modal big-data management for film productionabstractModern digital film production uses large quantities of data from videos, digital photographs, LIDAR scans, spherical photography and many other sources to create the final film frames. The processing and management of this massive amount of heterogeneous data consumes enormous resources. We propose an integrated pipeline for 2D/3D data registration for film production. We present the prototype application Jigsaw, which allows users to efficiently manage and process various data from digital photographs to 3D point clouds. A key requirement in the use of multi-modal 2D/3D data for content production is the registration into a common coordinate frame. 3D geometric information is reconstructed from 2D data and registered to the reference 3D models using 3D feature matching. We provide a public multi-modal database captured with a wide variety of devices in different environments to assist further research. An order of magnitude gain in efficiency is achieved with the proposed approach. Hansung Kim 0001, Simon Pabst, Justin Sneddon, Ted Waine, Jeff Clifford, Adrian Hilton 0001 |
ICIP | 6 |
| 2015 | Audio informed visual speaker tracking with SMC-PHD filterabstractSequential Monte Carlo probability hypothesis density (SMC-PHD) filter has received much interest in the field of nonlinear non-Gaussian visual tracking due to its ability to handle a variable number of speakers. The SMC-PHD filter employs surviving, spawned and born particles to model the state of the speakers and jointly estimates the variable number of speakers with their states. The born particles play a critical role in the detection of new speakers, which makes it necessary to propagate them in each frame. However, this increases the computational cost of the visual tracker. Here, we propose to use audio data to determine when to propagate the born particles and re-allocate the surviving and spawned particles. In our framework, we employ audio data as an aid to visual SMC-PHD (V-SMC-PHD) filter by using the direction of arrival (DOA) angles of the audio sources to reshape the distribution of the particles. Experimental results on the AV16:3 dataset with multi-speaker sequences show that our proposed audio-visual SMC-PHD (AV-SMC-PHD) filter improves the tracking performance in terms of estimation accuracy and computational efficiency. Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Adrian Hilton 0001, Josef Kittler |
ICME | 4 |
| 2015 | Analyzing Muscle Activity and Force with Skin Shape Captured by Non-contact Visual Sensor
Ryusuke Sagawa, Yusuke Yoshiyasu, Alexander Alspach, Ko Ayusawa, Katsu Yamane, Adrian Hilton 0001 |
PSIVT | 6 |
| 2015 | Modeling dynamic scenes by one-shot 3D acquisition system for moving humanoid robotabstractFor mobile robots, 3D acquisition is required to model the environment. Particularly for humanoid robots, a modeled environment is necessary to plan the walking control. This environment can include both static objects, such as a ground surface with obstacles, and dynamic objects, such as a person moving around the robot. This paper proposes a system for a robot to obtain a sufficiently accurate shape of the environment for walking on a ground surface with obstacles and a method to detect dynamic objects in the modeled environment, which is necessary for the robot to react to sudden changes in the scene. The 3D acquisition is achieved by a projector-camera system mounted on the robot head that uses a structured-light method to reconstruct the shapes of moving objects from a single frame. The acquired shapes are aligned and merged into a common coordinate system using the simultaneous localization and mapping method. Dynamic objects are detected as shapes that are inconsistent with the previous frames. Experiments were performed to evaluate the accuracy of the 3D acquisition and the robustness with regard to detecting dynamic objects when serving as the vision system of a humanoid robot. Ryusuke Sagawa, Charles Malleson, Mitsuharu Morisawa, Kenji Kaneko, Fumio Kanehiro, Yoshio Matsumoto, Adrian Hilton 0001 |
RO-MAN | 7 |
| 2015 | 4D Model Flow: Precomputed Appearance Alignment for Real-time 4D Video InterpolationabstractWe introduce the concept of 4D model flow for the precomputed alignment of dynamic surface appearance across 4D video sequences of different motions reconstructed from multi-view video. Precomputed 4D model flow allows the efficient parametrization of surface appearance from the captured videos, which enables efficient real-time rendering of interpolated 4D video sequences whilst accurately reproducing visual dynamics, even when using a coarse underlying geometry. We estimate the 4D model flow using an image-based approach that is guided by available geometry proxies. We propose a novel representation in surface texture space for efficient storage and online parametric interpolation of dynamic appearance. Our 4D model flow overcomes previous requirements for computationally expensive online optical flow computation for data-driven alignment of dynamic surface appearance by precomputing the appearance alignment. This leads to an efficient rendering technique that enables the online interpolation between 4D videos in real time, from arbitrary viewpoints and with visual quality comparable to the state of the art. Dan Casas, Christian Richardt, John P. Collomosse, Christian Theobalt, Adrian Hilton 0001 |
Comput. Graph. Forum | 5 |
| 2015 | Covariance estimation for minimal geometry solvers via scaled unscented transformation
Evren Imre, Adrian Hilton 0001 |
Comput. Vis. Image Underst. | 2 |
| 2015 | Block world reconstruction from spherical stereo image pairs
Hansung Kim 0001, Adrian Hilton 0001 |
Comput. Vis. Image Underst. | 2 |
| 2015 | Order Statistics of RANSAC and Their Practical Application
Evren Imre, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 2 |
| 2015 | Hybrid Skeletal-Surface Motion Graphs for Character Animation from 4D Performance CaptureabstractWe present a novel hybrid representation for character animation from 4D Performance Capture (4DPC) data which combines skeletal control with surface motion graphs. 4DPC data are temporally aligned 3D mesh sequence reconstructions of the dynamic surface shape and associated appearance from multiple-view video. The hybrid representation supports the production of novel surface sequences which satisfy constraints from user-specified key-frames or a target skeletal motion. Motion graph path optimisation concatenates fragments of 4DPC data to satisfy the constraints while maintaining plausible surface motion at transitions between sequences. Space-time editing of the mesh sequence using a learned part-based Laplacian surface deformation model is performed to match the target skeletal motion and transition between sequences. The approach is quantitatively evaluated for three 4DPC datasets with a variety of clothing styles. Results for key-frame animation demonstrate production of novel sequences that satisfy constraints on timing and position of less than 1% of the sequence duration and path length. Evaluation of motion-capture-driven animation over a corpus of 130 sequences shows that the synthesised motion accurately matches the target skeletal motion. The combination of skeletal control with the surface motion graph extends the range and style of motion which can be produced while maintaining the natural dynamics of shape and appearance from the captured performance. Peng Huang 0001, Margara Tejera, John P. Collomosse, Adrian Hilton 0001 |
ACM Trans. Graph. | 4 |
| 2014 | Influence of Colour and Feature Geometry on Multi-modal 3D Point Clouds Data RegistrationabstractWith the current transition of various digital contents from 2D to 3D, the problem of 3D data matching and registration is increasingly important. Registration of multi-modal 3D data acquired from different sensors remains a challenging problem due to the difference in types and characteristics of the data. In this paper, we evaluate the registration performance of 3D feature descriptors with different domains on datasets from various environments and modalities. Datasets are acquired in indoor and outdoor environments with 2D and 3D sensing devices including LIDAR, spherical imaging, digital camera and RGBD camera. FPFH, PFH and SHOT feature descriptors are applied to the 3D point clouds generated from the multi-modal datasets. Local neighbouring point distribution, key points distribution, colour information and their combinations are used for feature description. Finally we analyse their influences on the multi-modal 3D point clouds data registration. Hansung Kim 0001, Adrian Hilton 0001 |
3DV | 2 |
| 2014 | Structured Representation of Non-Rigid Surfaces from Single View 3D Point TracksabstractThis work considers the problem of structured representation of dynamic surfaces from incomplete 3D point tracks from a single viewpoint. The surface is segmented into a set of connected regions each of which can be represented by a fixed intrinsic shape and a parametrised rigid/non-rigid motion trajectory. Neither the model parameters nor the point-to-model assignments are known upfront. Motion and geometric shape parameters are estimated in alternation with a graph-cuts based point-to-model assignment. This modelling process facilitates in-filling of missing data as well as de-noising of measurements by temporal integration while adding meaningful structure to the geometry and reducing storage cost by an order of magnitude. Experiments are presented for real and synthetic sequences to validate the approach and show how a single tuning parameter can be used to trade modelling error with extrapolation level and storage cost. Charles Malleson, Martin Klaudiny, Jean-Yves Guillemaut, Adrian Hilton 0001 |
3DV | 4 |
| 2014 | A Layered Model of Human Body and Garment DeformationabstractIn this paper we present a framework for learning a three layered model of human shape, pose and garment deformation. The proposed deformation model provides intuitive control over the three parameters independently, while producing aesthetically pleasing deformations of both the garment and the human body. The shape and pose deformation layers of the model are trained on a rich dataset of full body 3D scans of human subjects in a variety of poses. The garment deformation layer is trained on animated mesh sequences of dressed actors and relies on a novel technique for human shape and posture estimation under clothing. The key contribution of this paper is that we consider garment deformations as the residual transformations between a naked mesh and the dressed mesh of the same subject. Alexandros Neophytou, Adrian Hilton 0001 |
3DV | 2 |
| 2014 | Optimal Representation of Multiple View Video
Marco Volino, Dan Casas, John P. Collomosse, Adrian Hilton 0001 |
BMVC | 4 |
| 2014 | Intrinsic Textures for Relightable Free-Viewpoint Video
James Imber, Jean-Yves Guillemaut, Adrian Hilton 0001 |
ECCV (2) | 3 |
| 2014 | Hybrid 3D feature description and matching for multi-modal data registrationabstractWe propose a robust 3D feature description and registration method for 3D models reconstructed from various sensor devices. General 3D feature detectors and descriptors generally show low distinctiveness and repeatability for matching between different data modalities due to differences in noise and errors in geometry. The proposed method considers not only local 3D points but also neighbouring 3D keypoints to improve keypoint matching. The proposed method is tested on various multi-modal datasets including LIDAR scans, multiple photos, spherical images and RGBD videos to evaluate the performance against existing methods. Hansung Kim 0001, Adrian Hilton 0001 |
ICIP | 2 |
| 2014 | Wide Baseline Multi-view Video Matting Using a Hybrid Markov Random FieldabstractWe describe a novel framework for segmenting a time- and view-coherent foreground matte sequence from synchronised multiple view video. We construct a Markov Random Field (MRF) comprising links between super pixels corresponded across views, and links between super pixels and their constituent pixels. Texture, colour and disparity cues are incorporated to model foreground appearance. We solve using a multi-resolution iterative approach enabling an eight view high definition (HD) frame to be processed in less than a minute. Furthermore we incorporate a temporal diffusion process introducing a prior on the MRF using information propagated from previous frames, and a facility for optional user correction. The result is a set of temporally coherent mattes solved for simultaneously across views for each frame, exploiting similarities across views and time. Tinghuai Wang, John P. Collomosse, Adrian Hilton 0001 |
ICPR | 3 |
| 2014 | 4D video textures for interactive character appearanceabstractAbstract 4D Video Textures (4DVT) introduce a novel representation for rendering video‐realistic interactive character animation from a database of 4D actor performance captured in a multiple camera studio. 4D performance capture reconstructs dynamic shape and appearance over time but is limited to free‐viewpoint video replay of the same motion. Interactive animation from 4D performance capture has so far been limited to surface shape only. 4DVT is the final piece in the puzzle enabling video‐realistic interactive animation through two contributions: a layered view‐dependent texture map representation which supports efficient storage, transmission and rendering from multiple view video capture; and a rendering approach that combines multiple 4DVT sequences in a parametric motion space, maintaining video quality rendering of dynamic surface appearance whilst allowing high‐level interactive control of character motion and viewpoint. 4DVT is demonstrated for multiple characters and evaluated both quantitatively and through a user‐study which confirms that the visual quality of captured video is maintained. The 4DVT representation achieves >90% reduction in size and halves the rendering cost. Dan Casas, Marco Volino, John P. Collomosse, Adrian Hilton 0001 |
Comput. Graph. Forum | 4 |
| 2014 | Error analysis of photometric stereo with colour lights
Martin Klaudiny, Adrian Hilton 0001 |
Pattern Recognit. Lett. | 2 |
| 2013 | Evaluation of 3D Feature Descriptors for Multi-modal Data RegistrationabstractWe propose a framework for 2D/3D multi-modal data registration and evaluate 3D feature descriptors for registration of 3D datasets from different sources. 3D datasets of outdoor environments can be acquired using a variety of active and passive sensor technologies. Registration of these datasets into a common coordinate frame is required for subsequent modelling and visualisation. 2D images are converted into 3D structure by stereo or multiview reconstruction techniques and registered to a unified 3D domain with other datasets in a 3D world. Multi-modal datasets have different density, noise, and types of errors in geometry. This paper provides a performance benchmark for existing 3D feature descriptors across multi-modal datasets. This analysis highlights the limitations of existing 3D feature detectors and descriptors which need to be addressed for robust multi-modal data registration. We analyse and discuss the performance of existing methods in registering various types of datasets then identify future directions required to achieve robust multi-modal data registration. Hansung Kim 0001, Adrian Hilton 0001 |
3DV | 2 |
| 2013 | Shape and Pose Space Deformation for Subject Specific AnimationabstractIn this paper we present a framework for generating arbitrary human models and animating them realistically given a few intuitive parameters. Shape and pose space deformation (SPSD) is introduced as a technique for modeling subject specific pose induced deformations from whole-body registered 3D scans. By exploiting examples of different people in multiple poses we are able to realistically animate a novel subject by interpolating and extrapolating in a joint shape and pose parameter space. Our results show that we can produce plausible animations of new people and that greater detail is achieved by incorporating subject specific pose deformations. We demonstrate the application of SPSD to produce subject specific animation sequences driven by RGB-Z performance capture. Alexandros Neophytou, Adrian Hilton 0001 |
3DV | 2 |
| 2013 | Learning Part-Based Models for Animation from Surface Motion CaptureabstractSurface motion capture (Surf Cap) enables 3D reconstruction of human performance with detailed cloth and hair deformation. However, there is a lack of tools that allow flexible editing of Surf Cap sequences. In this paper, we present a Laplacian editing technique that constrains the mesh deformation to plausible surface shapes learnt from a set of examples. A part-Based representation of the mesh enables learning of surface deformation locally in the space of Laplacian coordinates, avoiding correlations between body parts while preserving surface details. This extends the range of animation with natural surface deformation beyond the whole-body poses present in the Surf Cap data. We illustrate successful use of our tool on three different characters. Margara Tejera, Adrian Hilton 0001 |
3DV | 2 |
| 2013 | Global Non-rigid Alignment of Surface SequencesabstractThis paper presents a general approach based on the shape similarity tree for non-sequential alignment across databases of multiple unstructured mesh sequences from non-rigid surface capture. The optimal shape similarity tree for non-rigid alignment is defined as the minimum spanning tree in shape similarity space. Non-sequential alignment based on the shape similarity tree minimises the total non-rigid deformation required to register all frames in a database into a consistent mesh structure with surfaces in correspondence. This allows alignment across multiple sequences of different motions, reduces drift in sequential alignment and is robust to rapid non-rigid motion. Evaluation is performed on three benchmark databases of 3D mesh sequences with a variety of complex human and cloth motion. Comparison with sequential alignment demonstrates reduced errors due to drift and improved robustness to large non-rigid deformation, together with global alignment across multiple sequences which is not possible with previous sequential approaches. Chris Budd, Peng Huang 0001, Martin Klaudiny, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 4 |
| 2013 | 3D Scene Reconstruction from Multiple Spherical Stereo Pairs
Hansung Kim 0001, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 2 |
| 2013 | Animation Control of Surface Motion CaptureabstractSurface motion capture (SurfCap) of actor performance from multiple view video provides reconstruction of the natural nonrigid deformation of skin and clothing. This paper introduces techniques for interactive animation control of SurfCap sequences which allow the flexibility in editing and interactive manipulation associated with existing tools for animation from skeletal motion capture (MoCap). Laplacian mesh editing is extended using a basis model learned from SurfCap sequences to constrain the surface shape to reproduce natural deformation. Three novel approaches for animation control of SurfCap sequences, which exploit the constrained Laplacian mesh editing, are introduced: 1) space–time editing for interactive sequence manipulation; 2) skeleton-driven animation to achieve natural nonrigid surface deformation; and 3) hybrid combination of skeletal MoCap driven and SurfCap sequence to extend the range of movement. These approaches are combined with high-level parametric control of SurfCap sequences in a hybrid surface and skeleton-driven animation control framework to achieve natural surface deformation with an extended range of movement by exploiting existing MoCap archives. Evaluation of each approach and the integrated animation framework are presented on real SurfCap sequences for actors performing multiple motions with a variety of clothing styles. Results demonstrate that these techniques enable flexible control for interactive animation with the natural nonrigid surface dynamics of the captured performance and provide a powerful tool to extend current SurfCap databases by incorporating new motions from MoCap sequences. Margara Tejera, Dan Casas, Adrian Hilton 0001 |
IEEE Trans. Cybern. | 3 |
| 2013 | Interactive Animation of 4D Performance CaptureabstractA 4D parametric motion graph representation is presented for interactive animation from actor performance capture in a multiple camera studio. The representation is based on a 4D model database of temporally aligned mesh sequence reconstructions for multiple motions. High-level movement controls such as speed and direction are achieved by blending multiple mesh sequences of related motions. A real-time mesh sequence blending approach is introduced, which combines the realistic deformation of previous nonlinear solutions with efficient online computation. Transitions between different parametric motion spaces are evaluated in real time based on surface shape and motion similarity. Four-dimensional parametric motion graphs allow real-time interactive character animation while preserving the natural dynamics of the captured performance. Dan Casas, Margara Tejera, Jean-Yves Guillemaut, Adrian Hilton 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2012 | Through-the-Lens Synchronisation for Heterogeneous Camera NetworksabstractCamera synchronisation involves the temporal alignment of a set of video sequences, independently acquired by two or more cameras. Accurate synchronisation is crucial for a wide variety of applications requiring multi-camera setups, ranging from 3D modelling of dynamic scenes (e.g., featuring a performance, or a sports event) to video surveillance and superresolution. Conventional synchronisation methods, which typically rely on hardware or audio signals, have practical limitations, imposing constraints on the size and the span of the network [2][1]. Through-the-lens synchronisation offers a robust and flexible way to synchronise a camera network from the content it generates. In this paper, we propose a bottom-up synchronisation algorithm to estimate a frame rate and an offset for each member of a network composed of 2 or more cameras. Our approach involves the computation of a relative synchronisation estimate between each camera pair, from which the absolute synchronisation parameters of the individual cameras are calculated (Figure 1). The algorithm can handle hybrid networks of static and moving cameras with different resolutions and frame rates, and does not require rigid objects, long trajectories or overlapping fields-of-view beyond 2 cameras. It needs a set of image features on the dynamic scene elements, and the geometric relation between the images (which can be obtained from the static background features). Relative Synchronisation: The frame indices of the jth camera (t j) with respect to those of ith (ti) is defined by the line Evren Imre, Adrian Hilton 0001 |
BMVC | 2 |
| 2012 | Towards Optimal Non-rigid Surface Tracking
Martin Klaudiny, Chris Budd, Adrian Hilton 0001 |
ECCV (4) | 3 |
| 2012 | 4D parametric motion graphs for interactive animationabstractA 4D parametric motion graph representation is presented for interactive animation from actor performance capture in a multiple camera studio. The representation is based on a 4D model database of temporally aligned mesh sequence reconstructions for multiple motions. High-level movement controls such as speed and direction are achieved by blending multiple mesh sequences of related motions. A real-time mesh sequence blending approach is introduced which combines the realistic deformation of previous non-linear solutions with efficient online computation. Transitions between different parametric motion spaces are evaluated in real-time based on surface shape and motion similarity. 4D parametric motion graphs allow real-time interactive character animation while preserving the natural dynamics of the captured performance. Dan Casas, Margara Tejera, Jean-Yves Guillemaut, Adrian Hilton 0001 |
I3D | 4 |
| 2012 | Parametric animation of performance-captured mesh sequencesabstractABSTRACT In this paper, we introduce an approach to high‐level parameterisation of captured mesh sequences of actor performance for real‐time interactive animation control. High‐level parametric control is achieved by non‐linear blending between multiple mesh sequences exhibiting variation in a particular movement. For example, walking speed is parameterised by blending fast and slow walk sequences. A hybrid non‐linear mesh sequence blending approach is introduced to approximate the natural deformation of non‐linear interpolation techniques whilst maintaining the real‐time performance of linear mesh blending. Quantitative results show that the hybrid approach gives an accurate real‐time approximation of offline non‐linear deformation. An evaluation of the approach shows good performance not only for entire meshes but also with specific mesh areas. Results are presented for single and multi‐dimensional parametric control of walking (speed/direction), jumping (height/distance) and reaching (height) from captured mesh sequences. This approach allows continuous real‐time control of high‐level parameters such as speed and direction whilst maintaining the natural surface dynamics of captured movement. Copyright © 2012 John Wiley & Sons, Ltd. Dan Casas, Margara Tejera, Jean-Yves Guillemaut, Adrian Hilton 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2012 | Outdoor Dynamic 3-D Scene ReconstructionabstractExisting systems for 3-D reconstruction from multiple view video use controlled indoor environments with uniform illumination and backgrounds to allow accurate segmentation of dynamic foreground objects. In this paper, we present a portable system for 3-D reconstruction of dynamic outdoor scenes that require relatively large capture volumes with complex backgrounds and nonuniform illumination. This is motivated by the demand for 3-D reconstruction of natural outdoor scenes to support film and broadcast production. Limitations of existing multiple view 3-D reconstruction techniques for use in outdoor scenes are identified. Outdoor 3-D scene reconstruction is performed in three stages: 1) 3-D background scene modeling using spherical stereo image capture; 2) multiple view segmentation of dynamic foreground objects by simultaneous video matting across multiple views; and 3) robust 3-D foreground reconstruction and multiple view segmentation refinement in the presence of segmentation and calibration errors. Evaluation is performed on several outdoor productions with complex dynamic scenes including people and animals. Results demonstrate that the proposed approach overcomes limitations of previous indoor multiple view reconstruction approaches enabling high-quality free-viewpoint rendering and 3-D reference models for production. Hansung Kim 0001, Jean-Yves Guillemaut, Takeshi Takai, Muhammad Sarim, Adrian Hilton 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2011 | Global temporal registration of multiple non-rigid surface sequencesabstractIn this paper we consider the problem of aligning multiple non-rigid surface mesh sequences into a single temporally consistent representation of the shape and motion. A global alignment graph structure is introduced which uses shape similarity to identify frames for inter-sequence registration. Graph optimisation is performed to minimise the total non-rigid deformation required to register the input sequences into a common structure. The resulting global alignment ensures that all input sequences are resampled with a common mesh structure which preserves the shape and temporal correspondence. Results demonstrate temporally consistent representation of several public databases of mesh sequences for multiple people performing a variety of motions with loose clothing and hair. Peng Huang 0001, Chris Budd, Adrian Hilton 0001 |
CVPR | 3 |
| 2011 | 4D Performance Modelling and AnimationabstractVisual reconstruction of dynamic events as 3D video, such as an actor performance or sports action, has advanced to the stage where it is possible to achieve free-viewpoint replay with a quality approaching the captured video. In this talk we present research going beyond replay to allow the creation of 4D models which support interactive animation control from captured performance whilst maintaining the realism of video. 4D models are constructed by alignment of reconstructed mesh sequences into a temporally coherent structure. Recent work has introduced a non-sequential approach to non-rigid mesh sequence alignment which constructs a shape similarity tree to align across a database of multiple sequences. This avoids problems of drift and tracking failure associated with sequential alignment approaches. Temporally aligned 4D models provide the basis for parameterisation of multiple related sequences to give continuous interactive movement control. Representation of multiple sequences in a 4D parametric motion graph enables transition between multiple motions to achieve interactive character animation. Adrian Hilton 0001 |
DS-RT | 1 |
| 2011 | A FACS valid 3D dynamic action unit database with applications to 3D dynamic morphable facial modelingabstractThis paper presents the first dynamic 3D FACS data set for facial expression research, containing 10 subjects performing between 19 and 97 different AUs both individually and in combination. In total the corpus contains 519 AU sequences. The peak expression frame of each sequence has been manually FACS coded by certified FACS experts. This provides a ground truth for 3D FACS based AU recognition systems. In order to use this data, we describe the first framework for building dynamic 3D morphable models. This includes a novel Active Appearance Model (AAM) based 3D facial registration and mesh correspondence scheme. The approach overcomes limitations in existing methods that require facial markers or are prone to optical flow drift. We provide the first quantitative assessment of such 3D facial mesh registration techniques and show how our proposed method provides more reliable correspondence. Darren Cosker, Eva Krumhuber, Adrian Hilton 0001 |
ICCV | 3 |
| 2011 | Temporal trimap propagation for video matting using inferential statisticsabstractThis paper introduces a statistical inference framework to temporally propagate trimap labels from sparsely defined key frames to estimate trimaps for the entire video sequence. Trimap is a fundamental requirement for digital image and video matting approaches. Statistical inference is coupled with Bayesian statistics to allow robust trimap labelling in the presence of shadows, illumination variation and overlap between the foreground and background appearance. Results demonstrate that trimaps are sufficiently accurate to allow high quality video matting using existing natural image matting algorithms. Quantitative evaluation against ground-truth demonstrates that the approach achieves accurate matte estimation with less amount of user interaction compared to the state-of-the-art techniques. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut |
ICIP | 2 |
| 2011 | Parametric Control of Captured Mesh Sequences for Real-Time Animation
Dan Casas, Margara Tejera, Jean-Yves Guillemaut, Adrian Hilton 0001 |
MIG | 4 |
| 2011 | Special issue on 3D imaging and modelling
Adrian Hilton 0001, Guy Godin, Chang Shu 0001, Takeshi Masuda 0001 |
Comput. Vis. Image Underst. | 1 |
| 2011 | Joint Multi-Layer Segmentation and Reconstruction for Free-Viewpoint Video Applications
Jean-Yves Guillemaut, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 2 |
| 2010 | Moving Camera Registration for Multiple Camera Setups in Dynamic ScenesabstractThis paper describes a method to register a moving (principal) camera, given a set of fully calibrated static cameras (witnesses) viewing a dynamic scene, a common scenario in broadcasting and film production. Our ultimate aim is to equip the existing free-viewpoint video algorithms with the ability to exploit any available moving cameras in generic dynamic scenes, and to facilitate 3D content production by augmented reality and stereoscopic rendering. Evren Imre, Jean-Yves Guillemaut, Adrian Hilton 0001 |
BMVC | 3 |
| 2010 | Stereoscopic content production of complex dynamic scenes using a wide-baseline monoscopic camera set-upabstractConventional stereoscopic video content production requires use of dedicated stereo camera rigs which is both costly and lacking video editing flexibility. In this paper, we propose a novel approach which only requires a small number of standard cameras sparsely located around a scene to automatically convert the monocular inputs into stereoscopic streams. The approach combines a probabilistic spatio-temporal segmentation framework with a state-of-the-art multi-view graph-cut reconstruction algorithm, thus providing full control of the stereoscopic settings at render time. Results with studio sequences of complex human motion demonstrate the suitability of the method for high quality stereoscopic content generation with minimum user interaction. Jean-Yves Guillemaut, Muhammad Sarim, Adrian Hilton 0001 |
ICIP | 3 |
| 2010 | PDE-based disparity estimation with occlusion and texture handling for accurate depth recovery from a stereo image pairabstractThis paper presents a novel PDE-based method for floating-point disparity estimation which produces smooth disparity fields with sharp object boundaries for surface reconstruction. In order to avoid the over-segmentation problem of image-driven structure tensor and the blurred boundary problem of field-driven tensor, we propose a new anisotropic diffusivity function controlled by image and disparity gradients. We also embed a bi-directional disparity matching term to control the data term in occluded regions. We evaluate the proposed method on data sets from the Middlebury benchmarking site and real data sets with ground-truth models scanned by a LIDAR sensor. Hansung Kim 0001, Adrian Hilton 0001 |
ICIP | 2 |
| 2010 | Natural image matting for multiple wide-baseline viewsabstractIn this paper we present a novel approach to estimate the alpha mattes of a foreground object captured by a wide-baseline circular camera rig provided a single key frame trimap. Bayesian inference coupled with camera calibration information are used to propagate high confidence trimaps labels across the views. Recent techniques have been developed to estimate an alpha matte of an image using multiple views but they are limited to narrow baseline views with low foreground variation. The proposed wide-baseline trimap propagation is robust to inter-view foreground appearance changes, shadows and similarity in foreground/background appearance for cameras with opposing views enabling high quality alpha matte extraction using any state-of-the-art image matting algorithm. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut, Takeshi Takai, Hansung Kim 0001 |
ICIP | 2 |
| 2010 | ACM multimedia 2010 workshop on 3D video processingabstractResearch on 3D video processing has gained a tremendous amount of momentum due to advances in video communications, broadcasting and entertainment technology (e.g., animation blockbusters like Avatar and Up). There is an increasing need for reliable technologies capable of visualizing 3-D content from viewpoints decided by the user; the 2010 football World Cup in South Africa has made very evident the need to replay crucial football footage from new viewpoints to decide whether the ball has or has not crossed the goal line. Remote videoconferencing prototypes are introducing a sense of presence into large- and small-scale (PC-based) systems alike by manipulating single and multiple video sequences to improve eye contact and place participants in convincing virtual spaces. All this, and more, is pushing the introduction of 3D services and the development of high-quality 3D displays to be available in a future which is drawing nearer and nearer. Oliver Schreer, Adrian Hilton 0001, Emanuele Trucco |
ACM Multimedia | 2 |
| 2010 | Shape Similarity for 3D Video Sequences of People
Peng Huang 0001, Adrian Hilton 0001, Jonathan Starck |
Int. J. Comput. Vis. | 2 |
| 2009 | Non-parametric Patch based Video MattingabstractIn computer vision, matting is the process of accurate foreground estimation in images and videos. In this paper we presents a novel patch based approach to video matting relying on non-parametric statistics to represent image variations in appearance. This overcomes the limitation of parametric algorithms which only rely on strong colour correlation between the nearby pixels. Initially we construct a clean background by utilising the foreground object’s movement across the background. For a given frame, a trimap is constructed using the background and the last frame’s trimap. A patch-based approach is used to estimate the foreground colour for every unknown pixel and finally the alpha matte is extracted. Quantitative evaluation shows that the technique performs better, in terms of the accuracy and the required user interaction, than the current state-of-the-art parametric approaches. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut |
BMVC | 2 |
| 2009 | Human motion synthesis from 3D videoabstractMultiple view 3D video reconstruction of actor performance captures a level-of-detail for body and clothing movement which is time-consuming to produce using existing animation tools. In this paper we present a framework for concatenative synthesis from multiple 3D video sequences according to user constraints on movement, position and timing. Multiple 3D video sequences of an actor performing different movements are automatically constructed into a surface motion graph which represents the possible transitions with similar shape and motion between sequences without unnatural movement artifacts. Shape similarity over an adaptive temporal window is used to identify transitions between 3D video sequences. Novel 3D video sequences are synthesized by finding the optimal path in the surface motion graph between user specified key-frames for control of movement, location and timing. The optimal path which satisfies the user constraints whilst minimizing the total transition cost between 3D video sequences is found using integer linear programming. Results demonstrate that this framework allows flexible production of novel 3D video sequences which preserve the detailed dynamics of the captured movement for an actress with loose clothing and long hair without visible artifacts. Peng Huang 0001, Adrian Hilton 0001, Jonathan Starck |
CVPR | 2 |
| 2009 | Robust graph-cut scene segmentation and reconstruction for free-viewpoint video of complex dynamic scenesabstractCurrent state-of-the-art image-based scene reconstruction techniques are capable of generating high-fidelity 3D models when used under controlled capture conditions. However, they are often inadequate when used in more challenging outdoor environments with moving cameras. In this case, algorithms must be able to cope with relatively large calibration and segmentation errors as well as input images separated by a wide-baseline and possibly captured at different resolutions. In this paper, we propose a technique which, under these challenging conditions, is able to efficiently compute a high-quality scene representation via graph-cut optimisation of an energy function combining multiple image cues with strong priors. Robustness is achieved by jointly optimising scene segmentation and multiple view reconstruction in a view-dependent manner with respect to each input camera. Joint optimisation prevents propagation of errors from segmentation to reconstruction as is often the case with sequential approaches. View-dependent processing increases tolerance to errors in on-the-fly calibration compared to global approaches. We evaluate our technique in the case of challenging outdoor sports scenes captured with manually operated broadcast cameras and demonstrate its suitability for high-quality free-viewpoint video. Jean-Yves Guillemaut, Joe Kilner, Adrian Hilton 0001 |
ICCV | 3 |
| 2009 | Graph-based foreground extraction in extended color spaceabstractWe propose a region-based method to extract semantic foreground regions from color video sequences with static backgrounds. First, we introduce a new distance measure for background subtraction which is robust against shadows. Then the foreground region is extracted with a graph-based region segmentation method considering background difference and spatial homogeneity. For efficient computation, the graph structure is optimized by the minimum spanning tree before segmentation. The main contribution is that the proposed algorithm improves on conventional approaches especially in strong shadow regions and does not require manual initialization. We have verified through experiments and comparison to state of the art methods that the proposed algorithm works well with various cameras and environment. Hansung Kim 0001, Adrian Hilton 0001 |
ICIP | 2 |
| 2009 | Non-parametric natural image mattingabstractNatural image matting is an extremely challenging image processing problem due to its ill-posed nature. It often requires skilled user interaction to aid definition of foreground and background regions. Current algorithms use these predefined regions to build local foreground and background colour models. In this paper we propose a novel approach which uses non-parametric statistics to model image appearance variations. This technique overcomes the limitations of previous parametric approaches which are purely colour-based and thereby unable to model natural image structure. The proposed technique consists of three successive stages: (i) background colour estimation, (ii) foreground colour estimation, (iii) alpha estimation. Colour estimation uses patch-based matching techniques to efficiently recover the optimum colour by comparison against patches from the known regions. Quantitative evaluation against ground truth demonstrates that the technique produces better results and successfully recovers fine details such as hair where many other algorithms fail. Muhammad Sarim, Adrian Hilton 0001, Jean-Yves Guillemaut, Hansung Kim 0001 |
ICIP | 2 |
| 2009 | Objective quality assessment in free-viewpoint video production
Joe Kilner, Jonathan Starck, Jean-Yves Guillemaut, Adrian Hilton 0001 |
Signal Process. Image Commun. | 4 |
| 2009 | The Multiple-Camera 3-D Production StudioabstractMultiple-camera systems are currently widely used in research and development as a means of capturing and synthesizing realistic 3-D video content. Studio systems for 3-D production of human performance are reviewed from the literature, and the practical experience gained in developing prototype studios is reported across two research laboratories. System design should consider the studio backdrop for foreground matting, lighting for ambient illumination, camera acquisition hardware, the camera configuration for scene capture, and accurate geometric and photometric camera calibration. A ground-truth evaluation is performed to quantify the effect of different constraints on the multiple-camera system in terms of geometric accuracy and the requirement for high-quality view synthesis. As changing camera height has only a limited influence on surface visibility, multiple-camera sets or an active vision system may be required for wide area capture, and accurate reconstruction requires a camera baseline of 25deg, and the achievable accuracy is 5-10-mm at current camera resolutions. Accuracy is inherently limited, and view-dependent rendering is required for view synthesis with sub-pixel accuracy where display resolutions match camera resolutions. The two prototype studios are contrasted and state-of-the-art techniques for 3-D content production demonstrated. Jonathan Starck, Atsuto Maki, Shohei Nobuhara, Adrian Hilton 0001, Takashi Matsuyama |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2008 | Model-based human shape reconstruction from multiple views
Jonathan Starck, Adrian Hilton 0001 |
Comput. Vis. Image Underst. | 2 |
| 2007 | 2D face pose normalisation using a 3D morphable modelabstractThe ever growing need for improved security, surveillance and identity protection, calls for the creation of evermore reliable and robust face recognition technology that is scalable and can be deployed in all kinds of environments without compromising its effectiveness. In this paper we study the impact that pose correction has on the performance of 2D face recognition. To measure the effect, we use a state of the art 2D recognition algorithm. The pose correction is performed by means of 3D morphable model. Our results on the non frontal XM2VTS database showed that pose correction can improve recognition rates up to 30%. Jose Rafael Tena, Raymond S. Smith, Miroslav Hamouz, Josef Kittler, Adrian Hilton 0001, John Illingworth |
AVSS | 5 |
| 2007 | Correspondence labelling for wide-timeframe free-form surface matchingabstractThis paper addresses the problem of estimating dense correspondence between arbitrary frames from captured sequences of shape and appearance for surfaces undergoing free-form deformation. Previous techniques require either a prior model, limiting the range of surface deformations, or frame-to-frame surface tracking which suffers from stabilisation problems over complete motion sequences and does not provide correspondence between sequences. The primary contribution of this paper is the introduction of a system for wide-timeframe surface matching without the requirement for a prior model or tracking. Deformation- invariant surface matching is formulated as a locally isometric mapping at a discrete set of surface points. A set of feature descriptors are presented that are invariant to isometric deformations and a novel MAP-MRF framework is presented to label sparse-to-dense surface correspondence, preserving the relative distribution of surface features while allowing for changes in surface topology. Performance is evaluated on challenging data from a moving person with loose clothing. Ground-truth feature correspondences are manually marked and the recall-accuracy characteristic is quantified in matching. Results demonstrate an improved performance compared to non-rigid point-pattern matching using robust matching and graph-matching using relaxation labelling, with successful matching achieved across wide variations in human body pose and surface topology. Jonathan Starck, Adrian Hilton 0001 |
ICCV | 2 |
| 2007 | Visual analysis of lip coarticulation in VCV utterancesabstractThis paper presents an investigation of the visual variation on the bilabial plosive consonant /p/ in three coarticulation contexts. The aim is to provide detailed ensemble analysis to assist coarticulation modelling in visual speech synthesis. The underlying dynamics of labeled visual speech units, represented as lip shape, from symmetric VCV utterances, is investigated. Variation in lip dynamics is quantitively and qualitatively analyzed. This analysis shows that there are statistically significant differences in both the lip shape and trajectory during coarticulation. Aseel Turkmani, Adrian Hilton 0001, Philip J. B. Jackson, James D. Edge |
INTERSPEECH | 2 |
| 2006 | A Validated Method for Dense Non-rigid 3D Face RegistrationabstractDeformable surface fitting methods have been widely used to establish dense correspondence across different 3D objects of the same class. Dense correspondence is a critical step in constructing morphable face models for face recognition. In this paper a mainstream method for constructing dense correspondences is evaluated on 912 3D face scans from the Face Recognition Grand Challenge FRGC V1 database. A number of modifications to the standard deformable surface approach are introduced to overcome limitations identified in the evaluation. Proposed modifications include multi-resolution fitting, adaptive correspondence search range and enforcing symmetry constraints. The modified deformable surface approach is validated on the 912 FRGC 3D face scans and is shown to overcome limitations of the standard approach which resulted in gross fitting errors. The modified approach halves the rms fitting error with 98% of points within 0.5mm of their true position compared to 67% with the standard approach. Jose Rafael Tena, Miroslav Hamouz, Adrian Hilton 0001, John Illingworth |
AVSS | 3 |
| 2006 | Volumetric Stereo with Silhouette and Feature ConstraintsabstractThis paper presents a novel volumetric reconstruction technique that combines shape-from-silhouette with stereo photo-consistency in a global optimisation that enforces feature constraints across multiple views. Human shape reconstruction is considered where extended regions of uniform appearance, complex self-occlusions and sparse feature cues represent a challenging problem for conventional reconstruction techniques. A unified approach is introduced to first reconstruct the occluding contours and left-right consistent edge contours in a scene and then incorporate these contour constraints in a global surface optimisation using graph-cuts. The proposed technique maximises photo-consistency on the surface, while satisfying silhouette constraints to provide shape in the presence of uniform surface appearance and edge feature constraints to align key image features across views. 1 Jonathan Starck, Gregor Miller, Adrian Hilton 0001 |
BMVC | 3 |
| 2006 | Learnt inverse kinematics for animation synthesis
Eng-Jon Ong, Adrian Hilton 0001 |
Graph. Model. | 2 |
| 2006 | Modeling people: Vision-based understanding of a person's shape, appearance, movement, and behaviour
Adrian Hilton 0001, Pascal Fua, Rémi Ronfard |
Comput. Vis. Image Underst. | 1 |
| 2006 | A survey of advances in vision-based human motion capture and analysis
Thomas B. Moeslund, Adrian Hilton 0001, Volker Krüger |
Comput. Vis. Image Underst. | 2 |
| 2006 | Viewpoint invariant exemplar-based 3D human tracking
Eng-Jon Ong, Antonio S. Micilotta, Richard Bowden, Adrian Hilton 0001 |
Comput. Vis. Image Underst. | 4 |
| 2005 | 3D Assisted 2D Face Recognition: Methodology
Josef Kittler, Miroslav Hamouz, Jose Rafael Tena, Adrian Hilton 0001, John Illingworth, M. Ruiz |
CIARP | 4 |
| 2005 | Spherical Matching for Temporal Correspondence of Non-Rigid SurfacesabstractThis paper introduces spherical matching to estimate dense temporal correspondence of non-rigid surfaces with genus-zero topology. The spherical domain gives a consistent 1D parameterization of non-rigid surfaces for matching. Non-rigid 3D surface correspondence is formulated as the recovery of a bijective mapping between two surfaces in the 2D domain. Formulating matching as a 2D bijection guarantees a continuous one-to-one surface correspondence without overfolding. This overcomes limitations of direct estimation of non-rigid surface correspondence in the 3D domain. A multiple resolution coarse-to-fine algorithm is introduced to robustly estimate the dense correspondence which minimizes the disparity in shape and appearance between two surfaces. Spherical matching is applied to derive the temporal correspondence between non-rigid surfaces reconstructed at successive frames from multiple view video sequences of people. Dense surface correspondence is recovered across complete motion sequences for both textured and uniform regions, without the requirement for a prior model of human shape or kinematics structure for tracking Jonathan Starck, Adrian Hilton 0001 |
ICCV | 2 |
| 2005 | Virtual view synthesis of people from multiple view video sequences
Jonathan Starck, Adrian Hilton 0001 |
Graph. Model. | 2 |
| 2005 | Scene modelling from sparse 3D data
Adrian Hilton 0001 |
Image Vis. Comput. | 1 |
| 2003 | Simultaneous Pose Estimation of Multiple People using Multiple-View Cues with Hierarchical SamplingabstractWe present a novel method for dynamic estimation of pose of multiple people using multiple video cameras. Tracking is performed using a model-based approach and a set of cues which exploit both shape and colour information. For shape we propose a fast line-search method to incorporate multi-view constraints without the computational overhead of a voxel representation. The tracking algorithm is a new hierarchical stochastic sampling scheme. Results are presented using natural movements of up to two people sharing the same capture volume. Tracking is shown to be robust over a range of natural movements, including considerable occlusion. Processing times for tracking two people are at least as short as those for tracking one person using other stochastic schemes. The high performance and efficiency are attributed to the hierarchical search method and the accuracy of the cues in identifying suitable poses. 1 Joel R. Mitchelson, Adrian Hilton 0001 |
BMVC | 2 |
| 2003 | Model-Based Multiple View Reconstruction of PeopleabstractThis paper presents a framework to reconstruct a scene captured in multiple camera views based on a prior model of the scene geometry. The framework is applied to the capture of animated models of people. A multiple camera studio is used to simultaneously capture a moving person from multiple viewpoints. A humanoid computer graphics model is animated to match the pose at each time frame. Constrained optimisation is then used to recover the multiple view correspondence from silhouette, stereo and feature cues, updating the geometry and appearance of the model. The key contribution of this paper is a model-based computer vision framework for the reconstruction of shape and appearance from multiple views. This is compared to current model-free approaches for multiple view scene capture. The technique demonstrates improved scene reconstruction in the presence of visual ambiguities and provides the means to capture a dynamic scene with a consistent model that is instrumented with an animation structure to edit the scene dynamics or to synthesise new content. Jonathan Starck, Adrian Hilton 0001 |
ICCV | 2 |
| 2003 | Computer vision for human modelling and analysis
Adrian Hilton 0001 |
Mach. Vis. Appl. | 1 |
| 2003 | Animated statues
Jonathan Starck, Gordon Collins, Raymond S. Smith, Adrian Hilton 0001, John Illingworth |
Mach. Vis. Appl. | 4 |
| 2002 | A relaxation algorithm for real-time multiple view 3D-tracking
Adrian Hilton 0001, John Illingworth |
Image Vis. Comput. | 2 |
| 2001 | Human Shape Estimation in a Multi-Camera StudioabstractThis paper addresses the problem of estimating the shape of an actor in a multi-camera studio for arbitrarily positioned cameras and arbitrary human pose. We adopt a seamless articulated mesh model and introduce a novel shape matching technique to automatically transform the projected shape of the model to match multiple captured image silhouettes. Our approach treats the projected mesh as a deformable model and constrains the model to follow a smooth shape transformation to match each image silhouette. Multiple 2D transformations are integrated in 3D to update the shape of the model to match the actor. We assess the technique using virtual views generated for 3D scanned human data-sets and present preliminary results in a studio. 1 Jonathan Starck, Adrian Hilton 0001, John Illingworth |
BMVC | 2 |
| 2001 | Modeling People Toward Vision-Based Understanding of a Person's Shape, Appearance, and Movement
Adrian Hilton 0001, Pascal Fua |
Comput. Vis. Image Underst. | 1 |
| 2001 | Layered animation of captured data
Adrian Hilton 0001, Raymond S. Smith, John Illingworth |
Vis. Comput. | 2 |
| 2000 | A Statistical Geometric Framework for Reconstruction of Scene ModelsabstractThis paper addresses the problem of reconstructing surface models of indoor scenes from sparse 3D scene structure captured from N camera views. Sparse 3D measurements of real scenes are readily estimated from image sequences using structure-from-motion techniques. Currently there is no general method for reconstruction of 3D models of arbitrary scenes from sparse data. We previously introduced an algorithm for recursive integration of sparse 3D structure to obtain a consistent model. In this paper we focus on incorporating uncertainty information into model to achieve reliable reconstruction of real-scenes in the presence of noise. A statistical geometric framework is described that provides a unified approach to probabilistic scene reconstruction from sparse or even dense 3D scene structure. Anastasios Manessis, Adrian Hilton 0001, Philip F. McLauchlan, Phil Palmer |
BMVC | 2 |
| 2000 | Error Propogation from Camera Motion to Epipolar ConstraintabstractThis work investigates the propagation of errors from the camera motion to the epipolar constraint. A relation between the perturbation of the motion parameters and the error in the epipolar constraint is derived. Based on this relation, the sensitivity of the motion parameters to the epipolar constraint is characterised, and a constraint on the allowed perturbation in a motion parameter in response to a threshold for the produced error in the epipolar constraint is determined. The presented error propagation model is useful to vision systems such as a mobile robot where the camera motion is provided by an odometry sensor. Experimental results on real images are presented. Xinquan Shen, Phil Palmer, Philip F. McLauchlan, Adrian Hilton 0001 |
BMVC | 4 |
| 2000 | Layered Animation using Displacement MapsabstractThis paper presents a layered animation framework which uses displacement maps for efficient representation and animation of highly detailed surfaces. The model consists of three layers: a skeleton; low-resolution control model; and a displacement map image. The novel aspects of this approach are an automatic closed-form solution for displacement map generation and animation of the layered displacement map model. This approach provides an efficient representation of complex geometry which allows realistic deformable animation with multiple levels-of-detail. Raymond S. Smith, Adrian Hilton 0001, John Illingworth |
CA | 3 |
| 2000 | Reconstruction of Scene Models from Sparse 3D StructureabstractIn this paper we present a geometric theory for reconstruction of surface models from sparse 3D data captured from N camera views which are consistent with the data visibility. Sparse 3D measurements of real scenes are readily estimated from image sequences using structure-from-motion techniques. Currently there is no general method for reconstruction of 3D models of arbitrary scenes from sparse data. We introduce an algorithm for recursive integration of sparse 3D structure to obtain a consistent model. This algorithm is shown to converge to the real scene structure as the number of views increases and to have a computational cost which is linear in the number of views. Results are presented for real and synthetic image sequences which demonstrate correct reconstruction for scenes containing significant occlusions. Anastasios Manessis, Adrian Hilton 0001, Phil Palmer, Philip F. McLauchlan, Xinquan Shen |
CVPR | 2 |
| 2000 | Geometric fusion for a hand-held 3D sensor
Adrian Hilton 0001, John Illingworth |
Mach. Vis. Appl. | 1 |
| 2000 | Whole-body modelling of people from multiview images to populate virtual worlds
Adrian Hilton 0001, Daniel J. Beresford, Thomas Gentils, Raymond S. Smith, John Illingworth |
Vis. Comput. | 1 |
| 1999 | Virtual People: Capturing Human Models to Populate Virtual WorldsabstractA new technique is introduced for automatically building recognisable moving 3D models of individual people. Realistic modelling of people is essential for advanced multimedia, augmented reality and immersive virtual reality. Current systems for whole-body model capture are based on active 3D sensing to measure the shape of the body surface. Such systems are prohibitively expensive and do not enable capture of high-quality photo-realistic colour. This results in geometrically accurate but unrealistic human models. The goal of this research is to achieve automatic low cost modelling of people suitable for personalised avatars to populate virtual worlds. A model based approach is presented for automatic reconstruction of recognisable avatars from a set of low cost colour images of a person taken from four orthogonal views. A generic 3D human model represents both the human shape and kinematic joint structure. The shape of a specific person is captured by mapping 2D silhouette information from the orthogonal view colour images onto the generic 3D model. Colour texture mapping is achieved by projecting the set of images onto the deformed 3D model. This results in the capture of a recognisable 3D facsimile of an individual person suitable for articulated movement in a virtual world. The system is low cost, requires single shot capture, is reliable for large variations in shape and size and can cope with clothing of moderate complexity. Adrian Hilton 0001, Daniel J. Beresford, Thomas Gentils, Raymond S. Smith |
CA | 1 |
| 1998 | Implicit Surface-Based Geometric Fusion
Adrian Hilton 0001, Andrew J. Stoddart, John Illingworth, Terry Windeatt |
Comput. Vis. Image Underst. | 1 |
| 1998 | Estimating pose uncertainty for surface registration
Andrew J. Stoddart, S. Lemke, Adrian Hilton 0001, T. Renn |
Image Vis. Comput. | 3 |
| 1997 | Progress in arbitrary topology deformable surfaces
A. Saminathan, Andrew J. Stoddart, Adrian Hilton 0001, John Illingworth |
BMVC | 3 |
| 1996 | Estimating Pose Uncertainty for Surface RegistrationabstractAccurate registration of surfaces is a common problem in computer vision. Several algorithms exist to refine an approximate value for the pose to an accurate value. They are all more or less variants of the Iterated Closest Point algorithm of Besl and McKay (1992). Up to now the problem of determining the uncertainty in the pose estimate thus obtained has not been addressed in detail. In this paper we present a framework in which to quantify the uncertainty in pose. We introduce a new parameter called the registration index to give a simple means of quantifying the pose errors one might expect when registering a particular shape. 1 Introduction To appear at British Machine Vision Conference, Edinburgh, UK. (1996). The registration of surfaces and curves can be broken down into two separate tasks. A matching task in which an approximate estimate of the pose is obtained. This might also involve a large database of model shapes. Then a registration task in which an accurate po... Andrew J. Stoddart, S. Lemke, Adrian Hilton 0001, T. Renn |
BMVC | 3 |
| 1996 | Reliable Surface Reconstructiuon from Multiple Range Images
Adrian Hilton 0001, Andrew J. Stoddart, John Illingworth, Terry Windeatt |
ECCV (1) | 1 |
| 1996 | Marching triangles: range image fusion for complex object modellingabstractA new surface based approach to implicit surface polygonisation is introduced. This is applied to the reconstruction of 3D surface models of complex objects from multiple range images. Geometric fusion of multiple range images into an implicit surface representation was presented in previous work. This paper introduces an efficient algorithm to reconstruct a triangulated model of a manifold implicit surface, a local 3D constraint is derived which defines the Delaunay surface triangulation of a set of points on a manifold surface in 3D space. The 'marching triangles' algorithm uses the local 3D constraint to reconstruct a Delaunay triangulation of an arbitrary topology manifold surface. Computational and representational costs are both a factor of 3-5 lower than previous volumetric approaches such as marching cubes. Adrian Hilton 0001, Andrew J. Stoddart, John Illingworth, Terry Windeatt |
ICIP (2) | 1 |
| 1996 | Registration of multiple point setsabstractRegistering 3D point sets subject to rigid body motion is a common problem in computer vision. The optimal transformation is usually specified to be the minimum of a weighted least squares cost. The case of 2 point sets has been solved by several authors using analytic methods such as SVD. In this paper we present a numerical method for solving the problem when there are more than 2 point sets. Although of general applicability the new method is particularly aimed at the multiview surface registration problem. To date almost all authors have registered only two point sets at a time. This approach discards information and we show in quantitative terms the errors caused. Andrew J. Stoddart, Adrian Hilton 0001 |
ICPR | 2 |
| 1995 | Statistics of surface curvature estimates
Adrian Hilton 0001, John Illingworth, Terry Windeatt |
Pattern Recognit. | 1 |
| 1994 | SLIME: A new deformable surfaceabstractDeformable surfaces have many applications in surface reconstruction, tracking and segmentation of range or volumetric data. Many existing deformable surfaces connect control points in a predefined and inflexible way. This means that the surface topology is fixed in advance, and also imposes severe limitations on how a surface can be described. For example a rectangular grid of control points cannot be evenly distributed over a sphere, and singularities occur at the poles. In this paper we introduce a new (G 1 continuous) deformable surface. In contrast to other methods this method can represent a surface of arbitrary topology, and do so in an efficient way. The method is based on a generalization of biquadratic B-splines, and has a comparable computational cost to methods based on traditional tensor product B-splines. 1 Andrew J. Stoddart, Adrian Hilton 0001, John Illingworth |
BMVC | 2 |
| 1994 | Statistics of surface curvature estimatesabstractReliable curvature estimation is an important goal in image analysis to provide viewpoint independent cues for shape classification. This paper presents a model of the relationship between the variance of curvature estimates and the image noise. Agreement to within 10% is obtained for 3D range data. Previous models have only provided qualitative agreement with experimental observations. A perturbation error analysis is performed on the local least square surface fitting algorithm which is commonly used to obtain partial derivative estimates in the presence of noise. Adrian Hilton 0001, John Illingworth, Terry Windeatt |
ICPR (1) | 1 |