Lourdes Agapito

dblp:70/3697 · also Lourdes de Agapito · DBLP profile ↗
← Back
79ranked-venue papers
6as first author
22since 2021 · last 2025
0000-0002-6947-1092ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 4 first-author · 18 since 2021Systems, architecture and hardware · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 3D Reconstruction with Spatial Memory
abstract
We present Spann3R, a novel approach for dense 3D reconstruction from ordered or unordered image collections. Built on the DUSt3R paradigm, Spann3R uses a transformer-based architecture to directly regress pointmaps from images without any prior knowledge of the scene or camera parameters. Unlike DUSt3R, which pre-dicts per image-pair pointmaps expressed in a local coordinate frame, Spann3R predicts per-image pointmaps expressed in a global coordinate system, thus eliminating the need for optimization-based global alignment. The key idea behind Spann3R is to manage an external spa-tial memory that learns to keep track of all previous relevant 3D information. Spann3R then queries this spatial memory to predict the 3D structure of the next frame in a global coordinate system. Taking advantage of DUSt3R's pre-trained weights, and further fine-tuning on a subset of datasets, Spann3R shows competitive performance and generalization ability on various unseen datasets and can process ordered image collections in real-time. Project page: https://hengyiwang.github.io/projects/spanner
Hengyi Wang, Lourdes Agapito
3DV2
2025 Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors
abstract
We present Pow3R, a novel large 3D vision regression model that is highly versatile in the input modalities it accepts. Unlike previous feed-forward models that lack any mechanism to exploit known camera or scene priors at test time, Pow3R incorporates any combination of auxiliary information such as intrinsics, relative pose, dense or sparse depth, alongside input images, within a single network. Building upon the recent DUSt3R paradigm, a transformer-based architecture that leverages powerful pre-training, our lightweight and versatile conditioning acts as additional guidance for the network to predict more accurate estimates when auxiliary information is available. During training we feed the model with random subsets of modalities at each iteration, which enables the model to operate under different levels of known priors at test time. This in turn opens up new capabilities, such as performing inference in native image resolution, or point-cloud completion. Our experiments on 3D reconstruction, depth completion, multi-view depth prediction, multi-view stereo, and multi-view pose estimation tasks yield state-of-the-art results and confirm the effectiveness of Pow3R at exploiting all available information. The project webpage is https://europe.naverlabs.com/pow3r.
Wonbong Jang, Philippe Weinzaepfel, Vincent Leroy 0003, Lourdes Agapito, Jérôme Revaud
CVPR4
2025 Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
abstract
Rapid advancements in text-to-3D generation require robust and scalable evaluation metrics that align closely with human judgment, a need unmet by current metrics such as PSNR and CLIP, which require ground-truth data or focus only on prompt fidelity. To address this, we introduce Gen3DEval, a novel evaluation framework that leverages vision large language models (vLLMs) specifically fine-tuned for 3D object quality assessment. Gen3DEval evaluates text fidelity, appearance, and surface quality by analyzing 3D surface normals, without requiring ground-truth comparisons, bridging the gap between automated metrics and user preferences. Compared to state-of-the-art task-agnostic models, Gen3DEval demonstrates superior performance in user-aligned evaluations, placing it as a comprehensive and accessible benchmark for future research on text-to-3D generation. The project page can be found here: https://shalini-maiti.github.io/gen3deval.github.io/.
Shalini Maiti, Lourdes Agapito, Filippos Kokkinos
CVPR2
2025 BillBoard Splatting (BBSplat): Learnable Textured Primitives for Novel View Synthesis
abstract
We present billboard Splatting (BBSplat) - a novel approach for novel view synthesis based on textured geometric primitives. BBSplat represents the scene as a set of optimizable textured planar primitives with learnable RGB textures and alpha-maps to control their shape. BBSplat primitives can be used in any Gaussian Splatting pipeline as drop-in replacements for Gaussians. The proposed primitives close the rendering quality gap between 2D and 3D Gaussian Splatting (GS), enabling the accurate extraction of 3D mesh as in the 2DGS framework. Additionally, the explicit nature of planar primitives enables the use of the ray-tracing effects in rasterization. Our novel regularization term encourages textures to have a sparser structure, enabling an efficient compression that leads to a reduction in the storage space of the model up to x17 times compared to 3DGS. Our experiments show the efficiency of BBSplat on standard datasets of real indoor and outdoor scenes such as Tanks&Temples, DTU, and Mip-NeRF-360. Namely, we achieve a state-of-the-art PSNR of 29.72 for DTU at Full HD resolution.
David Svitov, Pietro Morerio, Lourdes Agapito, Alessio Del Bue
ICCV3
2025 Semantic Cross-Pose Correspondence from a Single Example
abstract
This article focuses on predicting how an object can be transformed to a semantically meaningful pose relative to another object, given only one or few examples. Current pose correspondence methods rely on vast 3D object datasets and do not actively consider semantic information, which limits the objects to which they can be applied. We present a novel method for learning cross-object pose correspondence. The proposed method detects interacting object parts, performs one-shot part correspondence, and uses geometric and visual-semantic features. Given one example of two objects posed relative to each other, the model can learn how to transfer the demonstrated relations to unseen object instances. Supplementary details can be found at https://sites.google.com/view/semantic-pose-correspondence
Denis Hadjivelichkov, Sicelukwanda Zwane, Marc Peter Deisenroth, Lourdes Agapito, Dimitrios Kanoulas
ICRA4
2024 DynamicSurf: Dynamic Neural RGB-D Surface Reconstruction With an Optimizable Feature Grid
abstract
We propose DynamicSurf, a model-free neural implicit surface reconstruction method for high-fidelity 3D modelling of non-rigid surfaces from monocular RGB-D video. To cope with the lack of multi-view cues in monocular sequences of deforming surfaces, one of the most challenging settings for 3D reconstruction, DynamicSurf exploits depth, surface normals, and RGB losses to improve reconstruction fidelity and optimisation time. DynamicSurf learns a neural deformation field that maps a canonical representation of the surface geometry to the current frame. We depart from current neural non-rigid surface reconstruction models by designing the canonical representation as a learned feature grid which leads to faster and more accurate surface reconstruction than competing approaches that use a single MLP. We demonstrate DynamicSurf on public datasets and show that it can optimize sequences of varying frames with 6× speedup over pure MLP-based approaches while achieving comparable results to the state-of-the-art methods.11Project is available at https://mirgahney.github.io//DynamicSurf.io/.
Mirgahney Mohamed, Lourdes Agapito
3DV2
2024 HAHA: Highly Articulated Gaussian Human Avatars with Textured Mesh Prior
David Svitov, Pietro Morerio, Lourdes Agapito, Alessio Del Bue
ACCV (9)3
2024 MonoNPHM: Dynamic Head Reconstruction from Monocular Videos
abstract
We present Monocular Neural Parametric Head Models (MonoNPHM) for dynamic 3D head reconstructions from monocular RGB videos. To this end, we propose a latent appearance space that parameterizes a texture field on top of a neural parametric model. We constrain predicted color values to be correlated with the underlying geometry such that gradients from RGB effectively influence latent geometry codes during inverse rendering. To increase the representational capacity of our expression space, we augment our backward deformation field with hyper-dimensions, thus improving color and geometry representation in topologically challenging expressions. Using MonoNPHM as a learned prior, we approach the task of 3D head reconstruction using signed distance field based volumetric rendering. By numerically inverting our backward deformation field, we incorporated a landmark loss using facial anchor points that are closely tied to our canonical geometry representation. To evaluate the task of dynamic face reconstruction from monocular RGB videos we record 20 challenging Kinect sequences under casual conditions. MonoNPHM outper-forms all baselines with a significant margin, and makes an important step towards easily accessible neural parametric face models through RGB tracking.
Simon Giebenhain, Tobias Kirschstein, Markos Georgopoulos, Martin Rünz, Lourdes Agapito, Matthias Nießner
CVPR5
2024 NViST: In the Wild New View Synthesis from a Single Image with Transformers
abstract
We propose NViST, a transformer-based model for efficient and generalizable novel-view synthesis from a single image for real-world scenes. In contrast to many methods that are trained on synthetic data, object-centred scenarios, or in a category-specific manner, NViST is trained on MVImgNet, a large-scale dataset of casually-captured real-world videos of hundreds of object categories with diverse backgrounds. NViST transforms image inputs directly into a radiance field, conditioned on camera parameters via adaptive layer normalisation. In practice, NViST exploits fine-tuned masked autoencoder (MAE) features and translates them to 3D output tokens via cross-attention, while addressing occlusions with self-attention. To move away from object-centred datasets and enable full scene synthesis, NViST adopts a 6-DOF camera pose model and only requires relative pose, dropping the need for canonicalization of the training data, which removes a substantial barrier to it being used on casually captured datasets. We show results on unseen objects and categories from MVImgNet and even generalization to casual phone captures. We conduct qualitative and quantitative evaluations on MVImgNet and ShapeNet to show that our model represents a step forward towards enabling true in-the-wild generalizable novel-view synthesis from a single image. Project webpage: https://wbjang.github.io/nvist_webpage.
Wonbong Jang, Lourdes Agapito
CVPR2
2024 MorpheuS: Neural Dynamic $360^{\circ}$ Surface Reconstruction from Monocular RGB-D Video
abstract
Neural rendering has demonstrated remarkable success in dynamic scene reconstruction. Thanks to the expressiveness of neural representations, prior works can accurately capture the motion and achieve high-fidelity reconstruction of the target object. Despite this, real-world video sce-narios often feature large unobserved regions where neural representations struggle to achieve realistic completion. To tackle this challenge, we introduce MorpheuS, a framework for dynamic$360^\circ$surface reconstruction from a casually captured RGB-D video. Our approach models the target scene as a canonical field that encodes its geometry and appearance, in conjunction with a defor-mation field that warps points from the current frame to the canonical space. We leverage a view-dependent diffusion prior and distill knowledge from it to achieve realistic completion of unobserved regions. Experimental results on various real-world and synthetic datasets show that our method can achieve high-fidelity 360° surface reconstruction of a deformable object from a monocular RGB-D video. Project page: https: / /hengyi wang. gi thub. io/ pro jects/morpheus.
Hengyi Wang, Jingwen Wang 0005, Lourdes Agapito
CVPR3
2024 RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation
abstract
For robots to be useful outside labs and specialized factories we need a way to teach them new useful behaviors quickly. Current approaches lack either the generality to onboard new tasks without task-specific engineering, or else lack the data-efficiency to do so in an amount of time that enables practical use. In this work we explore dense tracking as a representational vehicle to allow faster and more general learning from demonstration. Our approach utilizes Track-Any-Point (TAP) models to isolate the relevant motion in a demonstration, and parameterize a low-level controller to reproduce this motion across changes in the scene configuration. We show this results in robust robot policies that can solve complex object-arrangement tasks such as shape-matching, stacking, and even full path-following tasks such as applying glue and sticking objects together, all from demonstrations that can be collected in minutes.
Mel Vecerík, Carl Doersch, Yi Yang 0007, Todor Davchev, Yusuf Aytar, Raia Hadsell, Lourdes Agapito, Jonathan Scholz
ICRA8
2024 NPGA: Neural Parametric Gaussian Avatars
Simon Giebenhain, Tobias Kirschstein, Martin Rünz, Lourdes Agapito, Matthias Nießner
SIGGRAPH Asia4
2023 Learning Neural Parametric Head Models
abstract
We propose a novel 3D morphable model for complete human heads based on hybrid neural fields. At the core of our model lies a neural parametric representation that disentangles identity and expressions in disjoint latent spaces. To this end, we capture a person's identity in a canonical space as a signed distance field (SDF), and model facial expressions with a neural deformation field. In addition, our representation achieves high-fidelity local detail by introducing an ensemble of local fields centered around facial anchor points. To facilitate generalization, we train our model on a newly-captured dataset of over 3700 head scans from 203 different identities using a custom high-end 3D scanning setup. Our dataset significantly exceeds comparable existing datasets, both with respect to quality and completeness of geometry, averaging around 3.5M mesh faces per scan11We will publicly release our dataset along with a public benchmark for both neural head avatar construction as well as an evaluation on a hidden test-set for inference-time fitting.. Finally, we demonstrate that our approach outperforms state-of-the-art methods in terms of fitting error and reconstruction quality.
Simon Giebenhain, Tobias Kirschstein, Markos Georgopoulos, Martin Rünz, Lourdes Agapito, Matthias Nießner
CVPR5
2023 Co-SLAM: Joint Coordinate and Sparse Parametric Encodings for Neural Real-Time SLAM
abstract
We present Co-SLAM, a neural RGB-D SLAM system based on a hybrid representation, that performs robust camera tracking and high-fidelity surface reconstruction in real time. Co-SLAM represents the scene as a multi-resolution hash-grid to exploit its high convergence speed and ability to represent high-frequency local features. In addition, Co-SLAM incorporates one-blob encoding, to encourage surface coherence and completion in unobserved areas. This joint parametric-coordinate encoding enables real-time and robust performance by bringing the best of both worlds: fast convergence and surface hole filling. Moreover, our ray sampling strategy allows Co-SLAM to perform global bundle adjustment over all keyframes instead of requiring keyframe selection to maintain a small number of active keyframes as competing neural SLAM approaches do. Experimental results show that Co-SLAM runs at 10-17Hz and achieves state-of-the-art scene reconstruction results, and competitive tracking performance in various datasets and benchmarks (ScanNet, TUM, Replica, Synthetic RGBD). Project page: https://hengyiwang.github.io/projects/CoSLAM
Hengyi Wang, Jingwen Wang 0005, Lourdes Agapito
CVPR3
2023 SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted Camera
abstract
We present a new solution to egocentric 3D body pose estimation from monocular images captured from a downward looking fish-eye camera installed on the rim of a head mounted virtual reality device. This unusual viewpoint leads to images with unique visual appearance, characterized by severe self-occlusions and strong perspective distortions that result in a drastic difference in resolution between lower and upper body. We propose a new encoder-decoder architecture with a novel multi-branch decoder designed specifically to account for the varying uncertainty in 2D joint locations. Our quantitative evaluation, both on synthetic and real-world datasets, shows that our strategy leads to substantial improvements in accuracy over state of the art egocentric pose estimation approaches. To tackle the severe lack of labelled training data for egocentric 3D pose estimation we also introduced a large-scale photo-realistic synthetic dataset. xR-EgoPose offers 383K frames of high quality renderings of people with diverse skin tones, body shapes and clothing, in a variety of backgrounds and lighting conditions, performing a range of actions. Our experiments show that the high variability in our new synthetic training corpus leads to good generalization to real world footage and to state of the art results on real world datasets with ground truth. Moreover, an evaluation on the Human3.6M benchmark shows that the performance of our method is on par with top performing approaches on the more classic problem of 3D human pose from a third person viewpoint.
Denis Tomè, Thiemo Alldieck, Patrick Peluse, Gerard Pons-Moll, Lourdes Agapito, Hernán Badino, Fernando De la Torre
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 HumanRF: High-Fidelity Neural Radiance Fields for Humans in Motion
abstract
Representing human performance at high-fidelity is an essential building block in diverse applications, such as film production, computer games or videoconferencing. To close the gap to production-level quality, we introduce HumanRF 1 , a 4D dynamic neural scene representation that captures full-body appearance in motion from multi-view video input, and enables playback from novel, unseen viewpoints. Our novel representation acts as a dynamic video encoding that captures fine details at high compression rates by factorizing space-time into a temporal matrix-vector decomposition. This allows us to obtain temporally coherent reconstructions of human actors for long sequences, while representing high-resolution details even in the context of challenging motion. While most research focuses on synthesizing at resolutions of 4MP or lower, we address the challenge of operating at 12MP. To this end, we introduce ActorsHQ, a novel multi-view dataset that provides 12MP footage from 160 cameras for 16 sequences with high-fidelity, per-frame mesh reconstructions 2 . We demonstrate challenges that emerge from using such high-resolution data and show that our newly introduced HumanRF effectively leverages this data, making a significant step towards production-level quality novel view synthesis.
Mustafa Isik, Martin Rünz, Markos Georgopoulos, Taras Khakhulin, Jonathan Starck, Lourdes Agapito, Matthias Nießner
ACM Trans. Graph.6
2022 GNPM: Geometric-Aware Neural Parametric Models
abstract
We propose Geometric Neural Parametric Models (GNPM), a learned parametric model that takes into account the local structure of data to learn disentangled shape and pose latent spaces of 4D dynamics, using a geometric-aware architecture on point clouds. Temporally consistent 3D deformations are estimated without the need for dense correspondences at training time, by exploiting cycle consistency. Besides its ability to learn dense correspondences, GNPMs also enable latent-space manipulations such as interpolation and shape/pose transfer. We evaluate GNPMs on various datasets of clothed humans, and show that it achieves comparable performance to state of the art methods that require dense correspondences during training.
Mirgahney Mohamed, Lourdes Agapito
3DV2
2022 GO-Surf: Neural Feature Grid Optimization for Fast, High-Fidelity RGB-D Surface Reconstruction
abstract
We present GO-Surf, a direct feature grid optimization method for accurate and fast surface reconstruction from RGB-D sequences. We model the underlying scene with a learned hierarchical feature voxel grid that encapsulates multi-level geometric and appearance local information. Feature vectors are directly optimized such that after being tri-linearly interpolated, decoded by two shallow MLPs into signed distance and radiance values, and rendered via volume rendering, the discrepancy between synthesized and observed RGB/depth values is minimized. Our supervision signals - RGB, depth and approximate SDF - can be obtained directly from input images without any need for fusion or post-processing. We formulate a novel SDF gradient regularization term that encourages surface smoothness and hole filling while maintaining high frequency details. GO-Surf can optimize sequences of 1-2K frames in 15–45 minutes, a speedup of$\times$60 over NeuralRGB-D [1], the most related approach based on an MLP representation, while maintaining on par performance on standard benchmarks. Project page: https://jingwenwang95.github.io/go_surf.
Jingwen Wang 0005, Tymoteusz Bleja, Lourdes Agapito
3DV3
2022 Few-Shot Keypoint Detection as Task Adaptation via Latent Embeddings
abstract
Dense object tracking, the ability to localize specific object points with pixel-level accuracy, is an important computer vision task with numerous downstream applications in robotics. Existing approaches either compute dense keypoint embeddings in a single forward pass, meaning the model is trained to track everything at once, or allocate their full capacity to a sparse predefined set of points, trading generality for accuracy. In this paper we explore a middle ground based on the observation that the number of relevant points at a given time are typically relatively few, e.g. grasp points on a target object. Our main contribution is a novel architecture, inspired by few-shot task adaptation, which allows a sparse-style network to condition on a keypoint embedding that indicates which point to track. Our central finding is that this approach provides the generality of dense-embedding models, while offering accuracy significantly closer to sparse-keypoint approaches. We present results illustrating this capacity vs. accuracy trade-off, and demonstrate the ability to zero-shot transfer to new object instances (within-class) using a real-robot pick-and-place task.
Mel Vecerík, Jackie Kay, Raia Hadsell, Lourdes Agapito, Jonathan Scholz
ICRA4
2021 DSP-SLAM: Object Oriented SLAM with Deep Shape Priors
abstract
We propose DSP-SLAM, an object-oriented SLAM system that builds a rich and accurate joint map of dense 3D models for foreground objects, and sparse landmark points to represent the background. DSP-SLAM takes as input the 3D point cloud reconstructed by a feature-based SLAM system and equips it with the ability to enhance its sparse map with dense reconstructions of detected objects. Objects are detected via semantic instance segmentation, and their shape and pose are estimated using category-specific deep shape embeddings as priors, via a novel second order optimization. Our object-aware bundle adjustment builds a pose-graph to jointly optimize camera poses, object locations and feature points. DSP-SLAM can operate at 10 frames per second on 3 different input modalities: monocular, stereo, or stereo+LiDAR. We demonstrate DSP-SLAM operating at almost frame rate on monocular-RGB sequences from the Friburg and Redwood-OS datasets, and on stereo+LiDAR sequences on the KITTI odometry dataset showing that it achieves high-quality full object reconstructions, even from partial observations, while maintaining a consistent global map. Our evaluation shows improvements in object pose and shape reconstruction with respect to recent deep prior-based reconstruction methods and reductions in camera tracking drift on the KITTI dataset. More details and demonstrations are available at our project page: https://jingwenwang95.github.io/dsp-slam/
Jingwen Wang 0005, Martin Rünz, Lourdes Agapito
3DV3
2021 Multi-Person Implicit Reconstruction From a Single Image
abstract
We present a new end-to-end learning framework to obtain detailed and spatially coherent reconstructions of multiple people from a single image. Existing multi-person methods suffer from two main drawbacks: they are often model-based and therefore cannot capture accurate 3D models of people with loose clothing and hair; or they require manual intervention to resolve occlusions or interactions. Our method addresses both limitations by introducing the first end-to-end learning approach to perform model-free implicit reconstruction for realistic 3D capture of multiple clothed people in arbitrary poses (with occlusions) from a single image. Our network simultaneously estimates the 3D geometry of each person and their 6DOF spatial locations, to obtain a coherent multi-human reconstruction. In addition, we introduce a new synthetic dataset that depicts images with a varying number of inter-occluded humans and a variety of clothing and hair styles. We demonstrate robust, high-resolution reconstructions on images of multiple humans with complex occlusions, loose clothing and a large variety of poses and scenes. Our quantitative evaluation on both synthetic and real world datasets demonstrates state-of-the-art performance with significant improvements in the accuracy and completeness of the reconstructions over competing approaches.
Armin Mustafa, Akin Caliskan, Lourdes Agapito, Adrian Hilton 0001
CVPR3
2021 CodeNeRF: Disentangled Neural Radiance Fields for Object Categories
abstract
CodeNeRF is an implicit 3D neural representation that learns the variation of object shapes and textures across a category and can be trained, from a set of posed images, to synthesize novel views of unseen objects. Unlike the original NeRF, which is scene specific, CodeNeRF learns to disentangle shape and texture by learning separate embeddings. At test time, given a single unposed image of an unseen object, CodeNeRF jointly estimates camera viewpoint, and shape and appearance codes via optimization. Unseen objects can be reconstructed from a single image, and then rendered from new viewpoints or their shape and texture edited by varying the latent codes. We conduct experiments on the SRN benchmark, which show that CodeNeRF generalises well to unseen objects and achieves on-par performance with methods that require known camera pose at test time. Our results on real-world images demonstrate that CodeNeRF can bridge the sim-to-real gap. Project page: https://github.com/wayne1123/code-nerf
Wonbong Jang, Lourdes Agapito
ICCV2
2020 FroDO: From Detections to 3D Objects
abstract
Object-oriented maps are important for scene understanding since they jointly capture geometry and semantics, allow individual instantiation and meaningful reasoning about objects. We introduce FroDO, a method for accurate 3D reconstruction of object instances from RGB video that infers their location, pose and shape in a coarse to fine manner. Key to FroDO is to embed object shapes in a novel learnt shape space that allows seamless switching between sparse point cloud and dense DeepSDF decoding. Given an input sequence of localized RGB frames, FroDO first aggregates 2D detections to instantiate a 3D bounding box per object. A shape code is regressed using an encoder network before optimizing shape and pose further under the learnt shape priors using sparse or dense shape representations. The optimization uses multi-view geometric, photometric and silhouette losses. We evaluate on real-world datasets, including Pix3D, Redwood-OS, and ScanNet, for single-view, multi-view, and multi-object reconstruction.
Martin Rünz, Kejie Li, Meng Tang 0001, Lingni Ma, Chen Kong, Tanner Schmidt, Ian D. Reid 0001, Lourdes Agapito, Julian Straub, Steven Lovegrove, Richard A. Newcombe
CVPR8
2019 xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera
abstract
We present a new solution to egocentric 3D body pose estimation from monocular images captured from a downward looking fish-eye camera installed on the rim of a head mounted virtual reality device. This unusual viewpoint, just 2 cm away from the user's face, leads to images with unique visual appearance, characterized by severe self-occlusions and strong perspective distortions that result in a drastic difference in resolution between lower and upper body. Our contribution is two-fold. Firstly, we propose a new encoder-decoder architecture with a novel dual branch decoder designed specifically to account for the varying uncertainty in the 2D joint locations. Our quantitative evaluation, both on synthetic and real-world datasets, shows that our strategy leads to substantial improvements in accuracy over state of the art egocentric pose estimation approaches. Our second contribution is a new large-scale photorealistic synthetic dataset - xR-EgoPose - offering 383K frames of high quality renderings ofpeople with a diversity of skin tones, body shapes, clothing, in a variety of backgrounds and lighting conditions, performing a range of actions. Our experiments show that the high variability in our new synthetic training corpus leads to good generalization to real world footage and to state of the art results on real world datasets with ground truth. Moreover, an evaluation on the Human3.6M benchmark shows that the performance of our method is on par with top performing approaches on the more classic problem of 3D human pose from a third person viewpoint.
Denis Tomè, Patrick Peluse, Lourdes Agapito, Hernán Badino
ICCV3
2018 Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
abstract
We propose a CNN-based approach for multi-camera markerless motion capture of the human body. Unlike existing methods that first perform pose estimation on individual cameras and generate 3D models as post-processing, our approach makes use of 3D reasoning throughout a multi-stage approach. This novelty allows us to use provisional 3D models of human pose to rethink where the joints should be located in the image and to recover from past mistakes. Our principled refinement of 3D human poses lets us make use of image cues, even from images where we previously misdetected joints, to refine our estimates as part of an end-to-end approach. Finally, we demonstrate how the high-quality output of our multi-camera setup can be used as an additional training source to improve the accuracy of existing single camera models.
Denis Tomè, Matteo Toso, Lourdes Agapito, Chris Russell 0001
3DV3
2018 3D Pick & Mix: Object Part Blending in Joint Shape and Image Manifolds
Adrián Peñate Sánchez, Lourdes Agapito
ACCV (1)2
2018 Structured Uncertainty Prediction Networks
abstract
This paper is the first work to propose a network to predict a structured uncertainty distribution for a synthesized image. Previous approaches have been mostly limited to predicting diagonal covariance matrices [15]. Our novel model learns to predict a full Gaussian covariance matrix for each reconstruction, which permits efficient sampling and likelihood evaluation. We demonstrate that our model can accurately reconstruct ground truth correlated residual distributions for synthetic datasets and generate plausible high frequency samples for real face images. We also illustrate the use of these predicted covariances for structure preserving image denoising.
Garoe Dorta, Sara Vicente, Lourdes Agapito, Neill D. F. Campbell, Ivor J. A. Simpson
CVPR3
2018 DiverseNet: When One Right Answer Is Not Enough
abstract
Many structured prediction tasks in machine vision have a collection of acceptable answers, instead of one definitive ground truth answer. Segmentation of images, for example, is subject to human labeling bias. Similarly, there are multiple possible pixel values that could plausibly complete occluded image regions. State-of-the art supervised learning methods are typically optimized to make a single test-time prediction for each query, failing to find other modes in the output space. Existing methods that allow for sampling often sacrifice speed or accuracy. We introduce a simple method for training a neural network, which enables diverse structured predictions to be made for each test-time query. For a single input, we learn to predict a range of possible answers. We compare favorably to methods that seek diversity through an ensemble of networks. Such stochastic multiple choice learning faces mode collapse, where one or more ensemble members fail to receive any training signal. Our best performing solution can be deployed for various tasks, and just involves small modifications to the existing single-mode architecture, loss function, and training regime. We demonstrate that our method results in quantitative improvements across three challenging tasks: 2D image completion, 3D volume estimation, and flow prediction.
Michael Firman, Neill D. F. Campbell, Lourdes Agapito, Gabriel J. Brostow
CVPR3
2018 MaskFusion: Real-Time Recognition, Tracking and Reconstruction of Multiple Moving Objects
abstract
We present MaskFusion, a real-time, object-aware, semantic and dynamic RGB-D SLAM system that goes beyond traditional systems which output a purely geometric map of a static scene. MaskFusion recognizes, segments and assigns semantic class labels to different objects in the scene, while tracking and reconstructing them even when they move independently from the camera. As an RGB-D camera scans a cluttered scene, image-based instance-level semantic segmentation creates semantic object masks that enable realtime object recognition and the creation of an object-level representation for the world map. Unlike previous recognition-based SLAM systems, MaskFusion does not require known models of the objects it can recognize, and can deal with multiple independent motions. MaskFusion takes full advantage of using instance-level semantic segmentation to enable semantic labels to be fused into an object-aware map, unlike recent semantics enabled SLAM systems that perform voxel-level semantic segmentation. We show augmented-reality applications that demonstrate the unique features of the map output by MaskFusion: instance-aware, semantic and dynamic. Code will be made available.
Martin Rünz, Maud Buffier, Lourdes Agapito
ISMAR3
2017 3D Reconstruction of Dynamic Scenes from Monocular Video
Lourdes Agapito
BMVC1
2017 Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image
abstract
We propose a unified formulation for the problem of 3D human pose estimation from a single raw RGB image that reasons jointly about 2D joint estimation and 3D pose reconstruction to improve both tasks. We take an integrated approach that fuses probabilistic knowledge of 3D human pose with a multi-stage CNN architecture and uses the knowledge of plausible 3D landmark locations to refine the search for better 2D locations. The entire process is trained end-to-end, is extremely efficient and obtains state-of-the-art results on Human3.6M outperforming previous approaches both on 2D and 3D errors.
Denis Tomè, Chris Russell 0001, Lourdes Agapito
CVPR3
2017 Co-fusion: Real-time segmentation, tracking and fusion of multiple objects
abstract
In this paper we introduce Co-Fusion, a dense SLAM system that takes a live stream of RGB-D images as input and segments the scene into different objects (using either motion or semantic cues) while simultaneously tracking and reconstructing their 3D shape in real time. We use a multiple model fitting approach where each object can move independently from the background and still be effectively tracked and its shape fused over time using only the information from pixels associated with that object label. Previous attempts to deal with dynamic scenes have typically considered moving regions as outliers, and consequently do not model their shape or track their motion over time. In contrast, we enable the robot to maintain 3D models for each of the segmented objects and to improve them over time through fusion. As a result, our system can enable a robot to maintain a scene description at the object level which has the potential to allow interactions with its working environment; even in the case of dynamic scenes.
Martin Rünz, Lourdes Agapito
ICRA2
2016 Better Together: Joint Reasoning for Non-rigid 3D Reconstruction with Specularities and Shading
Chris Russell 0001, Lourdes Agapito, Andrew W. Fitzgibbon, Liu-Yin Qi
BMVC2
2016 Solving Jigsaw Puzzles with Linear Programming
Chris Russell 0001, Lourdes Agapito
BMVC3
2016 Lifting Object Detection Datasets into 3D
abstract
While data has certainly taken the center stage in computer vision in recent years, it can still be difficult to obtain in certain scenarios. In particular, acquiring ground truth 3D shapes of objects pictured in 2D images remains a challenging feat and this has hampered progress in recognition-based object reconstruction from a single image. Here we propose to bypass previous solutions such as 3D scanning or manual design, that scale poorly, and instead populate object category detection datasets semi-automatically with dense, per-object 3D reconstructions, bootstrapped from:(i) class labels, (ii) ground truth figure-ground segmentations and (iii) a small set of keypoint annotations. Our proposed algorithm first estimates camera viewpoint using rigid structure-from-motion and then reconstructs object shapes by optimizing over visual hull proposals guided by loose within-class shape similarity assumptions. The visual hull sampling process attempts to intersect an object's projection cone with the cones of minimal subsets of other similar objects among those pictured from certain vantage points. We show that our method is able to produce convincing per-object 3D reconstructions and to accurately estimate cameras viewpoints on one of the most challenging existing object-category detection datasets, PASCAL VOC. We hope that our results will re-stimulate interest on joint object recognition and 3D reconstruction from a single image.
João Carreira 0002, Sara Vicente, Lourdes Agapito, Jorge P. Batista
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 Part-based modelling of compound scenes from images
abstract
We propose a method to recover the structure of a compound scene from multiple silhouettes. Structure is expressed as a collection of 3D primitives chosen from a predefined library, each with an associated pose. This has several advantages over a volume or mesh representation both for estimation and the utility of the recovered model. The main challenge in recovering such a model is the combinatorial number of possible arrangements of parts. We address this issue by exploiting the intrinsic structure and sparsity of the problem, and show that our method scales to scenes constructed from large libraries of parts.
Anton van den Hengel, Chris Russell 0001, Anthony R. Dick, John W. Bastian, Daniel Pooley, Lachlan Fleming, Lourdes Agapito
CVPR7
2015 Direct, Dense, and Deformable: Template-Based Non-rigid 3D Reconstruction from RGB Video
abstract
In this paper we tackle the problem of capturing the dense, detailed 3D geometry of generic, complex non-rigid meshes using a single RGB-only commodity video camera and a direct approach. While robust and even real-time solutions exist to this problem if the observed scene is static, for non-rigid dense shape capture current systems are typically restricted to the use of complex multi-camera rigs, take advantage of the additional depth channel available in RGB-D cameras, or deal with specific shapes such as faces or planar surfaces. In contrast, our method makes use of a single RGB video as input, it can capture the deformations of generic shapes, and the depth estimation is dense, per-pixel and direct. We first compute a dense 3D template of the shape of the object, using a short rigid sequence, and subsequently perform online reconstruction of the non-rigid mesh as it evolves over time. Our energy optimization approach minimizes a robust photometric cost that simultaneously estimates the temporal correspondences and 3D deformations with respect to the template mesh. In our experimental evaluation we show a range of qualitative results on novel datasets, we compare against an existing method that requires multi-frame optical flow, and perform a quantitative evaluation against other template-based approaches on a ground truth dataset.
Chris Russell 0001, Neill D. F. Campbell, Lourdes Agapito
ICCV4
2014 Online Dense Non-Rigid 3D Shape and Camera Motion Recovery
Antonio Agudo, J. M. M. Montiel, Lourdes Agapito, Begoña Calvo
BMVC3
2014 Good Vibrations: A Modal Analysis Approach for Sequential Non-rigid Structure from Motion
abstract
We propose an online solution to non-rigid structure from motion that performs camera pose and 3D shape estimation of highly deformable surfaces on a frame-by-frame basis. Our method models non-rigid deformations as a linear combination of some mode shapes obtained using modal analysis from continuum mechanics. The shape is first discretized into linear elastic triangles, modelled by means of finite elements, which are used to pose the force balance equations for an undamped free vibrations model. The shape basis computation comes down to solving an eigenvalue problem, without the requirement of a learning step. The camera pose and time varying weights that define the shape at each frame are then estimated on the fly, in an online fashion, using bundle adjustment over a sliding window of image frames. The result is a low computational cost method that can run sequentially in real-time. We show experimental results on synthetic sequences with ground truth 3D data and real videos for different scenarios ranging from sparse to dense scenes. Our system exhibits a good trade-off between accuracy and computational budget, it can handle missing data and performs favourably compared to competing methods.
Antonio Agudo, Lourdes Agapito, Begoña Calvo, J. M. M. Montiel
CVPR2
2014 Reconstructing PASCAL VOC
abstract
We address the problem of populating object category detection datasets with dense, per-object 3D reconstructions, bootstrapped from class labels, ground truth figure-ground segmentations and a small set of keypoint annotations. Our proposed algorithm first estimates camera viewpoint using rigid structure-from-motion, then reconstructs object shapes by optimizing over visual hull proposals guided by loose within-class shape similarity assumptions. The visual hull sampling process attempts to intersect an object's projection cone with the cones of minimal subsets of other similar objects among those pictured from certain vantage points. We show that our method is able to produce convincing per-object 3D reconstructions on one of the most challenging existing object-category detection datasets, PASCAL VOC. Our results may re-stimulate once popular geometry-oriented model-based recognition approaches.
Sara Vicente, João Carreira 0002, Lourdes Agapito, Jorge P. Batista
CVPR3
2014 Video Pop-up: Monocular 3D Reconstruction of Dynamic Scenes
Chris Russell 0001, Lourdes Agapito
ECCV (7)3
2014 Real-time sequential model-based non-rigid SFM
abstract
Tracking non-rigid objects from video is useful in robotic systems such as HMIs or robotic manipulator arms which interact with deformable objects. This paper proposes a method for sequential model-based 3D reconstruction of deformable objects and camera localization in real time. Non-rigid SFM methods commonly process a video sequence offline in a batch way. While there are real-time methods for rigid models, reconstruction of deformable 3D shapes for real-time applications is still unsolved. Dense approaches offer promising results, but processing all frames in batch, offline. We propose a real-time non-rigid reconstruction method based on a known deformable model. Object shape and pose is tracked by real-time estimation of camera pose and deformation coefficients. An extensive evaluation of the algorithm on several data sets, and comparison with state-of-the-art techniques is performed. The tests include different outlier rates, noise levels and occlusions handling.
Sebastián Bronte, Marco Paladini, Luis Miguel Bergasa, Lourdes Agapito, Roberto Arroyo
IROS4
2014 Semi-supervised Learning Using an Unsupervised Atlas
Nikolaos Pitelis, Chris Russell 0001, Lourdes Agapito
ECML/PKDD (2)3
2013 Balloon Shapes: Reconstructing and Deforming Objects with Volume from Images
abstract
Reconstructing the shape of a deformable object from a single image is a challenging problem, even when a 3D template shape is available. Many different methods have been proposed for this problem, however what they have in common is that they are only able to reconstruct the part of the surface which is visible in a reference image. In contrast, we are interested in recovering the full shape of a deformable 3D object. We introduce a new method designed to reconstruct closed surfaces. This type of surface is better suited for representing objects with volume. Our method relies on recent advances in silhouette Based reconstruction methods to obtain the template from a reference image. This template is then deformed in order to fit the measurements of a new input image. We combine an inextensibility prior on the deformation with powerful image measurements, in the form of silhouette and area constraints, to make our method less reliant on point correspondences. We show reconstruction results for different object classes, such as animals or hands, that have not been previously attempted with existing template methods.
Sara Vicente, Lourdes Agapito
3DV2
2013 Dense Variational Reconstruction of Non-rigid Surfaces from Monocular Video
abstract
This paper offers the first variational approach to the problem of dense 3D reconstruction of non-rigid surfaces from a monocular video sequence. We formulate non-rigid structure from motion (nrsfm) as a global variational energy minimization problem to estimate dense low-rank smooth 3D shapes for every frame along with the camera motion matrices, given dense 2D correspondences. Unlike traditional factorization based approaches to nrsfm, which model the low-rank non-rigid shape using a fixed number of basis shapes and corresponding coefficients, we minimize the rank of the matrix of time-varying shapes directly via trace norm minimization. In conjunction with this low-rank constraint, we use an edge preserving total-variation regularization term to obtain spatially smooth shapes for every frame. Thanks to proximal splitting techniques the optimization problem can be decomposed into many point-wise sub-problems and simple linear systems which can be easily solved on GPU hardware. We show results on real sequences of different objects (face, torso, beating heart) where, despite challenges in tracking, illumination changes and occlusions, our method reconstructs highly deforming smooth surfaces densely and accurately directly from video, without the need for any prior models or shape templates.
Ravi Garg, Anastasios Roussos, Lourdes Agapito
CVPR3
2013 Learning a Manifold as an Atlas
abstract
In this work, we return to the underlying mathematical definition of a manifold and directly characterise learning a manifold as finding an atlas, or a set of overlapping charts, that accurately describe local structure. We formulate the problem of learning the manifold as an optimisation that simultaneously refines the continuous parameters defining the charts, and the discrete assignment of points to charts. In contrast to existing methods, this direct formulation of a manifold does not require "unwrapping" the manifold into a lower dimensional space and allows us to learn closed manifolds of interest to vision, such as those corresponding to gait cycles or camera pose. We report state-of-the-art results for manifold based nearest neighbour classification on vision datasets, and show how the same techniques can be applied to the 3D reconstruction of human motion from a single image.
Nikolaos Pitelis, Chris Russell 0001, Lourdes Agapito
CVPR3
2013 Looking Beyond the Image: Unsupervised Learning for Object Saliency and Detection
abstract
We propose a principled probabilistic formulation of object saliency as a sampling problem. This novel formulation allows us to learn, from a large corpus of unlabelled images, which patches of an image are of the greatest interest and most likely to correspond to an object. We then sample the object saliency map to propose object locations. We show that using only a single object location proposal per image, we are able to correctly select an object in over 42% of the images in the Pascal VOC 2007 dataset, substantially outperforming existing approaches. Furthermore, we show that our object proposal can be used as a simple unsupervised approach to the weakly supervised annotation problem. Our simple unsupervised approach to annotating objects of interest in images achieves a higher annotation accuracy than most weakly supervised approaches.
Parthipan Siva, Chris Russell 0001, Tao Xiang 0002, Lourdes Agapito
CVPR4
2013 A Variational Approach to Video Registration with Subspace Constraints
abstract
This paper addresses the problem of non-rigid video registration, or the computation of optical flow from a reference frame to each of the subsequent images in a sequence, when the camera views deformable objects. We exploit the high correlation between 2D trajectories of different points on the same non-rigid surface by assuming that the displacement of any point throughout the sequence can be expressed in a compact way as a linear combination of a low-rank motion basis. This subspace constraint effectively acts as a trajectory regularization term leading to temporally consistent optical flow. We formulate it as a robust soft constraint within a variational framework by penalizing flow fields that lie outside the low-rank manifold. The resulting energy functional can be decoupled into the optimization of the brightness constancy and spatial regularization terms, leading to an efficient optimization scheme. Additionally, we propose a novel optimization scheme for the case of vector valued images, based on the dualization of the data term. This allows us to extend our approach to deal with colour images which results in significant improvements on the registration results. Finally, we provide a new benchmark dataset, based on motion capture data of a flag waving in the wind, with dense ground truth optical flow for evaluation of multi-frame optical flow algorithms for non-rigid surfaces. Our experiments show that our proposed approach outperforms state of the art optical flow and dense non-rigid registration algorithms.
Ravi Garg, Anastasios Roussos, Lourdes Agapito
Int. J. Comput. Vis.3
2012 Soft Inextensibility Constraints for Template-Free Non-rigid Reconstruction
Sara Vicente, Lourdes Agapito
ECCV (3)2
2012 Dense multibody motion estimation and reconstruction from a handheld camera
abstract
Existing approaches to camera tracking and reconstruction from a single handheld camera for Augmented Reality (AR) focus on the reconstruction of static scenes. However, most real world scenarios are dynamic and contain multiple independently moving rigid objects. This paper addresses the problem of simultaneous segmentation, motion estimation and dense 3D reconstruction of dynamic scenes. We propose a dense solution to all three elements of this problem: depth estimation, motion label assignment and rigid transformation estimation directly from the raw video by optimizing a single cost function using a hill-climbing approach. We do not require prior knowledge of the number of objects present in the scene - the number of independent motion models and their parameters are automatically estimated. The resulting inference method combines the best techniques in discrete and continuous optimization: a state of the art variational approach is used to estimate the dense depth maps while the motion segmentation is achieved using discrete graph-cut based optimization. For the rigid motion estimation of the independently moving objects we propose a novel tracking approach designed to cope with the small fields of view they induce and agile motion. Our experimental results on real sequences show how accurate segmentations and dense depth maps can be obtained in a completely automated way and used in marker-free AR applications.
Anastasios Roussos, Chris Russell 0001, Ravi Garg, Lourdes Agapito
ISMAR4
2012 Optimal Metric Projections for Deformable and Articulated Structure-from-Motion
Marco Paladini, Alessio Del Bue, João M. F. Xavier, Lourdes Agapito, Marko Stosic, Marija Dodig
Int. J. Comput. Vis.4
2012 Bilinear Modeling via Augmented Lagrange Multipliers (BALM)
abstract
This paper presents a unified approach to solve different bilinear factorization problems in computer vision in the presence of missing data in the measurements. The problem is formulated as a constrained optimization where one of the factors must lie on a specific manifold. To achieve this, we introduce an equivalent reformulation of the bilinear factorization problem that decouples the core bilinear aspect from the manifold specificity. We then tackle the resulting constrained optimization problem via Augmented Lagrange Multipliers. The strength and the novelty of our approach is that this framework can seamlessly handle different computer vision problems. The algorithm is such that only a projector onto the manifold constraint is needed. We present experiments and results for some popular factorization problems in computer vision such as rigid, non-rigid, and articulated Structure from Motion, photometric stereo, and 2D-3D non-rigid registration.
Alessio Del Bue, João M. F. Xavier, Lourdes Agapito, Marco Paladini
IEEE Trans. Pattern Anal. Mach. Intell.3
2011 Efficient Second Order Multi-Target Tracking with Exclusion Constraints
Chris Russell 0001, Lourdes Agapito, Francesco Setti
BMVC2
2011 Energy based multiple model fitting for non-rigid structure from motion
abstract
In this paper we reformulate the 3D reconstruction of deformable surfaces from monocular video sequences as a labeling problem. We solve simultaneously for the assignment of feature points to multiple local deformation models and the fitting of models to points to minimize a geometric cost, subject to a spatial constraint that neighboring points should also belong to the same model. Piecewise reconstruction methods rely on features shared between models to enforce global consistency on the 3D surface. To account for this overlap between regions, we consider a super-set of the classic labeling problem in which a set of labels, instead of a single one, is assigned to each variable. We propose a mathematical formulation of this new model and show how it can be efficiently optimized with a variant of α-expansion. We demonstrate how this framework can be applied to Non-Rigid Structure from Motion and leads to simpler explanations of the same data. Compared to existing methods run on the same data, our approach has up to half the reconstruction error, and is more robust to over-fitting and outliers.
Chris Russell 0001, João Fayad, Lourdes Agapito
CVPR3
2011 Automated articulated structure and 3D shape recovery from point correspondences
abstract
In this paper we propose a new method for the simultaneous segmentation and 3D reconstruction of interest point based articulated motion. We decompose a set of point tracks into rigid-bodied overlapping regions which are associated with skeletal links, while joint centres can be derived from the regions of overlap. This allows us to formulate the problem of 3D reconstruction as one of model assignment, where each model corresponds to the motion and shape parameters of an articulated body part. We show how this labelling can be optimised using a combination of pre-existing graph-cut based inference, and robust structure from motion factorization techniques. The strength of our approach comes from viewing both the decomposition into parts, and the 3D reconstruction as the optimisation of a single cost function, namely the image re-projection error. We show results of full 3D shape recovery on challenging real-world sequences with one or more articulated bodies, in the presence of outliers and missing data.
João Fayad, Chris Russell 0001, Lourdes Agapito
ICCV3
2011 Reconstruction of non-rigid 3D shapes from stereo-motion
Xavier Lladó, Alessio Del Bue, Arnau Oliver, Joaquim Salvi, Lourdes Agapito
Pattern Recognit. Lett.5
2010 Dense Multi-frame Optic Flow for Non-rigid Objects Using Subspace Constraints
Ravi Garg, Luis Pizarro, Daniel Rueckert, Lourdes Agapito
ACCV (4)4
2010 Bilinear Factorization via Augmented Lagrange Multipliers
Alessio Del Bue, João M. F. Xavier, Lourdes Agapito, Marco Paladini
ECCV (4)3
2010 Piecewise Quadratic Reconstruction of Non-Rigid Surfaces from Monocular Sequences
João Fayad, Lourdes Agapito, Alessio Del Bue
ECCV (4)2
2010 Sequential Non-Rigid Structure-from-Motion with the 3D-Implicit Low-Rank Shape Model
Marco Paladini, Adrien Bartoli, Lourdes Agapito
ECCV (2)3
2010 Non-rigid metric reconstruction from perspective cameras
Xavier Lladó, Alessio Del Bue, Lourdes Agapito
Image Vis. Comput.3
2009 Non-rigid Structure from Motion using Quadratic Deformation Models
abstract
In this paper we present a new approach to the modelling of non-rigid 3D surfaces from the observation of 2D motion in images captured by an orthographic camera. Our aim is to characterize strong variations of the shape due, for instance, to bending motions. Such motions are hard to describe with previously used deformation models, such as the linear basis shapes model, which would tend to overestimate the dimensionality of the deformable data. Our approach uses a quadratic deformation model which is able to represent non-linear non-rigid motions such as bending, stretching, shearing and twisting. The model is bilinear and thus fits easily into previous schemes for Non-Rigid Structure from Motion (NRSfM). We formulate the NRSfM problem using a non-linear optimization scheme to minimize image reprojection error and recover the camera parameters, the 3D shape at rest and the quadratic deformation transformations. Our experiments with synthetic and real data show examples in which methods based on the linear basis shape model perform poorly or do not converge and instead the quadratic model is able to achieve accurate 3D reconstructions. © 2009. The copyright of this document resides with its authors.
João Fayad, Alessio Del Bue, Lourdes Agapito, Pedro Aguiar
BMVC3
2009 Factorization for non-rigid and articulated structure using metric projections
abstract
This paper describes a new algorithm for recovering the 3D shape and motion of deformable and articulated objects purely from uncalibrated 2D image measurements using an iterative factorization approach. Most solutions to non-rigid and articulated structure from motion require metric constraints to be enforced on the motion matrix to solve for the transformation that upgrades the solution to metric space. While in the case of rigid structure the metric upgrade step is simple since the motion constraints are linear, deformability in the shape introduces non-linearities. In this paper we propose an alternating least-squares approach associated with a globally optimal projection step onto the manifold of metric constraints. An important advantage of this new algorithm is its ability to handle missing data which becomes crucial when dealing with real video sequences with self-occlusions. We show successful results of our algorithms on synthetic and real sequences of both deformable and articulated data.
Marco Paladini, Alessio Del Bue, Marko Stosic, Marija Dodig, João M. F. Xavier, Lourdes Agapito
CVPR6
2008 Recovering Euclidean deformable models from stereo-motion
abstract
In this paper we present a novel Structure from Motion (SfM) approach able to infer 3D deformable models from uncalibrated stereo images. Using a stereo setup dramatically improves the 3D model estimation when the observed 3D shape is mostly deforming without undergoing strong rigid motion. Our approach first calibrates the stereo system automatically and then computes a single metric rigid structure for each frame. Afterwards, these 3D shapes are aligned to a reference view using a RANSAC method in order to compute the mean shape of the object and to select the subset of points on the object which have remained rigid throughout the sequence without deforming. The selected rigid points are then used to compute frame-wise shape registration and to extract the motion parameters robustly from frame to frame. Finally, all this information is used in a global optimization stage with bundle adjustment which allows to refine the frame-wise initial solution and also to recover the non-rigid 3D model. We show results on synthetic and real data that prove the performance of the proposed method even when there is no rigid motion in the original sequence.
Xavier Lladó, Alessio Del Bue, Lourdes Agapito
ICPR3
2007 Non-rigid structure from motion using ranklet-based tracking and non-linear optimization
Alessio Del Bue, Fabrizio Smeraldi, Lourdes Agapito
Image Vis. Comput.3
2006 Non-Rigid Metric Shape and Motion Recovery from Uncalibrated Images Using Priors
abstract
In this paper we focus on the estimation of the 3D Euclidean shape and motion of a non-rigid object which is moving rigidly while deforming and is observed by a perspective camera. Our method exploits the fact that it is often a reasonable assumption that some of the points are deforming throughout the sequence while others remain rigid. First we use an automatic segmentation algorithm to identify the set of rigid points which in turn is used to estimate the internal camera calibration parameters and the overall rigid motion. Finally we formalise the problem of non-rigid shape estimation as a constrained non-linear minimization adding priors on the degree of deformability of each point. We perform experiments on synthetic and real data which show firstly that even when using a minimal set of rigid points it is possible to obtain reliable metric information and secondly that the shape priors help to disambiguate the contribution to the image motion caused by the deformation and the perspective distortion.
Alessio Del Bue, Xavier Lladó, Lourdes Agapito
CVPR (1)3
2006 Non-Rigid Stereo Factorization
Alessio Del Bue, Lourdes Agapito
Int. J. Comput. Vis.2
2005 Non-rigid 3D Factorization for Projective Reconstruction
abstract
In this paper we address the problem of projective reconstruction for deformable objects. Recent work in non-rigid factorization has proved that it is possible to model deformations as a linear combination of basis shapes, allowing the recovery of camera motion and 3D shape under weak perspective viewing conditions. However, the performance of these methods degrades when the object of interest is close to the camera and strong perspective distortion is present in the data. The main contribution of this work is the proposal of a practical method for the recovery of projective depths, camera motion and non-rigid 3D shape from a sequence of images under strong perspective conditions. Our approach is based on minimizing 2D reprojection errors, solving the minimization as four weighted least squares problems. Results using synthetic and real data are given to illustrate the performance of our method. 1
Xavier Lladó, Alessio Del Bue, Lourdes Agapito
BMVC3
2005 Tracking points on deformable objects with ranklets
abstract
We present a robust algorithm for point tracking on deformable objects. The key elements are the use of orientation selective rank features (ranklets), local filter adaptation and dynamic model update. A multi-scale vector of ranklets is used to encode a neighbourhood of each tracked point. The shape of the filters is optimised for each neighbourhood independently. Substantial appearance variations are catered for by maintaining a stack of models for each tracked point. This enables the system to recalibrate whenever the object reverts to its original appearance.
Fabrizio Smeraldi, Alessio Del Bue, Lourdes Agapito
ICIP (3)3
2002 Self-Calibration of Rotating and Zooming Cameras
Lourdes Agapito, Eric Hayman, Ian D. Reid 0001
Int. J. Comput. Vis.1
2001 Self-Calibration of Rotating and Zooming Cameras
Lourdes Agapito, Eric Hayman, Ian D. Reid 0001
Int. J. Comput. Vis.1
2000 The Role of Self-Calibration in Euclidean Reconstruction from Two Rotating and Zooming Cameras
Eric Hayman, Lourdes Agapito, Ian D. Reid 0001, David William Murray 0001
ECCV (2)2
2000 Motion Estimation Using the Differential Epipolar Equation
abstract
We consider the motion estimation problem in the case of very closely spaced views. We revisit the differential epipolar equation providing an interpretation of it. On the basis of this interpretation we introduce a cost function to estimate the parameters of the differential epipolar equation, which enables us to compute the camera extrinsics and some intrinsics. In the synthetic tests performed we compare this continuous method with traditional discrete motion estimation and, contrary to previous findings by Vieville et al. (1996), show that the continuous method did not perceive any computational advantage.
Luis Baumela, Lourdes Agapito, Ian D. Reid 0001, Pablo Bustos
ICPR2
1999 Linear Self-Calibration of a Rotating and Zooming Camera
abstract
A linear self-calibration method is given for computing the calibration of a stationary but rotating camera. The internal parameters of the camera are allowed to vary from image to image, allowing for zooming (change of focal length) and possible variation of the principal point of the camera. In order for calibration to be possible some constraints must be placed on the calibration of each image. The method works under the minimal assumption of zero-skew (rectangular pixels), or the more restrictive but reasonable conditions of square pixels, known pixel aspect ratio, and known principal point. Being linear the algorithm is extremely rapid, and avoids the convergence problems characteristic of iterative algorithms.
Lourdes Agapito, Eric Hayman, Richard I. Hartley
CVPR1
1999 Camera Calibration and the Search for Infinity
abstract
This paper considers the problem of self-calibration of a camera from an image sequence in the case where the camera's internal parameters (most notably focal length) may change. The problem of camera self-calibration from a sequence of images has proven to be a difficult one in practice, due to the need ultimately to resort to non-linear methods, which have often proven to be unreliable. In a stratified approach to self-calibration, a projective reconstruction is obtained first and this is successively refined first to an affine and then to a Euclidean (or metric) reconstruction. It has been observed that the difficult step is to obtain the affine reconstruction, or equivalently to locate the plane at infinity in the projective coordinate frame. The problem is inherently non-linear and requires iterative methods that risk not finding the optimal solution. The present paper overcomes this difficulty by imposing chirality constraints to limit the search for the plane at infinity to a 3-dimensional cubic region of parameter space. It is then possible to carry out a dense search over this cube in reasonable time. For each hypothesised placement of the plane at infinity, the calibration problem is reduced to one of calibration of a nontranslating camera, for which fast non-iterative algorithms exist. A cost function based on the result of the trial calibration is used to determine the best placement of the plane at infinity. Because of the simplicity of each trial, speeds of over 10,000 trials per second are achieved on a 256 MHz processor. It is shown that this dense search allows one to avoid areas of local minima effectively and find global minima of the cost function.
Richard I. Hartley, Lourdes Agapito, Ian D. Reid 0001, Eric Hayman
ICCV2
1998 Self-Calibration of a Rotating Camera with Varying Intrinsic Parameters
abstract
We present a method for self-calibration of a camera which is free to rotate and change its intrinsic parameters, but which cannot translate. The method is based on the so-called infinite homography constraint which leads to a non-linear minimisation routine to find the unknown camera intrinsics over an extended sequence of images. We give experimental results using real image sequences for which ground truth data was available. 1
Lourdes Agapito, Eric Hayman, Ian D. Reid 0001
BMVC1
1998 Self-calibrating a Stereo Head: An Error Analysis in the Neighbourhood of Degenerate Configurations
abstract
We show that the self-calibration of a stereo head corresponding points in an image pair is in certain circumstances prone to considerable error. A novel error analysis reveals that the automated determination of relative orientation and focal length is adversely affected when the cameras verge inwards a similar amount, and when the principal point locations have a horizontal error. This analysis is facilitated by the adoption of closed-form solutions for self-calibration from previous work of the authors. It is also shown that estimation of the fundamental matrix associated with a stereo head image pair is improved when a domain-specific parameterisation and associated computational techniques are adopted. Experiments conducted with such image pairs suggest that, given cognisance of sensitive configurations and adoption of the revised method of fundamental matrix estimation, robust reconstructions are attainable. This is demonstrated on the problem of metrically reconstructing a scene from two pairs of images obtained by an uncalibrated stereo head undergoing unknown ground-plane motion.
Lourdes Agapito, Du Q. Huynh, Michael J. Brooks
ICCV1
1998 Towards robust metric reconstruction via a dynamic uncalibrated stereo head
Michael J. Brooks, Lourdes Agapito, Du Q. Huynh, Luis Baumela
Image Vis. Comput.2
1996 Direct Methods for Self-Calibration of a Moving Stereo Head
Michael J. Brooks, Lourdes Agapito, Du Q. Huynh, Luis Baumela
ECCV (2)2