VLDB 2026 Research / reviewers in the wild / expert
Antonio Agudo
dblp:12/10782
· DBLP profile ↗
56ranked-venue papers
26as first author
28since 2021 · last 2026
0000-0001-6845-4998ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 19 first-author · 24 since 2021Artificial intelligence and machine learning · 34 · 16 first-author · 11 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Modal Soccer Scene Analysis with Masked Pre-TrainingabstractIn this work we propose a multi-modal architecture for analyzing soccer scenes from tactical camera footage, with a focus on three core tasks: ball trajectory inference, ball state classification, and ball possessor identification. To this end, our solution integrates three distinct input modalities (player trajectories, player types and image crops of individual players) into a unified framework that processes spatial and temporal dynamics using a cascade of sociotemporal transformer blocks. Unlike prior methods, which rely heavily on accurate ball tracking or handcrafted heuristics, our approach infers the ball trajectory without direct access to its past or future positions, and robustly identifies the ball state and ball possessor under noisy or occluded conditions from real top league matches. We also introduce CropDrop, a modality-specific masking pre-training strategy that prevents over-reliance on image features and encourages the model to rely on cross-modal patterns during pre-training. We show the effectiveness of our approach on a large-scale dataset providing substantial improvements over state-of-the-art baselines in all tasks. Our results highlight the benefits of combining structured and visual cues in a transformer-based architecture, and the importance of realistic masking strategies in multi-modal learning. Marc Peral, Guillem Capellera, Luis Ferraz, Antonio Rubio 0001, Antonio Agudo |
WACV | 5 |
| 2026 | Unsupervised Modular Adaptive Region Growing and RegionMix Classification for Wind Turbine SegmentationabstractReliable operation of wind turbines requires frequent inspections, as even minor surface damages can degrade aerodynamic performance, reduce energy output, and accelerate blade wear. Central to automating these inspections is the accurate segmentation of turbine blades from visual data. This task is traditionally addressed through dense, pixel-wise deep learning models. However, such methods demand extensive annotated datasets, posing scalability challenges. In this work, we introduce an annotation-efficient segmentation approach that reframes the pixel-level task into a binary region classification problem. Image regions are generated using a fully unsupervised, interpretable Modular Adaptive Region Growing technique, guided by image-specific Adaptive Thresholding and enhanced by a Region Merging process that consolidates fragmented areas into coherent segments. To improve generalization and classification robustness, we introduce RegionMix, an augmentation strategy that synthesizes new training samples by combining distinct regions. Our framework demonstrates state-of-the-art segmentation accuracy and strong cross-site generalization by consistently segmenting turbine blades across distinct windfarms. Raül Pérez-Gonzalo, Riccardo Magro, Andreas Espersen, Antonio Agudo |
WACV | 4 |
| 2026 | PnLCalib: Sports field registration via points and lines optimizationabstractCamera calibration in broadcast sports videos presents numerous challenges for accurate sports field registration due to multiple camera angles, varying camera parameters, and frequent occlusions of the field. Traditional search-based methods depend on initial camera pose estimates, which can struggle in non-standard positions and dynamic environments. In response, we propose an optimization-based calibration pipeline that leverages a 3D soccer field model and a predefined set of keypoints to overcome these limitations. Our method also introduces a novel refinement module that improves initial calibration by using detected field lines in a non-linear optimization process. This approach outperforms existing techniques in both multi-view and single-view 3D camera calibration tasks, while maintaining competitive performance in homography estimation. Extensive experimentation on real-world soccer datasets, including SoccerNet-Calibration, WorldCup 2014, and TS-WorldCup, highlights the robustness and accuracy of our method across diverse broadcast scenarios. Our approach offers significant improvements in camera calibration precision and reliability. Our project is available at https://github.com/mguti97/PnLCalib . • Most sports field registration methods focus on homography rather than full 3D camera calibration. • A robust keypoint grid based on field geometry can outperform existing methods. • Field’s landmark scarcity is mitigated by optimizing calibration with field lines. Marc Gutiérrez-Pérez, Antonio Agudo |
Comput. Vis. Image Underst. | 2 |
| 2026 | Once Upon a Goal: Towards orientation-based shot metrics in footballabstractSports analytics has been revolutionized by advanced tracking technologies, yet the integration of human pose estimation into performance metrics remains underexplored in football. Estimating the probability of scoring from a shot is a central task in football analytics and is commonly approached through expected goals (xG) models. Progress in this area, however, is often constrained by the limited availability of publicly accessible datasets that include fine-grained biomechanical information. In this work, we present xGHub, an open-source dataset of football shots enriched with player pose estimation, body orientation, and contextual features extracted from broadcast video. The dataset is generated using an automated pipeline for player detection and tracking, followed by an external verification process to ensure annotation reliability. As a use case, we analyze how pose- and orientation-related features can be incorporated into a standard xG modeling framework. Our results indicate that 3D orientation information is informative for specific subsets of shots, while its contribution is limited in others, reflecting the inherently non-linear nature of angular representations. This analysis serves to illustrate the potential and limitations of the released annotations. By making this dataset publicly available, we aim to support future research on the role of player biomechanics in shot analysis and related football analytics tasks.Sports analytics has been revolutionized by advanced tracking technologies, yet the integration of human pose estimation into performance metrics remains underexplored in football. Estimating the probability of scoring from a shot is a central task in football analytics and is commonly approached through expected goals (xG) models. Progress in this area, however, is often constrained by the limited availability of publicly accessible datasets that include fine-grained biomechanical information. In this work, we present xGHub, an open-source dataset of football shots enriched with player pose estimation, body orientation, and contextual features extracted from broadcast video. The dataset is generated using an automated pipeline for player detection and tracking, followed by an external verification process to ensure annotation reliability. As a use case, we analyze how pose- and orientation-related features can be incorporated into a standard xG modeling framework. Our results indicate that 3D orientation information is informative for specific subsets of shots, while its contribution is limited in others, reflecting the inherently non-linear nature of angular representations. This analysis serves to illustrate the potential and limitations of the released annotations. By making this dataset publicly available, we aim to support future research on the role of player biomechanics in shot analysis and related football analytics tasks. Marc Gutiérrez-Pérez, Calvin Yeung 0001, Keisuke Fujii 0001, Antonio Agudo |
Comput. Vis. Image Underst. | 4 |
| 2025 | Unified Uncertainty-Aware Diffusion for Multi-Agent Trajectory ModelingabstractMulti-agent trajectory modeling has primarily focused on forecasting future states, often overlooking broader tasks like trajectory completion, which are crucial for real-world applications such as correcting tracking data. Existing methods also generally predict agents’ states without offering any state-wise measure of uncertainty. Moreover, popular multi-modal sampling methods lack any error probability estimates for each generated scene under the same prior observations, making it difficult to rank the predictions during inference time. We introduce U2Diff, a unified diffusion model designed to handle trajectory completion while providing state-wise uncertainty estimates jointly. This uncertainty estimation is achieved by augmenting the simple denoising loss with the negative log-likelihood of the predicted noise and propagating latent space uncertainty to the real state space. Additionally, we incorporate a Rank Neural Network in post-processing to enable error probability estimation for each generated mode, demonstrating a strong correlation with the error relative to ground truth. Our method outperforms the state-of-the-art solutions in trajectory completion and forecasting across four challenging sports datasets (NBA, Basketball-U, Football-U, Soccer-U), highlighting the effectiveness of uncertainty and error probability estimation. Guillem Capellera, Antonio Rubio 0001, Luis Ferraz, Antonio Agudo |
CVPR | 4 |
| 2025 | Dual-Space Augmented Intrinsic-LoRA for Wind Turbine SegmentationabstractAccurate segmentation of wind turbine blade (WTB) images is critical for effective assessments, as it directly influences the performance of automated damage detection systems. Despite advancements in large universal vision models, these models often underperform in domain-specific tasks like WTB segmentation. To address this, we extend Intrinsic LoRA for image segmentation, and propose a novel dual-space augmentation strategy that integrates both image-level and latent-space augmentations. The image-space augmentation is achieved through linear interpolation between image pairs, while the latent-space augmentation is accomplished by introducing a noise-based latent probabilistic model. Our approach significantly boosts segmentation accuracy, surpassing current state-of-the-art methods in WTB image segmentation. Shubh Singhal, Raül Pérez-Gonzalo, Andreas Espersen, Antonio Agudo |
ICASSP | 4 |
| 2025 | Recovering and Classifying Upper Limb Impairment Trajectories After StrokeabstractUpper limb impairment is a loss of motor function after a stroke, leading to difficulties in performing daily tasks. With a low remission rate six months after stroke, monitoring during this critical period is essential. Telemonitoring with inertial sensors has become a common approach that requires the identification and recognition of specific movements. However, data quality in clinical databases is a challenge, making data augmentation necessary. In this work, we propose a novel method that can learn a trajectory subspace from partial 3D signals in an unsupervised manner. Our method is simple yet effective, producing novel human-feasible motions compatible with the estimated trajectory subspace. To this end, three approaches are introduced from global to local models that can capture a wide variety of human motions. Our method outperforms the results in the state of the art in both control and patient subjects. Martín Méndez-López, Alicia Fornés, Antonio Agudo |
ICIP | 3 |
| 2024 | TranSPORTmer: A Holistic Approach to Trajectory Understanding in Multi-agent Sports
Guillem Capellera, Luis Ferraz, Antonio Rubio 0001, Antonio Agudo, Francesc Moreno-Noguer |
ACCV (2) | 4 |
| 2024 | 4DPV: 4D Pet from Videos by Coarse-to-Fine Non-rigid Radiance Fields
Sergio Montoya de Paco, Antonio Agudo |
ACCV (9) | 2 |
| 2024 | VQ-HPS: Human Pose and Shape Estimation in a Vector-Quantized Latent Space
Guénolé Fiche, Simon Leglaive, Xavier Alameda-Pineda, Antonio Agudo, Francesc Moreno-Noguer |
ECCV (52) | 4 |
| 2024 | Photovoltaic Power Forecasting Using Sky Images and Sun MotionabstractSolar energy adoption is moving at a rapid pace. The variability in solar energy production causes grid stability issues and hinders mass adoption. To solve these issues, more accurate photovoltaic power forecasting systems are needed. In intra-hour forecasting, the most challenging issue is high output fluctuations due to cloud motion, which can occlude the sun. Using ground-based sky images, this paper proposes two convolutional neural network models for intra-hour nowcasting and forecasting that incorporate physical information on sun motion and cloud coverage by means of the sun area mean pixel intensity. Particularly, our models exploit that information instead of relying exclusively on photovoltaic output history data as it is standard in state of the art. Taking advantage of sun position and cloud coverage information, we were able to reduce the overall root mean squared error for the nowcasting task, making the model more accurate especially during cloudy days, and obtaining competitive results on forecasting. Moreover, our models are more robust against artifacts such as occlusion and noisy observations. Arne Berresheim, Antonio Agudo |
ICASSP | 2 |
| 2024 | Footbots: A Transformer-Based Architecture for Motion Prediction in SoccerabstractMotion prediction in soccer involves capturing complex dynamics from player and ball interactions. We present FootBots, an encoder-decoder transformer-based architecture addressing motion prediction and conditioned motion prediction through equivariance properties. FootBots captures temporal and social dynamics using set attention blocks and multi-attention block decoder. Our evaluation utilizes two datasets: a real soccer dataset and a tailored synthetic one. Insights from the synthetic dataset highlight the effectiveness of FootBots’ social attention mechanism and the significance of conditioned motion prediction. Empirical results on real soccer data demonstrate that FootBots outperforms baselines in motion prediction and excels in conditioned tasks, such as predicting the players based on the ball position, predicting the offensive (defensive) team based on the ball and the defensive (offensive) team, and predicting the ball position based on all players. Our evaluation connects quantitative and qualitative findings. https://youtu.be/9kaEkfzG3L8 Guillem Capellera, Luis Ferraz, Antonio Rubio 0001, Antonio Agudo, Francesc Moreno-Noguer |
ICIP | 4 |
| 2024 | Uncalibrated and Unsupervised Photometric Stereo with Piecewise RegularizerabstractPhotometric stereo is a technique for recovering a rigid object’s 3D shape, reflectance properties, lighting conditions, and specular highlights from multiple images captured under varying lighting conditions. Variational, uncalibrated, and unsupervised formulations have recently provided detailed and robust solutions to the problem, reducing the need for prior knowledge about shape geometry or lighting conditions. However, uncalibrated methods, especially when applied to real-world data, may be susceptible to noise and depth errors near boundaries or self-occlusions, stemming from missing or noisy data, surface orientation ambiguity, and calibration issues. In this paper, we introduce a novel piecewise depth regularizer to mitigate these errors, enhancing stability and improving robustness against initialization errors. We demonstrate the effectiveness of our approach through evaluations on both synthetic and real-world data, showcasing its promise in enhancing the accuracy and reliability of photometric stereo for practical applications. Alejandro Casanova, Antonio Agudo |
ICIP | 2 |
| 2024 | Generalized Nested Latent Variable Models For Lossy Coding Applied To Wind Turbine ScenariosabstractRate-distortion optimization through neural networks has accomplished competitive results in compression efficiency and image quality. This learning-based approach seeks to minimize the compromise between compression rate and reconstructed image quality by automatically extracting and retaining crucial information, while discarding less critical details. A successful technique consists in introducing a deep hyperprior that operates within a 2-level nested latent variable model, enhancing compression by capturing complex data dependencies. This paper extends this concept by designing a generalized L-level nested generative model with a Markov chain structure. We demonstrate as L increases that a trainable prior is detrimental and explore a common dimensionality along the distinct latent variables to boost compression performance. As this structured framework can represent autoregressive coders, we outperform the hyperprior model and achieve state-of-the-art performance while reducing substantially the computational cost. Our experimental evaluation is performed on wind turbine scenarios to study its application on visual inspections. Raül Pérez-Gonzalo, Andreas Espersen, Antonio Agudo |
ICIP | 3 |
| 2023 | Detail-Aware Uncalibrated Photometric StereoabstractPhotometric stereo is the problem of jointly inferring the 3D reconstruction, reflectance, lighting and specularities of an object from a set of visual signals. Recently, some variational, uncalibrated, unsupervised and unified formulations have provided robust solutions to the problem while reducing the prior knowledge about the shape geometry or the lighting conditions. Unfortunately, these approaches cannot still produce solutions with an ample variety of details in the shape. That is mainly due to the non-convex and non-linear nature of the problem which requires the best initialization as possible. In this context, we propose a fully interpretable formulation that combines a physically-aware image formation model under perspective projection with a minimal detail-aware initialization and that it can handle general lighting. As a result, our formulation can consider multiple scenarios composed of unknown complex geometries and lighting patterns. Experimental results on challenging synthetic and real datasets show the effectiveness of our approach to capture more fine details, outperforming state-of-the-art techniques in terms of 3D reconstruction. Antonio Agudo |
ICASSP | 1 |
| 2023 | Robust Wind Turbine Blade Segmentation from RGB Images in the WildabstractWith the relentless growth of the wind industry, there is an imperious need to design automatic data-driven solutions for wind turbine maintenance. As structural health monitoring mainly relies on visual inspections, the first stage in any automatic solution is to identify the blade region on the image. Thus, we propose a novel segmentation algorithm that strengthens the U-Net results by a tailored loss, which pools the focal loss with a contiguity regularization term. To attain top performing results, a set of additional steps are proposed to ensure a reliable, generic, robust and efficient algorithm. First, we leverage our prior knowledge on the images by filling the holes enclosed by temporarily-classified blade pixels and by the image boundaries. Subsequently, the mislead classified pixels are successfully amended by training an on-the-fly random forest. Our algorithm demonstrates its effectiveness reaching a non-trivial 97.39% of accuracy. Raül Pérez-Gonzalo, Andreas Espersen, Antonio Agudo |
ICIP | 3 |
| 2022 | Conditional-Flow NeRF: Accurate 3D Modelling with Reliable Uncertainty Quantification
Jianxiong Shen, Antonio Agudo, Francesc Moreno-Noguer, Adria Ruiz |
ECCV (3) | 2 |
| 2022 | Safari from Visual Signals: Recovering Volumetric 3d ShapesabstractIn this paper we propose a convex approach for recovering a detailed 3D volumetric geometry of several objects from visual signals. To this end, we first present a minimal detailed surface energy that is optimized together with a volume constraint by considering some geometrical priors, and without requiring neither additional training data nor templates in order to constrain the solution. Our problem can be efficiently solved by means of a gradient descent, and be applied for single RGB images or monocular videos even with very small rigid motions. Temporal-aware solutions and driven by point correspondences are incorporated without assuming any 2D tracking data over time. Thanks to this formulation, both rigid and non-rigid objects can be considered. We have extensively validated our approach in a wide variety of scenarios in the wild, recovering challenging type of shapes that have not been previously attempted without assuming any training data. Antonio Agudo |
ICASSP | 1 |
| 2022 | Spline Human Motion RecoveryabstractSimultaneous camera pose, 4D reconstruction of an object and deformation clustering from incomplete 2D point tracks in a video is a challenging problem. To solve it, in this work we introduce a union of piecewise subspaces to encode the 4D shape, where two modalities based on B-splines and Catmull-Rom curves are considered. We demonstrate that formulating the problem in terms of B-spline or Catmull-Rom functions, allows for a better physical interpretation of the resulting priors while C1and C2continuities are automatically imposed without needing any additional constraint. An optimization framework is proposed to sort out the problem in a unified, accurate, unsupervised and efficient manner. We extensively validate our claims on a wide range of human motions, including articulated and continuous deformations as well as those cases with noisy and missing measurements where our approach provides competing joint solutions. Antonio Agudo |
ICIP | 1 |
| 2022 | An Adaptable Approach to Learn Realistic Legged Locomotion without ExamplesabstractLearning controllers that reproduce legged locomotion in nature has been a longtime goal in robotics and computer graphics. While yielding promising results, recent approaches are not yet flexible enough to be applicable to legged systems of different morphologies. This is partly because they often rely on precise motion capture references or elaborate learning environments that ensure the naturality of the emergent locomotion gaits but prevent generalization. This work proposes a generic approach for ensuring realism in locomotion by guiding the learning process with the spring-loaded inverted pendulum model as a reference. Leveraging on the exploration capacities of Reinforcement Learning (RL), we learn a control policy that fills in the information gap between the template model and full-body dynamics required to maintain stable and periodic locomotion. The proposed approach can be applied to robots of different sizes and morphologies and adapted to any RL technique and control architecture. We present experimental results showing that even in a model-free setup and with a simple reactive control architecture, the learned policies can generate realistic and energy-efficient locomotion gaits for a bipedal and a quadrupedal robot. And most importantly, this is achieved without using motion capture, strong constraints in the dynamics or kinematics of the robot, nor prescribing limb coordination. We provide supplemental videos for qualitative analysis of the naturality of the learned gaits4. Daniel Felipe Ordoñez Apraez, Antonio Agudo, Francesc Moreno-Noguer, Mario Martín |
ICRA | 2 |
| 2022 | Matching and Recovering 3D People from Multiple ViewsabstractThis paper introduces an approach to simultaneously match and recover 3D people from multiple calibrated cameras. To this end, we present an affinity measure between 2D detections across different views that enforces an uncertainty geometric consistency. This similarity is then exploited by a novel multi-view matching algorithm to cluster the detections, being robust against partial observations as well as bad detections and without assuming any prior about the number of people in the scene. After that, the multi-view correspondences are used in order to efficiently infer the 3D pose of each body by means of a 3D pictorial structure model in combination with physico-geometric constraints. Our algorithm is thoroughly evaluated on challenging scenarios where several human bodies are performing different activities which involve complex motions, producing large occlusions in some views and noisy observations. We outperform state-of-the-art results in terms of matching and 3D reconstruction. Alejandro Pérez-Yus, Antonio Agudo |
WACV | 2 |
| 2022 | Unsupervised 3D Reconstruction and Grouping of Rigid and Non-Rigid CategoriesabstractIn this paper we present an approach to jointly recover camera pose, 3D shape, and object and deformation type grouping, from incomplete 2D annotations in a multi-instance collection of RGB images. Our approach is able to handle indistinctly both rigid and non-rigid categories. This advances existing work, which only addresses the problem for one single object or, they assume the groups to be known a priori when multiple instances are handled. In order to address this broader version of the problem, we encode object deformation by means of multiple unions of subspaces, that is able to span from small rigid motion to complex deformations. The model parameters are learned via Augmented Lagrange Multipliers, in a completely unsupervised manner that does not require any training data at all. Extensive experimental evaluation is provided in a wide variety of synthetic and real scenarios, including rigid and non-rigid categories with small and large deformations. We obtain state-of-the-art solutions in terms of 3D reconstruction accuracy, while also providing grouping results that allow splitting the input images into object instances and their associated type of deformation. Antonio Agudo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Simultaneous completion and spatiotemporal grouping of corrupted motion tracksabstractAbstract Given an unordered list of 2D or 3D point trajectories corrupted by noise and partial observations, in this paper we introduce a framework to simultaneously recover the incomplete motion tracks and group the points into spatially and temporally coherent clusters. This advances existing work, which only addresses partial problems and without considering a unified and unsupervised solution. We cast this problem as a matrix completion one, in which point tracks are arranged into a matrix with the missing entries set as zeros. In order to perform the double clustering, the measurement matrix is assumed to be drawn from a dual union of spatiotemporal subspaces. The bases and the dimensionality for these subspaces, the affinity matrices used to encode the temporal and spatial clusters to which each point belongs, and the non-visible tracks, are then jointly estimated via augmented Lagrange multipliers in polynomial time. A thorough evaluation on incomplete motion tracks for multiple-object typologies shows that the accuracy of the matrix we recover compares favorably to that obtained with existing low-rank matrix completion methods, specially under noisy measurements. In addition, besides recovering the incomplete tracks, the point trajectories are directly grouped into different object instances, and a number of semantically meaningful temporal primitive actions are automatically discovered. Antonio Agudo, Vincent Lepetit, Francesc Moreno-Noguer |
Vis. Comput. | 1 |
| 2021 | Stochastic Neural Radiance Fields: Quantifying Uncertainty in Implicit 3D RepresentationsabstractNeural Radiance Fields (NeRF) has become a popular framework for learning implicit 3D representations and addressing different tasks such as novel-view synthesis or depth-map estimation. However, in downstream applications where decisions need to be made based on automatic predictions, it is critical to leverage the confidence associated with the model estimations. Whereas uncertainty quantification is a long-standing problem in Machine Learning, it has been largely overlooked in the recent NeRF literature. In this context, we propose Stochastic Neural Radiance Fields (S-NeRF), a generalization of standard NeRF that learns a probability distribution over all the possible radiance fields modeling the scene. This distribution allows to quantify the uncertainty associated with the scene information provided by the model. S-NeRF optimization is posed as a Bayesian learning problem that is efficiently addressed using the Variational Inference framework. Exhaustive experiments over benchmark datasets demonstrate that S-NeRF is able to provide more reliable predictions and confidence values than generic approaches previously proposed for uncertainty estimation in other domains. Jianxiong Shen, Adria Ruiz, Antonio Agudo, Francesc Moreno-Noguer |
3DV | 3 |
| 2021 | Body Size and Depth Disambiguation in Multi-Person Reconstruction from Single ImagesabstractWe address the problem of multi-person 3D body pose and shape estimation from a single image. While this problem can be addressed by applying single-person approaches multiple times for the same scene, recent works have shown the advantages of building upon deep architectures that simultaneously reason about all people in the scene in a holistic manner by enforcing, e.g., depth order constraints or minimizing interpenetration among reconstructed bodies. However, existing approaches are still unable to capture the size variability of people caused by the inherent body scale and depth ambiguity. In this work, we tackle this challenge by devising a novel optimization scheme that learns the appropriate body scale and relative camera pose, by enforcing the feet of all people to remain on the ground floor. A thorough evaluation on MuPoTS- 3D and 3DPW datasets demonstrates that our approach is able to robustly estimate the body translation and shape of multiple people while retrieving their spatial arrangement, consistently improving current state-of-the-art, especially in scenes with people of very different heights. Code can be found at: https://github.com/nicolasugrinovic/size_depth_disambiguation Nicolas Ugrinovic, Adria Ruiz, Antonio Agudo, Alberto Sanfeliu, Francesc Moreno-Noguer |
3DV | 3 |
| 2021 | Uncertainty-Aware Camera Pose Estimation From Points and LinesabstractPerspective-n-Point-and-Line (PnPL) algorithms aim at fast, accurate, and robust camera localization with respect to a 3D model from 2D-3D feature correspondences, being a major part of modern robotic and AR/VR systems. Current point-based pose estimation methods use only 2D feature detection uncertainties, and the line-based methods do not take uncertainties into account. In our setup, both 3D co-ordinates and 2D projections of the features are considered uncertain. We propose PnP(L) solvers based on EPnP [20] and DLS [14] for the uncertainty-aware pose estimation. We also modify motion-only bundle adjustment to take 3D uncertainties into account. We perform exhaustive synthetic and real experiments on two different visual odometry datasets. The new PnP(L) methods outperform the state-of-the-art on real data in isolation, showing an increase in mean translation accuracy by 18% on a representative subset of KITTI, while the new uncertain refinement improves pose accuracy for most of the solvers, e.g. decreasing mean translation error for the EPnP by 16% compared to the standard refinement on the same dataset. The code is available at https://alexandervakhitov.github.io/uncertain-pnp/. Alexander Vakhitov, Luis Ferraz, Antonio Agudo, Francesc Moreno-Noguer |
CVPR | 3 |
| 2021 | Generating Attribution Maps with Disentangled Masked BackpropagationabstractAttribution map visualization has arisen as one of the most effective techniques to understand the underlying inference process of Convolutional Neural Networks. In this task, the goal is to compute an score for each image pixel related to its contribution to the network output. In this paper, we introduce Disentangled Masked Backpropagation (DMBP), a novel gradient-based method that leverages on the piecewise linear nature of ReLU networks to decompose the model function into different linear mappings. This decomposition aims to disentangle the attribution maps into positive, negative and nuisance factors by learning a set of variables masking the contribution of each filter during back-propagation. A thorough evaluation over standard architectures (ResNet50 and VGG16) and benchmark datasets (PASCAL VOC and ImageNet) demonstrates that DMBP generates more visually interpretable attribution maps than previous approaches. Additionally, we quantitatively show that the maps produced by our method are more consistent with the true contribution of each pixel to the final network output. Adria Ruiz, Antonio Agudo, Francesc Moreno-Noguer |
ICCV | 2 |
| 2021 | Piecewise Bézier Space: Recovering 3D Dynamic Motion From VideoabstractIn this paper we address the problem of jointly retrieving a 3D dynamic shape, camera motion, and deformation grouping from partial 2D point trajectories in a monocular video. To this end, we introduce a union of piecewise Bézier subspaces with enforcing continuities to model 3D motion. We show that formulating the problem in terms of piecewise curves, allows for a better physical interpretation of the resulting priors and a more accurate representation of the motion. An energy based formulation is presented to solve the problem in an unsupervised, unified, accurate and efficient manner, by means of the use of augmented Lagrange multipliers. We thoroughly validate the approach on a wide variety of human video sequences, including those cases with noisy and missing observations, and providing more accurate joint estimations than state-of-the-art approaches. Antonio Agudo |
ICIP | 1 |
| 2020 | Neural Dense Non-Rigid Structure from Motion with Latent Space Constraints
Vikramjit Sidhu, Edith Tretschk, Vladislav Golyanik, Antonio Agudo, Christian Theobalt |
ECCV (16) | 4 |
| 2020 | Segmentation and 3D Reconstruction of NON-RIGID Shape from RGB VideoabstractIn this paper we propose a unsupervised and unified approach to simultaneously recover time-varying 3D shape, camera motion, and temporal clustering into deformations, all of them, from partial 2D point tracks in a RGB video and without assuming any pre-trained model. As the data are drawn from a sequentially ordered images, we fully exploit this information to constrain all model parameters we estimate. We present an energy-based formulation that is efficiently solved and allows to estimate all model parameters in the same loop via augmented Lagrange multipliers in polynomial time, enforcing similarities between images at any level. Validation is done in a wide variety of human video sequences, including articulated and continuous motion, and for dense and missing tracks. Our approach is shown to outperform state-of-the-art solutions in terms of 3D reconstruction and clustering. Antonio Agudo |
ICIP | 1 |
| 2020 | Total Estimation from RGB Video: On-line Camera Self-Calibration, Non-Rigid Shape and MotionabstractIn this paper we present a sequential approach to jointly retrieve camera auto-calibration, camera pose and the 3D reconstruction of a non-rigid object from an uncalibrated RGB image sequence, without assuming any prior information about the shape structure, nor the need for a calibration pattern, nor the use of training data at all. To this end, we propose a Bayesian filtering approach based on a sum-of-Gaussians filter composed of a bank of extended Kalman filters (EKF). For every EKF, we make use of dynamic models to estimate its state vector, which later will be Gaussianly combined to achieve a global solution. To deal with deformable objects, we incorporate a mechanical model solved by using the finite element method. Thanks to these ingredients, the resulting method is both efficient and robust to several artifacts such as missing and noisy observations as well as sudden camera motions, while being available for a wide variety of objects and materials, including isometric and elastic shape deformations. Experimental validation is proposed in real experiments, showing its strengths with respect to competing approaches. Antonio Agudo |
ICPR | 1 |
| 2020 | E-DNAS: Differentiable Neural Architecture Search for Embedded SystemsabstractDesigning optimal and light weight networks to fit in resource-limited platforms like mobiles, DSPs or GPUs is a challenging problem with a wide range of interesting applications, e.g. in embedded systems for autonomous driving. While most approaches are based on manual hyperparameter tuning, there exist a new line of research, the so-called NAS (Neural Architecture Search) methods, that aim to optimize several metrics during the design process, including memory requirements of the network, number of FLOPs, number of MACs (Multiply-ACcumulate operations) or inference latency. However, while NAS methods have shown very promising results, they are still significantly time and cost consuming. In this work we introduce E-DNAS, a differentiable architecture search method, which improves the efficiency of NAS methods in designing light-weight networks for the task of image classification. Concretely, E-DNAS computes, in a differentiable manner, the optimal size of a number of meta-kernels that capture patterns of the input data at different resolutions. We also leverage on the additive property of convolution operations to merge several kernels with different compatible sizes into a single one, reducing thus the number of operations and the time required to estimate the optimal configuration. We evaluate our approach on several datasets to perform classification. We report results in terms of the SoC (System on Chips) metric, typically used in the Texas Instruments TDA2x families for autonomous driving applications. The results show that our approach allows designing low latency architectures significantly faster than state-of-the-art. Javier García López, Antonio Agudo, Francesc Moreno-Noguer |
ICPR | 2 |
| 2020 | GANimation: One-Shot Anatomically Consistent Facial Animation
Albert Pumarola, Antonio Agudo, Aleix Martinez, Alberto Sanfeliu, Francesc Moreno-Noguer |
Int. J. Comput. Vis. | 2 |
| 2019 | Robust Spatio-Temporal Clustering and Reconstruction of Multiple Deformable BodiesabstractIn this paper we present an approach to reconstruct the 3D shape of multiple deforming objects from a collection of sparse, noisy and possibly incomplete 2D point tracks acquired by a single monocular camera. Additionally, the proposed solution estimates the camera motion and reasons about the spatial segmentation (i.e., identifies each of the deforming objects in every frame) and temporal clustering (i.e., splits the sequence into motion primitive actions). This advances competing work, which mainly tackled the problem for one single object and non-occluded tracks. In order to handle several objects at a time from partial observations, we model point trajectories as a union of spatial and temporal subspaces, and optimize the parameters of both modalities, the non-observed point tracks, the camera motion, and the time-varying 3D shape via augmented Lagrange multipliers. The algorithm is fully unsupervised and does not require any training data at all. We thoroughly validate the method on challenging scenarios with several human subjects performing different activities which involve complex motions and close interaction. We show our approach achieves state-of-the-art 3D reconstruction results, while it also provides spatial and temporal segmentation. Antonio Agudo, Francesc Moreno-Noguer |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Shape Basis Interpretation for Monocular Deformable 3-D ReconstructionabstractIn this paper, we propose a novel interpretable shape model to encode object nonrigidity. We first use the initial frames of a monocular video to recover a rest shape, used later to compute a dissimilarity measure based on a distance matrix measurement. Spectral analysis is then applied to this matrix to obtain a reduced shape basis, that in contrast to existing approaches, can be physically interpreted. In turn, these precomputed shape bases are used to linearly span the deformation of a wide variety of objects. We introduce the low-rank basis into a sequential approach to recover both camera motion and nonrigid shape from the monocular video, by simply optimizing the weights of the linear combination using bundle adjustment. Since the number of parameters to optimize per frame is relatively small, specially when physical priors are considered, our approach is fast and can potentially run in real time. Validation is done in a wide variety of real-world objects, undergoing both inextensible and extensible deformations. Our approach achieves remarkable robustness to artifacts such as noisy and missing measurements and shows an improved performance to competing methods. Antonio Agudo, Francesc Moreno-Noguer |
IEEE Trans. Multim. | 1 |
| 2018 | Image Collection Pop-Up: 3D Reconstruction and Clustering of Rigid and Non-Rigid CategoriesabstractThis paper introduces an approach to simultaneously estimate 3D shape, camera pose, and object and type of deformation clustering, from partial 2D annotations in a multi-instance collection of images. Furthermore, we can indistinctly process rigid and non-rigid categories. This advances existing work, which only addresses the problem for one single object or, if multiple objects are considered, they are assumed to be clustered a priori. To handle this broader version of the problem, we model object deformation using a formulation based on multiple unions of subspaces, able to span from small rigid motion to complex deformations. The parameters of this model are learned via Augmented Lagrange Multipliers, in a completely unsupervised manner that does not require any training data at all. Extensive validation is provided in a wide variety of synthetic and real scenarios, including rigid and non-rigid categories with small and large deformations. In all cases our approach outperforms state-of-the-art in terms of 3D reconstruction accuracy, while also providing clustering results that allow segmenting the images into object instances and their associated type of deformation (or action the object is performing). Antonio Agudo, Melcior Pijoan, Francesc Moreno-Noguer |
CVPR | 1 |
| 2018 | Geometry-Aware Network for Non-Rigid Shape Prediction From a Single ViewabstractWe propose a method for predicting the 3D shape of a deformable surface from a single view. By contrast with previous approaches, we do not need a pre-registered template of the surface, and our method is robust to the lack of texture and partial occlusions. At the core of our approach is a geometry-aware deep architecture that tackles the problem as usually done in analytic solutions: first perform 2D detection of the mesh and then estimate a 3D shape that is geometrically consistent with the image. We train this architecture in an end-to-end manner using a large dataset of synthetic renderings of shapes under different levels of deformation, material properties, textures and lighting conditions. We evaluate our approach on a test split of this dataset and available real benchmarks, consistently improving state-of-the-art solutions with a significantly lower computational time. Albert Pumarola, Antonio Agudo, Lorenzo Porzi, Alberto Sanfeliu, Vincent Lepetit, Francesc Moreno-Noguer |
CVPR | 2 |
| 2018 | Unsupervised Person Image Synthesis in Arbitrary PosesabstractWe present a novel approach for synthesizing photorealistic images of people in arbitrary poses using generative adversarial learning. Given an input image of a person and a desired pose represented by a 2D skeleton, our model renders the image of the same person under the new pose, synthesizing novel views of the parts visible in the input image and hallucinating those that are not seen. This problem has recently been addressed in a supervised manner [16, 35], i.e., during training the ground truth images under the new poses are given to the network. We go beyond these approaches by proposing a fully unsupervised strategy. We tackle this challenging scenario by splitting the problem into two principal subtasks. First, we consider a pose conditioned bidirectional generator that maps back the initially rendered image to the original pose, hence being directly comparable to the input image without the need to resort to any training image. Second, we devise a novel loss function that incorporates content and style terms, and aims at producing images of high perceptual quality. Extensive experiments conducted on the DeepFashion dataset demonstrate that the images rendered by our model are very close in appearance to those obtained by fully supervised approaches. Albert Pumarola, Antonio Agudo, Alberto Sanfeliu, Francesc Moreno-Noguer |
CVPR | 2 |
| 2018 | GANimation: Anatomically-Aware Facial Animation from a Single Image
Albert Pumarola, Antonio Agudo, Aleix Martinez, Alberto Sanfeliu, Francesc Moreno-Noguer |
ECCV (10) | 2 |
| 2018 | Deformable Motion 3D Reconstruction by Union of Regularized SubspacesabstractThis paper presents an approach to jointly retrieve camera pose, time-varying 3D shape, and automatic clustering based on motion primitives, from incomplete 2D trajectories in a monocular video. We introduce the concept of order-varying temporal regularization in order to exploit video data, that can be indistinctly applied to the 3D shape evolution as well as to the similarities between images. This results in a union of regularized subspaces which effectively encodes the 3D shape deformation. All parameters are learned via augmented Lagrange multipliers, in a unified and unsupervised manner that does not assume any training data at all. Experimental validation is reported on human motion from sparse to dense shapes, providing more robust and accurate solutions than state-of-the-art approaches in terms of 3D reconstruction, while also obtaining motion grouping results. Antonio Agudo, Francesc Moreno-Noguer |
ICIP | 1 |
| 2018 | 2D-to-3D Facial Expression TransferabstractAutomatically changing the expression and physical features of a face from an input image is a topic that has been traditionally tackled in a 2D domain. In this paper, we bring this problem to 3D and propose a framework that given an input RGB video of a human face under a neutral expression, initially computes his/her 3D shape and then performs a transfer to a new and potentially non-observed expression. For this purpose, we parameterize the rest shape -obtained from standard factorization approaches over the input video- using a triangular mesh which is further clustered into larger macro-segments. The expression transfer problem is then posed as a direct mapping between this shape and a source shape, such as the blend shapes of an off-the-shelf 3D dataset of human facial expressions. The mapping is resolved to be geometrically consistent between 3D models by requiring points in specific regions to map on semantic equivalent regions. We validate the approach on several synthetic and real examples of input faces that largely differ from the source shapes, yielding very realistic expression transfers even in cases with topology changes, such as a synthetic video sequence of a single-eyed cyclops. Gemma Rotger, Felipe Lumbreras, Francesc Moreno-Noguer, Antonio Agudo |
ICPR | 4 |
| 2018 | A scalable, efficient, and accurate solution to non-rigid structure from motion
Antonio Agudo, Francesc Moreno-Noguer |
Comput. Vis. Image Underst. | 1 |
| 2018 | Force-Based Representation for Non-Rigid Shape and Elastic Model EstimationabstractThis paper addresses the problem of simultaneously recovering 3D shape, pose and the elastic model of a deformable object from only 2D point tracks in a monocular video. This is a severely under-constrained problem that has been typically addressed by enforcing the shape or the point trajectories to lie on low-rank dimensional spaces. We show that formulating the problem in terms of a low-rank force space that induces the deformation and introducing the elastic model as an additional unknown, allows for a better physical interpretation of the resulting priors and a more accurate representation of the actual object's behavior. In order to simultaneously estimate force, pose, and the elastic model of the object we use an expectation maximization strategy, where each of these parameters are successively learned by partial M-steps. Once the elastic model is learned, it can be transfered to similar objects to code its 3D deformation. Moreover, our approach can robustly deal with missing data, and encode both rigid and non-rigid points under the same formalism. We thoroughly validate the approach on Mocap and real sequences, showing more accurate 3D reconstructions than state-of-the-art, and additionally providing an estimate of the full elastic model with no a priori information. Antonio Agudo, Francesc Moreno-Noguer |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | DUST: Dual Union of Spatio-Temporal Subspaces for Monocular Multiple Object 3D ReconstructionabstractWe present an approach to reconstruct the 3D shape of multiple deforming objects from incomplete 2D trajectories acquired by a single camera. Additionally, we simultaneously provide spatial segmentation (i.e., we identify each of the objects in every frame) and temporal clustering (i.e., we split the sequence into primitive actions). This advances existing work, which only tackled the problem for one single object and non-occluded tracks. In order to handle several objects at a time from partial observations, we model point trajectories as a union of spatial and temporal subspaces, and optimize the parameters of both modalities, the non-observed point tracks and the 3D shape via augmented Lagrange multipliers. The algorithm is fully unsupervised and results in a formulation which does not need initialization. We thoroughly validate the method on challenging scenarios with several human subjects performing different activities which involve complex motions and close interaction. We show our approach achieves state-of-the-art 3D reconstruction results, while it also provides spatial and temporal segmentation. Antonio Agudo, Francesc Moreno-Noguer |
CVPR | 1 |
| 2017 | PL-SLAM: Real-time monocular visual SLAM with points and linesabstractLow textured scenes are well known to be one of the main Achilles heels of geometric computer vision algorithms relying on point correspondences, and in particular for visual SLAM. Yet, there are many environments in which, despite being low textured, one can still reliably estimate line-based geometric primitives, for instance in city and indoor scenes, or in the so-called “Manhattan worlds”, where structured edges are predominant. In this paper we propose a solution to handle these situations. Specifically, we build upon ORB-SLAM, presumably the current state-of-the-art solution both in terms of accuracy as efficiency, and extend its formulation to simultaneously handle both point and line correspondences. We propose a solution that can even work when most of the points are vanished out from the input images, and, interestingly it can be initialized from solely the detection of line correspondences in three consecutive frames. We thoroughly evaluate our approach and the new initialization strategy on the TUM RGB-D benchmark and demonstrate that the use of lines does not only improve the performance of the original ORB-SLAM solution in poorly textured frames, but also systematically improves it in sequence frames combining points and lines, without compromising the efficiency. Albert Pumarola, Alexander Vakhitov, Antonio Agudo, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICRA | 3 |
| 2017 | Global Model with Local Interpretation for Dynamic Shape ReconstructionabstractThe most standard approach to resolve the inherent ambiguities of the non-rigid structure from motion problem is using low-rank models that approximate deforming shapes by a linear combination of rigid basis. These models are typically global, i.e., each shape basis contributes equally to all points of the surface. While this approach has been shown effective to represent smooth deformations, its performance degrades for surfaces composed of various regions, each following a different deformation rule. Piece-wise methods attempt to capture this type of behavior by locally modeling surface patches, although they subsequently require enforcing global constraints to assemble back the patches. In this paper we propose an approach that combines the best of global and local models: it locally considers low-rank models but, by construction, does not need to impose global constraints to guarantee local patch continuity. We achieve this by a simple expectation maximization strategy that besides learning global shape bases, it locally adapts their contribution to each specific surface region. Furthermore, as a side contribution, in order to split the surface into different local patches, we propose a novel physically-based mesh segmentation approach that obeys an energy criterion. The complete framework is evaluated in both synthetic and real datasets, and shows an improved performance to competing methods. Antonio Agudo, Francesc Moreno-Noguer |
WACV | 1 |
| 2017 | Combining Local-Physical and Global-Statistical Models for Sequential Deformable Shape from Motion
Antonio Agudo, Francesc Moreno-Noguer |
Int. J. Comput. Vis. | 1 |
| 2016 | Recovering Pose and 3D Deformable Shape from Multi-instance Image Ensembles
Antonio Agudo, Francesc Moreno-Noguer |
ACCV (4) | 1 |
| 2016 | Mode-shape interpretation: Re-thinking modal space for recovering deformable shapesabstractThis paper describes an on-line approach for estimating non-rigid shape and camera pose from monocular video sequences. We assume an initial estimate of the shape at rest to be given and represented by a triangulated mesh, which is encoded by a matrix of the distances between every pair of vertexes. By applying spectral analysis on this matrix, we are then able to compute a low-dimensional shape basis, that in contrast to standard approaches, has a very direct physical interpretation and requires a much smaller number of modes to span a large variety of deformations, either for inextensible or extensible configurations. Based on this low-rank model, we then sequentially retrieve both camera motion and non-rigid shape in each image, optimizing the model parameters with bundle adjustment over a sliding window of image frames. Since the number of these parameters is small, specially when considering physical priors, our approach may potentially achieve real-time performance. Experimental results on real videos for different scenarios demonstrate remarkable robustness to artifacts such as missing and noisy observations. Antonio Agudo, J. M. M. Montiel, Begoña Calvo, Francesc Moreno-Noguer |
WACV | 1 |
| 2016 | Real-time 3D reconstruction of non-rigid shapes with a single moving camera
Antonio Agudo, Francesc Moreno-Noguer, Begoña Calvo, J. M. M. Montiel |
Comput. Vis. Image Underst. | 1 |
| 2016 | Sequential Non-Rigid Structure from Motion Using Physical PriorsabstractWe propose a new approach to simultaneously recover camera pose and 3D shape of non-rigid and potentially extensible surfaces from a monocular image sequence. For this purpose, we make use of the Extended Kalman Filter based Simultaneous Localization And Mapping (EKF-SLAM) formulation, a Bayesian optimization framework traditionally used in mobile robotics for estimating camera pose and reconstructing rigid scenarios. In order to extend the problem to a deformable domain we represent the object's surface mechanics by means of Navier's equations, which are solved using a Finite Element Method (FEM). With these main ingredients, we can further model the material's stretching, allowing us to go a step further than most of current techniques, typically constrained to surfaces undergoing isometric deformations. We extensively validate our approach in both real and synthetic experiments, and demonstrate its advantages with respect to competing methods. More specifically, we show that besides simultaneously retrieving camera pose and non-rigid shape, our approach is adequate for both isometric and extensible surfaces, does not require neither batch processing all the frames nor tracking points over the whole sequence and runs at several frames per second. Antonio Agudo, Francesc Moreno-Noguer, Begoña Calvo, J. M. M. Montiel |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Simultaneous pose and non-rigid shape with particle dynamicsabstractIn this paper, we propose a sequential solution to simultaneously estimate camera pose and non-rigid 3D shape from a monocular video. In contrast to most existing approaches that rely on global representations of the shape, we model the object at a local level, as an ensemble of particles, each ruled by the linear equation of the Newton's second law of motion. This dynamic model is incorporated into a bundle adjustment framework, in combination with simple regularization components that ensure temporal and spatial consistency of the estimated shape and camera poses. The resulting approach is both efficient and robust to several artifacts such as noisy and missing data or sudden camera motions, while it does not require any training data at all. Validation is done in a variety of real video sequences, including articulated and non-rigid motion, both for continuous and discontinuous shapes. Our system is shown to perform comparable to competing batch, computationally expensive, methods and shows remarkable improvement with respect to the sequential ones. Antonio Agudo, Francesc Moreno-Noguer |
CVPR | 1 |
| 2015 | Learning Shape, Motion and Elastic Models in Force SpaceabstractIn this paper, we address the problem of simultaneously recovering the 3D shape and pose of a deformable and potentially elastic object from 2D motion. This is a highly ambiguous problem typically tackled by using low-rank shape and trajectory constraints. We show that formulating the problem in terms of a low-rank force space that induces the deformation, allows for a better physical interpretation of the resulting priors and a more accurate representation of the actual object's behavior. However, this comes at the price of, besides force and pose, having to estimate the elastic model of the object. For this, we use an Expectation Maximization strategy, where each of these parameters are successively learned within partial M-steps, while robustly dealing with missing observations. We thoroughly validate the approach on both mocap and real sequences, showing more accurate 3D reconstructions than state-of-the-art, and additionally providing an estimate of the full elastic model with no a priori information. Antonio Agudo, Francesc Moreno-Noguer |
ICCV | 1 |
| 2014 | Online Dense Non-Rigid 3D Shape and Camera Motion Recovery
Antonio Agudo, J. M. M. Montiel, Lourdes Agapito, Begoña Calvo |
BMVC | 1 |
| 2014 | Good Vibrations: A Modal Analysis Approach for Sequential Non-rigid Structure from MotionabstractWe propose an online solution to non-rigid structure from motion that performs camera pose and 3D shape estimation of highly deformable surfaces on a frame-by-frame basis. Our method models non-rigid deformations as a linear combination of some mode shapes obtained using modal analysis from continuum mechanics. The shape is first discretized into linear elastic triangles, modelled by means of finite elements, which are used to pose the force balance equations for an undamped free vibrations model. The shape basis computation comes down to solving an eigenvalue problem, without the requirement of a learning step. The camera pose and time varying weights that define the shape at each frame are then estimated on the fly, in an online fashion, using bundle adjustment over a sliding window of image frames. The result is a low computational cost method that can run sequentially in real-time. We show experimental results on synthetic sequences with ground truth 3D data and real videos for different scenarios ranging from sparse to dense scenes. Our system exhibits a good trade-off between accuracy and computational budget, it can handle missing data and performs favourably compared to competing methods. Antonio Agudo, Lourdes Agapito, Begoña Calvo, J. M. M. Montiel |
CVPR | 1 |
| 2012 | Finite Element based sequential Bayesian Non-Rigid Structure from MotionabstractNavier's equations modelling linear elastic solid deformations are embedded within an Extended Kalman Filter (EKF) to compute a sequential Bayesian estimate for the Non-Rigid Structure from Motion problem. The algorithm processes every single frame of a sequence gathered with a full perspective camera. No prior data association is assumed because matches are computed within the EKF prediction-match-update cycle. Scene is coded as a Finite Element Method (FEM) elastic thin-plate solid, where the discretization nodes are the sparse set of scene points salient in the image. It is assumed a set of Gaussian forces acting on solid nodes to cause scene deformation. The EKF combines in a feedback loop an approximate FEM model and the frame rate measurements from the camera, resulting in an efficient method to embed Navier's equations without resorting to expensive non-linear FEM models. Classical FEM modelling has implied an interactive identification of boundary points to constrain the scene rigid motion, in this work this dissatisfying prior knowledge is no longer needed. The scene and camer rigid motion are combined in a unique pose vector and the estimation is coded relative to the camera. Additionally, the deforming effect of the Gaussian forces on the thin-plate is computed by means of the Moore-Penrose pseudoinverse of the FEM stiffness matrix. The proposed algorithm is validated with three real sequences gathered with hand-held camera observing isometric and non-isometric deformations. It is also shown the consistency of the EKF estimation with respect to ground truth computed from stereo. Antonio Agudo, Begoña Calvo, J. M. M. Montiel |
CVPR | 1 |