VLDB 2026 Research / reviewers in the wild / expert
Luis Ferraz
dblp:13/1378 · also Luis Ferraz Colomina
· DBLP profile ↗
13ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0001-7851-9193ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 3 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Modal Soccer Scene Analysis with Masked Pre-TrainingabstractIn this work we propose a multi-modal architecture for analyzing soccer scenes from tactical camera footage, with a focus on three core tasks: ball trajectory inference, ball state classification, and ball possessor identification. To this end, our solution integrates three distinct input modalities (player trajectories, player types and image crops of individual players) into a unified framework that processes spatial and temporal dynamics using a cascade of sociotemporal transformer blocks. Unlike prior methods, which rely heavily on accurate ball tracking or handcrafted heuristics, our approach infers the ball trajectory without direct access to its past or future positions, and robustly identifies the ball state and ball possessor under noisy or occluded conditions from real top league matches. We also introduce CropDrop, a modality-specific masking pre-training strategy that prevents over-reliance on image features and encourages the model to rely on cross-modal patterns during pre-training. We show the effectiveness of our approach on a large-scale dataset providing substantial improvements over state-of-the-art baselines in all tasks. Our results highlight the benefits of combining structured and visual cues in a transformer-based architecture, and the importance of realistic masking strategies in multi-modal learning. Marc Peral, Guillem Capellera, Luis Ferraz, Antonio Rubio 0001, Antonio Agudo |
WACV | 3 |
| 2025 | Unified Uncertainty-Aware Diffusion for Multi-Agent Trajectory ModelingabstractMulti-agent trajectory modeling has primarily focused on forecasting future states, often overlooking broader tasks like trajectory completion, which are crucial for real-world applications such as correcting tracking data. Existing methods also generally predict agents’ states without offering any state-wise measure of uncertainty. Moreover, popular multi-modal sampling methods lack any error probability estimates for each generated scene under the same prior observations, making it difficult to rank the predictions during inference time. We introduce U2Diff, a unified diffusion model designed to handle trajectory completion while providing state-wise uncertainty estimates jointly. This uncertainty estimation is achieved by augmenting the simple denoising loss with the negative log-likelihood of the predicted noise and propagating latent space uncertainty to the real state space. Additionally, we incorporate a Rank Neural Network in post-processing to enable error probability estimation for each generated mode, demonstrating a strong correlation with the error relative to ground truth. Our method outperforms the state-of-the-art solutions in trajectory completion and forecasting across four challenging sports datasets (NBA, Basketball-U, Football-U, Soccer-U), highlighting the effectiveness of uncertainty and error probability estimation. Guillem Capellera, Antonio Rubio 0001, Luis Ferraz, Antonio Agudo |
CVPR | 3 |
| 2024 | TranSPORTmer: A Holistic Approach to Trajectory Understanding in Multi-agent Sports
Guillem Capellera, Luis Ferraz, Antonio Rubio 0001, Antonio Agudo, Francesc Moreno-Noguer |
ACCV (2) | 2 |
| 2024 | Footbots: A Transformer-Based Architecture for Motion Prediction in SoccerabstractMotion prediction in soccer involves capturing complex dynamics from player and ball interactions. We present FootBots, an encoder-decoder transformer-based architecture addressing motion prediction and conditioned motion prediction through equivariance properties. FootBots captures temporal and social dynamics using set attention blocks and multi-attention block decoder. Our evaluation utilizes two datasets: a real soccer dataset and a tailored synthetic one. Insights from the synthetic dataset highlight the effectiveness of FootBots’ social attention mechanism and the significance of conditioned motion prediction. Empirical results on real soccer data demonstrate that FootBots outperforms baselines in motion prediction and excels in conditioned tasks, such as predicting the players based on the ball position, predicting the offensive (defensive) team based on the ball and the defensive (offensive) team, and predicting the ball position based on all players. Our evaluation connects quantitative and qualitative findings. https://youtu.be/9kaEkfzG3L8 Guillem Capellera, Luis Ferraz, Antonio Rubio 0001, Antonio Agudo, Francesc Moreno-Noguer |
ICIP | 2 |
| 2021 | Uncertainty-Aware Camera Pose Estimation From Points and LinesabstractPerspective-n-Point-and-Line (PnPL) algorithms aim at fast, accurate, and robust camera localization with respect to a 3D model from 2D-3D feature correspondences, being a major part of modern robotic and AR/VR systems. Current point-based pose estimation methods use only 2D feature detection uncertainties, and the line-based methods do not take uncertainties into account. In our setup, both 3D co-ordinates and 2D projections of the features are considered uncertain. We propose PnP(L) solvers based on EPnP [20] and DLS [14] for the uncertainty-aware pose estimation. We also modify motion-only bundle adjustment to take 3D uncertainties into account. We perform exhaustive synthetic and real experiments on two different visual odometry datasets. The new PnP(L) methods outperform the state-of-the-art on real data in isolation, showing an increase in mean translation accuracy by 18% on a representative subset of KITTI, while the new uncertain refinement improves pose accuracy for most of the solvers, e.g. decreasing mean translation error for the EPnP by 16% compared to the standard refinement on the same dataset. The code is available at https://alexandervakhitov.github.io/uncertain-pnp/. Alexander Vakhitov, Luis Ferraz, Antonio Agudo, Francesc Moreno-Noguer |
CVPR | 2 |
| 2015 | Discriminative Learning of Deep Convolutional Feature Point DescriptorsabstractDeep learning has revolutionalized image-level tasks such as classification, but patch-level tasks, such as correspondence, still rely on hand-crafted features, e.g. SIFT. In this paper we use Convolutional Neural Networks (CNNs) to learn discriminant patch representations and in particular train a Siamese network with pairs of (non-)corresponding patches. We deal with the large number of potential pairs with the combination of a stochastic sampling of the training set and an aggressive mining strategy biased towards patches that are hard to classify. By using the L2 distance during both training and testing we develop 128-D descriptors whose euclidean distances reflect patch similarity, and which can be used as a drop-in replacement for any task involving SIFT. We demonstrate consistent performance gains over the state of the art, and generalize well against scaling and rotation, perspective transformation, non-rigid deformation, and illumination changes. Our descriptors are efficient to compute and amenable to modern GPUs, and are publicly available. Edgar Simo-Serra, Eduard Trulls, Luis Ferraz, Iasonas Kokkinos, Pascal Fua, Francesc Moreno-Noguer |
ICCV | 3 |
| 2015 | Efficient monocular pose estimation for complex 3D modelsabstractWe propose a robust and efficient method to estimate the pose of a camera with respect to complex 3D textured models of the environment that can potentially contain more than 100; 000 points. To tackle this problem we follow a top down approach where we combine high-level deep network classifiers with low level geometric approaches to come up with a solution that is fast, robust and accurate. Given an input image, we initially use a pre-trained deep network to compute a rough estimation of the camera pose. This initial estimate constrains the number of 3D model points that can be seen from the camera viewpoint. We then establish 3D-to-2D correspondences between these potentially visible points of the model and the 2D detected image features. Accurate pose estimation is finally obtained from the 2D-to-3D correspondences using a novel PnP algorithm that rejects outliers without the need to use a RANSAC strategy, and which is between 10 and 100 times faster than other methods that use it. Two real experiments dealing with very large and complex 3D models demonstrate the effectiveness of the approach. Antonio Rubio 0001, Michael Villamizar, Luis Ferraz, Adrián Peñate Sánchez, Arnau Ramisa, Edgar Simo-Serra, Alberto Sanfeliu, Francesc Moreno-Noguer |
ICRA | 3 |
| 2014 | Leveraging Feature Uncertainty in the PnP Problem
Luis Ferraz, Xavier Binefa, Francesc Moreno-Noguer |
BMVC | 1 |
| 2014 | Very Fast Solution to the PnP Problem with Algebraic Outlier RejectionabstractWe propose a real-time, robust to outliers and accurate solution to the Perspective-n-Point (PnP) problem. The main advantages of our solution are twofold: first, it in- tegrates the outlier rejection within the pose estimation pipeline with a negligible computational overhead, and sec- ond, its scalability to arbitrarily large number of correspon- dences. Given a set of 3D-to-2D matches, we formulate pose estimation problem as a low-rank homogeneous sys- tem where the solution lies on its 1D null space. Outlier correspondences are those rows of the linear system which perturb the null space and are progressively detected by projecting them on an iteratively estimated solution of the null space. Since our outlier removal process is based on an algebraic criterion which does not require computing the full-pose and reprojecting back all 3D points on the image plane at each step, we achieve speed gains of more than 100× compared to RANSAC strategies. An extensive exper- imental evaluation will show that our solution yields accu- rate results in situations with up to 50% of outliers, and can process more than 1000 correspondences in less than 5ms. Luis Ferraz, Xavier Binefa, Francesc Moreno-Noguer |
CVPR | 1 |
| 2012 | Fast and robust monocular 3D deformable shape estimation for inextensible and smooth surfaces
Luis Ferraz, Xavier Binefa |
ICPR | 1 |
| 2012 | Dissociating rigid and articulated motion for hand tracking
Oriol Martínez, Pol Cirujeda, Luis Ferraz, Xavier Binefa |
ICPR | 3 |
| 2012 | A sparse curvature-based detector of affine invariant blobs
Luis Ferraz, Xavier Binefa |
Comput. Vis. Image Underst. | 1 |
| 2006 | Multiple Kernel Two-Step TrackingabstractIn tracking tasks, representing a target region as a weighted histogram has opened possibilities which led to excellent results, as mean shift or camshift algorithms. This representation is extracted from the image by giving weights with kernels and it depends on the properties of the kernels. By a first order Taylor approximation of the histograms it is possible to perform a tracking using several kernels, interpreted as different sources of information. This representation improves the possibilities and gives more flexibility when facing problems of tracking, as occlusions, model variance or projective deformations of the image. In this paper we use this multi-kernel model representation to perform a simultaneous tracking of the entire object and also of each different part individually. This is performed in a new two-step process. In the first step we perform the multikernel estimation and in a second step we update the model representation taking into account single kernel estimations, representing the local movement of each part. From a probabilistic view of the Matusita metric we analyze the usefulness of this method against partial occlusions and some projective transformations like zooms or 3D rotations and articulated movements. Brais Martínez, Luis Ferraz, Xavier Binefa, Jose Díaz-Caro |
ICIP | 2 |