VLDB 2026 Research / reviewers in the wild / expert
Francisco Vicente 0001
dblp:165/8170-1 · also Francisco Vicente Carrasco 0001
· DBLP profile ↗
10ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0005-3680-3029ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Blurry to Believable: Enhancing Low-Quality Talking Heads with 3D Generative PriorsabstractCreating high-fidelity, animatable 3D talking heads is crucial for immersive applications, yet often hindered by the prevalence of low-quality image or video sources, which yield poor 3D reconstructions. In this paper, we introduce SuperHead, a novel framework for enhancing lowresolution, animatable 3D head avatars. The core challenge lies in synthesizing high-quality geometry and textures, while ensuring both 3D and temporal consistency during animation and preserving subject identity. Despite recent progress in image, video and 3D-based superresolution ($S R$), existing$S R$techniques are ill-equipped to handle dynamic 3D inputs. To address this, SuperHead leverages the rich priors from pre-trained 3D generative models via a novel dynamics-aware 3D inversion scheme. This process optimizes the latent representation of the generative model to produce a super-resolved 3D Gaussian Splatting (3DGS) head model, which is subsequently rigged to an underlying parametric head model (e.g., FLAME) for animation. The inversion is jointly supervised using a sparse collection of upscaled 2D face renderings and corresponding depth maps, captured from diverse facial expressions and camera viewpoints, to ensure realism under$d y$namic facial motions. Experiments demonstrate that SuperHead generates avatars with fine-grained facial details under dynamic motions, significantly outperforming baseline methods in visual quality. Ding-Jiun Huang, Yuanhao Wang 0011, Shao-Ji Yuan, Albert Mosella-Montoro, Francisco Vicente 0001, Cheng Zhang 0014, Fernando De la Torre |
3DV | 5 |
| 2025 | Echoes of the Coliseum: Towards 3D Live streaming of Sports EventsabstractHuman-centered live events have always played a pivotal role in shaping culture and fostering social connections. Traditional 2D live transmissions fail to replicate the immersive quality of physical attendance. Addressing this gap, this paper proposes LiveSplats , a framework towards real-time, photo-realistic 3D reconstructions of live events using high-performance 3D Gaussian Splatting. Our solution capitalizes on strong geometric priors to optimize through distributed processing and load balancing, enabling interactive, freely explorable 3D experiences. By dividing scene reconstruction into actor-centric and environment-specific tasks, we employ hierarchical coarse-to-fine optimization to rapidly and accurately reconstruct human actors based on pose data, refining their geometry and appearance with photometric loss. For static environments, we focus on view-dependent appearance changes, streamlining rendering efficiency and maximizing GPU performance. To facilitate evaluation, we introduce (and distribute) a synthetic benchmark dataset of basketball games, offering high visual fidelity as ground truth. In both our synthetic benchmark and publicly available benchmarks, LiveSplats consistently outperforms existing approaches. The dataset is available at https://humansensinglab.github.io/basket-multiview. Junkai Huang 0004, Saswat Subhajyoti Mallick, Alejandro Amat, Marc Ruiz Olle, Albert Mosella-Montoro, Bernhard Kerbl, Francisco Vicente 0001, Fernando De la Torre |
ACM Trans. Graph. | 7 |
| 2024 | Generalizable Human Gaussians for Sparse View Synthesis
Youngjoong Kwon, Baole Fang, Yixing Lu, Haoye Dong, Cheng Zhang 0014, Francisco Vicente 0001, Albert Mosella-Montoro, Jianjin Xu, Shingo Takagi 0001, Daeil Kim, Aayush Prakash, Fernando De la Torre |
ECCV (78) | 6 |
| 2024 | Doubly Hierarchical Geometric Representations for Strand-based Human Hairstyle GenerationabstractWe introduce a doubly hierarchical generative representation for strand-based 3D hairstyle geometry that progresses from coarse, low-pass filtered guide hair to densely populated hair strands rich in high-frequency details. We employ the Discrete Cosine Transform (DCT) to separate low-frequency structural curves from high-frequency curliness and noise, avoiding the Gibbs' oscillation issues associated with the standard Fourier transform in open curves. Unlike the guide hair sampled from the scalp UV map grids which may lose capturing details of the hairstyle in existing methods, our method samples optimal sparse guide strands by utilising $k$-medoids clustering centres from low-pass filtered dense strands, which more accurately retain the hairstyle's inherent characteristics. The proposed variational autoencoder-based generation network, with an architecture inspired by geometric deep learning and implicit neural representations, facilitates flexible, off-the-grid guide strand modelling and enables the completion of dense strands in any quantity and density, drawing on principles from implicit neural representations. Empirical evaluations confirm the capacity of the model to generate convincing guide hair and dense strands, complete with nuanced high-frequency details. Yunlu Chen, Francisco Vicente 0001, Christian Häne, Giljoo Nam, Jean-Charles Bazin, Fernando De la Torre |
NeurIPS | 2 |
| 2024 | Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning Mambaabstract3D Hand reconstruction from a single RGB image is challenging due to the articulated motion, self-occlusion, and interaction with objects. Existing SOTA methods employ attention-based transformers to learn the 3D hand pose and shape, yet they do not fully achieve robust and accurate performance, primarily due to inefficiently modeling spatial relations between joints. To address this problem, we propose a novel graph-guided Mamba framework, named Hamba, which bridges graph learning and state space modeling. Our core idea is to reformulate Mamba's scanning into graph-guided bidirectional scanning for 3D reconstruction using a few effective tokens. This enables us to efficiently learn the spatial relationships between joints for improving reconstruction performance. Specifically, we design a Graph-guided State Space (GSS) block that learns the graph-structured relations and spatial sequences of joints and uses 88.5\% fewer tokens than attention-based methods. Additionally, we integrate the state space features and the global features using a fusion module. By utilizing the GSS block and the fusion module, Hamba effectively leverages the graph-guided state space features and jointly considers global and local features to improve performance. Experiments on several benchmarks and in-the-wild tests demonstrate that Hamba significantly outperforms existing SOTAs, achieving the PA-MPVPE of 5.3mm and F@15mm of 0.992 on FreiHAND. At the time of this paper's acceptance, Hamba holds the top position, Rank 1, in two competition leaderboards on 3D hand reconstruction. Haoye Dong, Aviral Chharia, Wenbo Gou, Francisco Vicente 0001, Fernando De la Torre |
NeurIPS | 4 |
| 2024 | Taming 3DGS: High-Quality Radiance Fields with Limited Resources
Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl, Markus Steinberger, Francisco Vicente 0001, Fernando De la Torre |
SIGGRAPH Asia | 5 |
| 2024 | FabricDiffusion: High-Fidelity Texture Transfer for 3D Garments Generation from In-The-Wild Images
Cheng Zhang 0014, Yuanhao Wang 0011, Francisco Vicente 0001, Chenglei Wu, Thabo Beeler, Fernando De la Torre |
SIGGRAPH Asia | 3 |
| 2024 | Towards Realistic Generative 3D Face ModelsabstractIn recent years, there has been significant progress in 2D generative face models fueled by applications such as animation, synthetic data generation, and digital avatars. However, due to the absence of 3D information, these 2D models often struggle to accurately disentangle facial attributes like pose, expression, and illumination, limiting their editing capabilities. To address this limitation, this paper proposes a 3D controllable generative face model to produce high-quality albedo and precise 3D shapes by leveraging existing 2D generative models. By combining 2D face generative models with semantic face manipulation, this method enables editing of detailed 3D rendered faces. The proposed framework utilizes an alternating descent optimization approach over shape and albedo. Differentiable rendering is used to train high-quality shapes and albedo without 3D supervision. Moreover, this approach outperforms most state-of-the-art (SOTA) methods in the well-known NoW and REALY benchmarks for 3D face re construction. It also outperforms the SOTA reconstruction models in recovering rendered faces’ identities across novel poses. Additionally, the paper demonstrates direct control of expressions in 3D faces by exploiting latent space leading to text-based editing of 3D faces. Aashish Rai, Hiresh Gupta, Francisco Vicente 0001, Shingo Takagi 0001, Amaury Aubel, Daeil Kim, Aayush Prakash, Fernando De la Torre |
WACV | 4 |
| 2019 | Road Curb Detection and Localization With Monocular Forward-View Vehicle CameraabstractWe propose a robust method for estimating road curb 3-D parameters (size, location, and orientation) using a calibrated monocular camera equipped with a fisheye lens. Automatic curb detection and localization is particularly important in the context of an advanced driver assistance system, i.e., to prevent possible collision and damage to the vehicle's bumper during perpendicular and diagonal parking maneuvers. Combining 3-D geometric reasoning with advanced vision-based detection methods, our approach is able to estimate the vehicle to curb distance in real time with a mean accuracy of more than 90%, as well as its orientation, height, and depth. Our approach consists of two distinct components-curb detection in each individual video frame and temporal analysis. The first part is comprised of sophisticated curb edges extraction and parameterized 3-D curb template fitting. Using a few assumptions regarding the real-world geometry, we can thus retrieve the curb's height and its relative position with respect to the moving vehicle on which the camera is mounted. Support vector machine classifier fed with histograms of oriented gradients is used for appearance-based filtering out outliers. In the second part, the detected curb regions are tracked in the temporal domain, so as to perform a second pass of false positives rejection. We have validated our approach on a newly collected database of 11 videos under different conditions. We have used point-wise LIDAR measurements and manual exhaustive labels as a ground truth. Stanislav Panev, Francisco Vicente 0001, Fernando De la Torre, Véronique Prinet |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Driver Gaze Tracking and Eyes Off the Road Detection SystemabstractDistracted driving is one of the main causes of vehicle collisions in the United States. Passively monitoring a driver's activities constitutes the basis of an automobile safety system that can potentially reduce the number of accidents by estimating the driver's focus of attention. This paper proposes an inexpensive vision-based system to accurately detect Eyes Off the Road (EOR). The system has three main components: 1) robust facial feature tracking; 2) head pose and gaze estimation; and 3) 3-D geometric reasoning to detect EOR. From the video stream of a camera installed on the steering wheel column, our system tracks facial features from the driver's face. Using the tracked landmarks and a 3-D face model, the system computes head pose and gaze direction. The head pose estimation algorithm is robust to nonrigid face deformations due to changes in expressions. Finally, using a 3-D geometric analysis, the system reliably detects EOR. Francisco Vicente 0001, Zehua Huang, Xuehan Xiong, Fernando De la Torre, Wende Zhang, Dan Levi |
IEEE Trans. Intell. Transp. Syst. | 1 |