Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Adrián Peñate Sánchez

dblp:59/5985 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 3 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Visual content generation and editing · 54% Rendering · 46%
Artificial intelligence
4 papers
3D vision · 76% Segmentation and scene understanding · 12% Image recognition and object detection · 12%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing › image editing
color editing
0.812024
IReNe: Instant Recoloring of Neural Radiance Fields · CVPR 2024
Visual content generation and editing › 3d content editing
neural radiance field editing
0.812024
IReNe: Instant Recoloring of Neural Radiance Fields · CVPR 2024
Rendering
neural radiance fields
0.712023
NeRFLight: Fast and Light Neural Radiance Fields using a Shared Feature Grid · CVPR 2023
Rendering
real-time rendering
0.712023
NeRFLight: Fast and Light Neural Radiance Fields using a Shared Feature Grid · CVPR 2023
Computer vision › 3D vision
camera pose estimation
0.422015
Efficient monocular pose estimation for complex 3D models · ICRA 2015
Exhaustive Linearization for Robust Camera Pose and Focal Length Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Computer vision › Segmentation and scene understanding › boundary detection
object boundary segmentation
0.212024
IReNe: Instant Recoloring of Neural Radiance Fields · CVPR 2024
Computer vision › 3D vision › pose estimation
monocular pose estimation
0.212015
Efficient monocular pose estimation for complex 3D models · ICRA 2015
Computer vision › 3D vision
object pose estimation
0.212015
A dynamic programming approach for fast and robust object pose recognition from range images · CVPR 2015
Computer vision › Image recognition and object detection
object recognition
0.212015
A dynamic programming approach for fast and robust object pose recognition from range images · CVPR 2015
Computer vision › 3D vision › 3d object recognition
range image object recognition
0.212015
A dynamic programming approach for fast and robust object pose recognition from range images · CVPR 2015
Computer vision › 3D vision
camera calibration
0.212013
Exhaustive Linearization for Robust Camera Pose and Focal Length Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Computer vision › 3D vision › camera calibration
focal length estimation
0.212013
Exhaustive Linearization for Robust Camera Pose and Focal Length Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Algorithms and data structures
dynamic programming
0.112015
A dynamic programming approach for fast and robust object pose recognition from range images · CVPR 2015

Methods — techniques the papers use, named apart from their topics

neuron classification · 1.5neural radiance field retraining · 1.5feature grid · 0.7decoders · 0.7dynamic programming · 0.4data parallelism · 0.4belief propagation · 0.4exhaustive linearization · 0.3EPnP · 0.3pnp algorithm · 0.2outlier rejection · 0.2deep network classifier · 0.2exhaustive relinearization · 0.2
YearPublicationVenuePosition
2026 G1NDiff: Fast 3D scene editing with reprojection-conditioned GAN-diffusion training on 2D datasets
abstract
• G1NDiff: Reprojection-Aware One-Step Diffusion for Fast Multi-View Consistent 3D Scene Editing • G1NDiff: A one-step GAN-based image editing model that preserves consistency across successive edits. • Occlusion-aware Reprojection: An algorithm that transforms a single image into a pair of conditioning-target views. • Multi-stage Pipeline: A scene editing pipeline where G1NDiff is applied for gradually editing all views of a given 3D scene. • State of the art performance on 2DGS scenes, with higher MEt3R consistency and over 10× speedup.
Ivan Ojeda-Martin, Jorge Bustos-Sanchez, Alessio Mazzucchelli, Mario Alfonso-Arsuaga, Adrián Peñate Sánchez
Comput. Graph.5
2026 Automatic white shrimp (Penaeus vannamei) biometrical analysis from laboratory images using computer vision and deep learning
abstract
Manual morphological analysis for genetic selection in Penaeus vannamei aquaculture is a slow, error-prone bottleneck. We introduce Imashrimp, an automated system that uses colour and depth images to optimize this task by adapting deep learning and computer vision techniques to shrimp morphology. Imashrimp incorporates two discrimination modules to classify images by the point of view and determine rostrum integrity. These modules function as a “two-factor authentication” (human and Artificial Intelligence) system to validate annotations; this approach reduced metadata annotation errors, cutting point of view classification errors from 0.64% to 0% and rostrum integrity errors from 10.44% to 1.04%. A transformer-based pose estimation module predicts 23 keypoints on the shrimp’s skeleton, achieving a general Mean Average Precision of 96.84% and a Percentage of Correct Keypoints of 91.67%. The resulting Two-Dimensional measurements are transformed into Three-Dimensional measurements using a Support Vector Machine regression. By achieving a final Mean Absolute Error (MAE) of 0.08 ± 0.25 cm, IMASHRIMP demonstrates the potential to automate and accelerate shrimp morphological analysis, enhancing the efficiency of genetic selection and contributing to more sustainable aquaculture practices.
Abiam Remache González, Meriem Chagour, Timon Bijan Rüth, Raúl Trapiella Cañedo, Marina Martínez Soler, Álvaro Lorenzo Felipe, Hyun-Suk Shin, María-Jesús Zamorano Serrano, Ricardo Torres, Juan-Antonio Castillo Parra, Eduardo Reyes Abad, Miguel-Ángel Ferrer Ballester, Juan-Manuel Afonso López, Francisco-Mario Hernández Tejera, Adrián Peñate Sánchez
Eng. Appl. Artif. Intell.15
2026 An annotation assistant for monitoring the electrical grid using aerial images
abstract
Monitoring the electrical grid is essential to ensure reliable service and prevent accidents. This supervision is performed by aerial vehicles for image collection; later, these collected images are processed and analyzed by expert annotators. Due to the high costs of manually handling such as large datasets, we present a novel hybrid methodology that leverages deep learning to reduce and optimize annotation workload. The approach uses annotator-provided labels to train a neural network that makes annotation suggestions and gradually reduces the manual workload. Our work is closely related to active learning, but with a key difference: all data must be labeled and verified to guarantee correctness. Therefore, our methodology focuses on reducing the annotation time rather maximizing model performance. Our hybrid method assists annotators by suggesting annotations on high-confidence images that only need verification instead of being created from scratch. Using the proposed approach, annotators can complete their task at least 2.67x faster than with the previous fully manual labeling procedure.
Cristina Benlliure-Jimenez, Adrián Peñate Sánchez, Javier Lorenzo-Navarro, Modesto Castrillón-Santana, Francisco-Mario Hernández Tejera
Knowl. Based Syst.2
2024 IReNe: Instant Recoloring of Neural Radiance Fields
abstract
Advances in NERFs have allowed for 3D scene reconstructions and novel view synthesis. Yet, efficiently editing these representations while retaining photo realism is an emerging challenge. Recent methods face three primary limitations: they're slow for interactive use, lack precision at object boundaries, and struggle to ensure multi-view consistency. We introduce IReNe to address these limitations, enabling swift, near real-time color editing in NeRF. Leveraging a pre-trained NeRF model and a single training image with user-applied color edits, IReNe swiftly adjusts network parameters in seconds. This adjustment allows the model to generate new scene views, accurately representing the color changes from the training image while also controlling object boundaries and view-specific effects. Object boundary control is achieved by integrating a trainable segmentation module into the model. The process gains efficiency by retraining only the weights of the last network layer. We observed that neurons in this layer can be classified into those responsible for view-dependent appearance and those contributing to diffuse appearance. We introduce an automated classification approach to identify these neuron types and exclusively fine-tune the weights of the diffuse neurons. This further accelerates training and ensures consistent color edits across different views. A thorough validation on a new dataset, with edited object colors, shows significant quantitative and qualitative advancements over competitors, accelerating speeds by 5x to 500x.
Alessio Mazzucchelli, Adrian Garcia-Garcia, Elena Garces 0001, Fernando Rivas-Manzaneque, Francesc Moreno-Noguer, Adrián Peñate Sánchez
CVPR6
2023 NeRFLight: Fast and Light Neural Radiance Fields using a Shared Feature Grid
abstract
While original Neural Radiance Fields (NeRF) have shown impressive results in modeling the appearance of a scene with compact MLP architectures, they are not able to achieve real-time rendering. This has been recently addressed by either baking the outputs of NeRF into a data structure or arranging trainable parameters in an explicit feature grid. These strategies, however, significantly increase the memory footprint of the model which prevents their deployment on bandwidth-constrained applications. In this paper, we extend the grid-based approach to achieve real-time view synthesis at more than 150 FPS using a lightweight model. Our main contribution is a novel architecture in which the density field of NeRF-based representations is split into N regions and the density is modeled using N different decoders which reuse the same feature grid. This results in a smaller grid where each feature is located in more than one spatial position, forcing them to learn a compact representation that is valid for different parts of the scene. We further reduce the size of the final model by disposing of the features symmetrically on each region, which favors feature pruning after training while also allowing smooth gradient transitions between neighboring voxels. An exhaustive evaluation demonstrates that our method achieves real-time performance and quality metrics on a pair with state-of-the-art with an improvement of more than 2× in the FPS/MB ratio.
Fernando Rivas-Manzaneque, Jorge Sierra Acosta, Adrián Peñate Sánchez, Francesc Moreno-Noguer, Ángela Ribeiro
CVPR3
2023 A machine learning approach to design a DPSIR model: A real case implementation of evidence-based policy creation using AI
Adrián Peñate Sánchez, Carolina Peña Alonso, Emma Perez-Chacon Espino, Antonio Falcón-Martel
Adv. Eng. Informatics1
2021 Improving user verification in human-robot interaction from audio or image inputs through sample quality assessment
abstract
In this paper, we tackle the task of improving biometric verification in the context of Human-Robot Interaction (HRI). A robot that wants to identify a specific person to provide a service can do so by either image verification or, if light conditions are not favourable, through voice verification. In our approach, we will take advantage of the possibility a robot has of recovering further data until it is sure of the identity of the person. The key contribution is that we select from both image and audio signals the parts that are of higher confidence. For images we use a system that looks at the face of each person and selects frames in which the confidence is high while keeping those frames separate in time to avoid using very similar facial appearance . For audio our approach tries to find the parts of the signal that contain a person talking, avoiding those in which noise is present by segmenting the signal. Once the parts of interest are found, each input is described with an independent deep learning architecture that obtains a descriptor for each kind of input (face/voice). We also present in this paper fusion methods that improve performance by combining the features from both face and voice, results to validate this are shown for each independent input and for the fusion methods.
David Freire-Obregón, Kevin Rosales-Santana, Pedro A. Marín-Reyes, Adrián Peñate Sánchez, Javier Lorenzo-Navarro, Modesto Castrillón-Santana
Pattern Recognit. Lett.4
2020 TGCRBNW: A Dataset for Runner Bib Number Detection (and Recognition) in the Wild
abstract
Racing bib number (RBN) detection and recognition is a specific problem related to text recognition in natural scenes. In this paper, we present a novel dataset created after registering participants in a real ultrarunning competition which comprises a wide range of acquisition conditions in five different recording points, including nightlight and daylight. The dataset contains more than 3K samples of over 400 different individuals. The aim is to provide an “in the wild” benchmark for both RBN detection and recognition problems. To illustrate the present difficulties, the dataset is evaluated for RBN detection using different Faster R-CNN specific detection models, filtering its output with heuristics based on body detection to improve the overall detection performance. Initial results are promising, but there is still significant room for improvement. And detection is just the first step to accomplish “in the wild” RBN recognition.
Pablo Hernández-Carrascosa, Adrián Peñate Sánchez, Javier Lorenzo-Navarro, David Freire-Obregón, Modesto Castrillón-Santana
ICPR2
2020 TGC20ReId: A dataset for sport event re-identification in the wild
Adrián Peñate Sánchez, David Freire-Obregón, Adrián Lorenzo-Melián, Javier Lorenzo-Navarro, Modesto Castrillón-Santana
Pattern Recognit. Lett.1
2018 3D Pick & Mix: Object Part Blending in Joint Shape and Image Manifolds
Adrián Peñate Sánchez, Lourdes Agapito
ACCV (1)1
2017 Depth-aware convolutional neural networks for accurate 3D pose estimation in RGB-D images
abstract
Most recent approaches to 3D pose estimation from RGB-D images address the problem in a two-stage pipeline. First, they learn a classifier-typically a random forest-to predict the position of each input pixel on the object surface. These estimates are then used to define an energy function that is minimized w.r.t. the object pose. In this paper, we focus on the first stage of the problem and propose a novel classifier based on a depth-aware Convolutional Neural Network. This classifier is able to learn a scale-adaptive regression model that yields very accurate pixel-level predictions, allowing to finally estimate the pose using a simple RANSAC-based scheme, with no need to optimize complex ad hoc energy functions. Our experiments on publicly available datasets show that our approach achieves remarkable improvements over state-of-the-art methods.
Lorenzo Porzi, Adrián Peñate Sánchez, Elisa Ricci 0001, Francesc Moreno-Noguer
IROS2
2015 Matchability Prediction for Full-Search Template Matching Algorithms
abstract
While recent approaches have shown that it is possible to do template matching by exhaustively scanning the parameter space, the resulting algorithms are still quite demanding. In this paper we alleviate the computational load of these algorithms by proposing an efficient approach for predicting the match ability of a template, before it is actually performed. This avoids large amounts of unnecessary computations. We learn the match ability of templates by using dense convolutional neural network descriptors that do not require ad-hoc criteria to characterize a template. By using deep learning descriptions of patches we are able to predict match ability over the whole image quite reliably. We will also show how no specific training data is required to solve problems like panorama stitching in which you usually require data from the scene in question. Due to the highly parallelizable nature of this tasks we offer an efficient technique with a negligible computational cost at test time.
Adrián Peñate Sánchez, Lorenzo Porzi, Francesc Moreno-Noguer
3DV1
2015 A dynamic programming approach for fast and robust object pose recognition from range images
abstract
Joint object recognition and pose estimation solely from range images is an important task e.g. in robotics applications and in automated manufacturing environments. The lack of color information and limitations of current commodity depth sensors make this task a challenging computer vision problem, and a standard random sampling based approach is prohibitively time-consuming. We propose to address this difficult problem by generating promising inlier sets for pose estimation by early rejection of clear outliers with the help of local belief propagation (or dynamic programming). By exploiting data-parallelism our method is fast, and we also do not rely on a computationally expensive training phase. We demonstrate state-of-the art performance on a standard dataset and illustrate our approach on challenging real sequences.
Christopher Zach, Adrián Peñate Sánchez, Minh-Tri Pham
CVPR2
2015 Efficient monocular pose estimation for complex 3D models
abstract
We propose a robust and efficient method to estimate the pose of a camera with respect to complex 3D textured models of the environment that can potentially contain more than 100; 000 points. To tackle this problem we follow a top down approach where we combine high-level deep network classifiers with low level geometric approaches to come up with a solution that is fast, robust and accurate. Given an input image, we initially use a pre-trained deep network to compute a rough estimation of the camera pose. This initial estimate constrains the number of 3D model points that can be seen from the camera viewpoint. We then establish 3D-to-2D correspondences between these potentially visible points of the model and the 2D detected image features. Accurate pose estimation is finally obtained from the 2D-to-3D correspondences using a novel PnP algorithm that rejects outliers without the need to use a RANSAC strategy, and which is between 10 and 100 times faster than other methods that use it. Two real experiments dealing with very large and complex 3D models demonstrate the effectiveness of the approach.
Antonio Rubio 0001, Michael Villamizar, Luis Ferraz, Adrián Peñate Sánchez, Arnau Ramisa, Edgar Simo-Serra, Alberto Sanfeliu, Francesc Moreno-Noguer
ICRA4
2014 LETHA: Learning from High Quality Inputs for 3D Pose Estimation in Low Quality Images
abstract
We introduce LETHA (Learning on Easy data, Test on Hard), a new learning paradigm consisting of building strong priors from high quality training data, and combining them with discriminative machine learning to deal with low-quality test data. Our main contribution is an implementation of that concept for pose estimation. We first automatically build a 3D model of the object of interest from high-definition images, and devise from it a pose-indexed feature extraction scheme. We then train a single classifier to process these feature vectors. Given a low quality test image, we visit many hypothetical poses, extract features consistently and evaluate the response of the classifier. Since this process uses locations recorded during learning, it does not require matching points anymore. We use a boosting procedure to train this classifier common to all poses, which is able to deal with missing features, due in this context to self-occlusion. Our results demonstrate that the method combines the strengths of global image representations, discriminative even for very tiny images, and the robustness to occlusions of approaches based on local feature point descriptors.
Adrián Peñate Sánchez, Francesc Moreno-Noguer, Juan Andrade-Cetto, François Fleuret
3DV1
2013 Simultaneous Pose, Focal Length and 2D-to-3D Correspondences from Noisy Observations
abstract
Presentado al 24th BMVC celebrado en Bristol (UK) del 9 al 13 de septiembre 2013.-- The copyright of this document resides with its authors.
Adrián Peñate Sánchez, Eduard Serradell, Francesc Moreno-Noguer, Juan Andrade-Cetto
BMVC1
2013 Exhaustive Linearization for Robust Camera Pose and Focal Length Estimation
abstract
We propose a novel approach for the estimation of the pose and focal length of a camera from a set of 3D-to-2D point correspondences. Our method compares favorably to competing approaches in that it is both more accurate than existing closed form solutions, as well as faster and also more accurate than iterative ones. Our approach is inspired on the EPnP algorithm, a recent O(n) solution for the calibrated case. Yet we show that considering the focal length as an additional unknown renders the linearization and relinearization techniques of the original approach no longer valid, especially with large amounts of noise. We present new methodologies to circumvent this limitation termed exhaustive linearization and exhaustive relinearization which perform a systematic exploration of the solution space in closed form. The method is evaluated on both real and synthetic data, and our results show that besides producing precise focal length estimation, the retrieved camera pose is almost as accurate as the one computed using the EPnP, which assumes a calibrated camera.
Adrián Peñate Sánchez, Juan Andrade-Cetto, Francesc Moreno-Noguer
IEEE Trans. Pattern Anal. Mach. Intell.1