EDBT 2026 Demo / reviewers in the wild / expert
Renato Martins
dblp:164/8301
· DBLP profile ↗
20ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-0053-0004ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 9 since 2021Systems, architecture and hardware · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Domain Human Action Recognition from Multiview Motion and Textual Descriptions
Yannick Porto, Renato Martins, Thomas Chalumeau, Cédric Demonceaux |
ICPR (10) | 2 |
| 2026 | Gaussian Splatting Map Registration with Orthographic Bird's-Eye-View RenderingsabstractGaussian Splatting (GS) is a promising scene representation for visual localization and SLAM. Recent works have explored loop closure detection via Gaussian registration, improving map consistency and accuracy. However, achieving reliable registration given two GS representations from different acquisitions remains challenging. In this paper, we propose a complete pipeline to perform the matching and registration given two GS maps. The proposed method is grounded in generating orthographic bird’s-eye views (BEVs) of optimized Gaussian models. The proposed approach leverages photometric and geometric information extracted directly from the GS to provide a trade-off of accuracy and invariance to different viewing changes (e.g., as types of GS maps, seasons, or illumination). Unlike existing 3D registration methods, which become inefficient as the number of Gaussians grows, our approach leverages 2D orthographic renders thus considerably reducing the registration complexity. Experiments on two public datasets demonstrate that our method achieves higher accuracy than several existing baselines, while also maintaining better registration results when dealing with GS maps learned by different techniques (e.g., 3DGS to LightGaussian), or GS maps presenting viewing changes such as varying illumination conditions. Source code is available at: https://gitlab.inria.fr/tangram/bev-splatreg Hugo Leblond, Gilles Simon, Renato Martins, Cédric Demonceaux, Marie-Odile Berger |
WACV | 3 |
| 2025 | Dense Scene Reconstruction from Light-Field Images Affected by Rolling ShutterabstractThis paper presents a dense depth estimation approach from light-field (LF) images that is able to compensate for strong rolling shutter (RS) effects. Our method estimates RS compensated views and dense RS compensated disparity maps. We present a two-stage method based on a 2D Gaussians Splatting that allows for a “render and compare” strategy with a point cloud formulation. In the first stage, a subset of sub-aperture images is used to estimate an RS agnostic 3D shape that is related to the scene target shape “up to a motion”. In the second stage, the deformation of the 3D shape is computed by estimating an admissible camera motion. We demonstrate the effectiveness and advantages of this approach through several experiments conducted for different scenes and types of motions. Due to lack of suitable datasets for evaluation, we also present a new carefully designed synthetic dataset of RS LF images. The source code, trained models and dataset will be made publicly available at: https://github.com/ICB-Vision-AI/DenseRSLF. Hermes McGriff, Renato Martins, Nicolas Andreff, Cédric Demonceaux |
WACV | 2 |
| 2024 | XFeat: Accelerated Features for Lightweight Image MatchingabstractWe introduce a lightweight and accurate architecture for resource-efficient visual correspondence. Our method, dubbed XFeat (Accelerated Features), revisits fundamen-tal design choices in convolutional neural networks for de-tecting, extracting, and matching local features. Our new model satisfies a critical need for fast and robust algorithms suitable to resource-limited devices. In particular, accu-rate image matching requires sufficiently large image res-olutions -for this reason, we keep the resolution as large as possible while limiting the number of channels in the net-work. Besides, our model is designed to offer the choice of matching at the sparse or semi-dense levels, each of which may be more suitable for different downstream applications, such as visual navigation and augmented reality. Our model is the first to offer semi-dense matching efficiently, leveraging a novel match refinement module that relies on coarse local descriptors. XFeat is versatile and hardware-independent, surpassing current deep learning-based local features in speed (up to 5xfaster) with comparable or better accuracy, proven in pose estimation and visual localization. We showcase it running in real-time on an inexpensive lap-top CPU without specialized hardware optimizations. Code and weights are available at verlab.dcc.ufmg.br/descriptors/xfeat_cvpr24. Guilherme A. Potje, Felipe C. Chamone, André Araújo 0001, Renato Martins, Erickson R. Nascimento |
CVPR | 4 |
| 2024 | Joint 3D Shape and Motion Estimation from Rolling Shutter Light-Field ImagesabstractIn this paper, we propose an approach to address the problem of 3D reconstruction of scenes from a single image captured by a light-field camera equipped with a rolling shutter sensor. Our method leverages the 3D information cues present in the light-field and the motion information provided by the rolling shutter effect. We present a generic model for the imaging process of this sensor and a two-stage algorithm that minimizes the re-projection error while considering the position and motion of the camera in a motion-shape bundle adjustment estimation strategy. Thereby, we provide an instantaneous 3D shape-and-pose-and-velocity sensing paradigm. To the best of our knowledge, this is the first study to leverage this type of sensor for this purpose. We also present a new benchmark dataset composed of different light-fields showing rolling shutter effects, which can be used as a common base to improve the evaluation and tracking the progress in the field. We demonstrate the effectiveness and advantages of our approach through several experiments conducted for different scenes and types of motions. The source code and dataset are publicly available at: https://github.com/ICB-Vision-AI/RSLF. Hermes McGriff, Renato Martins, Nicolas Andreff, Cédric Demonceaux |
WACV | 2 |
| 2023 | Enhancing Deformable Local Features by Jointly Learning to Detect and Describe KeypointsabstractLocal feature extraction is a standard approach in computer vision for tackling important tasks such as image matching and retrieval. The core assumption of most methods is that images undergo affine transformations, disregarding more complicated effects such as non-rigid deformations. Furthermore, incipient works tailored for non-rigid correspondence still rely on keypoint detectors designed for rigid transformations, hindering performance due to the limitations of the detector. We propose DALF (Deformation-Aware Local Features), a novel deformation-aware network for jointly detecting and describing keypoints, to handle the challenging problem of matching deformable surfaces. All network components work cooperatively through a feature fusion approach that enforces the descriptors' distinctiveness and invariance. Experiments using real deforming objects showcase the superiority of our method, where it delivers 8% improvement in matching scores compared to the previous best results. Our approach also enhances the performance of two real-world applications: deformable object retrieval and non-rigid 3D surface registration. Code for training, inference, and applications are publicly available at verlab.dcc.ufmg.br/descriptors/dalf_cvpr23. Guilherme A. Potje, Felipe C. Chamone, André Araújo 0001, Renato Martins, Erickson R. Nascimento |
CVPR | 4 |
| 2023 | Improving the matching of deformable objects by learning to detect keypoints
Felipe C. Chamone, Welerson Melo, Vaishnavi Kanagasabapathi, Guilherme A. Potje, Renato Martins, Erickson R. Nascimento |
Pattern Recognit. Lett. | 5 |
| 2022 | Leveraging Semantic Cues from Foundation Vision Models for Enhanced Local Feature Correspondence
Felipe C. Chamone, Guilherme A. Potje, Renato Martins, Cédric Demonceaux, Erickson R. Nascimento |
ACCV (4) | 3 |
| 2022 | Creating and Reenacting Controllable 3D Humans with Differentiable RenderingabstractThis paper proposes a new end-to-end neural rendering architecture to transfer appearance and reenact human actors. Our method leverages a carefully designed graph convolutional network (GCN) to model the human body manifold structure, jointly with differentiable rendering, to synthesize new videos of people in different contexts from where they were initially recorded. Unlike recent appearance transferring methods, our approach can reconstruct a fully controllable 3D texture-mapped model of a person, while taking into account the manifold structure from body shape and texture appearance in the view synthesis. Specifically, our approach models mesh deformations with a three-stage GCN trained in a self-supervised manner on rendered silhouettes of the human body. It also infers texture appearance with a convolutional network in the texture domain, which is trained in an adversarial regime to reconstruct human texture from rendered images of actors in different poses. Experiments on different videos show that our method successfully infers specific body deformations and avoid creating texture artifacts while achieving the best values for appearance in terms of Structural Similarity (SSIM), Learned Perceptual Image Patch Similarity (LPIPS), Mean Squared Error (MSE), and Frchet Video Distance (FVD). By taking advantages of both differentiable rendering and the 3D parametric model, our method is fully controllable, which allows controlling the human synthesis from both pose and rendering parameters. The source code is available at https://www.verlab.dcc.ufmg.br/retargeting-motion/wacv2022. Thiago L. Gomes, Thiago M. Coutinho, Rafael Azevedo, Renato Martins, Erickson R. Nascimento |
WACV | 4 |
| 2022 | Learning geodesic-aware local features from RGB-D images
Guilherme A. Potje, Renato Martins, Felipe C. Chamone, Erickson R. Nascimento |
Comput. Vis. Image Underst. | 2 |
| 2021 | Extracting Deformation-Aware Local Features by Learning to DeformabstractDespite the advances in extracting local features achieved by handcrafted and learning-based descriptors, they are still limited by the lack of invariance to non-rigid transformations. In this paper, we present a new approach to compute features from still images that are robust to non-rigid deformations to circumvent the problem of matching deformable surfaces and objects. Our deformation-aware local descriptor, named DEAL, leverages a polar sampling and a spatial transformer warping to provide invariance to rotation, scale, and image deformations. We train the model architecture end-to-end by applying isometric non-rigid deformations to objects in a simulated environment as guidance to provide highly discriminative local features. The experiments show that our method outperforms state-of-the-art handcrafted, learning-based image, and RGB-D descriptors in different datasets with both real and realistic synthetic deformable objects in still images. The source code and trained model of the descriptor are publicly available at https://www.verlab.dcc.ufmg.br/descriptors/neurips2021. Guilherme A. Potje, Renato Martins, Felipe C. Chamone, Erickson R. Nascimento |
NeurIPS | 2 |
| 2021 | Learning to dance: A graph convolutional adversarial network to generate realistic dance motions from audio
João Pedro Moreira Ferreira, Thiago M. Coutinho, Thiago L. Gomes, José F. Neto, Rafael Azevedo, Renato Martins, Erickson R. Nascimento |
Comput. Graph. | 6 |
| 2021 | A Shape-Aware Retargeting Approach to Transfer Human Motion and Appearance in Monocular Videos
Thiago L. Gomes, Renato Martins, João Pedro Moreira Ferreira, Rafael Azevedo, Guilherme Torres, Erickson R. Nascimento |
Int. J. Comput. Vis. | 2 |
| 2020 | Do As I Do: Transferring Human Motion and Appearance between Monocular Videos with Spatial and Temporal ConstraintsabstractCreating plausible virtual actors from images of real actors remains one of the key challenges in computer vision and computer graphics. Marker-less human motion estimation and shape modeling from images in the wild bring this challenge to the fore. Although the recent advances on view synthesis and image-to-image translation, currently available formulations are limited to transfer solely style and do not take into account the character's motion and shape, which are by nature intermingled to produce plausible human forms. In this paper, we propose a unifying formulation for transferring appearance and retargeting human motion from monocular videos that regards all these aspects. Our method synthesizes new videos of people in a different context where they were initially recorded. Differently from recent appearance transferring methods, our approach takes into account body shape, appearance, and motion constraints. The evaluation is performed with several experiments using publicly available real videos containing hard conditions. Our method is able to transfer both human motion and appearance outperforming state-of-the-art methods, while preserving specific features of the motion that must be maintained (e.g., feet touching the floor, hands touching a particular object) and holding the best visual quality and appearance metrics such as Structural Similarity (SSIM) and Learned Perceptual Image Patch Similarity (LPIPS). Thiago L. Gomes, Renato Martins, João P. K. Ferreira, Erickson R. Nascimento |
WACV | 2 |
| 2019 | GEOBIT: A Geodesic-Based Binary Descriptor Invariant to Non-Rigid Deformations for RGB-D ImagesabstractAt the core of most three-dimensional alignment and tracking tasks resides the critical problem of point correspondence. In this context, the design of descriptors that efficiently and uniquely identifies keypoints, to be matched, is of central importance. Numerous descriptors have been developed for dealing with affine/perspective warps, but few can also handle non-rigid deformations. In this paper, we introduce a novel binary RGB-D descriptor invariant to isometric deformations. Our method uses geodesic isocurves on smooth textured manifolds. It combines appearance and geometric information from RGB-D images to tackle non-rigid transformations. We used our descriptor to track multiple textured depth maps and demonstrate that it produces reliable feature descriptors even in the presence of strong non-rigid deformations and depth noise. The experiments show that our descriptor outperforms different state-of-the-art descriptors in both precision-recall and recognition rate metrics. We also provide to the community a new dataset composed of annotated RGB-D images of different objects (shirts, cloths, paintings, bags), subjected to strong non-rigid deformations, to evaluate point correspondence algorithms. Erickson R. Nascimento, Guilherme A. Potje, Renato Martins, Felipe C. Chamone, Mario Fernando Montenegro Campos, Ruzena Bajcsy |
ICCV | 3 |
| 2018 | A New Metric for Evaluating Semantic Segmentation: Leveraging Global and Contour AccuracyabstractSemantic segmentation of images is an important issue for intelligent vehicles and mobile robotics because it offers basic information which can be used for complex reasoning and safe navigation. Different solutions have been proposed for this problem along the last two decades, where recent deep neural networks approaches have shown very promising results in the context of urban navigation. One of the main problems when comparing different semantic segmentation solutions is how to select an appropriate metric to evaluate their accuracy. On the one hand, classic metrics do not measure properly the accuracy on the object contours, which is important in urban driving to differentiate road from sidewalk for instance. On the other hand, contour-based metrics [1] disregard the information far from class contours. This paper explores the problem multi-modal image segmentation, and presents a new metric to leverage global and contour accuracy in a simple formulation. This metric is validated with the evaluation of several semantic segmentation solutions that exploit RGB-D images to rank these solutions taking into account the quality of the segmented contours. We also present a comparative analysis of several commonly used metrics together with a statistical analysis of their correlation. Eduardo Fernández-Moral, Renato Martins, Denis F. Wolf, Patrick Rives |
Intelligent Vehicles Symposium | 2 |
| 2017 | An efficient rotation and translation decoupled initialization from large field of view depth imagesabstractImage and point cloud registration methods compute the relative pose between two images. Commonly used registration algorithms are iterative and rely on the assumption that the motion between the images is small. In this work, we propose a fast pose estimation technique to compute a rough estimate of large motions between depth images, which can be used as initialization to dense registration methods. The main idea is to explore the properties given by planar surfaces with co-visibility and their normals from two distinct viewpoints. We present, in two decoupled stages, the rotation and then the translation estimation, both based on the normal vectors orientation and on the depth. These two stages are efficiently computed by using low resolution depth images and without any feature extraction/matching. We also analyze the limitations and observabilty of this approach, and its relationship to ICP point-to-plane. Notably, if the rotation is observable, at least five degrees of freedom can be estimated in the worst case. To demonstrate the effectiveness of the method, we evaluate the initialization technique in a set of challenging scenarios, comprising simulated spherical images from the Sponza Atrium model benchmark and real spherical indoor sequences. Renato Martins, Eduardo Fernández-Moral, Patrick Rives |
IROS | 1 |
| 2016 | Adaptive Direct RGB-D Registration and Mapping for Large Motions
Renato Martins, Eduardo Fernández-Moral, Patrick Rives |
ACCV (4) | 1 |
| 2015 | A compact spherical RGBD keyframe-based representationabstractThis paper proposes an environmental representation approach based on hybrid metric and topological maps as a key component for mobile robot navigation. Focus is made on an ego-centric pose graph structure by the use of Keyframes to capture the local properties of the scene. With the aim of reducing data redundancy, suppress sensor noise whilst maintaining a dense compact representation of the environment, neighbouring augmented spheres are fused in a single representation. To this end, an uncertainty error model propagation is formulated for outlier rejection and data fusion, enhanced with the notion of landmark stability over time. Finally, our algorithm is tested thoroughly on a newly developed wide angle 360° field of view (FOV) spherical sensor where improvements such as trajectory drift, compactness and reduced tracking error are demonstrated. Tawsif Gokhool, Renato Martins, Patrick Rives, Noela Despré |
ICRA | 2 |
| 2015 | Dense accurate urban mapping from spherical RGB-D imagesabstractThis paper presents a methodology to combine information from a sequence of RGB-D spherical views acquired by a home-made multi-stereo device in order to improve the computed depth images both in terms of accuracy and completeness. This methodology is embedded in a larger visual mapping framework aiming to produce accurate and dense topometric urban maps. Our method is based on two main filtering stages. Firstly, we perform a segmentation process considering both geometric and photometric image constraints, followed by a regularization step (spatial-integration). We then proceed to a fusion stage where the geometric information is further refined by considering the depth images of nearby frames (temporal integration). This methodology can be applied to other projective models, such as perspective stereo images. Our approach is evaluated within the frameworks of image registration, localization and mapping, demonstrating higher accuracy and larger convergence domains over different datasets. Renato Martins, Eduardo Fernández-Moral, Patrick Rives |
IROS | 1 |