Erickson R. Nascimento

dblp:92/7785 · also Erickson Rangel do Nascimento · DBLP profile ↗
← Back
54ranked-venue papers
5as first author
22since 2021 · last 2025
0000-0003-2973-2232ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 8 since 2021Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Captar-Libras: A Bidirectional Translator with a Photorealistic Avatar for Medical Pre-consultation of Deaf Patients
Natália Sales Santos, Lucas Almeida S. de Souza, Ari Gonçalves da Silva Filho, Paulo H. de S. Coelho, Maria Clara de A. Ferreira, Cauã Magalhães Pereira, Bruno de S. Lages, Raniere A. A. Cordeiro, Elidéa L. A. Bernardino, Milena Soriano Marcolino, Raphael Nepomuceno, Michel Melo Silva, Thiago L. Gomes, Erickson R. Nascimento, Douglas G. Macharet, Raquel Oliveira Prates, Mario Fernando Montenegro Campos
INTERACT (4)14
2024 XFeat: Accelerated Features for Lightweight Image Matching
abstract
We introduce a lightweight and accurate architecture for resource-efficient visual correspondence. Our method, dubbed XFeat (Accelerated Features), revisits fundamen-tal design choices in convolutional neural networks for de-tecting, extracting, and matching local features. Our new model satisfies a critical need for fast and robust algorithms suitable to resource-limited devices. In particular, accu-rate image matching requires sufficiently large image res-olutions -for this reason, we keep the resolution as large as possible while limiting the number of channels in the net-work. Besides, our model is designed to offer the choice of matching at the sparse or semi-dense levels, each of which may be more suitable for different downstream applications, such as visual navigation and augmented reality. Our model is the first to offer semi-dense matching efficiently, leveraging a novel match refinement module that relies on coarse local descriptors. XFeat is versatile and hardware-independent, surpassing current deep learning-based local features in speed (up to 5xfaster) with comparable or better accuracy, proven in pose estimation and visual localization. We showcase it running in real-time on an inexpensive lap-top CPU without specialized hardware optimizations. Code and weights are available at verlab.dcc.ufmg.br/descriptors/xfeat_cvpr24.
Guilherme A. Potje, Felipe C. Chamone, André Araújo 0001, Renato Martins, Erickson R. Nascimento
CVPR5
2024 Leveraging Synthetic Data to Learn Video Stabilization Under Adverse Conditions
abstract
Stabilization plays a central role in improving the quality of videos. However, current methods perform poorly under adverse conditions. In this paper, we propose a synthetic-aware adverse weather video stabilization algorithm that dispenses real data for training, relying solely on synthetic data. Our approach leverages specially generated synthetic data to avoid the feature extraction issues faced by current methods. To achieve this, we present a novel data generator to produce the required training data with an automatic ground-truth extraction procedure. We also propose a new dataset, VSAC105Real, and compare our method to five recent video stabilization algorithms using two benchmarks. Our method generalizes well on real-world videos across all weather conditions and does not require large-scale synthetic training data. Implementations for our proposed video stabilization algorithm, generator, and datasets are available at https://github.com/A-Kerim/SyntheticData4VideoStabilization_WACV_2024.
Abdulrahman Kerim, Washington L. S. Ramos, Leandro Soriano Marcolino, Erickson R. Nascimento, Richard Jiang 0001
WACV4
2024 Empowering sign language communication: Integrating sentiment and semantics for facial expression synthesis
Rafael Azevedo, Thiago M. Coutinho, João Pedro Moreira Ferreira, Thiago L. Gomes, Erickson R. Nascimento
Comput. Graph.5
2023 Enhancing Deformable Local Features by Jointly Learning to Detect and Describe Keypoints
abstract
Local feature extraction is a standard approach in computer vision for tackling important tasks such as image matching and retrieval. The core assumption of most methods is that images undergo affine transformations, disregarding more complicated effects such as non-rigid deformations. Furthermore, incipient works tailored for non-rigid correspondence still rely on keypoint detectors designed for rigid transformations, hindering performance due to the limitations of the detector. We propose DALF (Deformation-Aware Local Features), a novel deformation-aware network for jointly detecting and describing keypoints, to handle the challenging problem of matching deformable surfaces. All network components work cooperatively through a feature fusion approach that enforces the descriptors' distinctiveness and invariance. Experiments using real deforming objects showcase the superiority of our method, where it delivers 8% improvement in matching scores compared to the previous best results. Our approach also enhances the performance of two real-world applications: deformable object retrieval and non-rigid 3D surface registration. Code for training, inference, and applications are publicly available at verlab.dcc.ufmg.br/descriptors/dalf_cvpr23.
Guilherme A. Potje, Felipe C. Chamone, André Araújo 0001, Renato Martins, Erickson R. Nascimento
CVPR5
2023 Text-Driven Video Acceleration: A Weakly-Supervised Reinforcement Learning Method
abstract
The growth of videos in our digital age and the users' limited time raise the demand for processing untrimmed videos to produce shorter versions conveying the same information. Despite the remarkable progress that summarization methods have made, most of them can only select a few frames or skims, creating visual gaps and breaking the video context. This paper presents a novel weakly-supervised methodology based on a reinforcement learning formulation to accelerate instructional videos using text. A novel joint reward function guides our agent to select which frames to remove and reduce the input video to a target length without creating gaps in the final video. We also propose the Extended Visually-guided Document Attention Network (VDAN+), which can generate a highly discriminative embedding space to represent both textual and visual data. Our experiments show that our method achieves the best performance in Precision, Recall, and F1 Score against the baselines while effectively controlling the video's output length.
Washington L. S. Ramos, Michel Melo Silva, Edson Araujo, Victor Moura, Keller Oliveira, Leandro Soriano Marcolino, Erickson R. Nascimento
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Improving the matching of deformable objects by learning to detect keypoints
Felipe C. Chamone, Welerson Melo, Vaishnavi Kanagasabapathi, Guilherme A. Potje, Renato Martins, Erickson R. Nascimento
Pattern Recognit. Lett.6
2023 A multimodal hyperlapse method based on video and songs' emotion alignment
Diognei de Matos, Washington L. S. Ramos, Michel Melo Silva, Luiz Romanhol, Erickson R. Nascimento
Pattern Recognit. Lett.5
2022 Leveraging Semantic Cues from Foundation Vision Models for Enhanced Local Feature Correspondence
Felipe C. Chamone, Guilherme A. Potje, Renato Martins, Cédric Demonceaux, Erickson R. Nascimento
ACCV (4)5
2022 Semantic Segmentation under Adverse Conditions: A Weather and Nighttime-aware Synthetic Data-based Approach
Abdulrahman Kerim, Felipe C. Chamone, Washington L. S. Ramos, Leandro Soriano Marcolino, Erickson R. Nascimento, Richard Jiang 0001
BMVC5
2022 Creating and Reenacting Controllable 3D Humans with Differentiable Rendering
abstract
This paper proposes a new end-to-end neural rendering architecture to transfer appearance and reenact human actors. Our method leverages a carefully designed graph convolutional network (GCN) to model the human body manifold structure, jointly with differentiable rendering, to synthesize new videos of people in different contexts from where they were initially recorded. Unlike recent appearance transferring methods, our approach can reconstruct a fully controllable 3D texture-mapped model of a person, while taking into account the manifold structure from body shape and texture appearance in the view synthesis. Specifically, our approach models mesh deformations with a three-stage GCN trained in a self-supervised manner on rendered silhouettes of the human body. It also infers texture appearance with a convolutional network in the texture domain, which is trained in an adversarial regime to reconstruct human texture from rendered images of actors in different poses. Experiments on different videos show that our method successfully infers specific body deformations and avoid creating texture artifacts while achieving the best values for appearance in terms of Structural Similarity (SSIM), Learned Perceptual Image Patch Similarity (LPIPS), Mean Squared Error (MSE), and Frchet Video Distance (FVD). By taking advantages of both differentiable rendering and the 3D parametric model, our method is fully controllable, which allows controlling the human synthesis from both pose and rendering parameters. The source code is available at https://www.verlab.dcc.ufmg.br/retargeting-motion/wacv2022.
Thiago L. Gomes, Thiago M. Coutinho, Rafael Azevedo, Renato Martins, Erickson R. Nascimento
WACV5
2022 Learning geodesic-aware local features from RGB-D images
Guilherme A. Potje, Renato Martins, Felipe C. Chamone, Erickson R. Nascimento
Comput. Vis. Image Underst.4
2022 A reinforcement learning approach for single redundant view co-training text classification
Bruno B. M. Paiva, Erickson R. Nascimento, Marcos André Gonçalves, Fabiano Muniz Belém
Inf. Sci.2
2022 Learning to Detect Changes in Aerial Images in the Presence of Registration Errors
abstract
We present a CNN-based approach for scene-level change detection in aerial images with registration errors. Thousands of aerial images and long videos are routinely acquired for monitoring large areas such as forests and oil pipelines. Annotating changes in those videos and images can be tedious, error-prone, or even unfeasible for a human operator. Moreover, accurate pixel-wise registration is usually unavailable, and conventional descriptor-based registration methods are doomed to fail since they rely on similarities to establish correspondences that are impaired due to the latent changes in the scene. We introduce a new neural network architecture that can be trained end-to-end to simultaneously perform image registration and change detection to mitigate these issues. Our approach reduces the number of parameters required to optimize the process while steeping towards a more robust change detection pipeline for UAV images. We evaluated our method in two datasets containing registration errors of up to 96 pixels in translation and 30° in rotation. The results showed that our approach outperformed the state-of-the-art, achieving a 6% improvement in the AUC metric.
Daniel Balbino de Mesquita, Mario Fernando Montenegro Campos, Erickson R. Nascimento
IEEE Geosci. Remote. Sens. Lett.3
2021 Anytime Fault-tolerant Adaptive Routing for Multi-Robot Teams
abstract
The Correlated Team Orienteering Problem (CTOP) is a routing problem where the objective is to determine a set of routes that maximizes the summation of collected rewards in the environment while respecting the vehicles’ budget. However, solutions to this problem usually consider static instances and may produce poor results in dynamic real-world scenarios. In this paper, we propose an approach to deal with the execution of missions planned as a CTOP instance, especially when vehicles of the team are prone to failure and may not complete their routes. The main contribution of this paper is a novel anytime heuristic that iteratively adapts the initial set of routes, allowing to increase the overall robustness of the mission and still collect the most profitable rewards. The methodology was thoroughly evaluated considering different scenarios and in all the cases was able to achieve comparable or better results in terms of reward than planning a new set of routes, however, spending considerably less time.
Ronaldo F. dos Santos, Erickson R. Nascimento, Douglas G. Macharet
ICRA2
2021 Extracting Deformation-Aware Local Features by Learning to Deform
abstract
Despite the advances in extracting local features achieved by handcrafted and learning-based descriptors, they are still limited by the lack of invariance to non-rigid transformations. In this paper, we present a new approach to compute features from still images that are robust to non-rigid deformations to circumvent the problem of matching deformable surfaces and objects. Our deformation-aware local descriptor, named DEAL, leverages a polar sampling and a spatial transformer warping to provide invariance to rotation, scale, and image deformations. We train the model architecture end-to-end by applying isometric non-rigid deformations to objects in a simulated environment as guidance to provide highly discriminative local features. The experiments show that our method outperforms state-of-the-art handcrafted, learning-based image, and RGB-D descriptors in different datasets with both real and realistic synthetic deformable objects in still images. The source code and trained model of the descriptor are publicly available at https://www.verlab.dcc.ufmg.br/descriptors/neurips2021.
Guilherme A. Potje, Renato Martins, Felipe C. Chamone, Erickson R. Nascimento
NeurIPS4
2021 Learning to dance: A graph convolutional adversarial network to generate realistic dance motions from audio
João Pedro Moreira Ferreira, Thiago M. Coutinho, Thiago L. Gomes, José F. Neto, Rafael Azevedo, Renato Martins, Erickson R. Nascimento
Comput. Graph.7
2021 A Shape-Aware Retargeting Approach to Transfer Human Motion and Appearance in Monocular Videos
Thiago L. Gomes, Renato Martins, João Pedro Moreira Ferreira, Rafael Azevedo, Guilherme Torres, Erickson R. Nascimento
Int. J. Comput. Vis.6
2021 Introducing the structural bases of typicality effects in deep learning
Omar Vidal Pino, Erickson R. Nascimento, Mario Fernando Montenegro Campos
Image Vis. Comput.2
2021 Towards automatic diagnosis of rheumatic heart disease on echocardiographic exams through video-based deep learning
abstract
OBJECTIVE: Rheumatic heart disease (RHD) affects an estimated 39 million people worldwide and is the most common acquired heart disease in children and young adults. Echocardiograms are the gold standard for diagnosis of RHD, but there is a shortage of skilled experts to allow widespread screenings for early detection and prevention of the disease progress. We propose an automated RHD diagnosis system that can help bridge this gap. MATERIALS AND METHODS: Experiments were conducted on a dataset with 11 646 echocardiography videos from 912 exams, obtained during screenings in underdeveloped areas of Brazil and Uganda. We address the challenges of RHD identification with a 3D convolutional neural network (C3D), comparing its performance with a 2D convolutional neural network (VGG16) that is commonly used in the echocardiogram literature. We also propose a supervised aggregation technique to combine video predictions into a single exam diagnosis. RESULTS: The proposed approach obtained an accuracy of 72.77% for exam diagnosis. The results for the C3D were significantly better than the ones obtained by the VGG16 network for videos, showing the importance of considering the temporal information during the diagnostic. The proposed aggregation model showed significantly better accuracy than the majority voting strategy and also appears to be capable of capturing underlying biases in the neural network output distribution, balancing them for a more correct diagnosis. CONCLUSION: Automatic diagnosis of echo-detected RHD is feasible and, with further research, has the potential to reduce the workload of experts, enabling the implementation of more widespread screening programs worldwide.
Joao Francisco B. S. Martins, Erickson R. Nascimento, Bruno Ramos Nascimento, Craig A. Sable, Andrea Z. Beaton, Antônio L. P. Ribeiro, Wagner Meira Jr., Gisele L. Pappa
J. Am. Medical Informatics Assoc.2
2021 A Sparse Sampling-Based Framework for Semantic Fast-Forward of First-Person Videos
abstract
Technological advances in sensors have paved the way for digital cameras to become increasingly ubiquitous, which, in turn, led to the popularity of the self-recording culture. As a result, the amount of visual data on the Internet is moving in the opposite direction of the available time and patience of the users. Thus, most of the uploaded videos are doomed to be forgotten and unwatched stashed away in some computer folder or website. In this paper, we address the problem of creating smooth fast-forward videos without losing the relevant content. We present a new adaptive frame selection formulated as a weighted minimum reconstruction problem. Using a smoothing frame transition and filling visual gaps between segments, our approach accelerates first-person videos emphasizing the relevant segments and avoids visual discontinuities. Experiments conducted on controlled videos and also on an unconstrained dataset of First-Person Videos (FPVs) show that, when creating fast-forward videos, our method is able to retain as much relevant information and smoothness as the state-of-the-art techniques, but in less processing time.
Michel Melo Silva, Washington L. S. Ramos, Mario Fernando Montenegro Campos, Erickson R. Nascimento
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 On the Development of an Acoustic-Driven Method to Improve Driver's Comfort Based on Deep Reinforcement Learning
abstract
The safety and comfort of drivers have been improved over the decades as a result of our broadened understanding of driver modeling and behavior prediction. Despite these remarkable advances in autonomous and interactive systems, there is a significant lack of approaches that consider the passengers and the vehicle as components of a dynamical vibro-acoustical system. Sound in vehicles is not only informative of the state of the vehicle and the environment, but can also critically affect the driver's performance, attention, and comfort. This paper aims to investigate the interplay between the perceived sounds of a vehicle and psychoacoustic annoyance (PA) metrics. Our goal is to create an intelligent agent that would act to improve driving pleasantness through acoustic-driven learning. To tackle the problem of choosing the correct actions to reduce the acoustic annoyance, the paper presents a method based on reinforcement learning that learns from the environment, i.e., the vehicle interior. The method actively changes the state inside the vehicle (e.g., closing or opening the window and choosing the cruise speed) in order to minimize acoustic annoyance experienced by the driver. The results of this work, performed using the GTA V simulator, showed that the trained agent successfully learned to take the correct actions to reduce PA metrics. The paper also present to the community a new multi-modal dataset composed of several rides on a real vehicle and an in-depth analysis of the influence of vehicle's signal on the acoustic annoyance.
Erickson R. Nascimento, Ruzena Bajcsy, Michal Gregor, Isabella Huang, Ismael Villegas, Gregorij Kurillo
IEEE Trans. Intell. Transp. Syst.1
2020 Straight to the Point: Fast-Forwarding Videos via Reinforcement Learning Using Textual Data
abstract
The rapid increase in the amount of published visual data and the limited time of users bring the demand for processing untrimmed videos to produce shorter versions that convey the same information. Despite the remarkable progress that has been made by summarization methods, most of them can only select a few frames or skims, which creates visual gaps and breaks the video context. In this paper, we present a novel methodology based on a reinforcement learning formulation to accelerate instructional videos. Our approach can adaptively select frames that are not relevant to convey the information without creating gaps in the final video. Our agent is textually and visually oriented to select which frames to remove to shrink the input video. Additionally, we propose a novel network, called Visually-guided Document Attention Network (VDAN), able to generate a highly discriminative embedding space to represent both textual and visual data. Our experiments show that our method achieves the best performance in terms of F1 Score and coverage at the video segment level.
Washington L. S. Ramos, Michel Melo Silva, Edson Araujo, Leandro Soriano Marcolino, Erickson R. Nascimento
CVPR5
2020 Do As I Do: Transferring Human Motion and Appearance between Monocular Videos with Spatial and Temporal Constraints
abstract
Creating plausible virtual actors from images of real actors remains one of the key challenges in computer vision and computer graphics. Marker-less human motion estimation and shape modeling from images in the wild bring this challenge to the fore. Although the recent advances on view synthesis and image-to-image translation, currently available formulations are limited to transfer solely style and do not take into account the character's motion and shape, which are by nature intermingled to produce plausible human forms. In this paper, we propose a unifying formulation for transferring appearance and retargeting human motion from monocular videos that regards all these aspects. Our method synthesizes new videos of people in a different context where they were initially recorded. Differently from recent appearance transferring methods, our approach takes into account body shape, appearance, and motion constraints. The evaluation is performed with several experiments using publicly available real videos containing hard conditions. Our method is able to transfer both human motion and appearance outperforming state-of-the-art methods, while preserving specific features of the motion that must be maintained (e.g., feet touching the floor, hands touching a particular object) and holding the best visual quality and appearance metrics such as Structural Similarity (SSIM) and Learned Perceptual Image Patch Similarity (LPIPS).
Thiago L. Gomes, Renato Martins, João P. K. Ferreira, Erickson R. Nascimento
WACV4
2020 Personalizing Fast-Forward Videos Based on Visual and Textual Features from Social Network
abstract
The growth of Social Networks has fueled the habit of people logging their day-to-day activities, and long First-Person Videos (FPVs) are one of the main tools in this new habit. Semantic-aware fast-forward methods are able to decrease the watch time and select meaningful moments, which is key to increase the chances of these videos being watched. However, these methods can not handle semantics in terms of personalization. In this paper, we present a new approach to automatically creating personalized fast-forward videos for FPVs. Our approach explores the availability of text-centric data from the user's social networks such as status updates to infer her/his topics of interest and assigns scores to the input frames according to her/his preferences. Extensive experiments are conducted on three different datasets with simulated and real-world users as input. Our method achieved an average F1score of up to 12.8 percentage points higher than the best competitors. We also present a user study to demonstrate the effectiveness of our method.
Washington L. S. Ramos, Michel Melo Silva, Edson Araujo, Alan C. Neves, Erickson R. Nascimento
WACV5
2020 Fully Convolutional Siamese Autoencoder for Change Detection in UAV Aerial Images
abstract
Different applications in remote sensing, such as crop monitoring and visual surveillance, demand the automatic detection of changes from sets of images acquired over time. Most traditional approaches use satellite imagery, which, besides the known issues such as cloud cover and image acquisition frequency for nongeostationary satellites, are very costly. In this context, with the recent technological advances, unmanned aerial vehicles (UAVs) have become ubiquitous in numerous applications. In this letter, we present a fully convolutional Siamese autoencoder method for change detection in aerial images, in particular for those obtained with UAVs. We show that, by using an autoencoder, we can further reduce the number of labeled samples required to achieve competitive results. We evaluated the performance of our approach on two different data sets, and the results showed that our methodology outperforms the state of the art, while demanding less training data.
Daniel Balbino de Mesquita, Ronaldo F. dos Santos, Douglas G. Macharet, Mario Fernando Montenegro Campos, Erickson R. Nascimento
IEEE Geosci. Remote. Sens. Lett.5
2019 GEOBIT: A Geodesic-Based Binary Descriptor Invariant to Non-Rigid Deformations for RGB-D Images
abstract
At the core of most three-dimensional alignment and tracking tasks resides the critical problem of point correspondence. In this context, the design of descriptors that efficiently and uniquely identifies keypoints, to be matched, is of central importance. Numerous descriptors have been developed for dealing with affine/perspective warps, but few can also handle non-rigid deformations. In this paper, we introduce a novel binary RGB-D descriptor invariant to isometric deformations. Our method uses geodesic isocurves on smooth textured manifolds. It combines appearance and geometric information from RGB-D images to tackle non-rigid transformations. We used our descriptor to track multiple textured depth maps and demonstrate that it produces reliable feature descriptors even in the presence of strong non-rigid deformations and depth noise. The experiments show that our descriptor outperforms different state-of-the-art descriptors in both precision-recall and recognition rate metrics. We also provide to the community a new dataset composed of annotated RGB-D images of different objects (shirts, cloths, paintings, bags), subjected to strong non-rigid deformations, to evaluate point correspondence algorithms.
Erickson R. Nascimento, Guilherme A. Potje, Renato Martins, Felipe C. Chamone, Mario Fernando Montenegro Campos, Ruzena Bajcsy
ICCV1
2019 On Modeling the Effects of Auditory Annoyance on Driving Style and Passenger Comfort
abstract
Despite the impressive progress being made in autonomous vehicles, human drivers will remain ubiquitous in the imminent years. Therefore, intelligent hybrid vehicular systems must be aware of the interactions between humans and the environment (e.g., sound, vibration, speed, etc.). In this paper, we evaluate the effect of acoustic annoyance on drivers in a real-world driving study. We found significant differences in driving styles elicited by annoying acoustics and present an online classifier that uses onboard inertial measurement unit measurements to distinguish whether a driver is annoyed with 77% accuracy. Moreover, we directly measured the forces applied on the passenger with a pressure mat lined on the car seat, and empirically confirm that our proposed passenger dynamics model is reasonable. However, due to our acoustically induced driving styles not being polarizing enough, we were unable to show that passengers' self-reported ride comfort changed with acoustic annoyance.
Edson Araujo, Michal Gregor, Isabella Huang, Erickson R. Nascimento, Ruzena Bajcsy
IROS4
2019 Prototypicality Effects in Global Semantic Description of Objects
abstract
In this paper, we introduce a novel approach for semantic description of object features based on the prototypicality effects of the Prototype Theory. Our prototype-based description model encodes and stores the semantic meaning of an object, while describing its features using the semantic prototype computed by CNN-classifications models. Our method uses semantic prototypes to create discriminative descriptor signatures that describe an object highlighting its most distinctive features within the category. Our experiments show that: i) our descriptor preserves the semantic information used by the CNN-models in classification tasks; ii) our distance metric can be used as the object's typicality score; iii) our descriptor signatures are semantically interpretable and enables the simulation of the prototypical organization of objects within a category.
Omar Vidal Pino, Erickson R. Nascimento, Mario Fernando Montenegro Campos
WACV2
2018 A Weighted Sparse Sampling and Smoothing Frame Transition Approach for Semantic Fast-Forward First-Person Videos
abstract
Thanks to the advances in the technology of low-cost digital cameras and the popularity of the self-recording culture, the amount of visual data on the Internet is going to the opposite side of the available time and patience of the users. Thus, most of the uploaded videos are doomed to be forgotten and unwatched in a computer folder or website. In this work, we address the problem of creating smooth fast-forward videos without losing the relevant content. We present a new adaptive frame selection formulated as a weighted minimum reconstruction problem, which combined with a smoothing frame transition method accelerates first-person videos emphasizing the relevant segments and avoids visual discontinuities. The experiments show that our method is able to fast-forward videos to retain as much relevant information and smoothness as the state-of-the-art techniques in less time. We also present a new 80-hour multimodal (RGB-D, IMU, and GPS) dataset of first-person videos with annotations for recorder profile, frame scene, activities, interaction, and attention.
Michel Melo Silva, Washington L. S. Ramos, João P. K. Ferreira, Felipe C. Chamone, Mario Fernando Montenegro Campos, Erickson R. Nascimento
CVPR6
2018 Visual-Quality-Driven Learning for Underwater Vision Enhancement
abstract
The image processing community has witnessed remarkable advances in enhancing and restoring images. Nevertheless, restoring the visual quality of underwater images remains a great challenge. End-to-end frameworks might fail to enhance the visual quality of underwater images since in several scenarios it is not feasible to provide the ground truth of the scene radiance. In this work, we propose a CNN-based approach that does not require ground truth data since it uses a set of image quality metrics to guide the restoration learning process. The experiments showed that our method improved the visual quality of underwater images preserving their edges and also performed well considering the UCIQE metric.
Walysson V. Barbosa, Henrique G. B. Amaral, Thiago Lages Rocha, Erickson R. Nascimento
ICIP4
2018 A Two-Step Learning Method for Detecting Landmarks on Faces from Different Domains
abstract
The detection of fiducial points on faces has significantly been favored by the rapid progress in the field of machine learning, in particular in the convolution networks. However, the accuracy of most of the detectors strongly depends on an enormous amount of annotated data. In this work, we present a domain adaptation approach based on a two-step learning to detect fiducial points on human and animal faces. We evaluate our method on three different datasets composed of different animal faces (cats, dogs, and horses). The experiments show that our method performs better than state of the art and can use few annotated data to leverage the detection of landmarks reducing the demand for large volume of annotated data.
Bruna Vieira Frade, Erickson R. Nascimento
ICIP2
2018 A 3D modeling methodology based on a concavity-aware geometric test to create 3D textured coarse models from concept art and orthographic projections
Sergio N. Silva Junior, Felipe C. Chamone, Renato Ferreira 0001, Erickson R. Nascimento
Comput. Graph.4
2018 Single-shot underwater image restoration: A visual quality-aware method based on light propagation model
Wagner Barros, Erickson R. Nascimento, Walysson V. Barbosa, Mario Fernando Montenegro Campos
J. Vis. Commun. Image Represent.2
2018 Making a long story short: A multi-importance fast-forwarding egocentric videos with the emphasis on relevant objects
Michel Melo Silva, Washington L. S. Ramos, Felipe C. Chamone, João P. K. Ferreira, Mario Fernando Montenegro Campos, Erickson R. Nascimento
J. Vis. Commun. Image Represent.6
2017 A Robust Indoor Scene Recognition Method Based on Sparse Representation
Guilherme Nascimento, Camila Laranjeira, Vinicius Braz, Anísio Lacerda, Erickson R. Nascimento
CIARP5
2017 Towards an efficient 3D model estimation methodology for aerial and ground images
Guilherme A. Potje, Gabriel Resende, Mario Fernando Montenegro Campos, Erickson R. Nascimento
Mach. Vis. Appl.4
2017 KVD: Scale invariant keypoints by combining visual and depth data
Levi O. Vasconcelos, Erickson R. Nascimento, Mario Fernando Montenegro Campos
Pattern Recognit. Lett.2
2016 Fast-forward video based on semantic extraction
abstract
Thanks to the low operational cost and large storage capacity of smartphones and wearable devices, people are recording many hours of daily activities, sport actions and home videos. These videos, also known as egocentric videos, are generally long-running streams with unedited content, which make them boring and visually unpalatable, bringing up the challenge to make egocentric videos more appealing. In this work we propose a novel methodology to compose the new fast-forward video by selecting frames based on semantic information extracted from images. The experiments show that our approach outperforms the state-of-the-art as far as semantic information is concerned and that it is also able to produce videos that are more pleasant to be watched.
Washington L. S. Ramos, Michel Melo Silva, Mario Fernando Montenegro Campos, Erickson R. Nascimento
ICIP4
2016 Real-time monocular obstacle avoidance using Underwater Dark Channel Prior
abstract
In this paper we propose a new vision-based obstacle avoidance strategy using the Underwater Dark Channel Prior (UDCP) that can be applied to any Unmanned Underwater Vehicle (UUV) equipped with a simple monocular camera and minimal on-board processing capabilities. For each incoming image, our method first computes a relative depth map to estimate the obstacles nearby. Then, the map is segmented and the most promising Region of Interest (RoI) is identified. Finally, an escape direction is computed within the RoI and a control action is performed accordingly to avoid the obstacles. We tested our approach on a video sequence in a natural environment and compared it against a state-of-the-art method showing better performance, specially in light changing conditions. We also provide online results on a low-cost Remotely Operated Vehicle (ROV) in a controlled environment.
Paulo L. J. Drews-Jr, Emili Hernández, Alberto Elfes, Erickson R. Nascimento, Mario Fernando Montenegro Campos
IROS4
2016 High performance moves recognition and sequence segmentation based on key poses filtering
abstract
We present a discriminative key pose-based approach for moves recognition and segmentation of training sequences for high performance sports. Compared to daily human gestures, moves in high performance sports are faster and have low inter-class variability, which produce noisy features and ambiguity. Our approach combines a robust filtering strategy to select frames composed of discriminative poses (key poses) and the discriminative Latent-Dynamic Conditional Random Fields (LDCRF) model to predict a label for each frame from the training sequence. We evaluate our approach on unsegmented sequences of Taekwondo training. Experimental results indicate that our methodology outperforms the Decision Forests method in terms of efficiency and accuracy. Our average recognition rate was equal to 74.72% while Decision Forests achieves 58.29%. The experiments also show that our approach was able to recognize and segment high speed moves like roundhouse kicks, which can reach peak linear speeds up to 26 m/s.
Claudio Marcio de Souza Vicente, Erickson R. Nascimento, Luiz Eduardo C. Emery, Cristiano Arruda G. Flor, Thales Vieira, Leonardo B. Oliveira
WACV2
2015 A Non-parametric Approach to Detect Changes in Aerial Images
Marco Túlio Alves N. Rodrigues, Daniel Balbino de Mesquita, Erickson R. Nascimento, William Robson Schwartz
CIARP3
2015 A Scale Invariant Keypoint Detector Based on Visual and Geometrical Cues
Levi O. Vasconcelos, Erickson R. Nascimento, Mario Fernando Montenegro Campos
CIARP2
2015 Automatic restoration of underwater monocular sequences of images
abstract
Underwater environments present a considerable challenge for computer vision, since water is a scattering medium with substantial light absorption characteristics which is made even more severe by turbidity. This poses significant problems for visual underwater navigation, object detection, tracking and recognition. Previous works tackle the problem by using unreliable priors or expensive and complex devices. This paper adopts a physical underwater light attenuation model which is used to enhance the quality of images and enable the applicability of traditional computer vision techniques images acquired from underwater scenes. The proposed method simultaneously estimates the attenuation parameter of the medium and the depth map of the scene to compute the image irradiance thus reducing the effect of the medium in the images. Our approach is based on a novel optical flow method, which is capable of dealing with scattering media, and a new technique that robustly estimates the medium parameters. Combined with structure-from-motion techniques, the depth map is estimated and a model-based restoration is performed. The method was tested both with simulated and real sequences of images. The experimental images were acquired with a camera mounted on a Remotely Operated Vehicle (ROV) navigating in a naturally lit, shallow seawater. The results show that the proposed technique allows for substantial restoration of the images, thereby improving the ability to identify and match features, which in turn is an essential step for other computer vision algorithms such as object detection and tracking, and autonomous navigation.
Paulo L. J. Drews-Jr, Erickson R. Nascimento, Mario Fernando Montenegro Campos, Alberto Elfes
IROS2
2014 Change detection based on features invariant to monotonic transforms and spatial constrained matching
abstract
Discovering regions that have changed in a set of images acquired from a scene at different times and possibly from different view points and cameras is a crucial step for many image processing applications. Remote sensing, visual surveillance, medical diagnosis and treatment, civil infrastructure, and underwater sensing are some examples of such applications. This work proposes a novel approach to detect changes automatically without a learning step by using image analysis techniques and segmentation based on superpixels. Unlike most common approaches, which are pixel-based, we present an approach that combines super-pixel extraction, hierarchical clustering and segment matching. The experimental results show the effectiveness of the proposed approach comparing it a background subtraction technique, demonstrating the robustness of our algorithm to illumination variations, non-uniform attenuations, atmospheric absorption and swaying trees.
Marco Túlio Alves N. Rodrigues, Luciano O. Milen, Erickson R. Nascimento, William Robson Schwartz
ICASSP3
2014 Generalized Optical Flow Model for Scattering Media
abstract
This paper proposes a novel methodology to estimate the optical flow in scattering media, which consists on new formulation based on the classical Horn-Schunk approach and the optical image formation model. Our formulation is able to deal with the hard problem of tracking points in a medium where there is absorption and scattering effects. This approach generalizes assumptions of the Horn-Schunk model in order to tackle both non-scattering and scattering media. Our approach uses the Dark Channel Prior to estimate the scene transmission, which attains a significant improvement in the optical flow estimation in scattering media. We show that our approach outperformed state-of-the-art models and we provide a detailed analysis of our technique that shows its applicability to image sequences acquired both in simulated and real scenes.
Paulo L. J. Drews-Jr, Erickson R. Nascimento, Arthur Xavier, Mario Fernando Montenegro Campos
ICPR2
2014 On the improvement of human action recognition from depth map sequences using Space-Time Occupancy Patterns
Antônio Wilson Vieira, Erickson R. Nascimento, Gabriel L. Oliveira, Zicheng Liu 0001, Mario Fernando Montenegro Campos
Pattern Recognit. Lett.2
2014 Sparse Spatial Coding: A Novel Approach to Visual Recognition
abstract
Successful image-based object recognition techniques have been constructed founded on powerful techniques such as sparse representation, in lieu of the popular vector quantization approach. However, one serious drawback of sparse space-based methods is that local features that are quite similar can be quantized into quite distinct visual words. We address this problem with a novel approach for object recognition, called sparse spatial coding, which efficiently combines a sparse coding dictionary learning and spatial constraint coding stage. We performed experimental evaluation using the Caltech 101, Caltech 256, Corel 5000, and Corel 10000 data sets, which were specifically designed for object recognition evaluation. Our results show that our approach achieves high accuracy comparable with the best single feature method previously published on those databases. Our method outperformed, for the same bases, several multiple feature methods, and provided equivalent, and in few cases, slightly less accurate results than other techniques specifically designed to that end. Finally, we report state-of-the-art results for scene recognition on COsy Localization Dataset (COLD) and high performance results on the MIT-67 indoor scene recognition, thus demonstrating the generalization of our approach for such tasks.
Gabriel L. Oliveira, Erickson R. Nascimento, Antônio Wilson Vieira, Mario Fernando Montenegro Campos
IEEE Trans. Image Process.2
2013 On the development of a robust, fast and lightweight keypoint descriptor
Erickson R. Nascimento, Gabriel L. Oliveira, Antônio Wilson Vieira, Mario Fernando Montenegro Campos
Neurocomputing1
2012 STOP: Space-Time Occupancy Patterns for 3D Action Recognition from Depth Map Sequences
Antônio Wilson Vieira, Erickson R. Nascimento, Gabriel L. Oliveira, Zicheng Liu 0001, Mario Fernando Montenegro Campos
CIARP2
2012 EDVD - Enhanced descriptor for visual and depth data
Erickson R. Nascimento, William Robson Schwartz, Mario Fernando Montenegro Campos
ICPR1
2012 Sparse Spatial Coding: A novel approach for efficient and accurate object recognition
abstract
Successful state-of-the-art object recognition techniques from images have been based on powerful methods, such as sparse representation, in order to replace the also popular vector quantization (VQ) approach. Recently, sparse coding, which is characterized by representing a signal in a sparse space, has raised the bar on several object recognition benchmarks. However, one serious drawback of sparse space based methods is that similar local features can be quantized into different visual words. We present in this paper a new method, called Sparse Spatial Coding (SSC), which combines a sparse coding dictionary learning, a spatial constraint coding stage and an online classification method to improve object recognition. An efficient new off-line classification algorithm is also presented. We overcome the problem of techniques which make use of sparse representation alone by generating the final representation with SSC and max pooling, presented for an online learning classifier. Experimental results obtained on the Caltech 101, Caltech 256, Corel 5000 and Corel 10000 databases, show that, to the best of our knowledge, our approach supersedes in accuracy the best published results to date on the same databases. As an extension, we also show high performance results on the MIT-67 indoor scene recognition dataset.
Gabriel L. Oliveira, Erickson R. Nascimento, Antônio Wilson Vieira, Mario Fernando Montenegro Campos
ICRA2
2012 BRAND: A robust appearance and depth descriptor for RGB-D images
abstract
This work introduces a novel descriptor called Binary Robust Appearance and Normals Descriptor (BRAND), that efficiently combines appearance and geometric shape information from RGB-D images, and is largely invariant to rotation and scale transform. The proposed approach encodes point information as a binary string providing a descriptor that is suitable for applications that demand speed performance and low memory consumption. Results of several experiments demonstrate that as far as precision and robustness are concerned, BRAND achieves improved results when compared to state of the art descriptors based on texture, geometry and combination of both information. We also demonstrate that our descriptor is robust and provides reliable results in a registration task even when a sparsely textured and poorly illuminated scene is used.
Erickson R. Nascimento, Gabriel L. Oliveira, Mario Fernando Montenegro Campos, Antônio Wilson Vieira, William Robson Schwartz
IROS1
2007 Fully automatic coloring of grayscale images
Luiz Filipe M. Vieira, Erickson R. Nascimento, Fernando A. Fernandes Jr., Rodrigo L. Carceroni, Rafael D. Vilela, Arnaldo de Albuquerque Araújo
Image Vis. Comput.2