Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Marcos Roberto e Souza

dblp:226/2175 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
5since 2021 · last 2025
0000-0003-4342-5220ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 1 heaviest of 1, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing
video stabilization
0.912025
NAFT and SynthStab: A RAFT-Based Network and a Synthetic Dataset for Digital Video Stabilization · Int. J. Comput. Vis. 2025

Methods — techniques the papers use, named apart from their topics

optical flow · 0.9RAFT · 0.9
YearPublicationVenuePosition
2025 NAFT and SynthStab: A RAFT-Based Network and a Synthetic Dataset for Digital Video Stabilization
Marcos Roberto e Souza, Helena Almeida Maia, Hélio Pedrini
Int. J. Comput. Vis.1
2023 Variable-hyperparameter visual transformer for efficient image inpainting
Jose L. Flores-Campana, Luis G. L. Decker, Marcos Roberto e Souza, Helena Almeida Maia, Hélio Pedrini
Comput. Graph.3
2023 Rethinking two-dimensional camera motion estimation assessment for digital video stabilization: A camera motion field-based metric
Marcos Roberto e Souza, Helena Almeida Maia, Hélio Pedrini
Neurocomputing1
2021 Pyramidal Layered Scene Inference with Image Outpainting for Monocular View Synthesis
Marcos Roberto e Souza, Jhonatas Santos de Jesus Conceição, Jose L. Flores-Campana, Luis G. L. Decker, Diogo C. Luvizon, Gustavo Sutter 0002, Helena Almeida Maia, Hélio Pedrini
CAIP (1)1
2021 Adaptive Multiplane Image Generation from a Single Internet Picture
abstract
In the last few years, several works have tackled the problem of novel view synthesis from stereo images or even from a single picture. However, previous methods are computationally expensive, specially for high-resolution images. In this paper, we address the problem of generating a multiplane image (MPI) from a single high-resolution picture. We present the adaptive-MPI representation, which allows rendering novel views with low computational requirements. To this end, we propose an adaptive slicing algorithm that produces an MPI with a variable number of image planes. We present a new lightweight CNN for depth estimation, which is learned by knowledge distillation from a larger network. Occluded regions in the adaptive-MPI are inpainted also by a lightweight CNN. We show that our method is capable of producing high-quality predictions with one order of magnitude less parameters compared to previous approaches. The robustness of our method is evidenced on challenging pictures from the Internet.
Diogo C. Luvizon, Gustavo Sutter 0002, Andreza A. dos Santos, Jhonatas Santos de Jesus Conceição, Jose L. Flores-Campana, Luis G. L. Decker, Marcos Roberto e Souza, Hélio Pedrini, Antonio Joia, Otávio A. B. Penatti
WACV7
2020 Parallax Motion Effect Generation Through Instance Segmentation And Depth Estimation
abstract
Stereo vision is a growing topic in computer vision due to the innumerable opportunities and applications this technology offers for the development of modern solutions, such as virtual and augmented reality applications. To enhance the user's experience in three-dimensional virtual environments, the motion parallax estimation is a promising technique to achieve this objective. In this paper, we propose an algorithm for generating parallax motion effects from a single image, taking advantage of state-of-the-art instance segmentation and depth estimation approaches. This work also presents a comparison against such algorithms to investigate the trade-off between efficiency and quality of the parallax motion effects, taking into consideration a multi-task learning network capable of estimating instance segmentation and depth estimation at once. Experimental results and visual quality assessment indicate that the PyD-Net network (depth estimation) combined with Mask R-CNN or FBNet networks (instance segmentation) can produce parallax motion effects with good visual quality.
Allan Pinto, Manuel Alberto Cordova Neira, Luis G. L. Decker, Jose L. Flores-Campana, Marcos Roberto e Souza, Andreza A. dos Santos, Jhonatas Santos de Jesus Conceição, Henrique F. Gagliardi, Diogo C. Luvizon, Ricardo da Silva Torres, Hélio Pedrini
ICIP5
2020 Survey on visual rhythms: A spatio-temporal representation for video sequences
Marcos Roberto e Souza, Helena Almeida Maia, Marcelo Bernardes Vieira, Hélio Pedrini
Neurocomputing1
2019 Learnable Visual Rhythms Based on the Stacking of Convolutional Neural Networks for Action Recognition
abstract
Recent deep learning techniques have achieved satisfactory results for various image-related problems. However, many research questions remain open in tasks involving video sequences. Several applications demand the understanding of complex events in videos, such as traffic monitoring, person re-identification, security and surveillance. In this work, we address the problem of human action recognition in videos through a multi-stream network that incorporates both spatial and temporal information. The main contribution of our work is a stream based on a new variant of the visual rhythm, called Learnable Visual Rhythm (LVR). We employ a deep network to extract features from the video frames in order to generate the rhythm. The features are collected at multiple depths of the network to enable the analysis of different abstraction levels. This strategy significantly outperforms the handcrafted version on the UCF101 and HMDB51 datasets. Experiments conducted on these datasets show that our final multi-stream network achieved competitive results compared to state-of-the-art approaches.
Helena Almeida Maia, Marcos Roberto e Souza, Anderson Carlos Sousa e Santos, Hélio Pedrini, Hemerson Tacon, André de Souza Brito, Hugo de Lima Chaves, Marcelo Bernardes Vieira, Saulo Moraes Villela
ICMLA2
2019 Motion energy image for evaluation of video stabilization
Marcos Roberto e Souza, Hélio Pedrini
Vis. Comput.1
2018 Improvement of global motion estimation in two-dimensional digital video stabilisation methods
abstract
A large amount of video content has been produced by compact and portable cameras. Several applications have been benefited from such growth of multimedia data, such as telemedicine, business conferencing, surveillance and security, entertainment, distance learning, and robotics. Video stabilisation is the process of detecting and removing undesired motion or instabilities from a video stream caused during the acquisition stage when handling the camera. In this work, the authors introduce and analyse a novel approach that identifies failures in the global motion estimation of the camera by means of local features. Moreover, they propose an optimisation method for computing a new estimate of the corrected motion. Experiments conducted on different video sequences are performed to demonstrate the effectiveness of the developed method. Results obtained with the stabilisation process are compared against the state‐of‐the‐art YouTube method.
Marcos Roberto e Souza, Luiz Fernando Rodrigues da Fonseca, Hélio Pedrini
IET Image Process.1