Samuel Felipe dos Santos

dblp:232/1583 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-6061-5582ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Beyond Masking: Alternative Strategies for Generalizable Facial Expression Recognition
Sergio Neres Pereira Junior, Samuel Felipe dos Santos, Jurandy Almeida
CIARP (2)2
2025 A Comprehensive Evaluation of Deep Learning Architectures and Loss Functions for Lumbar Spine Segmentation in MRI
Claudio Leite, Samuel Felipe dos Santos, Jurandy Almeida
CIARP (2)2
2025 Transferable-Guided Attention Is All You Need for Video Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) in videos is a challenging task that remains not well explored compared to image-based UDA techniques. Although vision transformers (ViT) achieve state-of-the-art performance in many computer vision tasks, their use in video UDA has been little explored. Our key idea is to use transformer layers as a feature encoder and incorporate spatial and temporal transferability relationships into the attention mechanism. A Transferable-guided Attention (TransferAttn) framework is then developed to exploit the capacity of the transformer to adapt cross-domain knowledge across different backbones. To improve the transferability of ViT, we introduce a novel and effective module, named Domain Transferable-guided Attention Block (DTAB). DTAB compels ViT to focus on the spatio-temporal transferability relationship among video frames by changing the self-attention mechanism to a transferability attention mechanism. Extensive experiments were conducted on UCF-HMDB, Kinetics-Gameplay, and Kinetics-NEC Drone datasets, with different backbones, like ResNet101, I3D, and STAM, to verify the effectiveness of TransferAttn compared with state-of-the-art approaches. Also, we demonstrate that DTAB yields performance gains when applied to other state-of-the-art transformer-based UDA methods from both video and image domains. Our code is available at https://github.com/Andre-Sacilotti/transferattn-project-code.
André Sacilotti, Samuel Felipe dos Santos, Nicu Sebe, Jurandy Almeida
WACV2
2025 Budget-aware pruning: Handling multiple domains with less parameters
Samuel Felipe dos Santos, Rodrigo Ferreira Berriel, Thiago Oliveira-Santos, Nicu Sebe, Jurandy Almeida
Pattern Recognit.1
2025 Beyond the known: Enhancing Open Set Domain Adaptation with unknown exploration
Lucas Fernando Alvarenga e Silva, Samuel Felipe dos Santos, Nicu Sebe, Jurandy Almeida
Pattern Recognit. Lett.2
2022 Weakly supervised learning based on hypergraph manifold ranking
João Gabriel Camacho Presotto, Samuel Felipe dos Santos, Lucas Pascotti Valem, Fábio Augusto Faria, João Paulo Papa, Jurandy Almeida, Daniel C. G. Pedronette
J. Vis. Commun. Image Represent.2
2021 Less Is More: Accelerating Faster Neural Networks Straight from JPEG
abstract
Most image data available are often stored in a compressed format, from which JPEG is the most widespread. To feed this data on a convolutional neural network (CNN), a preliminary decoding process is required to obtain RGB pixels, demanding a high computational load and memory usage. For this reason, the design of CNNs for processing JPEG compressed data has gained attention in recent years. In most existing works, typical CNN architectures are adapted to facilitate the learning with the DCT coefficients rather than RGB pixels. Although they are effective, their architectural changes either raise the computational costs or neglect relevant information from DCT inputs. In this paper, we examine different ways of speeding up CNNs designed for DCT inputs, exploiting learning strategies to reduce the computational complexity by taking full advantage of DCT inputs. Our experiments were conducted on the ImageNet dataset. Results show that learning how to combine all DCT inputs in a data-driven fashion is better than discarding them by hand, and its combination with a reduction of layers has proven to be effective for reducing the computational costs while retaining accuracy.
Samuel Felipe dos Santos, Jurandy Almeida
CIARP1
2020 The Good, The Bad, and The Ugly: Neural Networks Straight From JPEG
abstract
Over the past decade, convolutional neural networks (CNNs) have achieved state-of-the-art performance in many computer vision tasks. They can learn robust representations of image data by processing RGB pixels. Since image data are often stored in a compressed format, from which JPEG is the most widespread, a preliminary decoding process is demanded. Recently, the design of CNNs for processing JPEG compressed data has gained attention from the research community. They process DCT coefficients instead of RGB pixels, saving computation for decoding JPEG images, however, at the cost of increasing the computational complexity of the network. In this paper, we examine how spatial resolution and JPEG quality impacts on the performance of a state-of-the-art CNN designed to operate directly on the JPEG compressed domain. To alleviate its computational complexity, we propose a Frequency Band Selection (FBS) technique to select the most relevant DCT coefficients before feeding them to the network. Experiments were conducted on a subset of the ImageNet dataset considering both fine- and coarse-grained image classification tasks. Results show that such networks are resilient to JPEG quality but are susceptible to spatial resolution. Also, our FBS can reduce the computational complexity of the network while retaining a similar accuracy.
Samuel Felipe dos Santos, Nicu Sebe, Jurandy Almeida
ICIP1