Silvio Jamil Ferzoli Guimarães

dblp:08/3856 · DBLP profile ↗
← Back
58ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0001-8522-2056ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hierarchical Multi-Scale Deep Neural Network for Schizophrenia Detection in Neuroimaging
abstract
Schizophrenia remains difficult to diagnose due to its reliance on subjective clinical assessment.This work proposes a pipeline for automated schizophrenia classification using functional MRI data from the UCLA CNP dataset.The method extracts multi-view slices from nine anatomical orientations using a hierarchical analysis and processes them with a Vision Transformer model (MultiSliceViT).Under stratified 5-fold cross-validation, the approach achieved 82.6% accuracy, outperforming models with fewer views.Interpretability analyses highlighted consistent attention to key regions, including the dorsolateral prefrontal cortex, hippocampus, and anterior cingulate.These results demonstrate the effectiveness of multi-view transformer architectures for identifying meaningful functional biomarkers.
Carlos Dias Maia, Gabriel Barbosa da Fonseca, Luis E. Zárate, Silvio Jamil Ferzoli Guimarães
ESANN4
2025 Multilingual Visual Understanding: Extending Visual Dialog to Portuguese and Spanish Through Cross-Modal Adaptation
Milena Menezes Adão, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
CIARP2
2025 Detecting Hierarchical Inconsistencies from an Image Segmentation Series
Fábio Kochem, Felipe Belém, Zenilton Kleber Gonçalves do Patrocínio Jr., Benjamin Perret, Jean Cousty, Alexandre Falcao, Silvio Jamil Ferzoli Guimarães
CIARP (2)7
2025 Lightweight Graph Neural Networks for 3D Shape Classification
Eduardo Felipe Lopes, João Pedro Oliveira Batisteli, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
CIARP (2)3
2025 Do Superpixel Segmentation Methods Influence Deforestation Image Classification?
Hugo Resende, Fábio Augusto Faria, Eduardo B. Neto, Isabela Borlido, Victor Sundermann, Silvio Jamil Ferzoli Guimarães, Alvaro Luiz Fazenda 0001
CIARP6
2025 Extending Visual Dialog Beyond English: An Analysis of Monolingual and Multilingual Models
abstract
Visual Dialog is a challenging multimodal task requiring models to answer questions about images through multi-turn conversations. Despite significant progress, research has predominantly focused on English, limiting applicability to the 850+ million speakers of Portuguese and Spanish worldwide. We present the first comprehensive study of monolingual and multilingual Visual Dialog models for Portuguese and Spanish, introducing novel insights into cross-lingual visual grounding mechanisms. Through extensive experiments on newly translated VisDial datasets, we compare language-specific encoders (BERTimbau for Portuguese, BETO for Spanish) against multilingual BERT, achieving competitive performance with monolingual models while revealing distinct cross-modal attention patterns. Our mechanistic interpretability analysis demonstrates that despite different tokenization strategies and pretraining objectives, both approaches converge to similar attention distributions in deeper layers, with divergence decreasing from 0.000832 (Layer 0) to 0.000490 (Layer 11). We find that monolingual models exhibit holistic attention strategies while multilingual models show more selective, fine-grained visual grounding. These findings have important implications for developing inclusive vision-language technologies and understanding cross-lingual transfer in multimodal contexts.
Milena M. Adão, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
ISM2
2025 Hierarchical Multiscale Representation in Remote Sensing Scene Classification
abstract
Remote sensing scene classification (RSSC) poses significant challenges due to high spatial variability, complex textures, and semantic ambiguity in remote sensing imagery. While convolutional neural networks (CNNs) and transformer-based models have achieved notable success in this domain, their performance often depends on large-scale pretraining and substantial computational resources. Graph neural networks (GNNs) have emerged as a promising alternative to traditional deep learning methods by explicitly modeling the relational structure of image regions through graph representations, which have already demonstrated promising results across various image-based tasks involving images. In this work, we explore two GNN architectures tailored for RSSC: BRMv2, a novel simplified graph model built on a base region adjacency graph (RAG), and modified hierarchical layered multigraph network (mHELMNet), a modified hierarchical multigraph model that encodes multiscale and spatial relationships through a multigraph representation. Both models were evaluated on the EUROSAT and RESISC45 datasets, achieving accuracy comparable to, or in some cases exceeding, that of state-of-the-art CNN-based, hybrid GNN-based, and transformer-based methods, while using significantly fewer parameters and without relying on pretraining. Experimental results demonstrated that the proposed GNN models, mHELMNet and BRMv2, achieved over 96% accuracy on EUROSAT and approximately 85% on RESISC45, while requiring only 0.14% and 0.03% of the parameters of the leading transformer-based approach, respectively.
João Pedro Oliveira Batisteli, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
IEEE Geosci. Remote. Sens. Lett.2
2025 Streamlined extended Long Short-Term Memory for video skimming
Leonardo Vilela Cardoso, Barbara Hellen P. Soraggi, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
Pattern Recognit. Lett.3
2024 Seed-Based Superpixel Re-Segmentation for Improving Object Delineation
Lucca S. P. Lacerda, Felipe Belém, Zenilton Kleber Gonçalves do Patrocínio Jr., Alexandre X. Falcão, Silvio Jamil Ferzoli Guimarães
CIARP (1)5
2024 Towards Interactive Video Segmentation by Dynamic and Iterative Spanning Forest
Danielle Vieira, Isabela Borlido Barcelos, Zenilton Kleber Gonçalves do Patrocínio Jr., Alexandre X. Falcão, Silvio Jamil Ferzoli Guimarães
CIARP (1)5
2024 How to Identify Good Superpixels for Deforestation Detection on Tropical Rainforests
abstract
The conservation of tropical forests is a topic of significant social and ecological relevance due to their crucial role in the global ecosystem. Unfortunately, deforestation and degradation impact millions of hectares annually, requiring government or private initiatives for effective forest monitoring. However, identifying deforested regions in satellite images is challenging due to data imbalance, image resolution, low-contrast regions, and occlusion. Superpixel segmentation can overcome these drawbacks, reducing workload and preserving important image boundaries. However, most works for remote-sensing images do not exploit recent superpixel methods. In this work, we evaluate 16 superpixel methods in satellite images to support a deforestation detection system in tropical forests. We also assess the performance of superpixel methods for the target task, establishing a relationship with segmentation methodological evaluation. According to our results, ERS, GMMSP, and DISF perform best on undersegmentation error (UE), boundary recall (BR), and similarity between image and reconstruction from superpixels (SIRSs), respectively, whereas ERS has the best tradeoff with compactness index (CO) and Reg. In classification, SH, DISF, and ISF perform best on RGB, UMDA, and PCA compositions, respectively. According to our experiments, superpixel methods with better tradeoffs among delineation, homogeneity, compactness, and regularity are more suitable for identifying good superpixels for deforestation detection tasks.
Isabela Borlido Barcelos, Eduardo Bouhid, Victor Sundermann, Hugo Resende, Alvaro Luiz Fazenda 0001, Fábio Augusto Faria, Silvio Jamil Ferzoli Guimarães
IEEE Geosci. Remote. Sens. Lett.7
2023 Graph-Based Feature Learning from Image Markers
Isabela Borlido Barcelos, Leonardo de Melo Joao, Zenilton Kleber Gonçalves do Patrocínio Jr., Ewa Kijak, Alexandre X. Falcão, Silvio Jamil Ferzoli Guimarães
CIARP6
2023 Filtering Safe Temporal Motifs in Dynamic Graphs for Dissemination Purposes
Carolina Stephanie Jerônimo de Almeida, Simon Malinowski, Zenilton Kleber Gonçalves do Patrocínio Jr., Guillaume Gravier, Silvio Jamil Ferzoli Guimarães
CIARP5
2023 Streaming Graph-Based Supervoxel Computation Based on Dynamic Iterative Spanning Forest
Danielle Vieira, Isabela Borlido Barcelos, Felipe Belém, Zenilton Kleber Gonçalves do Patrocínio Jr., Alexandre X. Falcão, Silvio Jamil Ferzoli Guimarães
CIARP6
2023 A Novel Method for Temporal Graph Classification based on Transitive Reduction
abstract
Domains such as bio-informatics, social network analysis, and computer vision, describe relations between entities and cannot be interpreted as vectors or fixed grids, instead, they are naturally represented by graphs. Often this kind of data evolves over time in a dynamic world, respecting a temporal order being known as temporal graphs. The latter became a challenge since subgraph patterns are very difficult to find and the distance between those patterns may change irregularly over time. While state-of-the-art methods are primarily designed for static graphs and may not capture temporal information, recent works have proposed mapping temporal graphs to static graphs to allow for the use of conventional static kernels and graph neural approaches. In this study, we compare the transitive reduction impact on these mappings in terms of accuracy and computational efficiency across different classification tasks. Furthermore, we introduce a novel mapping method using a transitive reduction approach that outperforms existing techniques in terms of classification accuracy. Our experimental results demonstrate the effectiveness of the proposed mapping method in improving the accuracy of supervised classification for temporal graphs while maintaining reasonable computational efficiency.
Carolina Stephanie Jerônimo de Almeida, Zenilton Kleber Gonçalves do Patrocínio Jr., Simon Malinowski, Silvio Jamil Ferzoli Guimarães, Guillaume Gravier
DSAA4
2023 Multi-Scale Image Graph Representation: A Novel GNN Approach for Image Classification through Scale Importance Estimation
abstract
Image representation as graphs can enhance the understanding of image semantics and facilitate multi-scale image representation. However, existing methods often overlook the significance of the relationship between elements at each scale or fail to encode the hierarchical relationship between graph elements. To cope with that, we introduce a novel approach for graph construction from images. This approach utilizes a hierarchical image segmentation technique to generate segmentation at multiple scales and incorporates edges to encode the relationships at each scale. We also propose a new readout function that weighs the importance of each scale when deriving a fixed-size graph representation. Furthermore, we present a new model incorporating those ideas – called Hierarchical Image Graph with Scale Importance (HIGSI). Experimental results on the CIFAR-10 database indicate that our proposed model outperforms (or closely matches) state-of-the-art and baseline models while utilizing smaller graphs.
João Pedro Oliveira Batisteli, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
ISM2
2023 Graph-based image gradients aggregated with random forests
Raquel Almeida 0001, Ewa Kijak, Simon Malinowski, Zenilton Kleber Gonçalves do Patrocínio Jr., Arnaldo de Albuquerque Araújo, Silvio Jamil Ferzoli Guimarães
Pattern Recognit. Lett.6
2021 Enhanced-Memory Transformer for Coherent Paragraph Video Captioning
abstract
A coherent description is an ultimate goal concerning video captioning through multiple sentences since it may directly impact consistency and intelligibility. A paragraph describing a video is affected by different events extracted from it. When generating a new event, it should produce a detailed narrative of the video content. But it also might provide some clues that may help reduce the textual repetition occurring in the final description. Recently, transformers have emerged as an appealing solution to several tasks, including video captioning. An augmented transformer with a memory module can somehow cope with text repetition. Thus, to further increase the coherence among the generated sentences, we propose the adoption of attention mechanisms to enhance memory data in a memory-augmented transformer. This new approach, called Enhanced-Memory Transformer (EMT), assesses the data importance (about the video segments) contained in the memory module and uses that to improve readability by reducing repetition. Experimental evaluation of EMT using the test split of the ActivityNet Captions dataset achieved 22.84 in CIDEr-D score, and 4.55 in Reduction-4 score (R@4), representing improvements of 1.03% and 16.36%, respectively (compared to the literature). The obtained results show the great potential of this new approach as it provides increased coherence among the various video segments, decreasing the repetition in the generated sentences and improving the description diversity.
Leonardo Vilela Cardoso, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
ICTAI2
2021 Hierarchical multi-label propagation using speaking face graphs for multimodal person discovery
Gabriel Barbosa da Fonseca, Gabriel Sargent, Ronan Sicre, Zenilton Kleber Gonçalves do Patrocínio Jr., Guillaume Gravier, Silvio Jamil Ferzoli Guimarães
Multim. Tools Appl.6
2020 Image segmentation using dense and sparse hierarchies of superpixels
Felipe L. Galvão, Silvio Jamil Ferzoli Guimarães, Alexandre X. Falcão
Pattern Recognit.2
2020 Learning to realign hierarchy for image segmentation
Milena M. Adão, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
Pattern Recognit. Lett.2
2020 Efficient hierarchical graph partitioning for image segmentation by optimum oriented cuts
abstract
In this work, a hierarchical graph partitioning based on optimum cuts in graphs is proposed for unsupervised image segmentation, that can be tailored to the target group of objects, according to their boundary polarity, by extending Oriented Image Foresting Transform (OIFT). The proposed method, named UOIFT, theoretically encompasses as a particular case the single-linkage algorithm by minimum spanning tree (MST) and gives superior segmentation results compared to other approaches commonly used in the literature, usually requiring a lower number of image partitions to accurately isolate the desired regions of interest with known polarity. The method is supported by new theoretical results involving the usage of non-monotonic-incremental cost functions in directed graphs and exploits the local contrast of image regions, being robust in relation to illumination variations and inhomogeneity effects. UOIFT is demonstrated using a region adjacency graph of superpixels in medical and natural images.
Hans Harley Ccacyahuillca Bejar, Silvio Jamil Ferzoli Guimarães, Paulo André Vechiatto Miranda
Pattern Recognit. Lett.2
2020 Hierarchical segmentation from a non-increasing edge observation attribute
Edward Cayllahua, Jean Cousty, Silvio Jamil Ferzoli Guimarães, Yukiko Kenmochi, Guillermo Cámara Chávez, Arnaldo de Albuquerque Araújo
Pattern Recognit. Lett.3
2020 Superpixel Segmentation Using Dynamic and Iterative Spanning Forest
abstract
As constituent parts of image objects, superpixels can improve several higher-level operations. However, image segmentation methods might have their accuracy severely compromised for reduced numbers of superpixels. To mitigate the problem, we introduce Dynamic Iterative Spanning Forest (DISF), a seed-based method that improves all components in the Iterative Spanning Forest (ISF) framework for superpixel segmentation. DISF relies on a new strategy for seed estimation that can find more relevant seeds, reconstruct relevant edges along with iterations, and guarantee the desired number of superpixels. DISF also assures optimal spanning forests for path costs based on dynamic arc-weight estimation, being faster as the desired number of superpixels grows. We show that DISF can improve effectiveness on three datasets with distinct object properties, requiring significantly fewer iterations than all seed-based baselines.
Felipe Belém, Silvio Jamil Ferzoli Guimarães, Alexandre X. Falcão
IEEE Signal Process. Lett.2
2019 BRIEF-Based Mid-Level Representations for Time Series Classification
Renato Augusto de Souza, Raquel Almeida 0001, Roberto Miranda, Zenilton Kleber Gonçalves do Patrocínio Jr., Simon Malinowski, Silvio Jamil Ferzoli Guimarães
CIARP6
2019 Combining convolutional side-outputs for road image segmentation
abstract
Image segmentation consists in creating partitions within an image into meaningful areas and objects. It can be used in scene understanding and recognition, in fields like biology, medicine, robotics, satellite imaging, amongst others. In this work we take advantage of the learned model in a deep architecture, by extracting side-outputs at different layers of the network for the task of image segmentation. We study the impact of the amount of side-outputs and evaluate strategies to combine them. A post-processing filtering based on mathematical morphology idempotent functions is also used in order to remove some undesirable noises. Experiments were performed on the publicly available KITTI Road Dataset for image segmentation. Our comparison shows that the use of multiples side outputs can increase the overall performance of the network, making it easier to train and more stable when compared with a single output in the end of the network. Also, for a small number of training epochs (500), we achieved a competitive performance when compared to the best algorithm in KITTI Evaluation Server.
Felipe A. L. Reis, Raquel Almeida 0001, Ewa Kijak, Simon Malinowski, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
IJCNN5
2019 Efficient Algorithms for Hierarchical Graph-Based Segmentation Relying on the Felzenszwalb-Huttenlocher Dissimilarity
abstract
Hierarchical image segmentation provides a region-oriented scale-space, i.e. a set of image segmentations at different detail levels in which the segmentations at finer levels are nested with respect to those at coarser levels. However, most image segmentation algorithms, among which a graph-based image segmentation method relying on a region merging criterion was proposed by Felzenszwalb–Huttenlocher in 2004, do not lead to a hierarchy. In order to cope with a demand for hierarchical segmentation, Guimarães et al. proposed in 2012 a method for hierarchizing the popular Felzenszwalb–Huttenlocher method, without providing an algorithm to compute the proposed hierarchy. This paper is devoted to providing a series of algorithms to compute the result of this hierarchical graph-based image segmentation method efficiently, based mainly on two ideas: optimal dissimilarity measuring and incremental update of the hierarchical structure. Experiments show that, for an image of size 321 × 481 pixels, the most efficient algorithm produces the result in half a second whereas the most naive one requires more than 4 h.
Edward Cayllahua, Jean Cousty, Yukiko Kenmochi, Arnaldo de Albuquerque Araújo, Guillermo Cámara Chávez, Silvio Jamil Ferzoli Guimarães
Int. J. Pattern Recognit. Artif. Intell.6
2019 Removing non-significant regions in hierarchical clustering and segmentation
abstract
We propose an efficient algorithm that removes unimportant regions from a hierarchical partition tree, while preserving the hierarchical partition structure. Various experiments demonstrate that applying this algorithm on various classification or segmentation problems does indeed improve the results by a large margin. Code is available online at https://github.com/higra/Higra. B. Perret, G. Chierchia, J. Cousty, S.J. F. Guimarães, Y. Kenmochi, L. Najman, Higra: Hierarchical Graph Analysis, SoftwareX, 10, 1--6, ISSN 2019, 2352-7110, 10.1016/j.softx.2019.100335.
Benjamin Perret, Jean Cousty, Silvio Jamil Ferzoli Guimarães, Yukiko Kenmochi, Laurent Najman
Pattern Recognit. Lett.3
2018 Evaluation of Scale-Aware Realignments of Hierarchical Image Segmentation
Milena M. Adão, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
CIARP2
2018 Evaluation of Bag-of-Word Performance for Time Series Classification Using Discriminative SIFT-Based Mid-Level Representations
Raquel Almeida 0001, Hugo Herlanin, Zenilton Kleber Gonçalves do Patrocínio Jr., Simon Malinowski, Silvio Jamil Ferzoli Guimarães
CIARP5
2018 Superpixel Segmentation by Object-Based Iterative Spanning Forest
Felipe Belém, Silvio Jamil Ferzoli Guimarães, Alexandre X. Falcão
CIARP2
2018 Image Inpainting Based on Local Patch Search Supported by Image Segmentation
Sarah Almeida Carneiro, Hélio Pedrini, Silvio Jamil Ferzoli Guimarães
CIARP3
2018 Hierarchy-Based Salient Regions: A Region Detector Based on Hierarchies of Partitions
Karla Otiniano-Rodríguez, Arnaldo de Albuquerque Araújo, Guillermo Cámara Chávez, Jean Cousty, Silvio Jamil Ferzoli Guimarães, Benjamin Perret
CIARP5
2018 Hierarchical Graph-Based Segmentation in Detection of Object-Related Regions
Rafael Machado Ribeiro, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
CIARP2
2018 Evaluation of Hierarchical Watersheds
abstract
This paper aims to understand the practical features of hierarchies of morphological segmentations, namely the quasi-flat zones hierarchy and watershed hierarchies, and to evaluate their potential in the context of natural image analysis. We propose a novel evaluation framework for the hierarchies of partitions designed to capture various aspects of those representations: precision of their regions and contours, possibility to extract high quality horizontal cuts and optimal non-horizontal cuts for image segmentation, and the ease of finding a set of regions representing a semantic object. This framework is used to assess and to optimize hierarchies with respect to the possible pre- and post-processing steps. We show that, used in conjunction with a state-of-the-art contour detector, watershed hierarchies are competitive with the complex state-of-the-art methods for hierarchy construction. In particular, the proposed framework allows us to identify a watershed hierarchy based on a novel extinction value, the number of parent nodes that outperforms the other hierarchies of morphological segmentations. This coupled with the fact that watershed hierarchies satisfy clear global optimality properties and can be efficiently computed on large data, make them valuable candidates for various computer vision tasks.
Benjamin Perret, Jean Cousty, Silvio Jamil Ferzoli Guimarães, Deise Santana Maia
IEEE Trans. Image Process.3
2017 Exploring quantization error to improve human action classification
abstract
In human action classification task, a video must be classified into a pre-determined class. To cope with this problem, we propose a mid-level representation, in which information about quantization errors is embedded together with the aggregated data on low level features. The main contributions of this article are twofold: (i) assembly of low-level features (dense trajectories) by a mid-level representation enriched with information about distances between descriptors and codewords; and (ii) a survey of the most common protocols for human action classification methods when applied to three different datasets. Regarding classification protocols, we have experimented the training and testing classification (called split), the leave-one-out cross-validation (LOOCV) and the leave-one-group-out cross-validation (25-fold CV). Experimental results demonstrated that our strategy either has improved the classification rates with respect to the state-of-the-art for KTH dataset, achieving 98%, or it is a competitive one, for UCF-11 with 90%, when compared with methods with no feature learning.
Raquel Almeida 0001, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães
IJCNN3
2017 Combining pixel domain and compressed domain index for sketch based image retrieval
Carlos A. F. Pimentel Filho, Benjamin Bustos, Arnaldo de Albuquerque Araújo, Silvio Jamil Ferzoli Guimarães
Multim. Tools Appl.4
2016 Summarizing video sequence using a graph-based hierarchical approach
Luciana dos Santos Belo, Carlos Antônio Caetano Jr., Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães
Neurocomputing4
2016 A mid-level video representation based on binary descriptors: A case study for pornography detection
Carlos Antônio Caetano Jr., Sandra Eliza Fontes de Avila, William Robson Schwartz, Silvio Jamil Ferzoli Guimarães, Arnaldo de Albuquerque Araújo
Neurocomputing4
2015 Hierarchical Combination of Semantic Visual Words for Image Classification and Clustering
Vinicius von Glehn De Filippo, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães
CIARP3
2015 Re-ranking of the Merging Order for Hierarchical Image Segmentation
Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães
CIARP2
2015 Kernel Combination Through Genetic Programming for Image Classification
Yuri H. Ribeiro, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães
CIARP3
2015 An efficient access method for multimodal video retrieval
Ricardo C. Sperandio, Zenilton Kleber Gonçalves do Patrocínio Jr., Hugo Bastos de Paula, Silvio Jamil Ferzoli Guimarães
Multim. Tools Appl.4
2014 Phenological Event Detection by Visual Rhythms Dissimilarity Analysis
abstract
Plant phenology has been exploited as an important research venue for assessing the impact of climate changes. One common approach for monitoring vegetation relies on the use of digital cameras. The employment of imaging techniques for phenological observation allows the extraction and analysis of visual characteristics based on color and texture information with the objective of determining plant life cycle changes, such as the beginning of the leaf flushing or the senescence period. This paper presents a novel approach for detecting phenological changes by analyzing image temporal series. Our method is based on the use of visual rhythm analysis and the adoption of a dissimilarity measure to detect visual changes in the time line. Experiments were conducted on a three-year data set composed of 3,538 vegetation images and 21 samples of 6 different species of interest. Results demonstrate that the proposed change detection approach is able to effectively identify phenological events.
Lilian Chaves Brandao dos Santos, Jurandy Almeida, Jefersson A. dos Santos, Silvio Jamil Ferzoli Guimarães, Arnaldo de Albuquerque Araújo, Bruna Alberton, Leonor Patricia C. Morellato, Ricardo da Silva Torres
eScience4
2014 Graph-Based Hierarchical Video Summarization Using Global Descriptors
abstract
Video summarization is a simplification of video content for compacting the video information. The video summarization problem can be transformed to a clustering problem, in which some frames are selected to saliently represent the video content. In this work, we use a hierarchical graph-based clustering method for computing a video summary. In fact, the proposed approach, called Summary, adopts a hierarchical clustering method to generate a weight map from the frame similarity graph in which the clusters (or connected components of the graph) can easily be inferred. Moreover, the use of this strategy allows to apply a similarity measure between clusters during graph partition, instead of considering only the similarity between isolated frames. Furthermore, a new evaluation measure that assesses the diversity of opinions of user summaries, called Covering, is also proposed. Experimental results provide quantitative and qualitative comparison between the new approach and other popular algorithms from the literature, showing that the new algorithm is robust and efficient. Concerning quality measures, Summary outperforms the compared methods regardless of the visual feature used in terms of F-measure.
Luciana dos Santos Belo, Carlos Antônio Caetano Jr., Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães
ICTAI4
2014 Graph-based hierarchical video segmentation based on a simple dissimilarity measure
Kleber Jacques Ferreira de Souza, Arnaldo de Albuquerque Araújo, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães
Pattern Recognit. Lett.4
2013 Searching for Near-Duplicate Video Sequences from a Scalable Sequence Aligner
abstract
Near-duplicate video sequence identification consists in identifying real positions of a specific video clip in a video stream stored in a database. To address this problem, we propose a new approach based on a scalable sequence aligner borrowed from proteomics. Sequence alignment is performed on symbolic representations of features extracted from the input videos, based on an algorithm originally applied to bio-informatics. Experimental results demonstrate that our method performance achieved 94% recall with 100% precision, with an average searching time of about 1 second.
Leonardo S. de Oliveira, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães, Guillaume Gravier
ISM3
2012 A two-step video subsequence identification based on bipartite graph matching
abstract
Subsequence identification consists in identifying real positions of a specific video clip in a video stream together with the operations that may be used to transform the former into a subsequence from the latter. In order to cope with this problem, we propose a two-step method. First, a clip filtering strategy based on the identification of dense segments is used, in order to decrease the number of video clip candidates. Then, for each dense segment, a graph matching approach is applied to identify video subsequences similar to the query video. Our main contribution is the use of a simple and efficient distance to solve subsequence identification problem along with the definition of a hit function that identifies precisely which operations were used in query transformation. Experimental results demonstrate good performance for our method (90% recall with 93% precision).
Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
SMC1
2011 A Simple Hierarchical Clustering Method for Improving Flame Pixel Classification
abstract
In this paper, we propose a new approach for color image simplification in order to improve flame pixel classification. The fire detection performance depends critically on the performance of the flame pixel classifier. Color image simplification is the process of simplifying an image in order to decrease the number of colors while preserving, as much as possible, shapes. In this work, a hierarchical clustering method in a given color space is used to map the original colors into a smaller set of representative ones, allowing the use of a simple heuristic rule for classifying the clusters related to candidate flame colors. Using reverse mapping, we identify possible flame colors in the image. Main contributions of our work are the application of a simple hierarchical clustering method to color simplification, that decreases the number of possible flame colors, and a filtering methodology to reduce the influence of outliers. Several color spaces and distance measures were used to evaluate the proposed method. Experimental results demonstrate that color simplification is essential to successfully employ heuristic classification of flame colors.
Kleber Jacques Ferreira de Souza, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr., Arnaldo de Albuquerque Araújo, Jean Cousty
ICTAI2
2011 Video text extraction based on image regularization and temporal analysis
abstract
Video text extraction is the process of identifying embedded text on video, which is usually on complex background. This paper proposes a new approach to cope with this problem considering image regularization and temporal information. The former helps us to decrease the number of gray values in order to simplify the image content, and the second one takes advantage of video text persistence in order to identify video segments ignoring text changes. According to our experiments, the proposed method presents better results than other. Moreover, we propose a post-processing step for improving the text results obtained by Otsu method.
Ângelo Magno de Jesus, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr.
ISM2
2010 A Static Video Summarization Method Based on Hierarchical Clustering
Silvio Jamil Ferzoli Guimarães, Willer Gomes
CIARP1
2010 An Unified Transition Detection Based on Bipartite Graph Matching Approach
Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães, Henrique Batista da Silva, Kleber Jacques Ferreira de Souza
CIARP2
2009 Gradual transition detection based on bipartite graph matching approach
abstract
This paper addresses gradual transition detection which is part of video segmentation problem, and consists in identifying the boundary between consecutive shots. In this work, we propose an approach to cope with gradual transition detection in which we define and use a new dissimilarity measure based on the size of the maximum cardinality matching calculated using a bipartite graph with respect to a specified window. The experiments have used a video dataset which presents a variety of different video genres with more than 500 gradual transitions and our method with a much simpler classification approach achieves more than 90% recall with almost 80% precision which is similar to the best results found.
Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr., Kleber Jacques Ferreira de Souza, Hugo Bastos de Paula
MMSP1
2008 An approach for video cut detection using bipartite graph matching as dissimilarity distance
abstract
The video segmentation problem consists in the identification of the boundary between consecutive shots. When two consecutive frames are similar, they are considered to be in the same shot. In this work, we use the maximum cardinality of the bipartite graph matching between two frames as the dissimilarity distance in order to identify the cut locations. Thus, if two frames are similar then the maximum cardinality is high. We present some experiments to show the high performance of this distance.
Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr., Hugo Bastos de Paula
ICPR1
2008 A Rotation and Translation Invariant Algorithm for Cut Detection Using Bipartite Graph Matching
abstract
Cut detection is part of the video segmentation problem, and consists in the identification of the boundary between consecutive shots. In this case, when two consecutive frames are similar, they are considered to be in the same shot. This work presents an approach to cut detection using a rotation and translation invariant algorithm based on the use of the maximum cardinality of a bipartite graph matching between two frames as the dissimilarity distance. Experimental results provides a comparison between the new approach and other popular algorithms from the literature, showing that the new algorithm is robust and has a high performance if compared to other methods of cut detection.
Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr., Hugo Bastos de Paula
ISM1
2006 Counting of Video Clip Repetitions using a Modified BMH Algorithm: Preliminary Results
abstract
In this work, we cope with the problem of identifying the number of repetitions of a specific video clip in a target video clip. Generally, the methods that deal with this problem can be subdivided into methods that use: (i) video signatures afterward the step of temporal video segmentation; and (ii) string matching algorithms afterward transformation of the video frame content into a feature values. Here, we propose a modification of the fastest exact string matching algorithm, the Boyer-Moore-Horspool, to count video clip repetitions. We also present some experiments to validate our approach, mainly if we are interested in found identical video clips according to spatial and temporal features
Silvio Jamil Ferzoli Guimarães, Renata Rodrigues Coelho, Anne Torres
ICME1
2003 An approach to detect video transitions based on mathematical morphology
abstract
The video segmentation problem can be regarded as a problem of detecting the fundamental video units (shots). Due to different ways of linking two consecutive shots this task turns out to be difficult. In this work, we propose a method to detect both cuts and gradual transitions by image segmentation tools instead of using dissimilarity measures or mathematical models. Firstly, the video is transformed into a 2D image, called visual rhythm by sub-sampling. Afterwards, we apply image processing tools to detect all vertical aligned transitions in this image. The main operator applied here is the morphological multiscale gradient. We also present some experimental results.
Silvio Jamil Ferzoli Guimarães, Michel Couprie, Arnaldo de Albuquerque Araújo, Neucimar J. Leite
ICIP (2)1
2003 Video segmentation based on 2D image analysis
Silvio Jamil Ferzoli Guimarães, Michel Couprie, Arnaldo de Albuquerque Araújo, Neucimar J. Leite
Pattern Recognit. Lett.1