EDBT 2026 Demo / reviewers in the wild / expert
Zenilton Kleber Gonçalves do Patrocínio Jr.
dblp:82/8677 · also Zenilton K. G. Patrocínio Jr., Zenilton K. G. do Patrocínio, Zenilton Kleber G. do Patrocínio Jr., Zenilton Patrocínio
· DBLP profile ↗
42ranked-venue papers
3as first author
18since 2021 · last 2026
0000-0003-0804-1790ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A comparative study of machine learning methods to predict the thermal and energy performance of a small-scale solar chimney
Matheus Augusto Ferreira Soares, Zenilton Kleber Gonçalves do Patrocínio Jr., Cristiana Brasil Maia |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Multilingual Visual Understanding: Extending Visual Dialog to Portuguese and Spanish Through Cross-Modal Adaptation
Milena Menezes Adão, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
CIARP | 3 |
| 2025 | Detecting Hierarchical Inconsistencies from an Image Segmentation Series
Fábio Kochem, Felipe Belém, Zenilton Kleber Gonçalves do Patrocínio Jr., Benjamin Perret, Jean Cousty, Alexandre Falcao, Silvio Jamil Ferzoli Guimarães |
CIARP (2) | 3 |
| 2025 | Lightweight Graph Neural Networks for 3D Shape Classification
Eduardo Felipe Lopes, João Pedro Oliveira Batisteli, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
CIARP (2) | 4 |
| 2025 | Extending Visual Dialog Beyond English: An Analysis of Monolingual and Multilingual ModelsabstractVisual Dialog is a challenging multimodal task requiring models to answer questions about images through multi-turn conversations. Despite significant progress, research has predominantly focused on English, limiting applicability to the 850+ million speakers of Portuguese and Spanish worldwide. We present the first comprehensive study of monolingual and multilingual Visual Dialog models for Portuguese and Spanish, introducing novel insights into cross-lingual visual grounding mechanisms. Through extensive experiments on newly translated VisDial datasets, we compare language-specific encoders (BERTimbau for Portuguese, BETO for Spanish) against multilingual BERT, achieving competitive performance with monolingual models while revealing distinct cross-modal attention patterns. Our mechanistic interpretability analysis demonstrates that despite different tokenization strategies and pretraining objectives, both approaches converge to similar attention distributions in deeper layers, with divergence decreasing from 0.000832 (Layer 0) to 0.000490 (Layer 11). We find that monolingual models exhibit holistic attention strategies while multilingual models show more selective, fine-grained visual grounding. These findings have important implications for developing inclusive vision-language technologies and understanding cross-lingual transfer in multimodal contexts. Milena M. Adão, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
ISM | 3 |
| 2025 | Preface to the special issue: SIBGRAPI 2024 tutorials
Soraia Raupp Musse, Ricardo Marroquim, Zenilton Kleber Gonçalves do Patrocínio Jr. |
Comput. Graph. | 3 |
| 2025 | Hierarchical Multiscale Representation in Remote Sensing Scene ClassificationabstractRemote sensing scene classification (RSSC) poses significant challenges due to high spatial variability, complex textures, and semantic ambiguity in remote sensing imagery. While convolutional neural networks (CNNs) and transformer-based models have achieved notable success in this domain, their performance often depends on large-scale pretraining and substantial computational resources. Graph neural networks (GNNs) have emerged as a promising alternative to traditional deep learning methods by explicitly modeling the relational structure of image regions through graph representations, which have already demonstrated promising results across various image-based tasks involving images. In this work, we explore two GNN architectures tailored for RSSC: BRMv2, a novel simplified graph model built on a base region adjacency graph (RAG), and modified hierarchical layered multigraph network (mHELMNet), a modified hierarchical multigraph model that encodes multiscale and spatial relationships through a multigraph representation. Both models were evaluated on the EUROSAT and RESISC45 datasets, achieving accuracy comparable to, or in some cases exceeding, that of state-of-the-art CNN-based, hybrid GNN-based, and transformer-based methods, while using significantly fewer parameters and without relying on pretraining. Experimental results demonstrated that the proposed GNN models, mHELMNet and BRMv2, achieved over 96% accuracy on EUROSAT and approximately 85% on RESISC45, while requiring only 0.14% and 0.03% of the parameters of the leading transformer-based approach, respectively. João Pedro Oliveira Batisteli, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Streamlined extended Long Short-Term Memory for video skimming
Leonardo Vilela Cardoso, Barbara Hellen P. Soraggi, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
Pattern Recognit. Lett. | 4 |
| 2024 | Seed-Based Superpixel Re-Segmentation for Improving Object Delineation
Lucca S. P. Lacerda, Felipe Belém, Zenilton Kleber Gonçalves do Patrocínio Jr., Alexandre X. Falcão, Silvio Jamil Ferzoli Guimarães |
CIARP (1) | 3 |
| 2024 | Towards Interactive Video Segmentation by Dynamic and Iterative Spanning Forest
Danielle Vieira, Isabela Borlido Barcelos, Zenilton Kleber Gonçalves do Patrocínio Jr., Alexandre X. Falcão, Silvio Jamil Ferzoli Guimarães |
CIARP (1) | 3 |
| 2023 | Graph-Based Feature Learning from Image Markers
Isabela Borlido Barcelos, Leonardo de Melo Joao, Zenilton Kleber Gonçalves do Patrocínio Jr., Ewa Kijak, Alexandre X. Falcão, Silvio Jamil Ferzoli Guimarães |
CIARP | 3 |
| 2023 | Filtering Safe Temporal Motifs in Dynamic Graphs for Dissemination Purposes
Carolina Stephanie Jerônimo de Almeida, Simon Malinowski, Zenilton Kleber Gonçalves do Patrocínio Jr., Guillaume Gravier, Silvio Jamil Ferzoli Guimarães |
CIARP | 3 |
| 2023 | Streaming Graph-Based Supervoxel Computation Based on Dynamic Iterative Spanning Forest
Danielle Vieira, Isabela Borlido Barcelos, Felipe Belém, Zenilton Kleber Gonçalves do Patrocínio Jr., Alexandre X. Falcão, Silvio Jamil Ferzoli Guimarães |
CIARP | 4 |
| 2023 | A Novel Method for Temporal Graph Classification based on Transitive ReductionabstractDomains such as bio-informatics, social network analysis, and computer vision, describe relations between entities and cannot be interpreted as vectors or fixed grids, instead, they are naturally represented by graphs. Often this kind of data evolves over time in a dynamic world, respecting a temporal order being known as temporal graphs. The latter became a challenge since subgraph patterns are very difficult to find and the distance between those patterns may change irregularly over time. While state-of-the-art methods are primarily designed for static graphs and may not capture temporal information, recent works have proposed mapping temporal graphs to static graphs to allow for the use of conventional static kernels and graph neural approaches. In this study, we compare the transitive reduction impact on these mappings in terms of accuracy and computational efficiency across different classification tasks. Furthermore, we introduce a novel mapping method using a transitive reduction approach that outperforms existing techniques in terms of classification accuracy. Our experimental results demonstrate the effectiveness of the proposed mapping method in improving the accuracy of supervised classification for temporal graphs while maintaining reasonable computational efficiency. Carolina Stephanie Jerônimo de Almeida, Zenilton Kleber Gonçalves do Patrocínio Jr., Simon Malinowski, Silvio Jamil Ferzoli Guimarães, Guillaume Gravier |
DSAA | 2 |
| 2023 | Multi-Scale Image Graph Representation: A Novel GNN Approach for Image Classification through Scale Importance EstimationabstractImage representation as graphs can enhance the understanding of image semantics and facilitate multi-scale image representation. However, existing methods often overlook the significance of the relationship between elements at each scale or fail to encode the hierarchical relationship between graph elements. To cope with that, we introduce a novel approach for graph construction from images. This approach utilizes a hierarchical image segmentation technique to generate segmentation at multiple scales and incorporates edges to encode the relationships at each scale. We also propose a new readout function that weighs the importance of each scale when deriving a fixed-size graph representation. Furthermore, we present a new model incorporating those ideas – called Hierarchical Image Graph with Scale Importance (HIGSI). Experimental results on the CIFAR-10 database indicate that our proposed model outperforms (or closely matches) state-of-the-art and baseline models while utilizing smaller graphs. João Pedro Oliveira Batisteli, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
ISM | 3 |
| 2023 | Graph-based image gradients aggregated with random forests
Raquel Almeida 0001, Ewa Kijak, Simon Malinowski, Zenilton Kleber Gonçalves do Patrocínio Jr., Arnaldo de Albuquerque Araújo, Silvio Jamil Ferzoli Guimarães |
Pattern Recognit. Lett. | 4 |
| 2021 | Enhanced-Memory Transformer for Coherent Paragraph Video CaptioningabstractA coherent description is an ultimate goal concerning video captioning through multiple sentences since it may directly impact consistency and intelligibility. A paragraph describing a video is affected by different events extracted from it. When generating a new event, it should produce a detailed narrative of the video content. But it also might provide some clues that may help reduce the textual repetition occurring in the final description. Recently, transformers have emerged as an appealing solution to several tasks, including video captioning. An augmented transformer with a memory module can somehow cope with text repetition. Thus, to further increase the coherence among the generated sentences, we propose the adoption of attention mechanisms to enhance memory data in a memory-augmented transformer. This new approach, called Enhanced-Memory Transformer (EMT), assesses the data importance (about the video segments) contained in the memory module and uses that to improve readability by reducing repetition. Experimental evaluation of EMT using the test split of the ActivityNet Captions dataset achieved 22.84 in CIDEr-D score, and 4.55 in Reduction-4 score (R@4), representing improvements of 1.03% and 16.36%, respectively (compared to the literature). The obtained results show the great potential of this new approach as it provides increased coherence among the various video segments, decreasing the repetition in the generated sentences and improving the description diversity. Leonardo Vilela Cardoso, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
ICTAI | 3 |
| 2021 | Hierarchical multi-label propagation using speaking face graphs for multimodal person discovery
Gabriel Barbosa da Fonseca, Gabriel Sargent, Ronan Sicre, Zenilton Kleber Gonçalves do Patrocínio Jr., Guillaume Gravier, Silvio Jamil Ferzoli Guimarães |
Multim. Tools Appl. | 4 |
| 2020 | Learning to realign hierarchy for image segmentation
Milena M. Adão, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
Pattern Recognit. Lett. | 3 |
| 2019 | BRIEF-Based Mid-Level Representations for Time Series Classification
Renato Augusto de Souza, Raquel Almeida 0001, Roberto Miranda, Zenilton Kleber Gonçalves do Patrocínio Jr., Simon Malinowski, Silvio Jamil Ferzoli Guimarães |
CIARP | 4 |
| 2019 | Combining convolutional side-outputs for road image segmentationabstractImage segmentation consists in creating partitions within an image into meaningful areas and objects. It can be used in scene understanding and recognition, in fields like biology, medicine, robotics, satellite imaging, amongst others. In this work we take advantage of the learned model in a deep architecture, by extracting side-outputs at different layers of the network for the task of image segmentation. We study the impact of the amount of side-outputs and evaluate strategies to combine them. A post-processing filtering based on mathematical morphology idempotent functions is also used in order to remove some undesirable noises. Experiments were performed on the publicly available KITTI Road Dataset for image segmentation. Our comparison shows that the use of multiples side outputs can increase the overall performance of the network, making it easier to train and more stable when compared with a single output in the end of the network. Also, for a small number of training epochs (500), we achieved a competitive performance when compared to the best algorithm in KITTI Evaluation Server. Felipe A. L. Reis, Raquel Almeida 0001, Ewa Kijak, Simon Malinowski, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
IJCNN | 6 |
| 2018 | Evaluation of Scale-Aware Realignments of Hierarchical Image Segmentation
Milena M. Adão, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
CIARP | 3 |
| 2018 | Evaluation of Bag-of-Word Performance for Time Series Classification Using Discriminative SIFT-Based Mid-Level Representations
Raquel Almeida 0001, Hugo Herlanin, Zenilton Kleber Gonçalves do Patrocínio Jr., Simon Malinowski, Silvio Jamil Ferzoli Guimarães |
CIARP | 3 |
| 2018 | Evaluating AdaBoost for Plagiarism Detection
Thiago V. Reginaldo, Magali R. G. Meireles, Zenilton Kleber Gonçalves do Patrocínio Jr. |
CIARP | 3 |
| 2018 | Hierarchical Graph-Based Segmentation in Detection of Object-Related Regions
Rafael Machado Ribeiro, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
CIARP | 3 |
| 2017 | Exploring quantization error to improve human action classificationabstractIn human action classification task, a video must be classified into a pre-determined class. To cope with this problem, we propose a mid-level representation, in which information about quantization errors is embedded together with the aggregated data on low level features. The main contributions of this article are twofold: (i) assembly of low-level features (dense trajectories) by a mid-level representation enriched with information about distances between descriptors and codewords; and (ii) a survey of the most common protocols for human action classification methods when applied to three different datasets. Regarding classification protocols, we have experimented the training and testing classification (called split), the leave-one-out cross-validation (LOOCV) and the leave-one-group-out cross-validation (25-fold CV). Experimental results demonstrated that our strategy either has improved the classification rates with respect to the state-of-the-art for KTH dataset, achieving 98%, or it is a competitive one, for UCF-11 with 90%, when compared with methods with no feature learning. Raquel Almeida 0001, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães |
IJCNN | 2 |
| 2016 | Summarizing video sequence using a graph-based hierarchical approach
Luciana dos Santos Belo, Carlos Antônio Caetano Jr., Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães |
Neurocomputing | 3 |
| 2015 | Hierarchical Combination of Semantic Visual Words for Image Classification and Clustering
Vinicius von Glehn De Filippo, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães |
CIARP | 2 |
| 2015 | Re-ranking of the Merging Order for Hierarchical Image Segmentation
Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães |
CIARP | 1 |
| 2015 | Kernel Combination Through Genetic Programming for Image Classification
Yuri H. Ribeiro, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães |
CIARP | 2 |
| 2015 | An efficient access method for multimodal video retrieval
Ricardo C. Sperandio, Zenilton Kleber Gonçalves do Patrocínio Jr., Hugo Bastos de Paula, Silvio Jamil Ferzoli Guimarães |
Multim. Tools Appl. | 2 |
| 2014 | Graph-Based Hierarchical Video Summarization Using Global DescriptorsabstractVideo summarization is a simplification of video content for compacting the video information. The video summarization problem can be transformed to a clustering problem, in which some frames are selected to saliently represent the video content. In this work, we use a hierarchical graph-based clustering method for computing a video summary. In fact, the proposed approach, called Summary, adopts a hierarchical clustering method to generate a weight map from the frame similarity graph in which the clusters (or connected components of the graph) can easily be inferred. Moreover, the use of this strategy allows to apply a similarity measure between clusters during graph partition, instead of considering only the similarity between isolated frames. Furthermore, a new evaluation measure that assesses the diversity of opinions of user summaries, called Covering, is also proposed. Experimental results provide quantitative and qualitative comparison between the new approach and other popular algorithms from the literature, showing that the new algorithm is robust and efficient. Concerning quality measures, Summary outperforms the compared methods regardless of the visual feature used in terms of F-measure. Luciana dos Santos Belo, Carlos Antônio Caetano Jr., Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães |
ICTAI | 3 |
| 2014 | Graph-based hierarchical video segmentation based on a simple dissimilarity measure
Kleber Jacques Ferreira de Souza, Arnaldo de Albuquerque Araújo, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães |
Pattern Recognit. Lett. | 3 |
| 2013 | Searching for Near-Duplicate Video Sequences from a Scalable Sequence AlignerabstractNear-duplicate video sequence identification consists in identifying real positions of a specific video clip in a video stream stored in a database. To address this problem, we propose a new approach based on a scalable sequence aligner borrowed from proteomics. Sequence alignment is performed on symbolic representations of features extracted from the input videos, based on an algorithm originally applied to bio-informatics. Experimental results demonstrate that our method performance achieved 94% recall with 100% precision, with an average searching time of about 1 second. Leonardo S. de Oliveira, Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães, Guillaume Gravier |
ISM | 2 |
| 2012 | A two-step video subsequence identification based on bipartite graph matchingabstractSubsequence identification consists in identifying real positions of a specific video clip in a video stream together with the operations that may be used to transform the former into a subsequence from the latter. In order to cope with this problem, we propose a two-step method. First, a clip filtering strategy based on the identification of dense segments is used, in order to decrease the number of video clip candidates. Then, for each dense segment, a graph matching approach is applied to identify video subsequences similar to the query video. Our main contribution is the use of a simple and efficient distance to solve subsequence identification problem along with the definition of a hit function that identifies precisely which operations were used in query transformation. Experimental results demonstrate good performance for our method (90% recall with 93% precision). Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
SMC | 2 |
| 2011 | A Simple Hierarchical Clustering Method for Improving Flame Pixel ClassificationabstractIn this paper, we propose a new approach for color image simplification in order to improve flame pixel classification. The fire detection performance depends critically on the performance of the flame pixel classifier. Color image simplification is the process of simplifying an image in order to decrease the number of colors while preserving, as much as possible, shapes. In this work, a hierarchical clustering method in a given color space is used to map the original colors into a smaller set of representative ones, allowing the use of a simple heuristic rule for classifying the clusters related to candidate flame colors. Using reverse mapping, we identify possible flame colors in the image. Main contributions of our work are the application of a simple hierarchical clustering method to color simplification, that decreases the number of possible flame colors, and a filtering methodology to reduce the influence of outliers. Several color spaces and distance measures were used to evaluate the proposed method. Experimental results demonstrate that color simplification is essential to successfully employ heuristic classification of flame colors. Kleber Jacques Ferreira de Souza, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr., Arnaldo de Albuquerque Araújo, Jean Cousty |
ICTAI | 3 |
| 2011 | Video text extraction based on image regularization and temporal analysisabstractVideo text extraction is the process of identifying embedded text on video, which is usually on complex background. This paper proposes a new approach to cope with this problem considering image regularization and temporal information. The former helps us to decrease the number of gray values in order to simplify the image content, and the second one takes advantage of video text persistence in order to identify video segments ignoring text changes. According to our experiments, the proposed method presents better results than other. Moreover, we propose a post-processing step for improving the text results obtained by Otsu method. Ângelo Magno de Jesus, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
ISM | 3 |
| 2010 | An Unified Transition Detection Based on Bipartite Graph Matching Approach
Zenilton Kleber Gonçalves do Patrocínio Jr., Silvio Jamil Ferzoli Guimarães, Henrique Batista da Silva, Kleber Jacques Ferreira de Souza |
CIARP | 1 |
| 2009 | Gradual transition detection based on bipartite graph matching approachabstractThis paper addresses gradual transition detection which is part of video segmentation problem, and consists in identifying the boundary between consecutive shots. In this work, we propose an approach to cope with gradual transition detection in which we define and use a new dissimilarity measure based on the size of the maximum cardinality matching calculated using a bipartite graph with respect to a specified window. The experiments have used a video dataset which presents a variety of different video genres with more than 500 gradual transitions and our method with a much simpler classification approach achieves more than 90% recall with almost 80% precision which is similar to the best results found. Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr., Kleber Jacques Ferreira de Souza, Hugo Bastos de Paula |
MMSP | 2 |
| 2008 | An approach for video cut detection using bipartite graph matching as dissimilarity distanceabstractThe video segmentation problem consists in the identification of the boundary between consecutive shots. When two consecutive frames are similar, they are considered to be in the same shot. In this work, we use the maximum cardinality of the bipartite graph matching between two frames as the dissimilarity distance in order to identify the cut locations. Thus, if two frames are similar then the maximum cardinality is high. We present some experiments to show the high performance of this distance. Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr., Hugo Bastos de Paula |
ICPR | 2 |
| 2008 | A Rotation and Translation Invariant Algorithm for Cut Detection Using Bipartite Graph MatchingabstractCut detection is part of the video segmentation problem, and consists in the identification of the boundary between consecutive shots. In this case, when two consecutive frames are similar, they are considered to be in the same shot. This work presents an approach to cut detection using a rotation and translation invariant algorithm based on the use of the maximum cardinality of a bipartite graph matching between two frames as the dissimilarity distance. Experimental results provides a comparison between the new approach and other popular algorithms from the literature, showing that the new algorithm is robust and has a high performance if compared to other methods of cut detection. Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr., Hugo Bastos de Paula |
ISM | 2 |
| 2003 | A Lagrangian-based heuristic for traffic grooming in WDM optical networksabstractTraffic grooming problem (TGP) deals with efficiently combining low-speed traffic streams into high-capacity wavelength channels in order to improve bandwidth utilization and minimize network cost. In this paper, we investigate TGP in WDM optical networks regardless of underlying physical topology. The problem is formulated as an integer linear program (ILP) and a Lagrangian-based heuristic is proposed. Numerical results for ring and mesh networks are presented and analyzed. Zenilton Kleber Gonçalves do Patrocínio Jr., Geraldo Robson Mateus |
GLOBECOM | 1 |