Rodrigo Minetto

dblp:00/1588 · DBLP profile ↗
← Back
25ranked-venue papers
10as first author
8since 2021 · last 2025
0000-0003-2277-4632ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 60% Geometric modeling and processing · 40%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval › affective computing › sentiment analysis
visual sentiment analysis
0.412020
OutdoorSent: Sentiment Analysis of Urban Outdoor Images by Using Semantic and Deep Features · ACM Trans. Inf. Syst. 2020
Geometric modeling and processing
mesh processing
0.312017
An optimal algorithm for 3D triangle mesh slicing · Comput. Aided Des. 2017

Methods — techniques the papers use, named apart from their topics

transfer learning · 0.4convolutional neural network · 0.4optimal algorithm · 0.3
YearPublicationVenuePosition
2025 Detection of Pavement Defects on Roads using a Multimodal YOLOv8 with Image and IMU Data
abstract
Monitoring and maintaining the conditions of roads is a challenging and essential task. Object detection models can accelerate the assessment of roads, making it more efficient and standardized. Recent studies have explored automating road assessment using images and inertial sensors, such as accelerometers. This work hypothesizes that a multimodal model, combining image and sensor data, improves the performance of pavement defect detection. Experimental results show that the proposed YOLOv8-MultiDetect (multimodal) model outperforms the YOLOv8-Detect (unimodal) model, with average [email protected] scores of 0.809 versus 0.759 and [email protected]:0.65 scores of 0.780 versus 0.726. These findings suggest that integrating features from multiple data sources enhances both object classification and localization in detecting road defects.
Arthur Polli, Ricardo Dutra da Silva, Rodrigo Minetto
ICIP3
2025 TSlicer: An optimal topology-based slicing algorithm for Z-monotone 3D meshes
Ricardo Dutra da Silva, Henrique Romaniuk Ramalho, Rodrigo Minetto, Neri Volpato, Jorge Stolfi
Comput. Graph.3
2025 DED-IM: A Novel Method for Mapping and Path Planning in Wire Arc Directed Energy Deposition
abstract
This paper introduces DED-IM (https://github.com/machipanski/DED-IM), a novel image-based mapping and tool-path planning method for Wire Arc Directed Energy Deposition (WA-DED). The method automatically segments 3D models into distinct regions such as thin walls, contours, bottlenecks, and large internal areas by processing binary images. These segments enable precise, geometry-tailored tool-path generation, significantly improving control over material deposition and reducing arc interruptions. Key innovations include the application of medial axis transforms to accurately map thin walls, the introduction of oscillatory paths for bottleneck areas, and an adaptive weaving pattern that minimizes voids while enhancing geometric precision. The method also integrates user-defined parameters, offering flexibility in region mapping and filling strategies to address specific manufacturing requirements. Experimental results demonstrate the efficacy of DED-IM, showing substantial reductions in common defects such as voids, insufficient material filling, and arc interruptions. This project, provides a flexible and scalable foundation baseline for automated WA-DED processes, enabling the integration of diverse path-planning techniques, fostering innovation and customization in tool-path generation for different manufacturing requirements.Note to Practitioners—This paper introduces DED-IM, a novel path-planning method for Wire Arc Directed Energy Deposition (WA-DED) processes. Developed in Python, the method takes a 3D model as input, along with machine-specific parameters, and generates G-code instructions that can be customized for different WA-DED machines. DED-IM automatically segments 3D models into regions like thin walls, contours, bottlenecks, and large internal areas by processing binary images. These parameters are fully adjustable, allowing users to tailor the tool-path strategy to their specific needs. The modular design also makes it easy to integrate new strategies, such as custom zigzag patterns. To use DED-IM in a production workflow, practitioners will need to adjust the software according to their machine specifications.
Matheus Antunes Chipanski, Tadeu Castro da Silva, Neri Volpato, Ricardo Dutra da Silva, Valdemar Rebelo R. Duarte, Telmo G. Santos, Rodrigo Minetto
IEEE Trans Autom. Sci. Eng.7
2024 Sentinel-2 Active Fire Segmentation: Analyzing Convolutional and Transformer Architectures, Knowledge Transfer, Fine-Tuning, and Seam Lines
abstract
Active fire segmentation in satellite imagery is a critical remote sensing task, providing essential support for planning, decision-making, and policy development. Several techniques have been proposed for this problem over the years, generally based on specific equations and thresholds, which are sometimes empirically chosen. Some satellites, such as MODIS and Landsat-8, have consolidated algorithms for this task. However, for other important satellites such as Sentinel-2, this is still an open problem. In this letter, we explore the possibility of using transfer learning to train convolutional and transformer-based deep architectures (U-Net, DeepLabV3+, and SegFormer) for active fire segmentation. We pretrain these architectures based on Landsat-8 images and automatically labeled samples and fine-tune them to Sentinel-2 images. The experiments show that the proposed method achieves$F1$-scores of up to 88.4% for Sentinel-2 images, outperforming three threshold-based algorithms by at least 19% while maintaining a low demand for manually labeled samples. We also address detection over seam-line regions that present a particular challenge for existing methods. The source code and trained models are available athttps://github.com/Minoro/l8tos2-transf-seamlines.
André Minoro Fusioka, Gabriel Henrique de Almeida Pereira, Bogdan Tomoyuki Nassu, Rodrigo Minetto
IEEE Geosci. Remote. Sens. Lett.4
2023 Leveraging Model Fusion for Improved License Plate Recognition
Rayson Laroca, Luiz Antonio Zanlorensi, Valter Estevam, Rodrigo Minetto, David Menotti
CIARP4
2023 Do We Train on Test Data? The Impact of Near-Duplicates on License Plate Recognition
abstract
This work draws attention to the large fraction of near-duplicates in the training and test sets of datasets widely adopted in License Plate Recognition (LPR) research. These duplicates refer to images that, although different, show the same license plate. Our experiments, conducted on the two most popular datasets in the field, show a substantial decrease in recognition rate when six well-known models are trained and tested under fair splits, that is, in the absence of duplicates in the training and test sets. Moreover, in one of the datasets, the ranking of models changed considerably when they were trained and tested under duplicate-free splits. These findings suggest that such duplicates have significantly biased the evaluation and development of deep learning-based models for LPR. The list of near-duplicates we have found and proposals for fair splits are publicly available for further research at https://raysonlaroca.github.io/supp/lpr-train-on-test/.
Rayson Laroca, Valter Estevam, Alceu S. Britto Jr., Rodrigo Minetto, David Menotti
IJCNN4
2023 PerceptSent - Exploring Subjectivity in a Novel Dataset for Visual Sentiment Analysis
abstract
Visual sentiment analysis is a challenging problem. Many datasets and approaches have been designed to foster breakthroughs in this trending research topic. However, most works scrutinize only subsymbolic models through visual attributes of the evaluated images, paying less attention to the subjectivity of viewers’ perceptions as a basis for neuro-symbolic systems. Aiming to fill this gap, we present PerceptSent, a novel dataset for visual sentiment analysis that spans 5,000 images shared by users on social networks. Besides the sentiment opinion (positive, slightly positive, neutral, slightly negative, negative) expressed by every evaluator about each image analyzed, the dataset contains evaluator's metadata (age, gender, socioeconomic status, education, and psychological hints) as well as perceptions observed by the evaluator about the image — such as the presence of nature, violence, lack of maintenance, etc. Deep architectures and different problem formulations are explored using our dataset to combine visual and extra attributes (external knowledge) for automatic sentiment analysis. We show evidence that evaluator's perceptionss, when correctly employed, are crucial in visual sentiment analysis, improving the F-score performance from 61% to an impressive rate above 97%. Although, at this point, we do not have automatic approaches to capture these perceptions, our results open up new investigation avenues.
Cesar Rafael Lopes, Rodrigo Minetto, Myriam Delgado, Thiago H. Silva 0001
IEEE Trans. Affect. Comput.2
2021 Measuring Human and Economic Activity From Satellite Imagery to Support City-Scale Decision-Making During COVID-19 Pandemic
abstract
The COVID-19 outbreak forced governments worldwide to impose lockdowns and quarantines to prevent virus transmission. As a consequence, there are disruptions in human and economic activities all over the globe. The recovery process is also expected to be rough. Economic activities impact social behaviors, which leave signatures in satellite images that can be automatically detected and classified. Satellite imagery can support the decision-making of analysts and policymakers by providing a different kind of visibility into the unfolding economic changes. In this article, we use a deep learning approach that combines strategic location sampling and an ensemble of lightweight convolutional neural networks (CNNs) to recognize specific elements in satellite images that could be used to compute economic indicators based on it, automatically. This CNN ensemble framework ranked third place in the US Department of Defense xView challenge, the most advanced benchmark for object detection in satellite images. We show the potential of our framework for temporal analysis using the US IARPA Function Map of the World (fMoW) dataset. We also show results on real examples of different sites before and after the COVID-19 outbreak to illustrate different measurable indicators. Our code and annotated high-resolution aerial scenes before and after the outbreak are available on GitHub.1.https://github.com/maups/covid19-satellite-analysis.
Rodrigo Minetto, Maurício Pamplona Segundo, Gilbert Rotich, Sudeep Sarkar
IEEE Trans. Big Data1
2020 OutdoorSent: Sentiment Analysis of Urban Outdoor Images by Using Semantic and Deep Features
abstract
Opinion mining in outdoor images posted by users during different activities can provide valuable information to better understand urban areas. In this regard, we propose a framework to classify the sentiment of outdoor images shared by users on social networks. We compare the performance of state-of-the-art ConvNet architectures and one specifically designed for sentiment analysis. We also evaluate how the merging of deep features and semantic information derived from the scene attributes can improve classification and cross-dataset generalization performance. The evaluation explores a novel dataset—namely, OutdoorSent—and other publicly available datasets. We observe that the incorporation of knowledge about semantic attributes improves the accuracy of all ConvNet architectures studied. Besides, we found that exploring only images related to the context of the study—outdoor, in our case—is recommended, i.e., indoor images were not significantly helpful. Furthermore, we demonstrated the applicability of our results in the United States city of Chicago, Illinois, showing that they can help to improve the knowledge of subjective characteristics of different areas of the city. For instance, particular areas of the city tend to concentrate more images of a specific class of sentiment, which are also correlated with median income, opening up opportunities in different fields.
Wyverson Bonasoli de Oliveira, Leyza Baldo Dorini, Rodrigo Minetto, Thiago H. Silva 0001
ACM Trans. Inf. Syst.3
2019 A Two-Stream Siamese Neural Network for Vehicle Re-Identification by Using Non-Overlapping Cameras
abstract
We describe in this paper a Two-Stream Siamese Neural Network for vehicle re-identification. The proposed network is fed simultaneously with small coarse patches of the vehicle shape's, with 96 × 96 pixels, in one stream, and fine features extracted from license plate patches, easily readable by humans, with 96 × 48 pixels, in the other one. Then, we combined the strengths of both streams by merging the Siamese distance descriptors with a sequence of fully connected layers, as an attempt to tackle a major problem in the field, false alarms caused by a huge number of car design and models with nearly the same appearance or by similar license plate strings. In our experiments, with 2 hours of videos containing 2982 vehicles, extracted from two low-cost cameras in the same roadway, 546 ft away, we achieved a F-measure and accuracy of 92.6% and 98.7%, respectively. We show that our network, available at https://github.com/icarofua/siamese-two-stream, outperforms other One-Stream architectures, even if they use higher resolution image features.
Icaro O. de Oliveira, Keiko Verônica Ono Fonseca, Rodrigo Minetto
ICIP3
2019 Hydra: An Ensemble of Convolutional Neural Networks for Geospatial Land Classification
abstract
In this paper, we describe Hydra, an ensemble of convolutional neural networks (CNNs) for geospatial land classification. The idea behind Hydra is to create an initial CNN that is coarsely optimized but provides a good starting pointing for further optimization, which will serve as the Hydra's body. Then, the obtained weights are fine-tuned multiple times with different augmentation techniques, crop styles, and classes weights to form an ensemble of CNNs that represent the Hydra's heads. By doing so, we prompt convergence to different endpoints, which is a desirable aspect for ensembles. With this framework, we were able to reduce the training time while maintaining the classification performance of the ensemble. We created ensembles for our experiments using two state-of-the-art CNN architectures, residual network (ResNet), and dense convolutional networks (DenseNet). We have demonstrated the application of our Hydra framework in two data sets, functional map of world (FMOW) and NWPU-RESISC45, achieving results comparable to the state-of-the-art for the former and the best-reported performance so far for the latter. Code and CNN models are available at https://github.com/maups/hydra-fmow.
Rodrigo Minetto, Maurício Pamplona Segundo, Sudeep Sarkar
IEEE Trans. Geosci. Remote. Sens.1
2018 Corrigendum to "An optimal algorithm for 3D triangle mesh slicing" [Comput. Aided Des. 97 (2017)]
Rodrigo Minetto, Jorge Stolfi
Comput. Aided Des.1
2017 Convolutional neural networks for license plate detection in images
abstract
License plate detection is a challenging task when dealing with open environments and images captured from a certain distance by low-cost cameras. In this paper, we propose an approach for detecting license plates based on a convolutional neural network which models a function that produces a score for each image sub-region, allowing us to estimate the locations of the detected license plates by combining the results obtained from sparse overlapping regions. Experiments were performed on a challenging benchmark, containing 4,070 license plates in 1,829 images, captured under several weather conditions. The proposed approach achieved a precision of 0.87 and recall of 0.83, outperforming a state-of-the-art detector - a promising result, given that the experiments were performed on single images, without any kind of preprocessing or temporal integration.
Francisco Delmar Kurpiel, Rodrigo Minetto, Bogdan Tomoyuki Nassu
ICIP2
2017 An optimal algorithm for 3D triangle mesh slicing
Rodrigo Minetto, Neri Volpato, Jorge Stolfi, Rodrigo M. M. H. Gregori, Murilo V. G. da Silva
Comput. Aided Des.1
2017 A Video-Based System for Vehicle Speed Measurement in Urban Roadways
abstract
In this paper, we propose a nonintrusive video-based system for vehicle speed measurement in urban roadways. Our system uses an optimized motion detector and a novel text detector to efficiently locate vehicle license plates in image regions containing motion. Distinctive features are then selected on the license plate regions, tracked across multiple frames, and rectified for perspective distortion. Vehicle speed is measured by comparing the trajectories of the tracked features to known real-world measures. The proposed system was tested on a data set containing approximately 5 h of videos recorded in different weather conditions by a single low-cost camera, with associated ground truth speeds obtained by an inductive loop detector. Our data set is freely available for research purposes. The measured speeds have an average error of -0.5 km/h, staying inside the [-3, +2] km/h limit determined by regulatory authorities in several countries in over 96.0% of the cases. To the authors' knowledge, there are no other video-based systems able to achieve results comparable to those produced by an inductive loop detector. We also show that our license plate detector outperforms two other published state-of-the-art text detectors, as well as a well-known license plate detector, achieving a precision of 0.93 and a recall of 0.87.
Diogo C. Luvizon, Bogdan Tomoyuki Nassu, Rodrigo Minetto
IEEE Trans. Intell. Transp. Syst.3
2014 Vehicle speed estimation by license plate detection and tracking
abstract
We describe a novel system for vehicle speed estimation from videos captured in urban roadways. Our system uses text detection to locate the license plates of passing vehicles, which are then used to select stable features for tracking. The tracked features are then filtered and rectified for perspective distortion. Vehicle speed is estimated by comparing the trajectory of the tracked features to known real world measures. In experiments performed on videos captured under real operation conditions, our system attained a precision of 0.87 and a recall of 0.92 for license plate detection. Vehicle speeds were estimated with an average error of 0.59 km/h, staying inside the +2/-3 km/h limit, determined by regulatory authorities in several countries, in over 75% of the cases.
Diogo C. Luvizon, Bogdan Tomoyuki Nassu, Rodrigo Minetto
ICASSP3
2014 SnooperText: A text detection system for automatic indexing of urban scenes
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J. Leite, Jorge Stolfi
Comput. Vis. Image Underst.1
2013 Text Line Detection in Document Images: Towards a Support System for the Blind
abstract
We introduce a novel approach for text line detection in document images, keeping in mind the requirements of a portable text recognition system designed to support the blind. Challenges include shadows, cluttered backgrounds, and perspective distortion. Different from previous approaches, the proposed method does not segment the image. A text model is created by clustering SIFT features extracted from positive and negative examples. Text regions are located by matching the features extracted from the input image to the clusters in the text model. Regions around the correspondences are then analyzed, and text lines are identified based on features such as gradients and histogram distribution. Experimental results show that our approach outperforms a state-of-the-art text detector in a text/non-text classification task.
Bogdan Tomoyuki Nassu, Rodrigo Minetto, Luiz Eduardo Soares de Oliveira
ICDAR2
2013 Adaptive edge-preserving image denoising using wavelet transforms
Ricardo Dutra da Silva, Rodrigo Minetto, William Robson Schwartz, Hélio Pedrini
Pattern Anal. Appl.2
2013 T-HOG: An effective gradient-based descriptor for single line text regions
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J. Leite, Jorge Stolfi
Pattern Recognit.1
2013 Semiautomatic White Blood Cell Segmentation Based on Multiscale Analysis
abstract
This paper approaches novel methods to segment the nucleus and cytoplasm of white blood cells (WBC). This information is the basis to perform higher level tasks such as automatic differential counting, which plays an important role in the diagnosis of different diseases. We explore the image simplification and contour regularization resulting from the application of the Self-Dual Multiscale Morphological Toggle (SMMT), an operator with scale-space properties. To segment the nucleus, the image preprocessing with SMMT has shown to be essential to ensure the accuracy of two well-known image segmentations techniques, namely, watershed transform and Level Set methods. To identify the cytoplasm region, we propose two different schemes, based on granulometric analysis and on morphological transformations. The proposed methods have been successfully applied to a large number of images, showing promising segmentation and classification results for varying cell appearance and image quality, encouraging future works.
Leyza Baldo Dorini, Rodrigo Minetto, Neucimar J. Leite
IEEE J. Biomed. Health Informatics2
2012 IFTrace: Video segmentation of deformable objects using the Image Foresting Transform
Rodrigo Minetto, Thiago Vallin Spina, Alexandre X. Falcão, Neucimar J. Leite, João Paulo Papa, Jorge Stolfi
Comput. Vis. Image Underst.1
2011 Snoopertrack: Text detection and tracking for outdoor videos
abstract
In this work we introduced SnooperTrack, an algorithm for the automatic detection and tracking of text objects - such as store names, traffic signs, license plates, and advertisements - in videos of out door scenes. The purpose is to improve the performances of text detection process in still images by taking advantage of the temporal coherence in videos. We first propose an efficient tracking algorithm using particle filtering framework with original region descriptors. The second contribution is our strategy to merge tracked regions and new detections. We also propose an improved version of our previously published text detection algorithm in still images. Tests indicate that SnooperTrack is fast, robust, enable false positive suppression, and achieved great performances in complex videos of outdoor scenes.
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J. Leite, Jorge Stolfi
ICIP1
2010 Snoopertext: A multiresolution system for text detection in complex visual scenes
abstract
Text detection in natural images remains a very challenging task. For instance, in an urban context, the detection is very difficult due to large variations in terms of shape, size, color, orientation, and the image may be blurred or have irregular illumination, etc. In this paper, we describe a robust and accurate multiresolution approach to detect and classify text regions in such scenarios. Based on generation/validation paradigm, we first segment images to detect character regions with a multiresolution algorithm able to manage large character size variations. The segmented regions are then filtered out using shape-based classification, and neighboring characters are merged to generate text hypotheses. A validation step computes a region signature based on texture analysis to reject false positives. We evaluate our algorithm in two challenging databases, achieving very good results.
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Jonathan Fabrizio, Beatriz Marcotegui
ICIP1
2009 AffTrack: Robust tracking of features in variable-zoom videos
abstract
We describe a robust and accurate algorithm, nicknamed AffTrack, to track selected features of a rigid 3D object in a video recording, given a canonical image of each feature and its position on the object. AffTrack uses a synergistic combination of a multiscale feature finder and a flexible camera calibrator. This synergy between the two modules allows, AffTrack to recover features after occlusions of arbitrary duration. Compared to other solutions to this problem, AffTrack can handle videos with variable zoom, and variable lens distortion, does not require a complete geometric model of the object, and does not require the selection of key frames. Tests indicate that AffTrack is more robust and accurate than the popular object trackers include in H. Kato's ARToolKit and in the OpenCV library.
Rodrigo Minetto, Neucimar J. Leite, Jorge Stolfi
ICIP1