EDBT 2026 Demo / reviewers in the wild / expert
Thiago Oliveira-Santos
dblp:83/9778 · also Thiago Oliveira Dos Santos
· DBLP profile ↗
72ranked-venue papers
1as first author
31since 2021 · last 2026
0000-0001-7607-635XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 64 · 1 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 since 2021Systems, architecture and hardware · 4Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring question answering: metric analysis and evaluation framework for enhanced interpretabilityabstractAbstract Evaluating open-ended question answering (QA) remains challenging, as traditional metrics often fail to reflect semantic correctness, especially in cases with paraphrastic variation or multiple valid answers. To deal with this challenge, we propose Score2Choice , a structured evaluation framework that reformulates QA evaluation as a multiple-choice selection task. This setup enables similarity-based metrics to be interpreted via accuracy, enhancing transparency and comparability. To support this approach, we introduce WikiTrapQA , a new MCQA dataset built from recent Wikipedia content and enriched with paraphrased and adversarial answers. Alongside a reformulated version of TruthfulQA, this dataset allows us to systematically compare lexical, semantic, and LLM-based metrics. A preliminary score distribution analysis reveals that many metrics struggle to distinguish correct from incorrect answers based on similarity scores alone. Experimental results show that LLM-based methods, in our case LLaMA 3, achieve the highest discriminative performance, while Sentence-BERT and BARTScore emerge as strong non-LLM alternatives. Our findings highlight the limitations of surface-level metrics and demonstrate the value of Score2Choice as a reproducible and interpretable framework for QA evaluation. Letícia C. Navarro, Sérgio Silva Mucciaccia, Filipe Wall Mutz, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
Neural Comput. Appl. | 7 |
| 2025 | Automatic Multiple-Choice Question Generation and Evaluation Systems Based on LLM: A Study Case With University ResolutionsabstractMultiple choice questions (MCQs) are often used in both employee selection and training, providing objectivity, efficiency, and scalability. However, their creation is resource-intensive, requiring significant expertise and financial investment. This study leverages large language models (LLMs) and prompt engineering techniques to automate the generation and validation of MCQs, particularly within the context of university regulations. Mainly, two novel approaches are proposed in this work: an automatic question generation system for university resolution and an automatic evaluation system to assess the performance of MCQ generation systems. The generation system combines different prompt engineering techniques and a review process to create well formulated questions. The evaluation system uses prompt engineering combined with an advanced LLM model to assess the integrity of the generated question. Experimental results demonstrate the effectiveness of both systems. The findings highlight the transformative potential of LLMs in educational assessment, reducing the burden on human resources and enabling scalable, cost-effective MCQ generation. Sérgio Silva Mucciaccia, Thiago Meireles Paixão, Filipe Wall Mutz, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
COLING | 6 |
| 2025 | Integrating Neural Language Models in Clinical Practice: A Case Study on Chronic Kidney Disease ConsultationsabstractThis study explores the integration of large language models (LLMs) into medical consultations, specifically in the care of chronic kidney disease (CKD) patients. Nephrology, the branch of medicine focused on kidney health, relies heavily on detailed patient histories and complex clinical reasoning. To support this process, we collected real-world consultation data in partnership with a hospital and developed an AI-driven framework to assist in transcription and clinical documentation. Our system automates speech-to-text conversion and structured text refinement, leveraging Whisper for transcription and GPT-4o for medical reasoning and summarization. The model achieved ROUGE-L F1 scores of 0.43 for patient phrases, 0.57 for medical team phrases, and 0.63 overall, with TF-IDF cosine similarity scores reaching 0.76, 0.89, and 0.91, respectively. Clinical outputs were evaluated by specialized LLMs, OpenBioLLM-70B and DeepSeek v3, with most predictions aligning with or providing plausible alternatives to physician assessments. These findings highlight the potential of LLMs to enhance clinical workflows by reducing documentation burdens while maintaining diagnostic accuracy. Victor Nascimento Neves, Claudine Badue, Joao Felipe Gobeti Calenzani, Lucas Thom Ramos, Thiago Oliveira-Santos, Alberto Ferreira de Souza |
IJCNN | 5 |
| 2025 | Decision-Making Algorithm Based on Multiple Context-Aware Agents for Human-Robot InteractionabstractDecision-making in robotic systems within real-world environments involves significant challenges, such as interpreting ambiguous user requests, integrating unstructured contextual data, and effectively utilizing sensory information. Current approaches in Human-Robot Interaction (HRI), which rely on Natural Language Processing (NLP) techniques, often encounter limitations such as scalability, ambiguity in communication, and the inability to relate unstructured data to structured data. These constraints reduce robots’ adaptability and limit their functional flexibility in dynamic environments. This study proposes a novel decision-making algorithm that integrates Generative Artificial Intelligence with context-aware systems and utilizes a modular framework comprising request validation, map validation, and response generation. By leveraging Large Language Models (LLMs) and Retrieval-Augmented Generation strategies, the system efficiently synthesizes and relates structured and unstructured data, enabling robots to navigate, guide users, and perform adaptive functions. Validation experiments demonstrated a 91% success rate, showcasing the system’s ability to process user requests and execute tasks typical of guide robots. This highlights the transformative role of LLMs in NLP, revolutionizing HRI and enhancing decision-making in real-world scenarios. Elio David Triana Rodríguez, Guilherme Goes Zanetti, Ricardo C. de Mello, Mario F. Jiménez, Thiago Oliveira-Santos, Anselmo Frizera-Neto |
IJCNN | 5 |
| 2025 | Blast furnace control by hierarchical similarity search in historical data
Lucas L. Amorim, Filipe Wall Mutz, Thiago Meireles Paixão, Vinicius Rampinelli, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos |
Eng. Appl. Artif. Intell. | 7 |
| 2025 | Electrical submersible pump fault diagnosis based on 2D transformation of vibration signals and transfer learning of image classification networks
Luciano Henrique Peixoto da Silva, Alexandre Rodrigues 0001, Flávio Miguel Varejão, Marcos Pellegrini Ribeiro, Thiago Oliveira-Santos |
Neural Comput. Appl. | 5 |
| 2025 | TransConv: a lightweight architecture based on transformers and convolutional neural networks for adenocarcinoma and Barrett's esophagus identification
Luis Souza 0001, André G. C. Pacheco, Alberto Ferreira de Souza, Thiago Oliveira-Santos, Claudine Badue, Christoph Palm, João Paulo Papa |
Neural Comput. Appl. | 4 |
| 2025 | Budget-aware pruning: Handling multiple domains with less parameters
Samuel Felipe dos Santos, Rodrigo Ferreira Berriel, Thiago Oliveira-Santos, Nicu Sebe, Jurandy Almeida |
Pattern Recognit. | 3 |
| 2024 | Few-Shot Copycat: Improving Performance of Black-Box Attack with Random Natural Images and Few Examples of Problem Domain
Jhonatan Machado Leão, Jacson Rodrigues Correia da Silva, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos |
ICPR (7) | 5 |
| 2024 | A Study on the Effectiveness of GPT-4V in Classifying Driver Behavior Captured on Video Using Just a Few Frames per VideoabstractThis paper introduces an innovative study that evaluates the effectiveness of GPT-4V vision processing technology in identifying risk events within driving scenarios. These scenarios are captured in a series of videos, with GPT-4V’s analysis focusing on only a few frames from each video. The study specifically targets risk behaviors such as yawning, smoking, phone usage, and distractions from the road. To achieve this, it utilizes a comprehensive collection of video recordings featuring drivers, which have been previously annotated by human evaluators to identify and tag instances of such risk behaviors. Our methodology involves a detailed analysis of GPT-4V’s performance, assessing its accuracy and consistency against human benchmarks across both private and public datasets. For the private dataset, GPT-4V demonstrated strong performance in identifying yawning events with a 98.9% accuracy, closely followed by a 98.4% accuracy in detecting smoking. In terms of recognizing driver distractions, it achieved a 91.7% accuracy, and for phone usage, it recorded a 95.7% accuracy. The "Face Not Visible" events achieved a 94.1% accuracy. For the public dataset, GPT-4V achieved an accuracy of 90.9% for the "Using Cellphone" category, with a recall of 76.6% and a precision of 92.1%. In identifying "Distraction" events, it achieved an accuracy of 91.0%, with a recall of 93.1% and a precision of 97.4%. For "Yawning" events, it achieved an accuracy of 98.2%, although its recall was lower at 43.7%, with a precision of 87.5%. These findings significantly enhance our understanding of how multimodal foundation models can be applied to improve road safety. Additionally, these results provide clear directions for future developments in autonomous monitoring systems. Joao Felipe Gobeti Calenzani, Victor Nascimento Neves, Lucas Thom Ramos, Lauro Jose Lyrio Junior, Luiz C. S. Magnago, Claudine Badue, Thiago Oliveira-Santos, Alberto Ferreira de Souza |
IJCNN | 7 |
| 2024 | Integrating Pretrained CNNs with One-Class Classifiers for Fault-Agnostic Electrical Submersible Pumps Anomaly DetectionabstractAs employed in industries, anomaly detection systems can be unstable due to the lack of training examples and often narrow feature extraction methods, both of which burden common models as abnormalities are rare and exceptionally unique. Since equipment used within oil and gas companies often tend to incur on great financial losses due to inherent unrecognized malfunctions, diagnosing components’ vibration signals beforehand becomes essential. However, as obtaining examples of a plethora of possible fault patterns can be difficult, developing a system based on the equipment’s normal behavior allows for better identification of irregularities within the signals. Additionally, as erroneous feature extraction methods can mask or dismiss a signal’s abnormal demeanor, approaching this issue with a system that has generalized knowledge over the dataset’s nature may grant a more thorough signal analysis. This paper advances the fault detection in the oil industry by investigating well-known feature extraction neural networks (such as ResNet, VGG and VGGish) together with one-class classifiers (such as Isolation Forest, Elliptic Envelope and OneClass SVM). With this approach, vibration signals are converted (using transformations like spectrogram) to images in order to leverage image-pretrained networks. Results show that features extracted with pretrained networks have similar performance to those hand-crafted with prior knowledge about the faults and are significantly better than standard statistical features. Furthermore, the audio-originated pretrained CNN, VGGish, tends to perform slightly better than those trained with natural images like ImageNET. Therefore, the proposed system can be applied as a fault-agnostic feature extraction method alternative to the problem specific hand-crafted feature extraction in one-class classification problems. Nilo Garcia Monteiro, Luciano Henrique Peixoto da Silva, Alexandre Rodrigues 0001, Flávio Miguel Varejão, Marcos Pellegrini Ribeiro, Thiago Oliveira-Santos |
IJCNN | 6 |
| 2024 | A Mobile Application for Speed Bump Sign Detection Using Neural NetworksabstractA significant number of traffic accidents are attributed to poor traffic signalization, including issues related to speed bump signs, which are a crucial element in the driving context to regulate speed limits. In these circumstances, object detection associated with deep learning methodology is rapidly gaining momentum over the Advanced Driver-Assistance Systems (ADAS) context. It has already showed to be very efficient in helping the driver to achieve a safer experience. Exploring this technology, this article presents a mobile application that detects speed bump signs in real time using the smartphone’s back camera. The application generates sound and visual alerts for the driver to increase the chances of a proper reaction to the speed bump. Since mobile phone usually have limited computational capacities, this article leverages the use of a compact deep learning model trained using an artificially generated dataset. Results show that the final model is able to operate with the limited mobile resources available, while achieving a 90.8 percent Average Precision Score. Matheus Meier Schreiber, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos |
IJCNN | 4 |
| 2024 | Analysis of Bias in GPT Language Models through Fine-tuning Containing Divergent DataabstractIn this study, we examined the effects of integrating data that contains divergent information, especially concerning anti-vaccination narratives, into the training of a GPT-2 language model. The model was fine-tuned using content sourced from anti-vaccination groups and channels on Telegram, aiming to analyze its ability to generate coherent and rationalized texts in comparison to a model pre-trained on OpenAI’s WebText dataset. The results demonstrate that fine-tuning a GPT-2 model with biased data leads the model to perpetuate these biases in its responses, albeit with a certain degree of rationalization. This finding underscores the importance of using high-quality and reliable data in training natural language processing models, highlighting the implications for information dissemination through these models. It also provides social scientists with a tool to explore and understand the complexities and challenges associated with public health misinformation via the use of language models, particularly in the context of vaccines. Leandro Furlam Turi, Athus Cavalini, Giovanni Comarela, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza |
IJCNN | 4 |
| 2024 | An open source experimental framework and public dataset for vibration-based fault diagnosis of electrical submersible pumps used on offshore oil exploration
Flávio Miguel Varejão, Lucas H. Sousa Mello, Marcos Pellegrini Ribeiro, Thiago Oliveira-Santos, Alexandre Rodrigues 0001 |
Knowl. Based Syst. | 4 |
| 2023 | Active learning for new-fault class sample recovery in electrical submersible pump fault diagnosis
Luciano Henrique Peixoto da Silva, Lucas H. Sousa Mello, Alexandre Rodrigues 0001, Flávio Miguel Varejão, Marcos Pellegrini Ribeiro, Thiago Oliveira-Santos |
Expert Syst. Appl. | 6 |
| 2022 | Unsupervised Domain Adaptation for Video Transformers in Action RecognitionabstractOver the last few years, Unsupervised Domain Adaptation (UDA) techniques have acquired remarkable importance and popularity in computer vision. However, when compared to the extensive literature available for images, the field of videos is still relatively unexplored. On the other hand, the performance of a model in action recognition is heavily affected by domain shift. In this paper, we propose a simple and novel UDA approach for video action recognition. Our approach leverages recent advances on spatio-temporal transformers to build a robust source model that better generalises to the target domain. Furthermore, our architecture learns domain invariant features thanks to the introduction of a novel alignment loss term derived from the Information Bottleneck principle. We report results on two video action recognition benchmarks for UDA, showing state-of-the-art performance on HMDB ↔ UCF, as well as on Kinetics→NEC-Drone, which is more challenging. This demonstrates the effectiveness of our method in handling different levels of domain shift. The source code is available at https://github.com/vturrisi/UDAVT. Victor G. T. da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, Elisa Ricci 0001 |
ICPR | 4 |
| 2022 | Lane Marking Detection and Classification using Spatial-Temporal Feature PoolingabstractThe lane detection problem has been extensively researched in the past decades, especially since the advent of deep learning. Despite the numerous works proposing solutions to the localization task (i.e., localizing the lane boundaries in an input image), the classification task has not seen the same focus. Nonetheless, knowing the type of lane boundary, particularly that of the ego lane, can be very useful for many applications. For instance, a vehicle might not be allowed by law to overtake depending on the type of the ego lane. Beyond that, very few works take advantage of the temporal information available in the videos captured by the vehicles: most methods employ a single-frame approach. In this work, building upon the recent LaneATT model, we propose an approach to exploit the temporal information and integrate the classification task into the model. Our results show that the proposed modifications can improve the detection performance on the most recent benchmark by 2.34 %, establishing a new state-of-the-art. Finally, an extensive evaluation shows that it enables a high classification performance (89.37 %) that serves as a future benchmark for the field. Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos |
IJCNN | 5 |
| 2022 | Dual-Head Contrastive Domain Adaptation for Video Action RecognitionabstractUnsupervised domain adaptation (UDA) methods have become very popular in computer vision. However, while several techniques have been proposed for images, much less attention has been devoted to videos. This paper introduces a novel UDA approach for action recognition from videos, inspired by recent literature on contrastive learning. In particular, we propose a novel two-headed deep architecture that simultaneously adopts cross-entropy and contrastive losses from different network branches to robustly learn a target classifier. Moreover, this work introduces a novel large-scale UDA dataset, Mixamo→Kinetics, which, to the best of our knowledge, is the first dataset that considers the domain shift arising when transferring knowledge from synthetic to real video sequences. Our extensive experimental evaluation conducted on three publicly available benchmarks and on our new Mixamo→Kinetics dataset demonstrate the effectiveness of our approach, which outperforms the current state-of-the-art methods. Code is available at https://github.com/vturrisi/CO2A. Victor G. T. da Costa, Giacomo Zara, Paolo Rota, Thiago Oliveira-Santos, Nicu Sebe, Vittorio Murino, Elisa Ricci 0001 |
WACV | 4 |
| 2022 | Cross-domain object detection using unsupervised image translation
Vinicius F. Arruda, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Nicu Sebe, Thiago Oliveira-Santos |
Expert Syst. Appl. | 7 |
| 2022 | Deep traffic sign detection and recognition without target domain real images
Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Alberto Ferreira de Souza, Claudine Badue, Nicu Sebe, Thiago Oliveira-Santos |
Mach. Vis. Appl. | 7 |
| 2022 | A human-in-the-loop recommendation-based framework for reconstruction of mechanically shredded documents
Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Maria Cláudia Silva Boeres, Alessandro L. Koerich, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
Pattern Recognit. Lett. | 7 |
| 2021 | Keep Your Eyes on the Lane: Real-Time Attention-Guided Lane DetectionabstractModern lane detection methods have achieved remarkable performances in complex real-world scenarios, but many have issues maintaining real-time efficiency, which is important for autonomous vehicles. In this work, we pro-pose LaneATT: an anchor-based deep lane detection model, which, akin to other generic deep object detectors, uses the anchors for the feature pooling step. Since lanes follow a regular pattern and are highly correlated, we hypothesize that in some cases global information may be crucial to infer their positions, especially in conditions such as occlusion, missing lane markers, and others. Thus, this work proposes a novel anchor-based attention mechanism that aggregates global information. The model was evaluated extensively on three of the most widely used datasets in the literature. The results show that our method outperforms the current state-of-the-art methods showing both higher efficacy and efficiency. Moreover, an ablation study is performed along with a discussion on efficiency trade-off options that are useful in practice. Code and models are available at https://github.com/lucastabelini/LaneATT. Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
CVPR | 6 |
| 2021 | Sisfrutos Papaya: A Dataset for Detection and Classification of Diseases in Papaya
Jairo Lucas de Moraes, Jorcy de Oliveira Neto, Jacson Rodrigues Correia da Silva, Thiago Meireles Paixão, Claudine Badue, Thiago Oliveira-Santos, Alberto Ferreira de Souza |
ICANN (2) | 6 |
| 2021 | Path Planning in Unstructured Urban Environments for Self-driving Cars
Anderson Mozart, Gabriel Moraes, Ranik Guidolini, Vinicius B. Cardoso, Thiago Oliveira-Santos, Alberto Ferreira de Souza, Claudine Badue |
ICINCO | 5 |
| 2021 | Visual Global Localization Based on Deep Neural Netwoks for Self-Driving CarsabstractIn this work, we present a visual global localization system based on Deep Neural Networks (DNNs) for self-driving cars, named DeepVgl(Deep Visual Global Localization). In training mode, DeepVglis trained with images and associated poses from datasets built during the mapping process; and, in operating mode, DeepVglreceives images captured online and infers the global poses of the self-driving car. To assess the performance of DeepVgl,we carried out experiments using datasets collected by experimental self-driving cars on trips made over long periods of time, thus including significant changes in the environment, traffic volume and weather conditions, as well as different times of the day and seasons of the year. Experimental results show that DeepVglis able to correctly locate the self-driving car up to 75% of the time for 0.2 m of accuracy and 96% of the time for 5 m of accuracy. Thiago Gonçalves Cavalcante, Thiago Oliveira-Santos, Alberto Ferreira de Souza, Claudine Badue, Avelino Forechi |
IJCNN | 2 |
| 2021 | Detecting Cancerous Tissue in Mammograms Using Deep Neural NetworksabstractBreast cancer cases are steadily increasing over the years. To improve the chances of a successful treatment, it is essential to diagnose and treat lesions in the initial stages. Screening methods, such as mammography, are one of the most effective strategies to achieve this goal. However, identifying some types of lesions can be challenging due to their morphological structure. For instance, calcifications typically have small diameters that range from 0.1mm to 1mm. Detecting calcifications require a careful and detailed analysis of the exams by radiologists. This work proposes a system for detecting malignant calcifications in mammograms using deep convolutional neural networks to assist in this task. The mammography image is split into patches that are classified by the deep convolutional neural network as containing malignant lesions (cancerous tissue) or not (healthy tissue). The patch-wise probabilities of cancer are then summarized so that a whole mammogram decision can be obtained. Moreover, a heatmap is built highlighting regions with a high probability of containing cancerous tissue. This heatmap can be used to draw the attention of radiologists to specific areas. Experiments performed with the CBIS-DDSM dataset indicate that the best model evaluated achieved the F-Measure of 0.99 in the classification task. These results indicate that applying the system in a real environment could improve the detection of cancer cases. Sabrina S. Panceri, Filipe Wall Mutz, Vinicius B. Cardoso, Raphael V. Carneiro, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza |
IJCNN | 5 |
| 2021 | Deep traffic light detection by overlaying synthetic context on arbitrary natural images
Jean Pablo Vieira de Mello, Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Alberto Ferreira de Souza, Claudine Badue, Nicu Sebe, Thiago Oliveira-Santos |
Comput. Graph. | 8 |
| 2021 | Self-driving cars: A survey
Claudine Badue, Ranik Guidolini, Raphael V. Carneiro, Pedro Azevedo, Vinicius B. Cardoso, Avelino Forechi, Luan F. R. Jesus, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Filipe Wall Mutz, Lucas de Paula Veronese, Thiago Oliveira-Santos, Alberto Ferreira de Souza |
Expert Syst. Appl. | 12 |
| 2021 | What is the best grid-map for self-driving cars localization? An evaluation under diverse types of illumination, traffic, and environment
Filipe Wall Mutz, Thiago Oliveira-Santos, Avelino Forechi, Karin Satie Komati, Claudine Badue, Felipe M. G. França, Alberto Ferreira de Souza |
Expert Syst. Appl. | 2 |
| 2021 | Copycat CNN: Are random non-Labeled data enough to steal knowledge from black-box models?
Jacson Rodrigues Correia da Silva, Rodrigo Ferreira Berriel, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
Pattern Recognit. | 5 |
| 2021 | Evaluating the Limits of a LiDAR for an Autonomous Driving LocalizationabstractIn general, proposed solutions for LiDAR-based localization used in autonomous cars require expensive sensors and computationally expensive mapping processes. Moreover, the global localization for autonomous driving is converging to the use of maps. Straightforward strategies to reduce the costs are to produce simpler sensors and use maps already available on the Internet. Here, an analysis is presented to show how simple can a LiDAR sensor be without degrading the localization accuracy that uses road and satellite maps together to globally pose the car. Three characteristics of the sensor are evaluated: the number of range readings, the amount of noise in the LiDAR readings, and the frame rate, with the aim of finding the minimum number of LiDAR lines, the maximum acceptable noise and the sensor frame rate needed to obtain an accurate position estimation. The analysis is performed using an autonomous car in complex field scenarios equipped with a 3D LiDAR Velodyne HDL-32E. Several experiments were conducted reducing the number of frames, the number of scans per 3D point-cloud and artificially adding up to 15% of error in the ray length. Among other results, we found that using only 4 vertical lines per scan and with an artificial error added up to 15% of the ray length, the car was capable to localize itself within 2.11 meters error average. All experimental results and the followed methodology are explained in detail herein. Lucas de Paula Veronese, Fernando Alfredo Auat Cheeín, Filipe Wall Mutz, Thiago Oliveira-Santos, José E. Guivant, Edilson de Aguiar, Claudine Badue, Alberto Ferreira de Souza |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2020 | Fast(er) Reconstruction of Shredded Text Documents via Self-Supervised Deep Asymmetric Metric LearningabstractThe reconstruction of shredded documents consists in arranging the pieces of paper (shreds) in order to reassemble the original aspect of such documents. This task is particularly relevant for supporting forensic investigation as documents may contain criminal evidence. As an alternative to the laborious and time-consuming manual process, several researchers have been investigating ways to perform automatic digital reconstruction. A central problem in automatic reconstruction of shredded documents is the pairwise compatibility evaluation of the shreds, notably for binary text documents. In this context, deep learning has enabled great progress for accurate reconstructions in the domain of mechanically-shredded documents. A sensitive issue, however, is that current deep model solutions require an inference whenever a pair of shreds has to be evaluated. This work proposes a scalable deep learning approach for measuring pairwise compatibility in which the number of inferences scales linearly (rather than quadratically) with the number of shreds. Instead of predicting compatibility directly, deep models are leveraged to asymmetrically project the raw shred content onto a common metric space in which distance is proportional to the compatibility. Experimental results show that our method has accuracy comparable to the state-of-the-art with a speed-up of about 22 times for a test instance with 505 shreds (20 mixed shredded-pages from different documents). Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Maria Cláudia Silva Boeres, Alessandro L. Koerich, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
CVPR | 7 |
| 2020 | Deep Learning-based Type Identification of Volumetric MRI SequencesabstractThe analysis of Magnetic Resonance Imaging (MRI) sequences enables clinical professionals to monitor the progression of a brain tumor. As the interest for automatizing brain volume MRI analysis increases, it becomes convenient to have each sequence well identified. However, the unstandardized naming of MRI sequences makes their identification difficult for automated systems, as well as makes it difficult for researches to generate or use datasets for machine learning research. In the face of that, we propose a system for identifying types of brain MRI sequences based on deep learning. By training a Convolutional Neural Network (CNN) based on 18-layer ResNet architecture, our system can classify a volumetric brain MRI as a FLAIR, Tl, T1c or T2 sequence, or whether it does not belong to any of these classes. The network was evaluated on publicly available datasets comprising both, pre-processed (BraTS dataset) and non-pre-processed (TCGA-GBM dataset), image types with diverse acquisition protocols, requiring only a few slices of the volume for training. Our system can classify among sequence types with an accuracy of 96.81 %. Jean Pablo Vieira de Mello, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Mauricio Reyes 0001, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
ICPR | 7 |
| 2020 | PolyLaneNet: Lane Estimation via Deep Polynomial RegressionabstractOne of the main factors that contributed to the large advances in autonomous driving is the advent of deep learning. For safer self-driving vehicles, one of the problems that has yet to be solved completely is lane detection. Since methods for this task have to work in real-time (+30 FPS), they not only have to be effective (i.e., have high accuracy) but they also have to be efficient (i.e., fast). In this work, we present a novel method for lane detection that uses as input an image from a forward-looking camera mounted in the vehicle and outputs polynomials representing each lane marking in the image, via deep polynomial regression. The proposed method is shown to be competitive with existing state-of-the-art methods in the TuSimple dataset while maintaining its efficiency (115 FPS). Additionally, extensive qualitative results on two additional public datasets are presented, alongside with limitations in the evaluation metrics used by recent works for lane detection. Finally, we provide source code and trained models that allow others to replicate all the results shown in this paper, which is surprisingly rare in state-of-the-art lane detection methods. Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
ICPR | 6 |
| 2020 | A Large-Scale Mapping Method Based on Deep Neural Networks Applied to Self-Driving Car LocalizationabstractWe propose a new approach for real time inference of occupancy maps for self-driving cars using deep neural networks (DNN) named NeuralMapper. NeuralMapper receives LiDAR sensor data as input and generates as output the occupancy grid map around the car. NeuralMapper infers the probability of each grid map cell from one of the three following classes: Occupied, Free and Unknown. The system was tested with two datasets and achieved an average accuracy of 76.48% and 73.81%. We also evaluated our approach for localization purposes in a self-driving car and most of the localization pose errors were less than 0.20m with an RMSE of 0.28 which are close to the results in the literature for methods using other grid mapping approaches. Vinicius B. Cardoso, André Seidel Oliveira, Avelino Forechi, Pedro Azevedo, Filipe Wall Mutz, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza |
IJCNN | 6 |
| 2020 | Metric Learning for Electrical Submersible Pump Fault DiagnosisabstractMachine learning classification algorithms are highly dependent of a dataset composed of high-level features. In this paper, a deep learning approach is combined with traditional machine learning classifiers in order to circumvent the need of a specialist for extracting relevant features from one dimensional frequency-domain vibration signals. Our approach relies on a convolutional architecture trained with a triplet loss function for extracting relevant features directly from the raw data. A previously hand-crafted feature set, created by a specialist over the course of many years of research, is compared with the newly extracted feature set. Six conventional classifiers models (K-Nearest Neighbors, Support Vector Machine, Decision Tree, Random Forest, Quadratic Discriminant Analysis and Naive Bayes) are trained in both features set separately and compared in terms of macro F-measure. Results shows statistical evidence towards to the acceptance that the extracted feature set is as good as or better than the hand-crafted feature set, for classification purposes. Lucas H. Sousa Mello, Marcos Pellegrini Ribeiro, Thiago Oliveira-Santos, Flávio Miguel Varejão, Alexandre Rodrigues 0001 |
IJCNN | 3 |
| 2020 | Image-Based Real-Time Path Generation Using Deep Neural NetworksabstractWe propose an image-based real-time path planner for the self-driving car IARA, named DeepPath. DeepPath uses a CNN for inferring paths from images. During the self-driving car operation, DeepPath receives an image and the current car pose. Then, it sends the image to a CNN trained to infer a model of the path. After that, DeepPath generates the path in the IARA's coordinate system using the path model. Subsequently, given the current IARA's pose, DeepPath transforms each pose of the path in the IARA's coordinate system into another pose in the world coordinate system. Finally, it sends the path to the IARA's Behavior Selector subsystem, the next subsystem in the IARA's Decision-Making system. We evaluated the performance of DeepPath in real world scenarios. Our results showed that DeepPath is able to correctly generate paths for IARA that differ only slightly from those defined by humans. Gabriel Moraes, Anderson Mozart, Pedro Azevedo, Marcos Piumbini, Vinicius B. Cardoso, Thiago Oliveira-Santos, Alberto Ferreira de Souza, Claudine Badue |
IJCNN | 6 |
| 2020 | Product Categorization by Title Using Deep Neural Networks as Feature ExtractorabstractNatural Language Processing (NLP) has been receiving increasing attention in the past few years. In part, this is related to the huge flow of data being made available everyday on the internet, which increased the need for automatic tools capable of analyzing and extracting relevant information, especially from the text. In this context, text classification became one of the most studied tasks on the NLP domain. The objective is to assign predefined categories or labels to text or sentences. Important applications include sentence classification, sentiment analysis, spam detection, among many others. This work proposes an automatic system for product categorization using only their titles. The proposed system employs a state-of-the-art deep neural network as a tool to extract features from the titles to be used as input in different machine learning models. The system is evaluated in the large-scale Mercado Libre dataset, which has the common characteristics of real-world problems such as imbalanced classes, unreliable labels, besides having a large number of samples: 20,000,000 in total. The results showed that the proposed system was able to correctly categorize the products with a balanced accuracy of 86.57% on the local test split of the Mercado Libre dataset. It also surpassed the fourth place on the public rank of the MeLi Data Challenge with 91.19% of balanced accuracy, which represents less than 1% of the difference to the winner. Leonardo S. Paulucio, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos |
IJCNN | 6 |
| 2020 | Self-supervised deep reconstruction of mixed strip-shredded text documents
Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Maria Cláudia Silva Boeres, Alessandro L. Koerich, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
Pattern Recognit. | 7 |
| 2019 | Budget-Aware Adapters for Multi-Domain LearningabstractMulti-Domain Learning (MDL) refers to the problem of learning a set of models derived from a common deep architecture, each one specialized to perform a task in a certain domain (e.g., photos, sketches, paintings). This paper tackles MDL with a particular interest in obtaining domain-specific models with an adjustable budget in terms of the number of network parameters and computational complexity. Our intuition is that, as in real applications the number of domains and tasks can be very large, an effective MDL approach should not only focus on accuracy but also on having as few parameters as possible. To implement this idea we derive specialized deep models for each domain by adapting a pre-trained architecture but, differently from other methods, we propose a novel strategy to automatically adjust the computational complexity of the network. To this aim, we introduce Budget-Aware Adapters that select the most relevant feature channels to better handle data from a novel domain. Some constraints on the number of active switches are imposed in order to obtain a network respecting the desired complexity budget. Experimentally, we show that our approach leads to recognition accuracy competitive with state-of-the-art approaches but with much lighter networks both in terms of storage and computation. Rodrigo Ferreira Berriel, Stéphane Lathuilière, Moin Nabi, Tassilo Klein, Thiago Oliveira-Santos, Nicu Sebe, Elisa Ricci 0001 |
ICCV | 5 |
| 2019 | Cross-Domain Car Detection Using Unsupervised Image-to-Image Translation: From Day to NightabstractDeep learning techniques have enabled the emergence of state-of-the-art models to address object detection tasks. However, these techniques are data-driven, delegating the accuracy to the training dataset which must resemble the images in the target task. The acquisition of a dataset involves annotating images, an arduous and expensive process, generally requiring time and manual effort. Thus, a challenging scenario arises when the target domain of application has no annotated dataset available, making tasks in such situation to lean on a training dataset of a different domain. Sharing this issue, object detection is a vital task for autonomous vehicles where the large amount of driving scenarios yields several domains of application requiring annotated data for the training process. In this work, a method for training a car detection system with annotated data from a source domain (day images) without requiring the image annotations of the target domain (night images) is presented. For that, a model based on Generative Adversarial Networks (GANs) is explored to enable the generation of an artificial dataset with its respective annotations. The artificial dataset (fake dataset) is created translating images from day-time domain to night-time domain. The fake dataset, which comprises annotated images of only the target domain (night images), is then used to train the car detector model. Experimental results showed that the proposed method achieved significant and consistent improvements, including the increasing by more than 10% of the detection performance when compared to the training with only the available annotated data (i.e., day images). Vinicius F. Arruda, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Alberto Ferreira de Souza, Claudine Badue, Nicu Sebe, Thiago Oliveira-Santos |
IJCNN | 7 |
| 2019 | Bio-Inspired Foveated Technique for Augmented-Range Vehicle Detection Using Deep Neural NetworksabstractWe propose a bio-inspired foveated technique to detect cars in a long range camera view using a deep convolutional neural network (DCNN) for the IARA self-driving car. The DCNN receives as input (i) an image, which is captured by a camera installed on IARA's roof; and (ii) crops of the image, which are centered in the waypoints computed by IARA's path planner and whose sizes increase with the distance from IARA. We employ an overlap filter to discard detections of the same car in different crops of the same image based on the percentage of overlap of detections' bounding boxes. We evaluated the performance of the proposed augmented-range vehicle detection system (ARVDS) using the hardware and software infrastructure available in the IARA self-driving car. Using IARA, we captured thousands of images of real traffic situations containing cars in a long range. Experimental results show that ARVDS increases the Average Precision (AP) of long range car detection from 29.51% (using a single whole image) to 63.15%. Pedro Azevedo, Sabrina S. Panceri, Ranik Guidolini, Vinicius B. Cardoso, Claudine Badue, Thiago Oliveira-Santos, Alberto Ferreira de Souza |
IJCNN | 6 |
| 2019 | Removing Movable Objects from Grid Maps of Self-Driving Cars Using Deep Neural NetworksabstractWe propose a technique for removing traces of movable objects from occupancy grid maps based on deep neural networks, dubbed enhanced occupancy grid map generation (E-OGM-G). In E-OGM-G, we capture camera images synchronized and aligned with LiDAR rays, semantically segment these images, and compute which laser rays of the LiDAR hit pixels segmented as belonging to movable objects. By clustering laser rays that are close together in a 2D projection, we are able to identify clusters that belong to movable objects and avoid using them in the process of generating the OGMs - this allows generating OGMs clean of movable objects. Clean OGMs are important for several aspects of self-driving cars' operation (i.e., localization). We tested E-OGM-G using data obtained in a real-world scenario - a 2.6 km stretch of a busy multi-lane urban road. Our results showed that E-OGM-G can achieve a precision of 81.19% considering the whole OGMs generated, of 89.76% considering a track in these OGMs of width of 12 m, and of 100.00% considering a track of width of 3.4 m. We then tested a self-driving car using the automatically cleaned OGMs. The self-driving car was able to properly localize itself and to autonomously drive itself in the world using the cleaned OGMs. These successful results showed that the proposed technique is effective in removing movable objects from static OGMs. Ranik Guidolini, Raphael V. Carneiro, Claudine Badue, Thiago Oliveira-Santos, Alberto Ferreira de Souza |
IJCNN | 4 |
| 2019 | Traffic Light Recognition Using Deep Learning and Prior Maps for Autonomous CarsabstractAutonomous terrestrial vehicles must be capable of perceiving traffic lights and recognizing their current states to share the streets with human drivers. Most of the time, human drivers can easily identify the relevant traffic lights. To deal with this issue, a common solution for autonomous cars is to integrate recognition with prior maps. However, additional solution is required for the detection and recognition of the traffic light. Deep learning techniques have showed great performance and power of generalization including traffic related problems. Motivated by the advances in deep learning, some recent works leveraged some state-of-the-art deep detectors to locate (and further recognize) traffic lights from 2D camera images. However, none of them combine the power of the deep learning-based detectors with prior maps to recognize the state of the relevant traffic lights. Based on that, this work proposes to integrate the power of deep learning-based detection with the prior maps used by our car platform IARA (acronym for Intelligent Autonomous Robotic Automobile) to recognize the relevant traffic lights of predefined routes. The process is divided in two phases: an offline phase for map construction and traffic lights annotation; and an online phase for traffic light recognition and identification of the relevant ones. The proposed system was evaluated on five test cases (routes) in the city of Vitória, each case being composed of a video sequence and a prior map with the relevant traffic lights for the route. Results showed that the proposed technique is able to correctly identify the relevant traffic light along the trajectory. Lucas C. Possatti, Ranik Guidolini, Vinicius B. Cardoso, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
IJCNN | 8 |
| 2019 | Effortless Deep Training for Traffic Sign Detection Using Templates and Arbitrary Natural ImagesabstractDeep learning has been successfully applied to several problems related to autonomous driving. Often, these solutions rely on large networks that require databases of real image samples of the problem (i.e., real world) for proper training. The acquisition of such real-world data sets is not always possible in the autonomous driving context, and sometimes their annotation is not feasible (e.g., takes too long or is too expensive). Moreover, in many tasks, there is an intrinsic data imbalance that most learning-based methods struggle to cope with. It turns out that traffic sign detection is a problem in which these three issues are seen altogether. In this work, we propose a novel database generation method that requires only (i) arbitrary natural images, i.e., requires no real image from the domain of interest, and (ii) templates of the traffic signs, i.e., templates synthetically created to illustrate the appearance of the category of a traffic sign. The effortlessly generated training database is shown to be effective for the training of a deep detector (such as Faster R-CNN) on German traffic signs, achieving 95.66% of mAP on average. In addition, the proposed method is able to detect traffic signs with an average precision, recall and F1-score of about 94%, 91% and 93%, respectively. The experiments surprisingly show that detectors can be trained with simple data generation methods and without problem domain data for the background, which is in the opposite direction of the common sense for deep learning. Lucas Tabelini Torres, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Alberto Ferreira de Souza, Claudine Badue, Nicu Sebe, Thiago Oliveira-Santos |
IJCNN | 7 |
| 2019 | Handling pedestrians in self-driving cars using image tracking and alternative path generation with Frenét frames
Renan Sarcinelli, Ranik Guidolini, Vinicius B. Cardoso, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Pedro Azevedo, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos |
Comput. Graph. | 9 |
| 2019 | Reducing power companies billing costs via empirical bayes and seasonality remover
Alexandre Rodrigues 0001, Lucas Martinuzzo, Flávio Miguel Varejão, Vítor E. Silva Souza, Thiago Oliveira-Santos |
Eng. Appl. Artif. Intell. | 5 |
| 2019 | Exploring Character Shapes for Unsupervised Reconstruction of Strip-Shredded Text DocumentsabstractDigital reconstruction of mechanically shredded documents has received increasing attention in the last years mainly for historical and forensics needs. Computational methods to solve this problem are highly desirable in order to mitigate the time-consuming human effort and to preserve document integrity. The reconstruction of strips-shredded documents is accomplished by horizontally splicing pieces so that the arising sequence (solution) is as similar as the original document. In this context, a central issue is the quantification of the fitting between the pieces (strips), which generally involves stating a function that associates a pair of strips to a real value indicating the fitting quality. This problem is also more challenging for text documents, such as business letters or legal documents, since they depict poor color information. The system proposed here addresses this issue by exploring character shapes as visual features for compatibility computation. Experiments conducted with real mechanically shredded documents showed that our approach outperformed in accuracy other popular techniques in the literature considering documents with (almost) only textual content. Thiago Meireles Paixão, Maria Cláudia Silva Boeres, Cinthia Obladen de Almendra Freitas, Thiago Oliveira-Santos |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2018 | Heading Direction Estimation Using Deep Learning with Automatic Large-scale Data AcquisitionabstractAdvanced Driver Assistance Systems (ADAS) have experienced major advances in the past few years. The main objective of ADAS includes keeping the vehicle in the correct road direction, and avoiding collision with other vehicles or obstacles around. In this paper, we address the problem of estimating the heading direction that keeps the vehicle aligned with the road direction. This information can be used in precise localization, road and lane keeping, lane departure warning, and others. To enable this approach, a large-scale database 1+ million images) was automatically acquired and annotated using publicly available platforms such as the Google Street View API and OpenStreetMap. After the acquisition of the database, a CNN model was trained to predict how much the heading direction of a car should change in order to align it to the road 4 meters ahead. To assess the performance of the model, experiments were performed using images from two different sources: a hidden test set from Google Street View (GSV) images and two datasets from our autonomous car (IARA). The model achieved a low mean average error of 2.359° and 2.524° for the GSV and IARA datasets, respectively; performing consistently across the different datasets. It is worth noting that the images from the IARA dataset are very different (camera, FOV, brightness, etc.) from the ones of the GSV dataset, which shows the robustness of the model. In conclusion, the model was trained effortlessly (using automatic processes) and showed promising results in real-world databases working in real-time (more than 75 frames per second). Rodrigo Ferreira Berriel, Lucas Tabelini Torres, Vinicius B. Cardoso, Ranik Guidolini, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
IJCNN | 7 |
| 2018 | Mapping Road Lanes Using Laser Remission and Deep Neural NetworksabstractWe propose the use of deep neural networks (DNN) for solving the problem of inferring the position and relevant properties of lanes of urban roads with poor or absent horizontal signalization, in order to allow the operation of autonomous cars in such situations. We take a segmentation approach to the problem and use the Efficient Neural Network (ENet) DNN for segmenting LiDAR remission grid maps into road maps. We represent road maps using what we called road grid maps. Road grid maps are square matrixes and each element of these matrixes represents a small square region of real-world space. The value of each element is a code associated with the semantics of the road map. Our road grid maps contain all information about the roads' lanes required for building the Road Definition Data Files (RDDFs) that are necessary for the operation of our autonomous car, IARA (Intelligent Autonomous Robotic Automobile). We have built a dataset of tens of kilometers of manually marked road lanes and used part of it to train ENet to segment road grid maps from remission grid maps. After being trained, ENet achieved an average segmentation accuracy of 83.7%. We have tested the use of inferred road grid maps in the real world using IARA on a stretch of 3.7 km of urban roads and it has shown performance equivalent to that of the previous IARA's subsystem that uses a manually generated RDDF. Raphael V. Carneiro, Rafael Correia Nascimento, Ranik Guidolini, Vinicius B. Cardoso, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza |
IJCNN | 5 |
| 2018 | Visual Global Localization with a Hybrid WNN-CNN ApproachabstractCurrently, self-driving cars rely greatly on the Global Positioning System (GPS) infrastructure, albeit there is an increasing demand for alternative methods for GPS-denied environments. One of them is known as place recognition, which associates images of places with their corresponding positions. We previously proposed systems based on Weightless Neural Networks (WNN) to address this problem as a classification task. This encompasses solely one part of the global localization, which is not precise enough for driverless cars. Instead of just recognizing past places and outputting their poses, it is desired that a global localization system estimates the pose of current place images. In this paper, we propose to tackle this problem as follows. Firstly, given a live image, the place recognition system returns the most similar image and its pose. Then, given live and recollected images, a visual localization system outputs the relative camera pose represented by those images. To estimate the relative camera pose between the recollected and the current images, a Convolutional Neural Network (CNN) is trained with the two images as input and a relative pose vector as output. Together, these systems solve the global localization problem using the topological and metric information to approximate the current vehicle pose. The full approach is compared to a Real- Time Kinematic GPS system and a Simultaneous Localization and Mapping (SLAM) system. Experimental results show that the proposed approach correctly localizes a vehicle 90% of the time with a mean error of 1.20m compared to 1.12m of the SLAM system and 0.37m of the GPS, 89% of the time. Avelino Forechi, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza |
IJCNN | 2 |
| 2018 | Handling Pedestrians in Crosswalks Using Deep Neural Networks in the IARA Autonomous CarabstractIn this work, we propose a subsystem to handle pedestrians in crosswalks using deep neural networks for the IARA autonomous car, which relies on camera and LIDAR data fusion. Crosswalks' positions were manually annotated in IARA's map. Pedestrians are detected in the camera image using a convolutional neural network (CNN). Then, pedestrians' positions in the map are obtained by fusing their positions in the image with the LIDAR point cloud. Subsequently, if a pedestrian position is inside the crosswalk area, the crosswalk is set as busy. Finally, a busy crosswalk message is published to the High-Level Decision Maker subsystem. This subsystem selects the car's behavior according to the crosswalk condition and propagates this decision down through the control pipeline, in order to make the car drive correctly through the crosswalk area. The Pedestrian Handler subsystem was evaluated on IARA, which was driven autonomously for various laps along a real and complex circuit with various crosswalks. In all passages through crosswalks, the Pedestrian Handler dealt with pedestrians as expected, i.e., without any human intervention. Ranik Guidolini, Lucas G. Scart, Luan F. R. Jesus, Vinicius B. Cardoso, Claudine Badue, Thiago Oliveira-Santos |
IJCNN | 6 |
| 2018 | Copycat CNN: Stealing Knowledge by Persuading Confession with Random Non-Labeled DataabstractIn the past few years, Convolutional Neural Networks (CNNs) have been achieving state-of-the-art performance on a variety of problems. Many companies employ resources and money to generate these models and provide them as an API, therefore it is in their best interest to protect them, i.e., to avoid that someone else copy them. Recent studies revealed that stateof-the-art CNNs are vulnerable to adversarial examples attacks, and this weakness indicates that CNNs do not need to operate in the problem domain (PD). Therefore, we hypothesize that they also do not need to be trained with examples of the PD in order to operate in it. Given these facts, in this paper, we investigate if a target blackbox CNN can be copied by persuading it to confess its knowledge through random non-labeled data. The copy is two-fold: i) the target network is queried with random data and its predictions are used to create a fake dataset with the knowledge of the network; and ii) a copycat network is trained with the fake dataset and should be able to achieve similar performance as the target network. This hypothesis was evaluated locally in three problems (facial expression, object, and crosswalk classification) and against a cloud-based API. In the copy attacks, images from both nonproblem domain and PD were used. All copycat networks achieved at least 93.7% of the performance of the original models with non-problem domain data, and at least 98.6% using additional data from the PD. Additionally, the copycat CNN successfully copied at least 97.3% of the performance of the Microsoft Azure Emotion API. Our results show that it is possible to create a copycat CNN by simply querying a target network as black-box with random non-labeled data. Jacson Rodrigues Correia da Silva, Rodrigo Ferreira Berriel, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
IJCNN | 5 |
| 2018 | Electricity Readers Routing Based on Clustering and Communities DetectionabstractElectric power distribution companies in Brazil assess the energy consumption of most of their costumers by reading the meters in loco. A human reader has an itinerary with clients that should be read on each working day. The number of meters per route tends to increase over time, as the number of customers is constantly growing. At a certain point, the route must be restructured so that it still runs in a single day. This paper proposes the use of clustering techniques integrated with community detection algorithms and a heuristic routing algorithm to solve the problem of globally restructuring the routes used by the readers of the energy distribution companies. Experimental results showed that the proposed method significantly reduced the number of routes required to perform the meters readings. Lucas Martinuzzo, Diego Lucchil, Flávio Miguel Varejão, Alexandre Rodngues, Filipe Lima, Thiago Oliveira-Santos |
INDIN | 6 |
| 2017 | A Model-Predictive Motion Planner for the IARA autonomous carabstractWe present the Model-Predictive Motion Planner (MPMP) of the Intelligent Autonomous Robotic Automobile (IARA). IARA is a fully autonomous car that uses a path planner to compute a path from its current position to the desired destination. Using this path, the current position, a goal in the path and a map, IARA's MPMP is able to compute smooth trajectories from its current position to the goal in less than 50 ms. MPMP computes the poses of these trajectories so that they follow the path closely and, at the same time, are at a safe distance of occasional obstacles. Our experiments have shown that MPMP is able to compute trajectories that precisely follow a path produced by a Human driver (distance of 0.15 m in average) while smoothly driving IARA at speeds of up to 32.4 km/h (9 m/s). Vinicius B. Cardoso, Josias Oliveira, Thomas Teixeira, Claudine Badue, Filipe Wall Mutz, Thiago Oliveira-Santos, Lucas de Paula Veronese, Alberto Ferreira de Souza |
ICRA | 6 |
| 2017 | Monthly energy consumption forecast: A deep learning approachabstractEvery year, energy consumption grows world widely. Therefore, power companies need to investigate models to better forecast and plan the energy use. One approach to address this problem is the estimation of energy consumption in the customer level. Energy consumption forecasting problem is a time series regression task. It consists of predicting the energy consumption for the next month given a finite history of a customer. Machine learning techniques have shown promising results in a variety of problems including time series and regression problems. Part of these promising results are attributed to deep neural networks. Although investigated in other domains, deep architectures have not been used to address the energy consumption prediction problem. In this work, we propose a system to predict monthly energy consumption using deep learning techniques. Three deep learning models were studied: Deep Fully Connected, Convolutional and Long Short-Term Memory Neural Networks. Due to the sensitivity of these models to the input range, normalization techniques were also investigated. The proposed system was validated with real data of almost a million customers (resulting in over 9 million samples). Results showed that our system can predict monthly energy consumption with an absolute error of 31.83 kWh and a relative error of 17.29%. Rodrigo Ferreira Berriel, Andre Teixeira Lopes, Alexandre Rodrigues 0001, Flávio Miguel Varejão, Thiago Oliveira-Santos |
IJCNN | 5 |
| 2017 | Kernel and random extreme learning machine applied to submersible motor pump fault diagnosisabstractThis paper presents an extension of a comparative study of classifier architectures for automatic fault diagnosis, with a special emphasis on the Extreme Learning Machine (ELM), with and without kernel mapping. Besides the explanation of the ELM model, an attempt is made to find theoretical hints of the excellent generalization capabilities of this model, based on the findings of Cover about dichotomies and the equivalence of Mean Squared Error minimization in the high-dimensional feature spaces induced by kernels, and spaces defined by a finite sample set. The field of application is a practical problem in the context of offshore petroleum exploration where sophisticated submersible motor pumps are extensively tested before being deployed. The work juxtaposes the performance of ELM to an existing statistically sound comparison of state of the art classifier methods for a hand-crafted feature model tailored specially to the spectra of the vibrational signals of the pump. The results suggest the remarkably good generalization capability of ELM, exhibiting the highest scores for the chosen F-measure performance criterion. Thomas W. Rauber, Thiago Oliveira-Santos, Francisco de Assis Boldt, Alexandre Rodrigues 0001, Flávio Miguel Varejão, Marcos Pellegrini Ribeiro |
IJCNN | 2 |
| 2017 | Automatic large-scale data acquisition via crowdsourcing for crosswalk classification: A deep learning approach
Rodrigo Ferreira Berriel, Franco Schmidt Rossi, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
Comput. Graph. | 4 |
| 2017 | Ego-Lane Analysis System (ELAS): Dataset and algorithms
Rodrigo Ferreira Berriel, Edilson de Aguiar, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
Image Vis. Comput. | 4 |
| 2017 | Deep Learning-Based Large-Scale Automatic Satellite Crosswalk ClassificationabstractHigh-resolution satellite imagery has been increasingly used on remote sensing classification problems. One of the main factors is the availability of this kind of data. Despite the high availability, very little effort has been placed on the zebra crossing classification problem. In this letter, crowdsourcing systems are exploited in order to enable the automatic acquisition and annotation of a large-scale satellite imagery database for crosswalks related tasks. Then, this data set is used to train deep-learning-based models in order to accurately classify satellite images that contain or not contain zebra crossings. A novel data set with more than 240000 images from 3 continents, 9 countries, and more than 20 cities was used in the experiments. The experimental results showed that freely available crowdsourcing data can be used to accurately (97.11%) train robust models to perform crosswalk classification on a global scale. Rodrigo Ferreira Berriel, Andre Teixeira Lopes, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2017 | Facial expression recognition with Convolutional Neural Networks: Coping with few data and the training sample order
Andre Teixeira Lopes, Edilson de Aguiar, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
Pattern Recognit. | 4 |
| 2016 | Submersible Motor Pump Fault Diagnosis System: A Comparative Study of Classification MethodsabstractIn this paper, an artificial intelligence solution to diagnose faults before acquisition of submersible petroleum motor pump systems is presented. Proper fault identification is time consuming and demands highly trained human experts. The diagnosis system is intended to facilitate the work of the human component of this important process by replicating the decision of highly trained experts through a classifier. To perform the automatic diagnosis, firstly intermediate features are extracted as the vibration spectra. Subsequently, high level features are extracted and fed into a classifier that outputs the final diagnose. To validate our proposal and to select the best classifier (among K-Nearest-Neighbour, Random Forest, Support Vector Machine and Decision Tree) for this problem, we performed a comparative study using real data acquired in tests accomplished before acquisition of submersible motor pumps. Our dataset comprises thousands of entries of accelerometer sensors (vertically distributed along the particular system components) data labelled by an human expert to one of the considered scenarios (normal pump, faulty sensor, faulty pump with rubbing, misalignment or unbalance). Results have showed that the evaluated classifiers have equivalent performance for the given problem, and that the standardization procedure can improve the performance of some classifiers. The performance of the classifiers is sufficient to facilitate the work performed by humans and consequently reduce the time spent in the pump fault diagnosis process. Thiago Oliveira-Santos, Thomas W. Rauber, Flávio Miguel Varejão, Lucas Martinuzzo, Willian Oliveira, Marcos Pellegrini Ribeiro, Alexandre Rodrigues 0001 |
ICTAI | 1 |
| 2016 | Sequential appearance-based Global Localization using an ensemble of kNN-DTW classifiersabstractThe human Episodic Memory system stores sequences of events (episodes encoded in time and in space) experienced or imagined by an individual for later access to the episodes in whole or in part. Such ability provided by our Episodic Memory system is important, among other functions, for our localization in space throughout time. Inspired by that, in this paper, it is presented a new Sequential Appearance-Based approach to the Global Localization problem, dubbed SABGL. SABGL is based on an ensemble of kNN classifiers, where each classifier uses the Dynamic Time Warping (DTW) and the Hamming distance to compare binary features extracted from sequences of images. SABGL is designed to solve the global localization problem in two phases: mapping and localization. During mapping, it is trained with a sequence of images and associated locations that represents episodes experienced by an autonomous robot. During localization, it receives subsequences of images of the same environment and compares them to its previous experienced episodes, trying to recollect the most similar “experience” in time and space at once. Then, the system outputs the positions where it “believes” these images were captured. SABGL is compared with a previous single-image approach, named VibGL. Experimental results show that SABGL consistently outperforms VibGL in classification accuracy with higher precision in all maximum distance tolerance analyzed. For instance, given a maximum tolerance of 10m, SABGL is able to correctly localize an autonomous car 98% of the time. Avelino Forechi, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos |
IJCNN | 4 |
| 2016 | Simulating robotic cars using time-delay neural networksabstractIn this paper, we propose a simulator for robotic cars based on two time-delay neural networks. These networks are intended to simulate the mechanisms that govern how a set of effort commands changes the car's velocity and the direction it is moving. The first neural network receives as input a temporal sequence of current and previous throttle and brake efforts, along with a temporal sequence of the previous car's velocities (estimated by the network), and outputs the velocity that the real car would reach in the next time interval given these inputs. The second neural network estimates the arctangent of curvature (a variable related to the steering wheel angle) that a real car would reach in the next time interval given a temporal sequence of current and previous steering efforts and previous arctangents of curvatures of the car estimated by the network. We evaluated the performance of our simulator using real-world datasets acquired using an autonomous robotic car. Experimental results showed that our simulator was able to simulate in real time how a set of efforts influences the car's velocity and arctangent of curvature. While navigating in a map of a real-world environment, our car simulator was able to emulate the velocity and arctangent of curvature of the real car with mean squared error of 2.2×10-3(m/s)2and 4.0×10-5rad2, respectively. Alberto Ferreira de Souza, Jacson Rodrigues Correia da Silva, Filipe Wall Mutz, Claudine Badue, Thiago Oliveira-Santos |
IJCNN | 5 |
| 2016 | Large-scale mapping in complex field scenarios using an autonomous car
Filipe Wall Mutz, Lucas de Paula Veronese, Thiago Oliveira-Santos, Edilson de Aguiar, Fernando Alfredo Auat Cheeín, Alberto Ferreira de Souza |
Expert Syst. Appl. | 3 |
| 2016 | Visual tracking with VG-RAM Weightless Neural Networks
Mariella Berger, Alberto Ferreira de Souza, Jorcy de Oliveira Neto, Edilson de Aguiar, Thiago Oliveira-Santos |
Neurocomputing | 5 |
| 2016 | Fat-Fast VG-RAM WNN: A high performance approach
Avelino Forechi, Alberto Ferreira de Souza, Jorcy de Oliveira Neto, Edilson de Aguiar, Claudine Badue, Artur S. d'Avila Garcez, Thiago Oliveira-Santos |
Neurocomputing | 7 |
| 2015 | Image-based mapping, global localization and position tracking using VG-RAM weightless neural networksabstractHumans can easily memorize images of places and labels (road names, addresses, etc.) associated with them, as well as trajectories defined by sequences of images and corresponding positions. Later, they are able to remember places' labels and relative positions when seeing the same images again. In this work, we present an image-based mapping, global localization and position tracking system based on Virtual Generalizing Random Access Memory (VG-RAM) weightless neural networks, dubbed VIBML. VIBML mimics humans ability of learning about a place and of recognizing the same place in a later moment, as well as of tracking self-movement through the environment using images. We evaluated the performance of VIBML on the precise localization of an autonomous car using real-world datasets. Our experimental results showed that VIBML is able to localize car-like robots on large maps of real world environments with accuracy equivalent to that of state-of-the-art methods - VIBML is able to localize an autonomous car with average positioning error of 1.12m and with 75% of the poses with error below 1.5m in a 3.75km path around the main campus of the Federal University of Espírito Santo. Lauro Jose Lyrio Junior, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza |
ICRA | 2 |
| 2015 | Re-emission and satellite aerial maps applied to vehicle localization on urban environmentsabstractVehicle localization in large-scale urban environments has been commonly addressed as a map-matching problem in the literature. Generally, the maps are 2D images of the world where each pixel covers a part of it. However, building maps for large-scale urban environments requires driving the vehicle along the desired path at least once. In order to simplify this task, in this work, we propose a new localization system that uses satellite aerial map-images available on the Internet to localize a vehicle in a complex urban environment. Satellite aerial map-images are compared against re-emission maps built from the infrared reflectance information of the vehicle's LiDAR. Normalized Mutual Information (NMI) is used to compare re-emission and aerial map images. A Particle Filter Localization strategy is applied for vehicle's localization. As a result, the system has an accuracy of 0.89m in a test course with 6.5km. Our system can be used continuously without losing track, and it works even in dark and partially occluded areas. Lucas de Paula Veronese, Edilson de Aguiar, Rafael Correia Nascimento, José E. Guivant, Fernando Alfredo Auat Cheeín, Alberto Ferreira de Souza, Thiago Oliveira-Santos |
IROS | 7 |
| 2014 | Compressing VG-RAM WNN memory for lightweight applicationsabstractThe Virtual Generalizing Random Access Memory Weightless Neural Network (VG-RAM WNN) is an effective machine learning technique that offers simple implementation and fast training. One disadvantage of VG-RAM WNN, however, is the test time for applications with many training samples, i.e. large multi-class classification applications. In such cases, the test time tends to be high, since it increases with the size of the memory of each neuron. In this paper, we present a new methodology for handling such applications using VG-RAM WNN. By employing data clustering techniques to reduce the overall size of the neurons' memory, we were able to reduce the network's memory footprint and the system's runtime, while maintaining a high and acceptable classification performance. We evaluated the performance of our VG-RAM WNN system with compressed memory on the problem of traffic sign recognition. Our experimental results showed that, after compression, the system was able to run at very fast response times in standard computers. Also, we were able to load and run the system at interactive rates in small low-power systems, experiencing only a small reduction in classification performance. Edilson de Aguiar, Avelino Forechi, Lucas de Paula Veronese, Mariella Berger, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos |
IJCNN | 7 |
| 2014 | Image-based global localization using VG-RAM Weightless Neural NetworksabstractMapping and localization are fundamental problems in autonomous robotics. Autonomous robots need to know where they are in their area of operation to navigate through it and to perform activities of interest. In this paper, we propose an Image-Based Global Localization (VibGL) system that uses Virtual Generalizing Random Access Memory Weightless Neural Networks (VG-RAM WNN). For mapping, we employ a VG-RAM WNN that learns the world positions associated with the images captured along a trajectory. During the localization, new images from the trajectory are presented to the VG-RAM WNN, which outputs their positions in the world. We performed experiments with our VibGL system applied to the problem of localizing an autonomous car. Our experimental results show that the system is able to learn large maps (several kilometers in length) of real world environments and perform global localization with median pose precision of about 3m. Considering a tolerance of 10m VibGL is able to localize the car 95% of the time. Lauro Jose Lyrio Junior, Thiago Oliveira-Santos, Avelino Forechi, Lucas de Paula Veronese, Claudine Badue, Alberto Ferreira de Souza |
IJCNN | 2 |
| 2014 | Programming a VG-RAM based Neural Network ComputerabstractWe propose a Virtual Generalizing Random Access Memory (VG-RAM) Weightless Neural Network (WNN) Computer (V'Ger Computer for short). VG-RAM WNNs are very effective pattern recognition tools, offering fast training (one shot training) and competitive recognition performance, if compared with other current techniques. The V'Ger Computer architecture was inspired on the organization of the human neocortex and is composed of hierarchically organized and recurrently interconnected layers of VG-RAM WNN neurons. One layer is connected to another in a way similar to cortico-cortical feed-forward and feedback connections between functionally adjacent and hierarchically organized areas. We have "programmed" the V'Ger Computer for counting from 0 to 9 three times. Our preliminary experimental results showed that V'Ger is capable of executing this sequence of actions in spite of strong interferences. Alberto Ferreira de Souza, Avelino Forechi, Filipe Wall Mutz, Mariella Berger, Thiago Oliveira-Santos, Claudine Badue |
IJCNN | 5 |