Claudine Badue

dblp:12/97 · also Claudine Santos Badue · DBLP profile ↗
← Back
64ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0003-1810-8581ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 2 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-authorSystems, architecture and hardware · 2
YearPublicationVenuePosition
2026 Exploring question answering: metric analysis and evaluation framework for enhanced interpretability
abstract
Abstract Evaluating open-ended question answering (QA) remains challenging, as traditional metrics often fail to reflect semantic correctness, especially in cases with paraphrastic variation or multiple valid answers. To deal with this challenge, we propose Score2Choice , a structured evaluation framework that reformulates QA evaluation as a multiple-choice selection task. This setup enables similarity-based metrics to be interpreted via accuracy, enhancing transparency and comparability. To support this approach, we introduce WikiTrapQA , a new MCQA dataset built from recent Wikipedia content and enriched with paraphrased and adversarial answers. Alongside a reformulated version of TruthfulQA, this dataset allows us to systematically compare lexical, semantic, and LLM-based metrics. A preliminary score distribution analysis reveals that many metrics struggle to distinguish correct from incorrect answers based on similarity scores alone. Experimental results show that LLM-based methods, in our case LLaMA 3, achieve the highest discriminative performance, while Sentence-BERT and BARTScore emerge as strong non-LLM alternatives. Our findings highlight the limitations of surface-level metrics and demonstrate the value of Score2Choice as a reproducible and interpretable framework for QA evaluation.
Letícia C. Navarro, Sérgio Silva Mucciaccia, Filipe Wall Mutz, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
Neural Comput. Appl.5
2025 Automatic Multiple-Choice Question Generation and Evaluation Systems Based on LLM: A Study Case With University Resolutions
abstract
Multiple choice questions (MCQs) are often used in both employee selection and training, providing objectivity, efficiency, and scalability. However, their creation is resource-intensive, requiring significant expertise and financial investment. This study leverages large language models (LLMs) and prompt engineering techniques to automate the generation and validation of MCQs, particularly within the context of university regulations. Mainly, two novel approaches are proposed in this work: an automatic question generation system for university resolution and an automatic evaluation system to assess the performance of MCQ generation systems. The generation system combines different prompt engineering techniques and a review process to create well formulated questions. The evaluation system uses prompt engineering combined with an advanced LLM model to assess the integrity of the generated question. Experimental results demonstrate the effectiveness of both systems. The findings highlight the transformative potential of LLMs in educational assessment, reducing the burden on human resources and enabling scalable, cost-effective MCQ generation.
Sérgio Silva Mucciaccia, Thiago Meireles Paixão, Filipe Wall Mutz, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
COLING4
2025 Integrating Neural Language Models in Clinical Practice: A Case Study on Chronic Kidney Disease Consultations
abstract
This study explores the integration of large language models (LLMs) into medical consultations, specifically in the care of chronic kidney disease (CKD) patients. Nephrology, the branch of medicine focused on kidney health, relies heavily on detailed patient histories and complex clinical reasoning. To support this process, we collected real-world consultation data in partnership with a hospital and developed an AI-driven framework to assist in transcription and clinical documentation. Our system automates speech-to-text conversion and structured text refinement, leveraging Whisper for transcription and GPT-4o for medical reasoning and summarization. The model achieved ROUGE-L F1 scores of 0.43 for patient phrases, 0.57 for medical team phrases, and 0.63 overall, with TF-IDF cosine similarity scores reaching 0.76, 0.89, and 0.91, respectively. Clinical outputs were evaluated by specialized LLMs, OpenBioLLM-70B and DeepSeek v3, with most predictions aligning with or providing plausible alternatives to physician assessments. These findings highlight the potential of LLMs to enhance clinical workflows by reducing documentation burdens while maintaining diagnostic accuracy.
Victor Nascimento Neves, Claudine Badue, Joao Felipe Gobeti Calenzani, Lucas Thom Ramos, Thiago Oliveira-Santos, Alberto Ferreira de Souza
IJCNN2
2025 Blast furnace control by hierarchical similarity search in historical data
Lucas L. Amorim, Filipe Wall Mutz, Thiago Meireles Paixão, Vinicius Rampinelli, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos
Eng. Appl. Artif. Intell.6
2025 TransConv: a lightweight architecture based on transformers and convolutional neural networks for adenocarcinoma and Barrett's esophagus identification
Luis Souza 0001, André G. C. Pacheco, Alberto Ferreira de Souza, Thiago Oliveira-Santos, Claudine Badue, Christoph Palm, João Paulo Papa
Neural Comput. Appl.5
2024 Few-Shot Copycat: Improving Performance of Black-Box Attack with Random Natural Images and Few Examples of Problem Domain
Jhonatan Machado Leão, Jacson Rodrigues Correia da Silva, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos
ICPR (7)4
2024 A Study on the Effectiveness of GPT-4V in Classifying Driver Behavior Captured on Video Using Just a Few Frames per Video
abstract
This paper introduces an innovative study that evaluates the effectiveness of GPT-4V vision processing technology in identifying risk events within driving scenarios. These scenarios are captured in a series of videos, with GPT-4V’s analysis focusing on only a few frames from each video. The study specifically targets risk behaviors such as yawning, smoking, phone usage, and distractions from the road. To achieve this, it utilizes a comprehensive collection of video recordings featuring drivers, which have been previously annotated by human evaluators to identify and tag instances of such risk behaviors. Our methodology involves a detailed analysis of GPT-4V’s performance, assessing its accuracy and consistency against human benchmarks across both private and public datasets. For the private dataset, GPT-4V demonstrated strong performance in identifying yawning events with a 98.9% accuracy, closely followed by a 98.4% accuracy in detecting smoking. In terms of recognizing driver distractions, it achieved a 91.7% accuracy, and for phone usage, it recorded a 95.7% accuracy. The "Face Not Visible" events achieved a 94.1% accuracy. For the public dataset, GPT-4V achieved an accuracy of 90.9% for the "Using Cellphone" category, with a recall of 76.6% and a precision of 92.1%. In identifying "Distraction" events, it achieved an accuracy of 91.0%, with a recall of 93.1% and a precision of 97.4%. For "Yawning" events, it achieved an accuracy of 98.2%, although its recall was lower at 43.7%, with a precision of 87.5%. These findings significantly enhance our understanding of how multimodal foundation models can be applied to improve road safety. Additionally, these results provide clear directions for future developments in autonomous monitoring systems.
Joao Felipe Gobeti Calenzani, Victor Nascimento Neves, Lucas Thom Ramos, Lauro Jose Lyrio Junior, Luiz C. S. Magnago, Claudine Badue, Thiago Oliveira-Santos, Alberto Ferreira de Souza
IJCNN6
2024 A Mobile Application for Speed Bump Sign Detection Using Neural Networks
abstract
A significant number of traffic accidents are attributed to poor traffic signalization, including issues related to speed bump signs, which are a crucial element in the driving context to regulate speed limits. In these circumstances, object detection associated with deep learning methodology is rapidly gaining momentum over the Advanced Driver-Assistance Systems (ADAS) context. It has already showed to be very efficient in helping the driver to achieve a safer experience. Exploring this technology, this article presents a mobile application that detects speed bump signs in real time using the smartphone’s back camera. The application generates sound and visual alerts for the driver to increase the chances of a proper reaction to the speed bump. Since mobile phone usually have limited computational capacities, this article leverages the use of a compact deep learning model trained using an artificially generated dataset. Results show that the final model is able to operate with the limited mobile resources available, while achieving a 90.8 percent Average Precision Score.
Matheus Meier Schreiber, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos
IJCNN3
2024 Analysis of Bias in GPT Language Models through Fine-tuning Containing Divergent Data
abstract
In this study, we examined the effects of integrating data that contains divergent information, especially concerning anti-vaccination narratives, into the training of a GPT-2 language model. The model was fine-tuned using content sourced from anti-vaccination groups and channels on Telegram, aiming to analyze its ability to generate coherent and rationalized texts in comparison to a model pre-trained on OpenAI’s WebText dataset. The results demonstrate that fine-tuning a GPT-2 model with biased data leads the model to perpetuate these biases in its responses, albeit with a certain degree of rationalization. This finding underscores the importance of using high-quality and reliable data in training natural language processing models, highlighting the implications for information dissemination through these models. It also provides social scientists with a tool to explore and understand the complexities and challenges associated with public health misinformation via the use of language models, particularly in the context of vaccines.
Leandro Furlam Turi, Athus Cavalini, Giovanni Comarela, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza
IJCNN5
2022 Lane Marking Detection and Classification using Spatial-Temporal Feature Pooling
abstract
The lane detection problem has been extensively researched in the past decades, especially since the advent of deep learning. Despite the numerous works proposing solutions to the localization task (i.e., localizing the lane boundaries in an input image), the classification task has not seen the same focus. Nonetheless, knowing the type of lane boundary, particularly that of the ego lane, can be very useful for many applications. For instance, a vehicle might not be allowed by law to overtake depending on the type of the ego lane. Beyond that, very few works take advantage of the temporal information available in the videos captured by the vehicles: most methods employ a single-frame approach. In this work, building upon the recent LaneATT model, we propose an approach to exploit the temporal information and integrate the classification task into the model. Our results show that the proposed modifications can improve the detection performance on the most recent benchmark by 2.34 %, establishing a new state-of-the-art. Finally, an extensive evaluation shows that it enables a high classification performance (89.37 %) that serves as a future benchmark for the field.
Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos
IJCNN4
2022 Extracting Knowledge from Pharmaceutical Package Inserts
Cristiano da Silveira Colombo, Claudine Badue, Elias de Oliveira
ISDA (3)2
2022 Cross-domain object detection using unsupervised image translation
Vinicius F. Arruda, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Nicu Sebe, Thiago Oliveira-Santos
Expert Syst. Appl.4
2022 Deep traffic sign detection and recognition without target domain real images
Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Alberto Ferreira de Souza, Claudine Badue, Nicu Sebe, Thiago Oliveira-Santos
Mach. Vis. Appl.5
2022 A human-in-the-loop recommendation-based framework for reconstruction of mechanically shredded documents
Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Maria Cláudia Silva Boeres, Alessandro L. Koerich, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
Pattern Recognit. Lett.5
2021 Keep Your Eyes on the Lane: Real-Time Attention-Guided Lane Detection
abstract
Modern lane detection methods have achieved remarkable performances in complex real-world scenarios, but many have issues maintaining real-time efficiency, which is important for autonomous vehicles. In this work, we pro-pose LaneATT: an anchor-based deep lane detection model, which, akin to other generic deep object detectors, uses the anchors for the feature pooling step. Since lanes follow a regular pattern and are highly correlated, we hypothesize that in some cases global information may be crucial to infer their positions, especially in conditions such as occlusion, missing lane markers, and others. Thus, this work proposes a novel anchor-based attention mechanism that aggregates global information. The model was evaluated extensively on three of the most widely used datasets in the literature. The results show that our method outperforms the current state-of-the-art methods showing both higher efficacy and efficiency. Moreover, an ablation study is performed along with a discussion on efficiency trade-off options that are useful in practice. Code and models are available at https://github.com/lucastabelini/LaneATT.
Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
CVPR4
2021 Sisfrutos Papaya: A Dataset for Detection and Classification of Diseases in Papaya
Jairo Lucas de Moraes, Jorcy de Oliveira Neto, Jacson Rodrigues Correia da Silva, Thiago Meireles Paixão, Claudine Badue, Thiago Oliveira-Santos, Alberto Ferreira de Souza
ICANN (2)5
2021 Path Planning in Unstructured Urban Environments for Self-driving Cars
Anderson Mozart, Gabriel Moraes, Ranik Guidolini, Vinicius B. Cardoso, Thiago Oliveira-Santos, Alberto Ferreira de Souza, Claudine Badue
ICINCO7
2021 Visual Global Localization Based on Deep Neural Netwoks for Self-Driving Cars
abstract
In this work, we present a visual global localization system based on Deep Neural Networks (DNNs) for self-driving cars, named DeepVgl(Deep Visual Global Localization). In training mode, DeepVglis trained with images and associated poses from datasets built during the mapping process; and, in operating mode, DeepVglreceives images captured online and infers the global poses of the self-driving car. To assess the performance of DeepVgl,we carried out experiments using datasets collected by experimental self-driving cars on trips made over long periods of time, thus including significant changes in the environment, traffic volume and weather conditions, as well as different times of the day and seasons of the year. Experimental results show that DeepVglis able to correctly locate the self-driving car up to 75% of the time for 0.2 m of accuracy and 96% of the time for 5 m of accuracy.
Thiago Gonçalves Cavalcante, Thiago Oliveira-Santos, Alberto Ferreira de Souza, Claudine Badue, Avelino Forechi
IJCNN4
2021 Detecting Cancerous Tissue in Mammograms Using Deep Neural Networks
abstract
Breast cancer cases are steadily increasing over the years. To improve the chances of a successful treatment, it is essential to diagnose and treat lesions in the initial stages. Screening methods, such as mammography, are one of the most effective strategies to achieve this goal. However, identifying some types of lesions can be challenging due to their morphological structure. For instance, calcifications typically have small diameters that range from 0.1mm to 1mm. Detecting calcifications require a careful and detailed analysis of the exams by radiologists. This work proposes a system for detecting malignant calcifications in mammograms using deep convolutional neural networks to assist in this task. The mammography image is split into patches that are classified by the deep convolutional neural network as containing malignant lesions (cancerous tissue) or not (healthy tissue). The patch-wise probabilities of cancer are then summarized so that a whole mammogram decision can be obtained. Moreover, a heatmap is built highlighting regions with a high probability of containing cancerous tissue. This heatmap can be used to draw the attention of radiologists to specific areas. Experiments performed with the CBIS-DDSM dataset indicate that the best model evaluated achieved the F-Measure of 0.99 in the classification task. These results indicate that applying the system in a real environment could improve the detection of cancer cases.
Sabrina S. Panceri, Filipe Wall Mutz, Vinicius B. Cardoso, Raphael V. Carneiro, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza
IJCNN6
2021 Named Entities as a Metadata Resource for Indexing and Searching Information
Flávio Izo, Elias de Oliveira, Claudine Badue
ISDA3
2021 Deep traffic light detection by overlaying synthetic context on arbitrary natural images
Jean Pablo Vieira de Mello, Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Alberto Ferreira de Souza, Claudine Badue, Nicu Sebe, Thiago Oliveira-Santos
Comput. Graph.6
2021 Self-driving cars: A survey
Claudine Badue, Ranik Guidolini, Raphael V. Carneiro, Pedro Azevedo, Vinicius B. Cardoso, Avelino Forechi, Luan F. R. Jesus, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Filipe Wall Mutz, Lucas de Paula Veronese, Thiago Oliveira-Santos, Alberto Ferreira de Souza
Expert Syst. Appl.1
2021 What is the best grid-map for self-driving cars localization? An evaluation under diverse types of illumination, traffic, and environment
Filipe Wall Mutz, Thiago Oliveira-Santos, Avelino Forechi, Karin Satie Komati, Claudine Badue, Felipe M. G. França, Alberto Ferreira de Souza
Expert Syst. Appl.5
2021 Copycat CNN: Are random non-Labeled data enough to steal knowledge from black-box models?
Jacson Rodrigues Correia da Silva, Rodrigo Ferreira Berriel, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
Pattern Recognit.3
2021 Evaluating the Limits of a LiDAR for an Autonomous Driving Localization
abstract
In general, proposed solutions for LiDAR-based localization used in autonomous cars require expensive sensors and computationally expensive mapping processes. Moreover, the global localization for autonomous driving is converging to the use of maps. Straightforward strategies to reduce the costs are to produce simpler sensors and use maps already available on the Internet. Here, an analysis is presented to show how simple can a LiDAR sensor be without degrading the localization accuracy that uses road and satellite maps together to globally pose the car. Three characteristics of the sensor are evaluated: the number of range readings, the amount of noise in the LiDAR readings, and the frame rate, with the aim of finding the minimum number of LiDAR lines, the maximum acceptable noise and the sensor frame rate needed to obtain an accurate position estimation. The analysis is performed using an autonomous car in complex field scenarios equipped with a 3D LiDAR Velodyne HDL-32E. Several experiments were conducted reducing the number of frames, the number of scans per 3D point-cloud and artificially adding up to 15% of error in the ray length. Among other results, we found that using only 4 vertical lines per scan and with an artificial error added up to 15% of the ray length, the car was capable to localize itself within 2.11 meters error average. All experimental results and the followed methodology are explained in detail herein.
Lucas de Paula Veronese, Fernando Alfredo Auat Cheeín, Filipe Wall Mutz, Thiago Oliveira-Santos, José E. Guivant, Edilson de Aguiar, Claudine Badue, Alberto Ferreira de Souza
IEEE Trans. Intell. Transp. Syst.7
2020 Fast(er) Reconstruction of Shredded Text Documents via Self-Supervised Deep Asymmetric Metric Learning
abstract
The reconstruction of shredded documents consists in arranging the pieces of paper (shreds) in order to reassemble the original aspect of such documents. This task is particularly relevant for supporting forensic investigation as documents may contain criminal evidence. As an alternative to the laborious and time-consuming manual process, several researchers have been investigating ways to perform automatic digital reconstruction. A central problem in automatic reconstruction of shredded documents is the pairwise compatibility evaluation of the shreds, notably for binary text documents. In this context, deep learning has enabled great progress for accurate reconstructions in the domain of mechanically-shredded documents. A sensitive issue, however, is that current deep model solutions require an inference whenever a pair of shreds has to be evaluated. This work proposes a scalable deep learning approach for measuring pairwise compatibility in which the number of inferences scales linearly (rather than quadratically) with the number of shreds. Instead of predicting compatibility directly, deep models are leveraged to asymmetrically project the raw shred content onto a common metric space in which distance is proportional to the compatibility. Experimental results show that our method has accuracy comparable to the state-of-the-art with a speed-up of about 22 times for a test instance with 505 shreds (20 mixed shredded-pages from different documents).
Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Maria Cláudia Silva Boeres, Alessandro L. Koerich, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
CVPR5
2020 Deep Learning-based Type Identification of Volumetric MRI Sequences
abstract
The analysis of Magnetic Resonance Imaging (MRI) sequences enables clinical professionals to monitor the progression of a brain tumor. As the interest for automatizing brain volume MRI analysis increases, it becomes convenient to have each sequence well identified. However, the unstandardized naming of MRI sequences makes their identification difficult for automated systems, as well as makes it difficult for researches to generate or use datasets for machine learning research. In the face of that, we propose a system for identifying types of brain MRI sequences based on deep learning. By training a Convolutional Neural Network (CNN) based on 18-layer ResNet architecture, our system can classify a volumetric brain MRI as a FLAIR, Tl, T1c or T2 sequence, or whether it does not belong to any of these classes. The network was evaluated on publicly available datasets comprising both, pre-processed (BraTS dataset) and non-pre-processed (TCGA-GBM dataset), image types with diverse acquisition protocols, requiring only a few slices of the volume for training. Our system can classify among sequence types with an accuracy of 96.81 %.
Jean Pablo Vieira de Mello, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Mauricio Reyes 0001, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
ICPR5
2020 PolyLaneNet: Lane Estimation via Deep Polynomial Regression
abstract
One of the main factors that contributed to the large advances in autonomous driving is the advent of deep learning. For safer self-driving vehicles, one of the problems that has yet to be solved completely is lane detection. Since methods for this task have to work in real-time (+30 FPS), they not only have to be effective (i.e., have high accuracy) but they also have to be efficient (i.e., fast). In this work, we present a novel method for lane detection that uses as input an image from a forward-looking camera mounted in the vehicle and outputs polynomials representing each lane marking in the image, via deep polynomial regression. The proposed method is shown to be competitive with existing state-of-the-art methods in the TuSimple dataset while maintaining its efficiency (115 FPS). Additionally, extensive qualitative results on two additional public datasets are presented, alongside with limitations in the evaluation metrics used by recent works for lane detection. Finally, we provide source code and trained models that allow others to replicate all the results shown in this paper, which is surprisingly rare in state-of-the-art lane detection methods.
Lucas Tabelini Torres, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
ICPR4
2020 A Large-Scale Mapping Method Based on Deep Neural Networks Applied to Self-Driving Car Localization
abstract
We propose a new approach for real time inference of occupancy maps for self-driving cars using deep neural networks (DNN) named NeuralMapper. NeuralMapper receives LiDAR sensor data as input and generates as output the occupancy grid map around the car. NeuralMapper infers the probability of each grid map cell from one of the three following classes: Occupied, Free and Unknown. The system was tested with two datasets and achieved an average accuracy of 76.48% and 73.81%. We also evaluated our approach for localization purposes in a self-driving car and most of the localization pose errors were less than 0.20m with an RMSE of 0.28 which are close to the results in the literature for methods using other grid mapping approaches.
Vinicius B. Cardoso, André Seidel Oliveira, Avelino Forechi, Pedro Azevedo, Filipe Wall Mutz, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza
IJCNN7
2020 Image-Based Real-Time Path Generation Using Deep Neural Networks
abstract
We propose an image-based real-time path planner for the self-driving car IARA, named DeepPath. DeepPath uses a CNN for inferring paths from images. During the self-driving car operation, DeepPath receives an image and the current car pose. Then, it sends the image to a CNN trained to infer a model of the path. After that, DeepPath generates the path in the IARA's coordinate system using the path model. Subsequently, given the current IARA's pose, DeepPath transforms each pose of the path in the IARA's coordinate system into another pose in the world coordinate system. Finally, it sends the path to the IARA's Behavior Selector subsystem, the next subsystem in the IARA's Decision-Making system. We evaluated the performance of DeepPath in real world scenarios. Our results showed that DeepPath is able to correctly generate paths for IARA that differ only slightly from those defined by humans.
Gabriel Moraes, Anderson Mozart, Pedro Azevedo, Marcos Piumbini, Vinicius B. Cardoso, Thiago Oliveira-Santos, Alberto Ferreira de Souza, Claudine Badue
IJCNN8
2020 Product Categorization by Title Using Deep Neural Networks as Feature Extractor
abstract
Natural Language Processing (NLP) has been receiving increasing attention in the past few years. In part, this is related to the huge flow of data being made available everyday on the internet, which increased the need for automatic tools capable of analyzing and extracting relevant information, especially from the text. In this context, text classification became one of the most studied tasks on the NLP domain. The objective is to assign predefined categories or labels to text or sentences. Important applications include sentence classification, sentiment analysis, spam detection, among many others. This work proposes an automatic system for product categorization using only their titles. The proposed system employs a state-of-the-art deep neural network as a tool to extract features from the titles to be used as input in different machine learning models. The system is evaluated in the large-scale Mercado Libre dataset, which has the common characteristics of real-world problems such as imbalanced classes, unreliable labels, besides having a large number of samples: 20,000,000 in total. The results showed that the proposed system was able to correctly categorize the products with a balanced accuracy of 86.57% on the local test split of the Mercado Libre dataset. It also surpassed the fourth place on the public rank of the MeLi Data Challenge with 91.19% of balanced accuracy, which represents less than 1% of the difference to the winner.
Leonardo S. Paulucio, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos
IJCNN5
2020 Finding Entities and Related Facts in Newspaper
Jaimel de Oliveira Lima, Cristiano da Silveira Colombo, Flávio Izo, Elias de Oliveira, Claudine Badue
ISDA5
2020 Self-supervised deep reconstruction of mixed strip-shredded text documents
Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Maria Cláudia Silva Boeres, Alessandro L. Koerich, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
Pattern Recognit.5
2019 Cross-Domain Car Detection Using Unsupervised Image-to-Image Translation: From Day to Night
abstract
Deep learning techniques have enabled the emergence of state-of-the-art models to address object detection tasks. However, these techniques are data-driven, delegating the accuracy to the training dataset which must resemble the images in the target task. The acquisition of a dataset involves annotating images, an arduous and expensive process, generally requiring time and manual effort. Thus, a challenging scenario arises when the target domain of application has no annotated dataset available, making tasks in such situation to lean on a training dataset of a different domain. Sharing this issue, object detection is a vital task for autonomous vehicles where the large amount of driving scenarios yields several domains of application requiring annotated data for the training process. In this work, a method for training a car detection system with annotated data from a source domain (day images) without requiring the image annotations of the target domain (night images) is presented. For that, a model based on Generative Adversarial Networks (GANs) is explored to enable the generation of an artificial dataset with its respective annotations. The artificial dataset (fake dataset) is created translating images from day-time domain to night-time domain. The fake dataset, which comprises annotated images of only the target domain (night images), is then used to train the car detector model. Experimental results showed that the proposed method achieved significant and consistent improvements, including the increasing by more than 10% of the detection performance when compared to the training with only the available annotated data (i.e., day images).
Vinicius F. Arruda, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Alberto Ferreira de Souza, Claudine Badue, Nicu Sebe, Thiago Oliveira-Santos
IJCNN5
2019 Bio-Inspired Foveated Technique for Augmented-Range Vehicle Detection Using Deep Neural Networks
abstract
We propose a bio-inspired foveated technique to detect cars in a long range camera view using a deep convolutional neural network (DCNN) for the IARA self-driving car. The DCNN receives as input (i) an image, which is captured by a camera installed on IARA's roof; and (ii) crops of the image, which are centered in the waypoints computed by IARA's path planner and whose sizes increase with the distance from IARA. We employ an overlap filter to discard detections of the same car in different crops of the same image based on the percentage of overlap of detections' bounding boxes. We evaluated the performance of the proposed augmented-range vehicle detection system (ARVDS) using the hardware and software infrastructure available in the IARA self-driving car. Using IARA, we captured thousands of images of real traffic situations containing cars in a long range. Experimental results show that ARVDS increases the Average Precision (AP) of long range car detection from 29.51% (using a single whole image) to 63.15%.
Pedro Azevedo, Sabrina S. Panceri, Ranik Guidolini, Vinicius B. Cardoso, Claudine Badue, Thiago Oliveira-Santos, Alberto Ferreira de Souza
IJCNN5
2019 Removing Movable Objects from Grid Maps of Self-Driving Cars Using Deep Neural Networks
abstract
We propose a technique for removing traces of movable objects from occupancy grid maps based on deep neural networks, dubbed enhanced occupancy grid map generation (E-OGM-G). In E-OGM-G, we capture camera images synchronized and aligned with LiDAR rays, semantically segment these images, and compute which laser rays of the LiDAR hit pixels segmented as belonging to movable objects. By clustering laser rays that are close together in a 2D projection, we are able to identify clusters that belong to movable objects and avoid using them in the process of generating the OGMs - this allows generating OGMs clean of movable objects. Clean OGMs are important for several aspects of self-driving cars' operation (i.e., localization). We tested E-OGM-G using data obtained in a real-world scenario - a 2.6 km stretch of a busy multi-lane urban road. Our results showed that E-OGM-G can achieve a precision of 81.19% considering the whole OGMs generated, of 89.76% considering a track in these OGMs of width of 12 m, and of 100.00% considering a track of width of 3.4 m. We then tested a self-driving car using the automatically cleaned OGMs. The self-driving car was able to properly localize itself and to autonomously drive itself in the world using the cleaned OGMs. These successful results showed that the proposed technique is effective in removing movable objects from static OGMs.
Ranik Guidolini, Raphael V. Carneiro, Claudine Badue, Thiago Oliveira-Santos, Alberto Ferreira de Souza
IJCNN3
2019 Traffic Light Recognition Using Deep Learning and Prior Maps for Autonomous Cars
abstract
Autonomous terrestrial vehicles must be capable of perceiving traffic lights and recognizing their current states to share the streets with human drivers. Most of the time, human drivers can easily identify the relevant traffic lights. To deal with this issue, a common solution for autonomous cars is to integrate recognition with prior maps. However, additional solution is required for the detection and recognition of the traffic light. Deep learning techniques have showed great performance and power of generalization including traffic related problems. Motivated by the advances in deep learning, some recent works leveraged some state-of-the-art deep detectors to locate (and further recognize) traffic lights from 2D camera images. However, none of them combine the power of the deep learning-based detectors with prior maps to recognize the state of the relevant traffic lights. Based on that, this work proposes to integrate the power of deep learning-based detection with the prior maps used by our car platform IARA (acronym for Intelligent Autonomous Robotic Automobile) to recognize the relevant traffic lights of predefined routes. The process is divided in two phases: an offline phase for map construction and traffic lights annotation; and an online phase for traffic light recognition and identification of the relevant ones. The proposed system was evaluated on five test cases (routes) in the city of Vitória, each case being composed of a video sequence and a prior map with the relevant traffic lights for the route. Results showed that the proposed technique is able to correctly identify the relevant traffic light along the trajectory.
Lucas C. Possatti, Ranik Guidolini, Vinicius B. Cardoso, Rodrigo Ferreira Berriel, Thiago Meireles Paixão, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
IJCNN6
2019 Effortless Deep Training for Traffic Sign Detection Using Templates and Arbitrary Natural Images
abstract
Deep learning has been successfully applied to several problems related to autonomous driving. Often, these solutions rely on large networks that require databases of real image samples of the problem (i.e., real world) for proper training. The acquisition of such real-world data sets is not always possible in the autonomous driving context, and sometimes their annotation is not feasible (e.g., takes too long or is too expensive). Moreover, in many tasks, there is an intrinsic data imbalance that most learning-based methods struggle to cope with. It turns out that traffic sign detection is a problem in which these three issues are seen altogether. In this work, we propose a novel database generation method that requires only (i) arbitrary natural images, i.e., requires no real image from the domain of interest, and (ii) templates of the traffic signs, i.e., templates synthetically created to illustrate the appearance of the category of a traffic sign. The effortlessly generated training database is shown to be effective for the training of a deep detector (such as Faster R-CNN) on German traffic signs, achieving 95.66% of mAP on average. In addition, the proposed method is able to detect traffic signs with an average precision, recall and F1-score of about 94%, 91% and 93%, respectively. The experiments surprisingly show that detectors can be trained with simple data generation methods and without problem domain data for the background, which is in the opposite direction of the common sense for deep learning.
Lucas Tabelini Torres, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Alberto Ferreira de Souza, Claudine Badue, Nicu Sebe, Thiago Oliveira-Santos
IJCNN5
2019 Handling pedestrians in self-driving cars using image tracking and alternative path generation with Frenét frames
Renan Sarcinelli, Ranik Guidolini, Vinicius B. Cardoso, Thiago Meireles Paixão, Rodrigo Ferreira Berriel, Pedro Azevedo, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos
Comput. Graph.8
2018 Heading Direction Estimation Using Deep Learning with Automatic Large-scale Data Acquisition
abstract
Advanced Driver Assistance Systems (ADAS) have experienced major advances in the past few years. The main objective of ADAS includes keeping the vehicle in the correct road direction, and avoiding collision with other vehicles or obstacles around. In this paper, we address the problem of estimating the heading direction that keeps the vehicle aligned with the road direction. This information can be used in precise localization, road and lane keeping, lane departure warning, and others. To enable this approach, a large-scale database 1+ million images) was automatically acquired and annotated using publicly available platforms such as the Google Street View API and OpenStreetMap. After the acquisition of the database, a CNN model was trained to predict how much the heading direction of a car should change in order to align it to the road 4 meters ahead. To assess the performance of the model, experiments were performed using images from two different sources: a hidden test set from Google Street View (GSV) images and two datasets from our autonomous car (IARA). The model achieved a low mean average error of 2.359° and 2.524° for the GSV and IARA datasets, respectively; performing consistently across the different datasets. It is worth noting that the images from the IARA dataset are very different (camera, FOV, brightness, etc.) from the ones of the GSV dataset, which shows the robustness of the model. In conclusion, the model was trained effortlessly (using automatic processes) and showed promising results in real-world databases working in real-time (more than 75 frames per second).
Rodrigo Ferreira Berriel, Lucas Tabelini Torres, Vinicius B. Cardoso, Ranik Guidolini, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
IJCNN5
2018 Mapping Road Lanes Using Laser Remission and Deep Neural Networks
abstract
We propose the use of deep neural networks (DNN) for solving the problem of inferring the position and relevant properties of lanes of urban roads with poor or absent horizontal signalization, in order to allow the operation of autonomous cars in such situations. We take a segmentation approach to the problem and use the Efficient Neural Network (ENet) DNN for segmenting LiDAR remission grid maps into road maps. We represent road maps using what we called road grid maps. Road grid maps are square matrixes and each element of these matrixes represents a small square region of real-world space. The value of each element is a code associated with the semantics of the road map. Our road grid maps contain all information about the roads' lanes required for building the Road Definition Data Files (RDDFs) that are necessary for the operation of our autonomous car, IARA (Intelligent Autonomous Robotic Automobile). We have built a dataset of tens of kilometers of manually marked road lanes and used part of it to train ENet to segment road grid maps from remission grid maps. After being trained, ENet achieved an average segmentation accuracy of 83.7%. We have tested the use of inferred road grid maps in the real world using IARA on a stretch of 3.7 km of urban roads and it has shown performance equivalent to that of the previous IARA's subsystem that uses a manually generated RDDF.
Raphael V. Carneiro, Rafael Correia Nascimento, Ranik Guidolini, Vinicius B. Cardoso, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza
IJCNN6
2018 Visual Global Localization with a Hybrid WNN-CNN Approach
abstract
Currently, self-driving cars rely greatly on the Global Positioning System (GPS) infrastructure, albeit there is an increasing demand for alternative methods for GPS-denied environments. One of them is known as place recognition, which associates images of places with their corresponding positions. We previously proposed systems based on Weightless Neural Networks (WNN) to address this problem as a classification task. This encompasses solely one part of the global localization, which is not precise enough for driverless cars. Instead of just recognizing past places and outputting their poses, it is desired that a global localization system estimates the pose of current place images. In this paper, we propose to tackle this problem as follows. Firstly, given a live image, the place recognition system returns the most similar image and its pose. Then, given live and recollected images, a visual localization system outputs the relative camera pose represented by those images. To estimate the relative camera pose between the recollected and the current images, a Convolutional Neural Network (CNN) is trained with the two images as input and a relative pose vector as output. Together, these systems solve the global localization problem using the topological and metric information to approximate the current vehicle pose. The full approach is compared to a Real- Time Kinematic GPS system and a Simultaneous Localization and Mapping (SLAM) system. Experimental results show that the proposed approach correctly localizes a vehicle 90% of the time with a mean error of 1.20m compared to 1.12m of the SLAM system and 0.37m of the GPS, 89% of the time.
Avelino Forechi, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza
IJCNN3
2018 Handling Pedestrians in Crosswalks Using Deep Neural Networks in the IARA Autonomous Car
abstract
In this work, we propose a subsystem to handle pedestrians in crosswalks using deep neural networks for the IARA autonomous car, which relies on camera and LIDAR data fusion. Crosswalks' positions were manually annotated in IARA's map. Pedestrians are detected in the camera image using a convolutional neural network (CNN). Then, pedestrians' positions in the map are obtained by fusing their positions in the image with the LIDAR point cloud. Subsequently, if a pedestrian position is inside the crosswalk area, the crosswalk is set as busy. Finally, a busy crosswalk message is published to the High-Level Decision Maker subsystem. This subsystem selects the car's behavior according to the crosswalk condition and propagates this decision down through the control pipeline, in order to make the car drive correctly through the crosswalk area. The Pedestrian Handler subsystem was evaluated on IARA, which was driven autonomously for various laps along a real and complex circuit with various crosswalks. In all passages through crosswalks, the Pedestrian Handler dealt with pedestrians as expected, i.e., without any human intervention.
Ranik Guidolini, Lucas G. Scart, Luan F. R. Jesus, Vinicius B. Cardoso, Claudine Badue, Thiago Oliveira-Santos
IJCNN5
2018 Copycat CNN: Stealing Knowledge by Persuading Confession with Random Non-Labeled Data
abstract
In the past few years, Convolutional Neural Networks (CNNs) have been achieving state-of-the-art performance on a variety of problems. Many companies employ resources and money to generate these models and provide them as an API, therefore it is in their best interest to protect them, i.e., to avoid that someone else copy them. Recent studies revealed that stateof-the-art CNNs are vulnerable to adversarial examples attacks, and this weakness indicates that CNNs do not need to operate in the problem domain (PD). Therefore, we hypothesize that they also do not need to be trained with examples of the PD in order to operate in it. Given these facts, in this paper, we investigate if a target blackbox CNN can be copied by persuading it to confess its knowledge through random non-labeled data. The copy is two-fold: i) the target network is queried with random data and its predictions are used to create a fake dataset with the knowledge of the network; and ii) a copycat network is trained with the fake dataset and should be able to achieve similar performance as the target network. This hypothesis was evaluated locally in three problems (facial expression, object, and crosswalk classification) and against a cloud-based API. In the copy attacks, images from both nonproblem domain and PD were used. All copycat networks achieved at least 93.7% of the performance of the original models with non-problem domain data, and at least 98.6% using additional data from the PD. Additionally, the copycat CNN successfully copied at least 97.3% of the performance of the Microsoft Azure Emotion API. Our results show that it is possible to create a copycat CNN by simply querying a target network as black-box with random non-labeled data.
Jacson Rodrigues Correia da Silva, Rodrigo Ferreira Berriel, Claudine Badue, Alberto Ferreira de Souza, Thiago Oliveira-Santos
IJCNN3
2017 A Model-Predictive Motion Planner for the IARA autonomous car
abstract
We present the Model-Predictive Motion Planner (MPMP) of the Intelligent Autonomous Robotic Automobile (IARA). IARA is a fully autonomous car that uses a path planner to compute a path from its current position to the desired destination. Using this path, the current position, a goal in the path and a map, IARA's MPMP is able to compute smooth trajectories from its current position to the goal in less than 50 ms. MPMP computes the poses of these trajectories so that they follow the path closely and, at the same time, are at a safe distance of occasional obstacles. Our experiments have shown that MPMP is able to compute trajectories that precisely follow a path produced by a Human driver (distance of 0.15 m in average) while smoothly driving IARA at speeds of up to 32.4 km/h (9 m/s).
Vinicius B. Cardoso, Josias Oliveira, Thomas Teixeira, Claudine Badue, Filipe Wall Mutz, Thiago Oliveira-Santos, Lucas de Paula Veronese, Alberto Ferreira de Souza
ICRA4
2017 Neural-based model predictive control for tackling steering delays of autonomous cars
abstract
We propose a Neural Based Model Predictive Control (N-MPC) approach to tackle delays in the steering plant of autonomous cars. We examined the N-MPC approach as an alternative for the implementation of the Intelligent and Autonomous Robotic Automobile (IARA) steering control subsystem. For that, we compared the standard solution, based on the Proportional Integral Derivative (PID) control approach, with a N-MPC approach. For speeds of up to 25 km/h, the IARA's steering plant delay is not a problem for the PID control approach. However, in higher speeds, it causes large steering oscillations, which prevent proper operation. For this, we modeled the IARA's steering plant using a neural network and employed the neural model in the N-MPC. Our experimental results showed N-MPC can drastically reduce the impact of IARA's steering plant delays, which allowed its autonomous operation at speeds of up to 37 km/h.
Ranik Guidolini, Alberto Ferreira de Souza, Filipe Wall Mutz, Claudine Badue
IJCNN4
2017 Following the leader using a tracking system based on pre-trained deep neural networks
abstract
In this work, we present a software architecture to solve, at some level, the follow the leader problem. This problem consists of an autonomous vehicle trying to track and follow a leader vehicle. To track the leader position in consecutive camera images, we employed the Generic Object Tracking Using Regression Networks (GOTURN). GOTURN is a pre-trained Deep Neural Network capable of tracking generic objects, without application-specific training or fine-tuning. The proposed software architecture was evaluated using a real autonomous vehicle, in four stretches of a University ring road. In all experiments, the autonomous vehicle was able to follow the leader's path with maximum root mean square error of 0.28m.
Filipe Wall Mutz, Vinicius B. Cardoso, Thomas Teixeira, Luan F. R. Jesus, Michael André Golçalves, Ranik Guidolini, Josias Oliveira, Claudine Badue, Alberto Ferreira de Souza
IJCNN8
2016 Sequential appearance-based Global Localization using an ensemble of kNN-DTW classifiers
abstract
The human Episodic Memory system stores sequences of events (episodes encoded in time and in space) experienced or imagined by an individual for later access to the episodes in whole or in part. Such ability provided by our Episodic Memory system is important, among other functions, for our localization in space throughout time. Inspired by that, in this paper, it is presented a new Sequential Appearance-Based approach to the Global Localization problem, dubbed SABGL. SABGL is based on an ensemble of kNN classifiers, where each classifier uses the Dynamic Time Warping (DTW) and the Hamming distance to compare binary features extracted from sequences of images. SABGL is designed to solve the global localization problem in two phases: mapping and localization. During mapping, it is trained with a sequence of images and associated locations that represents episodes experienced by an autonomous robot. During localization, it receives subsequences of images of the same environment and compares them to its previous experienced episodes, trying to recollect the most similar “experience” in time and space at once. Then, the system outputs the positions where it “believes” these images were captured. SABGL is compared with a previous single-image approach, named VibGL. Experimental results show that SABGL consistently outperforms VibGL in classification accuracy with higher precision in all maximum distance tolerance analyzed. For instance, given a maximum tolerance of 10m, SABGL is able to correctly localize an autonomous car 98% of the time.
Avelino Forechi, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos
IJCNN3
2016 Simulating robotic cars using time-delay neural networks
abstract
In this paper, we propose a simulator for robotic cars based on two time-delay neural networks. These networks are intended to simulate the mechanisms that govern how a set of effort commands changes the car's velocity and the direction it is moving. The first neural network receives as input a temporal sequence of current and previous throttle and brake efforts, along with a temporal sequence of the previous car's velocities (estimated by the network), and outputs the velocity that the real car would reach in the next time interval given these inputs. The second neural network estimates the arctangent of curvature (a variable related to the steering wheel angle) that a real car would reach in the next time interval given a temporal sequence of current and previous steering efforts and previous arctangents of curvatures of the car estimated by the network. We evaluated the performance of our simulator using real-world datasets acquired using an autonomous robotic car. Experimental results showed that our simulator was able to simulate in real time how a set of efforts influences the car's velocity and arctangent of curvature. While navigating in a map of a real-world environment, our car simulator was able to emulate the velocity and arctangent of curvature of the real car with mean squared error of 2.2×10-3(m/s)2and 4.0×10-5rad2, respectively.
Alberto Ferreira de Souza, Jacson Rodrigues Correia da Silva, Filipe Wall Mutz, Claudine Badue, Thiago Oliveira-Santos
IJCNN4
2016 Fat-Fast VG-RAM WNN: A high performance approach
Avelino Forechi, Alberto Ferreira de Souza, Jorcy de Oliveira Neto, Edilson de Aguiar, Claudine Badue, Artur S. d'Avila Garcez, Thiago Oliveira-Santos
Neurocomputing5
2015 Image-based mapping, global localization and position tracking using VG-RAM weightless neural networks
abstract
Humans can easily memorize images of places and labels (road names, addresses, etc.) associated with them, as well as trajectories defined by sequences of images and corresponding positions. Later, they are able to remember places' labels and relative positions when seeing the same images again. In this work, we present an image-based mapping, global localization and position tracking system based on Virtual Generalizing Random Access Memory (VG-RAM) weightless neural networks, dubbed VIBML. VIBML mimics humans ability of learning about a place and of recognizing the same place in a later moment, as well as of tracking self-movement through the environment using images. We evaluated the performance of VIBML on the precise localization of an autonomous car using real-world datasets. Our experimental results showed that VIBML is able to localize car-like robots on large maps of real world environments with accuracy equivalent to that of state-of-the-art methods - VIBML is able to localize an autonomous car with average positioning error of 1.12m and with 75% of the poses with error below 1.5m in a 3.75km path around the main campus of the Federal University of Espírito Santo.
Lauro Jose Lyrio Junior, Thiago Oliveira-Santos, Claudine Badue, Alberto Ferreira de Souza
ICRA3
2014 Compressing VG-RAM WNN memory for lightweight applications
abstract
The Virtual Generalizing Random Access Memory Weightless Neural Network (VG-RAM WNN) is an effective machine learning technique that offers simple implementation and fast training. One disadvantage of VG-RAM WNN, however, is the test time for applications with many training samples, i.e. large multi-class classification applications. In such cases, the test time tends to be high, since it increases with the size of the memory of each neuron. In this paper, we present a new methodology for handling such applications using VG-RAM WNN. By employing data clustering techniques to reduce the overall size of the neurons' memory, we were able to reduce the network's memory footprint and the system's runtime, while maintaining a high and acceptable classification performance. We evaluated the performance of our VG-RAM WNN system with compressed memory on the problem of traffic sign recognition. Our experimental results showed that, after compression, the system was able to run at very fast response times in standard computers. Also, we were able to load and run the system at interactive rates in small low-power systems, experiencing only a small reduction in classification performance.
Edilson de Aguiar, Avelino Forechi, Lucas de Paula Veronese, Mariella Berger, Alberto Ferreira de Souza, Claudine Badue, Thiago Oliveira-Santos
IJCNN6
2014 Image-based global localization using VG-RAM Weightless Neural Networks
abstract
Mapping and localization are fundamental problems in autonomous robotics. Autonomous robots need to know where they are in their area of operation to navigate through it and to perform activities of interest. In this paper, we propose an Image-Based Global Localization (VibGL) system that uses Virtual Generalizing Random Access Memory Weightless Neural Networks (VG-RAM WNN). For mapping, we employ a VG-RAM WNN that learns the world positions associated with the images captured along a trajectory. During the localization, new images from the trajectory are presented to the VG-RAM WNN, which outputs their positions in the world. We performed experiments with our VibGL system applied to the problem of localizing an autonomous car. Our experimental results show that the system is able to learn large maps (several kilometers in length) of real world environments and perform global localization with median pose precision of about 3m. Considering a tolerance of 10m VibGL is able to localize the car 95% of the time.
Lauro Jose Lyrio Junior, Thiago Oliveira-Santos, Avelino Forechi, Lucas de Paula Veronese, Claudine Badue, Alberto Ferreira de Souza
IJCNN5
2014 Programming a VG-RAM based Neural Network Computer
abstract
We propose a Virtual Generalizing Random Access Memory (VG-RAM) Weightless Neural Network (WNN) Computer (V'Ger Computer for short). VG-RAM WNNs are very effective pattern recognition tools, offering fast training (one shot training) and competitive recognition performance, if compared with other current techniques. The V'Ger Computer architecture was inspired on the organization of the human neocortex and is composed of hierarchically organized and recurrently interconnected layers of VG-RAM WNN neurons. One layer is connected to another in a way similar to cortico-cortical feed-forward and feedback connections between functionally adjacent and hierarchically organized areas. We have "programmed" the V'Ger Computer for counting from 0 to 9 three times. Our preliminary experimental results showed that V'Ger is capable of executing this sequence of actions in spite of strong interferences.
Alberto Ferreira de Souza, Avelino Forechi, Filipe Wall Mutz, Mariella Berger, Thiago Oliveira-Santos, Claudine Badue
IJCNN6
2013 Real-time road surface mapping using stereo matching, v-disparity and machine learning
abstract
We present and evaluate a computer vision approach for real-time mapping of traversable road surfaces ahead of an autonomous vehicle that relies only on a stereo camera. Our system first determines the camera position with respect to the ground plane using stereo vision algorithms and probabilistic methods, and then reprojects the camera raw image to a bidimensional grid map that represents the ground plane in world coordinates. After that, it generates a road surface grid map from the bidimensional grid map using an online trained pixel classifier based on mixture of Gaussians. Finally, to build a high quality map, each road surface grid map is integrated to a probabilistic bidimensional grid map using a binary Bayes filter for estimating the occupancy probability of each grid cell. We evaluated the performance of our approach for road surface mapping in comparison to manually classified images. Our experimental results show that our approach is able to correctly map regions at 50 m ahead of an autonomous vehicle, with True Positive Rate (TPR) of 90.32% for regions between 20 and 35 m ahead and False Positive Rate (FPR) not superior to 4.23% for any range.
Vitor B. Azevedo, Alberto Ferreira de Souza, Lucas de Paula Veronese, Claudine Badue, Mariella Berger
IJCNN4
2013 Traffic sign detection with VG-RAM weightless neural networks
abstract
We present a biologically inspired approach to traffic sign detection based on Virtual Generalizing Random Access Memory Weightless Neural Networks (VG-RAM WNN). VG-RAM WNN are effective machine learning tools that offer simple implementation and fast training and test. Our VG-RAM WNN architecture models the saccadic eye movement system and the transformations suffered by the images captured by the eyes from the retina to the superior colliculus in the mammalian brain. We evaluated the performance of our VG-RAM WNN system on traffic sign detection using the German Traffic Sign Detection Benchmark (GTSDB). Using only 12 traffic sign images for training, our system was ranked between the first 16 methods for the prohibitory category in the German Traffic Sign Detection Competition, part of the IJCNN'2013. Our experimental results showed that our approach is capable of reliably and efficiently detect a large variety of traffic sign categories using a few training samples.
Alberto Ferreira de Souza, Cayo Fontana, Filipe Wall Mutz, Tiago Alves de Oliveira, Mariella Berger, Avelino Forechi, Jorcy de Oliveira Neto, Edilson de Aguiar, Claudine Badue
IJCNN9
2012 Traffic sign recognition with VG-RAM Weightless Neural Networks
abstract
Virtual Generalizing Random Access Memory Weightless Neural Networks (VG-RAM WNN) is an effective machine learning technique that offers simple implementation and fast training and test. In this paper, we present a new approach for traffic sign recognition based on VG-RAM WNN. We evaluate its performance using the German Traffic Sign Recognition Benchmark (GTSRB), a large multi-class classification benchmark. Our experimental results showed that our VG-RAM WNN architecture for traffic sign recognition was able to rank at 4th position in the GTSRB evaluation system, with a recognition rate of 98.73%, and was overcome by only one automatic approach.
Mariella Berger, Avelino Forechi, Alberto Ferreira de Souza, Jorcy de Oliveira Neto, Lucas de Paula Veronese, Claudine Badue
ISDA6
2010 The Dynamic Block Remapping Cache
abstract
In this paper we present a new architecture of Level 2 (L2) cache – the Dynamic Block Remapping Cache (DBRC). DBRC mimics important characteristics of virtual memory systems to reduce the impact of L2 in system performance. Similar to virtual memory systems, the DBRC uses a hierarchy of tables to map blocks of L2 cache into blocks of physical memory. It also uses a Block-TLB to speedup accesses to previously performed block translations. We verified that the benefits of full associativity and the consequent possibility of employment of global block replacement algorithms allow hit rates higher than those of equivalent standard caches. We compare DBRC with standard caches in terms of miss rate, energy consumption and impact on the instruction-level parallelism (ILP) of a simulated superscalar processor. Our results show that DBRC outperforms standard caches in terms of miss rate, energy consumption and impact on ILP.
Felipe Pedroni, Alberto Ferreira de Souza, Claudine Badue
SBAC-PAD3
2009 Automated multi-label text categorization with VG-RAM weightless neural networks
Alberto Ferreira de Souza, Felipe Pedroni, Elias de Oliveira, Patrick Marques Ciarelli, Wallace Favoreto Henrique, Lucas de Paula Veronese, Claudine Badue
Neurocomputing7
2008 Face Recognition with VG-RAM Weightless Neural Networks
Alberto Ferreira de Souza, Claudine Badue, Felipe Pedroni, Elias de Oliveira, Stiven S. Dias, Hallysson Oliveira, Sotério Ferreira de Souza
ICANN (1)2
2007 Analyzing imbalance among homogeneous index servers in a web search system
Claudine Badue, Ricardo Baeza-Yates, Berthier A. Ribeiro-Neto, Artur Ziviani, Nivio Ziviani
Inf. Process. Manag.1
2006 Modeling performance-driven workload characterization of web search systems
abstract
No abstract available.
Claudine Badue, Ricardo Baeza-Yates, Berthier A. Ribeiro-Neto, Artur Ziviani, Nivio Ziviani
CIKM1
2005 Basic issues on the processing of web queries
abstract
In this paper we study three basic and key issues related to Web query processing: load balance, broker behavior, and performance by individual index servers. Our study, while preliminary, does reveal interesting tradeoffs: (1) load unbalance at low query arrival rates can be controlled with a simple measure of randomizing the distribution of documents among the index servers, (2) the broker is not a bottleneck, and (3) disk utilization is higher than CPU utilization.
Claudine Badue, Ramurti A. Barbosa, Paulo Braz Golgher, Berthier A. Ribeiro-Neto, Nivio Ziviani
SIGIR1
2001 Distributed Query Processing Using Partitioned Inverted Files
abstract
In this paper, we study query processing in a distributed text database. The novelty is a real distributed architecture implementation that offers concurrent query service. The distributed system adopts a network of workstations model and the client-server paradigm. The document collection is indexed with an inverted file. We adopt two distinct strategies of index partitioning in the distributed system, namely local index partitioning and global index partitioning. In both strategies, documents are ranked using the vector space model along with a document filtering technique for fast ranking. We evaluate and compare the impact of the two index partitioning strategies on query processing performance. Experimental results on retrieval efficiency show that, within our framework, the global index partitioning outperforms the local index partitioning. 1.
Claudine Badue, Ricardo Baeza-Yates, Berthier A. Ribeiro-Neto, Nivio Ziviani
SPIRE1