VLDB 2026 Research / reviewers in the wild / expert
Luiz Eduardo Soares de Oliveira
dblp:o/LuizSOliveira · also Luiz E. S. Oliveira, Luiz E. Soares de Oliveira, Luiz S. Oliveira
· DBLP profile ↗
139ranked-venue papers
12as first author
26since 2021 · last 2025
0000-0002-0595-5370ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 103 · 12 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 16 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 6 since 2021Databases, data management, data science and information retrieval · 14 · 4 first-author · 3 since 2021Security and privacy · 5 · 1 since 2021Computer networks · 3Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Wrecking SQL Incremental Validation Methodology
Ruanitto Docini, Eduardo C. de Almeida, Luiz Eduardo Soares de Oliveira |
DEXA (2) | 3 |
| 2025 | Predicting Heart Failure Hospitalizations with LLMs from Health Insurance DataabstractHeart failure (HF) represents a global clinical and economic challenge, with hospitalizations accounting for 65% of disease-related costs. This study proposes an approach to predict HF hospitalizations using Large Language Models (LLMs) trained on chronological data from Brazilian health insurance beneficiaries. By converting administrative records (consultations, medications, diagnoses) into temporal narratives, models like RoBERTa and Open-Cabrita3B were fine-tuned to identify clinical deterioration patterns. The HealthHistoryRoBERTapt model, trained with historical health insurance data and specifically adjusted for HF, achieved an AUC-ROC of 0.93-0.95 in prediction windows from 5 to 180 days, significantly outperforming other studies (AUC 0.63-0.76) using static clinical data or basic demographics, and those combining clinical-administrative data (AUC 0.82). It is noteworthy that the ability of the model to maintain an F1-score greater than 0.85 and sensitivity of 0.87 in predictions of up to 150 days, revealing that administrative variables (e.g., history of hospitalizations, frequency of consultations) function as effective proxies for socioeconomic and behavioral factors, traditionally neglected. Compared to other works, this study demonstrates that longitudinal health insurance data combined with NLP techniques capture non-linear risk trajectories, enabling precise predictions for strategic health planning. Everton F. Baro, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr. |
SMC | 2 |
| 2025 | Drift-Aware Machine Learning for Operational State Classification in Biogas Dry ReformingabstractThis paper presents a machine learning-based approach for classifying the operational states in the biogas Dry Reforming (DR) reactor, focusing on catalyst activation, reaction, and irregularity detection. A key challenge in DR processes is the formation of coke, which can lead to reactor clogging. To address this, we propose incorporating the virtual drifts number, which are changes in the input data distribution, as an additional feature to enhance model performance. Different drift detection algorithms and classifiers were evaluated on a dataset comprising nine distinct DR reactions. Experimental results demonstrate that integrating virtual drift counts improves the average accuracy from 84.74% to 88.01% and the average F1 scores from 81.25% to 85.39% across all models, with RF achieving the highest performance (accuracy from 88.40% to 92.35% and F1 score from 85.95% to 91.59%). Our results highlight the potential of drift-aware features for real-time monitoring and fault detection in DR systems, offering a scalable solution to optimize reactor operations. Marcos A. Schreiner, Renan Akira Escribano, Heitor Murilo Gomes, Paulo R. L. de Almeida, Luiz Eduardo Soares de Oliveira |
SMC | 5 |
| 2025 | Improving open set recognition with dissimilarity-based metric learning
Lucas O. Teixeira, Diego Bertolini, Luiz Eduardo Soares de Oliveira, George D. C. Cavalcanti, Yandre M. G. Costa |
Knowl. Based Syst. | 3 |
| 2025 | A database for automatic identification of herbarium specimens in Piperaceae family
Alexandre Yuji Kajihara, George Azevedo de Queiroz, Marcelo Galeazzi Caxambú, Luiz Eduardo Soares de Oliveira, Diego Bertolini, André Luís Schwerz |
Multim. Tools Appl. | 4 |
| 2025 | Triplet dissimilarity: a texture classification approach using dissimilarity and siamese networks
Lucas O. Teixeira, Diego Bertolini, Luiz Eduardo Soares de Oliveira, George D. C. Cavalcanti, Yandre M. G. Costa |
Soft Comput. | 3 |
| 2024 | Optimizing Parking Space Classification: Distilling Ensembles into Lightweight ClassifiersabstractWhen deploying large-scale machine learning models for smart city applications, such as image-based parking lot monitoring, data often must be sent to a central server to perform classification tasks. This is challenging for the city's infrastructure, where image-based applications require transmitting large volumes of data, necessitating complex network and hardware infrastructures to process the data. To address this issue in image-based parking space classification, we propose creating a robust ensemble of classifiers to serve as Teacher models. These Teacher models are distilled into lightweight and specialized Student models that can be deployed directly on edge devices. The knowledge is distilled to the Student models through pseudo-labeled samples generated by the Teacher model, which are utilized to fine-tune the Student models on the target scenario. Our results show that the Student models, with 26 times fewer parameters than the Teacher models, achieved an average accuracy of 96.6 % on the target test datasets, surpassing the Teacher models, which attained an average accuracy of 95.3 %. Paulo Luza Alves, Andre G. Hochuli, Luiz Eduardo Soares de Oliveira, Paulo R. L. Almeida |
ICMLA | 3 |
| 2024 | Using Deep Neural Networks to Quantify Parking Dwell TimeabstractIn smart cities, it is common practice to define a maximum length of stay for a given parking space to increase the space's rotativity and discourage the usage of individual transportation solutions. However, automatically determining individual car dwell times from images faces challenges, such as images collected from low-resolution cameras, lighting variations, and weather effects. In this work, we propose a method that combines two deep neural networks to compute the dwell time of each car in a parking lot. The proposed method first defines the parking space status between occupied and empty using a deep classification network. Then, it uses a Siamese network to check if the parked car is the same as the previous image. Using an experimental protocol that focuses on a cross-dataset scenario, we show that if a perfect classifier is used, the proposed system generates 75% of perfect dwell time predictions, where the predicted value matched exactly the time the car stayed parked. Nevertheless, our experiments show a drop in prediction quality when a real-world classifier is used to predict the parking space statuses, reaching 49% of perfect predictions, showing that the proposed Siamese network is promising but impacted by the quality of the classifier used at the beginning of the pipeline. Marcelo Eduardo Marques Ribas, Heloisa Benedet Mendes, Luiz Eduardo Soares de Oliveira, Luiz Antonio Zanlorensi, Paulo R. L. Almeida |
ICMLA | 3 |
| 2024 | Fault distance estimation for transmission lines with dynamic regressor selection
Leandro Augusto Ensina, Luiz Eduardo Soares de Oliveira, Rafael M. O. Cruz, George D. C. Cavalcanti |
Neural Comput. Appl. | 2 |
| 2024 | Contrastive dissimilarity: optimizing performance on imbalanced and limited data sets
Lucas O. Teixeira, Diego Bertolini, Luiz Eduardo Soares de Oliveira, George D. C. Cavalcanti, Yandre M. G. Costa |
Neural Comput. Appl. | 3 |
| 2023 | Vehicle Occurrence-Based Parking Space DetectionabstractSmart-parking solutions use sensors, cameras, and data analysis to improve parking efficiency and reduce traffic congestion. Computer vision-based methods have been used extensively in recent years to tackle the problem of parking lot management, but most of the works assume that the parking spots are manually labeled, impacting the cost and feasibility of deployment. To fill this gap, this work presents an automatic parking space detection method, which receives a sequence of images of a parking lot and returns a list of coordinates identifying the detected parking spaces. The proposed method employs instance segmentation to identify cars and, using vehicle occurrence, generate a heat map of parking spaces. The results using twelve different subsets from the PKLot and CNRPark-EXT parking lot datasets show that the method achieved an AP25 score up to 95.60% and AP50 score up to 79.90%. Paulo R. L. Almeida, Jeovane Honório Alves, Luiz Eduardo Soares de Oliveira, Andre G. Hochuli, João V. Fröhlich, Rodrigo A. Krauel |
SMC | 3 |
| 2023 | A Small Claims Court for the NLP: Judging Legal Text Classification Strategies With Small DatasetsabstractRecent advances in language modelling has significantly decreased the need of labelled data in text classification tasks. Transformer-based models, pre-trained on unlabeled data, can outmatch the performance of models trained from scratch for each task. However, the amount of labelled data need to fine-tune such type of model is still considerably high for domains requiring expert-level annotators, like the legal domain. This paper investigates the best strategies for optimizing the use of a small labeled dataset and large amounts of unlabeled data and perform a classification task in the legal area with 50 predefined topics. More specifically, we use the records of demands to a Brazilian Public Prosecutor's Office aiming to assign the descriptions in one of the subjects, which currently demands deep legal knowledge for manual filling. The task of optimizing the performance of classifiers in this scenario is especially challenging, given the low amount of resources available regarding the Portuguese language, especially in the legal domain. Our results demonstrate that classic supervised models such as logistic regression and SVM and the ensembles random forest and gradient boosting achieve better performance along with embeddings extracted with word2vec when compared to BERT language model. The latter demonstrates superior performance in association with the architecture of the model itself as a classifier, having surpassed all previous models in that regard. The best result was obtained with Unsupervised Data Augmentation (UDA), which jointly uses BERT, data augmentation, and strategies of semi-supervised learning, with an accuracy of 80.7% in the aforementioned task. Mariana Y. Noguti, Eduardo Vellasques, Luiz Eduardo Soares de Oliveira |
SMC | 3 |
| 2023 | Fast & Furious: On the modelling of malware detection as an evolving data streamabstractMalware is a major threat to computer systems and imposes many challenges to cyber security. Targeted threats, such as ransomware, cause millions of dollars in losses every year. The constant increase of malware infections has been motivating popular antiviruses (AVs) to develop dedicated detection strategies, which include meticulously crafted machine learning (ML) pipelines. However, malware developers unceasingly change their samples' features to bypass detection. This constant evolution of malware samples causes changes to the data distribution (i.e., concept drifts) that directly affect ML model detection rates, something not considered in the majority of the literature work. In this work, we evaluate the impact of concept drift on malware classifiers for two Android datasets: DREBIN (about 130K apps) and a subset of AndroZoo (about 285K apps). We used these datasets to train an Adaptive Random Forest (ARF) classifier, as well as a Stochastic Gradient Descent (SGD) classifier. We also ordered all datasets samples using their VirusTotal submission timestamp and then extracted features from their textual attributes using two algorithms (Word2Vec and TF-IDF). Then, we conducted experiments comparing both feature extractors, classifiers, as well as four drift detectors (DDM, EDDM, ADWIN, and KSWIN) to determine the best approach for real environments. Finally, we compare some possible approaches to mitigate concept drift and propose a novel data stream pipeline that updates both the classifier and the feature extractor. To do so, we conducted a longitudinal evaluation by (i) classifying malware samples collected over nine years (2009-2018), (ii) reviewing concept drift detection algorithms to attest its pervasiveness, (iii) comparing distinct ML approaches to mitigate the issue, and (iv) proposing an ML data stream pipeline that outperformed literature approaches. Fabricio Ceschin, Marcus Botacin, Heitor Murilo Gomes, Felipe Azevedo Pinage, Luiz Eduardo Soares de Oliveira, André Ricardo Abed Grégio |
Expert Syst. Appl. | 5 |
| 2023 | Large-margin representation learning for texture classification
Jonathan de Matos, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Alessandro L. Koerich |
Pattern Recognit. Lett. | 2 |
| 2022 | Fault Classification in Transmission Lines with Generalization CompetenceabstractTransmission lines are crucial components of the electric power system and are exposed to several conditions that can disrupt the transmission of electrical power. In this scenario, a protection system must detect and classify a fault, for example, to enable the quick repair and restoration of a faulty line. This paper presents a method for fault type classification with two main characteristics: (1) independence of the sampling rate of the protection system; and (2) capacity to classify failures for different transmission lines than the one used to train the algorithm (generalization competence). Initially, the first post-fault cycle of the three-phase current signal is used to get the groundDetection feature, which aims to indicate the action or not of the ground in the failure. Next, maximum and minimum values are obtained from a pre-fault cycle of the current waveform to normalize the post-fault cycle by the MinMax technique, separately for each phase and individually for each example. Lastly, we get energy-based attributes together with maximum and minimum values from these normalized cycles for each phase to create the feature vector along with the groundDetection attribute. The extracted features are then used as input to the Random Forest algorithm to predict the fault type. The results demonstrate average accuracy higher than 99% for diversified simulated events for both characteristics previously mentioned. Our method also manifested the capacity to classify real fault events even being trained with synthetic examples with an accuracy of 96%. Leandro Augusto Ensina, Luiz Eduardo Soares de Oliveira, Eduardo C. de Almeida, Signie Laureano França Santos, Leandro Silva Bernardino |
IECON | 2 |
| 2022 | Assessing Batch and Online Learning for Delivery in Full and On Time PredictionsabstractImproving results by optimizing process execution is one objective of major companies. For these corporations, the main point for achieving better results is the good maintenance of supply chain management. The most important supply chain metric is Delivery in Full and On Time (DIFOT). DIFOT measures how well a supply chain delivers value to the customer. In this work, we bring forward an analysis of DIFOT prediction from large Brazilian food company. More specifically, we compare a batch and online learning algorithm for DIFOT prediction and depict why the latter is suitable for this problem. Furthermore, we report a feature drift analysis to identify whether there are considerable shifts along with the dataset timespan. As a byproduct of this research, we make the dataset used in this analysis publicly available for future research in DIFOT prediction. Adriano Alves de Lima, Márcie Venâncio Batista, Jean Paul Barddal, Danilo Sipoli Sanches, Luiz Eduardo Soares de Oliveira |
IJCNN | 5 |
| 2022 | Predicting Hospitalization from Health Insurance DataabstractHospitalizations represent an expressive part of total health costs and, therefore, reducing the number of hospitalizations, when possible, can generate both economic gains and enhanced quality of life of patients. Several works have been striving to use machine learning to create models for hospitalization predictions. Most of them require specialized knowledge in the health area, mainly in the stages of data preparation and selection of features. This feature engineering is not always perfect and may fail to select relevant features for the model training process. In this paper, to fill this gap, we explore three sources of information to extract features, i.e., medical specialty, event description, and the International Classification of Diseases. In addition, we introduce a dataset composed of 38,524 records of medical events from 34,930 patients. To assess and set a baseline for this new dataset, we have used two well-known ensemble methods (Random Forest and Gradient Boosting). The best results, AUC = 0.82, were achieved by combining the models generated from the three feature set tested and gradient boosting. We believe that researchers will find this dataset a valuable tool in their work on hospitalization prediction. It will also make future benchmarking and evaluation possible. Everton F. Baro, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr. |
SMC | 2 |
| 2022 | Two-view fine-grained classification of plant species
Voncarlos Araujo, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich |
Neurocomputing | 3 |
| 2021 | Data Augmentation for Writer Identification Using a Cognitive Inspired Model
Fabio Pignelli, Yandre M. G. Costa, Luiz Eduardo Soares de Oliveira, Diego Bertolini |
ICDAR (4) | 3 |
| 2021 | On the Evaluation of Competence Measures for Time Series ForecastingabstractDynamic selection systems work by selecting the most competent models from an ensemble. The key issue in these systems is to define the competence of the models. The models’ accuracy is commonly used to select the models or define the weights to be used in their combination. This competence is calculated using the feature space region, known as the region of competence, around the test pattern. The literature of dynamic classifier systems presents a good variety of competence measures, but some are not suitable for time series forecasting. However, some dynamic regression selection works present measures that can be used with time series forecasting problems. Such measures are extracted from the region of competence, and this work aims to evaluate these competence measures to calculate the weights of the models to combine them dynamically. The experiments are performed using three different machine learning algorithms with ten time series datasets and compare dynamic weighting algorithm, single model, and classic statistical combination techniques as Mean and Median. Our results show that the models’ combination performs better than the single model, Mean and Median, but the competence measure’s choice is time series and model-dependent. Thiago J. M. Moura, George D. C. Cavalcanti, Luiz Eduardo Soares de Oliveira |
SMC | 3 |
| 2021 | A comprehensive comparison of end-to-end approaches for handwritten digit string recognition
Andre G. Hochuli, Alceu S. Britto Jr., David A. Saji, José M. Saavedra, Robert Sabourin, Luiz Eduardo Soares de Oliveira |
Expert Syst. Appl. | 6 |
| 2021 | MINE: A framework for dynamic regressor selection
Thiago J. M. Moura, George D. C. Cavalcanti, Luiz Eduardo Soares de Oliveira |
Inf. Sci. | 3 |
| 2021 | Dynamic selection and combination of one-class classifiers for multi-class classification
Rogerio C. P. Fragoso, George D. C. Cavalcanti, Roberto H. W. Pinheiro, Luiz Eduardo Soares de Oliveira |
Knowl. Based Syst. | 4 |
| 2021 | Automatic chronic degenerative diseases identification using enteric nervous system images
Gustavo Zanoni Felipe, Jacqueline Nelisis Zanoni, Camila C. Sehaber-Sierakowski, Gleison D. P. Bossolani, Sara R. G. Souza, Franklin César Flores, Luiz Eduardo Soares de Oliveira, Rodolfo Miranda Pereira, Yandre M. G. Costa |
Neural Comput. Appl. | 7 |
| 2021 | A database for automatic classification of gender in Araucaria angustifolia plants
Jefferson G. Martins, Luiz Eduardo Soares de Oliveira, Daniel Weingaertner, Andersson Barison, Gerlon A. R. Oliveira, Luciano M. Lião |
Soft Comput. | 2 |
| 2021 | Intrapersonal Parameter Optimization for Offline Handwritten Signature AugmentationabstractUsually, in a real-world scenario, few signature samples are available to train an automatic signature verification system (ASVS). However, such systems do indeed need a lot of signatures to achieve an acceptable performance. Neuromotor signature duplication methods and feature space augmentation methods may be used to meet the need for an increase in the number of samples. Such techniques manually or empirically define a set of parameters to introduce a degree of writer variability. Therefore, in the present study, a method to automatically model the most common writer variability traits is proposed. The method is used to generate offline signatures in the image and the feature space and train an ASVS. We also introduce an alternative approach to evaluate the quality of samples considering their feature vectors. We evaluated the performance of an ASVS with the generated samples using three well-known offline signature datasets: GPDS, MCYT-75, and CEDAR. In GPDS-300, when the SVM classifier was trained using one genuine signature per writer and the duplicates generated in the image space, the Equal Error Rate (EER) decreased from 5.71% to 1.08%. Under the same conditions, the EER decreased to 1.04% using the feature space augmentation technique. We also verified that the model that generates duplicates in the image space reproduces the most common writer variability traits in the three different datasets. Teruo M. Maruyama, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | A Framework for Analyzing the Impact of Missing Data in Predictive ModelsabstractWe propose a stochastic framework to evaluate the impact of missing data on the performance of predictive models. The framework allows full control of important aspects of the data set structure. These include the number and type of the input variables, the correlation between the input variables and their general predictive power, and sample size. The missing process is generated from a multivariate Bernoulli distribution, which allows us to simulate missing patterns corresponding to the MCAR, MAR and MNAR mechanisms. Although the framework may be applied to virtually all types of predictive models, in this article, we focus on the logistic regression model and choose the accuracy as the predictive measure. The simulation results show that the effects of missing data disappear for large sample sizes, as expected. On the other hand, as the number of input variables increases, the accuracy decreases mainly for binary inputs. Fabiola Santore, Eduardo C. de Almeida, Wagner Hugo Bonat, Eduardo H. M. Pena, Luiz Eduardo Soares de Oliveira |
CIKM | 5 |
| 2020 | Classifier Pool Generation based on a Two-level Diversity ApproachabstractThis paper describes a classifier pool generation method guided by the diversity estimated on the data complexity and classifier decisions. First, the behavior of complexity measures is assessed by considering several subsamples of the dataset. The complexity measures with high variability across the subsamples are selected for posterior pool adaptation, where an evolutionary algorithm optimizes diversity in both complexity and decision spaces. A robust experimental protocol with 28 datasets and 20 replications is used to evaluate the proposed method. Results show significant accuracy improvements in 69.4% of the experiments when Dynamic Classifier Selection and Dynamic Ensemble Selection methods are applied. Marcos Monteiro 0001, Alceu S. Britto Jr., Jean Paul Barddal, Luiz Eduardo Soares de Oliveira, Robert Sabourin |
ICPR | 4 |
| 2020 | Data Augmentation for Histopathological Images Based on Gaussian-Laplacian Pyramid BlendingabstractData imbalance is a major problem that affects several machine learning (ML) algorithms. Such a problem is troublesome because most of the ML algorithms attempt to optimize a loss function that does not take into account the data imbalance. Accordingly, the ML algorithm simply generates a trivial model that is biased toward predicting the most frequent class in the training data. In the case of histopathologic images (HIs), both low-level and high-level data augmentation (DA) techniques still present performance issues when applied in the presence of inter-patient variability; whence the model tends to learn color representations, which is related to the staining process. In this paper, we propose a novel approach capable of not only augmenting HI dataset but also distributing the inter-patient variability by means of image blending using the Gaussian-Laplacian pyramid. The proposed approach consists of finding the Gaussian pyramids of two images of different patients and finding the Laplacian pyramids thereof. Afterwards, the left-half side and the right-half side of different HIs are joined in each level of the Laplacian pyramid, and from the joint pyramids, the original image is reconstructed. This composition combines the stain variation of two patients, avoiding that color differences mislead the learning process. Experimental results on the BreakHis dataset have shown promising gains vis-à-vis the majority of DA techniques presented in the literature. Steve Ataky, Jonathan de Matos, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich |
IJCNN | 4 |
| 2020 | An End-to-End Approach for Recognition of Modern and Historical Handwritten Numeral StringsabstractAn end-to-end solution for handwritten numeral string recognition is proposed, in which the numeral string is considered as composed of objects automatically detected and recognized by a YoLo-based model. The main contribution of this paper is to avoid heuristic-based methods for string preprocessing and segmentation, the need for task-oriented classifiers, and also the use of specific constraints related to the string length. A robust experimental protocol based on several numeral string datasets, including one composed of historical documents, has shown that the proposed method is a feasible end-to-end solution for numeral string recognition. Besides, it reduces the complexity of the string recognition task considerably since it drops out classical steps, in special preprocessing, segmentation, and a set of classifiers devoted to strings with a specific length. Andre G. Hochuli, Alceu S. Britto Jr., Jean Paul Barddal, Robert Sabourin, Luiz Eduardo Soares de Oliveira |
IJCNN | 5 |
| 2020 | Legal Document Classification: An Application to Law Area Prediction of Petitions to Public Prosecution ServiceabstractIn recent years, there has been an increased interest in the application of Natural Language Processing (NLP) to legal documents. The use of convolutional and recurrent neural networks along with word embedding techniques have presented promising results when applied to textual classification problems, such as sentiment analysis and topic segmentation of documents. This paper proposes the use of NLP techniques for textual classification, with the purpose of categorizing the descriptions of the services provided by the Public Prosecutor's Office of the State of Paraná to the population in one of the areas of law covered by the institution. Our main goal is to automate the process of assigning petitions to their respective areas of law, with a consequent reduction in costs and time associated with such process while allowing the allocation of human resources to more complex tasks. In this paper, we compare different approaches to word representations in the aforementioned task: including document-term matrices and a few different word embeddings. With regards to the classification models, we evaluated three different families: linear models, boosted trees and neural networks. The best results were obtained with a combination of Word2Vec trained on a domain-specific corpus and a Recurrent Neural Network (RNN) architecture (more specifically, LSTM), leading to an accuracy of 90% and F1-Score of 85% in the classification of eighteen categories (law areas). Mariana Y. Noguti, Eduardo Vellasques, Luiz Eduardo Soares de Oliveira |
IJCNN | 3 |
| 2020 | Naïve Approaches to Deal With Concept DriftsabstractA common problem in machine learning is to find representative real-world labeled datasets to put the methods to test. When developing approaches to deal with concept drifts, some datasets such as the Forest Covertype and Nebraska Weather are common choices for testing, even though there is no consensus on whether these exhibit concept drifts or not. We argue that some well-known real-world concept drift datasets present a high serial dependence in the target class and may have only minor changes. With this in mind, we propose the use of Naïve methods that should be used for comparison with methods that deal with concept drifts. The experimental results using six real-world well-known concept drift datasets show that the Naïve approaches can be better than some methods to deal with possible concept drifts in datasets such as the Forest Covertype, Electricity, and Nebraska Weather. These results suggest that some widely used datasets may be trivial from the concept drift standpoint, and thus, should be avoided, or at least the results should be compared with the proposed Naïve methods. Paulo R. L. Almeida, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Jean Paul Barddal |
SMC | 2 |
| 2020 | On the Selection of the Competence Measure for Dynamic Regressor SelectionabstractDynamic regressor selection (DRS) systems work by selecting the most competent regressors from an ensemble to predict the target value of a query pattern. This competence is calculated using the performance of the regressors in a local region of the feature space around the query pattern that is called the region of competence. Nonetheless, defining the correct measure to compute the degree of competence of the regressors is a hard task. In this work, we propose a new technique to DRS that selects the best competence measure for a given dataset. To validate our technique, we perform a set of comprehensive experiments on 15 regression datasets. The proposed technique can operate in three different fashions: (i) selection of the most competent regressor; (ii) combination of all regressors from the ensemble; and (iii) selection of a subset composed of the most competent ones and combine them. The proposals are compared against DRS algorithms, individual regressors, and static systems that use the Mean and the Median as a fusion strategy. The results show that the proposed technique, which chooses a different competence measure per task, outperforms literature techniques. Thiago J. M. Moura, George D. C. Cavalcanti, Luiz Eduardo Soares de Oliveira |
SMC | 3 |
| 2020 | CSBF: A static ensemble fusion method based on the centrality score of complex networksabstractAbstract Ensemble of classifiers can improve classification accuracy by combining several models. The fusion method plays an important role in the ensemble performance. Usually, a criterion for weighting the decision of each ensemble member is adopted. Frequently, this can be done using some heuristic based on accuracy or confidence. Then, the used fusion rule must consider the established criterion for providing a most reliable ensemble output through a kind of competition among the ensemble members. This article presents a new ensemble fusion method, named centrality score‐based fusion, which uses the centrality concept in the context of social network analysis (SNA) as a criterion for the ensemble decision. Centrality measures have been applied in the SNA to measure the importance of each person inside of a social network, taking into account the relationship of each person with all others. Thus, the idea is to derive the classifier weight considering the overall classifier prominence inside the ensemble network, which reflects the relationships among pairs of classifiers. We hypothesized that the prominent position of a classifier based on its pairwise relationship with the other ensemble members could be its weight in the fusion process. A robust experimental protocol has confirmed that centrality measures represent a promising strategy to weight the classifiers of an ensemble, showing that the proposed fusion method performed well against the literature. Ronan Assumpção Silva, Alceu S. Britto Jr., Fabrício Enembreck, Robert Sabourin, Luiz Eduardo Soares de Oliveira |
Comput. Intell. | 5 |
| 2020 | Meta-Learning for Fast Classifier Adaptation to New Users of Signature Verification SystemsabstractOffline Handwritten Signature verification presents a challenging Pattern Recognition problem, where only knowledge of the positive class is available for training. While classifiers have access to a few genuine signatures for training, during generalization they also need to discriminate forgeries. This is particularly challenging for skilled forgeries, where a forger practices imitating the user's signature, and often is able to create forgeries visually close to the original signatures. Most work in the literature address this issue by training for a surrogate objective: discriminating genuine signatures of a user and random forgeries (signatures from other users). In this work, we propose a solution for this problem based on meta-learning, where there are two levels of learning: a task-level (where a task is to learn a classifier for a given user) and a meta-level (learning across tasks). In particular, the meta-learner guides the adaptation (learning) of a classifier for each user, which is a lightweight operation that only requires genuine signatures. The meta-learning procedure learns what is common for the classification across different users. In a scenario where skilled forgeries from a subset of users are available, the meta-learner can guide classifiers to be discriminative of skilled forgeries even if the classifiers themselves do not use skilled forgeries for learning. Experiments conducted on the GPDS-960 dataset show improved performance compared to Writer-Independent systems, and achieve results comparable to state-of-the-art Writer-Dependent systems in the regime of few samples per user (5 reference signatures). Luiz G. Hafemann, Robert Sabourin, Luiz Eduardo Soares de Oliveira |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | Texture CNN for Histopathological Image ClassificationabstractBiopsies are the gold standard for breast cancer diagnosis. This task can be improved by the use of Computer Aided Diagnosis (CAD) systems, reducing the time of diagnosis and reducing the inter and intra-observer variability. The advances in computing have brought this type of system closer to reality. However, datasets of Histopathological Images (HI) from biopsies are quite small and unbalanced what makes difficult to use modern machine learning techniques such as deep learning. In this paper we propose a compact architecture based on texture filters that has fewer parameters than traditional deep models but is able to capture the difference between malignant and benign tissues with relative accuracy. The experimental results on the BreakHis dataset have show that the proposed texture CNN achieves almost 90% of accuracy for classifying benign and malignant tissues. Jonathan de Matos, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich |
CBMS | 3 |
| 2019 | Decoupling Direction and Norm for Efficient Gradient-Based L2 Adversarial Attacks and DefensesabstractResearch on adversarial examples in computer vision tasks has shown that small, often imperceptible changes to an image can induce misclassification, which has security implications for a wide range of image processing systems. Considering L2 norm distortions, the Carlini and Wagner attack is presently the most effective white-box attack in the literature. However, this method is slow since it performs a line-search for one of the optimization terms, and often requires thousands of iterations. In this paper, an efficient approach is proposed to generate gradient-based attacks that induce misclassifications with low L2 norm, by decoupling the direction and the norm of the adversarial perturbation that is added to the image. Experiments conducted on the MNIST, CIFAR-10 and ImageNet datasets indicate that our attack achieves comparable results to the state-of-the-art (in terms of L2 norm) with considerably fewer iterations (as few as 100 iterations), which opens the possibility of using these attacks for adversarial training. Models trained with our attack achieve state-of-the-art robustness against white-box gradient-based L2 attacks on the MNIST and CIFAR-10 datasets, outperforming the Madry defense when the attacks are limited to a maximum norm. Jérôme Rony, Luiz G. Hafemann, Luiz Eduardo Soares de Oliveira, Ismail Ben Ayed, Robert Sabourin, Eric Granger |
CVPR | 3 |
| 2019 | Binarization of Degraded Document Images using Convolutional Neural Networks Based on Predicted Two-Channel ImagesabstractDue to the poor condition of most of historical documents, binarization is difficult to separate document image background pixels from foreground pixels. This paper proposes Convolutional Neural Networks (CNNs) based on predicted two-channel images in which CNNs are trained to classify the foreground pixels. The promising results from the use of multispectral images for semantic segmentation inspired our efforts to create a novel prediction-based two-channel image. In our method, the original image is binarized by the structural symmetric pixels (SSPs) method, and the two-channel image is constructed from the original image and its binarized image. In order to explore impact of proposed two-channel images as network inputs, we use two popular CNNs architectures, namely SegNet and U-net. The results presented in this work show that our approach fully outperforms SegNet and U-net when trained by the original images and demonstrates competitiveness and robustness compared with state-of-the-art results using the DIBCO database. Younes Akbari, Alceu S. Britto Jr., Somaya Al-Máadeed, Luiz Eduardo Soares de Oliveira |
ICDAR | 4 |
| 2019 | Adaptive Random Forests with Resampling for Imbalanced data StreamsabstractThe large volume of data generated by computer networks, smartphones, wearables and a wide range of sensors, which produce real-time data, are only useful if they can be efficiently processed so that individuals can make timely decisions based on them. In this context, machine learning techniques are widely used. While it performs better than humans in such tasks, every machine learning algorithm has a certain intrinsic bias, which means they assume that the data have specific characteristics, such as having a balanced distribution between classes. As many real-world applications present imbalanced traits in their data, this topic is gaining repercussion over time. In this work, we present the Adaptive Random Forest with Resampling (ARFRE), which is a classifier designed to deal with imbalanced datasets. ARFREresample the instances based on the current class label distribution. We show through a set of extensive experiments on seven datasets that the proposed method can considerably improve the performance of the minority class(es) while avoiding degrading the performance in the majority class. On top of that, ARFREis more efficient regarding execution time in comparison to the standard ARF algorithm. Luis Eduardo Boiko Ferreira, Heitor Murilo Gomes, Albert Bifet, Luiz Eduardo Soares de Oliveira |
IJCNN | 4 |
| 2019 | Double Transfer Learning for Breast Cancer Histopathologic Image ClassificationabstractThis work proposes a classification approach for breast cancer histopathologic images (HI) that uses transfer learning to extract features from HI using an Inception-v3 CNN pre-trained with ImageNet dataset. We also use transfer learning on training a support vector machine (SVM) classifier on a tissue labeled colorectal cancer dataset aiming to filter the patches from a breast cancer HI and remove the irrelevant ones. We show that removing irrelevant patches before training a second SVM classifier, improves the accuracy for classifying malign and benign tumors on breast cancer images. We are able to improve the classification accuracy in 3.7% using the feature extraction transfer learning and an additional 0.7% using the irrelevant patch elimination. The proposed approach outperforms the state-of-the-art in three out of the four magnification factors of the breast cancer dataset. Jonathan de Matos, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich |
IJCNN | 3 |
| 2019 | Evaluating Competence Measures for Dynamic Regressor SelectionabstractDynamic regressor selection (DRS) systems work by selecting the most competent regressors from an ensemble to estimate the target value of a given test pattern. This competence is usually quantified using the performance of the regressors in local regions of the feature space around the test pattern. However, choosing the best measure to calculate the level of competence correctly is not straightforward. The literature of dynamic classifier selection presents a wide variety of competence measures, which cannot be used or adapted for DRS. In this paper, we review eight measures used with regression problems, and adapt them to test the performance of the DRS algorithms found in the literature. Such measures are extracted from a local region of the feature space around the test pattern, called region of competence, therefore competence measures. To better compare the competence measures, we perform a set of comprehensive experiments of 15 regression datasets. Three DRS systems were compared against individual regressor and static systems that use the Mean and the Median to combine the outputs of the regressors from the ensemble. The DRS systems were assessed varying the competence measures. Our results show that DRS systems outperform individual regressors and static systems but the choice of the competence measure is problem-dependent. Thiago J. M. Moura, George D. C. Cavalcanti, Luiz Eduardo Soares de Oliveira |
IJCNN | 3 |
| 2019 | Representation Learning vs. Handcrafted Features for Music Genre ClassificationabstractIn this work we present a comprehensive set of experiments aiming to perform music genre classification using learned and handcrafted features plus the fusion of them. Handcrafted features were obtained from the audio signal itself, lyrics, chords and spectrogram images extracted from the audio. The rationale behind this investigation is based on the assumption that one can find some complementarity between classifiers created from these different resources. The experimental protocol was conducted on the Brazilian Music Dataset using the artist filter restriction and they confirm the power of non-handcrafted features to perform audio classification tasks. The experimental results have shown a significant complementarity among the handcrafted features for which the evaluated fusion strategies allowed an improvement in the classification accuracy up to 4 percent points. On the other hand, the fusion of learned and handcrafted features provided similar accuracy than the best individual CNN (0.7815). Rodolfo Miranda Pereira, Yandre M. G. Costa, Rafael de Lima Aguiar, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Carlos Nascimento Silla Jr. |
IJCNN | 5 |
| 2019 | Image Retrieval and Pattern Spotting using Siamese Neural NetworkabstractThis paper presents a novel approach for image retrieval and pattern spotting in document image collections. The manual feature engineering is avoided by learning a similarity-based representation using a Siamese Neural Network trained on a previously prepared subset of image pairs from the ImageNet dataset. The learned representation is used to provide the similarity-based feature maps used to find relevant image candidates in the data collection given an image query. A robust experimental protocol based on the public Tobacco800 document image collection shows that the proposed method compares favor-ably against state-of-the-art document image retrieval methods, reaching 0.94 and 0.83 of mean average precision (mAP) for retrieval and pattern spotting (IoU=0.7), respectively. Besides, we have evaluated the proposed method considering feature maps of different sizes, showing the impact of reducing the number of features in the retrieval performance and time-consuming. Kelly Lais Wiggers, Alceu S. Britto Jr., Laurent Heutte, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira |
IJCNN | 5 |
| 2019 | L(a)ying in (Test)Bed - How Biased Datasets Produce Impractical Results for Actual Malware Families' Classification
Tamy Beppler, Marcus Botacin, Fabricio Ceschin, Luiz Eduardo Soares de Oliveira, André Ricardo Abed Grégio |
ISC | 4 |
| 2019 | Multiple instance learning for histopathological breast cancer image classification
P. J. Sudharshan, Caroline Petitjean, Fabio A. Spanhol, Luiz Eduardo Soares de Oliveira, Laurent Heutte, Paul Honeine |
Expert Syst. Appl. | 4 |
| 2019 | Characterizing and Evaluating Adversarial Examples for Offline Handwritten Signature VerificationabstractThe phenomenon of adversarial examples is attracting increasing interest from the machine learning community, due to its significant impact on the security of machine learning systems. Adversarial examples are similar (from a perceptual notion of similarity) to samples from the data distribution, that “fool” a machine learning classifier. For computer vision applications, these are images with carefully crafted but almost imperceptible changes, which are misclassified. In this paper, we characterize this phenomenon under an existing taxonomy of threats to biometric systems, in particular identifying new attacks for offline handwritten signature verification systems. We conducted an extensive set of experiments on four widely used datasets: MCYT-75, CEDAR, GPDS-160, and the Brazilian PUC-PR, considering both a CNN-based system and a system using a handcrafted feature extractor. We found that attacks that aim to get a genuine signature rejected are easy to generate, even in a limited knowledge scenario, where the attacker does not have access to the trained classifier nor the signatures used for training. Attacks that get a forgery to be accepted are harder to produce, and often require a higher level of noise-in most cases, no longer “imperceptible” as previous findings in object recognition. We also evaluated the impact of two countermeasures on the success rate of the attacks and the amount of noise required for generating successful attacks. Luiz G. Hafemann, Robert Sabourin, Luiz Eduardo Soares de Oliveira |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Exploring Textures in Traffic Matrices to Classify Data Center CommunicationsabstractData analytics and scientific computing are two modern applications that in recent years have substantially changed their computation and communication needs, requiring additional processing capability and bandwidth to be able to keep pace with current demands. These applications are commonly processed within data centers, exchanging enormous volumes of data, rapidly stressing existing network infrastructures. Thus, it is crucial for data center operations and management to be able to understand and classify the communication demands of these applications. The traditional approaches for classifying application traffic are port-based and Deep Packet Inspection, both presenting issues with current network technology. Some recent works propose using machine learning plus statistical information collected from application flows to classify traffic. Applications running in data centers present communication patterns which can be recognized through their traffic matrices. So, the main contribution of this paper is a method that explores the textural information extracted from these matrices to classify the data center traffic using machine learning techniques. As a proof-of-concept, we implemented this method in a system named DCTraCS. The experimental dataset was gathered from two real data centers, collecting the traffic matrices of MapReduce and a set of scientific applications every second for a period of 30 minutes. For assessing our proposal, we compared it with other machine learning techniques for classifying application traffic found in current literature. Results show that our approach achieved the highest accuracy, classifying correctly over 99% of our data center applications. Celio Trois, Luis C. E. Bona, Luiz Eduardo Soares de Oliveira, Magnos Martinello, Douglas Harewood-Gill, Marcos Didonet Del Fabro, Reza Nejabati, Dimitra Simeonidou, João Carlos D. Lima, Benhur de Oliveira Stein |
AINA | 3 |
| 2018 | Confusion Matrix-Based Building of Hierarchical Classification
Paulo Rodrigo Cavalin, Luiz Eduardo Soares de Oliveira |
CIARP | 2 |
| 2018 | Fine-Grained Hierarchical Classification of Plant Leaf Images Using Fusion of Deep ModelsabstractA fine-grained plant leaf classification method based on the fusion of deep models is described. Complementary global and patch-based leaf features are combined at each hierarchical level (genus and species) by pre-trained CNNs. The deep models are adapted for plant recognition by using data augmentation techniques to face the problem of plant classes with very few samples for training in the available imbalanced dataset. Experimental results have shown that the proposed coarse-to-fine classification strategy is a very promising alternative to deal with the low inter-class and high intra-class variability inherent to the problem of plant identification. The proposed method was able to surpass other state-of-the-art approaches on the ImageCLEF 2015 plant recognition dataset in terms of average classification scores. Voncarlos Araujo, Alceu S. Britto Jr., Andre L. Brun, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira |
ICTAI | 5 |
| 2018 | Forest Species Recognition Based on Ensembles of ClassifiersabstractRecognition of forest species is a very challenging task thanks to the great intra-class variability. To cope with such a variability, we propose a multiple classifier system based on a two-level classification strategy and microscopic images. By using a divide-and-conquer approach, an image is first divided into several sub-images which are classified independently by each classifier. In a first fusion level, partial decisions for the sub-images are combined to generate a new partial decision for the original image. Then, the second fusion level combines all these new partial decisions to produce the final classification of the original image. To generate the pool of diverse classifiers, we used classical texture-based features as well as keypoint-based features. A series of experiments shows that the proposed strategy achieves compelling results. Compared to the best single classifier, a Support Vector Machine (SVM) trained with a keypoint based feature set, the divide-and-conquer strategy improves the recognition rate in about 4 and 6 percentage points in the first and second fusion levels, respectively. The best recognition rate achieved by this proposed method is 98.47%. Jefferson G. Martins, Luiz Eduardo Soares de Oliveira, Robert Sabourin, Alceu S. Britto Jr. |
ICTAI | 2 |
| 2018 | A Brazilian Speech DatabaseabstractThis work introduces a Brazilian Speech Database (BrSD), a novel dataset freely available created to support the development of speech-based recognition tasks. As far as we know, this is the first Portuguese language based database with these characteristics created and made available to the research community. We also describe experiments accomplished on BrSD exploring its different possibilities of classification tasks, i.e., age group and gender classification. We use four well-known acoustic features extracted directly from the audio signal and one texture-based feature extracted from a visual representation of the audio signal, the spectrogram. We considered three different classification scenarios: each feature individually, early fusion of the features, and late fusion of the features. Experiments were conducted using Support Vector Machine (SVM) and Multi-layer Perceptron (MLP) classifiers. The obtained results showed that SVM classifier achieved the best recognition rates both in early and late fusion scenarios. The best recognition rates achieved were 91.25%, 88.75%, and 80.25% for gender, age group, and age-gender classification tasks, respectively. Marco Aurelio Deoldoto Paulino, Yandre M. G. Costa, Alceu S. Britto Jr., Alisson Renan Svaigen, Linnyer B. Ruiz, Luiz Eduardo Soares de Oliveira |
ICTAI | 6 |
| 2018 | Dynamic Ensemble Selection by K-Nearest Local Oracles with Discrimination IndexabstractThis work describes a new oracle based Dynamic Ensemble Selection (DES) method in which an Ensemble of Classifiers (EoC) is selected to predict the class of a given test instance (xt). The competence of each classifier is estimated on a local region (LR) of the feature space (Region of Competence - RoC) represented by the most promising k-nearest neighbors (or advisors) related to xt according to a discrimination index (D) originally proposed in the Item and Test Analysis (ITA) theory. The D value is used to better define the advisors of the RoC since they will suggest the classifiers (local oracles) to compose the EoC. A robust experimental protocol based on 30 classification problems and 20 replications have shown that the proposed DES compares favorably with 15 state-of-the-art dynamic selection methods and the combination of all classifiers in the pool. Marcelo Pereira, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Robert Sabourin |
ICTAI | 3 |
| 2018 | Segmentation-Free Approaches For Handwritten Numeral String RecognitionabstractThis paper presents segmentation-free strategies for the recognition of handwritten numeral strings of unknown length. A synthetic dataset of touching numeral strings of sizes 2-, 3- and 4-digits was created to train end-to-end solutions based on Convolutional Neural Networks. A robust experimental protocol is used to show that the proposed segmentation-free methods may reach the state-of-the-art performance without suffering the heavy burden of over-segmentation based methods. In addition, they confirmed the importance of introducing contextual information in the design of end-to-end solutions, such as the proposed length classifier when recognizing numeral strings. Andre G. Hochuli, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
IJCNN | 2 |
| 2018 | A Robust Real-Time Automatic License Plate Recognition Based on the YOLO DetectorabstractAutomatic License Plate Recognition (ALPR) has been a frequent topic of research due to many practical applications. However, many of the current solutions are still not robust in real-world situations, commonly depending on many constraints. This paper presents a robust and efficient ALPR system based on the state-of-the-art YOLO object detector. The Convolutional Neural Networks (CNNs) are trained and finetuned for each ALPR stage so that they are robust under different conditions (e.g., variations in camera, lighting, and background). Specially for character segmentation and recognition, we design a two-stage approach employing simple data augmentation tricks such as inverted License Plates (LPs) and flipped characters. The resulting ALPR approach achieved impressive results in two datasets. First, in the SSIG dataset, composed of 2,000 frames from 101 vehicle videos, our system achieved a recognition rate of 93.53% and 47 Frames Per Second (FPS), performing better than both Sighthound and OpenALPR commercial systems (89.80% and 93.03%, respectively) and considerably outperforming previous results (81.80%). Second, targeting a more realistic scenario, we introduce a larger public dataset1dataset, designed to ALPR. This dataset contains 150 videos and 4,500 frames captured when both camera and vehicles are moving and also contains different types of vehicles (cars, motorcycles, buses and trucks). In our proposed dataset, the trial versions of commercial systems achieved recognition rates below 70%. On the other hand, our system performed better, with recognition rate of 78.33% and 35 FPS.The UFPR-ALPR dataset is publicly available to the research community at https://web.inf.ufpr.br/vri/databases/ufpr-alpr/ subject to privacy restrictions. Rayson Laroca, Evair Severo, Luiz Antonio Zanlorensi, Luiz Eduardo Soares de Oliveira, Gabriel Resende Gonçalves, William Robson Schwartz, David Menotti |
IJCNN | 4 |
| 2018 | Document Image Retrieval Using Deep FeaturesabstractThis paper proposes a novel approach for content based graphical object retrieval in document images. The challenge is to search for occurrences of a queried graphical objects in document images that can vary in terms of color, shape, texture and quality, increasing considerably the level of difficulty of the retrieval process. To that end, the manual feature engineering is avoided by learning the image representation for the retrieval task using a Convolutional Neural Network (CNN). However, such a representation should be as compact as possible to allow a fast document image retrieval and storage. Thus, a pretrained CNN model is used to cope with the lack of training data, which is fine tuned to achieve a compact yet discriminant representation of the graphical objects. From experiments conducted on the public Tobacco800 document image collection, we show that the proposed method compares favorably against state-of-the-art document image retrieval methods, reaching 0.72 of average precision (mAP). In addition, an increase of 4 percentage points in the average precision is observed using a compact deep representation in which the number of features is reduced by 16 times, thus allowing a reduction of 47% in terms of computation time by the image retrieval task. Kelly Lais Wiggers, Alceu S. Britto Jr., Laurent Heutte, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira |
IJCNN | 5 |
| 2018 | Enabling Anomaly-based Intrusion Detection Through Model GeneralizationabstractAnomaly-based intrusion detection by the means of machine learning techniques is extensively studied in the literature mainly due to its promise to detect new attacks. However, despite the promising reported results, it is hardly deployed to real world environments. The main challenge in its adoption is the discrepancy between the accuracy rates obtained during the classifier development process and the rates obtained during its use in production environments. Such a discrepancy is mainly caused by non-representative training databases and nongeneralizable (scenario-specific) classifier's model. This paper presents a method to create intrusion databases, which aims at mimicking the production environments characteristics by using well-known tools. Moreover, we present and evaluate a new validation technique, which aims at ensuring the generalization capacity of the obtained models, reached using cross-validating with different intrusion databases. The evaluation tests showed the feasibility of the proposed method. The feature selection technique ensured the model generalization capacity, improving its accuracy rate by 13%, while testing in different intrusion databases. Finally, the proposed anomaly-based approach was compared with Snort, reaching an accuracy rate of 99% against 27% of Snort for detecting DoS attacks. Eduardo Viegas 0001, Altair Olivo Santin, Vilmar Abreu, Luiz Eduardo Soares de Oliveira |
ISCC | 4 |
| 2018 | Adapting dynamic classifier selection for concept drift
Paulo R. L. Almeida, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
Expert Syst. Appl. | 2 |
| 2018 | Fixed-sized representation learning from offline handwritten signatures of different sizes
Luiz G. Hafemann, Luiz Eduardo Soares de Oliveira, Robert Sabourin |
Int. J. Document Anal. Recognit. | 2 |
| 2018 | A framework for dynamic classifier selection oriented by the classification problem difficulty
Andre L. Brun, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Fabrício Enembreck, Robert Sabourin |
Pattern Recognit. | 3 |
| 2018 | Handwritten digit segmentation: Is it still necessary?
Andre G. Hochuli, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
Pattern Recognit. | 2 |
| 2018 | Learning Deep Off-the-Person Heart Biometrics RepresentationsabstractSince the beginning of the new millennium, the electrocardiogram (ECG) has been studied as a biometric trait for security systems and other applications. Recently, with devices such as smartphones and tablets, the acquisition of ECG signal in the off-the-person category has made this biometric signal suitable for real scenarios. In this paper, we introduce the usage of deep learning techniques, specifically convolutional networks, for extracting useful representation for heart biometrics recognition. Particularly, we investigate the learning of feature representations for heart biometrics through two sources: on the raw heartbeat signal and on the heartbeat spectrogram. We also introduce heartbeat data augmentation techniques, which are very important to generalization in the context of deep learning approaches. Using the same experimental setup for six methods in the literature, we show that our proposal achieves state-of-the-art results in the two off-the-person publicly available databases. Eduardo José da S. Luz, Gladston J. P. Moreira, Luiz Eduardo Soares de Oliveira, William Robson Schwartz, David Menotti |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | Knowledge Transfer for Writer Identification
Diego Bertolini, Luiz Eduardo Soares de Oliveira, Yandre M. G. Costa, Lucas Georges Helal |
CIARP | 2 |
| 2017 | A two-step cascade classification methodabstractThis paper proposes a classification approach in which monolithic and multiple classifier systems are combined in a cascading fashion. The rationale behind that is to deal with the existing trade-off between the need for increasing the accuracy, while reducing the complexity of the classification method. In other words, the idea is to offer an interesting strategy to conciliate the different levels of efforts necessary to deal with easy and hard patterns usually observed in a classification problem. The experimental results have shown that for some problems more than 90% of the instances can be processed in the first step of the cascade, saving efforts by avoiding the use of the second step in which a more complex classification method is used. It means that for some problems the reduction of the classification cost achieved more than 70% when compared to the use of an MCS. In addition to this interesting classification cost reduction, the cascade approach has shown to be able of improving the classification accuracy up to 15.19 percentage points. Eunelson Jose da Silva Junior, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Fabrício Enembreck, Robert Sabourin, Alessandro L. Koerich |
IJCNN | 3 |
| 2017 | Stream learning and anomaly-based intrusion detection in the adversarial settingsabstractDespite existing many anomaly-based intrusion detection studies in the literature, they are not frequently adopted by the industry in production environments (products). Such a usage gap occurs mainly due to the difficulty to maintain the detection rate in acceptable level, given the occurrence of false alarms. In general, the literature does not consider the adversarial settings, when an opponent attempt to evade the detection system, thus possibly rendering the system unreliable over time. In this paper, we propose and evaluate a new approach to reliably perform real time stream learning for anomaly-based intrusion detection. We employ a class-specific stream outlier detector to automatically update the intrusion detection engine over the time, and a rejection mechanism, which makes it possible to obtain indications that an evasion attempt might being happening. Furthermore, the proposal is resilient to causative attacks, providing a secure intrusion detection mechanism even when the attacker can inject misclassified instances in the training dataset. The evaluation tests show that the proposed approach is resilient to exploratory attacks, allowing the system administrator to know when an evasion attempt might be occurring. Eduardo Viegas 0001, Altair Olivo Santin, Vilmar Abreu, Luiz Eduardo Soares de Oliveira |
ISCC | 4 |
| 2017 | Multi-scale texture recognition systems with reduced cost: A case study on forest speciesabstractThis work focuses on cost reduction methods, applied on forest species recognition systems as a case-study. Current state-of-the-art shows that the accuracy of these systems, generally employing texture recognition approaches, have increased considerably in the past years. However, the cost in time to perform the recognition of input samples has also increased proportionally. By taking into account previous research that demonstrated that cost reduction at classification level can provide much faster systems, in this work we focus on proposing metrics to measure the impact of cost reduction at another important module of image recognition system, i.e the feature extraction stage, and on how to measure cost reduction at global level, i.e. combining cost reduction at both feature extraction and classification. The evaluation of the proposed metrics on a forest species dataset demonstrated that, with global cost reduction, not only the cost of the system can be reduced to less than 1/20, but also the recognition rates can be improved. Paulo Rodrigo Cavalin, Marcelo N. Kapp, Luiz Eduardo Soares de Oliveira |
SMC | 3 |
| 2017 | Deep features for breast cancer histopathological image classificationabstractBreast cancer (BC) is a deadly disease, killing millions of people every year. Developing automated malignant BC detection system applied on patient's imagery can help dealing with this problem more efficiently, making diagnosis more scalable and less prone to errors. Not less importantly, such kind of research can be extended to other types of cancer, making even more impact to help saving lives. Recent results on BC recognition show that Convolution Neural Networks (CNN) can achieve higher recognition rates than hand-crafted feature descriptors, but the price to pay is an increase in complexity to develop the system, requiring longer training time and specific expertise to fine-tune the architecture of the CNN. DeCAF (or deep) features consist of an in-between solution it is based on reusing a previously trained CNN only as feature vectors, which is then used as input for a classifier trained only for the new classification task. In the light of this, we present an evaluation of DeCaf features for BC recognition, in order to better understand how they compare to the other approaches. The experimental evaluation shows that these features can be a viable alternative to fast development of high-accuracy BC recognition systems, generally achieving better results than traditional hand-crafted textural descriptors and outperforming task-specific CNNs in some cases. Fabio A. Spanhol, Luiz Eduardo Soares de Oliveira, Paulo Rodrigo Cavalin, Caroline Petitjean, Laurent Heutte |
SMC | 2 |
| 2017 | Toward a reliable anomaly-based intrusion detection in real-world environments
Eduardo Viegas 0001, Altair Olivo Santin, Luiz Eduardo Soares de Oliveira |
Comput. Networks | 3 |
| 2017 | Bias effect on predicting market trends with EMD
Dennis Carnelossi Furlaneto, Luiz Eduardo Soares de Oliveira, David Menotti, George D. C. Cavalcanti |
Expert Syst. Appl. | 2 |
| 2017 | Learning features for offline handwritten signature verification using deep convolutional neural networks
Luiz G. Hafemann, Robert Sabourin, Luiz Eduardo Soares de Oliveira |
Pattern Recognit. | 3 |
| 2017 | Towards an Energy-Efficient Anomaly-Based Intrusion Detection Engine for Embedded SystemsabstractNowadays, a significant part of all network accesses comes from embedded and battery-powered devices, which must be energy efficient. This paper demonstrates that a hardware (HW) implementation of network security algorithms can significantly reduce their energy consumption compared to an equivalent software (SW) version. The paper has four main contributions: (i) a new feature extraction algorithm, with low processing demands and suitable for hardware implementation; (ii) a feature selection method with two objectives - accuracy and energy consumption; (iii) detailed energy measurements of the feature extraction engine and three machine learning (ML) classifiers implemented in SW and HW-Decision Tree (DT), Naive-Bayes (NB), and k-Nearest Neighbors (kNN); and (iv) a detailed analysis of the tradeoffs in implementing the feature extractor and ML classifiers in SW and HW. The new feature extractor demands significantly less computational power, memory, and energy. Its SW implementation consumes only 22 percent of the energy used by a commercial product and its HW implementation only 12 percent. The dual-objective feature selection enabled an energy saving of up to 93 percent. Comparing the most energy-efficient SW implementation (new extractor and DT classifier) with an equivalent HW implementation, the HW version consumes only 5.7 percent of the energy used by the SW version. Eduardo Viegas 0001, Altair Olivo Santin, André França 0001, Ricardo P. Jasinski, Volnei A. Pedroni, Luiz Eduardo Soares de Oliveira |
IEEE Trans. Computers | 6 |
| 2016 | Multi-script writer identification using dissimilarityabstractMulti-script writer identification consists in identifying a person of a given text written in one script from the samples of the same person written in another script. The rationale behind this is that the writing style of an individual remains constant across different scripts. While this hypothesis may hold, recent results on a multi-script writer identification competition show that classical writer-dependent classifiers fail in this task. In this work we investigate the efficacy of a writer-independent classifier based on dissimilarity for multi-script writer identification. The classifiers were trained using two different texture descriptors (LBP and LPQ). Our experiments on 475 writers of the QUWI dataset, which is composed of Arabic and English samples, show that the proposed strategy surpasses the results published in the literature by a large margin, achieving error rates similar to single-script writer identification systems. Diego Bertolini, Luiz Eduardo Soares de Oliveira, Robert Sabourin |
ICPR | 2 |
| 2016 | Analyzing features learned for Offline Signature Verification using Deep CNNsabstractResearch on Offline Handwritten Signature Verification explored a large variety of handcrafted feature extractors, ranging from graphology, texture descriptors to interest points. In spite of advancements in the last decades, performance of such systems is still far from optimal when we test the systems against skilled forgeries - signature forgeries that target a particular individual. In previous research, we proposed a formulation of the problem to learn features from data (signature images) in a Writer-Independent format, using Deep Convolutional Neural Networks (CNNs), seeking to improve performance on the task. In this research, we push further the performance of such method, exploring a range of architectures, and obtaining a large improvement in state-of-the-art performance on the GPDS dataset, the largest publicly available dataset on the task. In the GPDS-160 dataset, we obtained an Equal Error Rate of 2.74%, compared to 6.97% in the best result published in literature (that used a combination of multiple classifiers). We also present a visual analysis of the feature space learned by the model, and an analysis of the errors made by the classifier. Our analysis shows that the model is very effective in separating signatures that have a different global appearance, while being particularly vulnerable to forgeries that very closely resemble genuine signatures, even if their line quality is bad, which is the case of slowly-traced forgeries. Luiz G. Hafemann, Robert Sabourin, Luiz Eduardo Soares de Oliveira |
ICPR | 3 |
| 2016 | Handling Concept Drifts Using Dynamic Selection of ClassifiersabstractThis work describes the Dynse framework, which uses dynamic selection of classifiers to deal with concept drift. Basically, classifiers trained on new supervised batches available over time are add to a pool, from which is elected a custom ensemble for each test instance during the classification time. The Dynse framework is highly customizable, and can be adapted to use any method for dynamic selection of classifiers given a test instance. In this work we propose a default configuration for the framework which has provided promising results in a range of problems. The experimental results have shown that the proposed framework achieved the best average rank when considering all datasets, and outperformed the state-of-the-art in three of four tested datasets. Paulo R. L. Almeida, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
ICTAI | 2 |
| 2016 | Contribution of data complexity features on dynamic classifier selectionabstractDifferent dynamic classifier selection techniques have been proposed in the literature to determine among diverse classifiers available in a pool which should be used to classify a test instance. The individual competence of each classifier in the pool is usually evaluated taking into account its accuracy on the neighborhood of the test instance in a validation dataset. In this work we investigate the possible contribution of considering during the classifier evaluation the use of features related to the problem complexity. Since usually the pool generation technique does not assure diversity, the idea is to consider diversity during the selection. Basically, we select a classifier trained in subset of data showing similar complexity than that observed in neighborhood of the test instance. We expect that this similarity in terms of complexity allow us to select a more competent classifier. Experiments on 30 classification problems representing different levels of difficulty have shown that the proposed selection method is comparable to well known dynamic selection strategies. When compared with other DS approaches it was able to win on 123 over 150 experiments. This promising results indicate that further investigation must be done to increase diversity in terms of data complexity during the process of pool generation. Andre L. Brun, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Fabrício Enembreck, Robert Sabourin |
IJCNN | 3 |
| 2016 | Writer-independent feature learning for Offline Signature Verification using Deep Convolutional Neural NetworksabstractAutomatic Offline Handwritten Signature Verification has been researched over the last few decades from several perspectives, using insights from graphology, computer vision, signal processing, among others. In spite of the advancements on the field, building classifiers that can separate between genuine signatures and skilled forgeries (forgeries made targeting a particular signature) is still hard. We propose approaching the problem from a feature learning perspective. Our hypothesis is that, in the absence of a good model of the data generation process, it is better to learn the features from data, instead of using hand-crafted features that have no resemblance to the signature generation process. To this end, we use Deep Convolutional Neural Networks to learn features in a writer-independent format, and use this model to obtain a feature representation on another set of users, where we train writer-dependent classifiers. We tested our method in two datasets: GPDS-960 and Brazilian PUC-PR. Our experimental results show that the features learned in a subset of the users are discriminative for the other users, including across different datasets, reaching close to the state-of-the-art in the GPDS dataset, and improving the state-of-the-art in the Brazilian PUC-PR dataset. Luiz G. Hafemann, Robert Sabourin, Luiz Eduardo Soares de Oliveira |
IJCNN | 3 |
| 2016 | Breast cancer histopathological image classification using Convolutional Neural NetworksabstractThe performance of most conventional classification systems relies on appropriate data representation and much of the efforts are dedicated to feature engineering, a difficult and time-consuming process that uses prior expert domain knowledge of the data to create useful features. On the other hand, deep learning can extract and organize the discriminative information from the data, not requiring the design of feature extractors by a domain expert. Convolutional Neural Networks (CNNs) are a particular type of deep, feedforward network that have gained attention from research community and industry, achieving empirical successes in tasks such as speech recognition, signal processing, object recognition, natural language processing and transfer learning. In this paper, we conduct some preliminary experiments using the deep learning approach to classify breast cancer histopathological images from BreaKHis, a publicly dataset available at http://web.inf.ufpr.br/vri/breast-cancer-database. We propose a method based on the extraction of image patches for training the CNN and the combination of these patches for final classification. This method aims to allow using the high-resolution histopathological images from BreaKHis as input to existing CNN, avoiding adaptations of the model that can lead to a more complex and computationally costly architecture. The CNN performance is better when compared to previously reported results obtained by other machine learning models trained with hand-crafted textural descriptors. Finally, we also investigate the combination of different CNNs using simple fusion rules, achieving some improvement in recognition rates. Fabio A. Spanhol, Luiz Eduardo Soares de Oliveira, Caroline Petitjean, Laurent Heutte |
IJCNN | 2 |
| 2016 | Combining diversity measures for ensemble pruning
George D. C. Cavalcanti, Luiz Eduardo Soares de Oliveira, Thiago J. M. Moura, Guilherme V. Carvalho |
Pattern Recognit. Lett. | 2 |
| 2015 | Improving Writer Identification Through Writer Selection
Diego Bertolini, Luiz Eduardo Soares de Oliveira, Robert Sabourin |
CIARP | 2 |
| 2015 | Towards a SignWriting recognition systemabstractSignWriting is a writing system for sign languages. It is based on visual symbols to represent the hand shapes, movements and facial expressions, among other elements. It has been adopted by more than 40 countries, but to ensure the social integration of the deaf community, writing systems based on sign languages should be properly incorporated into the Information Technology. This article reports our first efforts toward the implementation of an automatic reading system for SignWiring. This would allow converting the SignWriting script into text so that one can store, retrieve, and index information in an efficient way. In order to make this work possible, we have been collecting a database of hand configurations, which at the present moment sums up to 7,994 images divided into 103 classes of symbols. To classify such symbols, we have performed a comprehensive set of experiments using different features, classifiers, and combination strategies. The best result, 94.4% of recognition rate, was achieved by a Convolutional Neural Network. D. Stiehl, L. Addams, Luiz Eduardo Soares de Oliveira, Cayley Guimaraes, Alceu S. Britto Jr. |
ICDAR | 3 |
| 2015 | Transfer learning between texture classification tasks using Convolutional Neural NetworksabstractConvolutional Neural Networks (CNNs) have set the state-of-the-art in many computer vision tasks in recent years. For this type of model, it is common to have millions of parameters to train, commonly requiring large datasets. We investigate a method to transfer learning across different texture classification problems, using CNNs, in order to take advantage of this type of architecture to problems with smaller datasets. We use a Convolutional Neural Network trained on a source dataset (with lots of data) to project the data of a target dataset (with limited data) onto another feature space, and then train a classifier on top of this new representation. Our experiments show that this technique can achieve good results in tasks with small datasets, by leveraging knowledge learned from tasks with larger datasets. Testing the method on the the Brodatz-32 dataset, we achieved an accuracy of 97.04% - superior to models trained with popular texture descriptors, such as Local Binary Patterns and Gabor Filters, and increasing the accuracy by 6 percentage points compared to a CNN trained directly on the Brodatz-32 dataset. We also present a visual analysis of the projected dataset, showing that the data is projected to a space where samples from the same class are clustered together - suggesting that the features learned by the CNN in the source task are relevant for the target task. Luiz G. Hafemann, Luiz Eduardo Soares de Oliveira, Paulo Rodrigo Cavalin, Robert Sabourin |
IJCNN | 2 |
| 2015 | Combining overall and local class accuracies in an oracle-based method for dynamic ensemble selectionabstractThis paper presents a k-nearest oracle-based dynamic ensemble selection method in which overall local accuracy (OLA) and local class accuracy (LCA) are combined into a twostep selection scheme. The OLA and LCA are computed on the neighborhood of the test pattern in a validation set to filter out the classifiers selected by the k-nearest oracles. The complementary information of OLA and LCA has shown to be an interesting alternative to approximate the classification performance to that estimated for the oracle of the initial pool of classifiers. The results were compared with the recognition rates of the majority voting of all classifiers in the initial pool, and also with the recognition rates of related classifier and ensemble selection methods which have inspired the proposed method and its variants. The proposed method achieved the best results on 5 out of 8 experiments using small and large datasets of different applications. Leila Maria Vriesmann, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Robert Sabourin |
IJCNN | 3 |
| 2015 | SIFT Applied to Perceptual Zoning for Trademark RetrievalabstractThe paper contributes to the CBIR systems applied to trademark retrieval. The proposed model uses Scale Invariant Feature Transform (SIFT) and includes aspects from visual perception of the shapes, by means of feature extractor associated to a non-symmetrical perceptual zoning mechanism based on the Principles of Gestalt. We carried out experiments using four different zonings strategies for matching and retrieval tasks. The proposed method achieved the normalized recall (Rn) equal to 0.84. Experiments show that the non-symmetrical zoning could be considered as a tool to build more reliable trademark retrieval systems. Simone B. K. Aires, Cinthia Obladen de Almendra Freitas, Luiz Eduardo Soares de Oliveira |
SMC | 3 |
| 2015 | PKLot - A robust dataset for parking lot classification
Paulo R. L. Almeida, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Eunelson Jose da Silva Junior, Alessandro L. Koerich |
Expert Syst. Appl. | 2 |
| 2015 | Forest species recognition based on dynamic classifier selection and dissimilarity feature vector representation
Jefferson G. Martins, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
Mach. Vis. Appl. | 2 |
| 2015 | A new algorithm for number of holes attribute filtering of grey-level images
Juan Climent, Luiz Eduardo Soares de Oliveira |
Pattern Recognit. Lett. | 2 |
| 2014 | ICFHR 2014 Competition on Handwritten Digit String Recognition in Challenging Datasets (HDSRC 2014)abstractThis paper presents the results of the HDSRC 2014 competition on handwritten digit string recognition in challenging datasets organized in conjunction with ICFHR 2014. The general objective of this competition is to identify, evaluate and compare recent developments in Western Arabic digit string recognition with varying length. In addition, this competition introduces two new challenging datasets for benchmarking. We describe competition details including the datasets and evaluation measures used, and give a comparative performance analysis of six (6) participating methods along with a short description of the respective methodologies. Markus Diem, Stefan Fiel, Florian Kleber, Robert Sablatnig, José M. Saavedra, David Contreras, Juan Manuel Barrios, Luiz Eduardo Soares de Oliveira |
ICFHR | 8 |
| 2014 | Assessing Textural Features for Writer Identification on Different Writing Styles and ForgeriesabstractIn this study we assess the performance of textural descriptors for writer identification on different writing styles and also on forgeries. To do that, we have performed a series of experiments using the Fire maker database, which provides for the same writer texts written on three different writing styles and also copied forged text. Our experimental protocol is based on the dissimilarity framework and SVM classifiers, which were trained with LBP (Local Binary Pattern) and LPQ (Local Phase Quantization). The 250 writers of the database were divided into different configurations to observe the impacts of different sizes of the training set on the performance of the system. Our experimental results corroborates the fact that the texture is an interesting alternative for writer identification. The classifier trained with LPQ was able to produce error rates 23 percentage points smaller than those reported in the literature for upper-case and free writing styles. Regarding the forgeries, the LPQ-based classifier goes further reducing the error rate up to 44 percentage points depending on the writing style used for training. Diego Bertolini, Luiz Eduardo Soares de Oliveira, Edson José Rodrigues Justino, Robert Sabourin |
ICPR | 2 |
| 2014 | Forest Species Recognition Using Deep Convolutional Neural NetworksabstractForest species recognition has been traditionally addressed as a texture classification problem, and explored using standard texture methods such as Local Binary Patterns (LBP), Local Phase Quantization (LPQ) and Gabor Filters. Deep learning techniques have been a recent focus of research for classification problems, with state-of-the art results for object recognition and other tasks, but are not yet widely used for texture problems. This paper investigates the usage of deep learning techniques, in particular Convolutional Neural Networks (CNN), for texture classification in two forest species datasets - one with macroscopic images and another with microscopic images. Given the higher resolution images of these problems, we present a method that is able to cope with the high-resolution texture images so as to achieve high accuracy and avoid the burden of training and defining an architecture with a large number of free parameters. On the first dataset, the proposed CNN-based method achieves 95.77% of accuracy, compared to state-of-the-art of 97.77%. On the dataset of microscopic images, it achieves 97.32%, beating the best published result of 93.2%. Luiz G. Hafemann, Luiz Eduardo Soares de Oliveira, Paulo Rodrigo Cavalin |
ICPR | 2 |
| 2014 | An HMM-Based Gesture Recognition Method Trained on Few SamplesabstractThis paper addresses the problem of recognizing gestures which are captured using the Kinect sensor in a educational game devoted to the deaf community. Different strategies are evaluated to deal with the problem of having few samples for training. We have experimented a Leave One Out Training and Testing (LOOT) strategy and an HMM-based ensemble of classifiers. A dataset containing 181 videos of gestures related to nine signs commonly used in educational games is introduced, which is available for research purposes. The experimental results have shown that the proposed ensemble-based method is a promising strategy to deal with problems where few training samples are available. Vinicius Godoy, Alceu S. Britto Jr., Alessandro L. Koerich, Jacques Facon, Luiz Eduardo Soares de Oliveira |
ICTAI | 5 |
| 2014 | Automatic forest species recognition based on multiple feature setsabstractIn this paper we investigate the use of multiple feature sets for automatic forest species recognition. In order to accomplish this, different feature sets are extracted, evaluated, and combined into a framework based on two approaches: image segmentation and multiple feature sets. The experimental results on microscopic and macroscopic images of wood indicate that the recognition rates can be improved from 74.58% to about 95.68% and from 68.69% to 88.90%, respectively. In addition, they reveal us the importance of exploring different window sizes and appropriate local estimation functions for the LPQ descriptor, further than the classical uniform and gaussian functions. Marcelo N. Kapp, Rodrigo Bloot, Paulo Rodrigo Cavalin, Luiz Eduardo Soares de Oliveira |
IJCNN | 4 |
| 2014 | Forest species recognition using macroscopic images
Pedro Luiz de Paula Filho, Luiz Eduardo Soares de Oliveira, Silvana Nisgoski, Alceu S. Britto Jr. |
Mach. Vis. Appl. | 2 |
| 2014 | Dynamic selection of classifiers - A comprehensive review
Alceu S. Britto Jr., Robert Sabourin, Luiz Eduardo Soares de Oliveira |
Pattern Recognit. | 3 |
| 2013 | Music Genre Recognition Using Gabor Filters and LPQ Texture Descriptors
Yandre M. G. Costa, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Fabien Gouyon |
CIARP (2) | 2 |
| 2013 | Text Line Detection in Document Images: Towards a Support System for the BlindabstractWe introduce a novel approach for text line detection in document images, keeping in mind the requirements of a portable text recognition system designed to support the blind. Challenges include shadows, cluttered backgrounds, and perspective distortion. Different from previous approaches, the proposed method does not segment the image. A text model is created by clustering SIFT features extracted from positive and negative examples. Text regions are located by matching the features extracted from the input image to the clusters in the text model. Regions around the correspondences are then analyzed, and text lines are identified based on features such as gradients and histogram distribution. Experimental results show that our approach outperforms a state-of-the-art text detector in a text/non-text classification task. Bogdan Tomoyuki Nassu, Rodrigo Minetto, Luiz Eduardo Soares de Oliveira |
ICDAR | 3 |
| 2013 | Parking Space Detection Using Textural DescriptorsabstractIn this paper we assess the use of textural de-scriptors for the problem of parking space detection. We focus our experiments on two descriptors (Local Binary Patterns and Local Phase Quantization) that have attracted a great deal of attention because of their outstanding performance in a number of applications. We show through a series of comprehensive experiments that both descriptors are able to achieve very low error rates on a database composed of 105,837 images of parking spaces. We also show that the combination of the diverse classifiers developed in this work can bring further improvement achieving an error rate of 0.16%. The results reached in this work compare favorably to other published methods. Paulo R. L. Almeida, Luiz Eduardo Soares de Oliveira, Eunelson Jose da Silva Junior, Alceu S. Britto Jr., Alessandro L. Koerich |
SMC | 2 |
| 2013 | LIBRAS Sign Language Hand Configuration Recognition Based on 3D MeshesabstractThis paper presents a method for recognizing hand configurations of the Brazilian sign language (LIBRAS) using 3D meshes and 2D projections of the hand. Five actors performing 61 different hand configurations of the LIBRAS language were recorded twice, and the videos were manually segmented to extract one frame with a frontal and one with a lateral view of the hand. For each frame pair, a 3D mesh of the hand was constructed using the Shape from Silhouette method, and the rotation, translation and scale invariant Spherical Harmonics method was used to extract features for classification. A Support Vector Machine (SVM) achieved a correct classification of Rank1 = 86.06% and Rank3 = 96.83% on a database composed of 610 meshes. SVM classification was also performed on a database composed of 610 image pairs using 2D horizontal and vertical projections as features, resulting in Rank1 = 88.69% and Rank3 = 98.36%. Results encourage the use of 3D meshes as opposed to videos or images, given that their direct, real time acquisition is becoming possible due to devices like Leap Motion® or high resolution depth cameras. Andres Jessé Porfirio, Kelly Lais Wiggers, Luiz Eduardo Soares de Oliveira, Daniel Weingaertner |
SMC | 3 |
| 2013 | Texture-based descriptors for writer identification and verification
Diego Bertolini, Luiz Eduardo Soares de Oliveira, Edson José Rodrigues Justino, Robert Sabourin |
Expert Syst. Appl. | 2 |
| 2013 | Fusion of feature sets and classifiers for facial expression recognition
Thiago H. H. Zavaschi, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich |
Expert Syst. Appl. | 3 |
| 2013 | Handwritten digit segmentation: a comparative study
F. C. Ribas, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
Int. J. Document Anal. Recognit. | 2 |
| 2013 | A database for automatic classification of forest species
Jefferson G. Martins, Luiz Eduardo Soares de Oliveira, Silvana Nisgoski, Robert Sabourin |
Mach. Vis. Appl. | 2 |
| 2012 | Authorship Attribution of Electronic Documents Comparing the Use of Normalized Compression Distance and Support Vector Machine in Authorship Attribution
Walter Ribeiro de Oliveira Jr., Edson José Rodrigues Justino, Luiz Eduardo Soares de Oliveira |
ICONIP (1) | 3 |
| 2012 | Comparing textural features for music genre classificationabstractIn this paper we compare two different textural feature sets for automatic music genre classification. The idea is to convert the audio signal into spectrograms and then extract features from this visual representation. Two textural descriptors are explored in this work: the Gray Level Co-Occurrence Matrix (GLCM) and Local Binary Patterns (LBP). Besides, two different strategies of extracting features are considered: a global approach where the features are extracted from the entire spectrogram image and then classified by a single classifier; a local approach where the spectrogram image is split into several zones which are classified independently and final decision is then obtained by combining all the partial results. The database used in our experiments was the Latin Music Database, which contains music pieces categorized into 10 musical genres, and has been used for MIREX (Music Information Retrieval Evaluation eXchange) competitions. After a comprehensive series of experiments we show that the SVM classifier trained with LBP is able to achieve a recognition rate of 80%. This rate not only outperforms the GLCM by a fair margin but also is slightly better than the results reported in the literature. Yandre M. G. Costa, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Fabien Gouyon |
IJCNN | 2 |
| 2012 | Writer verification using texture-based features
Regiane Kowalek Hanusiak, Luiz Eduardo Soares de Oliveira, Edson José Rodrigues Justino, Robert Sabourin |
Int. J. Document Anal. Recognit. | 2 |
| 2012 | Music genre classification using LBP textural features
Yandre M. G. Costa, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Fabien Gouyon, Jefferson G. Martins |
Signal Process. | 2 |
| 2011 | Facial expression recognition using ensemble of classifiersabstractThis paper presents a novel method for facial expression classification that employs the combination of two different feature sets in an ensemble approach. A pool of base classifiers is created using two feature sets: Gabor filters and local binary patterns (LBP). Then a multi-objective genetic algorithm is used to search for the best ensemble using as objective functions the accuracy and the size of the ensemble. The experimental results on two databases have shown the efficiency of the proposed strategy by finding powerful ensembles, which improves the recognition rates between 5% and 10%. Thiago H. H. Zavaschi, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira |
ICASSP | 3 |
| 2011 | Selecting syntactic attributes for authorship attributionabstractIn this work we present a methodology to select syntactic attributes for authorship attribution. The approach takes into account a multi-objective genetic algorithm and a Support Vector Machine classifier and it operates in a wrapper mode. Through a series of comprehensive experiments on a database composed of 3000 short articles written in Portuguese we show that the proposed methodology is able to provide a concise subset of attributes, which increases the recognition rate in about 15 percentage points. Paulo Varela, Edson José Rodrigues Justino, Luiz Eduardo Soares de Oliveira |
IJCNN | 3 |
| 2011 | A framework to support development of Sign Language human-computer interaction: Building tools for effective information access and inclusion of the deafabstractSign Languages are tools the deaf use for their communication, education, information access needs, among others. Information Systems, whose role should be to facilitate those processes, still do not present a natural interaction for the deaf. There are many attempts by Computer Vision researches that are limited in their approach, their object of study, their lack of end use results etc. The challenge is to devise a framework with which to work towards addressing those shortcomings. The present study presents such a framework to support sign language recognition and interaction to serve as “de facto” standard that should be used by Computer Vision in order to claim back the field's noble task of developing effective technological services that take the deaf's needs into consideration towards social inclusion. Diego R. Antunes, Cayley Guimaraes, Laura Sánchez García, Luiz Eduardo Soares de Oliveira, Sueli Fernandes |
RCIS | 4 |
| 2010 | Verification of Unconstrained Handwritten Words at Character LevelabstractIn this paper we present a verification module that has as input the output provided by a word recognizer which is based on the segmentation-recognition paradigm. The word recognizer models words as the concatenation of character hidden Markov models (HMMs) and it provides at the output a list with the Top N best word hypotheses, including their likelihoods and the segmentation points of the words into sub words, which ideally should be characters. The verification module uses the segmentation points provided by the word recognizer for each word hypothesis to extract different features from each sub word. A classifier based on a multilayer perceptron neural network assigns a character class (A-Z) and estimates the a posteriori probability to each sub word that make up a word. Further, both the character class and the a posteriori probabilities are combined with the original output of the word recognizer to re-rank the word hypothesis into the Top N list. Experimental results show that the verification module improves the Top 1 recognition rate in 3.9% for an 85,092-word recognition task. Alessandro L. Koerich, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira |
ICFHR | 3 |
| 2010 | Selection of Training Instances for Music Genre ClassificationabstractIn this paper we present a method for the selection of training instances based on the classification accuracy of a SVM classifier. The instances consist of feature vectors representing short-term, low-level characteristics of music audio signals. The objective is to build, from only a portion of the training data, a music genre classifier with at least similar performance as when the whole data is used. The particularity of our approach lies in a pre-classification of instances prior to the main classifier training: i.e. we select from the training data those instances that show better discrimination with respect to class memberships. On a very challenging dataset of 900 music pieces divided among 10 music genres, the instance selection method slightly improves the music genre classification in 2.4 percentage points. On the other hand, the resulting classification model is significantly reduced, permitting much faster classification over test data. Miguel Lopes, Fabien Gouyon, Alessandro L. Koerich, Luiz Eduardo Soares de Oliveira |
ICPR | 4 |
| 2010 | Forest Species Recognition Using Color-Based FeaturesabstractIn this work we address the problem of forest species recognition which is a very challenging task and has several potential applications in the wood industry. The first contribution of this work is a database composed of 22 different species of the Brazilian flora that has been carefully labeled by expert in wood anatomy. In addition, in this work we demonstrate through a series of comprehensive experiments that color-based features are quite useful to increase the discrimination power for this kind of application. Last but not least, we propose a segmentation approach so that a wood can be locally processed to mitigate the intra-class variability featured in some classes. Such an approach also brings important contribution to improve the final performance in terms of classification. Pedro Luiz de Paula Filho, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
ICPR | 2 |
| 2010 | Reducing forgeries in writer-independent off-line signature verification through ensemble of classifiers
Diego Bertolini, Luiz Eduardo Soares de Oliveira, Edson José Rodrigues Justino, Robert Sabourin |
Pattern Recognit. | 2 |
| 2009 | Document reconstruction using dynamic programmingabstractIn this work we propose a methodology for document reconstruction based on dynamic programming and a modified version of the Prim's algorithm. Firstly, we use polygonal approximation to reduce the complexity of the boundaries and extract features from them. Thereafter, these features are used to feed the LCS dynamic programming algorithm. The scores yielded by the LCS algorithm are then used into a modified Prim's algorithm to find the best match among all pieces. Comprehensive experiments on a database composed of 100 shredded documents support the efficiency of the proposed methodology. When compared to global search algorithms, this approach brings an improvement of 18% in the number of fragments reconstructed. Andre Pimenta, Edson José Rodrigues Justino, Luiz Eduardo Soares de Oliveira, Robert Sabourin |
ICASSP | 3 |
| 2009 | Author Identification Using Compression ModelsabstractIn this paper we discuss the use of compression algorithms for author identification. We present the basic background about compression algorithms and introduce the prediction by partial matching algorithm, which has been used in our experiments. To better compare the results produced by the PPM algorithm, we present some experiments using stylometric features used very often by forensic examiners. In this case the authors are modeled using support vector machines. Comprehensive experiments performed on a database composed of 20 different authors show that the PPM algorithm is an interesting alternative for author identification, since all the process of feature definition, extraction, and selection can be avoided. Daniel Pavelec, Luiz Eduardo Soares de Oliveira, Edson José Rodrigues Justino, Francisco D. Nobre Neto, Leonardo Vidal Batista |
ICDAR | 2 |
| 2009 | Evaluation of Different Strategies to Optimize an HMM-Based Character Recognition SystemabstractDifferent strategies for combination of complementary features in an HMM-based method for handwritten character recognition are evaluated. In addition, a noise reduction method is proposed to deal with the negative impact of low probability symbols in the training database. New sequences of observations are generated based on the original ones, but considering a noise reduction process. The experimental results based on 52 classes of alphabetic characters and more than 23,000 samples have shown that the strategies proposed to optimize the HMM-based recognition method are very promising. Murilo Santos, Albert Hung-Ren Ko, Luiz Eduardo Soares de Oliveira, Robert Sabourin, Alessandro L. Koerich, Alceu S. Britto Jr. |
ICDAR | 3 |
| 2009 | Compression and stylometry for author identificationabstractIn this paper we compare two different paradigms for author identification. The first one is based on compression algorithms where the entire process of defining and extracting features and training a classifier is avoided. The second paradigm, on the other hand, takes into account the classical pattern recognition framework, where linguistic features proposed by forensic experts are used to train a Support Vector Machine classifier. Comprehensive experiments performed on a database composed of 20 writers show that both strategies achieve similar performance but with an interesting degree of complementarity demonstrated through the confusion matrices. Advantages and drawback of both paradigms are also discussed. Daniel Pavelec, Luiz Eduardo Soares de Oliveira, Edson José Rodrigues Justino, Francisco D. Nobre Neto, Leonardo Vidal Batista |
IJCNN | 2 |
| 2009 | Combining different biometric traits with one-class classification
Cheila Bergamini, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Robert Sabourin |
Signal Process. | 2 |
| 2008 | Overfitting in the selection of classifier ensembles: a comparative study between PSO and GAabstractClassifier ensemble selection may be formulated as a learning task since the search algorithm operates by minimizing/maximizing the objective function. As a consequence, the selection process may be prone to overfitting. The objectives of this paper are: (1) to show how overfitting can be detected when the selection is performed by two classical search algorithms: Genetic Algorithm and Particle Swarm Optimization; and (2) to verify which algorithm is more prone to overfitting. The experimental results demonstrate that GA appears to be more affected by overfitting. Eulanda M. dos Santos, Luiz Eduardo Soares de Oliveira, Robert Sabourin, Patrick Maupin |
GECCO | 2 |
| 2008 | The implication of data diversity for a classifier-free ensemble selection in random subspacesabstractEnsemble of Classifiers (EoC) has been shown effective in improving the performance of single classifiers by combining their outputs. By using diverse data subsets to train classifiers, the ensemble creation methods can create diverse classifiers for the EoC. In this work, we propose a scheme to measure the data diversity directly from random subspaces and we explore the possibility of using the data diversity directly to select the best data subsets for the construction of the EoC. The applicability is tested on NIST SD19 handwritten numerals. Albert Hung-Ren Ko, Robert Sabourin, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr. |
ICPR | 3 |
| 2008 | Fusion of biometric systems using one-class classificationabstractOne of the main requirements of biometric systems is the ability of producing very low false acceptation rate, which very often can be achieved only by combining different biometric traits. The literature has shown that the pattern classification approach usually surpasses the classifier combination approach for this task. In this work we take into account the pattern classification approach, but considering the one-class classification approach. We show that one-class classification could be considered as an alternative for biometric fusion specially when the data is highly unbalanced or data from a single class is available. The results for one-class classification reported in this paper compares to the standard two-class SVM and surpasses all the conventional classifier combination rules tested. Cheila Bergamini, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Robert Sabourin |
IJCNN | 2 |
| 2008 | Ensemble of classifiers for off-line signature verificationabstractIn this work we address two important issues of off-line signature verification. The first one regards feature extraction. We introduce a new graphometric feature set that considers the curvature of the main strokes of the signature. The idea is to simulate the shape of the signature by using Bezier curves and then extract features from these curves. The second important aspect is the use of an ensemble of classifiers based on graphometric features to improve the reliability of the classification, hence reducing the false acceptance. The ensemble was built using a standard genetic algorithm and different fitness functions were assessed to drive the search. Thorough experiments were conduct on a database composed of 100 writers and the results compare favorably. Diego Bertolini, Luiz Eduardo Soares de Oliveira, Edson José Rodrigues Justino, Robert Sabourin |
SMC | 2 |
| 2008 | Filtering segmentation cuts for digit string recognition
Eduardo Vellasques, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Alessandro L. Koerich, Robert Sabourin |
Pattern Recognit. | 2 |
| 2007 | Off-line Signature Verification Using Writer-Independent ApproachabstractIn this work we present a strategy for off-line signature verification. It takes into account a writer-independent model which reduces the pattern recognition problem to a 2-class problem, hence, makes it possible to build robust signature verification systems even when few signatures per writer are available. Receiver operating characteristic (ROC) curves are used to improve the performance of the proposed system . The contribution of this paper is two-fold. First of all, we analyze the impacts of choosing different fusion strategies to combine the partial decisions yielded by the SVM classifiers. Then ROC produced by different classifiers are combined using maximum likelihood analysis, producing an ROC combined classifier. Through comprehensive experiments on a database composed of 100 writers, we demonstrate that the ROC combined classifier based on the writer-independent approach can reduce considerably false rejection rate while keeping false acceptance rates at acceptable levels. Luiz Eduardo Soares de Oliveira, Edson José Rodrigues Justino, Robert Sabourin |
IJCNN | 1 |
| 2007 | Detection and Classification of Human Movements in Video Scenes
Andre G. Hochuli, Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Alessandro L. Koerich |
PSIVT | 2 |
| 2007 | People Counting in Low Density Video Sequences
Jaime Dalla Valle, Luiz Eduardo Soares de Oliveira, Alessandro L. Koerich, Alceu S. Britto Jr. |
PSIVT | 2 |
| 2007 | Handwritten Character Recognition Using Nonsymmetrical Perceptual ZoningabstractIn this paper we present an alternative strategy to define zoning for handwriting recognition, which is based on nonsymmetrical perceptual zoning. The idea is to extract some knowledge from the confusion matrices in order to make the zoning process less empirical. The feature set considered in this work is based on concavities/convexities deficiencies, which are obtained by labeling the background pixels of the input image. To better assess the nonsymmetrical zoning we carried out experiments using four different zonings strategies. Experiments show that the nonsymmetrical zoning could be considered as a tool to build more reliable handwriting recognition systems. Cinthia Obladen de Almendra Freitas, Luiz Eduardo Soares de Oliveira, Flávio Bortolozzi, Simone B. K. Aires |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2007 | Pairwise fusion matrix for combining classifiers
Albert Hung-Ren Ko, Robert Sabourin, Alceu S. Britto Jr., Luiz Eduardo Soares de Oliveira |
Pattern Recognit. | 4 |
| 2006 | Particle Swarm Optimization of Fuzzy ARTMAP ParametersabstractIn this paper a particle swarm optimization (PSO)-based training strategy is introduced for fuzzy ARTMAP that minimizes generalization error while optimizing parameter values. Through a comprehensive set simulations, it has been shown that this training strategy allows fuzzy ARTMAP to achieve a significantly lower generalization error than when it uses typical training strategies. Furthermore, the PSO strategy eliminates degradation of generalization error due to overtraining resulting from the training set size, number of training epochs, and data set structure. Overall results obtained with the PSO strategy reveal the importance of optimizing parameters and weights using a consistent objective function. In fact, the parameters found using this strategy vary significantly according to, e.g., training set size and data set structure, and always differ considerably from the popular choice of parameters that allows to minimize resources. Eric Granger, Philippe Henniges, Luiz Eduardo Soares de Oliveira, Robert Sabourin |
IJCNN | 3 |
| 2006 | Feature selection for ensembles applied to handwriting recognition
Luiz Eduardo Soares de Oliveira, Marisa E. Morita, Robert Sabourin |
Int. J. Document Anal. Recognit. | 1 |
| 2005 | Multi-objective Genetic Algorithms to Create Ensemble of Classifiers
Luiz Eduardo Soares de Oliveira, Marisa E. Morita, Robert Sabourin, Flávio Bortolozzi |
EMO | 1 |
| 2005 | A Synthetic Database to Assess Segmentation AlgorithmsabstractIn this paper we describe a synthetic database composed of 273,452 handwritten touching digits pairs to assess segmentation algorithms. It contains several different kinds of touching and it was generated by connecting 2,000 images of isolated digits extracted from the NIST SD19. In order to get a better insight on the proposed database and establish some parameters for further comparisons, we carried out experiments using four state-of-the-art segmentation algorithms. Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
ICDAR | 1 |
| 2005 | Improving Cascading Classifiers with Particle Swarm OptimizationabstractThis paper addresses the issue of class related reject thresholds for cascading classifier systems. It has been demonstrated in the literature that class related reject thresholds provide an error-reject tradeoff better than a single global threshold. In this work we argue that the error-reject tradeoff yielded by class-related reject thresholds can be further improved if a proper algorithm is used to find the thresholds. In light of this, we propose using a recently developed optimization algorithm called particle swarm optimization. It has been proved to be very effective in solving real valued global optimization problems. In order to show the benefits of such an algorithm, we have applied it to optimize the thresholds of a cascading classifier system devoted to recognize handwritten digits. Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
ICDAR | 1 |
| 2005 | Optimizing class-related thresholds with particle swarm optimizationabstractIn this paper we address the issue of class-related reject thresholds for classification systems. It has been demonstrated in the literature that class related reject thresholds provide an error-reject tradeoff better than a single global threshold. In this work we argue that the error-reject tradeoff yielded by class related reject thresholds can be further improved if a proper algorithm is used to find the thresholds. In light of this, we propose using a recently developed optimization algorithm called particle swarm optimization. It has been proved to be very effective in solving real valued global optimization problems. In order to show the benefits of such an algorithm, we have applied it to optimize the thresholds of a cascading classifier system devoted to recognize handwritten digits. Luiz Eduardo Soares de Oliveira, Alceu S. Britto Jr., Robert Sabourin |
IJCNN | 1 |
| 2003 | Feature Selection for Ensembles: A Hierarchical Multi-Objective Genetic Algorithm ApproachabstractFeature selection for ensembles has shown to be an effective strategy for ensemble creation. In this paper we present an ensemble feature selection approach based on a hierarchical multi-objective genetic algorithm. The first level performs feature selection in order to generate a set of good classifiers while the second one combines them to provide a set of powerful ensembles. The proposed method is evaluated in the context of handwritten digit recognition, using three different feature sets and neural networks (MLP) as classifiers. Experiments conducted on NIST SD19 demonstrated the effectiveness of the proposed strategy. Luiz Eduardo Soares de Oliveira, Robert Sabourin, Flávio Bortolozzi, Ching Y. Suen |
ICDAR | 1 |
| 2003 | Intelligent Zoning Design Using Multi-Objective Evolutionary AlgorithmsabstractThis paper discusses the use of multi objective evolutionary algorithms applied to the engineering of zoning for handwriten recognition. Usually a task fulfilled by an human expert, zoning design relies on specific domain knowledge and a trial and error process to select an adequate design. Our proposed approach to automatically define the zone design was tested and was able to define zoning strategies that performed better than our former strategy defined manually. Paulo Vinicius Wolski Radtke, Luiz Eduardo Soares de Oliveira, Robert Sabourin, Tony Wong |
ICDAR | 2 |
| 2003 | A Methodology for Feature Selection Using Multiobjective Genetic Algorithms for Handwritten Digit String RecognitionabstractIn this paper a methodology for feature selection for the handwritten digit string recognition is proposed. Its novelty lies in the use of a multiobjective genetic algorithm where sensitivity analysis and neural network are employed to allow the use of a representative database to evaluate fitness and the use of a validation database to identify the subsets of selected features that provide a good generalization. Some advantages of this approach include the ability to accommodate multiple criteria such as number of features and accuracy of the classifier, as well as the capacity to deal with huge databases in order to adequately represent the pattern recognition problem. Comprehensive experiments on the NIST SD19 demonstrate the feasibility of the proposed methodology. Luiz Eduardo Soares de Oliveira, Robert Sabourin, Flávio Bortolozzi, Ching Y. Suen |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2003 | Impacts of verification on a numeral string recognition system
Luiz Eduardo Soares de Oliveira, Robert Sabourin, Flávio Bortolozzi, Ching Y. Suen |
Pattern Recognit. Lett. | 1 |
| 2002 | Automatic Recognition of Handwritten Numerical Strings: A Recognition and Verification StrategyabstractA modular system to recognize handwritten numerical strings is proposed. It uses a segmentation-based recognition approach and a recognition and verification strategy. The approach combines the outputs from different levels such as segmentation, recognition, and postprocessing in a probabilistic model. A new verification scheme which contains two verifiers to deal with the problems of oversegmentation and undersegmentation is presented. A new feature set is also introduced to feed the oversegmentation verifier. A postprocessor based on a deterministic automaton is used and the global decision module makes an accept/reject decision. Finally, experimental results on two databases are presented: numerical amounts on Brazilian bank checks and NIST SD19. The latter aims at validating the concept of modular system and showing the robustness of the system using a well-known database. Luiz Eduardo Soares de Oliveira, Robert Sabourin, Flávio Bortolozzi, Ching Y. Suen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | A Modular System to Recognize Numerical Amounts on Brazilian Bank ChecksabstractThe paper presents a modular system to recognize numerical amounts on Brazilian bank cheques. The system uses a segmentation-based recognition approach and the recognition function is based on a recognition and verification strategy. Our approach consists of combining the outputs from different levels such as segmentation, recognition and post-processing in a probabilistic model. A new feature set is introduced to the verifier module in order to detect segmentation effects such as over-segmentation and under-segmentation. Finally, we present experimental results on two databases: numerical amounts and NIST SD19. The latter aims at validating the concept of modular system and showing the robustness of the system over a well-known database. Luiz Eduardo Soares de Oliveira, Robert Sabourin, Flávio Bortolozzi, Ching Y. Suen |
ICDAR | 1 |
| 2000 | A New Segmentation Approach for Handwritten Digits
Luiz Eduardo Soares de Oliveira, Edouard Lethelier, Flávio Bortolozzi, Robert Sabourin |
ICPR | 1 |