Cleber Zanchettin

dblp:75/5629 · DBLP profile ↗
← Back
90ranked-venue papers
15as first author
31since 2021 · last 2026
0000-0001-6421-9747ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 69 · 14 first-author · 22 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive salience autoencoder for unsupervised anomaly detection in computer network traffic
Gabriel Coelho, Luís Gonçalves, Cleber Zanchettin, Lígia F. Braga, Diana R. Silva, Klarissa A. Morais
Eng. Appl. Artif. Intell.3
2025 Incremental learning approach using fuzzy logic to mitigate catastrophic forgetting
abstract
In this work, we propose a new approach that integrates Elastic Weight Consolidation (EWC) with fuzzy reasoning to address catastrophic forgetting in continual learning scenarios. The EWC-Fuzzy approach mitigates the challenge of forgetting previously learned knowledge while enabling the model to adapt to new tasks by balancing neural network weight regularization with fuzzy rule adaptation. Initially, the model learns from the first task without EWC regularization, allowing for standard backpropagation-based learning. For subsequent tasks, EWC is introduced to prevent significant parameter changes in the neural network that are critical for previous tasks. Meanwhile, the fuzzy rule parameters-such as the centers, widths, and outputs-dynamically evolve according to the new data without EWC regularization, allowing them to self-organize in response to the data distribution. This dual mechanism ensures that model preserves learned knowledge while remaining flexible and adaptable in the face of new tasks. Our approach addresses the gap in current research, which often treats EWC and fuzzy reasoning independently. By integrating these techniques, we provide a promising solution to the challenge of catastrophic forgetting and enhance the model’s adaptability in dynamic environments. This study lays the groundwork for further exploration into the fusion of EWC and fuzzy systems in continual learning.
Lívia de Souza Alexandre, Cleber Zanchettin
AICCSA2
2025 The Di2Win Document Intelligence Platform
abstract
We present the Di2Win Document Intelligence Platform (DIP). This modular AI-driven pipeline transforms raw document images --- captured by scanners or mobile phones --- into structured data and business actions in a single pass. The system comprises five loosely-coupled micro-services: (1) image-quality verification using a contrast-invariant model that flags blur, skew, and illumination issues above 100 ms per page; (2) document classification via a Transformer-base model with layout embeddings, delivering top-k types with calibrated confidence; (3) information extraction through i) Dilbert, a multimodal Token-Layout-Language model fine-tuned on weakly-labeled forms or ii) Delfos, a Large Language Model Mixture of Experts fine-tuned with well-defined prompts; (4) DataDrift, a powerful rules engine to avoid inconsistent outputs concerning the business process; and (5) process automation orchestrated by a Business Process Model Notation (BPMN) plus a Robot Process Automation (RPA) engine that routes results to databases, APIs, or human-review queues. All AI components are orchestrated through a messaging service to control the information flow, and the application exposes REST/gRPC endpoints to communicate with outside consumers. This enables the hot-swapping of models without downstream code changes by plugging a new message consumer into the messaging system. This also provides horizontal scalability since to increase the application throughput, we only need to add new AI engine consumers to the messaging system. Deployed in banking, insurance, and healthcare, the Di2Win DIP has processed more than 30 million pages, reducing average handling time by 79% and re-keying errors by 86 %, speeding up the workflows up to ten times. Our DocEng demonstration allows attendees to upload documents, observe live quality and confidence dashboards, and edit extracted fields with immediate feedback to the active-learning loop.
Afonso Ferreira, Cleber Zanchettin, Romulo Andrade, Byron L. D. Bezerra
DocEng2
2025 Breaking the Barrier of Hard Samples: A Data-Centric Approach to Synthetic Data for Medical Tasks
abstract
Data scarcity and quality issues remain significant barriers to developing robust predictive models in medical research. Traditional reliance on real-world data often leads to biased models with poor generalizability across diverse patient populations. Synthetic data generation has emerged as a promising solution, yet challenges related to these sample’s representativeness and effective utilization persist. This paper introduces Profile2Gen, a novel data-centric framework designed to guide the generation and refinement of synthetic data, focusing on addressing hard-to-learn samples in regression tasks. We conducted approximately 18,000 experiments to validate its effectiveness across six medical datasets, utilizing seven state-of-the-art generative models. Results demonstrate that refined synthetic samples can reduce predictive errors and enhance model reliability. Additionally, we generalize the DataIQ framework to support regression tasks, enabling its application in broader contexts. Statistical analyses confirm that our approach achieves equal or superior performance compared to models trained exclusively on real data.
Maynara Donato de Souza, Cleber Zanchettin
ICML2
2025 Improving COVID-19 Detection in Chest X-Rays Using EfficientNet with Self-Supervised Contrastive Learning
abstract
We propose a self-supervised framework for COVID-19 detection from chest X-rays (CXRs) that combines EfficientNet with four contrastive learning techniques: DINOCXR, BYOL, SimCLR, and VICReg. By pretraining on over 13,000 unlabeled CXRs from the ChestX-ray14-v3 dataset and fine-tuning on the COVIDGR dataset, our approach addresses the challenge of scarce labeled data in medical imaging. Experimental results show that SimCLR achieves the highest recall (77.78%), making it well-suited for initial screening, while EfficientNet finetuned directly provides the best precision$(87.18 \%)$, suggesting its use for confirmatory diagnosis. Our findings reveal that this twostage combination outperforms state-of-the-art models such as COVIDNet-CXR and DINO-CXR in terms of F1-score and data efficiency, using only 6% of the labeled data. This hybrid strategy offers a compelling balance between sensitivity and specificity, making it particularly valuable for clinical deployment.
Tales T. Alves, Ing Ren Tsang, Cleber Zanchettin
ICTAI3
2025 TE-CNN-AAE: Learning Robust Financial Time Series Representations with Trend-Enhanced Adversarial Autoencoders
abstract
This paper introduces the Trend-Enhanced CNN Adversarial Autoencoder (TE-CNN-AAE), a novel approach for generating low-dimensional representations of intraday stock market activity from 5-minute interval quotes. The proposed model extends traditional autoencoders by incorporating an adversarial component that enhances the quality of embeddings by reconstructing input sequences while simultaneously estimating market trends. Using five years of intraday data from Dow Jones Industrial Average (DJIA) assets, we demonstrate that TE-CNN-AAE outperforms baseline methods in capturing intraday patterns and predicting daily price movements. Qualitative evaluations via UMAP visualizations show improved class separability in the latent space, further confirmed by the Silhouette Score and Davies-Bouldin Index, while quantitative analysis using an LSTM classifier validates the superior predictive utility of the embeddings. Ablation tests confirm the importance of the adversarial discriminator in generating robust representations. Our results suggest that TE-CNN-AAE effectively captures the complex dynamics of financial time series and holds the potential for improving decision-support systems in trading.
Jefferson O. A. de Araujo, Adriano Lorena Inácio de Oliveira, Cleber Zanchettin
SMC3
2025 Deep contrastive variational subspace clustering
Marcos de Souza Oliveira, Sérgio Ricardo de Melo Queiroz, Cleber Zanchettin, Francisco de A. T. de Carvalho
Neurocomputing3
2024 DocLightDetect: A New Algorithm for Occlusion Classification in Identification Documents
Ricardo Batista das Neves Junior, Byron L. D. Bezerra, Cleber Zanchettin
DAS3
2024 How Does Changing the Optical Character Recognition System Impact the Layout-Aware Named Entity Recognition Models?
João Macedo, Byron L. D. Bezerra, Cleber Zanchettin
DAS3
2024 Classification of Dehiscence Defects in Titanium and Zirconium Dental Implants
Antonio Barros da Silva Netto, Willian Farias Carvalho Oliveira, Cleber Zanchettin
ICANN (8)3
2024 SG-RSRNN - Score Guided Robust Subspace Recovery-based Neural Network for Network Intrusion Detection
abstract
The current era of ubiquitous Internet connectivity has made network security a vital concern. A robust and safe network environment is crucial to protect users from malicious activities. Intrusion detection techniques play a valuable role in safeguarding IT infrastructure from malicious actors. Research teams have developed methods to identify patterns in network traffic that deviate from normal or expected behaviors, often considered outliers or anomalies. Network-based anomaly detection techniques can construct behavioral models to detect anomalous or suspicious activities in the network by leveraging machine learning approaches. Autoencoders are widely used in network anomaly detection models, utilizing reconstruction errors to identify attacks accurately. However, these models primarily rely on reconstruction errors, which might miss anomalies if the model learns to recreate the anomalous traffic precisely, adapting to reconstructing anomalous traffic could render the autoencoders ineffective in such cases. Moreover, the presence of both normal and abnormal samples within the data can obscure the identification of anomalies, particularly those residing in the transitional zones between data distributions. Introducing regularization methods can enforce robustness against anomalies, encouraging the model to learn fundamental underlying regularities capable of improving the distinction between normal and abnormal data. This work incorporates a special regularization method into the latent space of an autoencoder to enforce anomaly-robust structure. A scoring neural network is then used to improve the detection capabilities in the data transition zone by amplifying the differences between normal and abnormal data. The experimental results in relevant datasets show our proposal outperforms strong competitors in detecting anomalies in network traffic data.
Luís Gonçalves, Cleber Zanchettin
IJCNN2
2024 Planning the Path with Reinforcement Learning: Optimal Robot Motion Planning in RoboCup Small Size League Environments
Mateus G. Machado, João G. Melo, Cleber Zanchettin, Pedro H. M. Braga, Pedro V. Cunha, Edna Barros, Hansenclever de F. Bassani
RoboCup3
2024 Detecting abnormal logins by discovering anomalous links via graph transformers
Luís Gonçalves, Cleber Zanchettin
Comput. Secur.2
2024 PrAACT: Predictive Augmentative and Alternative Communication with Transformers
Jayr Pereira, Jaylton Alencar Pereira, Cleber Zanchettin, Robson do Nascimento Fidalgo
Expert Syst. Appl.3
2024 On the improvement of handwritten text line recognition with octave convolutional recurrent neural networks
Dayvid Castro, Cleber Zanchettin, Luis A. Nunes Amaral
Int. J. Document Anal. Recognit.2
2023 An Augmentative and Alternative Communication Synthetic Corpus for Brazilian Portuguese
abstract
In recent years, Augmentative and Alternative Communication (AAC) systems have grown significantly in Brazil, particularly for individuals with cognitive disorders who rely on high-tech AAC tools. Artificial Intelligence (AI) has significantly improved high-tech AAC systems by enhancing accessibility, increasing output generation speed, and improving AAC interfaces' customization and adaptability. This study investigates the use of Large Language Models (LLMs) to generate synthetic text data to augment a corpus for AAC in Brazilian Portuguese. A three-step method was used to augment an initial corpus of 667 AAC-like sentences produced by specialists to a corpus of 13k sentences, comprising sentence collection, corpus augmentation using GPT-3 in a few-shot setting, and corpus cleaning. The quality and reliability of the generated corpus were assessed through a coverage analysis, comparing the content of the generated sentences with the original human-composed sentences. The results provide insights into the methods' strengths and limitations and inform future efforts to improve the generation of synthetic text data for the AAC domain in Brazilian Portuguese.
Jayr Pereira, Rodrigo Nogueira 0001, Cleber Zanchettin, Robson do Nascimento Fidalgo
ICALT3
2023 Towards Background and Foreground Color Robustness with Adversarial Right for the Right Reasons
Flávio Arthur O. Santos, Maynara Donato de Souza, Cleber Zanchettin
ICANN (5)3
2023 Improving small object detection with DETRAug
abstract
Small object detection is a challenge for computer vision models due to a shortage of image details, textures, and varying distances from the camera, resulting in objects of different scales and partial occlusion issues. In this paper, we present a new method for enhancing the robustness of image detection models using AUGMIX. Our approach involves applying various augmentations to the input images in a stochastic manner, resulting in a single output image after all transformations have been applied. In addition, we used the Jensen-Shannon loss to maintain a more stable model. In our experiments, we observed a decrease in the number of “no-object” detections, which refers to the detection of unrelated or background objects. The new approach was evaluated using the Deformable DETR, a model known for detecting small objects accurately, and compared to DETR and EfficientDet. We verified an improvement of at least 4.15% using the proposed technique and a more stable loss error.
Evair Cunha, David Macedo, Cleber Zanchettin
IJCNN3
2023 Learning What, Where and Which to Transfer
abstract
Deep learning models often require large datasets to perform well from scratch. Transfer learning methods solve this issue by using a pre-trained source network to improve a target network training. Recent approaches involve using feature maps from the source network to guide the target network training. The latest transfer learning methods use meta-networks to enhance the knowledge transfer process. These meta-networks bridge the source and target networks, deciding which pairs of feature map layers and channels should be matched for optimal knowledge transfer. This paper improves this approach by using pixel-level information, in addition to layers and channels, for better knowledge transfer. Our experiments on multiple datasets show that the proposed approach outperforms previous baselines in scenarios with limited labels per class. The source code is available at https://github.com/lucasdelimanogueira/L2T-www.
Lucas de Lima Nogueira, David Macedo, Cleber Zanchettin, Fernando M. de Paula Neto, Adriano Lorena Inácio de Oliveira
IJCNN3
2023 Self-calibrated U-Net for Document Segmentation
abstract
Based on the need to digitalize identification documents per several institutions and companies, the segmentation task of textual information acquired importance. Commonly, Convolutional Neural Networks are applied to solve such problems. Vanilla convolution in deep learning models can provide sub-optimal performance in some learning tasks, such as semantic segmentation. Due to this, more specific proposals may be applied to better fit this context, such as self-calibrated convolutions, which consider unique characteristics of different feature maps. Considering this context, we propose a new fully convolutional network architecture based on U-Net for segmentation associated with self-calibrated convolution. We consider changing all vanilla convolution layers in this new neural network by self-calibrating convolutions. In this work, we evaluate our proposal in three different tasks of document segmentation. All experiments demonstrate an increase in the proposed neural network's segmentation performance compared with traditional U-Net in the analyzed context, but the inference time performance decreases.
Iago Richard Rodrigues, Leylane Ferreira, David Macedo, Cleber Zanchettin, Patricia Takako Endo, Djamel Fawzi Hadj Sadok
IJCNN4
2022 LogBERT-BiLSTM: Detecting Malicious Web Requests
Levi S. Ramos Júnior, David Macedo, Adriano Lorena Inácio de Oliveira, Cleber Zanchettin
ICANN (3)4
2022 Unsupervised Multi-view Multi-person 3D Pose Estimation Using Reprojection Error
Diógenes Wallis de França Silva, Joao Paulo Silva do Monte Lima, David Macedo, Cleber Zanchettin, Diego Thomas, Hideaki Uchiyama, Veronica Teichrieb
ICANN (3)4
2022 An Adapted GRASP Approach for Hyperparameter Search on Deep Networks Applied to Tabular Data
abstract
The robustness and resilience of the deep learning models offer consistent and competitive results in real-world applications. Despite its adaptability, the training and adjustment of the hyperparameters still demand knowledge and time from the designer. This paper proposes a simple and effective approach based on the Greedy Randomized Adaptive Search Procedure (GRASP) algorithm that we adapt to optimize deep neural networks models. We evaluated the performance of the proposed approach using the models Deep Feedforward Neural Network (DFNN) and TabNet, considering the Tabu Search algorithm as a baseline in five tabular datasets. Both optimization algorithms showed high performance regarding the (i) quality of the best solution, (ii) convergence, and (iii) local search. However, the adapted GRASP approach showed better results, optimizing the deep models in all datasets with statistical significance.
Andersson A. Da Silva, Amanda S. Xavier, David Macedo, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira
IJCNN4
2022 Multi-human Fall Detection and Localization in Videos
Mouglas Eugênio Nasário Gomes, David Macedo, Cleber Zanchettin, Paulo S. G. de Mattos Neto, Adriano Lorena Inácio de Oliveira
Comput. Vis. Image Underst.3
2022 PictoBERT: Transformers for next pictogram prediction
Jayr Pereira, David Macedo, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira, Robson do Nascimento Fidalgo
Expert Syst. Appl.3
2022 Entropic Out-of-Distribution Detection: Seamless Detection of Unknown Examples
abstract
In this article, we argue that the unsatisfactory out-of-distribution (OOD) detection performance of neural networks is mainly due to the SoftMax loss anisotropy and propensity to produce low entropy probability distributions in disagreement with the principle of maximum entropy. On the one hand, current OOD detection approaches usually do not directly fix the SoftMax loss drawbacks, but rather build techniques to circumvent it. Unfortunately, those methods usually produce undesired side effects (e.g., classification accuracy drop, additional hyperparameters, slower inferences, and collecting extra data). On the other hand, we propose replacing SoftMax loss with a novel loss function that does not suffer from the mentioned weaknesses. The proposed IsoMax loss is isotropic (exclusively distance-based) and provides high entropy posterior probability distributions. Replacing the SoftMax loss by IsoMax loss requires no model or training changes. Additionally, the models trained with IsoMax loss produce as fast and energy-efficient inferences as those trained using SoftMax loss. Moreover, no classification accuracy drop is observed. The proposed method does not rely on outlier/background data, hyperparameter tuning, temperature calibration, feature extraction, metric learning, adversarial training, ensemble procedures, or generative models. Our experiments showed that IsoMax loss works as a seamless SoftMax loss drop-in replacement that significantly improves neural networks' OOD detection performance. Hence, it may be used as a baseline OOD detection approach to be combined with current or future OOD detection techniques to achieve even higher results.
David Macedo, Ing Ren Tsang, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira, Teresa Bernarda Ludermir
IEEE Trans. Neural Networks Learn. Syst.3
2021 Construction of Brazilian Regulatory Traffic Sign Recognition Dataset
Bruno O. Prado, Leonardo N. Matos, Flávio Arthur O. Santos, Cleber Zanchettin, Paulo Novais
CIARP5
2021 Entropic Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection approaches usually present special requirements (e.g., hyperparameter validation, collection of outlier data) and produce side effects (e.g., classification accuracy drop, slower energy-inefficient inferences). We argue that these issues are a consequence of the SoftMax loss anisotropy and disagreement with the maximum entropy principle. Thus, we propose the IsoMax loss and the entropic score. The seamless drop-in replacement of the SoftMax loss by IsoMax loss requires neither additional data collection nor hyperparameter validation. The trained models do not exhibit classification accuracy drop and produce fast energy-efficient inferences. Moreover, our experiments show that training neural networks with IsoMax loss significantly improves their OOD detection performance. The IsoMax loss exhibits state-of-the-art performance under the mentioned conditions (fast energy-efficient inference, no classification accuracy drop, no collection of outlier data, and no hyperparameter validation), which we call the seamless OOD detection task. In future work, current OOD detection methods may replace the SoftMax loss with the IsoMax loss to improve their performance on the commonly studied non-seamless OOD detection problem.
David Macedo, Ing Ren Tsang, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira, Teresa Bernarda Ludermir
IJCNN3
2021 Multi-Class Mobile Money Service Financial Fraud Detection by Integrating Supervised Learning with Adversarial Autoencoders
abstract
Given the actual volume and speed of financial transactions, financial fraud detection systems are constantly evolving based on new computational intelligence algorithms. Therefore, transaction monitoring and analysis prevent monetary losses caused by fraudsters. Since the fraud detection process is a labor-intensive task for human auditors given the huge amount of daily transactions processed by financial services information systems. Credit card is the financial product most explored in the financial fraud detection literature, while mobile money service is becoming a popular option for payments, fraud detection for such financial product has not yet been deeply explored. Therefore, it is interesting to optimize the auditing process and test new quantitative techniques, such as deep learning, to support human auditors before double-checking a suspicious transaction. Thus, we propose an integration of adversarial autoencoders and machine learning methods to perform an objective classification among three transaction types: regular, local, and global anomaly. The integration consists of using the autoencoder's generated latent vectors as features for the supervised learning algorithms. The experiments considered different latent vector space forms concerning their dimensionality and the clusters generated by a prior Gaussian mixture. The results show that some classifiers may accept latent characteristics well, getting better or similar performance when using all the original characteristics.
Julio Cezar Soares Silva, David Macedo, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira, Adiel Almeida Filho
IJCNN3
2021 Identification of Microorganism Colony Odor Signature using InceptionTime
abstract
Microorganisms that cause infectious diseases are defined as pathogens, as they multiply and cause tissue damage. All microorganisms isolated in culture from a location on the body should be considered potential pathogens. The infectious processes demonstrate physiological responses to the multiplication invasion of the aggressor microorganism. The disease’s development is influenced by the patient’s general health, defense mechanisms, and previous contact with the offending agent. When an infectious disease is suspected, cultures should be performed. This article uses an electronic nose to collect and analyze volatile organic compounds VOCs expelled by colonies of microorganisms. We propose signature identification of these colony odors from microorganisms using InceptionTime. The InceptionTime model is a set of models of the deep convolutional neural network, inspired by the Inception-v4 architecture. The results were excellent, with an average accuracy in the test set above 98%. The aim of our research is to propose a faster, cheaper and more accurate method of detecting these pathogens and the encouraging results of this stage encourage further research.
Paulo M. Vasconcelos, David Macedo, Leandro M. Almeida, Reginaldo G. L. Neto, Clayton A. Benevides, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira
SMC6
2021 Intrusion Detection for Cyber-Physical Systems Using Generative Adversarial Networks in Fog Environment
abstract
Cyber-attacks cyber-physical systems (CPSs) can lead to sensing and actuation misbehavior, severe damages to physical objects, and safety risks. Machine learning algorithms have been proposed for hindering cyber-attacks on CPSs, but the absence of labeled data from novel attacks makes their detection quite challenging. In this context, generative adversarial networks (GANs) are a promising unsupervised approach to detect cyber-attacks by implicitly modeling the system. However, the detection of cyber-attacks on CPSs has strict latency requirements, since the attacks need to be stopped before the system is compromised. In this article, we propose FID-GAN, a novel fog-based, unsupervised intrusion detection system (IDS) for CPSs using GANs. The IDS is proposed for a fog architecture, which brings computation resources closer to the end nodes and thus contributes to meeting low-latency requirements. In order to achieve higher detection rates, the proposed architecture computes a reconstruction loss based on the reconstruction of data samples mapped to the latent space. Other works that follow a similar approach struggle with the time required to compute the reconstruction loss, which renders them impractical for latency constrained applications. We address this problem by training an encoder that accelerates the reconstruction loss computation. Experiments show that the proposed solution achieves higher detection rates and is at least 5.5 times faster than a baseline approach in the three studied data sets.
Paulo Freitas de Araujo-Filho, Georges Kaddoum, Divanilson Campelo, Aline Gondim Santos, David Macedo, Cleber Zanchettin
IEEE Internet Things J.6
2020 On Analysing Similarity Knowledge Transfer by Ensembles
Danilo Pereira, Flávio Arthur O. Santos, Leonardo N. Matos, Paulo Novais, Cleber Zanchettin, Teresa Bernarda Ludermir
IDEAL (2)5
2020 KutralNet: A Portable Deep Learning Model for Fire Recognition
abstract
Most of the automatic fire alarm systems detect the fire presence through sensors like thermal, smoke, or flame. One of the new approaches to the problem is the use of images to perform the detection. The image approach is promising since it does not need specific sensors and can be easily embedded in different devices. However, besides the high performance, the computational cost of the used deep learning methods is a challenge to their deployment in portable devices. In this work, we propose a new deep learning architecture that requires fewer floating-point operations (flops) for fire recognition. Additionally, we propose a portable approach for fire recognition and the use of modern techniques such as inverted residual block, convolutions like depth-wise, and octave, to reduce the model's computational cost. The experiments show that our model keeps high accuracy while substantially reducing the number of parameters and flops. One of our models presents 71% fewer parameters than FireNet, while still presenting competitive accuracy and AUROC performance. The proposed methods are evaluated on FireNet and FiSmo datasets. The obtained results are promising for the implementation of the model in a mobile device, considering the reduced number of flops and parameters acquired.
Angel Ayala, Bruno J. T. Fernandes, Francisco Cruz 0002, David Macedo, Adriano Lorena Inácio de Oliveira, Cleber Zanchettin
IJCNN6
2020 Squeezed Deep 6DoF Object Detection using Knowledge Distillation
abstract
The detection of objects considering a 6DoF pose is a common requirement to build virtual and augmented reality applications. It is usually a complex task which requires real-time processing and high precision results for adequate user experience. Recently, different deep learning techniques have been proposed to detect objects in 6DoF in RGB images. However, they rely on high complexity networks, requiring a computational power that prevents them from working on mobile devices. In this paper, we propose an approach to reduce the complexity of 6DoF detection networks while maintaining accuracy. We used Knowledge Distillation to teach portables Convolutional Neural Networks (CNN) to learn from a real-time 6DoF detection CNN. The proposed method allows real-time applications using only RGB images while decreasing the hardware requirements. We used the LINEMOD dataset to evaluate the proposed method, and the experimental results show that the proposed method reduces the memory requirement by almost 99% in comparison to the original architecture with the cost of reducing half the accuracy in one of the metrics. Code is available at https://github.com/heitorcfelix/singleshot6Dpose.
Heitor Felix, Walber M. Rodrigues, David Macedo, Francisco Simões, Adriano Lorena Inácio de Oliveira, Veronica Teichrieb, Cleber Zanchettin
IJCNN7
2020 A Fast Fully Octave Convolutional Neural Network for Document Image Segmentation
abstract
The Know Your Customer (KYC) and Anti Money Laundering (AML) are worldwide practices to online customer identification based on personal identification documents, similarity and liveness checking, and proof of address. To answer the basic regulation question: are you whom you say you are? The customer needs to upload valid identification documents (ID). This task imposes some computational challenges since these documents are diverse, may present different and complex backgrounds, some occlusion, partial rotation, poor quality, or damage. Advanced text and document segmentation algorithms were used to process the ID images. In this context, we investigated a method based on U-Net to detect the document edges and text regions in ID images. Besides the promising results on image segmentation, the U-Net based approach is computationally expensive for a real application, since the image segmentation is a customer device task. We propose a model optimization based on Octave Convolutions to qualify the method to situations where storage, processing, and time resources are limited, such as in mobile and robotic applications. We conducted the evaluation experiments in two new datasets CDPhotoDataset and DTDDataset, which are composed of real ID images of Brazilian documents. Our results showed that the proposed models are efficient to document segmentation tasks and portable.
Ricardo Batista das Neves Junior, Luiz Felipe Verçosa, David Macedo, Byron L. D. Bezerra, Cleber Zanchettin
IJCNN5
2020 Distantly-Supervised Neural Relation Extraction with Side Information using BERT
abstract
Relation extraction (RE) consists in categorizing the relationship between entities in a sentence. A recent paradigm to develop relation extractors is Distant Supervision (DS), which allows the automatic creation of new datasets by taking an alignment between a text corpus and a Knowledge Base (KB). KBs can sometimes also provide additional information to the RE task. One of the methods that adopt this strategy is the RESIDE model, which proposes a distantly-supervised neural relation extraction using side information from KBs. Considering that this method outperformed state-of-the-art baselines, in this paper, we propose a related approach to RESIDE also using additional side information, but simplifying the sentence encoding with BERT embeddings. Through experiments, we show the effectiveness of the proposed method in Google Distant Supervision and Riedel datasets concerning the BGWA and RESIDE baseline methods. Although Area Under the Curve is decreased because of unbalanced datasets, P@N results have shown that the use of BERT as sentence encoding allows superior performance to baseline methods.
Johny Moreira, Chaina Santos Oliveira, David Macedo, Cleber Zanchettin, Luciano Barbosa
IJCNN4
2020 AM-MobileNet1D: A Portable Model for Speaker Recognition
abstract
Speaker Recognition and Speaker Identification are challenging tasks with essential applications such as automation, authentication, and security. Deep learning approaches like SincNet and AM-SincNet presented great results on these tasks. The promising performance took these models to real-world applications that becoming fundamentally end-user driven and mostly mobile. The mobile computation requires applications with reduced storage size, non-processing and memory intensive and efficient energy-consuming. The deep learning approaches, in contrast, usually are energy expensive, demanding storage, processing power, and memory. To address this demand, we propose a portable model called Additive Margin MobileNet1D (AM-MobileNet1D) to Speaker Identification on mobile devices. We evaluated the proposed approach on TIMIT and MIT datasets obtaining equivalent or better performances concerning the baseline methods. Additionally, the proposed model takes only 11.6 megabytes on disk storage against 91.2 from SincNet and AM-SincNet architectures, making the model seven times faster, with eight times fewer parameters.
João Antônio Chagas Nunes, David Macedo, Cleber Zanchettin
IJCNN3
2020 Energy Consumption Optimization for CSMA/CA Protocol Employing Machine Learning
abstract
The algorithms commonly used for energy control in systems with Carrier Sense Multiple Access with Collision Avoidance (CSMA/CA) protocol involve optimization functions with considerable computational complexity and need rigorous control of the test environment. These restrictions create a gap among design, theoretical analysis, real-time processing of network devices and the dependence on human support for parameters setting. In this paper, we propose a novel approach to reach energy saving based on machine learning which considers the input and output of a power consumption control algorithm in CSMA networks, taking into account multiple physical (PHY) layer variables. The results show that the proposed approach obtained a better performance regarding processing time, computational cost and self-adaptation of the parameters currently defined by greedy search energy control algorithms.
Paulo Filipe Candido Barbosa, Bruna Alves da Silva, Cleber Zanchettin, Renato M. de Moraes
VTC Spring3
2020 HU-PageScan: a fully convolutional neural network for document page crop
abstract
November The offer of online, automated, and impersonal services demand users to upload scanned copies of their documents to the organisations. As a consequence of this decentralisation, the documents present more challenges to the already complex process of image processing and information extraction. To address this problem, the authors presented an optimised fully convolutional neural network model for document segmentation that works on mobile devices to detect the region of the document in the captured image. They performed experiments in three representative datasets comparing the proposed method with the Geodesic object Proposals, U‐net, Mask R‐CNN, and OctHU‐PageScan algorithms. They also compared the proposed model with all competitors of the ICDAR2015 Competition on smartphone document capture. Furthermore, they performed a qualitative and comparative analysis with the CamScanner software, a popular app for Android and iOS smartphones used for more than 100 million users in over 200 countries. The proposed approach achieved a significant performance compared with the current state‐of‐the‐art methods, providing a powerful approach for document segmentation in photos and scanned images.
Ricardo Batista das Neves Junior, Estanislau Lima, Byron L. D. Bezerra, Cleber Zanchettin, Alejandro H. Toselli
IET Image Process.4
2019 Squeezed Very Deep Convolutional Neural Networks for Text Classification
Andréa B. Duque, Luã Lázaro J. Santos, David Macedo, Cleber Zanchettin
ICANN (1)4
2019 Dynamic Centroid Insertion and Adjustment for Data Sets with Multiple Imbalanced Classes
Evandro J. R. Silva, Cleber Zanchettin
ICANN (2)2
2019 Improving Deep Image Clustering with Spatial Transformer Layers
Thiago V. M. Souza, Cleber Zanchettin
ICANN (4)2
2019 Speeding-up the Handwritten Signature Segmentation Process through an Optimized Fully Convolutional Neural Network
abstract
The handwritten signature is the most used method of identity authentication. Due to their nature, signatures can be used as an agreement in many types of documentation with legal repercussions. The validation of the firmed signature is used to prevent frauds, fake documents, and identity checking. However, working with automated signature verification is a challenging task because it can appear in any part of documents with complex backgrounds, with logos, handwritten texts, and many different patterns. Besides, the application needs to consider a real-time response. In this paper, we propose an optimized architecture of a fully convolutional neural network based on the U-Net architecture for handwritten signature segmentation. Furthermore, we used data augmentation in order to increase the diversity of the available dataset and prevent the overfitting problem when training the proposed model. We conducted experiments with DSSigDataset, and we used four different data augmentation techniques to increase the dataset size. The experimental results show that our proposed approach speed-up the handwritten signature segmentation task, at the same time, achieving higher accuracy and lower variance than previous works.
Paloma G. S. Silva, Celso A. M. Lopes Junior, Estanislau Lima, Byron L. D. Bezerra, Cleber Zanchettin
ICDAR5
2019 Towards Optimizing Convolutional Neural Networks for Robotic Surgery Skill Evaluation
abstract
In medicine courses, improve the skills of surgery students is an essential part of the program. For training the surgeon residents the institutions normally using a standard checklist to evaluate the student evolution. However, the checklist evaluation is susceptible to evaluator bias, inter-evaluator variability, besides being time-consuming. The automation of this process is an important evolution in medical training. An alternative to the instructor checklist is capturing and evaluation of kinematic data regarding the surgical motion. We propose a novel CNN architecture for automated robot-assisted skill assessment. We explore the use of the SELU activation function and a global mixed pooling approach based on the average and max-pooling layers. Finally, we examine two types of convolutional layers: real-value and quaternion-valued. The results suggest that our model presents a higher average accuracy across the three surgical subtasks of the JIGSAWS dataset.
Dayvid Castro, Danilo Pereira, Cleber Zanchettin, David Macedo, Byron L. D. Bezerra
IJCNN3
2019 Heartbeat Anomaly Detection using Adversarial Oversampling
abstract
Cardiovascular diseases are one of the most common causes of death in the world. Prevention, knowledge of previous cases in the family, and early detection is the best strategy to reduce this fact. Different machine learning approaches to automatic diagnostic are being proposed to this task. As in most health problems, the imbalance between examples and classes is predominant in this problem and affects the performance of the automated solution. In this paper, we address the classification of heartbeats images in different cardiovascular diseases. We propose a two-dimensional Convolutional Neural Network for classification after using a InfoGAN architecture for generating synthetic images to unbalanced classes. We call this proposal Adversarial Oversampling and compare it with the classical oversampling methods as SMOTE, ADASYN, and Random Oversampling. The results show that the proposed approach improves the classifier performance for the minority classes without harming the performance in the balanced classes.
Jefferson L. P. Lima, David Macedo, Cleber Zanchettin
IJCNN3
2019 Additive Margin SincNet for Speaker Recognition
abstract
Speaker Recognition is a challenging task with essential applications such as authentication, automation, and security. The SincNet is a new deep learning based model which has produced promising results to tackle the mentioned task. To train deep learning systems, the loss function is essential to the network performance. The Softmax loss function is a widely used function in deep learning methods, but it is not the best choice for all kind of problems. For distance-based problems, one new Softmax based loss function called Additive Margin Softmax (AM-Softmax) is proving to be a better choice than the traditional Softmax. The AM-Softmax introduces a margin of separation between the classes that forces the samples from the same class to be closer to each other and also maximizes the distance between classes. In this paper, we propose a new approach for speaker recognition systems called AM-SincNet, which is based on the SincNet but uses an improved AM-Softmax layer. The proposed method is evaluated in the TIMIT dataset and obtained an improvement of approximately 40% in the Frame Error Rate when compared to SincNet.
João Antônio Chagas Nunes, David Macedo, Cleber Zanchettin
IJCNN3
2019 On the Influence of the Color Model for Image Boundary Detection Algorithms based on Convolutional Neural Networks
abstract
Image analysis and understanding are challenging tasks, usually having segmentation as a major step. Boundary detection is a type of segmentation which aims to highlight the boundaries of the objects in a scene. Models based on Convolutional Neural Networks (CNN) have presented promising results for boundary detection, where the input usually is the entire image or some patches, often described in the RGB color model. In this paper, we provide a qualitative analysis of boundary detection algorithms based on CNN but considering images in different color models. We have used the color models RGB, Lab, Luv, dRdGdB, YO1O2 and HSV for this analysis. The Holistically-Nested Edge Detection (HED) and Convolutional Encoder Decoder Network (CEDN) are the CNN's chosen due to their high performance. The benchmark BSDS is the boundary detection evaluator. Experiments show that the results of the edge detection process tend to be similar when training the CNN with weights randomly initialized, regardless of the color model used. For the HED architecture, the use of Lab and Luv color models has resulted in a significant improvement to the case of transfer learning and fine-tuning of weights.
T. J. dos Santos, Carlos A. B. Mello, Cleber Zanchettin, Thiago V. M. Souza
IJCNN3
2019 Improving Universal Language Model Fine-Tuning using Attention Mechanism
abstract
Inductive transfer learning is widespread in computer vision applications. However, in natural language processing (NLP) applications is still an under-explored area. The most common transfer learning method in NLP is the use of pre-trained word embeddings. The Universal Language Model Fine-Tuning (ULMFiT) is a recent approach which proposes to train a language model and transfer its knowledge to a final classifier. During the classification step, ULMFiT uses a max and average pooling layer to select the useful information of an embedding sequence. We propose to replace max and average pooling layers with a soft attention mechanism. The goal is to learn the most important information of the embedding sequence rather than assuming that they are max and average values. We evaluate the proposed approach in six datasets and achieve the best performance in all of them against literature approaches.
Flávio Arthur O. Santos, Karina L. Ponce-Guevara, David Macedo, Cleber Zanchettin
IJCNN4
2019 Enhancing batch normalized convolutional networks using displaced rectifier linear units: A systematic comparative study
David Macedo, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira, Teresa Bernarda Ludermir
Expert Syst. Appl.2
2018 SegNetRes-CRF: A Deep Convolutional Encoder-Decoder Architecture for Semantic Image Segmentation
abstract
Semantic segmentation is an essential task in computer vision that aims to label each image pixel. Several of the actual best approaches in this context are based on deep neural networks. For example, SegNet is a deep encoder-decoder architecture approach whose results were disruptive because it is fast and performs well. However, this architecture fails to fine-delineating the edges between the objects of interest in the image. We propose some modifications in the SegNet-Basic architecture by using a post-processing segmentation layer (using Conditional Random Fields) and by transferring high resolution features combined to the decoder network. The proposed method was evaluated in the dataset CamVid. Moreover, it was compared with important variants of SegNet and showed to be able to improve the overall accuracy of SegNet-Basic by up to 17.5%.
Luiz A. Oliveira, Heitor R. Medeiros, David Macedo, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira, Teresa Bernarda Ludermir
IJCNN4
2018 Reducing SqueezeNet Storage Size with Depthwise Separable Convolutions
abstract
Current research in the field of convolutional neural networks usually focuses on improving network accuracy, regardless of the network size and inference time. In this paper, we investigate the effects of storage space reduction in SqueezeNet as it relates to inference time when processing single test samples. In order to reduce the storage space, we suggest adjusting SqueezeNet's Fire Modules to include Depthwise Separable Convolutions (DSC). The resulting network, referred to as SqueezeNet-DSC, is compared to different convolutional neural networks such as MobileNet, AlexNet, VGG19, and the original SqueezeNet itself. When analyzing the models, we consider accuracy, the number of parameters, parameter storage size and processing time of a single test sample on CIFAR-10 and CIFAR-100 databases. The SqueezeNet-DSC exhibited a considerable size reduction (37% the size of SqueezeNet), while experiencing a loss in network accuracy of 1,07% in CIFAR-10 and 3,06% in top 1 CIFAR-100.
Aline Gondim Santos, Camila Oliveira de Souza, Cleber Zanchettin, David Macedo, Adriano Lorena Inácio de Oliveira, Teresa Bernarda Ludermir
IJCNN3
2018 Gesture recognition: A review focusing on sign language in a mobile context
Davi Hirafuji Neiva, Cleber Zanchettin
Expert Syst. Appl.2
2017 The Impact of Dataset Complexity on Transfer Learning over Convolutional Neural Networks
Miguel D. de S. Wanderley, Leonardo de A. e Bueno, Cleber Zanchettin, Adriano Lorena Inácio de Oliveira
ICANN (2)3
2017 QRNN: q -Generalized Random Neural Network
abstract
Artificial neural networks (ANNs) are widely used in applications with complex decision boundaries. A large number of activation functions have been proposed in the literature to achieve better representations of the observed data. However, only a few works employ Tsallis statistics, which has successfully been applied to various other fields. This paper presents a random neural network (RNN) with q -Gaussian activation functions [ q -generalized RNN (QRNN)] based on Tsallis statistics. The proposed method employs an additional parameter q (called the entropic index) which reflects the degree of nonextensivity. This approach has the flexibility to model complex decision boundaries of different shapes by varying the entropic index. We conduct numerical experiments to analyze the efficiency of QRNN compared with RNNs and several other classical methods. Statistical tests (Wilcoxon and Friedman) are used to validate our results and show that the QRNN performs significantly better than RNNs with different activation functions. In addition, we find that QRNN outperforms many of the compared classical methods, with the exception of support vector machines, in which case it still exhibits a substantial advantage in terms of implementation simplicity and speed.
Dusan Stosic, Darko Stosic, Cleber Zanchettin, Teresa Bernarda Ludermir, Borko D. Stosic
IEEE Trans. Neural Networks Learn. Syst.3
2016 Objective Video Quality Assessment Based on Neural Networks
abstract
Image/Video Quality Assessment (IQA/VQA) plays a significant role in image and video processing, as it can directly predict the impact of distortions on the video in the quality of experience (QoE) of the user. For this propose, in this paper, it is presented a new method for objective video quality assessment using an artificial neural network to predict the subjective evaluation of the video as if it were observed by a human user. The network was trained using degradation indicators extracted from the VQEG Phase I video database, which describe the level of distortion suffered by the original video under spatial and temporal scopes. The proposed method obtained an excellent correlation with the subjective scores over this same database.
Diego P. A. Menor, Carlos A. B. Mello, Cleber Zanchettin
KES3
2016 An efficient static gesture recognizer embedded system based on ELM pattern recognition algorithm
Lucas F. S. Cambuim, Rafael M. Macieira, Fernando M. de Paula Neto, Edna Barros, Teresa Bernarda Ludermir, Cleber Zanchettin
J. Syst. Archit.6
2015 Extreme Learning Machine for Real Time Recognition of Brazilian Sign Language
abstract
The quantity of computing application that interacts with users through gesture or body motion has been growing. Among these applications is the sign language recognizer used to help hearing impaired people. This work proposes an architecture able to recognize Brazilian sign language (LIBRAS) in an embedded platform. The system focuses on a simple feature from 'finger spelling expressions' represented by a series of hands gestural images, and uses the Extreme Learning Machine network to classify them. The proposed structure uses camera images only and does not need any gloves or sensors. The obtained results are 5 times faster and 16 times better than classical approaches.
Fernando M. de Paula Neto, Lucas F. S. Cambuim, Rafael M. Macieira, Teresa Bernarda Ludermir, Cleber Zanchettin, Edna Barros
SMC5
2015 On the Existence of a Threshold in Class Imbalance Problems
abstract
One common approach to class imbalance problem is the resampling of data. However this strategy has some drawbacks, e.g., Unnecessary noise or the possibility of throwing out useful information. These inconveniences may be avoided or minimized by using a class proportion threshold allowing to identify when the imbalance data represent a problem to the classifier performance. In this paper we investigate the existence of this threshold and evaluate the performance of different classifiers in imbalanced problems. Results showed that for classifiers sensible to imbalanced data this threshold exists.
Evandro J. R. Silva, Cleber Zanchettin
SMC2
2014 Handwriting recognition system for mobile accessibility to the visually impaired people
abstract
This paper proposes the combination of preprocessing and handwriting recognition approaches aiming to develop a flexible and assistive mobile tool to help visually impaired people to understand and interact with handwriting text. The proposed system is described and its performance is evaluated in the IAM Handwriting Database. The system presented promising results and may aid visually impaired to overcome an everyday accessibility barrier as handwriting recognition.
Felipe Mendonca Gouveia, Byron L. D. Bezerra, Cleber Zanchettin, Joao Raul Jardim Meneses
SMC3
2014 Advances in intelligent systems
Teresa Bernarda Ludermir, Cleber Zanchettin, Ana Carolina Lorena
Neurocomputing2
2013 An adaptive thresholding algorithm based on edge detection and morphological operations for document images
abstract
This paper presents a new algorithm to threshold document images. The proposed algorithm deal with complex background images, illumination and aspect variants, back-to-front interference, variation of brightness and different positioned shadows. The algorithm have two phases. The first one uses edge detection and morphological operations to identify the text on the image. The second phase uses the positions of the text to define the threshold value in an adaptive process. Our approach presents promising results in images with complex background released from the Document Image Binarization Contest (DIBCO) when compared with other literature and competition thresholding algorithms.
Renata Freire de Paiva Neves, Cleber Zanchettin, Carlos A. B. Mello
ACM Symposium on Document Engineering2
2013 Metaclasses and zoning for handwritten document recognition
abstract
This work presents a complete method for improving the handwritten document recognition. In this task some characters are confused with others because of their visual/structural similarity. A SOM and TreeSOM neural network were used to sort different characters in metaclasses. In each metaclass a zoning approach was applied trying to get particular features to improve the character classification. The experiments with this new approach were performed in the NIST database with the classic MLP and a fast neural network RBF-DDA.
V. Macário, G. F. P. Silva, Milena R. P. Souza, Cleber Zanchettin, George D. C. Cavalcanti
IJCNN4
2012 A MDRNN-SVM Hybrid Model for Cursive Offline Handwriting Recognition
Byron L. D. Bezerra, Cleber Zanchettin, Vinícius Braga de Andrade
ICANN (2)2
2012 An Efficient Way of Combining SVMs for Handwritten Digit Recognition
Renata F. P. Neves, Cleber Zanchettin, Alberto N. G. Lopes Filho
ICANN (2)2
2012 Offline handwritten signature verification through network radial basis functions optimized by Differential Evolution
abstract
The handwritten signature is present in all important documents. In law, if the signature on a document is false, this document is also considered a fraud. This paper uses a neural network of radial basis function optimized by Differential Evolution Algorithm with features that best discriminates between a genuine signature of a simulated forgery. The experiments with this promising technique were made with a GPDS-300gray images base and the results subjected to statistical tests with the performance of technical literature.
Saulo Henrique Leoncio de Medeiros Napoles, Cleber Zanchettin
IJCNN2
2012 Odor recognition systems for natural gas odorization monitoring
abstract
This paper presents a system consisting of physical sensors and intelligent software for the automatic identification of the concentration of natural gas odorants and details the development of the sensor and pattern recognition systems. The sensor system uses spectroscopic technology and the pattern recognition system uses wavelet and artificial neural network technology. The aim is to determine the concentration of a natural gas odorant in the environment and associate this concentration with the benchmark index, which measures the degree of human perception to the presence of gas in the environment. Experiments were conducted comparing the performance of the system with human performance, which is normally used to deal with this problem. The proposed system demonstrated promising results.
Cleber Zanchettin, Leandro M. Almeida, Frederico D. Menezes, Teresa Bernarda Ludermir, Walter M. Azevedo
IJCNN1
2012 A KNN-SVM hybrid model for cursive handwriting recognition
abstract
This paper presents a hybrid KNN-SVM method for cursive character recognition. Specialized Support Vector Machines (SVMs) are introduced to significantly improve the performance of KNN in handwrite recognition. This hybrid approach is based on the observation that when using KNN in the task of handwritten characters recognition, the correct class is almost always one of the two nearest neighbors of the KNN. Specialized local SVMs are introduced to detect the correct class among these two different classification hypotheses. The hybrid KNN-SVM recognizer showed significant improvement in terms of recognition rate compared with MLP, KNN and a hybrid MLP-SVM approach for a task of character recognition.
Cleber Zanchettin, Byron L. D. Bezerra, Washington Wagner Azevedo da Silva
IJCNN1
2011 Evolving Clonal Adaptive Resonance Theory based on ECOS theory
abstract
The present work describes an evolution of the hybrid immune approach called Clonart (Clonal Adaptive Resonance Theory) using ECOS (Evolving Connectionist Systems) architectures. Some improvements were developed to allow the control of the growth of the clusters. Clonart's architecture is an Evolutionary Algorithm biologically inspired on the use of the Clonal Selection Principle. Therefore, a technique inspired on ART 1 network was combined to store the best antibodies. However, these strategies may create a lot of clusters due to the ART behavior. For that reason, techniques of insertion, aggregation and pruning inspired on ECOS operation were used to control the amount of clusters in Clonart. In this way, old and unnecessary clusters may confuse the Clonart and increase the learning error rate. This behavior was especially important, because many problems need constant retraining. The effectiveness of this approach was evaluated using ten databases from UCI Machine Learning Repository.
Jose Lima Alexandrino, Cleber Zanchettin, Edson Costa de Barros Carvalho Filho
IJCNN2
2011 A MLP-SVM hybrid model for cursive handwriting recognition
abstract
This paper presents a hybrid MLP-SVM method for cursive characters recognition. Specialized Support Vector Machines (SVMs) are introduced to significantly improve the performance of Multilayer Perceptron (MLP) in the local areas around the surfaces of separation between each pair of characters in the space of input patterns. This hybrid architecture is based on the observation that when using MLPs in the task of handwritten characters recognition, the correct class is almost always one of the two maximum outputs of the MLP. The second observation is that most of the errors consist of pairs of classes in which the characters have similarities (e.g. (U, V), (m, n), (O, Q), among others). Specialized local SVMs are introduced to detect the correct class among these two classification hypotheses. The hybrid MLP-SVM recognizer showed improvement, significant, in performance in terms of recognition rate compared with an MLP for a task of character recognition.
Washington Wagner Azevedo da Silva, Cleber Zanchettin
IJCNN2
2011 A SVM based off-line handwritten digit recognizer
abstract
This paper presents an efficient method for handwritten digit recognition. The proposed method makes use of Support Vector Machines (SVM), benefitting from its generalization power. The method presents improved recognition rates when compared to Multi-Layer Perceptron (MLP) classifiers, other SVM classifiers and hybrid classifiers. Experiments and comparisons were done using a digit set extracted from the NIST SD19 digit database. The proposed SVM method achieved higher recognition rates and it outperformed other methods. It is also shown that although using solely SVMs for the task, the new method does not suffer when considering processing time.
Renata F. P. Neves, Alberto N. G. Lopes Filho, Carlos A. B. Mello, Cleber Zanchettin
SMC4
2011 A Multi-Layer Perceptron approach to threshold documents with complex background
abstract
This paper describes a thresholding method based on a Multi-Layer Perceptron approach for documents with complex backgrounds. The study case is focused on two regions of Brazilian bank checks: the courtesy amount and the character magnetic code. Those images have complex backgrounds with different patterns, which is a problem for an automatic recognition system. The new approach is based on a connectionistic approach to find the best threshold value. The proposed method is compared to ten thresholding algorithms (classical and specific for bank checks) in three different real bank checks databases, according to different evaluation metrics (recognition rate, peak signal-to-noise ratio, mean square error, precision, recall, accuracy, specificity, negative rate metric, misclassification penalty metric and f-measure). Based on the results, we may conclude the proposed method is more robust to variations in the image acquiring process, which influences the contrast, bright, hue and amount of noise verified in the image.
Juliano Rabelo 0001, Cleber Zanchettin, Carlos A. B. Mello, Byron L. D. Bezerra
SMC2
2011 Hybrid Training Method for MLP: Optimization of Architecture and Training
abstract
The performance of an artificial neural network (ANN) depends upon the selection of proper connection weights, network architecture, and cost function during network training. This paper presents a hybrid approach (GaTSa) to optimize the performance of the ANN in terms of architecture and weights. GaTSa is an extension of a previous method (TSa) proposed by the authors. GaTSa is based on the integration of the heuristic simulated annealing (SA), tabu search (TS), genetic algorithms (GA), and backpropagation, whereas TSa does not use GA. The main advantages of GaTSa are the following: a constructive process to add new nodes in the architecture based on GA, the ability to escape from local minima with uphill moves (SA feature), and faster convergence by the evaluation of a set of solutions (TS feature). The performance of GaTSa is investigated through an empirical evaluation of 11 public-domain data sets using different cost functions in the simultaneous optimization of the multilayer perceptron ANN architecture and weights. Experiments demonstrated that GaTSa can also be used for relevant feature selection. GaTSa presented statistically relevant results in comparison with other global and local optimization techniques.
Cleber Zanchettin, Teresa Bernarda Ludermir, Leandro M. Almeida
IEEE Trans. Syst. Man Cybern. Part B1
2010 Design of Experiments in Neuro-Fuzzy Systems
abstract
Interest in hybrid methods that combine artificial neural networks and fuzzy inference systems has grown in recent years. These systems are robust solutions that search for representations of domain knowledge, reasoning on uncertainty, automatic learning and adaptation. However, the design and definition of the parameter effectiveness of such systems is still a hard task. In the present work, we perform a statistical analysis to verify interactions and interrelations between parameters in the design of neuro-fuzzy systems. The analysis is carried out using a powerful statistical tool, namely, Design of Experiments (DOE), in two neuro-fuzzy models — Adaptive Neuro Fuzzy Inference System (ANFIS) and Evolving Fuzzy Neural Networks (EFuNN). The results show that, for ANFIS, input MFs number and output MFs shape are usually the factors with the largest influence on the system's RMSE. For EFFuNN, the MF shape and the interaction between MF shape and number usually have the largest effect size.
Cleber Zanchettin, Leandro L. Minku, Teresa Bernarda Ludermir
Int. J. Comput. Intell. Appl.1
2009 Hybrid intelligent immune system using Radial Basis Function applied to Time Series Analysis
abstract
The present work proposes an integration of Clonal Adaptive Resonance Theory framework (Clonart) with Radial Basis Function (RBF) called ClonalRBF. This framework was already used in a handwritten digit classification problem, a forecasting for the Brazilian Energy Distribution System and now a Time Series Analysis in Gas Furnace and Mackey-Glass databases. In Clonart, the population memory was organized using an ART 1 network and in the new approach it was organized using a RBF network. This framework has biologically inspired characteristics like the grouping of similar antibodies and memory antibodies. It was studied to allow the evolution of the artificial immune system. The focus of this study was to evaluate the ClonalRBF and to compare with Clonart using these two databases.
Jose Lima Alexandrino, Cleber Zanchettin, Edson Costa de Barros Carvalho Filho
IJCNN2
2008 A hybrid intelligent system clonart for short and mid-term forecasting for the Brazilian Energy Distribution System
abstract
The present work describes an application of Clonart (Clonal Adaptive Resonance Theory) for forecasting of amount of precipitation for the Brazilian Energy Distribution System. The effectiveness of the Brazilian electricity system directly depends on the difference between hydroelectric energy production and consumer use. Production depends upon the volume of water stored in the reservoirs. A forecasting system for the amount of rainfall throughout the year contributes significantly to the analysis. The plasticity of the Clonart ensures that a new piece of knowledge does not overshadow previous knowledge. This is especially important for forecast problems because this type of problem needs constants training.
Jose Lima Alexandrino, Cleber Zanchettin, Edson Costa de Barros Carvalho Filho
IJCNN2
2008 Feature subset selection in a methodology for training and improving artificial neural network weights and connections
abstract
This paper investigates the problem of feature subset selection as part of a methodology that integrates heuristic tabu search, simulated annealing, genetic algorithms and backpropagation. This technique combines both global and local search strategies for the simultaneous optimization of the number of connections and connection values of multi-layer perceptron neural networks. We compare the performance of the proposed method for feature subset selection to five classical feature selection methods in three different classification problems.
Cleber Zanchettin, Teresa Bernarda Ludermir
IJCNN1
2007 Artificial Immune System with ART Memory Hibridization
abstract
The present work proposes the architecture Clonart (Clonal Adaptive Resonance Theory) that employs many different techniques like intelligent operators, clonal selection principle, local search, memory antibodies and ART clusterization in order to increase the performance of the algorithm. The approach uses a mechanism similar to the ART 1 network for storing a population of memory antibodies that will be responsible for the acquired knowledge of the algorithm. This characteristic allows the algorithm a self-organization of the antibodies in accordance with the complexity of the database used.
Jose Lima Alexandrino, Cleber Zanchettin, Edson Costa de Barros Carvalho Filho
HIS2
2007 An Efficient Thresholding Algorithm for Brazilian Bank Checks
abstract
It is present herein an algorithm for thresholding images of bank checks. These images have complex background elements. Some of these patterns make very hard to distinguish between the text and the texture pattern defined by the bank. For the binarizing process, an adaptive global thresholding algorithm is proposed based on ROC curves and it is compared to several well-known algorithms. The images generated by the new algorithm achieved a hit rate of 97% for recognition of the CMC7 code.
Carlos A. B. Mello, Byron L. D. Bezerra, Cleber Zanchettin, V. Macário
ICDAR3
2007 Comparison of the Effectiveness of Different Cost Functions in Global Optimization Techniques
abstract
It is present herein an evaluation of the effect of different cost functions on a methodology that integrates heuristic tabu search, simulated annealing, genetic algorithms and backpropagation. We investigated five cost function approaches: the average method, weighted average, weight-decay, multi- objective optimization, combined multi-objective and weight- decay. The weight-decay approach presented promising results in the optimization process. The experiments were performed in four classifications and one prediction problem.
Cleber Zanchettin, Teresa Bernarda Ludermir
IJCNN1
2006 A Heuristic Binarization Algorithm for Documents with Complex Background
abstract
This paper proposes a new method for binarization of digital documents. The proposed approach performs binarization by using a heuristic algorithm with two different thresholds and the combination of the thresholded images. The method is suitable for binarization of complex background document images. In experiments, it obtained better results than classical techniques in the binarization of real bank checks.
George D. C. Cavalcanti, Eduardo F. A. Silva, Cleber Zanchettin, Byron L. D. Bezerra, Rodrigo C. Doria, Juliano Rabelo 0001
ICIP3
2006 A neural architecture to identify courtesy amount delimiters
abstract
This paper deals with automatic recognition of real bank checks. A new approach is proposed to read the numerical amount field from bank checks, considering the numeric value and the different delimiters that might exist in that field. The proposal combines different neural networks classifiers to perform the recognition. Experimental results have shown that this approach is robust and efficient for automatic recognition of real Brazilian bank checks.
Cleber Zanchettin, George D. C. Cavalcanti, Rodrigo C. Doria, Eduardo F. A. Silva, Juliano Rabelo 0001, Byron L. D. Bezerra
IJCNN1
2006 A methodology to train and improve artificial neural networks' weights and connections
abstract
This work presents a new methodology that integrates the heuristics Tabu search, simulated annealing, genetic algorithms and backpropagation in a pruning and constructive way. The approach obtained promising results in the simultaneous optimization of artificial neural network architecture and weights. The experiments were performed in four classification and one prediction problem.
Cleber Zanchettin, Teresa Bernarda Ludermir
IJCNN1
2006 An Optimization Methodology for Neural Network Weights and Architectures
abstract
This paper introduces a methodology for neural network global optimization. The aim is the simultaneous optimization of multilayer perceptron (MLP) network weights and architectures, in order to generate topologies with few connections and high classification performance for any data sets. The approach combines the advantages of simulated annealing, tabu search and the backpropagation training algorithm in order to generate an automatic process for producing networks with high classification performance and low complexity. Experimental results obtained with four classification problems and one prediction problem has shown to be better than those obtained by the most commonly used optimization techniques.
Teresa Bernarda Ludermir, Akio Yamazaki, Cleber Zanchettin
IEEE Trans. Neural Networks3
2005 Design of Experiments in Neuro-Fuzzy Systems
abstract
Interest in hybrid methods that combine artificial neural networks and fuzzy inference systems has grown. These systems are robust solutions that search for representation of domain knowledge, reasoning on uncertainty, automatic learning and adaptation. However, the design and the definition of parameters effectiveness of these systems is a hard task yet. In this paper we perform a statistical analysis to verify the interactions and interrelations between parameters in the design of neuro-fuzzy systems. The analysis carries out using a powerful statistical tool, the design of experiments (DOE) in two neuro-fuzzy models, adaptive neuro fuzzy inference system (ANFIS) and evolving fuzzy neural networks (EFuNN).
Cleber Zanchettin, Fernanda L. Minku, Teresa Bernarda Ludermir
HIS1
2005 Hybrid Technique for Artificial Neural Network Architecture and Weight Optimization
Cleber Zanchettin, Teresa Bernarda Ludermir
PKDD1
2005 Hybrid neural systems for pattern recognition in artificial noses
abstract
This work examines the use of Hybrid Intelligent Systems in the pattern recognition system of an artificial nose. The connectionist approaches Multi-Layer Perceptron and Time Delay Neural Networks, and the hybrid approaches Feature-Weighted Detector and Evolving Neural Fuzzy Networks were investigated. A Wavelet Filter is evaluated as a preprocessing method for odor signals. The signals generated by an artificial nose were composed by an array of conducting polymer sensors and exposed to two different odor databases.
Cleber Zanchettin, Teresa Bernarda Ludermir
Int. J. Neural Syst.1
2004 Evolving Fuzzy Neural Networks Applied to Odor Recognition
Cleber Zanchettin, Teresa Bernarda Ludermir
ICONIP1
2004 Evolving fuzzy neural networks applied to odor recognition in an artificial nose
abstract
A pattern recognition system using evolving fuzzy neural networks for an artificial nose is presented. The artificial nose is composed of an adaptive and on-line learning method. For the classification of gases derived from the petroliferous industry, the method presented achieves better results (mean classification error of 0.88%) than those obtained by time delay neural networks (10.54%).
Cleber Zanchettin, Teresa Bernarda Ludermir
IJCNN1
2003 Wavelet Filter for Noise Reduction and Signal Compression in an Artificial Nose
Cleber Zanchettin, Teresa Bernarda Ludermir
HIS1
2003 A Neuro-Fuzzy Model Applied to Odor Recognition in an Artificial Nose
Cleber Zanchettin, Teresa Bernarda Ludermir
HIS1