Byron L. D. Bezerra

dblp:75/1673 · also Byron Leite Dantas Bezerra · DBLP profile ↗
← Back
34ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-8327-9734ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 14 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5Graphics, computer vision, multimedia, augmented reality and games · 4Human-computer interaction and ubiquitous computing · 4Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 A Proposal of Post-OCR Spelling Correction Using Monolingual Byte-level Language Models
abstract
This work presents a proposal for a spelling corrector using monolingual byte-level language models (Monobyte) for the post-OCR task in texts produced by Handwritten Text Recognition (HTR) systems. We evaluate three Monobyte models, based on Google's ByT5, trained separately on English, French, and Brazilian Portuguese. The experiments evaluated three datasets with 21st century manuscripts: IAM, RIMES, and BRESSAY. In the IAM, Monobyte achieves reductions of 2.24% in character error rate (CER) and 26.37% in word error rate (WER). In RIMES, reductions are 13.48% (CER) and 33.34% (WER), while in BRESSAY, Monobyte improves CER by 12.78% and WER by 40.62%. The BRESSAY results surpass results reported in previous works using a multilingual ByT5 model. Our findings demonstrate the effectiveness of byte-level tokenization in noisy text and underscore the potential of computationally efficient, monolingual models. Code is availabled at https://github.com/savi8sant8s/monobyte-spelling-corrector.
Sávio S. Araújo, Byron L. D. Bezerra, Arthur Flor de Sousa Neto
DocEng2
2025 The Di2Win Document Intelligence Platform
abstract
We present the Di2Win Document Intelligence Platform (DIP). This modular AI-driven pipeline transforms raw document images --- captured by scanners or mobile phones --- into structured data and business actions in a single pass. The system comprises five loosely-coupled micro-services: (1) image-quality verification using a contrast-invariant model that flags blur, skew, and illumination issues above 100 ms per page; (2) document classification via a Transformer-base model with layout embeddings, delivering top-k types with calibrated confidence; (3) information extraction through i) Dilbert, a multimodal Token-Layout-Language model fine-tuned on weakly-labeled forms or ii) Delfos, a Large Language Model Mixture of Experts fine-tuned with well-defined prompts; (4) DataDrift, a powerful rules engine to avoid inconsistent outputs concerning the business process; and (5) process automation orchestrated by a Business Process Model Notation (BPMN) plus a Robot Process Automation (RPA) engine that routes results to databases, APIs, or human-review queues. All AI components are orchestrated through a messaging service to control the information flow, and the application exposes REST/gRPC endpoints to communicate with outside consumers. This enables the hot-swapping of models without downstream code changes by plugging a new message consumer into the messaging system. This also provides horizontal scalability since to increase the application throughput, we only need to add new AI engine consumers to the messaging system. Deployed in banking, insurance, and healthcare, the Di2Win DIP has processed more than 30 million pages, reducing average handling time by 79% and re-keying errors by 86 %, speeding up the workflows up to ten times. Our DocEng demonstration allows attendees to upload documents, observe live quality and confidence dashboards, and edit extracted fields with immediate feedback to the active-learning loop.
Afonso Ferreira, Cleber Zanchettin, Romulo Andrade, Byron L. D. Bezerra
DocEng4
2024 DocLightDetect: A New Algorithm for Occlusion Classification in Identification Documents
Ricardo Batista das Neves Junior, Byron L. D. Bezerra, Cleber Zanchettin
DAS2
2024 How Does Changing the Optical Character Recognition System Impact the Layout-Aware Named Entity Recognition Models?
João Macedo, Byron L. D. Bezerra, Cleber Zanchettin
DAS2
2024 BRESSAY: A Brazilian Portuguese Dataset for Offline Handwritten Text Recognition
Arthur Flor de Sousa Neto, Byron L. D. Bezerra, Sávio S. Araújo, Wiliane M. A. S. Souza, Kléberson F. Alves, Macileide F. Oliveira, Samara V. S. Lins, Hugo J. F. Hazin, Pedro H. V. Rocha, Alejandro H. Toselli
ICDAR (2)2
2024 ICDAR 2024 Competition on Handwritten Text Recognition in Brazilian Essays - BRESSAY
Arthur Flor de Sousa Neto, Byron L. D. Bezerra, Sávio S. Araújo, Wiliane M. A. S. Souza, Kléberson F. Alves, Macileide F. Oliveira, Samara V. S. Lins, Hugo J. F. Hazin, Pedro H. V. Rocha, Alejandro H. Toselli
ICDAR (6)2
2022 A robust handwritten recognition system for learning on different data restriction scenarios
Arthur Flor de Sousa Neto, Byron L. D. Bezerra, Alejandro H. Toselli, Estanislau Lima
Pattern Recognit. Lett.2
2021 ICDAR 2021 Competition on Components Segmentation Task of Document Photos
Celso A. M. Lopes Junior, Ricardo Batista das Neves Junior, Byron L. D. Bezerra, Alejandro H. Toselli, Donato Impedovo
ICDAR (4)3
2021 A Handwritten Signature Segmentation Approach for Multi-resolution and Complex Documents Acquired by Multiple Sources
Celso A. M. Lopes Junior, Murilo C. Stodolni, Byron L. D. Bezerra, Donato Impedovo
ICDAR (3)3
2020 HTR-Flor++: A Handwritten Text Recognition System Based on a Pipeline of Optical and Language Models
abstract
Offline Handwritten Text Recognition (HTR) is a task that offers a challenge in computer vision, where images are the only source of information. In fact, several approaches to optical models have been developed, such as through of Hidden Markov Model (HMM) or recurrent Bidirectional/Multidimensional layers. The current state-of-the-art consists of combined deep learning techniques, the Convolutional Recurrent Neural Networks (CRNN), in which recurrent layers still suffer from vanishing gradient problem when processing very long texts. In a way, high-performance models generally have millions of trainable parameters and a high computational cost. However, recently a new optical model architecture, Gated-CNN, demonstrated improvements to complement CRNN modeling. Thus, in this work, we present a new small architecture for HTR (based on Gated-CNN) integrated with two steps of language model at the character and word levels, respectively. Therefore, we used 9 state-of-the-art approaches and validated the results using the IAM public dataset. Finally, the proposed model surpasses the results obtained by different approaches in the literature, reaching recognition rates of CER 2.7% and WER 5.6%, which means an improvement of 13% over the best results on IAM dataset.
Arthur Flor de Sousa Neto, Byron L. D. Bezerra, Alejandro H. Toselli, Estanislau Lima
DocEng2
2020 FCN+RL: A Fully Convolutional Network followed by Refinement Layers to Offline Handwritten Signature Segmentation
abstract
Although secular, handwritten signature is one of the most reliable biometric method used by most countries. In the last ten years, the application of technology for verification of handwritten signatures has evolved strongly, including forensic aspects. Some factors, such as the complexity of the background and the small size of the region of interest - signature pixels - increase the difficulty of the targeting task. Other factors that make it challenging are the various variations present in handwritten signatures such as location, type of ink, color and type of pen, and the type of stroke. In this work, we propose an approach to locate and extract the pixels of handwritten signatures on identification documents, without any prior information on the location of the signatures. The technique used is based on a fully convolutional encoder-decoder network combined with a block of refinement layers for the alpha channel of the predicted image. The experimental results demonstrate that the technique outputs a clean signature with higher fidelity in the lines than the traditional approaches and preservation of the pertinent characteristics to the signer's spelling. To evaluate the quality of our proposal, we use the following image similarity metrics: SSIM, SIFT, and Dice Coefficient. The qualitative and quantitative results show a significant improvement in comparison with the baseline system.
Celso A. M. Lopes Junior, Matheus Henrique M. da Silva, Byron L. D. Bezerra, Bruno J. T. Fernandes, Donato Impedovo
IJCNN3
2020 A Fast Fully Octave Convolutional Neural Network for Document Image Segmentation
abstract
The Know Your Customer (KYC) and Anti Money Laundering (AML) are worldwide practices to online customer identification based on personal identification documents, similarity and liveness checking, and proof of address. To answer the basic regulation question: are you whom you say you are? The customer needs to upload valid identification documents (ID). This task imposes some computational challenges since these documents are diverse, may present different and complex backgrounds, some occlusion, partial rotation, poor quality, or damage. Advanced text and document segmentation algorithms were used to process the ID images. In this context, we investigated a method based on U-Net to detect the document edges and text regions in ID images. Besides the promising results on image segmentation, the U-Net based approach is computationally expensive for a real application, since the image segmentation is a customer device task. We propose a model optimization based on Octave Convolutions to qualify the method to situations where storage, processing, and time resources are limited, such as in mobile and robotic applications. We conducted the evaluation experiments in two new datasets CDPhotoDataset and DTDDataset, which are composed of real ID images of Brazilian documents. Our results showed that the proposed models are efficient to document segmentation tasks and portable.
Ricardo Batista das Neves Junior, Luiz Felipe Verçosa, David Macedo, Byron L. D. Bezerra, Cleber Zanchettin
IJCNN4
2020 HU-PageScan: a fully convolutional neural network for document page crop
abstract
November The offer of online, automated, and impersonal services demand users to upload scanned copies of their documents to the organisations. As a consequence of this decentralisation, the documents present more challenges to the already complex process of image processing and information extraction. To address this problem, the authors presented an optimised fully convolutional neural network model for document segmentation that works on mobile devices to detect the region of the document in the captured image. They performed experiments in three representative datasets comparing the proposed method with the Geodesic object Proposals, U‐net, Mask R‐CNN, and OctHU‐PageScan algorithms. They also compared the proposed model with all competitors of the ICDAR2015 Competition on smartphone document capture. Furthermore, they performed a qualitative and comparative analysis with the CamScanner software, a popular app for Android and iOS smartphones used for more than 100 million users in over 200 countries. The proposed approach achieved a significant performance compared with the current state‐of‐the‐art methods, providing a powerful approach for document segmentation in photos and scanned images.
Ricardo Batista das Neves Junior, Estanislau Lima, Byron L. D. Bezerra, Cleber Zanchettin, Alejandro H. Toselli
IET Image Process.3
2019 Speeding-up the Handwritten Signature Segmentation Process through an Optimized Fully Convolutional Neural Network
abstract
The handwritten signature is the most used method of identity authentication. Due to their nature, signatures can be used as an agreement in many types of documentation with legal repercussions. The validation of the firmed signature is used to prevent frauds, fake documents, and identity checking. However, working with automated signature verification is a challenging task because it can appear in any part of documents with complex backgrounds, with logos, handwritten texts, and many different patterns. Besides, the application needs to consider a real-time response. In this paper, we propose an optimized architecture of a fully convolutional neural network based on the U-Net architecture for handwritten signature segmentation. Furthermore, we used data augmentation in order to increase the diversity of the available dataset and prevent the overfitting problem when training the proposed model. We conducted experiments with DSSigDataset, and we used four different data augmentation techniques to increase the dataset size. The experimental results show that our proposed approach speed-up the handwritten signature segmentation task, at the same time, achieving higher accuracy and lower variance than previous works.
Paloma G. S. Silva, Celso A. M. Lopes Junior, Estanislau Lima, Byron L. D. Bezerra, Cleber Zanchettin
ICDAR4
2019 Towards Optimizing Convolutional Neural Networks for Robotic Surgery Skill Evaluation
abstract
In medicine courses, improve the skills of surgery students is an essential part of the program. For training the surgeon residents the institutions normally using a standard checklist to evaluate the student evolution. However, the checklist evaluation is susceptible to evaluator bias, inter-evaluator variability, besides being time-consuming. The automation of this process is an important evolution in medical training. An alternative to the instructor checklist is capturing and evaluation of kinematic data regarding the surgical motion. We propose a novel CNN architecture for automated robot-assisted skill assessment. We explore the use of the SELU activation function and a global mixed pooling approach based on the average and max-pooling layers. Finally, we examine two types of convolutional layers: real-value and quaternion-valued. The results suggest that our model presents a higher average accuracy across the three surgical subtasks of the JIGSAWS dataset.
Dayvid Castro, Danilo Pereira, Cleber Zanchettin, David Macedo, Byron L. D. Bezerra
IJCNN5
2018 Boosting the Deep Multidimensional Long-Short-Term Memory Network for Handwritten Recognition Systems
abstract
One of the main challenges in the handwriting recognition area lies in identifying complete lines of handwritten text. In this paper, we propose a handwriting recognition system based on a deep multidimensional long-short-term memory (MDLSTM) network within a hybrid hidden Markov model framework. The MDLSTM architecture was elaborated to enhance the recognition performance and decrease the recognition time. Accordingly, we present modifications regarding the layers order and the number of pooling layers compared to a standard MDLSTM model. Since the results reported in the literature for deeper MDLSTM architectures relies on optimizing the network width with a fixed depth, we investigate the trade-off between both these properties to obtain an optimal topology. The system was evaluated with English handwritten text lines from the IAM database and the experiments demonstrated that the proposed MDLSTM architecture was able to maintain a robust recognition performance (around 3.6% CER and 10.5% WER) and present significant speedups, approximately 48% and 32% faster than the state-of-the-art MDLSTM optical model, regarding the learning and classification times, respectively. The full system including a decoder with linguistic knowledge presents competitive results with the state-of-the-art.
Dayvid Castro, Byron L. D. Bezerra, Mêuser Valença
ICFHR2
2018 A Fully Convolutional Network for Signature Segmentation from Document Images
abstract
Handwritten signatures can be employed as a sign of confirmation in a wide variety of documents, namely, bank checks, identification documents and a variety of business certificates and contracts. Since those documents present complex backgrounds, the automatic extraction of handwritten signature from documents remains as an open task in the Offline Signature Verification field. In this paper we propose a method for the stroke-based extraction of signatures from document images. The approach is based on a Fully Convolutional Network trained to learn an end-to-end nonlinear mapping to extract the signatures from documents. Due to the lack of publicly available datasets containing the ground truth of signatures on the stroke level, we trained and evaluated our model on a dataset we created synthetically from real documents. It contains the stroke-based ground truth of signatures in a variety of documents with complex backgrounds. As a contribution of this work, the dataset will be made publicly available. Our method shows promising results on the test set, 89.8% recall and 66.9% precision.
Victor Kleber Santos Leite Melo, Byron L. D. Bezerra
ICFHR2
2017 A dynamic gesture recognition and prediction system using the convexity approach
Pablo V. A. Barros, Nestor T. Maciel-Junior, Bruno J. T. Fernandes, Byron L. D. Bezerra, Sérgio Murilo Maciel Fernandes
Comput. Vis. Image Underst.4
2015 Adoption of Software Product Line Development to an Environment of Voice User Interface
abstract
Software Product Line is a software development paradigm created to meet different market segments.This paradigm has shown great acceptance in the corporate environment (Motorola, Nokia, and Hewlett Packard) to allow the construction of more efficiently through reusing common components applications, besides being extensively researched by academics.The segment of voice interface, in turn, came up with the demand for systems capable of interacting with users, but in the application development process for this domain there is a lack of tools that make the task more productively.The FIVE (Framework for an Integrated Voice Environment) is a development environment for Voice Interface products designed to increase productivity in this segment.This paper aims to apply a SPL approach to FIVE.For this, a comparative evaluation of the process of construction of FIVE and SPL platforms was performed.Then adjustments in order to correct structural problems and, finally, the framework was validated using a set of experiments which sought to ensure the confirmation of such changes have been made.
Diógenes R. F. Oliveira, Byron L. D. Bezerra, Elyda L. S. X. Freitas, Alexandre M. A. Maciel
SEKE2
2014 A collaborative filtering framework based on local and global similarities with similarity tie-breaking criteria
abstract
Collaborative Filtering is the most commonly used technique in Recommender Systems, based on the users ratings in order to identify similar profiles and suggest them items. However, because it depends essentially on direct similarity measures between users or items, it usually suffers from the sparsity problem. Upon this situation, a good alternative is using global similarities to enrich the users neighborhood by transitively connecting them together, even when they do not share any common ratings. In this paper, we investigated the use of both local and global similarity measures with the maximin distance algorithm, along with tie-breaking criteria for neighbors with equal similarity. Our experiments showed that the maximin distance algorithm in fact produces many equally similar global neighbors, and that the criteria set for deciding between them severely improved the results of the recommendation process.
Andre R. S. Lopes, Ricardo B. C. Prudêncio, Byron L. D. Bezerra
IJCNN3
2014 Handwriting recognition system for mobile accessibility to the visually impaired people
abstract
This paper proposes the combination of preprocessing and handwriting recognition approaches aiming to develop a flexible and assistive mobile tool to help visually impaired people to understand and interact with handwriting text. The proposed system is described and its performance is evaluated in the IAM Handwriting Database. The system presented promising results and may aid visually impaired to overcome an everyday accessibility barrier as handwriting recognition.
Felipe Mendonca Gouveia, Byron L. D. Bezerra, Cleber Zanchettin, Joao Raul Jardim Meneses
SMC2
2013 An Effective Dynamic Gesture Recognition System Based on the Feature Vector Reduction for SURF and LCS
Pablo V. A. Barros, Nestor T. Maciel-Junior, Juvenal M. M. Bisneto, Bruno J. T. Fernandes, Byron L. D. Bezerra, Sérgio Murilo Maciel Fernandes
ICANN5
2012 A MDRNN-SVM Hybrid Model for Cursive Offline Handwriting Recognition
Byron L. D. Bezerra, Cleber Zanchettin, Vinícius Braga de Andrade
ICANN (2)1
2012 A KNN-SVM hybrid model for cursive handwriting recognition
abstract
This paper presents a hybrid KNN-SVM method for cursive character recognition. Specialized Support Vector Machines (SVMs) are introduced to significantly improve the performance of KNN in handwrite recognition. This hybrid approach is based on the observation that when using KNN in the task of handwritten characters recognition, the correct class is almost always one of the two nearest neighbors of the KNN. Specialized local SVMs are introduced to detect the correct class among these two different classification hypotheses. The hybrid KNN-SVM recognizer showed significant improvement in terms of recognition rate compared with MLP, KNN and a hybrid MLP-SVM approach for a task of character recognition.
Cleber Zanchettin, Byron L. D. Bezerra, Washington Wagner Azevedo da Silva
IJCNN2
2012 A recommender system architecture for an inter-application environment
abstract
Recommender systems help to show possible items of interest, but these recommendations are closely linked to the amount of available user information. This information is generated by user interaction with the system, which is generally absent, especially in the case of a new user. In this contribution, we intend to check the methods of filtering applied in the context of traditional recommendation systems by adapting them to an inter-application context by using collaborative filtering techniques. To this end, we propose the architecture and rules for building an environment of a recommendation system for inter-application. The resulting database was developed for evaluations of inter-analysis applications and it will be made available to the scientific community for future research.
Eduardo Jose Marcelino Vicente dos Santos, Byron L. D. Bezerra, Jefferson Silva de Amorim, Arthur Inacio do Nascimento
ISDA2
2011 A Multi-Layer Perceptron approach to threshold documents with complex background
abstract
This paper describes a thresholding method based on a Multi-Layer Perceptron approach for documents with complex backgrounds. The study case is focused on two regions of Brazilian bank checks: the courtesy amount and the character magnetic code. Those images have complex backgrounds with different patterns, which is a problem for an automatic recognition system. The new approach is based on a connectionistic approach to find the best threshold value. The proposed method is compared to ten thresholding algorithms (classical and specific for bank checks) in three different real bank checks databases, according to different evaluation metrics (recognition rate, peak signal-to-noise ratio, mean square error, precision, recall, accuracy, specificity, negative rate metric, misclassification penalty metric and f-measure). Based on the results, we may conclude the proposed method is more robust to variations in the image acquiring process, which influences the contrast, bright, hue and amount of noise verified in the image.
Juliano Rabelo 0001, Cleber Zanchettin, Carlos A. B. Mello, Byron L. D. Bezerra
SMC4
2011 Symbolic data analysis tools for recommendation systems
Byron L. D. Bezerra, Francisco de A. T. de Carvalho
Knowl. Inf. Syst.1
2009 A Filtering Algorithm for Highly Noisy Images of Brazilian ATM Bank Checks
abstract
This paper presents a new algorithm for filtering images of Brazilian bank checks. These images were generated by ATM machines and they present several kind of noise imposed by the digitization process. A new wavelet-based filtering algorithm is proposed for these images allowing a more efficient binarization process using percentage of black thresholding method.
Carlos A. B. Mello, Byron L. D. Bezerra, Afonso Ferreira, Juliano Rabelo 0001
SMC2
2008 A new algorithm to threshold the courtesy amount of Brazilian bank checks
abstract
This paper describes a new algorithm for thresholding the courtesy amount of Brazilian bank checks. These images have complex backgrounds which is a problem for an automatic recognition system. The new approach is based on Tsallis entropy to find the best threshold value. Histogram specification is also used for preprocessing some images. The bi-level images are analyzed through several quantitative measures and it achieved the best results when compared with the images produced by other classical thresholding algorithms.
Renata F. P. Neves, Carlos A. B. Mello, Mara S. Silva, Byron L. D. Bezerra
SMC4
2007 An Efficient Thresholding Algorithm for Brazilian Bank Checks
abstract
It is present herein an algorithm for thresholding images of bank checks. These images have complex background elements. Some of these patterns make very hard to distinguish between the text and the texture pattern defined by the bank. For the binarizing process, an adaptive global thresholding algorithm is proposed based on ROC curves and it is compared to several well-known algorithms. The images generated by the new algorithm achieved a hit rate of 97% for recognition of the CMC7 code.
Carlos A. B. Mello, Byron L. D. Bezerra, Cleber Zanchettin, V. Macário
ICDAR2
2006 A Heuristic Binarization Algorithm for Documents with Complex Background
abstract
This paper proposes a new method for binarization of digital documents. The proposed approach performs binarization by using a heuristic algorithm with two different thresholds and the combination of the thresholded images. The method is suitable for binarization of complex background document images. In experiments, it obtained better results than classical techniques in the binarization of real bank checks.
George D. C. Cavalcanti, Eduardo F. A. Silva, Cleber Zanchettin, Byron L. D. Bezerra, Rodrigo C. Doria, Juliano Rabelo 0001
ICIP4
2006 A neural architecture to identify courtesy amount delimiters
abstract
This paper deals with automatic recognition of real bank checks. A new approach is proposed to read the numerical amount field from bank checks, considering the numeric value and the different delimiters that might exist in that field. The proposal combines different neural networks classifiers to perform the recognition. Experimental results have shown that this approach is robust and efficient for automatic recognition of real Brazilian bank checks.
Cleber Zanchettin, George D. C. Cavalcanti, Rodrigo C. Doria, Eduardo F. A. Silva, Juliano Rabelo 0001, Byron L. D. Bezerra
IJCNN6
2006 C^2: : A Collaborative Recommendation System Based on Modal Symbolic User Profile
abstract
Recommendation Systems have become an important tool to cope with the information overload problem by acquiring information about the user behavior. However, the process of getting user personal data may vary in many different ways, and can be done implicitly (through actions) or explicitly (through rates). After tracing actions or getting rates of the user, Computational Recommendation Technologies use information filtering techniques to recommend items. In this paper we describe an approach to improve the recommendation quality in the first moments the user interacts with the system. The main idea is: (1) first of all, we describe the items with the general users opinion about them; and (2) after this, we use modal symbolic structures to save this content in the user profile. The proposed methodology outperforms, concerning the Find Good Items task measured by half-life utility metric, other approaches based on the following techniques: Cognitive Filtering, Social Filtering and hybrid methods.
Byron L. D. Bezerra, Francisco de A. T. de Carvalho, Valmir Macario
Web Intelligence1
2004 A symbolic approach for content-based information filtering
Byron L. D. Bezerra, Francisco de A. T. de Carvalho
Inf. Process. Lett.1