Gabriel Pereira e Silva

dblp:48/2434 · also Gabriel de F. P. e Silva, Gabriel de França Pereira e Silva · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
3since 2021 · last 2025
0009-0004-1214-0913ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 11 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
YearPublicationVenuePosition
2025 Binarizing Photographed Document Images 2025 Quality, Time and Space Assessment
abstract
Image binarization is fundamental for document image processing. The performance of binarization algorithms depends on several factors that range from the quality of the digitalization devices to the intrinsic features of the document itself and the kind and intensity of the noises present in the image. This assessment on binarizing photographed documents evaluated the quality, time, space, and performance of five new algorithms and ninety-eight "classical" algorithms. The test data set is composed of laser and deskjet printed documents, photographed using six widely used mobile devices with the strobe flash on, off, and in auto modes under three different angles and places of capture.
Gustavo P. Chaves, Thaylor Vieira, Gabriel Pereira e Silva, Rafael Dueire Lins, Steven J. Simske
DocEng3
2024 Competition on Binarizing Photographed Document Images 2024 Quality, Time and Space Report
abstract
Many document processing platforms have image binarization as a key step. The performance of binarization algorithms depends on several factors that span from the quality of the digitalization devices to the intrinsic features of the document itself and the kind and intensity of the noises present in the image. Besides that, the processing time is also an important factor that may restrict the applicability of a given algorithm. This competition on binarizing photographed documents assessed the quality, time, space, and performance of twenty three new algorithms and seventy-five "classical" algorithms. The test dataset is composed of offset and deskjet printed documents, photographed using six widely-used mobile devices with the strobe flash on, off, and in auto modes under two different angles and places of capture.
Rafael Dueire Lins, Gustavo P. Chaves, Gabriel Pereira e Silva, Thaylor Vieira, Ricardo da Silva Barboza, Steven J. Simske
DocEng3
2023 Quality, Space and Time Competition on Binarizing Photographed Document Images
abstract
Document image binarization is a fundamental step in many document processes. No binarization algorithm performs well on all types of document images, as the different kinds of digitalization devices and the physical noises present in the document and acquired in the digitalization process alter their performance. Besides that, the processing time is also an important factor that may restrict its applicability. This competition on binarizing photographed documents assessed the quality, time, space, and performance of five new algorithms and sixty-four "classical" and alternative algorithms. The evaluation dataset is composed of laser and deskjet printed documents, photographed using six widely-used mobile devices with the strobe flash on and off, under two different angles and places of capture.
Rafael Dueire Lins, Gabriel Pereira e Silva, Gustavo P. Chaves, Ricardo da Silva Barboza, Rodrigo Barros Bernardino, Steven J. Simske
DocEng2
2019 The CNN-Corpus: A Large Textual Corpus for Single-Document Extractive Summarization
abstract
This paper details the features and the methodology adopted in the construction of the CNN-corpus, a test corpus for single document extractive text summarization of news articles. The current version of the CNN-corpus encompasses 3,000 texts in English, and each of them has an abstractive and an extractive summary. The corpus allows quantitative and qualitative assessments of extractive summarization strategies.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng8
2019 The CNN-Corpus in Spanish: a Large Corpus for Extractive Text Summarization in the Spanish Language
abstract
This paper details the development and features of the CNN-corpus in Spanish, possibly the largest test corpus for single document extractive text summarization in the Spanish language. Its current version encompasses 1,117 well-written texts in Spanish, each of them has an abstractive and an extractive summary. The development methodology adopted allows good-quality qualitative and quantitative assessments of summarization strategies for tools developed in the Spanish language.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Diego A. Salcedo, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng9
2019 Generating Digital Libraries of M.Sc. and Ph.D. Theses
abstract
Postgraduate degrees are one of the most important propellers of all areas of science. M.Sc. and Ph.D. theses witness the important developments and provide a solid and global account of research projects. This paper describes a platform developed with the aim of generating digital libraries of theses and dissertations. Printed theses have to be scanned and then processed by the platform for marginal border removal, skew and orientation correction, image segmentation and enhancement, compression and pdf file generation. Both scanned and digitally generated theses are processed in the platform to extract relevant indexing information.
Rafael Dueire Lins, Paulo Hugo Espirito Santo, Gabriel Pereira e Silva
DocEng3
2015 A Quantitative and Qualitative Assessment of Automatic Text Summarization Systems
abstract
Text summarization is the process of automatically creating a shorter version of one or more text documents. This paper presents a qualitative and quantitative assessment of the 22 state-of-the-art extractive summarization systems using the CNN corpus, a dataset of 3,000 news articles.
Jamilson Batista, Rodolfo Ferreira, Hilário Tomaz, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Steven J. Simske, Gabriel Pereira e Silva, Marcelo Riss
DocEng7
2015 Automatic Text Document Summarization Based on Machine Learning
abstract
The need for automatic generation of summaries gained importance with the unprecedented volume of information available in the Internet. Automatic systems based on extractive summarization techniques select the most significant sentences of one or more texts to generate a summary. This article makes use of Machine Learning techniques to assess the quality of the twenty most referenced strategies used in extractive summarization, integrating them in a tool. Quantitative and qualitative aspects were considered in such assessment demonstrating the validity of the proposed scheme. The experiments were performed on the CNN-corpus, possibly the largest and most suitable test corpus today for benchmarking extractive summarization strategies.
Gabriel Pereira e Silva, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Luciano de Souza Cabral, Hilário Oliveira, Steven J. Simske, Marcelo Riss
DocEng1
2015 Binarizing complex scanned documents
abstract
Most binarization algorithms are suitable for scanned text documents and do not work adequately with complex documents that encompass photos, charts and text simultaneously. This paper presents a new binarization algorithm that works on complex documents. Each of the elements in the image is processed depending on its nature, thus their binarization takes that into account to preserve the original content.
Rafael Dueire Lins, Gabriel Pereira e Silva, Marcos Martins de Almeida
ICDAR2
2014 A Context Based Text Summarization System
abstract
Text summarization is the process of creating a shorter version of one or more text documents. Automatic text summarization has become an important way of finding relevant information in large text libraries or in the Internet. Extractive text summarization techniques select entire sentences from documents according to some criteria to form a summary. Sentence scoring is the technique most used for extractive text summarization, today. Depending on the context, however, some techniques may yield better results than some others. This paper advocates the thesis that the quality of the summary obtained with combinations of sentence scoring methods depend on text subject. Such hypothesis is evaluated using three different contexts: news, blogs and articles. The results obtained show the validity of the hypothesis formulated and point at which techniques are more effective in each of those contexts studied.
Rafael Ferreira Leite de Mello, Fred Freitas, Luciano de Souza Cabral, Rafael Dueire Lins, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro
Document Analysis Systems6
2014 Automatic Training Set Generation for Better Historic Document Transcription and Compression
abstract
The more complete the training set of an optical character recognition platform, the greater the chances of obtaining a better precision in transcription. The development of a database for such purpose is a task of paramount effort as it is performed manually and must be as extensive as possible in order to potentially cover all words in a language. Dealing with historic documents either handwritten, typed, or printed is even a harder effort as documents are often degraded by time and storage conditions. The recent work of Silva-Lins showed how to automatically generate training sets of isolated characters for cursive writing of one specific person. This is particularly important in the transcription of historic files of important people. The present work improves that strategy by analyzing letter ligature patterns. The improvement in OCR transcription accuracy both of printed, typed and handwritten documents is borne out by experimental evidence.
Gabriel Pereira e Silva, Rafael Dueire Lins, Cesar Gomes
Document Analysis Systems1
2014 A multi-document summarization system based on statistics and linguistic treatment
Rafael Ferreira Leite de Mello, Luciano de Souza Cabral, Fred Freitas, Rafael Dueire Lins, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro
Expert Syst. Appl.5
2013 A Four Dimension Graph Model for Automatic Text Summarization
abstract
Text summarization is the process of automatically creating a shorter version of one or more text documents. In this context, word-based, sentence-based and graph-based methods approaches are largely used. Among these, graph based methods for automatic text summarization produce summaries based on the relationships between sentences. These relationships may also support the creation of several text processing applications such as extractive and abstractive summaries, question-answering and information retrieval systems, among others. A new graph model for text processing applications is proposed in this paper. It relies on four dimensions (similarity, semantic similarity, co reference, discourse information) to create the graph. The rationale behind the proposal presented here is resorting to more dimensions than previous works, and taking into account co reference resolution, taking into account to the role of pronouns in connecting the sentences. Co reference was not used in any previous graph based summarization technique. An experiment was performed using the Text Rank algorithm with the presented approach, on the CNN corpus. The results show that the model proposed here outperforms the current approaches both quantitatively and qualitatively.
Rafael Ferreira Leite de Mello, Fred Freitas, Luciano de Souza Cabral, Rafael Dueire Lins, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro
Web Intelligence6
2013 Assessing sentence scoring techniques for extractive text summarization
Rafael Ferreira Leite de Mello, Luciano de Souza Cabral, Rafael Dueire Lins, Gabriel Pereira e Silva, Fred Freitas, George D. C. Cavalcanti, Rinaldo Lima, Steven J. Simske, Luciano Favaro
Expert Syst. Appl.4
2012 Automatic content recognition of teaching boards in the Tableau platform
Gabriel Pereira e Silva, Rafael Dueire Lins
ICPR1
2011 Automatically Discriminating between Digital and Scanned Photographs
abstract
True digital photos and the digital images of scanned photographs have very different properties. The illumination pattern and palette of the two kinds of images are different. Being able to distinguish between them is important, as each of these should be handled during printing with a class-specific pipeline of image transformation algorithms, and misclassification results in detrimental imaging effects. This paper presents an automatic classifier to discriminate between the two sources. The classifier proposed is fast enough to be embedded in the driver of any printing device today.
Rafael Dueire Lins, Gabriel Pereira e Silva, Steven J. Simske
ICDAR2
2011 Correcting Specular Noise in Multiple Images of Photographed Documents
abstract
Portable digital cameras have become omnipresent. Their low-lost, simplicity to use, flexibility, and good quality images have widened their applicability far beyond their original purpose of taking personal photos. Every day people discover new uses for them from photographing teaching boards to documents. One of the difficulties of using cameras is the occurrence of specular noise whenever the photographed object is glossy. This paper presents an efficient algorithm for removing the specular noise of photographed documents by taking multiple images with different illumination sources.
Ednardo Mariano, Rafael Dueire Lins, Gabriel Pereira e Silva, Jian Fan, Peter Majewicz, Marcelo Thielo
ICDAR3
2011 An Automatic Method for Enhancing Character Recognition in Degraded Historical Documents
abstract
Automatic optical character recognition is an important research area in document processing. There are several commercial tools for such purpose, which are becoming more efficient every day. There is still a lot to be improved, in the case of historical documents, however, due to the presence of noise and degradation. This paper presents a new approach for enhancing the character recognition in degraded historical documents. The system proposed consists in identifying regions in which there is information loss due to physical document degradation and process the document with possible candidates for the correct text transcription.
Gabriel Pereira e Silva, Rafael Dueire Lins
ICDAR1
2010 PDF profiling for B&W versus color pages cost estimation for efficient on-demand book printing
abstract
Today, the way books, magazines and newspapers are published is undergoing a democratic revolution. Digital Presses have enabled the on-demand model, which provides individuals with the opportunity to produce and publish their own books with very low upfront cost. With these new markets, opportunities, and challenges have arisen. In a traditional environment, black-and-white and color pages were printed using different presses. Later on, the book was assembled combining the pages accordingly. In a digital workflow all the pages are printed with the same press, although the page cost varies significantly between color and b/w pages. Having an accurate printing cost profiler for pdf-files is fundamental for the print-on-demand business, as jobs often have a mix of color and b/w pages. To meet the expectations of some of HP customers in the large Print Service Providers (PSPs) business, a profiler was developed which yielded a reasonable cost estimate. The industrial use of such a tool showed some discrepancies between estimated and printer log, however. The new profiler presented herein provides a more accurate account of pdf jobs to be printed. Tested on 79 "real world" pdf jobs, totaling 7,088 pages, the new profiler made only one page misclassification, while the previous one yielded 54 classification errors.
Fabio Giannetti, Gary Dispoto, Rafael Dueire Lins, Gabriel Pereira e Silva, Alexis Cabeda Faria
ACM Symposium on Document Engineering4
2010 Enhancing the Filtering-Out of the Back-to-Front Interference in Color Documents with a Neural Classifier
abstract
Back-to-front, show-through, or bleeding are the names given to the interference that appears whenever one writes or prints on both sides of translucent paper. Such interference degrades image binarization and document transcription via OCR. The technical literature presents several algorithms to remove the back-to-front noise, but no algorithm is good enough in all cases. This article presents a new technique to remove such noise in color documents which makes use of neural classifiers to evaluate the degree of intensity of the interference and besides that to indicate the existence of blur. Such classifier allows tuning the parameters of an algorithm for back-to-front interference and document enhancement.
Gabriel Pereira e Silva, Rafael Dueire Lins, J. M. Silva, Serene Banerjee, A. Kuchibhotla, Marcelo Thielo
ICPR1
2009 Image Classification to Improve Printing Quality of Mixed-Type Documents
abstract
Functional image classification is the assignment of different image types to separate classes to optimize their rendering for reading or other specific end task, and is an important area of research in the publishing and multi-Average industries. This paper presents recent research on optimizing the simultaneous classification of documents, photos and logos. Each of these is handled during printing with a class-specific pipeline of image transformation algorithms, and misclassification results in pejorative imaging effects. This paper reports on replacing an existing classifier with a Weka-based classifier that simultaneously improves accuracy (from 85.3% to 90.8%) and performance (from 1458 msec to 418 msec/image). Generic subsampling of the images further improved the performance (to 199 msec/image) with only a modest impact on accuracy (to 90.4%). A staggered subsampling approach, finally, improved both accuracy (to 96.4%) and performance (to 147 msec/image) for the Weka-base classifier. This approach did not appreciable benefit the HP classifier (85.4% accuracy, 497 msec/image). These data indicate staggered subsampling using the optimized Weka classifier substantially improves the classification accuracy and performance without resulting in additional “egregious” misclassifications (assigning photos or logos to the “document” class).
Rafael Dueire Lins, Gabriel Pereira e Silva, Steven J. Simske, Jian Fan, Mark Q. Shaw, Paulo Sá, Marcelo Thielo
ICDAR2
2007 Assessing and Improving the Quality of Document Images Acquired with Portable Digital Cameras
abstract
Professionals and students of many different areas start to use portable digital cameras to take photos of documents, instead of photocopying them. This article analyses the quality of such documents for optical character recognition and proposes ways of improving their transcription and readability.
Rafael Dueire Lins, Gabriel Pereira e Silva, André R. Gomes e Silva
ICDAR2