Rafael Dueire Lins

dblp:44/4611 · DBLP profile ↗
← Back
59ranked-venue papers in the field
28as first author
12since 2021 · last 2025
0000-0003-3497-5044ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 36 (15 first)Other / Interdisciplinary · 23 (13 first)
YearPublicationVenuePosition
2025 Binarizing Photographed Document Images 2025 Quality, Time and Space Assessment
abstract
Image binarization is fundamental for document image processing. The performance of binarization algorithms depends on several factors that range from the quality of the digitalization devices to the intrinsic features of the document itself and the kind and intensity of the noises present in the image. This assessment on binarizing photographed documents evaluated the quality, time, space, and performance of five new algorithms and ninety-eight "classical" algorithms. The test data set is composed of laser and deskjet printed documents, photographed using six widely used mobile devices with the strobe flash on, off, and in auto modes under three different angles and places of capture.
Gustavo P. Chaves, Thaylor Vieira, Gabriel Pereira e Silva, Rafael Dueire Lins, Steven J. Simske
DocEng4
2024 Texture-based Document Binarization
abstract
Image binarization, the conversion of a color image into its monochromatic version, plays a key role in many document processing pipelines. The technical literature presents over a hundred different algorithms for document image binarization yielding different quality results, depending on the intrinsic features of the document image. In addition, their processing times vary significantly, influencing their applicability. The fundamental question here is: Which binarization scheme should one choose, from all the available ones, to provide a reasonable quality-time monochromatic image? A recently published paper shows that the analysis of the texture of the paper background of a scanned document may offer elements to select an appropriate binarization algorithm. This work extends such results by analyzing twelve texture descriptors using three distinct distance measures to evaluate 315 binarization schemes, providing solid evidence that the analysis of the paper texture may guide the selection of a suitable binarization algorithm for a given scanned document image.
Rodrigo Barros Bernardino, Rafael Dueire Lins, Ricardo da Silva Barboza
DocEng2
2024 Competition on Binarizing Photographed Document Images 2024 Quality, Time and Space Report
abstract
Many document processing platforms have image binarization as a key step. The performance of binarization algorithms depends on several factors that span from the quality of the digitalization devices to the intrinsic features of the document itself and the kind and intensity of the noises present in the image. Besides that, the processing time is also an important factor that may restrict the applicability of a given algorithm. This competition on binarizing photographed documents assessed the quality, time, space, and performance of twenty three new algorithms and seventy-five "classical" algorithms. The test dataset is composed of offset and deskjet printed documents, photographed using six widely-used mobile devices with the strobe flash on, off, and in auto modes under two different angles and places of capture.
Rafael Dueire Lins, Gustavo P. Chaves, Gabriel Pereira e Silva, Thaylor Vieira, Ricardo da Silva Barboza, Steven J. Simske
DocEng1
2024 Which is the most suitable scanner resolution for documents? Detailing the answer given to the question raised by Professor George Nagy
abstract
Defining the correct image resolution is a fundamental issue to preserve all the information in a document, keeping the minimum image acquisition and processing times, as well as the storage space and computer bandwidth for network transmission, allowing for best-quality information extraction and automatic transcription. How exactly to find it? That important question was recently raised by Professor George Nagy, one of the pioneer researchers in Document Engineering. That fundamental problem in document engineering was never properly addressed. The digitalization standards and recommendations widely adopted today are based on "experimental evidences". This paper presents an answer to such an important question, based on the Nyquist-Shannon theorem from Information Theory.
Rafael Dueire Lins, Daniela Raposo Nunes de Mello, Raimundo Oliveira
DocEng1
2024 Assessing the Reliability and Validity of the Measures for Automatic Text Summarization
abstract
Automatic Text Summarization (ATS) is a research area that originated in the late 1950s and has gained increasing importance with the surging amount of text data available today. One of the key challenges in this area is how to quantitatively assess the quality of the summaries produced. The three most widely quantitative measures used for this task are: ROUGE, BLEU and BERTScore. This paper attempts to comparatively evaluate the validity and reliability of such measures. The concept of Shannon' entropy from information theory served as background for this work. Experiments were conducted using the CNN corpus, focusing on news articles written in English.
Rafael Dueire Lins, Hilário Oliveira, Steven J. Simske
DocEng1
2024 Assessing Abstractive and Extractive Methods for Automatic News Summarization
abstract
Automatic Text Summarization (ATS) is a research area that originated in the late 1950s and has gained increasing importance with the surge of text data available today. ATS approaches are generally classified into extractive and abstractive methods. Extractive summarization selects the most relevant sentences from a text and copies them verbatim into the summary. Abstractive text summarization aims to generate a more cohesive and concise version of the original text. Recently, pre-trained and large language models such as BERT, BART, Pegasus, Llama, Gemma, and GPT-3 have revolutionized numerous natural language tasks, including creating more humanlike summaries. Despite the progress in recent years, there is room for improvement in directly comparing different summarization tools using a corpus containing high-quality reference summaries. This paper evaluates the quality of summaries generated by such tools, comparing the results obtained from abstractive and extractive summarization methods. Experiments were conducted using the CNN-corpus, focusing on news articles written in English and employing four evaluation measures. The experimental results unveiled that the Pegasus model outperformed others in abstractive summarization. These findings highlight the potential for further advancements in the field, suggesting that using more specialized models explicitly tailored for the summarization task may yield superior results compared to large general-purpose language models.
Hilário Oliveira, Rafael Dueire Lins
DocEng2
2023 Quality, Space and Time Competition on Binarizing Photographed Document Images
abstract
Document image binarization is a fundamental step in many document processes. No binarization algorithm performs well on all types of document images, as the different kinds of digitalization devices and the physical noises present in the document and acquired in the digitalization process alter their performance. Besides that, the processing time is also an important factor that may restrict its applicability. This competition on binarizing photographed documents assessed the quality, time, space, and performance of five new algorithms and sixty-four "classical" and alternative algorithms. The evaluation dataset is composed of laser and deskjet printed documents, photographed using six widely-used mobile devices with the strobe flash on and off, under two different angles and places of capture.
Rafael Dueire Lins, Gabriel Pereira e Silva, Gustavo P. Chaves, Ricardo da Silva Barboza, Rodrigo Barros Bernardino, Steven J. Simske
DocEng1
2022 The Winner Takes It All: Choosing the "best" Binarization Algorithm for Photographed Documents
Rafael Dueire Lins, Rodrigo Barros Bernardino, Ricardo da Silva Barboza, Raimundo Oliveira
DAS1
2022 Binarization of photographed documents image quality, processing time and size assessment
abstract
Today, over eighty percent of the world's population owns a smart-phone with an in-built camera, and they are very often used to photograph documents. Document binarization is a key process in many document processing platforms. This competition on binarizing photographed documents assessed the quality, time, space, and performance of five new algorithms and sixty-four "classical" and alternative algorithms. The evaluation dataset is composed of offset, laser, and deskjet printed documents, photographed using six widely-used mobile devices with the strobe flash on and off, under two different angles and places of capture.
Rafael Dueire Lins, Rodrigo Barros Bernardino, Ricardo da Silva Barboza, Steven J. Simske
DocEng1
2021 Direct binarization a quality-and-time efficient binarization strategy
abstract
Most of the best known binarization algorithms have grayscale conversion as a pre-processing step, before applying the binarization strategy itself. Many algorithms produce equally good or even better quality images if fed with only one component of the image, instead of its gray-scale/luminance equivalent. The time-gain here is obtained in avoiding the several floating-point calculations in converting a RGB-color image into grayscale. More than 60 binarization algorithms were tested using "real-world" images.
Rafael Dueire Lins, Rodrigo Barros Bernardino, Ricardo da Silva Barboza, Zanoni Dueire Lins
DocEng1
2021 Binarisation of photographed documents image quality and processing time assessment
abstract
Smartphones with cameras are omnipresent in today's world and are very often used to photograph documents. Document binarization is a key process in many document processing platforms. This competition on binarizing photographed documents assessed the quality and time performance of 13 new algorithms and 50 existing algorithms. The evaluation dataset is composed of offset, laser, and deskjet printed documents, photographed using four widely-used mobile devices with the strobe flash on and off, under two different angles and places of capture.
Rafael Dueire Lins, Steven J. Simske, Rodrigo Barros Bernardino
DocEng1
2021 ICDAR 2021 Competition on Time-Quality Document Image Binarization
Rafael Dueire Lins, Rodrigo Barros Bernardino, Elisa H. Barney Smith, Ergina Kavallieratou
ICDAR (4)1
2020 DocEng'2020 Competition on Extractive Text Summarization
abstract
The DocEng'2020 Competition on Extractive Text Summarization assessed the performance of six new methods and fourteen classical algorithms for extractive text sumarization. The systems were evaluated using the CNN-Corpus, the largest test set available today for single document extractive summarization using two different strategies and the ROUGE and the direct match measures.
Rafael Dueire Lins, Rafael Ferreira Leite de Mello, Steven J. Simske
DocEng1
2020 DocEng'2020 Time-Quality Competition on Binarizing Photographed Documents
abstract
Document image binarization is a key process in many document processing platforms. The DocEng'2020 Time-Quality Competition on Binarizing Photographed Documents assessed the performance of eight new algorithms and also 41 other "classical" algorithms. Besides the quality of the binary image, the execution time of the algorithms was assessed. The evaluation dataset is composed of 32 documents photographed using four widely-used mobile devices with the strobe flash on and off, under several different angles of capture.
Rafael Dueire Lins, Steven J. Simske, Rodrigo Barros Bernardino
DocEng1
2020 An Assessment of Sentence Simplification Methods in Extractive Text Summarization
abstract
The unprecedented growth of textual content on the Web made essential the development of automatic or semi-automatic techniques to help people to find valuable information in such a huge heap of text data. Automatic text summarization is one of such techniques that is being pointed out as offering a viable solution in such a chaotic scenario. Extractive text summarization, in particular, selects a set of sentences from a text according to specific criteria. Strategies for extractive summarization can benefit from preprocessing techniques that emphasize the relevance or infor-mativeness of sentences with respect to the selection criteria. This paper tests such a hypothesis using sentence simplification methods. Four methods are used to simplify a corpus of news articles in English: a rule-based method, an optimization method, a supervised deep learning model and an unsupervised deep learning model. The simplified outputs are summarized using 14 sentence selection strategies. The combinations of simplification and summarization methods are compared with the baseline --- the summarized corpus without previous simplification --- with a quantitative analysis, which suggests sentence compression with restrictions and models learned from large parallel corpora tend to perform better and yield gains over summarization without prior simplification.
Rafaella F. Vale, Rafael Dueire Lins, Rafael Ferreira Leite de Mello
DocEng2
2019 DocEng'19 Competition on Extractive Text Summarization
abstract
The DocEng'19 Competition on Extractive Text Summarization assessed the performance of two new and fourteen previously published extractive text sumarization methods. The competitors were evaluated using the CNN-Corpus, the largest test set available today for single document extractive summarization.
Rafael Dueire Lins, Rafael Ferreira Leite de Mello, Steven J. Simske
DocEng1
2019 The CNN-Corpus: A Large Textual Corpus for Single-Document Extractive Summarization
abstract
This paper details the features and the methodology adopted in the construction of the CNN-corpus, a test corpus for single document extractive text summarization of news articles. The current version of the CNN-corpus encompasses 3,000 texts in English, and each of them has an abstractive and an extractive summary. The corpus allows quantitative and qualitative assessments of extractive summarization strategies.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng1
2019 The CNN-Corpus in Spanish: a Large Corpus for Extractive Text Summarization in the Spanish Language
abstract
This paper details the development and features of the CNN-corpus in Spanish, possibly the largest test corpus for single document extractive text summarization in the Spanish language. Its current version encompasses 1,117 well-written texts in Spanish, each of them has an abstractive and an extractive summary. The development methodology adopted allows good-quality qualitative and quantitative assessments of summarization strategies for tools developed in the Spanish language.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Diego A. Salcedo, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng1
2019 Generating Digital Libraries of M.Sc. and Ph.D. Theses
abstract
Postgraduate degrees are one of the most important propellers of all areas of science. M.Sc. and Ph.D. theses witness the important developments and provide a solid and global account of research projects. This paper describes a platform developed with the aim of generating digital libraries of theses and dissertations. Printed theses have to be scanned and then processed by the platform for marginal border removal, skew and orientation correction, image segmentation and enhancement, compression and pdf file generation. Both scanned and digitally generated theses are processed in the platform to extract relevant indexing information.
Rafael Dueire Lins, Paulo Hugo Espirito Santo, Gabriel Pereira e Silva
DocEng1
2019 Enhancing Document-Camera Images
abstract
Document Camera digitalization devices are low-cost, easy to use, produce good quality images, are able to digitalize pages of bound books without damaging their spine, etc. On the other hand, they may bring two serious problems. The first one appears if the document to be digitalized is printed on glossy paper. The paper reflects the different illumination sources from the environment producing a specular noise in the document image. The second problem occurs when the document to be digitalized does not lay flat on the digitalization surface. This paper presents solutions to both problems. The results obtained in almost 400 test images may be considered satisfactory.
Ednardo Mariano, Rafael Dueire Lins, Jian Fan
DocEng2
2019 A Quality and Time Assessment of Binarization Algorithms
abstract
Binarization algorithms are an important step in most document analysis and recognition applications. Many aspects of the document affect the performance of binarization algorithms, such as paper texture and color, noises such as the back-to-front interference, stains, and even the type and color of the ink. This work focuses on determining how each document characteristic impacts the time to process and the quality of the binarized image. This paper assesses thirty of the most widely used document binarization algorithms.
Rafael Dueire Lins, Rodrigo Barros Bernardino, Darlisson Marinho de Jesus
ICDAR1
2019 ICDAR 2019 Time-Quality Binarization Competition
abstract
The ICDAR 2019 Time-Quality Binarization Competition assessed the performance of seventeen new together with thirty previously published binarization algorithms. The quality of the resulting two-tone image and the execution time were assessed. Comparisons were on both in "real-world" and synthetic scanned images, and in documents photographed with four models of widely used portable phones. Most of the submitted algorithms employed machine learning techniques and performed best on the most complex images. Traditional algorithms provided very good results at a fraction of the time.
Rafael Dueire Lins, Ergina Kavallieratou, Elisa H. Barney Smith, Rodrigo Barros Bernardino, Darlisson Marinho de Jesus
ICDAR1
2018 Automatic Text Summarization and Classification
abstract
In this tutorial, we consider important aspects (algorithms, approaches, considerations) for tagging both unstructured and structured text for downstream use. This includes summarization, in which text information is compressed for more efficient archiving, searching, and clustering. In the tutorial, we focus on the topic of automatic text summarization, covering the most important milestones of the six decades of research in this area.
Steven J. Simske, Rafael Dueire Lins
DocEng2
2017 Assessing Binarization Techniques for Document Images
abstract
Image binarization is a technique widely used for documents as monochromatic documents claim for far less space for storage and computer bandwidth for network transmission than their color or even grayscale equivalent. Paper color, texture, aging, translucidity, kind and color of ink used in handwritting, printing process, digitalization process, etc., are some of the factors that affect binarization. No algorithm is good enough to be a winner in the binarization of all kinds of documents. This paper presents a methodology to assess the performance of binarization algorithms for a wide variety of text documents, allowing a judicious quantitative choice of the best algorithms and their parameters.
Rafael Dueire Lins, Marcos Martins de Almeida, Rodrigo Barros Bernardino, Darlisson Marinho de Jesus, José Mário Oliveira
DocEng1
2016 Towards Cohesive Extractive Summarization through Anaphoric Expression Resolution
abstract
This paper presents a new method for improving the cohesiveness of summaries generated by extractive summarization systems. The solution presented attempts to improve the legibility and cohesion of the generated summaries through coreference resolution. It is based on a post-processing step that binds dangling coreference to the most important entity in a given coreference chain. The proposed solution was evaluated on the CNN corpus of 3,000 news articles, using four state-of-the-art summarization systems and seventeen techniques for sentence scoring proposed in the literature. The experimental results may be considered encouraging, as the final summaries reached better ROUGE scores, besides being more cohesive.
Jamilson Batista, Rafael Dueire Lins, Rinaldo Lima, Steven J. Simske, Marcelo Riss
DocEng2
2016 Mobile Summarizer and News Summary Navigator: Two Multilingual News Article Summarization Tools for Mobile Devices
abstract
Mobile devices such as smart phones and tablets are omnipresent in modern societies. Such devices allow browsing the Internet. This paper briefly describes two tools for news article summarization in mobile devices that attempts to automatically collect and sieve the most important information of news article in WebPages.
Luciano de Souza Cabral, Manoel Neto, Artur Borges, Rafael Dueire Lins, Rinaldo Lima, Rafael Ferreira Leite de Mello, Marcelo Riss, Steven J. Simske
DocEng4
2016 Appling Link Target Identification and Content Extraction to improve Web News Summarization
abstract
The existing automatic text summarization systems whenever applied to web-pages of news articles show poor performance as the text is encapsulated within a HTML page. This paper takes advantage of the link identification and content extraction techniques. The results show the validity of such a strategy.
Rodolfo Ferreira, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Hilário Oliveira, Marcelo Riss, Steven J. Simske
DocEng3
2016 Assessing Concept Weighting in Integer Linear Programming based Single-document Summarization
abstract
Some of the recent state-of-the-art systems for Automatic Text Summarization rely on the concept-based approach using Integer Linear Programming (ILP), mainly for multi-document summarization. A study on the suitability of such an approach to single-document summarization is still missing, however. This work presents an assessment of several methods of concept weighing for a concept-based ILP approach on the single-document summarization scenario. The unigram and bigram representations for concepts are also investigated. The experimental results obtained on the DUC 2001-2002 and the CNN corpora show that bigrams are more suitable than unigrams for the representation of concepts. Among the concept scoring methods investigated, the sentence position method presented the best performance on all evaluation corpora.
Hilário Oliveira, Rinaldo Lima, Rafael Dueire Lins, Fred Freitas, Marcelo Riss, Steven J. Simske
DocEng3
2016 W-tree: A Compact External Memory Representation for Webgraphs
abstract
World Wide Web applications need to use, constantly update, and maintain large webgraphs for executing several tasks, such as calculating the web impact factor, finding hubs and authorities, performing link analysis by webometrics tools, and ranking webpages by web search engines. Such webgraphs need to use a large amount of main memory, and, frequently, they do not completely fit in, even if compressed. Therefore, applications require the use of external memory. This article presents a new compact representation for webgraphs, called w-tree , which is designed specifically for external memory. It supports the execution of basic queries (e.g., full read, random read, and batch random read), set-oriented queries (e.g., superset, subset, equality, overlap, range, inlink, and co-inlink), and some advanced queries, such as edge reciprocal and hub and authority. Furthermore, a new layout tree designed specifically for webgraphs is also proposed, reducing the overall storage cost and allowing the random read query to be performed with an asymptotically faster runtime in the worst case. To validate the advantages of the w-tree, a series of experiments are performed to assess an implementation of the w-tree comparing it to a compact main memory representation. The results obtained show that w-tree is competitive in compression time and rate and in query time, which may execute several orders of magnitude faster for set-oriented queries than its competitors. The results provide empirical evidence that it is feasible to use a compact external memory representation for webgraphs in real applications, contradicting the previous assumptions made by several researchers.
Bruno Tenório Ávila, Rafael Dueire Lins
ACM Trans. Web2
2015 A Quantitative and Qualitative Assessment of Automatic Text Summarization Systems
abstract
Text summarization is the process of automatically creating a shorter version of one or more text documents. This paper presents a qualitative and quantitative assessment of the 22 state-of-the-art extractive summarization systems using the CNN corpus, a dataset of 3,000 news articles.
Jamilson Batista, Rodolfo Ferreira, Hilário Tomaz, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Steven J. Simske, Gabriel Pereira e Silva, Marcelo Riss
DocEng5
2015 Automatic Document Classification using Summarization Strategies
abstract
An efficient way to automatically classify documents may be provided by automatic text summarization, the task of creating a shorter text from one or several documents. This paper presents an assessment of the 15 most widely used methods for automatic text summarization from the text classification perspective. A naive Bayes classifier was used showing that some of the methods tested are better suited for such a task.
Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Luciano de Souza Cabral, Fred Freitas, Steven J. Simske, Marcelo Riss
DocEng2
2015 Automatic Text Document Summarization Based on Machine Learning
abstract
The need for automatic generation of summaries gained importance with the unprecedented volume of information available in the Internet. Automatic systems based on extractive summarization techniques select the most significant sentences of one or more texts to generate a summary. This article makes use of Machine Learning techniques to assess the quality of the twenty most referenced strategies used in extractive summarization, integrating them in a tool. Quantitative and qualitative aspects were considered in such assessment demonstrating the validity of the proposed scheme. The experiments were performed on the CNN-corpus, possibly the largest and most suitable test corpus today for benchmarking extractive summarization strategies.
Gabriel Pereira e Silva, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Luciano de Souza Cabral, Hilário Oliveira, Steven J. Simske, Marcelo Riss
DocEng3
2015 Binarizing complex scanned documents
abstract
Most binarization algorithms are suitable for scanned text documents and do not work adequately with complex documents that encompass photos, charts and text simultaneously. This paper presents a new binarization algorithm that works on complex documents. Each of the elements in the image is processed depending on its nature, thus their binarization takes that into account to preserve the original content.
Rafael Dueire Lins, Gabriel Pereira e Silva, Marcos Martins de Almeida
ICDAR1
2014 A Context Based Text Summarization System
abstract
Text summarization is the process of creating a shorter version of one or more text documents. Automatic text summarization has become an important way of finding relevant information in large text libraries or in the Internet. Extractive text summarization techniques select entire sentences from documents according to some criteria to form a summary. Sentence scoring is the technique most used for extractive text summarization, today. Depending on the context, however, some techniques may yield better results than some others. This paper advocates the thesis that the quality of the summary obtained with combinations of sentence scoring methods depend on text subject. Such hypothesis is evaluated using three different contexts: news, blogs and articles. The results obtained show the validity of the hypothesis formulated and point at which techniques are more effective in each of those contexts studied.
Rafael Ferreira Leite de Mello, Fred Freitas, Luciano de Souza Cabral, Rafael Dueire Lins, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro
Document Analysis Systems4
2014 Automatic Training Set Generation for Better Historic Document Transcription and Compression
abstract
The more complete the training set of an optical character recognition platform, the greater the chances of obtaining a better precision in transcription. The development of a database for such purpose is a task of paramount effort as it is performed manually and must be as extensive as possible in order to potentially cover all words in a language. Dealing with historic documents either handwritten, typed, or printed is even a harder effort as documents are often degraded by time and storage conditions. The recent work of Silva-Lins showed how to automatically generate training sets of isolated characters for cursive writing of one specific person. This is particularly important in the transcription of historic files of important people. The present work improves that strategy by analyzing letter ligature patterns. The improvement in OCR transcription accuracy both of printed, typed and handwritten documents is borne out by experimental evidence.
Gabriel Pereira e Silva, Rafael Dueire Lins, Cesar Gomes
Document Analysis Systems2
2014 A platform for language independent summarization
abstract
The text data available on the Internet is not only huge in volume, but also in diversity of subject, quality and idiom. Such factors make it infeasible to efficiently scavenge useful information from it. Automatic text summarization is a possible solution for efficiently addressing such a problem, because it aims to sieve the relevant information in documents by creating shorter versions of the text. However, most of the techniques and tools available for automatic text summarization are designed only for the English language, which is a severe restriction. There are multilingual platforms that support, at most, 2 languages. This paper proposes a language independent summarization platform that provides corpus acquisition, language classification, translation and text summarization for 25 different languages.
Luciano de Souza Cabral, Rafael Dueire Lins, Rafael Ferreira Leite de Mello, Fred Freitas, Bruno Tenório Ávila, Steven J. Simske, Marcelo Riss
ACM Symposium on Document Engineering2
2014 A new sentence similarity assessment measure based on a three-layer sentence representation
abstract
Sentence similarity is used to measure the degree of likelihood between sentences. It is used in many natural language applications, such as text summarization, information retrieval, text categorization, and machine translation. The current methods for assessing sentence similarity represent sentences as vectors of bag of words or the syntactic information of the words in the sentence. The degree of likelihood between phrases is calculated by composing the similarity between the words in the sentences. Two important concerns in the area, the meaning problem and the word order, are not handled, however. This paper proposes a new sentence similarity assessment measure that largely improves and refines a recently published method that takes into account the lexical, syntactic and semantic components of sentences. The new method proposed here was benchmarked using a publically available standard dataset. The results obtained show that the new similarity assessment measure proposed outperforms the state of the art systems and achieve results comparable to the evaluation made by humans.
Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Fred Freitas, Steven J. Simske, Marcelo Riss
ACM Symposium on Document Engineering2
2014 Transforming graph-based sentence representations to alleviate overfitting in relation extraction
abstract
Relation extraction (RE) aims at finding the way entities, such as person, location, organization, date, etc., depend upon each other in a text document. Ontology Population, Automatic Summarization, and Question Answering are fields in which relation extraction offers valuable solutions. A relation extraction method based on inductive logic programming that induces extraction rules suitable to identify semantic relations between entities was proposed by the authors in a previous work. This paper proposes a method to simplify graph-based representations of sentences that replaces dependency graphs of sentences by simpler ones, keeping the target entities in it. The goal is to speed up the learning phase in a RE framework, by applying several rules for graph simplification that constrain the hypothesis space for generating extraction rules. Moreover, the direct impact on the extraction performance results is also investigated. The proposed techniques outperformed some other state-of-the-art systems when assessed on two standard datasets for relation extraction in the biomedical domain.
Rinaldo Lima, Jamilson Batista, Rafael Ferreira Leite de Mello, Fred Freitas, Rafael Dueire Lins, Steven J. Simske, Marcelo Riss
ACM Symposium on Document Engineering5
2013 A Color-Based Model to Determine the Age of Documents for Forensic Purposes
abstract
Detecting the age of a document is an important subject for forensic purposes. As the paper ages, its color changes depending on a number of factors such as its original color, storage conditions, environment temperature, humidity, etc. In Brazil, documents such and birth and wedding certificates during the second half of the 20th century used standardized preprinted forms. This paper proposes a model for determining the age of such documents based on the color components of the background of its scanned image.
Ricardo da Silva Barboza, Rafael Dueire Lins, Darlisson Marinho de Jesus
ICDAR2
2013 An Efficient Algorithm for Segmenting Warped Text-Lines in Document Images
abstract
Warped text-lines often appear whenever one performs the digitalization of bound documents using flatbed scanners or digital cameras. Compensating such distortion is an important pre-processing step in document transcription via OCR, for instance. This paper presents an efficient algorithm for text-line segmentation for document images. A typographic study and parameter tuning are done yielding into high values for precision, recall and f-measure metrics. The method presented outperforms the competing algorithms using a public available dataset.
Daniel M. Oliveira, Rafael Dueire Lins, Gabriel Torreão, Jian Fan, Marcelo Thielo
ICDAR2
2013 A Four Dimension Graph Model for Automatic Text Summarization
abstract
Text summarization is the process of automatically creating a shorter version of one or more text documents. In this context, word-based, sentence-based and graph-based methods approaches are largely used. Among these, graph based methods for automatic text summarization produce summaries based on the relationships between sentences. These relationships may also support the creation of several text processing applications such as extractive and abstractive summaries, question-answering and information retrieval systems, among others. A new graph model for text processing applications is proposed in this paper. It relies on four dimensions (similarity, semantic similarity, co reference, discourse information) to create the graph. The rationale behind the proposal presented here is resorting to more dimensions than previous works, and taking into account co reference resolution, taking into account to the role of pronouns in connecting the sentences. Co reference was not used in any previous graph based summarization technique. An experiment was performed using the Text Rank algorithm with the presented approach, on the CNN corpus. The results show that the model proposed here outperforms the current approaches both quantitatively and qualitatively.
Rafael Ferreira Leite de Mello, Fred Freitas, Luciano de Souza Cabral, Rafael Dueire Lins, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro
Web Intelligence4
2012 A Strategy for Automatically Extracting References from PDF Documents
abstract
Every day the number of citations an author receives is becoming more important than the size of his list of publications. The automatic extraction of bibliographic references in scientific articles is still a difficult problem in Document Engineering, even if the document is originally in digital form. This paper presents a strategy for extracting references of scientific documents in PDF format. The scheme proposed was validated in Live Memory platform, developed to generate digital libraries of proceedings of technical events.
Neide Ferreira Alves, Rafael Dueire Lins, Maria Lencastre 0001
Document Analysis Systems2
2011 Using Readers' Highlighting on Monochromatic Documents for Automatic Text Transcription and Summarization
abstract
Very often interested readers highlight documents with felt pens. Such marking may be seen as personal view of the most important aspects of the document, which are used to instantly draw the readers' attention at. This article addresses ways of using the highlighting made by the reader of a book or a paper documents to automatically generate its text summary.
Ricardo da Silva Barboza, Rafael Dueire Lins, Victor Matheus de S. Pereira
ICDAR2
2011 Automatically Discriminating between Digital and Scanned Photographs
abstract
True digital photos and the digital images of scanned photographs have very different properties. The illumination pattern and palette of the two kinds of images are different. Being able to distinguish between them is important, as each of these should be handled during printing with a class-specific pipeline of image transformation algorithms, and misclassification results in detrimental imaging effects. This paper presents an automatic classifier to discriminate between the two sources. The classifier proposed is fast enough to be embedded in the driver of any printing device today.
Rafael Dueire Lins, Gabriel Pereira e Silva, Steven J. Simske
ICDAR1
2011 Correcting Specular Noise in Multiple Images of Photographed Documents
abstract
Portable digital cameras have become omnipresent. Their low-lost, simplicity to use, flexibility, and good quality images have widened their applicability far beyond their original purpose of taking personal photos. Every day people discover new uses for them from photographing teaching boards to documents. One of the difficulties of using cameras is the occurrence of specular noise whenever the photographed object is glossy. This paper presents an efficient algorithm for removing the specular noise of photographed documents by taking multiple images with different illumination sources.
Ednardo Mariano, Rafael Dueire Lins, Gabriel Pereira e Silva, Jian Fan, Peter Majewicz, Marcelo Thielo
ICDAR2
2011 An Automatic Method for Enhancing Character Recognition in Degraded Historical Documents
abstract
Automatic optical character recognition is an important research area in document processing. There are several commercial tools for such purpose, which are becoming more efficient every day. There is still a lot to be improved, in the case of historical documents, however, due to the presence of noise and degradation. This paper presents a new approach for enhancing the character recognition in degraded historical documents. The system proposed consists in identifying regions in which there is information loss due to physical document degradation and process the document with possible candidates for the correct text transcription.
Gabriel Pereira e Silva, Rafael Dueire Lins
ICDAR2
2010 PDF profiling for B&W versus color pages cost estimation for efficient on-demand book printing
abstract
Today, the way books, magazines and newspapers are published is undergoing a democratic revolution. Digital Presses have enabled the on-demand model, which provides individuals with the opportunity to produce and publish their own books with very low upfront cost. With these new markets, opportunities, and challenges have arisen. In a traditional environment, black-and-white and color pages were printed using different presses. Later on, the book was assembled combining the pages accordingly. In a digital workflow all the pages are printed with the same press, although the page cost varies significantly between color and b/w pages. Having an accurate printing cost profiler for pdf-files is fundamental for the print-on-demand business, as jobs often have a mix of color and b/w pages. To meet the expectations of some of HP customers in the large Print Service Providers (PSPs) business, a profiler was developed which yielded a reasonable cost estimate. The industrial use of such a tool showed some discrepancies between estimated and printer log, however. The new profiler presented herein provides a more accurate account of pdf jobs to be printed. Tested on 79 "real world" pdf jobs, totaling 7,088 pages, the new profiler made only one page misclassification, while the previous one yielded 54 classification errors.
Fabio Giannetti, Gary Dispoto, Rafael Dueire Lins, Gabriel Pereira e Silva, Alexis Cabeda Faria
ACM Symposium on Document Engineering3
2009 Image Classification to Improve Printing Quality of Mixed-Type Documents
abstract
Functional image classification is the assignment of different image types to separate classes to optimize their rendering for reading or other specific end task, and is an important area of research in the publishing and multi-Average industries. This paper presents recent research on optimizing the simultaneous classification of documents, photos and logos. Each of these is handled during printing with a class-specific pipeline of image transformation algorithms, and misclassification results in pejorative imaging effects. This paper reports on replacing an existing classifier with a Weka-based classifier that simultaneously improves accuracy (from 85.3% to 90.8%) and performance (from 1458 msec to 418 msec/image). Generic subsampling of the images further improved the performance (to 199 msec/image) with only a modest impact on accuracy (to 90.4%). A staggered subsampling approach, finally, improved both accuracy (to 96.4%) and performance (to 147 msec/image) for the Weka-base classifier. This approach did not appreciable benefit the HP classifier (85.4% accuracy, 497 msec/image). These data indicate staggered subsampling using the optimized Weka classifier substantially improves the classification accuracy and performance without resulting in additional “egregious” misclassifications (assigning photos or logos to the “document” class).
Rafael Dueire Lins, Gabriel Pereira e Silva, Steven J. Simske, Jian Fan, Mark Q. Shaw, Paulo Sá, Marcelo Thielo
ICDAR1
2008 Cyclic reference counting
Rafael Dueire Lins
Inf. Process. Lett.1
2007 Assessing and Improving the Quality of Document Images Acquired with Portable Digital Cameras
abstract
Professionals and students of many different areas start to use portable digital cameras to take photos of documents, instead of photocopying them. This article analyses the quality of such documents for optical character recognition and proposes ways of improving their transcription and readability.
Rafael Dueire Lins, Gabriel Pereira e Silva, André R. Gomes e Silva
ICDAR1
2006 New Algorithms and Applications of Cyclic Reference Counting
Rafael Dueire Lins
ICGT1
2005 A fast orientation and skew detection algorithm for monochromatic document images
abstract
Very often in the digitization process, documents are either not placed with the correct orientation or are rotated of small angles in relation to the original image axis. These factors make more difficult the visualization of images by human users, increase the complexity of any sort of automatic image recognition, degrade the performance of OCR tools, increase the space needed for image storage, etc. This paper presents a fast algorithm for orientation and skew detection for complex monochromatic document images, which is capable of detecting any document rotation at a high precision.
Bruno Tenório Ávila, Rafael Dueire Lins
ACM Symposium on Document Engineering2
2005 A new rotation algorithm for monochromatic images
abstract
The classical rotation algorithm applied to monochromatic images introduces white holes in black areas, making edges uneven and disconnecting neighboring elements. Several algorithms in the literature address only the white hole problem. This paper proposes a new algorithm that solves those three problems, producing better quality images.
Bruno Tenório Ávila, Rafael Dueire Lins, Lamberto Oliveira
ACM Symposium on Document Engineering2
2005 BigBatch: a toolbox for monochromatic documents
abstract
BigBatch is a tool designed to automatically process thousands of monochromatic images of documents generated by production line scanners. It removes noisy borders, checks and corrects orientation, calculates and compensates the skew angle, crops the image standardizing document sizes, and finally compresses it according to user defined file format. BigBatch encompasses the best and recently developed algorithms for such kind of document images. BigBatch may work either in standalone or operator assisted modes. Besides that, BigBatch in standalone mode is able to process in clusters of workstations.
Rafael Dueire Lins, Bruno Tenório Ávila
ACM Symposium on Document Engineering1
2002 Generation of images of historical documents by composition
abstract
This paper describes a system for efficient storage, indexing and network transmission of images of historical documents. The documents are first decomposed into their features such as paper texture, colours, typewritten parts, pictures, etc. Document retrieval forces the re-assembling of the document, synthetising an image visually close to the original document. The information needed to build the final image occupies, in average, 2 Kbytes performing a very efficient compression scheme.
Carlos A. B. Mello, Rafael Dueire Lins
ACM Symposium on Document Engineering2
2002 An efficient algorithm for cyclic reference counting
Rafael Dueire Lins
Inf. Process. Lett.1
1993 Generational Cyclic Reference Counting
Rafael Dueire Lins
Inf. Process. Lett.1
1992 Cyclic Reference Counting with Lazy Mark-Scan
Rafael Dueire Lins
Inf. Process. Lett.1
1990 Cycle Reference Counting with Local Mark-Scan
Alejandro D. Martínez, Rosita Wachenchauzer, Rafael Dueire Lins
Inf. Process. Lett.3