Bruno Tenório Ávila

dblp:82/6847 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0002-5409-8918ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 3 first-authorTheory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
1 paper
Coding theory · 77% Automata and formal languages · 23%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Automata and formal languages
number systems
0.312017
Meta-Fibonacci Codes: Efficient Universal Coding of Natural Numbers · IEEE Trans. Inf. Theory 2017
Coding theory › source coding › variable-length codes
prefix codes
0.312017
Meta-Fibonacci Codes: Efficient Universal Coding of Natural Numbers · IEEE Trans. Inf. Theory 2017
Coding theory › source coding
universal coding
0.312017
Meta-Fibonacci Codes: Efficient Universal Coding of Natural Numbers · IEEE Trans. Inf. Theory 2017
Coding theory › source coding › universal coding
universal coding of integers
0.312017
Meta-Fibonacci Codes: Efficient Universal Coding of Natural Numbers · IEEE Trans. Inf. Theory 2017
Coding theory › source coding › variable-length codes
kraft inequality
0.112017
Meta-Fibonacci Codes: Efficient Universal Coding of Natural Numbers · IEEE Trans. Inf. Theory 2017

Methods — techniques the papers use, named apart from their topics

meta-fibonacci sequences · 0.3
YearPublicationVenuePosition
2019 The CNN-Corpus: A Large Textual Corpus for Single-Document Extractive Summarization
abstract
This paper details the features and the methodology adopted in the construction of the CNN-corpus, a test corpus for single document extractive text summarization of news articles. The current version of the CNN-corpus encompasses 3,000 texts in English, and each of them has an abstractive and an extractive summary. The corpus allows quantitative and qualitative assessments of extractive summarization strategies.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng5
2019 The CNN-Corpus in Spanish: a Large Corpus for Extractive Text Summarization in the Spanish Language
abstract
This paper details the development and features of the CNN-corpus in Spanish, possibly the largest test corpus for single document extractive text summarization in the Spanish language. Its current version encompasses 1,117 well-written texts in Spanish, each of them has an abstractive and an extractive summary. The development methodology adopted allows good-quality qualitative and quantitative assessments of summarization strategies for tools developed in the Spanish language.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Diego A. Salcedo, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng5
2017 Meta-Fibonacci Codes: Efficient Universal Coding of Natural Numbers
abstract
In this paper, we address the problem of the universal coding of natural numbers. A new numeration system is introduced, which is based on variable-r meta-Fibonacci sequences and it is a generalization of the Zeckendorf numeration system. This new numeration system is used to construct binary, prefix-free, uniquely decodable universal codes called meta-Fibonacci codes. The main advantage of these codes is that they are parametrized by a sequence of numbers, the sequence o. By controlling the growth of the values of this sequence, we can control the length of the code word. This means that we can provide a general framework for building efficient universal coders for natural numbers. Such framework is applied to the upper bounds of the code word length defined by Leung-Yan-Cheong and Cover (1978), Levenshtein (1968), and Ahlswede (1997). There is no other code meeting these bounds. In each case, we build meta-Fibonacci codes and demonstrate that the upper bound of their code word length is satisfied up to an additive constant, thereby solving these open problems. The framework may be applied to other upper bounds that satisfy Kraft inequality.
Bruno Tenório Ávila, Ricardo M. Campello de Souza
IEEE Trans. Inf. Theory1
2016 W-tree: A Compact External Memory Representation for Webgraphs
abstract
World Wide Web applications need to use, constantly update, and maintain large webgraphs for executing several tasks, such as calculating the web impact factor, finding hubs and authorities, performing link analysis by webometrics tools, and ranking webpages by web search engines. Such webgraphs need to use a large amount of main memory, and, frequently, they do not completely fit in, even if compressed. Therefore, applications require the use of external memory. This article presents a new compact representation for webgraphs, called w-tree , which is designed specifically for external memory. It supports the execution of basic queries (e.g., full read, random read, and batch random read), set-oriented queries (e.g., superset, subset, equality, overlap, range, inlink, and co-inlink), and some advanced queries, such as edge reciprocal and hub and authority. Furthermore, a new layout tree designed specifically for webgraphs is also proposed, reducing the overall storage cost and allowing the random read query to be performed with an asymptotically faster runtime in the worst case. To validate the advantages of the w-tree, a series of experiments are performed to assess an implementation of the w-tree comparing it to a compact main memory representation. The results obtained show that w-tree is competitive in compression time and rate and in query time, which may execute several orders of magnitude faster for set-oriented queries than its competitors. The results provide empirical evidence that it is feasible to use a compact external memory representation for webgraphs in real applications, contradicting the previous assumptions made by several researchers.
Bruno Tenório Ávila, Rafael Dueire Lins
ACM Trans. Web1
2014 A platform for language independent summarization
abstract
The text data available on the Internet is not only huge in volume, but also in diversity of subject, quality and idiom. Such factors make it infeasible to efficiently scavenge useful information from it. Automatic text summarization is a possible solution for efficiently addressing such a problem, because it aims to sieve the relevant information in documents by creating shorter versions of the text. However, most of the techniques and tools available for automatic text summarization are designed only for the English language, which is a severe restriction. There are multilingual platforms that support, at most, 2 languages. This paper proposes a language independent summarization platform that provides corpus acquisition, language classification, translation and text summarization for 25 different languages.
Luciano de Souza Cabral, Rafael Dueire Lins, Rafael Ferreira Leite de Mello, Fred Freitas, Bruno Tenório Ávila, Steven J. Simske, Marcelo Riss
ACM Symposium on Document Engineering5
2009 Merge source coding
abstract
We show that any comparison-based merging algorithm can be naturally mapped into a source coder via a conversion function introduced here. By applying this function over some well known merging algorithms, namely binary merging and recursive merging, we show that they are closely related to a runlength-based coder with rice coding and to the binary interpolative coder, respectively. Furthermore, by applying the conversion function over the probabilistic merging algorithm we obtain a runlength-based coder that uses a new variant of the rice code, namely randomized rice code. This new code uses a random source of bits with the aim of reducing its average redundancy.
Bruno Tenório Ávila, Eduardo Sany Laber
ISIT1
2005 A fast orientation and skew detection algorithm for monochromatic document images
abstract
Very often in the digitization process, documents are either not placed with the correct orientation or are rotated of small angles in relation to the original image axis. These factors make more difficult the visualization of images by human users, increase the complexity of any sort of automatic image recognition, degrade the performance of OCR tools, increase the space needed for image storage, etc. This paper presents a fast algorithm for orientation and skew detection for complex monochromatic document images, which is capable of detecting any document rotation at a high precision.
Bruno Tenório Ávila, Rafael Dueire Lins
ACM Symposium on Document Engineering1
2005 A new rotation algorithm for monochromatic images
abstract
The classical rotation algorithm applied to monochromatic images introduces white holes in black areas, making edges uneven and disconnecting neighboring elements. Several algorithms in the literature address only the white hole problem. This paper proposes a new algorithm that solves those three problems, producing better quality images.
Bruno Tenório Ávila, Rafael Dueire Lins, Lamberto Oliveira
ACM Symposium on Document Engineering1
2005 BigBatch: a toolbox for monochromatic documents
abstract
BigBatch is a tool designed to automatically process thousands of monochromatic images of documents generated by production line scanners. It removes noisy borders, checks and corrects orientation, calculates and compensates the skew angle, crops the image standardizing document sizes, and finally compresses it according to user defined file format. BigBatch encompasses the best and recently developed algorithms for such kind of document images. BigBatch may work either in standalone or operator assisted modes. Besides that, BigBatch in standalone mode is able to process in clusters of workstations.
Rafael Dueire Lins, Bruno Tenório Ávila
ACM Symposium on Document Engineering2