Sérgio D. Canuto

dblp:135/1259 · also Sérgio Daniel Canuto, Sérgio Daniel Carvalho Canuto · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
1since 2021 · last 2021
0000-0003-2973-4158ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Data mining · 67% Information retrieval · 18% Data integration and cleaning · 9%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 50% Parallel and multicore computing · 50%

Topics — the 25 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › text mining
text classification
1.342019
CluWords: Exploiting Semantic Word Clustering Representation for Enhanced Topic Modeling · WSDM 2019
Similarity-Based Synthetic Document Representations for Meta-Feature Generation in Text Classification · SIGIR 2019
A Thorough Evaluation of Distance-Based Meta-Features for Automated Text Classification · IEEE Trans. Knowl. Data Eng. 2018
Data mining › predictive modeling
classification
0.522017
Stacking Bagged and Boosted Forests for Effective Automated Classification · SIGIR 2017
An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing · SIGIR 2015
Data mining
sampling
0.522016
A practical and effective sampling selection strategy for large scale deduplication · ICDE 2016
A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015
Information retrieval › document processing › document analysis
document representation
0.412019
CluWords: Exploiting Semantic Word Clustering Representation for Enhanced Topic Modeling · WSDM 2019
Information retrieval › ranking › learning to rank
feature selection for ranking
0.412019
Risk-Sensitive Learning to Rank with Evolutionary Multi-Objective Feature Selection · ACM Trans. Inf. Syst. 2019
Information retrieval › ranking
learning to rank
0.412019
Risk-Sensitive Learning to Rank with Evolutionary Multi-Objective Feature Selection · ACM Trans. Inf. Syst. 2019
Data mining › text mining
topic modeling
0.412019
CluWords: Exploiting Semantic Word Clustering Representation for Enhanced Topic Modeling · WSDM 2019
Data mining › dimensionality reduction
feature selection
0.312018
A Thorough Evaluation of Distance-Based Meta-Features for Automated Text Classification · IEEE Trans. Knowl. Data Eng. 2018
Data mining › predictive modeling › classification › ensemble learning
boosting
0.312017
Stacking Bagged and Boosted Forests for Effective Automated Classification · SIGIR 2017
Data mining › predictive modeling › classification
ensemble learning
0.312017
Stacking Bagged and Boosted Forests for Effective Automated Classification · SIGIR 2017
Data mining › predictive modeling › classification › ensemble learning
random forest
0.312017
Stacking Bagged and Boosted Forests for Effective Automated Classification · SIGIR 2017
Data mining › predictive modeling › classification › ensemble learning
stacking
0.312017
Stacking Bagged and Boosted Forests for Effective Automated Classification · SIGIR 2017
Data mining
feature engineering
0.212016
Exploiting New Sentiment-Based Meta-level Features for Effective Sentiment Analysis · WSDM 2016
Data integration and cleaning › entity resolution
record deduplication
0.212016
A practical and effective sampling selection strategy for large scale deduplication · ICDE 2016
Data mining › text mining
sentiment analysis
0.212016
Exploiting New Sentiment-Based Meta-level Features for Effective Sentiment Analysis · WSDM 2016
Information retrieval
text analysis
0.212016
Exploiting New Sentiment-Based Meta-level Features for Effective Sentiment Analysis · WSDM 2016
Machine learning and data management
active learning
0.212015
A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015
Machine learning and data management
data selection
0.212015
A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015
Data integration and cleaning › entity resolution
deduplication
0.212015
A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015
Data integration and cleaning
entity resolution
0.212015
A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication · IEEE Trans. Knowl. Data Eng. 2015
Data mining › predictive modeling › classification › nearest neighbor classification
k-nearest neighbor classification
0.212015
An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing · SIGIR 2015
Data mining › text mining › sentiment analysis
sentiment classification
0.112017
Stacking Bagged and Boosted Forests for Effective Automated Classification · SIGIR 2017
Data mining › text mining › text classification
topic classification
0.112017
Stacking Bagged and Boosted Forests for Effective Automated Classification · SIGIR 2017
GPUs and heterogeneous computing
GPU computing
0.112015
An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing · SIGIR 2015
Parallel and multicore computing › parallel algorithms
massively parallel algorithms
0.112015
An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing · SIGIR 2015

Methods — techniques the papers use, named apart from their topics

multi-objective optimization · 0.7word embeddings · 0.4matrix factorization · 0.4local and global meta-features · 0.4hyperplane distance · 0.4evolutionary algorithm · 0.4TF-IDF · 0.4SPEA2 · 0.4distance-based meta-features · 0.3bagging · 0.3parallel kNN · 0.2GPU acceleration · 0.2
YearPublicationVenuePosition
2021 On the cost-effectiveness of neural and non-neural approaches and representations for text classification: A comprehensive comparative study
Washington Cunha, Vítor Mangaravite, Christian Gomes, Sérgio D. Canuto, Elaine Resende, Cecilia Nascimento, Felipe Viegas, Celso França, Wellington Santos Martins, Jussara M. Almeida, Thierson Couto, Leonardo Rocha 0001, Marcos André Gonçalves
Inf. Process. Manag.4
2020 Extended pre-processing pipeline for text classification: On the role of meta-feature representations, sparsification and selective sampling
Washington Cunha, Sérgio D. Canuto, Felipe Viegas, Thiago Salles, Christian Gomes, Vítor Mangaravite, Elaine Resende, Thierson Couto, Marcos André Gonçalves, Leonardo Rocha 0001
Inf. Process. Manag.2
2020 Exploiting semantic relationships for unsupervised expansion of sentiment lexicons
Felipe Viegas, Mário S. Alvim, Sérgio D. Canuto, Thierson Couto, Marcos André Gonçalves, Leonardo Rocha 0001
Inf. Syst.3
2019 Document Performance Prediction for Automatic Text Classification
Gustavo Penha, Raphael R. Campos, Sérgio D. Canuto, Marcos André Gonçalves, Rodrygo L. T. Santos
ECIR (2)3
2019 Similarity-Based Synthetic Document Representations for Meta-Feature Generation in Text Classification
abstract
We propose new solutions that enhance and extend the already very successful application of meta-features to text classification. Our newly proposed meta-features are capable of: (1) improving the correlation of small pieces of evidence shared by neighbors with labeled categories by means of synthetic document representations and (local and global) hyperplane distances; and (2) estimating the level of error introduced by these newly proposed and the existing meta-features in the literature, specially for hard-to-classify regions of the feature space. Our experiments with large and representative number of datasets show that our new solutions produce the best results in all tested scenarios, achieving gains of up to 12% over the strongest meta-feature proposal of the literature.
Sérgio D. Canuto, Thiago Salles, Thierson Couto, Marcos André Gonçalves
SIGIR1
2019 CluWords: Exploiting Semantic Word Clustering Representation for Enhanced Topic Modeling
abstract
In this paper, we advance the state-of-the-art in topic modeling by means of a new document representation based on pre-trained word embeddings for non-probabilistic matrix factorization. Specifically, our strategy, called CluWords, exploits the nearest words of a given pre-trained word embedding to generate meta-words capable of enhancing the document representation, in terms of both, syntactic and semantic information. The novel contributions of our solution include: (i)the introduction of a novel data representation for topic modeling based on syntactic and semantic relationships derived from distances calculated within a pre-trained word embedding space and (ii)the proposal of a new TF-IDF-based strategy, particularly developed to weight the CluWords. In our extensive experimentation evaluation, covering 12 datasets and 8 state-of-the-art baselines, we exceed (with a few ties) in almost cases, with gains of more than 50% against the best baselines (achieving up to 80% against some runner-ups). Finally, we show that our method is able to improve document representation for the task of automatic text classification.
Felipe Viegas, Sérgio D. Canuto, Christian Gomes, Washington Cunha, Thierson Couto, Sabir Ribas, Leonardo Rocha 0001, Marcos André Gonçalves
WSDM2
2019 Quality assessment of collaboratively-created web content with no manual intervention based on soft multi-view generation
Luiz Gonçalves 0001, Marcos André Gonçalves, Sérgio D. Canuto, Daniel Hasan Dalip, Marco Cristo, Pável Calado
Expert Syst. Appl.3
2019 Risk-Sensitive Learning to Rank with Evolutionary Multi-Objective Feature Selection
abstract
Learning to Rank (L2R) is one of the main research lines in Information Retrieval. Risk-sensitive L2R is a sub-area of L2R that tries to learn models that are good on average while at the same time reducing the risk of performing poorly in a few but important queries (e.g., medical or legal queries). One way of reducing risk in learned models is by selecting and removing noisy, redundant features, or features that promote some queries to the detriment of others. This is exacerbated by learning methods that usually maximize an average metric (e.g., mean average precision (MAP) or Normalized Discounted Cumulative Gain (NDCG)). However, historically, feature selection (FS) methods have focused only on effectiveness and feature reduction as the main objectives. Accordingly, in this work, we propose to evaluate FS for L2R with an additional objective in mind, namely risk-sensitiveness . We present novel single and multi-objective criteria to optimize feature reduction, effectiveness, and risk-sensitiveness, all at the same time. We also introduce a new methodology to explore the search space, suggesting effective and efficient extensions of a well-known Evolutionary Algorithm (SPEA2) for FS applied to L2R. Our experiments show that explicitly including risk as an objective criterion is crucial to achieving a more effective and risk-sensitive performance. We also provide a thorough analysis of our methodology and experimental results.
Daniel Xavier de Sousa, Sérgio D. Canuto, Marcos André Gonçalves, Thierson Couto, Wellington Santos Martins
ACM Trans. Inf. Syst.2
2018 Semantically-Enhanced Topic Modeling
abstract
In this paper, we advance the state-of-the-art in topic modeling by means of the design and development of a novel (semi-formal) general topic modeling framework. The novel contributions of our solution include: (i) the introduction of new semantically-enhanced data representations for topic modeling based on pooling, and (ii) the proposal of a novel topic extraction strategy - ASToC - that solves the difficulty in representing topics in our semantically-enhanced information space. In our extensive experimentation evaluation, covering 12 datasets and 12 state-of-the-art baselines, totalizing 108 tests, we exceed (with a few ties) in almost 100 cases, with gains of more than 50% against the best baselines (achieving up to 80% against some runner-ups). We provide qualitative and quantitative statistical analyses of why our solutions work so well. Finally, we show that our method is able to improve document representation in automatic text classification.
Felipe Viegas, Washington Cunha, Christian Gomes, Amir Khatibi, Sérgio D. Canuto, Fernando Mourão, Thiago Salles, Leonardo Rocha 0001, Marcos André Gonçalves
CIKM5
2018 A Thorough Evaluation of Distance-Based Meta-Features for Automated Text Classification
abstract
We address the problem of automatically learning to classify texts by exploiting information derived from meta-features, i.e., features derived from the original bag-of-words representation. Specifically, we provide an in-depth analysis on the recently proposed distance-based meta-features, a data engineering technique that relies on the distance between documents to transform the original feature space into a new one, potentially smaller and more informed. Despite its potential, the meta-feature space may be unnecessarily complex and highly dimensional, which increases the tendency of overfitting, limits the application of meta-features in different contexts, and increases computational costs. In this work, we propose the use of multi-objective strategies to reduce the number of meta-features while maximizing the classification effectiveness, when considering the adequacy of the selected meta-features to a particular dataset or classification method. We present effective and efficient proposals for meta-feature selection that can substantially reduce the number of meta-features by up to 89 percent while keeping or improving the classification effectiveness, something not possible with any of the evaluated baselines. We also use our selection strategies as evaluation tools to analyze different combinations of meta-features. We found very compact combinations of meta-features that can achieve high classification effectiveness in most datasets, despite their peculiarities.
Sérgio D. Canuto, Daniel Xavier de Sousa, Marcos André Gonçalves, Thierson Couto
IEEE Trans. Knowl. Data Eng.1
2017 Automatic Hierarchical Categorization of Research Expertise Using Minimum Information
Gustavo Oliveira de Siqueira, Sérgio D. Canuto, Marcos André Gonçalves, Alberto H. F. Laender
TPDL2
2017 Stacking Bagged and Boosted Forests for Effective Automated Classification
abstract
Random Forest (RF) is one of the most successful strategies for automated classification tasks. Motivated by the RF success, recently proposed RF-based classification approaches leverage the central RF idea of aggregating a large number of low-correlated trees, which are inherently parallelizable and provide exceptional generalization capabilities. In this context, this work brings several new contributions to this line of research. First, we propose a new RF-based strategy (BERT) that applies the boosting technique in bags of extremely randomized trees. Second, we empirically demonstrate that this new strategy, as well as the recently proposed BROOF and LazyNN_RF classifiers do complement each other, motivating us to stack them to produce an even more effective classifier. Up to our knowledge, this is the first strategy to effectively combine the three main ensemble strategies: stacking, bagging (the cornerstone of RFs) and boosting. Finally, we exploit the efficient and unbiased stacking strategy based on out-of-bag (OOB) samples to considerably speedup the very costly training process of the stacking procedure. Our experiments in several datasets covering two high-dimensional and noisy domains of topic and sentiment classification provide strong evidence in favor of the benefits of our RF-based solutions. We show that BERT is among the top performers in the vast majority of analyzed cases, while retaining the unique benefits of RF classifiers (explainability, parallelization, easiness of parameterization). We also show that stacking only the recently proposed RF-based classifiers and BERT using our OOB-based strategy is not only significantly faster than recently proposed stacking strategies (up to six times) but also much more effective, with gains up to 21% and 17% on MacroF1 and MicroF1, respectively, over the best base method, and of 5% and 6% over a stacking of traditional methods, performing no worse than a complete stacking of methods at a much lower computational effort.
Raphael R. Campos, Sérgio D. Canuto, Thiago Salles, Clebson C. A. de Sá, Marcos André Gonçalves
SIGIR2
2017 Ranked batch-mode active learning
Thiago N. C. Cardoso, Rodrigo M. Silva, Sérgio D. Canuto, Mirella M. Moro, Marcos André Gonçalves
Inf. Sci.3
2016 Incorporating Risk-Sensitiveness into Feature Selection for Learning to Rank
abstract
Learning to Rank (L2R) is currently an essential task in basically all types of information systems given the huge and ever increasing amount of data made available. While many solutions have been proposed to improve L2R functions, relatively little attention has been paid to the task of improving the quality of the feature space. L2R strategies usually rely on dense feature representations, which contain noisy or redundant features, increasing the cost of the learning process, without any benefits. Although feature selection (FS) strategies can be applied to reduce dimensionality and noise, side effects of such procedures have been neglected, such as the risk of getting very poor predictions in a few (but important) queries. In this paper we propose multi-objective FS strategies that optimize both aspects at the same time: ranking performance and risk-sensitive evaluation. For this, we approximate the Pareto-optimal set for multi-objective optimization in a new and original application to L2R. Our contributions include novel FS methods for L2R which optimize multiple, potentially conflicting, criteria. In particular, one of the objectives (risk-sensitive evaluation) has never been optimized in the context of FS for L2R before. Our experimental evaluation shows that our proposed methods select features that are more effective (ranking performance) and low-risk than those selected by other state-of-the-art FS methods.
Daniel Xavier de Sousa, Sérgio D. Canuto, Thierson Couto, Wellington Santos Martins, Marcos André Gonçalves
CIKM2
2016 A practical and effective sampling selection strategy for large scale deduplication
abstract
Record deduplication aims at identifying entities that are potentially the same in a data repository. A set of pairs that is manually labeled is generally used to tune the deduplication process, as each dataset has a particular dirtiness pattern. However, producing an informative set of pairs is a very costly task, especially in very large datasets (even for expert users). We propose a new sampling strategy that is able to select a very small and informative set of pairs from large datasets. Our results show that our approach reduces user effort substantially while achieving a competitive or superior matching quality.
Guilherme Dal Bianco, Renata Galante, Carlos Alberto Heuser, Marcos André Gonçalves, Sérgio D. Canuto
ICDE5
2016 Exploiting New Sentiment-Based Meta-level Features for Effective Sentiment Analysis
abstract
In this paper we address the problem of automatically learning to classify the sentiment of short messages/reviews by exploiting information derived from meta-level features i.e., features derived primarily from the original bag-of-words representation. We propose new meta-level features especially designed for the sentiment analysis of short messages such as: (i) information derived from the sentiment distribution among the k nearest neighbors of a given short test document x, (ii) the distribution of distances of x to their neighbors and (iii) the document polarity of these neighbors given by unsupervised lexical-based methods. Our approach is also capable of exploiting information from the neighborhood of document x regarding (highly noisy) data obtained from 1.6 million Twitter messages with emoticons. The set of proposed features is capable of transforming the original feature space into a new one, potentially smaller and more informed. Experiments performed with a substantial number of datasets (nineteen) demonstrate that the effectiveness of the proposed sentiment-based meta-level features is not only superior to the traditional bag-of-word representation (by up to 16%) but is also superior in most cases to state-of-art meta-level features previously proposed in the literature for text classification tasks that do not take into account some idiosyncrasies of sentiment analysis. Our proposal is also largely superior to the best lexicon-based methods as well as to supervised combinations of them. In fact, the proposed approach is the only one to produce the best results in all tested datasets in all scenarios.
Sérgio D. Canuto, Marcos André Gonçalves, Fabrício Benevenuto
WSDM1
2015 An Efficient and Scalable MetaFeature-based Document Classification Approach based on Massively Parallel Computing
abstract
The unprecedented growth of available data nowadays has stimulated the development of new methods for organizing and extracting useful knowledge from this immense amount of data. Automatic Document Classification (ADC) is one of such methods, that uses machine learning techniques to build models capable of automatically associating documents to well-defined semantic classes. ADC is the basis of many important applications such as language identification, sentiment analysis, recommender systems, spam filtering, among others. Recently, the use of meta-features has been shown to substantially improve the effectiveness of ADC algorithms. In particular, the use of meta-features that make a combined use of local information (through kNN-based features) and global information (through category centroids) has produced promising results. However, the generation of these meta-features is very costly in terms of both, memory consumption and runtime since there is the need to constantly call the kNN algorithm. We take advantage of the current manycore GPU architecture and present a massively parallel version of the kNN algorithm for highly dimensional and sparse datasets (which is the case for ADC). Our experimental results show that we can obtain speedup gains of up to 15x while reducing memory consumption in more than 5000x when compared to a state-of-the-art parallel baseline. This opens up the possibility of applying meta-features based classification in large collections of documents, that would otherwise take too much time or require the use of an expensive computational platform.
Sérgio D. Canuto, Marcos André Gonçalves, Wisllay M. V. dos Santos, Thierson Couto, Wellington Santos Martins
SIGIR1
2015 A Practical and Effective Sampling Selection Strategy for Large Scale Deduplication
abstract
The data deduplication task has attracted a considerable amount of attention from the research community in order to provide effective and efficient solutions. The information provided by the user to tune the deduplication process is usually represented by a set of manually labeled pairs. In very large datasets, producing this kind of labeled set is a daunting task since it requires an expert to select and label a large number of informative pairs. In this article, we propose a two-stage sampling selection strategy (T3S) that selects a reduced set of pairs to tune the deduplication process in large datasets. T3S selects the most representative pairs by following two stages. In the first stage, we propose a strategy to produce balanced subsets of candidate pairs for labeling. In the second stage, an active selection is incrementally invoked to remove the redundant pairs in the subsets created in the first stage in order to produce an even smaller and more informative training set. This training set is effectively used both to identify where the most ambiguous pairs lie and to configure the classification approaches. Our evaluation shows that T3S is able to reduce the labeling effort substantially while achieving a competitive or superior matching quality when compared with state-of-the-art deduplication methods in large datasets.
Guilherme Dal Bianco, Renata Galante, Marcos André Gonçalves, Sérgio D. Canuto, Carlos Alberto Heuser
IEEE Trans. Knowl. Data Eng.4
2014 On Efficient Meta-Level Features for Effective Text Classification
abstract
This paper addresses the problem of automatically learning to classify texts by exploiting information derived from meta-level features (i.e., features derived from the original bag-of-words representation). We propose new meta-level features derived from the class distribution, the entropy and the within-class cohesion observed in the k nearest neighbors of a given test document x, as well as from the distribution of distances of x to these neighbors. The set of proposed features is capable of transforming the original feature space into a new one, potentially smaller and more informed. Experiments performed with several standard datasets demonstrate that the effectiveness of the proposed meta-level features is not only much superior than the traditional bag-of-word representation but also superior to other state-of-art meta-level features previously proposed in the literature. Moreover, the proposed meta-features can be computed about three times faster than the existing meta-level ones, making our proposal much more scalable. We also demonstrate that the combination of our meta features and the original set of features produce significant improvements when compared to each feature set used in isolation.
Sérgio D. Canuto, Thiago Salles, Marcos André Gonçalves, Leonardo Rocha 0001, Gabriel Spada Ramos, Luiz Gonçalves 0001, Thierson Couto, Wellington Santos Martins
CIKM1