VLDB 2026 Research / reviewers in the wild / expert
Solange Oliveira Rezende
dblp:90/5551 · also Solange O. Rezende
· DBLP profile ↗
58ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-5233-7639ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 7 since 2021Databases, data management, data science and information retrieval · 19 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021Software engineering, systems software and programming languages · 5Graphics, computer vision, multimedia, augmented reality and games · 5Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EPHG-CR: embedding propagation for heterogeneous graphs with class refinementabstractAbstract Heterogeneous graphs can represent real-world problems in a way close to reality, supporting diverse types of vertices and edges. However, their inherent heterogeneity poses challenges in interpreting problem semantics. To address this, heterogeneous graph embedding, aiming to map graph elements to low-dimensional vectors, simplifies subsequent machine learning analysis. This approach has gained prominence in machine learning, fueling classification, recommendation, and similarity search applications. Embedding diverse data is essential for efficient data processing. Incorporating language models, like BERT, into heterogeneous graphs enhances semantic context capture, which is particularly useful when one vertex type represents text. Language models stand out in contextual representation, enriching graph vertex embeddings for various tasks. This paper proposes a novel approach to enhancing heterogeneous graph embeddings by combining language models and task class data. Our approach increases vector quality, accounting for graph structure, semantic textual information, and task labels. We compared our proposal with a language model in the aspect-based sentiment analysis task, demonstrating competitive results and, in some cases, a slight superiority. Furthermore, we explore applications of embeddings from auxiliary vertices in another task, highlighting another advantage of the approach over the language model. Brucce Neves dos Santos, Ricardo M. Marcacini, Alípio Mário Jorge, Ricardo Campos 0001, Solange Oliveira Rezende |
Appl. Intell. | 5 |
| 2026 | Explainable visual emotion recognition via modular reasoningabstractAffective computing still faces significant challenges in mapping complex visual features to emotional states with high interpretability. Although Multimodal Large Language Models (MLLMs) offer strong generalization, their high computational cost for fine-tuning constitutes a limitation. Moreover, MLLMs often function as “black boxes”, resulting in a lack of transparency in how they prioritize visual cues, especially in zero-shot settings. This work investigates these limitations by proposing a modular reasoning strategy to enhance interpretability and enable more disentangled affective reasoning under noisy cues and resource-constrained conditions. We introduce the Chain-of-Responsibility (CoR), a multi-agent framework that decomposes the affective inference process into specialized agents (Facial, Body, and Contextual). Unlike traditional end-to-end models or single-prompt reasoning, CoR employs modular decomposition designed to expose intermediate modality-specific analyses. A Synthesizer agent finally integrates these analyses using an explicit priority policy enforced through prompt design. This design enhances transparency and auditability of the decision process, allowing researchers to inspect how facial, bodily, and contextual cues contribute to the final prediction, while maintaining competitive performance in zero-shot settings. Magaly Lika Fujimoto, Ricardo M. Marcacini, Solange Oliveira Rezende |
Pattern Recognit. | 3 |
| 2025 | Content-Based Macroscopic Microbial Image Retrieval
Antonio Rafael Sabino Parmezan, Angela Patricia Mestas Muñante, Diego Minatel, Solange Oliveira Rezende |
IEEE Big Data | 4 |
| 2025 | A Spatio-Temporal Approach for Identifying Microorganisms in Short Image Sequences
Antonio Rafael Sabino Parmezan, João Pedro Ribeiro da Silva, Diego Minatel, Solange Oliveira Rezende |
IEEE Big Data | 4 |
| 2025 | How do financial time series enhance the detection of news significance in market movements? A study using graph neural networks with heterogeneous representations
Ivan J. Reis Filho, Marcos P. S. Gôlo, Ricardo M. Marcacini, Solange Oliveira Rezende |
Neural Comput. Appl. | 4 |
| 2024 | Keywords attention for fake news detection using few positive labels
Mariana Caravanti de Souza, Marcos P. S. Gôlo, Alípio Mário Jorge, Evelin Amorim, Ricardo Campos 0001, Ricardo M. Marcacini, Solange Oliveira Rezende |
Inf. Sci. | 7 |
| 2023 | One-class learning for fake news detection through multimodal variational autoencoders
Marcos P. S. Gôlo, Mariana Caravanti de Souza, Rafael Geraldeli Rossi, Solange Oliveira Rezende, Bruno M. Nogueira 0001, Ricardo M. Marcacini |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | A network-based positive and unlabeled learning approach for fake news detectionabstractFake news can rapidly spread through internet users and can deceive a large audience. Due to those characteristics, they can have a direct impact on political and economic events. Machine Learning approaches have been used to assist fake news identification. However, since the spectrum of real news is broad, hard to characterize, and expensive to label data due to the high update frequency, One-Class Learning (OCL) and Positive and Unlabeled Learning (PUL) emerge as an interesting approach for content-based fake news detection using a smaller set of labeled data than traditional machine learning techniques. In particular, network-based approaches are adequate for fake news detection since they allow incorporating information from different aspects of a publication to the problem modeling. In this paper, we propose a network-based approach based on Positive and Unlabeled Learning by Label Propagation (PU-LP), a one-class and transductive semi-supervised learning algorithm that performs classification by first identifying potential interest and non-interest documents into unlabeled data and then propagating labels to classify the remaining unlabeled documents. A label propagation approach is then employed to classify the remaining unlabeled documents. We assessed the performance of our proposal considering homogeneous (only documents) and heterogeneous (documents and terms) networks. Our comparative analysis considered four OCL algorithms extensively employed in One-Class text classification ( k -Means, k -Nearest Neighbors Density-based, One-Class Support Vector Machine, and Dense Autoencoder), and another traditional PUL algorithm (Rocchio Support Vector Machine). The algorithms were evaluated in three news collections, considering balanced and extremely unbalanced scenarios. We used Bag-of-Words and Doc2Vec models to transform news into structured data. Results indicated that PU-LP approaches are more stable and achieve better results than other PUL and OCL approaches in most scenarios, performing similarly to semi-supervised binary algorithms. Also, the inclusion of terms in the news network activate better results, especially when news are distributed in the feature space considering veracity and subject. News representation using the Doc2Vec achieved better results than the Bag-of-Words model for both algorithms based on vector-space model and document similarity network. Mariana Caravanti de Souza, Bruno M. Nogueira 0001, Rafael Geraldeli Rossi, Ricardo M. Marcacini, Brucce Neves dos Santos, Solange Oliveira Rezende |
Mach. Learn. | 6 |
| 2020 | Dropout through Extended Association Rule Netwoks: A Complementary View
Maicon Dall'Agnol, Leandro Rondado de Souza, Renan de Padua, Veronica Oliveira de Carvalho, Solange Oliveira Rezende |
CSEDU (1) | 5 |
| 2020 | A context-aware recommender method based on text and opinion miningabstractAbstract A recommender system is an information filtering technology that can be used to recommend items that may be of interest to users. Additionally, there are the context‐aware recommender systems that consider contextual information to generate the recommendations. Reviews can provide relevant information that can be used by recommender systems, including contextual and opinion information. In a previous work, we proposed a context‐aware recommendation method based on text mining (CARM‐TM). The method includes two techniques to extract context from reviews: CIET.5embed, a technique based on word embeddings; and RulesContext, a technique based on association rules. In this work, we have extended our previous method by including CEOM, a new technique which extracts context by using aspect‐based opinions. We call our extension of CARM‐TOM (context‐aware recommendation method based on text and opinion mining). To generate recommendations, our method makes use of the CAMF algorithm, a context‐aware recommender based on matrix factorization. To evaluate CARM‐TOM, we ran an extensive set of experiments in a dataset about restaurants, comparing CARM‐TOM against the MF algorithm, an uncontextual recommender system based on matrix factorization; and against a context extraction method proposed in literature. The empirical results strongly indicate that our method is able to improve a context‐aware recommender system. Camila Vaccari Sundermann, Renan de Padua, Vítor Rodrigues Tonon, Ricardo M. Marcacini, Marcos Aurélio Domingues, Solange Oliveira Rezende |
Expert Syst. J. Knowl. Eng. | 6 |
| 2020 | A two-stage regularization framework for heterogeneous event networksabstractEvent analysis from news and social networks is a promising way to understand complex social phenomena. Each event consists of different components, which indicate what happened, when, where, and the people and organizations involved. Heterogeneous networks are useful for modeling large event datasets, where we map different types of objects (e.g. events and their components), as well as the different relationships between objects. Such networks enable the identification of related events, in which users label some events in categories and then use the network's topological structure to find other events of interest. Although this process can be automated, there is a lack of machine learning methods to properly handle event classification from heterogeneous networks. In this paper, we present the framework named Heterogeneous Event Network Regularization in Two-stages (HENR2). The first stage of HENR2 aims to learn the importance level of each relationship between events and their components. In the second stage, the regularization process considers the importance levels of each relationship to propagate labels on the network. Thus, the classification process is improved by considering the domain characteristics of the event dataset, such as temporal seasonality and geographical distribution. In both stages, our approach also deals with noisy data through parameters that define the confidence level of labeled events during label propagation. Experimental results involving twelve event networks from different application domains show that our proposal outperforms existing regularization frameworks. Brucce Neves dos Santos, Rafael Geraldeli Rossi, Solange Oliveira Rezende, Ricardo M. Marcacini |
Pattern Recognit. Lett. | 3 |
| 2019 | Sentiment classification improvement using semantically enriched informationabstractThe emergence of new and challenging text mining applications is demanding the development of novel text processing and knowledge extraction techniques. One important challenge of text mining is the proper treatment of text meaning, which may be addressed by incorporating different types of information (e.g., syntactic or semantic) into the text representation model. Sentiment classification is one of the challenging text mining applications. It may be considered more complex than the traditional topic classification since, although sentiment words are important, they may not be enough to correctly classify the sentiment expressed in a document. In this work, we propose a novel and straightforward method to improve sentiment classification performance, with the use of semantically enriched information derived from domain expressions. We also propose a superior scheme for generating these expressions. We conducted an experimental evaluation applying different classification algorithms to three datasets composed by reviews of different products and services. The results indicate that the proposed method enables the improvement of classification accuracy when dealing with reviews of a narrow domain. Ricardo B. Scheicher, Roberta Akemi Sinoara, Jonas C. Felinto, Solange Oliveira Rezende |
DocEng | 4 |
| 2019 | Knowledge-enhanced document embeddings for text classificationabstractAccurate semantic representation models are essential in text mining applications. For a successful application of the text mining process, the text representation adopted must keep the interesting patterns to be discovered. Although competitive results for automatic text classification may be achieved with traditional bag of words, such representation model cannot provide satisfactory classification performances on hard settings where richer text representations are required. In this paper, we present an approach to represent document collections based on embedded representations of words and word senses. We bring together the power of word sense disambiguation and the semantic richness of word- and word-sense embedded vectors to construct embedded representations of document collections. Our approach results in semantically enhanced and low-dimensional representations. We overcome the lack of interpretability of embedded vectors, which is a drawback of this kind of representation, with the use of word sense embedded vectors. Moreover, the experimental evaluation indicates that the use of the proposed representations provides stable classifiers with strong quantitative results, especially in semantically-complex classification scenarios Roberta Akemi Sinoara, José Camacho-Collados, Rafael Geraldeli Rossi, Roberto Navigli, Solange Oliveira Rezende |
Knowl. Based Syst. | 5 |
| 2018 | Asymmetric Objective Measures Applied to Filter Association Rules NetworksabstractIn this paper, the Filtered-Association Rules Network (Filtered-ARN) is presented to structure, prune, and analyze a set of association rules to construct candidate hypotheses. The Filtered-ARN algorithm selects association rules with the use of asymmetric objective measures, Added Value and Gain, then builds a network allowing more exploration information. The Filtered-ARN was validated using three datasets: Lenses and Soybean Large, both available online for a text and a real dataset with data on organic fertilization (Green Manure). The results were validated by comparing the Filtered-ARN with the conventional ARN and also comparing the results with the decision tree. The approach presented promising results, showing its ability to explain a set of objective items and the aid to build more consolidated hypotheses by guaranteeing statistical dependence with the use of objective measures. Dario Brito Calcada, Renan de Padua, Solange Oliveira Rezende |
CLEI | 3 |
| 2018 | Agribusiness Time Series Forecasting using Perceptually Important EventsabstractModern agribusiness management incorporates instruments for risk management with the objective of mitigating uncertainties to the producer. In this context, the producer (risk averse) transfer the risk of price oscillation to companies or individuals that operate in the futures market and who expect to receive a payment (risk premium) for assuming such risk. Defining the adequate strategies for risk management depends on the knowledge about the problem to determine prices ranges in the future. Recent studies demonstrate that time series forecasting can be significantly improved by considering additional information about the problem. In particular, besides the historical time series, textual knowledge extracted from the news portals, social networking and other public data sources available in the web may also be used. This paper presents an approach for agribusiness time series forecasting that allows incorporating external knowledge in the form of events extracted from news about agribusiness, without the need to previously label textual information. In this case, periods of significant uptrends and downtrends of time series are automatically identified - known in the literature as perceptually important points (PIP). We extend the concept of PIP to news events, where similar events published with a certain regularity in periods of uptrends and downtrends are selected as perceptually important events to improve time series forecasting models. An experimental evaluation based on price prediction on ten corn futures contracts (derivatives) provides evidence that the proposed approach is promising. Lusas S. Rodrigues, Solange Oliveira Rezende, Maria Fernanda Moura, Ricardo M. Marcacini |
CLEI | 2 |
| 2018 | A Semantic Approach to Uncovering Implicit Relationships in Textual DatabasesabstractThe discovery of knowledge in textual databases is an approach that basically seeks for implicit relationships between different concepts in different documents written in natural language, in order to identify new useful knowledge. To assist in this process, this approach can count on the help of Text Mining techniques. Despite all the progress made, researchers in this area must still deal with a large number of false relationships generated by most of the available processes. A semantic approach that supports the understanding of the relationships may bridge this gap. Thus, the objective of this work is to support the identification of implicit relationships between concepts present in different texts, considering the verbal semantics of relationships. To this end, analysis based on association rules were used together with metrics from complex networks and a verbal semantics approach. Through a case study, a set of texts from alternative medicine was selected and the different extractions showed that the proposed approach facilitates the identification of implicit causal relationships. Dildre Georgiana Vasques, Paulo Sérgio Martins, Solange Oliveira Rezende |
CLEI | 3 |
| 2018 | An Analysis on Community Detection and Clustering Algorithms on the Post-Processing of Association RulesabstractAssociation rules are widely used to extract patterns from a given database. The association rules are capable of finding correlations among items, making it possible for the user to learn which items are present in the transactions and which of them have a significant correlation. One of the major problems with association rules is that the number of extracted rules usually exceeds the number of transactions present in the database, also surpassing the user's capability to explore the obtained knowledge. To overcome this problem, the post-processing phase was proposed with the objective of directing the user to the rules that potentially have the most interesting knowledge. One of the used approaches is to divide the association rules into groups (or clusters), so that rules behave similarly are on the same group, facilitating the rule set understanding. In the literature, there are some works that uses clustering algorithms to split the rules while some other works use community detection algorithms. As both approaches obtain groups of association rules, but using different premises, different results can be obtained. No study has been done on the differences among clustering and community detection algorithms, which makes the selection of the algorithm hard, once their behavior is not well known in the association rule post-processing phase. This paper presents an analysis on both approaches, aiming to find the differences and the similarities among them, making it easier to select an approach by knowing its behavior. Renan de Padua, Laís Pessine do Carmo, Solange Oliveira Rezende, Veronica Oliveira de Carvalho |
IJCNN | 3 |
| 2018 | Transforming Geo-Referenced Data in Contextual Information for Context-Aware Recommender SystemsabstractA recommender system can be defined as an information filtering technology which can be used to output a ranking of items (e.g. products, places, etc) that are likely to be of interest to a user. Context-aware recommender systems makes recommendations by incorporating contextual information into the recommendation process. However, there is a lack of automatic methods to obtain contextual information for such systems. In this work, we have proposed to apply clustering techniques to transform geo-referenced data (i.e. latitude and longitude) in contextual information (i.e. regions) to feed the contextual systems. We have evaluated our proposal in the Yelp dataset, which showed evidences that our contextual information can provide better recommendations. Igor André Pegoraro Santana, Abner Suniga, Juliano Donini, Camila Vaccari Sundermann, Solange Oliveira Rezende, Marcos Aurélio Domingues |
WI | 5 |
| 2018 | Exploration of Word Embedding Model to Improve Context-Aware Recommender SystemsabstractRecommender systems aim to assist users by recommending items that may be of interest to them. Traditionally, these systems use only user and item information. Over time, new information is being used, such as contextual information, which has improved the accuracy of the generated recommendations. In this work, we propose a context-aware recommender method that extracts contextual information from textual reviews using a word embedding based model. In addition, we propose two ways of considering textual contexts in recommender systems, the "Context of Reviews" and the "Context of Items". We evaluated our proposal by using the Yelp dataset (RecSysChallenge 2013); three baselines; and four context-aware recommender systems. In general, our proposal seems to be superior to the three baselines, mainly considering the "Context of Items", and the results were promising, allowing some lines of future work. Camila Vaccari Sundermann, João Antunes, Marcos Aurélio Domingues, Solange Oliveira Rezende |
WI | 4 |
| 2018 | Cross-domain aspect extraction for sentiment analysis: A transductive learning approach
Ricardo M. Marcacini, Rafael Geraldeli Rossi, Ivone Penque Matsuno, Solange Oliveira Rezende |
Decis. Support Syst. | 4 |
| 2018 | Latent association rule cluster based model to extract topics for classification and recommendation applications
Fabiano Fernandes dos Santos, Marcos Aurélio Domingues, Camila Vaccari Sundermann, Veronica Oliveira de Carvalho, Maria Fernanda Moura, Solange Oliveira Rezende |
Expert Syst. Appl. | 6 |
| 2017 | Unsupervised active learning techniques for labeling training sets: An experimental evaluation on sequential dataabstractMany real-world applications, such as those related to sensors, allow collecting large amounts of inexpensive unlabeled sequential data. However, the use of supervised machine learning methods is frequently hindered by the high costs involved in gath Vinícius M. A. de Souza, Rafael Geraldeli Rossi, Gustavo Batista, Solange Oliveira Rezende |
Intell. Data Anal. | 4 |
| 2017 | Using bipartite heterogeneous networks to speed up inductive semi-supervised learning and improve automatic text categorizationabstractDue to the volume of texts available in digital form, the organization, management and knowledge extraction are laborious and frequently impossible to be handled. To automatically cope with these tasks, usually classification models are generated through supervised learning techniques. Unfortunately, this type of learning usually demands a huge human effort to label large volume of texts to build accurate classification models. Since collecting unlabeled texts is easy and inexpensive in several domains, the generation of classification models through inductive semi-supervised learning has been highlighted in recent years. Inductive semi-supervised learning allows to build a classification model using labeled and unlabeled texts. In this scenario, the goal is to augment the set of labeled documents with unlabeled documents to better discriminate class patterns. Hence, fewer texts must be previously labeled. However, semi-supervised learning algorithms that consider texts represented in a vector space model usually obtain unsatisfactory classification performances and are surpassed by semi-supervised learning algorithms that consider texts represented in a network. Nevertheless, despite the classification performances, effective approaches based on networks are generated through the similarities among documents and the classification of a new document are also based on the computation of similarities. This implies to set parameters and compute similarities to both generation the networks and classification of new documents. This approach is not feasible to generate fast responses and consequently to classify a huge volume of texts. In this article, we propose an approach to induce a classification model through semi-supervised learning considering text collections represented by bipartite heterogeneous networks. Bipartite networks are easily and quickly generated, leading to classification performance equivalent or better than other approaches based on network or vector space model and allows a fast classification of new documents. The results presented in this article demonstrate that the proposed approach is able to (i) speed up semi-supervised learning, (ii) speed up the classification of new documents and (iii) surpass classification performance of other existing inductive semi-supervised learning techniques. Rafael Geraldeli Rossi, Alneu de Andrade Lopes, Solange Oliveira Rezende |
Knowl. Based Syst. | 3 |
| 2016 | Flexible document organization by mixing fuzzy and possibilistic clustering algorithmsabstractA powerful and flexible organization of documents can be obtained by mixing fuzzy and possibilistic clustering. In such organization, documents can belong to more than one cluster simultaneously with different compatibility degrees. Clusters represent topics, which are identified by one or more descriptors extracted by a proposed method. In this manuscript, we investigated whether or not the descriptors extracted after applying possibilistic fuzzy clustering improve the flexible organization of documents. Experiments were carried out on real-world document collections and we evaluated the ability of descriptors to capture the essential information in every dataset. Results have shown the effectiveness of extracting possibilistic fuzzy cluster descriptors, improving the flexible organization of documents. Nilton V. Carvalho, Solange Oliveira Rezende, Heloisa A. Camargo, Tatiane N. Rios |
FUZZ-IEEE | 2 |
| 2016 | Semantic role-based representations in text classificationabstractAlthough good results for automatic text classification can be achieved with the use of bag-of-words representation, this model is not suitable for all classification problems and richer text representations can be required. In this paper, we proposed two text representation models based on semantic role labels and analyzed them in text classification scenarios. We also evaluated the combination of bag-of-words with a semantic representation considering ensemble multi-view strategies. We explored different classification problems for two text collections and pointed out situations that require more than a bag-of-words. The experimental evaluation indicates that the combination of bag-of-words and a text representation based on semantic role labels can improve text classification accuracies. Roberta Akemi Sinoara, Rafael Geraldeli Rossi, Solange Oliveira Rezende |
ICPR | 3 |
| 2016 | Solving the Problem of Selecting Suitable Objective Measures by Clustering Association Rules Through the Measures Themselves
Veronica Oliveira de Carvalho, Renan de Padua, Solange Oliveira Rezende |
SOFSEM | 3 |
| 2016 | Post-processing Association Rules: A Network Based Label Propagation Approach
Renan de Padua, Veronica Oliveira de Carvalho, Solange Oliveira Rezende |
SOFSEM | 3 |
| 2016 | Privileged contextual information for context-aware recommender systemsabstractA recommender system is used in various fields to recommend items of interest to the users. Most recommender approaches focus only on the users and items to make the recommendations. However, in many applications, it is also important to incorporate contextual information into the recommendation process. Although the use of contextual information has received great focus in recent years, there is a lack of automatic methods to obtain such information for context-aware recommender systems. Some works address this problem by proposing supervised methods, which require greater human effort and whose results are not so satisfactory. In this scenario, we propose an unsupervised method to extract contextual information from web page content. Our method builds topic hierarchies from page textual content considering, besides the traditional bag-of-words, valuable information of texts as named entities and domain terms (privileged information). The topics extracted from the hierarchies are used as contextual information in context-aware recommender systems. We conducted experiments by using two data sets and two baselines: the first baseline is a recommendation system that does not use contextual information and the second baseline is a method proposed in literature to extract contextual information. The results are, in general, very good and present significant gains. In conclusion, our method has advantages and innovations:(i) it is unsupervised; (ii) it considers the context of the item (Web page), instead of the context of the user as in most of the few existing methods, which is an innovation; (iii) it uses privileged information in addition to the existing technical information from pages; and (iv) it presented good and promising empirical results. This work represents an advance in the state-of-the-art in context extraction, which means an important contribution to context-aware recommender systems, a kind of specialized and intelligent system. Camila Vaccari Sundermann, Marcos Aurélio Domingues, Merley da Silva Conrado, Solange Oliveira Rezende |
Expert Syst. Appl. | 4 |
| 2016 | Optimization and label propagation in bipartite heterogeneous networks to improve transductive classification of textsabstractTransductive classification is a useful way to classify texts when labeled training examples are insufficient. Several algorithms to perform transductive classification considering text collections represented in a vector space model have been proposed. However, the use of these algorithms is unfeasible in practical applications due to the independence assumption among instances or terms and the drawbacks of these algorithms. Network-based algorithms come up to avoid the drawbacks of the algorithms based on vector space model and to improve transductive classification. Networks are mostly used for label propagation, in which some labeled objects propagate their labels to other objects through the network connections. Bipartite networks are useful to represent text collections as networks and perform label propagation. The generation of this type of network avoids requirements such as collections with hyperlinks or citations, computation of similarities among all texts in the collection, as well as the setup of a number of parameters. In a bipartite heterogeneous network, objects correspond to documents and terms, and the connections are given by the occurrences of terms in documents. The label propagation is performed from documents to terms and then from terms to documents iteratively. Nevertheless, instead of using terms just as means of label propagation, in this article we propose the use of the bipartite network structure to define the relevance scores of terms for classes through an optimization process and then propagate these relevance scores to define labels for unlabeled documents. The new document labels are used to redefine the relevance scores of terms which consequently redefine the labels of unlabeled documents in an iterative process. We demonstrated that the proposed approach surpasses the algorithms for transductive classification based on vector space model or networks. Moreover, we demonstrated that the proposed algorithm effectively makes use of unlabeled documents to improve classification and it is faster than other transductive algorithms. Rafael Geraldeli Rossi, Alneu de Andrade Lopes, Solange Oliveira Rezende |
Inf. Process. Manag. | 3 |
| 2016 | Mining unstructured content for recommender systems: an ensemble approach
Marcelo G. Manzato, Marcos Aurélio Domingues, Arthur F. Da Costa, Camila Vaccari Sundermann, Rafael Martins D'Addio, Merley da Silva Conrado, Solange Oliveira Rezende, Maria da Graça Campos Pimentel |
Inf. Retr. J. | 7 |
| 2015 | Term Network Approach for Transductive Classification
Rafael Geraldeli Rossi, Solange Oliveira Rezende, Alneu de Andrade Lopes |
CICLing (2) | 2 |
| 2015 | Flexible document organization: Comparing fuzzy and possibilistic approachesabstractSystem flexibility means the ability of a system to manage imprecise and/or uncertain information. A lot of commercially available Information Retrieval Systems (IRS) address this issue at the level of query formulation. Another way to make the flexibility of an IRS possible is by means of the flexible organization of documents. Such organization can be carried out using clustering algorithms by which documents can be automatically organized in multiple clusters simultaneously. Fuzzy and possibilistic clustering algorithms are examples of methods by which documents can belong to more than one cluster simultaneously with different membership degrees. The interpretation of these membership degrees can be used to quantify the compatibility of a document with a particular topic. The topics are represented by clusters and the clusters are identified by one or more descriptors extracted by a proposed method. We aim to investigate if the performance of each clustering algorithm can affect the extraction of meaningful overlapping cluster descriptors. Experiments were carried using well-known collections of documents and the predictive power of the descriptors extracted from both fuzzy and possibilistic document clustering was evaluated. The results prove that descriptors extracted after both fuzzy and possibilistic clustering are effective and can improve the flexible organization of documents. Tatiane N. Rios, Solange Oliveira Rezende, Heloisa A. Camargo |
FUZZ-IEEE | 2 |
| 2015 | Interactive textual feature selection for consensus clustering
Geraldo N. Correa, Ricardo M. Marcacini, Eduardo R. Hruschka, Solange Oliveira Rezende |
Pattern Recognit. Lett. | 4 |
| 2014 | Semi-Supervised Learning to Support the Exploration of Association Rules
Veronica Oliveira de Carvalho, Renan de Padua, Solange Oliveira Rezende |
DaWaK | 3 |
| 2014 | Post-Processing Association Rules Using Networks and Transductive LearningabstractAssociation is widely used to find relations among items in a given database. However, finding the interesting patterns is a challenging task due to the large number of rules that are generated. Traditionally, this task is done by post-processing approaches that explore and direct the user to the interesting rules of the domain. Some of these approaches use the user's knowledge to guide the exploration according to what is defined (thought) as interesting by the user. However, this definition is done before the process starts. Therefore, the user must know what may be and what may not be interesting to him/her. This work proposes a general association rule post-processing approach that extracts the user's knowledge during the post-processing phase. That way, the user does not need to have a prior knowledge in the database. For that, the proposed approach models the association rules in a network, uses its measures to suggest rules to be classified by the user and, then, propagates these classifications to the entire network using transductive learning algorithms. Therefore, this approach treats the post-processing problem as a classification task. Experiments were carried out to demonstrate that the proposed approach reduces the number of rules to be explored by the user and directs him/her to the potentially interesting rules of the domain. Renan de Padua, Solange Oliveira Rezende, Veronica Oliveira de Carvalho |
ICMLA | 2 |
| 2014 | Using Contextual Information from Topic Hierarchies to Improve Context-Aware Recommender SystemsabstractUnlike the traditional recommender systems, that make recommendations only by using the relation between user and item, a context-aware recommender system makes recommendations by incorporating available contextual information into the recommendation process as explicit additional categories of data to improve the recommendation process. In this paper, we propose to use contextual information from topic hierarchies to improve the accuracy of context-aware recommender systems. Additionally, we also propose two context-aware recommender algorithms for item recommendation. These are extensions from algorithms proposed in literature for rating prediction. The empirical results demonstrate that by using topic hierarchies our technique can provide better recommendations. Marcos Aurélio Domingues, Marcelo G. Manzato, Ricardo M. Marcacini, Camila Vaccari Sundermann, Solange Oliveira Rezende |
ICPR | 5 |
| 2014 | Improving Personalized Ranking in Recommender Systems with Topic Hierarchies and Implicit FeedbackabstractThe knowledge of semantic information about the content and user's preferences is an important issue to improve recommender systems. However, the extraction of such meaningful metadata needs an intense and time-consuming human effort, which is impractical specially with large databases. In this paper, we mitigate this problem by proposing a recommendation model based on latent factors and implicit feedback which uses an unsupervised topic hierarchy constructor algorithm to organize and collect metadata at different granularities from unstructured textual content. We provide an empirical evaluation using a dataset of web pages written in Portuguese language, and the results show that personalized ranking with better quality can be generated using the extracted topics at medium granularity. Marcelo G. Manzato, Marcos Aurélio Domingues, Ricardo M. Marcacini, Solange Oliveira Rezende |
ICPR | 4 |
| 2014 | Privileged Information for Hierarchical Document Clustering: A Metric Learning ApproachabstractTraditional hierarchical text clustering methods assume that the documents are represented only by "technical information", i.e., keywords, phrases, expressions and named entities that can be directly extracted from the texts. However, in many scenarios there is an additional and valuable information about the documents which is usually disregarded during the clustering task, such as user-validated tags, annotations and comments from experts, dictionaries and domain ontologies. Recently, Vapnik introduced a new learning paradigm, called LUPI - Learning Using Privileged Information, which allows the incorporation of this additional (privileged) information in a supervised learning setting. We investigated the incorporation of privileged information in unsupervised setting. The key idea in our proposed approach is to extract important relationships among documents represented in the privileged information dimensional space to learn a more accurate metric for text clustering in the technical information space. A thorough experimental evaluation indicates that the incorporation of privileged information through metric learning significantly improves the hierarchical clustering accuracy. Ricardo M. Marcacini, Marcos Aurélio Domingues, Eduardo R. Hruschka, Solange Oliveira Rezende |
ICPR | 4 |
| 2014 | Named entities as privileged information for hierarchical text clusteringabstractText clustering is a text mining task which is often used to aid the organization, knowledge extraction, and exploratory search of text collections. Nowadays, the automatic text clustering becomes essential as the volume and variety of digital text documents increase, either in social networks and the Web or inside organizations. This paper explores the use of named entities as privileged information in a hierarchical clustering process, so as to improve clusters quality and interpretation. We carried out an experimental evaluation on three text collections (one written in Portuguese and two written in English) and the results show that named entities can be applied as privileged information to power clustering solution in dynamic text collection scenarios. Roberta Akemi Sinoara, Camila Vaccari Sundermann, Ricardo M. Marcacini, Marcos Aurélio Domingues, Solange Oliveira Rezende |
IDEAS | 5 |
| 2014 | Inductive Model Generation for Text Classification Using a Bipartite Heterogeneous Network
Rafael Geraldeli Rossi, Alneu de Andrade Lopes, Thiago de Paulo Faleiros, Solange Oliveira Rezende |
J. Comput. Sci. Technol. | 4 |
| 2013 | Metrics to Support the Evaluation of Association Rule Clustering
Veronica Oliveira de Carvalho, Fabiano Fernandes dos Santos, Solange Oliveira Rezende |
DaWaK | 3 |
| 2013 | Incremental hierarchical text clustering with privileged informationabstractIn many text clustering tasks, there is some valuable knowledge about the problem domain, in addition to the original textual data involved in the clustering process. Traditional text clustering methods are unable to incorporate such additional (privileged) information into data clustering. Recently, a new paradigm called LUPI - Learning Using Privileged Information - was proposed by Vapnik to incorporate privileged information in classification tasks. In this paper, we extend the LUPI paradigm to deal with text clustering tasks. In particular, we show that the LUPI paradigm is potentially promising for incremental hierarchical text clustering, being very useful for organizing large textual databases. In our method, the privileged information about the text documents is applied to refine an initial clustering model by means of consensus clustering. The initial model is used for incremental clustering of the remaining text documents. We carried out an experimental evaluation on two benchmark text collections and the results showed that our method significantly improves the clustering accuracy when compared to a traditional hierarchical clustering method. Ricardo M. Marcacini, Solange Oliveira Rezende |
ACM Symposium on Document Engineering | 2 |
| 2013 | A Machine Learning Approach to Automatic Term Extraction using a Rich Feature Set
Merley da Silva Conrado, Thiago A. S. Pardo, Solange Oliveira Rezende |
HLT-NAACL | 3 |
| 2013 | Influence of Graph Construction on Semi-supervised Learning
Celso André R. de Sousa, Solange Oliveira Rezende, Gustavo Batista |
ECML/PKDD (3) | 2 |
| 2012 | HCAC: Semi-supervised Hierarchical Clustering Using Confidence-Based Active Learning
Bruno M. Nogueira 0001, Alípio Mário Jorge, Solange Oliveira Rezende |
Discovery Science | 3 |
| 2012 | Evaluation of Normalization Techniques in Text Classification for Portuguese
Merley da Silva Conrado, Víctor Antonio Laguna Gutiérrez, Solange Oliveira Rezende |
ICCSA (3) | 3 |
| 2012 | Inductive Model Generation for Text Categorization Using a Bipartite Heterogeneous NetworkabstractUsually, algorithms for categorization of numeric data have been applied for text categorization after a preprocessing phase which assigns weights for textual terms deemed as attributes. However, due to characteristics of textual data, some algorithms for data categorization are not efficient for text categorization. Characteristics of textual data such as sparsity and high dimensionality sometimes impair the quality of general purpose classifiers. Here, we propose a text classifier based on a bipartite heterogeneous network used to represent textual document collections. Such algorithm induces a classification model assigning weights to objects that represents terms of the textual document collection. The induced weights correspond to the influence of the terms in the classification of documents they appear. The least-mean-square algorithm is used in the inductive process. Empirical evaluation using a large amount of textual document collections shows that the proposed IMBHN algorithm produces significantly better results than the k-NN, C4.5, SVM and Naïve Bayes algorithms. Rafael Geraldeli Rossi, Thiago de Paulo Faleiros, Alneu de Andrade Lopes, Solange Oliveira Rezende |
ICDM | 4 |
| 2012 | An active learning approach to frequent itemset-based text clustering
Ricardo M. Marcacini, Geraldo N. Correa, Solange Oliveira Rezende |
ICPR | 3 |
| 2012 | Fuzzy cluster descriptors improve flexible organization of documentsabstractSystem flexibility means the ability of a system to manage imprecise and/or uncertain information. There are two ways to address the Information Retrieval Systems (IRS) flexibility: through methods that improve the query formulation and through methods that improve the document organization. Since the query formulation has obtained more attention in retrieval process, we aim to investigate the flexibility in document organization. When a document organization is carried out using fuzzy clustering, the documents can belong to more than one cluster simultaneously with different membership degrees, allowing the management of imprecise and/or uncertain information in the collection organization. Clusters represent topics and are identified by one or more descriptors. In this work we use an unsupervised method to extract cluster descriptors for a specific database and investigate whether the quality of the fuzzy cluster descriptors improves the flexible organization of documents. Tatiane N. Rios, Solange Oliveira Rezende, Heloisa A. Camargo |
ISDA | 2 |
| 2011 | Building a topic hierarchy using the bag-of-related-words representationabstractA simple and intuitive way to organize a huge document collection is by a topic hierarchy. Generally two steps are carried out to build a topic hierarchy automatically: 1) hierarchical document clustering and 2) cluster labeling. For both steps, a good textual document representation is essential. The bag-of-words is the common way to represent text collections. In this representation, each document is represented by a vector where each word in the document collection represents a dimension (feature). This approach has well known problems as the high dimensionality and sparsity of data. Besides, most of the concepts are composed by more than one word, as "document engineering" or "text mining". In this paper an approach called bag-of-related-words is proposed to generate features compounded by a set of related words with a dimensionality smaller than the bag-of-words. The features are extracted from each textual document of a collection using association rules. Different ways to map the document into transactions in order to allow the extraction of association rules and interest measures to prune the number of features are analyzed. To evaluate how much the proposed approach can aid the topic hierarchy building, we carried out an objective evaluation for the clustering structure, and a subjective evaluation for topic hierarchies. All the results were compared with the bag-of-words. The obtained results demonstrated that the proposed representation is better than the bag-of-words for the topic hierarchy building. Rafael Geraldeli Rossi, Solange Oliveira Rezende |
ACM Symposium on Document Engineering | 2 |
| 2011 | Fuzzy cluster descriptor extraction for flexible organization of documentsabstractThe text mining process and its set of techniques have been widely used in order to look for new knowledge in textual documents, which can be recovered by information retrieval systems. There is a variety of methods developed to automatically organize documents based on the knowledge extracted from their content. The management of imprecision and uncertainty is very important to improve these methods. Therefore, this work proposes a method to manage imprecision and uncertainty in document organization, by clustering the documents in fuzzy clusters and extracting cluster descriptors. The proposed method was evaluated and the obtained results showed that this is a promising approach to deal with the problem of imprecision and uncertainty when organizing textual documents. Tatiane N. Rios, Solange Oliveira Rezende, Heloisa A. Camargo |
HIS | 2 |
| 2010 | On the use of fuzzy rules to text document classificationabstractThis work presents the integration of a fuzzy method and text mining to obtain an approach that enables the text documents classification to be closer to the user needs. The aim of this work is to develop a mechanism to reduce the high dimensionality of the attribute-value matrix obtained from the documents and, with this, to manage the imprecision and uncertainty using fuzzy rules to classify text documents. Some experiments have been run using different domains in order to validate the proposed approach and to compare the results with the ones obtained with the Ibk, J48, Naive Bayes and OneR classification methods. The advantages of the method, the experiments and the results obtained are discussed. Tatiane N. Rios, Solange Oliveira Rezende, Heloisa A. Camargo |
HIS | 2 |
| 2010 | Incremental Construction of Topic Hierarchies using Hierarchical Term Clustering
Ricardo M. Marcacini, Solange Oliveira Rezende |
SEKE | 2 |
| 2010 | MMWA-ae: boosting knowledge from Multimodal Interface Design, Reuse and Usability Evaluation
Américo Talarico Neto, Renata Pontin de Mattos Fortes, Rafael Geraldeli Rossi, Solange Oliveira Rezende |
SEKE | 4 |
| 2007 | An Analytical Evaluation of Objective Measures Behavior for Generalized Association RulesabstractThe association rule mining task identifies all the intrinsic associations among the items contained in data and leads to only specialized knowledge. To overcome this problem the generalized association rules appeared. This type of rule associates not only the items contained in data, but also some items encoded into a given taxonomy. Therefore, the techniques used to obtain generalized association rules are very useful since they provide a more general view of the domain. However, a problem found when using these techniques is how to identify the most useful rules to avoid overload the user with a huge amount of patterns. Nowadays, the researches use objective evaluation measures to evaluate and select the most interesting knowledge to the user. Despite the fact these measures have been studied by many researches to evaluate many types of rules (for example, classification and traditional association rules), it is important to study these measures in the context of generalized rules. Thus, this paper presents an analytical evaluation to understand the behavior of some objective measures when applied in a set of generalized rules. Many relations were obtained to express the behavior of these measures, what represents a meaningful contribution to the post-processing data mining area Veronica Oliveira de Carvalho, Solange Oliveira Rezende, Mário de Castro |
CIDM | 2 |
| 2007 | An Itemset-Driven Cluster-Oriented Approach to Extract Compact and Meaningful Sets of Association RulesabstractExtracting association rules from large datasets typically results in a huge amount of rules. An approach to tackle this problem is to filter the resulting rule set, which reduces the rules, at the cost of also eliminating potentially interesting ones. In exploring a new dataset in search of relevant associations, it may be more useful for miners to have an overview of the space of rules obtainable from the dataset, rather than getting an arbitrary set satisfying high values for given interest measures. We describe a rule extraction approach that favors rule diversity, allowing miners to gain an overview of the rule space while reducing semantic redundancy within the rule set. This approach adopts an itemset-driven rule generation coupled with a cluster-based filtering process. The set of rules so obtained provides a starting point for a user-driven exploration of it. Claudio Haruo Yamamoto, Maria Cristina Ferreira de Oliveira, Magaly Lika Fujimoto, Solange Oliveira Rezende |
ICMLA | 4 |
| 2004 | Combining Intelligent Techniques for Sensor Fusion
Katti Faceli, André C. P. L. F. de Carvalho, Solange Oliveira Rezende |
Appl. Intell. | 3 |
| 1999 | Applying neural networks to determine vibration parameters in a turbineabstractVibration signals analysis is considered as an appropriate diagnosis method for detecting faults. Several techniques have been used for detecting vibration signals. In this work, artificial neural networks (ANN) were used to predict vibration signals using process parameters measured in a turbine. The ANN models can be viewed as "black-boxes". One way to improve their comprehensibility is to use a symbolic model. As a first step in this direction, a hybrid rule-based regression model was also tested. Chandler W. Caulkins, Robson B. T. Oliveira, André C. P. L. F. de Carvalho, Solange Oliveira Rezende, Maria Carolina Monard |
IJCNN | 4 |