VLDB 2026 Research / reviewers in the wild / expert
Ricardo M. Marcacini
dblp:69/8767 · also Ricardo Marcondes Marcacini
· DBLP profile ↗
39ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-2309-3487ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 5 first-author · 16 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EPHG-CR: embedding propagation for heterogeneous graphs with class refinementabstractAbstract Heterogeneous graphs can represent real-world problems in a way close to reality, supporting diverse types of vertices and edges. However, their inherent heterogeneity poses challenges in interpreting problem semantics. To address this, heterogeneous graph embedding, aiming to map graph elements to low-dimensional vectors, simplifies subsequent machine learning analysis. This approach has gained prominence in machine learning, fueling classification, recommendation, and similarity search applications. Embedding diverse data is essential for efficient data processing. Incorporating language models, like BERT, into heterogeneous graphs enhances semantic context capture, which is particularly useful when one vertex type represents text. Language models stand out in contextual representation, enriching graph vertex embeddings for various tasks. This paper proposes a novel approach to enhancing heterogeneous graph embeddings by combining language models and task class data. Our approach increases vector quality, accounting for graph structure, semantic textual information, and task labels. We compared our proposal with a language model in the aspect-based sentiment analysis task, demonstrating competitive results and, in some cases, a slight superiority. Furthermore, we explore applications of embeddings from auxiliary vertices in another task, highlighting another advantage of the approach over the language model. Brucce Neves dos Santos, Ricardo M. Marcacini, Alípio Mário Jorge, Ricardo Campos 0001, Solange Oliveira Rezende |
Appl. Intell. | 2 |
| 2026 | Explainable visual emotion recognition via modular reasoningabstractAffective computing still faces significant challenges in mapping complex visual features to emotional states with high interpretability. Although Multimodal Large Language Models (MLLMs) offer strong generalization, their high computational cost for fine-tuning constitutes a limitation. Moreover, MLLMs often function as “black boxes”, resulting in a lack of transparency in how they prioritize visual cues, especially in zero-shot settings. This work investigates these limitations by proposing a modular reasoning strategy to enhance interpretability and enable more disentangled affective reasoning under noisy cues and resource-constrained conditions. We introduce the Chain-of-Responsibility (CoR), a multi-agent framework that decomposes the affective inference process into specialized agents (Facial, Body, and Contextual). Unlike traditional end-to-end models or single-prompt reasoning, CoR employs modular decomposition designed to expose intermediate modality-specific analyses. A Synthesizer agent finally integrates these analyses using an explicit priority policy enforced through prompt design. This design enhances transparency and auditability of the decision process, allowing researchers to inspect how facial, bodily, and contextual cues contribute to the final prediction, while maintaining competitive performance in zero-shot settings. Magaly Lika Fujimoto, Ricardo M. Marcacini, Solange Oliveira Rezende |
Pattern Recognit. | 2 |
| 2025 | MuPe Life Stories Dataset: Spontaneous Speech in Brazilian Portuguese with a Case Study Evaluation on ASR Bias against Speakers Groups and Topic ModelingabstractRecently, several public datasets for automatic speech recognition (ASR) in Brazilian Portuguese (BP) have been released, improving ASR systems performance. However, these datasets lack diversity in terms of age groups, regional accents, and education levels. In this paper, we present a new publicly available dataset consisting of 289 life story interviews (365 hours), featuring a broad range of speakers varying in age, education, and regional accents. First, we demonstrated the presence of bias in current BP ASR models concerning education levels and age groups. Second, we showed that our dataset helps mitigate these biases. Additionally, an ASR model trained on our dataset performed better during evaluation on a diverse test set. Finally, the ASR model trained with our dataset was extrinsically evaluated through a topic modeling task that utilized the automatically transcribed output. Sidney Evaldo Leal, Arnaldo Cândido Jr., Ricardo M. Marcacini, Edresson Casanova, Odilon Gonçalves, Anderson da Silva Soares, Rodrigo Lima 0004, Lucas Gris, Sandra M. Aluísio |
COLING | 3 |
| 2025 | Advancing Multi-step Mathematical Reasoning in Large Language Models Through Multi-layered Self-reflection with Auto-prompting
André de Souza Loureiro, Jorge Carlos Valverde-Rebaza, Julieta Noguez 0001, David Escarcega, Ricardo M. Marcacini |
ECML/PKDD (4) | 5 |
| 2025 | LLM-based approaches for automated vocabulary mapping between SIGTAP and OMOP CDM concepts
Vinícius João de Barros Vanzin, Dilvan de Abreu Moreira, Ricardo M. Marcacini |
Artif. Intell. Medicine | 3 |
| 2025 | One-class graph autoencoder: A new end-to-end, low-dimensional, and interpretable approach for node classification
Marcos P. S. Gôlo, José Gilberto Barbosa de Medeiros Júnior, Diego Furtado Silva, Ricardo M. Marcacini |
Inf. Sci. | 4 |
| 2025 | How do financial time series enhance the detection of news significance in market movements? A study using graph neural networks with heterogeneous representations
Ivan J. Reis Filho, Marcos P. S. Gôlo, Ricardo M. Marcacini, Solange Oliveira Rezende |
Neural Comput. Appl. | 3 |
| 2025 | Issue detection and prioritization based on mobile application reviews
Vitor Mesaque Alves de Lima, Jacson Rodrigues Barbosa, Ricardo M. Marcacini |
Softw. Qual. J. | 3 |
| 2024 | Monitoring Temporal Dynamics of Issues in Crowdsourced User Reviews and their Impact on Mobile App UpdatesabstractAnalyzing user feedback from app stores through opinion mining aims to support software engineering activities, specifically in software maintenance and evolution. It is essential to promptly detect emerging app issues and facilitate the software's ongoing development. Manual analysis is impractical due to the large volume of textual data, necessitating machine learning methods for automation. Current methods lack mechanisms for trend detection and monitoring temporal dynamics, considering the relationship between issues and app release dates. This paper presents a two-fold approach: (i) identifying app issues and (ii) monitoring their evolution through temporal dynamic modeling using time series, release dates, and alerts. We present the MApp-TIME (Monitoring App by Temporal dynamic of Issues for app Maintenance and Evolution) approach, a microservices architecture designed to detect and monitor the temporal dynamics of issues and app releases. The goal is to reduce the time between issue detection and resolution, facilitating better software maintenance and evolution. We analyzed 13 million reviews across 20 domains and the findings revealed that about 75% of app releases correspond with issue peaks in the analyzed time series. Monitoring the temporal dynamics of crowdsourced user reviews can allow us to detect and prioritize issues early, sianificantly mitigating their impact. Vitor Mesaque Alves de Lima, Jacson Rodrigues Barbosa, Ricardo M. Marcacini |
ICSME | 3 |
| 2024 | iRisk: A Scalable Microservice for Classifying Issue Risks Based on Crowdsourced App ReviewsabstractAnalyzing mobile app reviews is essential for identifying trends and issue patterns that affect user experience and app reputation in app stores. A risk matrix provides a straightforward, intuitive method to prioritize software maintenance actions to mitigate negative ratings. However, manually constructing a risk matrix is time-consuming, and stakeholders often struggle to understand the context of risks due to varied descriptions and the sheer volume of reviews. Therefore, machine learning-based methods are needed to extract risks and classify their priority effectively. While existing studies have automated risk matrix generation in software development, they have not explored app reviews or utilized Large Language Models (LLMs) in a scalable architecture. To address this gap, we present iRisk (scalable microservice for classifying issue Risks), a tool for generating a risk matrix based on crowdsourced app reviews using LLM. We present i-LLAMA, a fine-tuned version of LLaMA 3, optimized to detect and prioritize app-related issues using a risk analysis dataset of reviews categorized by severity and likelihood of occurrence. This dataset is also publicly available. Our contributions include the open-source resources to support the software maintenance and evolution industry, fine-tuning of LLaMA 3, and a scalable microservice architecture to handle large volumes of data. The iRisk can manage app issues and risks and provide an automated dashboard and visualizations for decision-making, monitoring, and risk mitigation. The tool is available on GitHub11https://github.com/vitormesaque/iRisk, and a presentation about the tool can be found in this video22https://irisk.mappidea.com. Vitor Mesaque Alves de Lima, Jacson Rodrigues Barbosa, Ricardo M. Marcacini |
ICSME | 3 |
| 2024 | DODFMiner: An automated tool for Named Entity Recognition from Official Gazettes
Gabriel M. C. Guimarães, Felipe X. B. da Silva, Andrei L. Queiroz, Ricardo M. Marcacini, Thiago de Paulo Faleiros, Vinicius Ruela Pereira Borges, Luís Paulo F. Garcia |
Neurocomputing | 4 |
| 2024 | Keywords attention for fake news detection using few positive labels
Mariana Caravanti de Souza, Marcos P. S. Gôlo, Alípio Mário Jorge, Evelin Amorim, Ricardo Campos 0001, Ricardo M. Marcacini, Solange Oliveira Rezende |
Inf. Sci. | 6 |
| 2024 | Artist Similarity Based on Heterogeneous Graph Neural NetworksabstractMusic streaming platforms rely on recommending similar artists to maintain user engagement, with artists benefiting from these suggestions to boost their popularity. Another important feature is music information retrieval, allowing users to explore new content. In both scenarios, performance depends on how to compute the similarity between musical content. This is a challenging process since musical data is inherently multimodal, containing textual and audio data. We propose a novel graph-based artist representation that integrates audio, lyrics features, and artist relations. Thus, a multimodal representation on a heterogeneous graph is proposed, along with a network regularization process followed by a GNN model to aggregate multimodal information into a more robust unified representation. The proposed method explores this final multimodal representation for the task of artist similarity as a link prediction problem. Our method introduces a new importance matrix to emphasize related artists in this multimodal space. We compare our approach with other strong baselines based on combining input features, importance matrix construction, and GNN models. Experimental results highlight the superiority of multimodal representation through the transfer learning process and the value of the importance matrix in enhancing GNN models for artist similarity. Angelo Cesar Mendes da Silva, Diego Furtado Silva, Ricardo M. Marcacini |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Weak Supervision for Question and Answering Sentiment AnalysisabstractCompanies and government agencies are keen on comprehending their customers' sentiments regarding their products and services. This has given rise to the concept of Social Customer Relationship Management (Social CRM). Leveraging sentiment analysis methods, Social CRM intelligent systems aim to extract the overall sentiment concerning a product. Never-theless, conventional sentiment analysis methods face limitations when dealing with human queries that often pose questions. As a promising alternative, question-and-answer (QA) systems for sentiment analysis have emerged. However, they typically necessitate an extensive amount of annotated question-answer pairs, specifically focused on sentiment analysis, which can prove impractical. To tackle this challenge, this paper proposes an innovative approach termed Weak Supervision for Question and Answering Sentiment Analysis (WSQASA), which fine-tunes and extracts sentiment through QA models in an unsupervised manner. We explore question-generation models and sentiment filters to achieve weak supervision, thereby generating domain-specific question-and-answer pairs for fine-tuning the QA model. Our method enables the creation of domain-specific question-and-answer pairs, significantly enhancing the results of QA-based sentiment analysis, even in the absence of labeled data. Victor Akihito Kamada Tomita, Fábio M. F. Lobato, Ricardo M. Marcacini |
ICMLA | 3 |
| 2023 | On the Use of Aggregation Functions for Semi-Supervised Network EmbeddingabstractNetwork embedding methods map nodes into vector representations, aiming to preserve important properties of relationships between nodes through similarities in a latent vector-space model. Graph Neural Networks (GNNs) based on aggregate functions have received significant attention among different network embedding methods. In general, the embeddings of a node are recursively generated by aggregating embeddings from neighboring nodes. Aggregation is a crucial step in these methods, and different aggregation functions have been proposed, from simple averaging and max pooling operations to complex functions based on attention mechanisms. However, we note that there is a lack of studies comparing aggregate functions, especially in more practical real-world scenarios involving semi-supervised tasks. This paper introduces a methodology to evaluate different aggregation functions for semi-supervised learning through a model selection strategy guided by a statistical significance analysis framework. We show that Transformers-based aggregation functions are competitive for semi-supervised scenarios and obtain relevant results in different domains. Furthermore, we also discuss scenarios where “less is more”, mainly when there are constraints on the availability of computational resources. Marcelo Isaias de Moraes, Ricardo M. Marcacini |
IJCNN | 2 |
| 2023 | One-class learning for fake news detection through multimodal variational autoencoders
Marcos P. S. Gôlo, Mariana Caravanti de Souza, Rafael Geraldeli Rossi, Solange Oliveira Rezende, Bruno M. Nogueira 0001, Ricardo M. Marcacini |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | Sentence Similarity Recognition in Portuguese from Multiple Embedding ModelsabstractDistinct pre-trained embedding models perform differently in sentence similarity recognition tasks. The current assumption is that they encode different features due to differences in algorithm design and characteristics of the datasets employed in the pre-trained process. The perspective of benefiting from different encoded features to generate more suitable representations motivated the assembly of multiple embedding models, so-called meta-embedding. Meta-embedding methods combine different pre-trained embedding models to perform a task. Recently, multiple pre-trained language representations derived from Transformers architecture-based systems have been shown to be effective in many downstream tasks. This paper introduces a supervised meta-embedding neural network to combine contextualized pre-trained models for sentence similarity recognition in Portuguese. Our results show that combining multiple sentence pre-trained embedding models outperforms single models and can be a promising alternative to improve performance sentence similarity. Moreover, we also discuss the results considering our simple extension of a model explainability method to the meta-embedding context, allowing the visual identification of the impact of each token on the sentence similarity score. Ana Rodrigues 0007, Ricardo M. Marcacini |
ICMLA | 2 |
| 2022 | Opinion mining for app reviews: an analysis of textual representation and predictive models
Adailton Ferreira de Araújo, Marcos P. S. Gôlo, Ricardo M. Marcacini |
Autom. Softw. Eng. | 3 |
| 2022 | Multimodal representation learning over heterogeneous networks for tag-based music retrieval
Angelo Cesar Mendes da Silva, Diego Furtado Silva, Ricardo M. Marcacini |
Expert Syst. Appl. | 3 |
| 2022 | Detecting relevant app reviews for software evolution and maintenance through multimodal one-class learning
Marcos P. S. Gôlo, Adailton Ferreira de Araújo, Rafael Geraldeli Rossi, Ricardo M. Marcacini |
Inf. Softw. Technol. | 4 |
| 2022 | A network-based positive and unlabeled learning approach for fake news detectionabstractFake news can rapidly spread through internet users and can deceive a large audience. Due to those characteristics, they can have a direct impact on political and economic events. Machine Learning approaches have been used to assist fake news identification. However, since the spectrum of real news is broad, hard to characterize, and expensive to label data due to the high update frequency, One-Class Learning (OCL) and Positive and Unlabeled Learning (PUL) emerge as an interesting approach for content-based fake news detection using a smaller set of labeled data than traditional machine learning techniques. In particular, network-based approaches are adequate for fake news detection since they allow incorporating information from different aspects of a publication to the problem modeling. In this paper, we propose a network-based approach based on Positive and Unlabeled Learning by Label Propagation (PU-LP), a one-class and transductive semi-supervised learning algorithm that performs classification by first identifying potential interest and non-interest documents into unlabeled data and then propagating labels to classify the remaining unlabeled documents. A label propagation approach is then employed to classify the remaining unlabeled documents. We assessed the performance of our proposal considering homogeneous (only documents) and heterogeneous (documents and terms) networks. Our comparative analysis considered four OCL algorithms extensively employed in One-Class text classification ( k -Means, k -Nearest Neighbors Density-based, One-Class Support Vector Machine, and Dense Autoencoder), and another traditional PUL algorithm (Rocchio Support Vector Machine). The algorithms were evaluated in three news collections, considering balanced and extremely unbalanced scenarios. We used Bag-of-Words and Doc2Vec models to transform news into structured data. Results indicated that PU-LP approaches are more stable and achieve better results than other PUL and OCL approaches in most scenarios, performing similarly to semi-supervised binary algorithms. Also, the inclusion of terms in the news network activate better results, especially when news are distributed in the feature space considering veracity and subject. News representation using the Doc2Vec achieved better results than the Bag-of-Words model for both algorithms based on vector-space model and document similarity network. Mariana Caravanti de Souza, Bruno M. Nogueira 0001, Rafael Geraldeli Rossi, Ricardo M. Marcacini, Brucce Neves dos Santos, Solange Oliveira Rezende |
Mach. Learn. | 4 |
| 2021 | Embedding propagation over heterogeneous event networks for link predictionabstractEvents can be defined as phenomena that occur at a specific time and place. Social networks and news portals publish thousands of events daily, and this knowledge is beneficial for many social, political, and economic studies. Recently, heterogeneous networks have been used successfully for modeling large event datasets since they model different event components as nodes (e.g., events, actors, locations, people, and organizations), and network links express different relationships between these nodes. However, event analysis from heterogeneous networks is a research challenge due to several factors: (1) event nodes are usually associated with textual (unstructured) and high dimensional data; and (2) inapplicability of several machine learning methods that assume input data represented by independent vectors in a vector space. In this paper, we present a language model-based embedding propagation method for heterogeneous event networks. While most of the existing network embedding methods mainly explore the network’s topology, our method maps both (i) textual information about events and (ii) the complex relationships between events and their components to a low dimensional vector space in order to use several machine learning algorithms, such as clustering and classification. We carried out an extensive experimental evaluation involving link prediction tasks in heterogeneous event networks, such as event forecasting, prediction of event locations, and event actors. Our approach proved competitive compared to the state-of-the-art network embedding methods for link prediction tasks in different real-world event datasets, in addition to allowing dynamic and incremental updating of the embeddings as new events arise. Paulo do Carmo, Ricardo M. Marcacini |
IEEE BigData | 2 |
| 2021 | Semi-Supervised Graph Attention Networks for Event Representation LearningabstractEvent analysis from news and social networks is very useful for a wide range of social studies and real-world applications. Recently, event graphs have been explored to model event datasets and their complex relationships, where events are vertices connected to other vertices representing locations, people’s names, dates, and various other event metadata. Graph representation learning methods are promising for extracting latent features from event graphs to enable the use of different classification algorithms. However, existing methods fail to meet essential requirements for event graphs, such as (i) dealing with semi-supervised graph embedding to take advantage of some labeled events, (ii) automatically determining the importance of the relationships between event vertices and their metadata vertices, as well as (iii) dealing with the graph heterogeneity. This paper presents GNEE (GAT Neural Event Embeddings), a method that combines Graph Attention Networks and Graph Regularization. First, an event graph regularization is proposed to ensure that all graph vertices receive event features, thereby mitigating the graph heterogeneity drawback. Second, semi-supervised graph embedding with self-attention mechanism considers existing labeled events, as well as learns the importance of relationships in the event graph during the representation learning process. A statistical analysis of experimental results with five real-world event graphs and six graph embedding methods shows that our GNEE outperforms state-of-the-art semi-supervised graph embedding methods. João Mattos, Ricardo M. Marcacini |
ICDM | 2 |
| 2020 | A context-aware recommender method based on text and opinion miningabstractAbstract A recommender system is an information filtering technology that can be used to recommend items that may be of interest to users. Additionally, there are the context‐aware recommender systems that consider contextual information to generate the recommendations. Reviews can provide relevant information that can be used by recommender systems, including contextual and opinion information. In a previous work, we proposed a context‐aware recommendation method based on text mining (CARM‐TM). The method includes two techniques to extract context from reviews: CIET.5embed, a technique based on word embeddings; and RulesContext, a technique based on association rules. In this work, we have extended our previous method by including CEOM, a new technique which extracts context by using aspect‐based opinions. We call our extension of CARM‐TOM (context‐aware recommendation method based on text and opinion mining). To generate recommendations, our method makes use of the CAMF algorithm, a context‐aware recommender based on matrix factorization. To evaluate CARM‐TOM, we ran an extensive set of experiments in a dataset about restaurants, comparing CARM‐TOM against the MF algorithm, an uncontextual recommender system based on matrix factorization; and against a context extraction method proposed in literature. The empirical results strongly indicate that our method is able to improve a context‐aware recommender system. Camila Vaccari Sundermann, Renan de Padua, Vítor Rodrigues Tonon, Ricardo M. Marcacini, Marcos Aurélio Domingues, Solange Oliveira Rezende |
Expert Syst. J. Knowl. Eng. | 4 |
| 2020 | A two-stage regularization framework for heterogeneous event networksabstractEvent analysis from news and social networks is a promising way to understand complex social phenomena. Each event consists of different components, which indicate what happened, when, where, and the people and organizations involved. Heterogeneous networks are useful for modeling large event datasets, where we map different types of objects (e.g. events and their components), as well as the different relationships between objects. Such networks enable the identification of related events, in which users label some events in categories and then use the network's topological structure to find other events of interest. Although this process can be automated, there is a lack of machine learning methods to properly handle event classification from heterogeneous networks. In this paper, we present the framework named Heterogeneous Event Network Regularization in Two-stages (HENR2). The first stage of HENR2 aims to learn the importance level of each relationship between events and their components. In the second stage, the regularization process considers the importance levels of each relationship to propagate labels on the network. Thus, the classification process is improved by considering the domain characteristics of the event dataset, such as temporal seasonality and geographical distribution. In both stages, our approach also deals with noisy data through parameters that define the confidence level of labeled events during label propagation. Experimental results involving twelve event networks from different application domains show that our proposal outperforms existing regularization frameworks. Brucce Neves dos Santos, Rafael Geraldeli Rossi, Solange Oliveira Rezende, Ricardo M. Marcacini |
Pattern Recognit. Lett. | 4 |
| 2018 | Agribusiness Time Series Forecasting using Perceptually Important EventsabstractModern agribusiness management incorporates instruments for risk management with the objective of mitigating uncertainties to the producer. In this context, the producer (risk averse) transfer the risk of price oscillation to companies or individuals that operate in the futures market and who expect to receive a payment (risk premium) for assuming such risk. Defining the adequate strategies for risk management depends on the knowledge about the problem to determine prices ranges in the future. Recent studies demonstrate that time series forecasting can be significantly improved by considering additional information about the problem. In particular, besides the historical time series, textual knowledge extracted from the news portals, social networking and other public data sources available in the web may also be used. This paper presents an approach for agribusiness time series forecasting that allows incorporating external knowledge in the form of events extracted from news about agribusiness, without the need to previously label textual information. In this case, periods of significant uptrends and downtrends of time series are automatically identified - known in the literature as perceptually important points (PIP). We extend the concept of PIP to news events, where similar events published with a certain regularity in periods of uptrends and downtrends are selected as perceptually important events to improve time series forecasting models. An experimental evaluation based on price prediction on ten corn futures contracts (derivatives) provides evidence that the proposed approach is promising. Lusas S. Rodrigues, Solange Oliveira Rezende, Maria Fernanda Moura, Ricardo M. Marcacini |
CLEI | 4 |
| 2018 | Improving Instance Selection via Metric LearningabstractThe k-Nearest Neighbor (k-NN) rule is widely used for classification tasks because of its simplicity and efficiency. However, a well-known drawback of k-NN is its dependence on the quality of the training set, since the k-NN makes no assumption about the importance of each instance. In fact, the existence of noisy and superfluous instances in the training set tends to increase the classification error rate. Thus, instance selection methods are useful to identify which instances belonging to the training set will be considered in the k-NN classifier. Our proposal shows a simple and effective way to improve instance selection methods using metric learning. The idea of our proposal relies on a pure geometric intuition that metric learning transforms the input space where points in the same class are simultaneously near each other and far from points in the other classes. In a more “organised” space, we show that instance selection methods can benefit from this transformed space. We carried out an experimental evaluation to compare the instance selection with and without metric learning on UCI benchmark data sets. The results reveals that the combination of metric and instance selection is very welcome. All tested instance selection methods improved significantly. Eduardo Zarate Max, Ricardo M. Marcacini, Edson Takashi Matsubara |
IJCNN | 2 |
| 2018 | Cross-domain aspect extraction for sentiment analysis: A transductive learning approach
Ricardo M. Marcacini, Rafael Geraldeli Rossi, Ivone Penque Matsuno, Solange Oliveira Rezende |
Decis. Support Syst. | 1 |
| 2017 | Constrained Hierarchical Clustering for News EventsabstractKnowledge discovery from web news events has received great attention in recent years. In practice, this knowledge is a digital representation (virtual world) of various phenomena that occur in our physical world. Hierarchical clustering algorithms are used to organize related events into groups and subgroups according to some similarity measure. The main motivation for this organization is based on the hypothesis that if the user is interested in a specific event of a certain cluster, then the user may also be interested in other related events of this same cluster. However, existing event clustering methods do not effectively use the different types of information about events, such as temporal information, geographical data, name of people and organizations. In this paper, we propose the COH-KMeans algorithm (Constrained Hierarchical K-Means) that obtains a hierarchical clustering structure considering certain conditions imposed by the users, for example, events of similar content that occurred in nearby geographic locations or that occurred within a predefined time window. A statistical analysis of the experimental results reveals that the incorporation of constraints performed by COH-KMeans allows to obtain higher quality clusters when compared to a state-of-the-art unsupervised hierarchical clustering method. Moreover, we present our tool for exploratory analysis of events and we discuss how event clustering can be used to support the decision-making process from the perspective of a Data Analytics System. Ronaldo Florence, Bruno M. Nogueira 0001, Ricardo M. Marcacini |
IDEAS | 3 |
| 2017 | Integrating distance metric learning and cluster-level constraints in semi-supervised clusteringabstractSemi-supervised clustering has been widely explored in the last years. In this paper, we present HCAC-ML (Hierarchical Confidence-based Active Clustering with Metric Learning), an innovative approach for this task which employs distance metric learning through cluster-level constraints. HCAC-ML is based on the HCAC algorithm, an state-of-the-art algorithm for hierarchical semi-supervised clustering that uses an active learning approach for inserting cluster-level constraints. These constraints are presented to a variation of ITML (Information-theoretic Metric Learning) algorithm to learn a Mahalanobis-like distance function. We compared HCAC-ML with other semi-supervised clustering algorithms in 26 different datasets. Results indicate that HCAC-ML outperforms other algorithms in most of the scenarios, but specially when the number of constraints is small. This makes HCAC-ML useful in practical applications. Bruno M. Nogueira 0001, Yuri Karan Benevides Tomas, Ricardo M. Marcacini |
IJCNN | 3 |
| 2016 | On combining Websensors and DTW distance for kNN Time Series ForecastingabstractIn the pattern recognition field, different approaches have been proposed to improve time series forecasting models. In this sense, k-Nearest-Neighbour (kNN) with DTW (Dynamic Time Warping) distance is one of the most representative methods, due to its effectiveness, simplicity and intuitiveness. The great advantage of the DTW distance is the robustness to distortions in the time axis by allowing stretching and squeezing (time warping) of the time series, while traditional measures require a linear alignment between each data point. However, as well as other traditional measures, the DTW distance has the limitation of focusing only on historical time series data to predict future values, thereby not considering additional external knowledge of the problem domain. In this paper, we propose an approach called TSFW (Time Series Forecasting with Websensors) that incorporates Websensors into DTW distance to improve kNN time series forecasting. Websensors are models that represent knowledge extracted from news about the problem domain as well as the temporal evolution of this knowledge. In our proposed TSFW approach, we show that Websensors allow a more robust non-linear alignment of the time series by using similar events (extracted from news) that have occurred in the both time series. Thus, distortions in the time axis among the time series can be corrected more accurately compared to the traditional technique that uses only the original values of the time series. Ricardo M. Marcacini, Julio C. Carnevali, Joao Domingos |
ICPR | 1 |
| 2015 | Interactive textual feature selection for consensus clustering
Geraldo N. Correa, Ricardo M. Marcacini, Eduardo R. Hruschka, Solange Oliveira Rezende |
Pattern Recognit. Lett. | 2 |
| 2014 | Using Contextual Information from Topic Hierarchies to Improve Context-Aware Recommender SystemsabstractUnlike the traditional recommender systems, that make recommendations only by using the relation between user and item, a context-aware recommender system makes recommendations by incorporating available contextual information into the recommendation process as explicit additional categories of data to improve the recommendation process. In this paper, we propose to use contextual information from topic hierarchies to improve the accuracy of context-aware recommender systems. Additionally, we also propose two context-aware recommender algorithms for item recommendation. These are extensions from algorithms proposed in literature for rating prediction. The empirical results demonstrate that by using topic hierarchies our technique can provide better recommendations. Marcos Aurélio Domingues, Marcelo G. Manzato, Ricardo M. Marcacini, Camila Vaccari Sundermann, Solange Oliveira Rezende |
ICPR | 3 |
| 2014 | Improving Personalized Ranking in Recommender Systems with Topic Hierarchies and Implicit FeedbackabstractThe knowledge of semantic information about the content and user's preferences is an important issue to improve recommender systems. However, the extraction of such meaningful metadata needs an intense and time-consuming human effort, which is impractical specially with large databases. In this paper, we mitigate this problem by proposing a recommendation model based on latent factors and implicit feedback which uses an unsupervised topic hierarchy constructor algorithm to organize and collect metadata at different granularities from unstructured textual content. We provide an empirical evaluation using a dataset of web pages written in Portuguese language, and the results show that personalized ranking with better quality can be generated using the extracted topics at medium granularity. Marcelo G. Manzato, Marcos Aurélio Domingues, Ricardo M. Marcacini, Solange Oliveira Rezende |
ICPR | 3 |
| 2014 | Privileged Information for Hierarchical Document Clustering: A Metric Learning ApproachabstractTraditional hierarchical text clustering methods assume that the documents are represented only by "technical information", i.e., keywords, phrases, expressions and named entities that can be directly extracted from the texts. However, in many scenarios there is an additional and valuable information about the documents which is usually disregarded during the clustering task, such as user-validated tags, annotations and comments from experts, dictionaries and domain ontologies. Recently, Vapnik introduced a new learning paradigm, called LUPI - Learning Using Privileged Information, which allows the incorporation of this additional (privileged) information in a supervised learning setting. We investigated the incorporation of privileged information in unsupervised setting. The key idea in our proposed approach is to extract important relationships among documents represented in the privileged information dimensional space to learn a more accurate metric for text clustering in the technical information space. A thorough experimental evaluation indicates that the incorporation of privileged information through metric learning significantly improves the hierarchical clustering accuracy. Ricardo M. Marcacini, Marcos Aurélio Domingues, Eduardo R. Hruschka, Solange Oliveira Rezende |
ICPR | 1 |
| 2014 | Named entities as privileged information for hierarchical text clusteringabstractText clustering is a text mining task which is often used to aid the organization, knowledge extraction, and exploratory search of text collections. Nowadays, the automatic text clustering becomes essential as the volume and variety of digital text documents increase, either in social networks and the Web or inside organizations. This paper explores the use of named entities as privileged information in a hierarchical clustering process, so as to improve clusters quality and interpretation. We carried out an experimental evaluation on three text collections (one written in Portuguese and two written in English) and the results show that named entities can be applied as privileged information to power clustering solution in dynamic text collection scenarios. Roberta Akemi Sinoara, Camila Vaccari Sundermann, Ricardo M. Marcacini, Marcos Aurélio Domingues, Solange Oliveira Rezende |
IDEAS | 3 |
| 2013 | Incremental hierarchical text clustering with privileged informationabstractIn many text clustering tasks, there is some valuable knowledge about the problem domain, in addition to the original textual data involved in the clustering process. Traditional text clustering methods are unable to incorporate such additional (privileged) information into data clustering. Recently, a new paradigm called LUPI - Learning Using Privileged Information - was proposed by Vapnik to incorporate privileged information in classification tasks. In this paper, we extend the LUPI paradigm to deal with text clustering tasks. In particular, we show that the LUPI paradigm is potentially promising for incremental hierarchical text clustering, being very useful for organizing large textual databases. In our method, the privileged information about the text documents is applied to refine an initial clustering model by means of consensus clustering. The initial model is used for incremental clustering of the remaining text documents. We carried out an experimental evaluation on two benchmark text collections and the results showed that our method significantly improves the clustering accuracy when compared to a traditional hierarchical clustering method. Ricardo M. Marcacini, Solange Oliveira Rezende |
ACM Symposium on Document Engineering | 1 |
| 2012 | An active learning approach to frequent itemset-based text clustering
Ricardo M. Marcacini, Geraldo N. Correa, Solange Oliveira Rezende |
ICPR | 1 |
| 2010 | Incremental Construction of Topic Hierarchies using Hierarchical Term Clustering
Ricardo M. Marcacini, Solange Oliveira Rezende |
SEKE | 1 |