Gisele L. Pappa

dblp:59/701 · also Gisele Lobo Pappa · DBLP profile ↗
← Back
15ranked-venue papers in the field
2as first author
3since 2021 · last 2025
0000-0002-0349-4494ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 9 (2 first)Information Retrieval & Web Search · 2Knowledge Engineering, Semantic Web & Information Systems · 2Other / Interdisciplinary · 2
YearPublicationVenuePosition
2025 A CNN-Based Local-Global Self-attention via Averaged Window Embeddings for Hierarchical ECG Analysis
Arthur Buzelin, Pedro Robles Dutenhefner, Turi Rezende, Luisa G. Porfírio, Pedro Bento, Yan Aquino, Jose Fernandes, Caio Santana, Gabriela Miana, Gisele L. Pappa, Antônio L. P. Ribeiro, Wagner Meira Jr.
ECML/PKDD (3)10
2022 Counterfactual inference with latent variable and its application in mental health care
Guilherme F. Marchezini, Anísio Lacerda, Gisele L. Pappa, Wagner Meira Jr., Débora M. Miranda, Marco Aurélio Romano-Silva, Danielle S. Costa, Leandro Malloy Diniz
Data Min. Knowl. Discov.3
2021 An Instance Space Analysis of Regression Problems
abstract
The quest for greater insights into algorithm strengths and weaknesses, as revealed when studying algorithm performance on large collections of test problems, is supported by interactive visual analytics tools. A recent advance is Instance Space Analysis, which presents a visualization of the space occupied by the test datasets, and the performance of algorithms across the instance space. The strengths and weaknesses of algorithms can be visually assessed, and the adequacy of the test datasets can be scrutinized through visual analytics. This article presents the first Instance Space Analysis of regression problems in Machine Learning, considering the performance of 14 popular algorithms on 4,855 test datasets from a variety of sources. The two-dimensional instance space is defined by measurable characteristics of regression problems, selected from over 26 candidate features. It enables the similarities and differences between test instances to be visualized, along with the predictive performance of regression algorithms across the entire instance space. The purpose of creating this framework for visual analysis of an instance space is twofold: one may assess the capability and suitability of various regression techniques; meanwhile the bias, diversity, and level of difficulty of the regression problems popularly used by the community can be visually revealed. This article shows the applicability of the created regression instance space to provide insights into the strengths and weaknesses of regression algorithms, and the opportunities to diversify the benchmark test instances to support greater insights.
Mario A. Muñoz, Matheus R. Leal, Kate Smith-Miles, Ana Carolina Lorena, Gisele L. Pappa, Rômulo Madureira Rodrigues
ACM Trans. Knowl. Discov. Data6
2020 Is Rank Aggregation Effective in Recommender Systems? An Experimental Analysis
abstract
Recommender Systems are tools designed to help users find relevant information from the myriad of content available online. They work by actively suggesting items that are relevant to users according to their historical preferences or observed actions. Among recommender systems, top- N recommenders work by suggesting a ranking of N items that can be of interest to a user. Although a significant number of top- N recommenders have been proposed in the literature, they often disagree in their returned rankings, offering an opportunity for improving the final recommendation ranking by aggregating the outputs of different algorithms. Rank aggregation was successfully used in a significant number of areas, but only a few rank aggregation methods have been proposed in the recommender systems literature. Furthermore, there is a lack of studies regarding rankings’ characteristics and their possible impacts on the improvements achieved through rank aggregation. This work presents an extensive two-phase experimental analysis of rank aggregation in recommender systems. In the first phase, we investigate the characteristics of rankings recommended by 15 different top- N recommender algorithms regarding agreement and diversity. In the second phase, we look at the results of 19 rank aggregation methods and identify different scenarios where they perform best or worst according to the input rankings’ characteristics. Our results show that supervised rank aggregation methods provide improvements in the results of the recommended rankings in six out of seven datasets. These methods provide robustness even in the presence of a big set of weak recommendation rankings. However, in cases where there was a set of non-diverse high-quality input rankings, supervised and unsupervised algorithms produced similar results. In these cases, we can avoid the cost of the former in favor of the latter.
Samuel E. L. Oliveira, Victor Diniz, Anísio Lacerda, Luiz H. C. Merschmann, Gisele L. Pappa
ACM Trans. Intell. Syst. Technol.5
2018 Reddit Weight Loss Communities: Do They Have What It Takes for Effective Health Interventions?
abstract
Online social networks are an important tool for people to share information and have been extensively used for people to achieve beneficial changes in health. Obesity is a major public health concern that affects about one third of the world's population. In order to alleviate this problem, health professionals are focusing on health interventions, which can be performed online. In this study we analyze three distinct online communities about weight and diet in Reddit. We model our data as 3 directed and weighted graphs of the posts and comments and evaluate the interaction between users of each community. We also analyze specific characteristics of each community, the habits of daily activity of the users and the formation of implicit bonds of friendship through the formation of communities. Our main results show that Reddit is a content-centered social network, in which what matters is what is posted and not who posts. In addition, users tend to create implicit friendship relationships through denser regions of interactions. Our results show that, contrary to expectations, the three communities present the same behavior pattern in a general point of view, which facilitates the development of non-directed online weight loss intervention strategies.
Karen Braga Enes, Pedro Paulo Valadares Brum, Tiago Oliveira Cunha, Fabricio Murai, Ana Paula Couto da Silva, Gisele L. Pappa
WI6
2018 Selective harvesting over networks
Fabricio Murai, Diogo Rennó, Bruno Ribeiro 0001, Gisele L. Pappa, Don Towsley, Krista Gile
Data Min. Knowl. Discov.4
2018 Strategies for combining Twitter users geo-location methods
Sílvio S. Ribeiro Jr., Gisele L. Pappa
GeoInformatica2
2017 A general framework to expand short text for topic modeling
Paulo Viana Bicalho, Marcelo Pita, Gabriel Pedrosa, Anísio Lacerda, Gisele L. Pappa
Inf. Sci.5
2017 H3AD: A hybrid hyper-heuristic for algorithm design
Péricles B. C. Miranda, Ricardo B. C. Prudêncio, Gisele L. Pappa
Inf. Sci.3
2015 Twitter Population Sample Bias and its impact on predictive outcomes: a case study on elections
abstract
In the past years a lot of effort has been spent analyzing online social network data to understand how the world reality is reflected in the "virtual" world. Twitter is by far the network most used in these studies, given its policy of public data availability. However, a big discussion is still on on whether the data available is enough to make user characterization or event outcomes prediction, and what are the pitfalls people do not usually account for. In this direction, we propose a new methodology for drawing representative samples from Twitter data, which is divided into four phases: (i) user filtering, (ii) user demographic characterization, (iii) user sampling, and (iv) event prediction. The methodology is tested into a common scenario in Twitter event outcome prediction: elections. The methodology was tested with municipality elections from six different Brazilian cities, and compared to official election results. Results show it is worth further investigating the topic, but that a very hight number of messages is required to match real data distributions.
Renato Miranda, Jussara M. Almeida, Gisele L. Pappa
ASONAM3
2010 Exploiting co-occurrence and information quality metrics to recommend tags in web 2.0 applications
abstract
This work addresses the task of recommending high quality tags by exploiting not only previously assigned tags, but also terms extracted from other textual features (e.g., title and description) associated with the target object.To estimate the quality of a candidate tag recommendation, we use several metrics related to both tag co-occurrence and information quality. We also propose a heuristic function to combine the metrics to produce a final ranking of the recommended tags. We evaluate our heuristic function in various scenarios, for three popular Web 2.0 applications. Our experimental results indicate that our heuristic function significantly outperforms two state-of-the-art tag recommendation algorithms.
Fabiano Muniz Belém, Eder Ferreira Martins, Jussara M. Almeida, Marcos André Gonçalves, Gisele L. Pappa
CIKM5
2010 Demand-Driven Tag Recommendation
Guilherme Vale Menezes, Jussara M. Almeida, Fabiano Muniz Belém, Marcos André Gonçalves, Anísio Lacerda, Edleno Silva de Moura, Gisele L. Pappa, Adriano Veloso, Nivio Ziviani
ECML/PKDD (2)7
2010 Temporally-aware algorithms for document classification
abstract
Automatic Document Classification (ADC) is still one of the major information retrieval problems. It usually employs a supervised learning strategy, where we first build a classification model using pre-classified documents and then use this model to classify unseen documents. The majority of supervised algorithms consider that all documents provide equally important information. However, in practice, a document may be considered more or less important to build the classification model according to several factors, such as its timeliness, the venue where it was published in, its authors, among others. In this paper, we are particularly concerned with the impact that temporal effects may have on ADC and how to minimize such impact. In order to deal with these effects, we introduce a temporal weighting function (TWF) and propose a methodology to determine it for document collections. We applied the proposed methodology to ACM-DL and Medline and found that the TWF of both follows a lognormal. We then extend three ADC algorithms (namely kNN, Rocchio and Naïve Bayes) to incorporate the TWF. Experiments showed that the temporally-aware classifiers achieved significant gains, outperforming (or at least matching) state-of-the-art algorithms.
Thiago Salles, Leonardo Rocha 0001, Gisele L. Pappa, Fernando Mourão, Wagner Meira Jr., Marcos André Gonçalves
SIGIR3
2009 Evolving rule induction algorithms with multi-objective grammar-based genetic programming
Gisele L. Pappa, Alex Alves Freitas
Knowl. Inf. Syst.1
2006 Automatically Evolving Rule Induction Algorithms
Gisele L. Pappa, Alex Alves Freitas
ECML1