Gisele L. Pappa

dblp:59/701 · also Gisele Lobo Pappa · DBLP profile ↗
← Back
70ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0002-0349-4494ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 57 · 3 first-author · 17 since 2021Databases, data management, data science and information retrieval · 15 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Evo-Reasoner: Evolutionary Optimization of Structured Reasoning in LLMs
Pedro Bento, Arthur Buzelin, Arthur Chagas, Yan Aquino, Wagner Meira Jr., Gisele L. Pappa
EvoApplications6
2025 Evolutionary Bias Identification with Embeddings
Arthur Buzelin, Yan Aquino, Victoria Estanislau, Pedro Bento, Lucas Dayrell, Samira Malaquias, Caio Santana, Guilherme H. G. Evangelista, Caio Souza Grossi, Pedro B. Rigueira, Luisa G. Porfírio, Marcelo Sartori Locatelli, Wagner Meira Jr., Gisele L. Pappa
EvoApplications (2)14
2025 Transformers as Surrogate Models for Genetic Programming in AutoML Tasks
abstract
In applications where the fitness function has a high computational cost, one of the main drawbacks of Evolutionary Algorithms when compared to other search methods is a prohibitive computational cost. The use of surrogates as proxies for fitness function calculation to alleviate this problem is not new, but addressing the problem as a binary relation learning, i.e., evaluating if one individual is better or worse than another without estimating the actual value of the fitness, is a recent trend.
Matheus Cândido Teixeira, Gisele L. Pappa
GECCO2
2025 Evaluation of Medical Large Language Models: Taxonomy, Review, and Directions
abstract
The integration of Large Language Models (LLMs) into medicine presents both great opportunities and significant challenges, particularly in ensuring these models are accurate, reliable, and safe. While LLMs have shown impressive capabilities in understanding and generating human language, their application in the medical domain requires careful evaluation due to the critical nature of medical applications which are inherently linked to patient life and health. Current evaluations of LLMs in medicine are often fragmented and insufficient, with a lack of standardized performance metrics, limited use of real patient data, and insufficient attention to important applications, such as documentation, education, and research. Furthermore, traditional NLP-based evaluations are often inadequate for assessing the text generated by LLMs. Therefore, a robust evaluation is essential to ensure the responsible and effective use of LLMs in medical settings, and to address the inherent challenges associated with their implementation. This paper explores the various dimensions of LLM evaluation in the medical domain, proposes a new taxonomy for categorizing medical applications, and discusses directions for future research in this critical area.
Anísio Lacerda, Gisele L. Pappa, Adriano C. M. Pereira, Wagner Meira Jr., Alexandre Guimarães de Almeida Barros
IJCAI2
2025 A CNN-Based Local-Global Self-attention via Averaged Window Embeddings for Hierarchical ECG Analysis
Arthur Buzelin, Pedro Robles Dutenhefner, Turi Rezende, Luisa G. Porfírio, Pedro Bento, Yan Aquino, Jose Fernandes, Caio Santana, Gabriela Miana, Gisele L. Pappa, Antônio L. P. Ribeiro, Wagner Meira Jr.
ECML/PKDD (3)10
2025 Novel applications of item response theory for analysing data set complexity and benchmark selection
João Luiz Junho Pereira, Alfredo Antonio Alencar Exposito de Queiroz, Telmo de Menezes e Silva Filho, Ana Carolina Lorena, Rafael Gomes Mantovani, Gisele L. Pappa, Ricardo B. C. Prudêncio
Mach. Learn.6
2024 Unsupervised Grouping of Public Procurement Similar Items: Which Text Representation Should I Use?
abstract
In public procurement, establishing reference prices is essential to guide competitors in setting product prices. Group-purchased products, which are not standardized by default, are necessary to estimate reference prices. Text clustering techniques can be used to group similar items based on their descriptions, enabling the definition of reference prices for specific products or services. However, selecting an appropriate representation for text is challenging. This paper introduces a framework for text cleaning, extraction, and representation. We test eight distinct sentence representations tailored for public procurement item descriptions. Among these representations, we propose an approach that captures the most important components of item descriptions. Through extensive evaluation of a dataset comprising over 2 million items, our findings show that using sophisticated supervised methods to derive vectors for unsupervised tasks offers little advantages over leveraging unsupervised methods. Our results also highlight that domain-specific contextual knowledge is crucial for representation improvement.
Pedro Paulo Valadares Brum, Mariana O. Silva, Gabriel P. Oliveira, Lucas G. L. Costa, Anísio Lacerda, Gisele L. Pappa
LREC/COLING6
2023 On the Effect of Solution Representation and Neighborhood Definition in AutoML Fitness Landscapes
Matheus Cândido Teixeira, Gisele L. Pappa
EvoCOP2
2023 Symbolic Regression Trees as Embedded Representations
abstract
Representation learning is an area responsible for learning data representations that makes it easier for machine learning algorithms to extract useful information from them. Deep learning currently has the most effective methods for this task and can learn distributed representations - also known as embeddings - able to represent different properties of the data and their relationship. In this direction, this paper introduces a new way to look at tree-like GP individuals for symbolic regression. Given a set of predefined operators and a sufficiently large number of solutions sampled from the space, we train a transformer to learn an encoding/decoding function. By transforming a tree representation into a distributed representation, we are able to measure distances between trees in a much more efficient way and, more importantly, generate the potential for these representations to capture semantics. We show the distance accounting for embedding presents results very similar to those of a tree-edition, which reflects their syntactic similarity. Although the model as it stands is not able to capture semantics yet, we show its potential by using the generated tree-representation model in a simple task: measuring distances between trees in a fitness-sharing scenario.
Victor Caetano, Matheus Cândido Teixeira, Gisele L. Pappa
GECCO3
2023 Algorithmic Recourse in Mental Healthcare
abstract
This paper explores using algorithmic recourse as a tool in mental healthcare. Algorithmic recourse provides explanations and recommendations to individuals who want to reverse a machine learning prediction and has been widely used in various domains such as finance and marketing. However, its application in mental healthcare has been restricted. This paper addresses this by examining the potential benefits and challenges of using algorithmic recourse in mental healthcare, specifically in how changing one's behavior may affect their quality of life and well-being. The paper proposes a new classification-based framework for algorithmic recourse in mental healthcare. The proposed framework considers both observed and latent variables to account for the individuality of individuals and provides a more comprehensive understanding of mental health outcomes. The research results can provide valuable insights for future work and help bridge the gap between machine learning and mental healthcare.
Anísio Lacerda, Claudio Almeida, Leonardo Augusto Ferreira, Adriano C. M. Pereira, Gisele L. Pappa, Wagner Meira Jr., Débora M. Miranda, Marco Aurélio Romano-Silva, Leandro Malloy Diniz
IJCNN5
2022 Understanding AutoML search spaces with local optima networks
abstract
AutoML tackles the problem of automatically configuring machine learning pipelines to specific data analysis problems. These pipelines may include methods for preprocessing and classifying data. There are many strategies to find the best pipeline configuration, and most of them address the problem from an optimization perspective, using Bayesian optimization and evolutionary algorithms. However, little is known about the shape of the search space these methods work upon. What is known for a fact is that these spaces include both categorical, continuous and conditional hyperparameters, and dealing with them may not be trivial. In this direction, this paper performs an analysis of AutoML search spaces using local optimal networks (LON) to better understand the global properties of the search space. Knowing these properties allow us to improve current optimization methods. By analyzing the metrics extracted from the LON we got many insights on the difficulty of the problem.
Matheus Cândido Teixeira, Gisele L. Pappa
GECCO2
2022 Probabilistic topic modeling for short text based on word embedding networks
Marcelo Pita, Matheus Nunes, Gisele L. Pappa
Appl. Intell.3
2022 Counterfactual inference with latent variable and its application in mental health care
Guilherme F. Marchezini, Anísio Lacerda, Gisele L. Pappa, Wagner Meira Jr., Débora M. Miranda, Marco Aurélio Romano-Silva, Danielle S. Costa, Leandro Malloy Diniz
Data Min. Knowl. Discov.3
2022 Explainable Regression Via Prototypes
abstract
Model interpretability/explainability is increasingly a concern when applying machine learning to real-world problems. In this article, we are interested in explaining regression models by exploiting prototypes, which are exemplar cases in the problem domain. Previous works focused on finding prototypes that are representative of all training data but ignore the model predictions, i.e., they explain the data distribution but not necessarily the predictions. We propose a two-level model-agnostic method that considers prototypes to provide global and local explanations for regression problems and that account for both the input features and the model output. M-PEER (Multiobjective Prototype-basEd Explanation for Regression) is based on a multi-objective evolutionary method that optimizes both the error of the explainable model and two other “semantics”-based measures of interpretability adapted from the context of classification, namely, model fidelity and stability. We compare the proposed method with the state-of-the-art method based on prototypes for explanation—ProtoDash—and with other methods widely used in correlated areas of machine learning, such as instance selection and clustering. We conduct experiments on 25 datasets, and results demonstrate significant gains of M-PEER over other strategies, with an average of 12% improvement in the proposed metrics (i.e., model fidelity and stability) and 17% in root mean squared error (RMSE) when compared to ProtoDash.
Renato Miranda, Anísio Lacerda, Gisele L. Pappa
ACM Trans. Evol. Learn. Optim.3
2021 Fitness landscape analysis of graph neural network architecture search spaces
abstract
Neural Architecture Search (NAS) is the name given to a set of methods designed to automatically configure the layout of neural networks. Their success on Convolutional Neural Networks inspired its use on optimizing other types of neural network architectures, including Graph Neural Networks (GNNs). GNNs have been extensively applied over several collections of real-world data, achieving state-of-the-art results in tasks such as circuit design, molecular structure generation and anomaly detection. Many GNN models have been recently proposed, and choosing the best model for each problem has become a cumbersome and error-prone task. Aiming to alleviate this problem, recent works have proposed strategies for applying NAS to GNN models. However, different search methods converge relatively fast in the search for a good architecture, which raises questions about the structure of the problem. In this work we use Fitness Landscape Analysis (FLA) measures to characterize the search space explored by NAS methods for GNNs. We sample almost 90k different architectures that cover most of the fitness range, and represent them using both a one-hot encoding and an embedding representation. Results of the fitness distance correlation and dispersion metrics show the fitness landscape is easy to be explored, and presents low neutrality.
Matheus Nunes, Paulo M. Fraga, Gisele L. Pappa
GECCO3
2021 Automatic Drone Identification Through Rhythm-based Features for the Internet of Drones
abstract
The Internet of Drones (IoD) refers to a robust mobile network with well-defined airways where drones perform interoperable services, enhancing the deployment of Smart Cities. Automatic Drone Identification (ADI) is a protection mechanism to detect and avoid malicious drones, where different techniques have been used, such as acoustic signals. In this field, the sound generated by propellers and motors has particular characteristics, being a potential aspect to explore; however, this investigation is still missing. This study examines the use of rhythm-based descriptors as input features to ADI, based on the hypothesis that the acoustic signal generated by different drones has different rhythmic properties. Aiming to explore and validate our approach, we formulate an ADI methodology using rhythm-based features. We use a freely available drone audio dataset, comparing our results with a baseline study. As a result, our classification model improves 3.47% the baseline binary classification and 2.97% the multiclass classification, reaching accuracy rates of 0.9985 and 0.9591, respectively. Although the improvements are narrow, they point out that acoustic features have a great potential to enhance ADI mainly in dense drone-based environments, where drone identification is an essential task.
Alisson Renan Svaigen, Lailla M. Siqueira Bine, Gisele L. Pappa, Linnyer B. Ruiz, Antonio Alfredo Ferreira Loureiro
ICTAI3
2021 Deep Thompson Sampling for Length of Stay Prediction
abstract
Length of stay (LoS) in the intensive care unit (ICU) is a general outcome measure used as an indicator of both quality of care and resource use. Existing LoS prediction methods usually model the problem as a binary classification or regression problem. This paper proposes to address this problem under a sequential learning framework using an embedded representation of patients. First, we exploit electronic health records (EHR) to learn patient distributed representation. Second, we formulate the LoS prediction task as a Deep Bayesian contextual bandit problem, which sequentially predicts LoS based on contextual distributed information about patients, with the goal of maximizing overall prediction outcomes. Experimental results show that our model has superior performance than other sequential learning approaches in the LoS task using the publicly available MIMIC-III ICU dataset. We also evaluate the robustness of our framework under different patient representations.
Anísio Lacerda, Gisele L. Pappa
IJCNN2
2021 Neural Architecture Search for Resource-Constrained Internet of Things Devices
abstract
The traditional process of extracting knowledge from the Internet of Things (IoT) happens through Cloud Computing by offloading the data generated in the IoT device to processing in the cloud. However, this regime significantly increases data transmission and monetary costs and may have privacy issues. Therefore, it is paramount to find solutions that achieve good results and can be processed as close as possible to an IoT object. In this scenario, we developed a Neural Architecture Search (NAS) solution to generate models small enough to be deployed to IoT devices without significantly losing inference performance. We based our approach on Evolutionary Algorithms, such as Grammatical Evolution and NSGA-II. Using model size and accuracy as fitness, our proposal generated a Convolutional Neural Network model with less than 2 MB, achieving an accuracy of about 81 % in the CIFAR-10 and 99 % in MNIST, with only 150 thousand parameters approximately.
Isadora Cardoso, Gisele L. Pappa, Heitor S. Ramos
ISCC2
2021 Towards automatic diagnosis of rheumatic heart disease on echocardiographic exams through video-based deep learning
abstract
OBJECTIVE: Rheumatic heart disease (RHD) affects an estimated 39 million people worldwide and is the most common acquired heart disease in children and young adults. Echocardiograms are the gold standard for diagnosis of RHD, but there is a shortage of skilled experts to allow widespread screenings for early detection and prevention of the disease progress. We propose an automated RHD diagnosis system that can help bridge this gap. MATERIALS AND METHODS: Experiments were conducted on a dataset with 11 646 echocardiography videos from 912 exams, obtained during screenings in underdeveloped areas of Brazil and Uganda. We address the challenges of RHD identification with a 3D convolutional neural network (C3D), comparing its performance with a 2D convolutional neural network (VGG16) that is commonly used in the echocardiogram literature. We also propose a supervised aggregation technique to combine video predictions into a single exam diagnosis. RESULTS: The proposed approach obtained an accuracy of 72.77% for exam diagnosis. The results for the C3D were significantly better than the ones obtained by the VGG16 network for videos, showing the importance of considering the temporal information during the diagnostic. The proposed aggregation model showed significantly better accuracy than the majority voting strategy and also appears to be capable of capturing underlying biases in the neural network output distribution, balancing them for a more correct diagnosis. CONCLUSION: Automatic diagnosis of echo-detected RHD is feasible and, with further research, has the potential to reduce the workload of experts, enabling the implementation of more widespread screening programs worldwide.
Joao Francisco B. S. Martins, Erickson R. Nascimento, Bruno Ramos Nascimento, Craig A. Sable, Andrea Z. Beaton, Antônio L. P. Ribeiro, Wagner Meira Jr., Gisele L. Pappa
J. Am. Medical Informatics Assoc.8
2021 Multi-region symbolic regression: combining functions under a multi-objective approach
Felipe Casadei, Gisele L. Pappa
Nat. Comput.2
2021 An Instance Space Analysis of Regression Problems
abstract
The quest for greater insights into algorithm strengths and weaknesses, as revealed when studying algorithm performance on large collections of test problems, is supported by interactive visual analytics tools. A recent advance is Instance Space Analysis, which presents a visualization of the space occupied by the test datasets, and the performance of algorithms across the instance space. The strengths and weaknesses of algorithms can be visually assessed, and the adequacy of the test datasets can be scrutinized through visual analytics. This article presents the first Instance Space Analysis of regression problems in Machine Learning, considering the performance of 14 popular algorithms on 4,855 test datasets from a variety of sources. The two-dimensional instance space is defined by measurable characteristics of regression problems, selected from over 26 candidate features. It enables the similarities and differences between test instances to be visualized, along with the predictive performance of regression algorithms across the entire instance space. The purpose of creating this framework for visual analysis of an instance space is twofold: one may assess the capability and suitability of various regression techniques; meanwhile the bias, diversity, and level of difficulty of the regression problems popularly used by the community can be visually revealed. This article shows the applicability of the created regression instance space to provide insights into the strengths and weaknesses of regression algorithms, and the opportunities to diversify the benchmark test instances to support greater insights.
Mario A. Muñoz, Matheus R. Leal, Kate Smith-Miles, Ana Carolina Lorena, Gisele L. Pappa, Rômulo Madureira Rodrigues
ACM Trans. Knowl. Discov. Data6
2020 Explaining Symbolic Regression Predictions
abstract
The outgrowing application of machine learning methods has raised a discussion in the artificial intelligence community on model transparency. In the center of this discussion is the question of model explanation and interpretability. The genetic programming (GP) community has systematically pointed out as one of the major advantages of GP the fact that it produces models that can be interpreted by humans. However, as other interpretable supervised models, the more complex the model becomes, the less interpretable it is. This work focuses on post-hoc interpretability of GP for symbolic regression. This approach does not explain the process followed by a model to reach a decision. Instead, it justifies the predictions it makes. The proposed approach, named Explanation by Local Approximation (ELA), is simple and model agnostic: it finds the nearest neighbors of the point we want to explain and performs a linear regression using this subset of points. The coefficients of this linear regression are then used to generate a local explanation to the model. Results show that the errors of ELA are similar to those of the regression performed with all points. It also shows that simple visualizations can provide insights to the users about the most relevant attributes.
Renato Miranda, Anísio Lacerda, Gisele L. Pappa
CEC3
2020 Instance Selection for Geometric Semantic Genetic Programming
abstract
Geometric Semantic Genetic Programming (GSGP) is a method that exploits the geometric properties describing the spatial relationship between possible solutions to a problem in an n-dimensional semantic space. In symbolic regression problems, n is equal to the number of training instances. Although very effective, the GSGP semantic space can become excessively big in most real applications, where the value of n is high, having a negative impact on the effectiveness of the GSGP search process. This paper tackles this problem by reducing the dimensionality of GSGP semantic space in symbolic regression problems using instance selection methods. Our approach relies on weighting functions-to estimate the relative importance of each instance based on its position with respect to its nearest neighbours-and on dimensionality reduction techniques-to improve the notion of closeness between instances, generating datasets with simplified input spaces. Experiments were performed on a set of 15 datasets and our experimental analysis shows that using instance selection by instance weighting and dimensionality reduction does improve the effectiveness of the search with almost no impact on root mean square error results.
Luis Fernando Miranda, Luiz Otávio Vilas Boas Oliveira, Joao Francisco Barreto da Silva Martins, Gisele L. Pappa
CEC4
2020 Fitness Landscape Analysis of Automated Machine Learning Search Spaces
Cristiano Guimarães Pimenta, Alex Guimarães Cardoso de Sá, Gabriela Ochoa, Gisele L. Pappa
EvoCOP4
2020 A robust experimental evaluation of automated multi-label classification methods
abstract
Automated Machine Learning (AutoML) has emerged to deal with the selection and configuration of algorithms for a given learning task. With the progression of AutoML, several effective methods were introduced, especially for traditional classification and regression problems. Apart from the AutoML success, several issues remain open. One issue, in particular, is the lack of ability of AutoML methods to deal with different types of data. Based on this scenario, this paper approaches AutoML for multi-label classification (MLC) problems. In MLC, each example can be simultaneously associated to several class labels, unlike the standard classification task, where an example is associated to just one class label. In this work, we provide a general comparison of five automated multi-label classification methods - two evolutionary methods, one Bayesian optimization method, one random search and one greedy search - on 14 datasets and three designed search spaces. Overall, we observe that the most prominent method is the one based on a canonical grammar-based genetic programming (GGP) search method, namely Auto-MEKAGGP. Auto-MEKAGGP presented the best average results in our comparison and was statistically better than all the other methods in different search spaces and evaluated measures, except when compared to the greedy search method.
Alex Guimarães Cardoso de Sá, Cristiano Guimarães Pimenta, Gisele L. Pappa, Alex Alves Freitas
GECCO3
2020 Is Rank Aggregation Effective in Recommender Systems? An Experimental Analysis
abstract
Recommender Systems are tools designed to help users find relevant information from the myriad of content available online. They work by actively suggesting items that are relevant to users according to their historical preferences or observed actions. Among recommender systems, top- N recommenders work by suggesting a ranking of N items that can be of interest to a user. Although a significant number of top- N recommenders have been proposed in the literature, they often disagree in their returned rankings, offering an opportunity for improving the final recommendation ranking by aggregating the outputs of different algorithms. Rank aggregation was successfully used in a significant number of areas, but only a few rank aggregation methods have been proposed in the recommender systems literature. Furthermore, there is a lack of studies regarding rankings’ characteristics and their possible impacts on the improvements achieved through rank aggregation. This work presents an extensive two-phase experimental analysis of rank aggregation in recommender systems. In the first phase, we investigate the characteristics of rankings recommended by 15 different top- N recommender algorithms regarding agreement and diversity. In the second phase, we look at the results of 19 rank aggregation methods and identify different scenarios where they perform best or worst according to the input rankings’ characteristics. Our results show that supervised rank aggregation methods provide improvements in the results of the recommended rankings in six out of seven datasets. These methods provide robustness even in the presence of a big set of weak recommendation rankings. However, in cases where there was a set of non-diverse high-quality input rankings, supervised and unsupervised algorithms produced similar results. In these cases, we can avoid the cost of the former in favor of the latter.
Samuel E. L. Oliveira, Victor Diniz, Anísio Lacerda, Luiz H. C. Merschmann, Gisele L. Pappa
ACM Trans. Intell. Syst. Technol.5
2018 Multi-objective Evolutionary Rank Aggregation for Recommender Systems
abstract
Recommender systems help users to overcome the information overload problem by selecting relevant items according to their preferences. This paper deals with the problem of rank aggregation in recommender systems, where we want to generate a single consensus ranking from a given set of input rankings generated by different recommendation algorithms. This problem is NP-hard, and hence the use of meta-heuristics to solve it is appealing. Although accurate suggestions are mandatory for effective recommender systems, other recommendation quality measures need to be taken into account for delivering high-quality suggestions. This paper proposes Multi-objective Evolutionary Rank Aggregation (MERA), a genetic programming algorithm following the concepts of SPEA2 that considers three measures when suggesting items to users, namely mean average precision, diversity, and novelty. The method was tested in 3 realworld recommendation datasets, and the results show MERA can indeed find a balance for these metrics while generating a diverse set of solutions to the problem. MERA was able to return solutions with improvements of up to 15% in diversity (for the Movielens 1M dataset) and 7% in novelty (for the Filmtrust dataset) while maintaining, or even improving, the values of precision.
Samuel E. L. Oliveira, Victor Diniz, Anísio Lacerda, Gisele L. Pappa
CEC4
2018 Improving Energy Efficiency of Field-Coupled Nanocomputing Circuits by Evolutionary Synthesis
abstract
Moore's law provoked decades of advances in computer's performance due to transistor's evolution. Despite all success in its improvement, current technology is reaching its physical limits and some replacements are the focus of investigations, such as the Field-Coupled Nanocomputing devices. These devices achieve information transfer and computation via local field interactions, reaching ultra-low power consumption. Nevertheless, there exists a hard energy limit related to the Laws of Thermodynamics that bounds any digital evaluation. To reduce the impact of this restriction, we propose a fitness function to improve the energy efficiency on a given circuit implementation. We embed our fitness function on an evolutionary method known as Cartesian Genetic Programming to iteratively modify the circuit, searching for a new valid configuration that dissipates less energy. To assess our method, we use it on pre-optimized circuits from benchmarks and compare our results with the ones from the classic Cartesian Genetic Programming. Based on the outcome, we show that our method outperforms the latter, achieving, on average, 15% gain in energy efficiency.
Marco A. Ribeiro, Iago A. Carvalho, Jeferson F. Chaves, Gisele L. Pappa, Omar P. Vilela Neto
CEC4
2018 Solving the exponential growth of symbolic regression trees in geometric semantic genetic programming
abstract
Advances in Geometric Semantic Genetic Programming (GSGP) have shown that this variant of Genetic Programming (GP) reaches better results than its predecessor for supervised machine learning problems, particularly in the task of symbolic regression. However, by construction, the geometric semantic crossover operator generates individuals that grow exponentially with the number of generations, resulting in solutions with limited use. This paper presents a new method for individual simplification named GSGP with Reduced trees (GSGP-Red). GSGP-Red works by expanding the functions generated by the geometric semantic operators. The resulting expanded function is guaranteed to be a linear combination that, in a second step, has its repeated structures and respective coefficients aggregated. Experiments in 12 real-world datasets show that it is not only possible to create smaller and completely equivalent individuals in competitive computational time, but also to reduce the number of nodes composing them by 58 orders of magnitude, on average.
Joao Francisco B. S. Martins, Luiz Otávio Vilas Boas Oliveira, Luis Fernando Miranda, Felipe Casadei, Gisele L. Pappa
GECCO5
2018 An Approximative Bayes-Optimal Kernel Classifier Based on Version Space Reduction
abstract
The Bayes-optimal classifier is defined as a classifier that induces an hypothesis able to minimize the prediction error for any given sample in binary classification problems. Finding the Bayes-optimal classifier is an intractable problem. It is known that it is approximately equivalent to the center of mass of the version space, which is given by the set of all classifiers consistent with the training set. Previously solutions to find the center of mass are not feasible, as they present a high computational cost, and do not work properly in non-linear separable problems. Aiming to solve these problems, this paper presents the Dual Version Space Reduction Machine (Dual VSRM), an effective kernel method to approximate the center of mass of the version space. The Dual VSRM algorithm employs successive reductions of the version space based on an oracle's decision. As an oracle, we propose the Ensemble of Dissimilar Balanced Kernel Perceptrons (EBPK). EBPK enhances the accuracy of each individual classifier by balancing the final hyperplane solution while maximizing the diversity of its components by applying a dissimilarity measure. In order to evaluate the proposed methods, we conduct an experimental evaluation on 7 datasets. We compare the performance of our proposed methods against several baselines. Our results for EBKP indicate the strategies for improving individual accuracy and diversity of the ensemble components work properly. Also, the Dual VSRM consistently outperforms the baselines, indicating that the proposed method generates a better approximation to the center of mass.
Karen Braga Enes, Saulo Moraes Villela, Gisele L. Pappa, Raul Fonseca Neto
ICMLA3
2018 Tutorials at PPSN 2018
Gisele L. Pappa, Michael T. M. Emmerich, Ana L. C. Bazzan, Will N. Browne, Kalyanmoy Deb, Carola Doerr, Marko Durasevic, Michael G. Epitropakis, Saemundur O. Haraldsson, Domagoj Jakobovic, Pascal Kerschke, Krzysztof Krawiec, Per Kristian Lehre, Xiaodong Li 0001, Andrei Lissovoi, Pekka Malo, Luis Martí, Yi Mei 0001, Juan Julián Merelo Guervós, Julian Francis Miller, Alberto Moraglio, Antonio J. Nebro, Su Nguyen, Gabriela Ochoa, Pietro S. Oliveto, Stjepan Picek, Nelishia Pillay, Mike Preuss, Marc Schoenauer, Roman Senkerik, Ankur Sinha 0001, Ofer M. Shir, Dirk Sudholt, L. Darrell Whitley, Mark Wineberg, John R. Woodward, Mengjie Zhang 0001
PPSN (2)1
2018 Automated Selection and Configuration of Multi-Label Classification Algorithms with Grammar-Based Genetic Programming
Alex Guimarães Cardoso de Sá, Alex Alves Freitas, Gisele L. Pappa
PPSN (2)3
2018 Reddit Weight Loss Communities: Do They Have What It Takes for Effective Health Interventions?
abstract
Online social networks are an important tool for people to share information and have been extensively used for people to achieve beneficial changes in health. Obesity is a major public health concern that affects about one third of the world's population. In order to alleviate this problem, health professionals are focusing on health interventions, which can be performed online. In this study we analyze three distinct online communities about weight and diet in Reddit. We model our data as 3 directed and weighted graphs of the posts and comments and evaluate the interaction between users of each community. We also analyze specific characteristics of each community, the habits of daily activity of the users and the formation of implicit bonds of friendship through the formation of communities. Our main results show that Reddit is a content-centered social network, in which what matters is what is posted and not who posts. In addition, users tend to create implicit friendship relationships through denser regions of interactions. Our results show that, contrary to expectations, the three communities present the same behavior pattern in a general point of view, which facilitates the development of non-directed online weight loss intervention strategies.
Karen Braga Enes, Pedro Paulo Valadares Brum, Tiago Oliveira Cunha, Fabricio Murai, Ana Paula Couto da Silva, Gisele L. Pappa
WI6
2018 Selective harvesting over networks
Fabricio Murai, Diogo Rennó, Bruno Ribeiro 0001, Gisele L. Pappa, Don Towsley, Krista Gile
Data Min. Knowl. Discov.4
2018 A customized classification algorithm for credit card fraud detection
Alex Guimarães Cardoso de Sá, Adriano C. M. Pereira, Gisele L. Pappa
Eng. Appl. Artif. Intell.3
2018 Strategies for combining Twitter users geo-location methods
Sílvio S. Ribeiro Jr., Gisele L. Pappa
GeoInformatica2
2017 Strategies for Improving the Distribution of Random Function Outputs in GSGP
Luiz Otávio Vilas Boas Oliveira, Felipe Casadei, Gisele L. Pappa
EuroGP3
2017 RECIPE: A Grammar-Based Framework for Automatically Evolving Classification Pipelines
Alex Guimarães Cardoso de Sá, Walter José G. S. Pinto, Luiz Otávio Vilas Boas Oliveira, Gisele L. Pappa
EuroGP4
2017 How noisy data affects geometric semantic genetic programming
abstract
Noise is a consequence of acquiring and pre-processing data from the environment, and shows fluctuations from different sources---e.g., from sensors, signal processing technology or even human error. As a machine learning technique, Genetic Programming (GP) is not immune to this problem, which the field has frequently addressed. Recently, Geometric Semantic Genetic Programming (GSGP), a semantic-aware branch of GP, has shown robustness and high generalization capability. Researchers believe these characteristics may be associated with a lower sensibility to noisy data. However, there is no systematic study on this matter. This paper performs a deep analysis of the GSGP performance over the presence of noise. Using 15 synthetic datasets where noise can be controlled, we added different ratios of noise to the data and compared the results obtained with those of a canonical GP. The results show that, as we increase the percentage of noisy instances, the generalization performance degradation is more pronounced in GSGP than GP. However, in general, GSGP is more robust to noise than GP in the presence of up to 10% of noise, and presents no statistical difference for values higher than that in the test bed.
Luis Fernando Miranda, Luiz Otávio Vilas Boas Oliveira, Joao Francisco B. S. Martins, Gisele L. Pappa
GECCO4
2017 Top-down strategies for hierarchical classification of transposable elements with neural networks
abstract
Transposable Elements are DNA sequences that can move from one place to another inside the genome of a cell. They are important for genetic variability, and can modify the functionality of genes. The correct classification of these elements is crucial to understand their role in the evolution of species. In this paper, we investigate Transposable Elements classification as a Hierarchical Classification problem using Machine Learning. We present new hierarchical datasets suitable to be used by Machine Learning methods, and also new hierarchical top-down classification strategies using neural networks. We compared our strategies with existing ones in the literature, and evaluated them using measures specific for hierarchical problems. Experiments showed that our proposal achieved better or competitive results than those found by other methods in the literature.
Felipe Kenji Nakano, Walter José G. S. Pinto, Gisele L. Pappa, Ricardo Cerri
IJCNN3
2017 BRACIS 2015: Progress in computation intelligence in Brazil
Gisele L. Pappa, Kate Revoredo, Teresa Bernarda Ludermir
Neurocomputing1
2017 A general framework to expand short text for topic modeling
Paulo Viana Bicalho, Marcelo Pita, Gabriel Pedrosa, Anísio Lacerda, Gisele L. Pappa
Inf. Sci.5
2017 H3AD: A hybrid hyper-heuristic for algorithm design
Péricles B. C. Miranda, Ricardo B. C. Prudêncio, Gisele L. Pappa
Inf. Sci.3
2016 Evolutionary rank aggregation for recommender systems
abstract
Recommender systems are methods built to actively suggest personalized items to users based on their explicit declared preferences (ratings of movies in Netflix), or implicitly observed actions (purchase history). Although a great number of recommendation methods have been previously proposed in the literature, in many problems these methods present a high degree of disagreement in their recommendations. In this scenario, rank aggregation methods are an interesting solution. They can help finding a consensus on which items should be recommended to the user by taking into account the opinion of all available methods. In this direction, this paper proposes ERA (Evolutionary Rank Aggregation), a genetic programming method that outputs an aggregated ranking function built from information extracted from individual input rankings. ERA was tested in four large scale datasets, and obtained better results than other rank aggregation methods in three datasets, improving the results of mean average ranking precision in up to 9.5%.
Samuel E. L. Oliveira, Victor Diniz, Anísio Lacerda, Gisele L. Pappa
CEC4
2016 A Dispersion Operator for Geometric Semantic Genetic Programming
abstract
Recent advances in geometric semantic genetic programming (GSGP) have shown that the results obtained by these methods can outperform those obtained by classical genetic programming algorithms, in particular in the context of symbolic regression. However, there are still many open issues on how to improve their search mechanism. One of these issues is how to get around the fact that the GSGP crossover operator cannot generate solutions that are placed outside the convex hull formed by the individuals of the current population. Although the mutation operator alleviates this problem, we cannot guarantee it will find promising regions of the search space within feasible computational time. In this direction, this paper proposes a new geometric dispersion operator that uses multiplicative factors to move individuals to less dense areas of the search space around the target solution before applying semantic genetic operators. Experiments in sixteen datasets show that the results obtained by the proposed operator are statistically significantly better than those produced by GSGP and that the operator does indeed spread the solutions around the target solution.
Luiz Otávio Vilas Boas Oliveira, Fernando E. B. Otero, Gisele L. Pappa
GECCO3
2016 Discovering Combos in Fighting Games with Evolutionary Algorithms
abstract
In fighting games, players can perform many different actions at each instant of time, leading to an exponential number of possible sequences of actions. Some of these combinations can lead to unexpected behaviors, which can compromise the game design. One example of these unexpected behaviors is the occurrence of long or infinite combos, a long sequence of actions that does not allow any reactions from the opponent. Finding these sequences is essential to ensure fairness in fighting games, but evaluating all possible sequences is a time consuming task. In this paper, we propose the use of an evolutionary algorithm to find combos on a fighting game. The main idea is to use a genetic algorithm to evolve a population composed of sequences of inputs and, using an adequate fitness function, select the ones that are more suitable to be considered combos. We performed a series of experiments and the results show that the proposed approach was not only successful in finding combos, managing to find unexpected sequences, but also superior to previous methods.
Gianlucca L. Zuin, Yuri P. A. Macedo, Luiz Chaimowicz, Gisele L. Pappa
GECCO4
2016 Reducing Dimensionality to Improve Search in Semantic Genetic Programming
Luiz Otávio Vilas Boas Oliveira, Luis Fernando Miranda, Gisele L. Pappa, Fernando E. B. Otero, Ricardo H. C. Takahashi
PPSN3
2016 Exploring multiple evidence to infer users' location in Twitter
Erica C. Rodrigues, Renato Assunção, Gisele L. Pappa, Diogo Rennó, Wagner Meira Jr.
Neurocomputing3
2015 Twitter Population Sample Bias and its impact on predictive outcomes: a case study on elections
abstract
In the past years a lot of effort has been spent analyzing online social network data to understand how the world reality is reflected in the "virtual" world. Twitter is by far the network most used in these studies, given its policy of public data availability. However, a big discussion is still on on whether the data available is enough to make user characterization or event outcomes prediction, and what are the pitfalls people do not usually account for. In this direction, we propose a new methodology for drawing representative samples from Twitter data, which is divided into four phases: (i) user filtering, (ii) user demographic characterization, (iii) user sampling, and (iv) event prediction. The methodology is tested into a common scenario in Twitter event outcome prediction: elections. The methodology was tested with municipality elections from six different Brazilian cities, and compared to official election results. Results show it is worth further investigating the topic, but that a very hight number of messages is required to match real data distributions.
Renato Miranda, Jussara M. Almeida, Gisele L. Pappa
ASONAM3
2015 The Effect of Distinct Geometric Semantic Crossover Operators in Regression Problems
Julio Albinati, Gisele L. Pappa, Fernando E. B. Otero, Luiz Otávio Vilas Boas Oliveira
EuroGP2
2015 GASS: identifying enzyme active sites with genetic algorithms
abstract
MOTIVATION: Currently, 25% of proteins annotated in Pfam have their function unknown. One way of predicting proteins function is by looking at their active site, which has two main parts: the catalytic site and the substrate binding site. The active site is more conserved than the other residues of the protein and can be a rich source of information for protein function prediction. This article presents a new heuristic method, named genetic active site search (GASS), which searches for given active site 3D templates in unknown proteins. The method can perform non-exact amino acid matches (conservative mutations), is able to find amino acids in different chains and does not impose any restrictions on the active site size. RESULTS: GASS results were compared with those catalogued in the catalytic site atlas (CSA) in four different datasets and compared with two other methods: amino acid pattern search for substructures and motif and catalytic site identification. The results show GASS can correctly identify >90% of the templates searched. Experiments were also run using data from the substrate binding sites prediction competition CASP 10, and GASS is ranked fourth among the 18 methods considered.
Sandro C. Izidoro, Raquel Cardoso de Melo Minardi, Gisele L. Pappa
Bioinform.3
2015 An Extensive Evaluation of Decision Tree-Based Hierarchical Multilabel Classification Methods and Performance Measures
abstract
Hierarchical multilabel classification is a complex classification problem where an instance can be assigned to more than one class simultaneously, and these classes are hierarchically organized with superclasses and subclasses, that is, an instance can be classified as belonging to more than one path in the hierarchical structure. This article experimentally analyses the behavior of different decision tree–based hierarchical multilabel classification methods based on the local and global classification approaches. The approaches are compared using distinct hierarchy‐based and distance‐based evaluation measures, when they are applied to a variation of real multilabel and hierarchical datasets' characteristics. Also, the different evaluation measures investigated are compared according to their degrees of consistency, discriminancy, and indifferency. As a result of the experimental analysis, we recommend the use of the global classification approach and suggest the use of the Hierarchical Precision and Hierarchical Recall evaluation measures.
Ricardo Cerri, Gisele L. Pappa, André C. P. L. F. de Carvalho, Alex Alves Freitas
Comput. Intell.2
2013 A genetic algorithm for the minimum cost localization problem in wireless sensor networks
abstract
Localization is a paramount concern in wireless sensor networks. Beacon nodes, which have their position defined a priori, might be used in the process, serving as references to find the position of other nodes. Many studies focused on finding the location of as many nodes as possible, given a set of beacons and distance measurements. In this work, we determine the set of beacon nodes in order to localize all nodes in the network. This can reduce the overall cost involved in the network localization process, i.e., reducing the number of nodes in a WSN with GSP. We present a new approach to this problem using Genetic Algorithms. Our simulations results show the efficiency of the proposed approach, which has results up to 50% better than the best greedy algorithm found in the literature.
Angelo Ferreira Assis, Luiz Filipe M. Vieira, Marco Tulio Reis Rodrigues, Gisele L. Pappa
IEEE Congress on Evolutionary Computation4
2013 A new representation for instance-based clonal selection algorithms
abstract
This work borrows the traditional Pittsburgh-style representation from Genetic-Based Machine Learning and evaluates its performance in artificial immune systems (AIS) for classification. Our main goal is to select as few instances as possible to represent the data from the training set without losing accuracy. The new representation is tested in a modified version of a clonal selection algorithm, where the antibodies represent lists of prototypes instead of a single one. The generated method, named Clonal Selection Prototypes Generator, was tested in 10 UCI datasets and compared to other seven methods that execute the same task. Results showed that the proposed method is very good at considering a trade-off between the number of prototypes generated and the accuracy of the system.
Luiz Otávio Vilas Boas Oliveira, Isabela Drummond, Gisele L. Pappa
IEEE Congress on Evolutionary Computation3
2012 A methodology for photometric validation in vehicles visual interactive systems
Alexandre W. C. Faria, David Menotti, Gisele L. Pappa, Daniel S. D. Lara, Arnaldo de Albuquerque Araújo
Expert Syst. Appl.3
2011 HCGA: A genetic algorithm for hierarchical classification
abstract
Hierarchical classification (HC) is a specialization of the well-known flat classification task. The main difference among them is that in HC examples have to be assigned to classes organized in a previously defined class hierarchy, while in traditional flat classification no class order is imposed. There are two main approaches commonly used to tackle HC: the top down or local approach, which is classifier independent, and the big-bang or global approach, which usually is the product of a modification of a well-known flat classifier. Although evolutionary algorithms have been successfully applied to flat classification, they are underexplored in HC. In this direction, this paper pro poses HCGA (Hierarchical Classification Genetic Algorithm), a method that takes both local and global information into account. HCGA uses a top-down approach for building a classification model and also for classifying new examples. This is in contrast with current top-down methods, which make use of this strategy only for test, using flat classifiers for training models. The method was applied to four GPCR (G protein-coupled receptor) activity datasets, obtaining results statistically equal or better than five baseline classifiers run using a top-down approach.
Rafael V. Carvalho, Gustavo Brunoro, Gisele L. Pappa
IEEE Congress on Evolutionary Computation3
2011 A PSO algorithm for improving multi-view classification
abstract
The Multi-view or multi-modality learning approach is becoming popular for providing different representations of a problem from which classifiers can learn from. Examples of these representations are, for instance, sound and image for the case of the video classification problem. The main idea behind multi-view learning is that learning from these representations separately can lead to better gains than merging them into a single dataset. In the same way as ensembles combine results from different classifiers, the outputs given by classifiers in different views have to be combined in order to provide a final class for an example. This paper proposes a PSO algorithm to combine the outputs coming from different views. It also considers that some views may be better at classifying specific classes, and provides weighting schemes for both views and classes. Experiments were performed in two datasets with three views each, and compared with all views in a single dataset, a majority voting scheme and a scheme based on the Dempster-Shafer theory. Experimental results show that the PSO obtains statistically better results than the other approaches evaluated.
Zilton Cordeiro Junior, Gisele L. Pappa
IEEE Congress on Evolutionary Computation2
2011 Assessing documents' credibility with genetic programming
abstract
The concept of example credibility evaluates how much a classifier can trust an example when building a classification model. It is given by a credibility function, which is application dependent and estimated according to a series of factors that influence the credibility of the examples. Here we deal with automatic document classification and study the credibility of a document according to three factors: content, authorship and citations. We propose a genetic programming algorithm to estimate the credibility of training examples, and then add this estimation to a credibility-aware classifier. For that, we model the authorship and citation data as a complex network, and select a set of structural metrics that can be used to estimate credibility. These metrics are then merged with other content-related ones, and used as terminals for the GP. The GP was tested in a subset of the ACM-DL, and results showed that the credibility-aware classifier obtained results of micro and macroF1from 5% to 8% better than the traditional classifiers.
João R. M. Palotti, Thiago Salles, Gisele L. Pappa, Marcos André Gonçalves, Wagner Meira Jr.
IEEE Congress on Evolutionary Computation3
2011 Semi-supervised genetic programming for classification
abstract
Learning from unlabeled data provides innumerable advantages to a wide range of applications where there is a huge amount of unlabeled data freely available. Semi-supervised learning, which builds models from a small set of labeled examples and a potential large set of unlabeled examples, is a paradigm that may effectively use those unlabeled data. Here we propose KGP, a semi-supervised transductive genetic programming algorithm for classification. Apart from being one of the first semi-supervised algorithms, it is transductive (instead of inductive), i.e., it requires only a training dataset with labeled and unlabeled examples, which should represent the complete data domain. The algorithm relies on the three main assumptions on which semi-supervised algorithms are built, and performs both global search on labeled instances and local search on unlabeled instances. Periodically, unlabeled examples are moved to the labeled set after a weighted voting process performed by a committee. Results on eight UCI datasets were compared with Self-Training and KNN, and showed KGP as a promising method for semi-supervised learning.
Filipe de Lima Arcanjo, Gisele L. Pappa, Paulo Viana Bicalho, Wagner Meira Jr., Altigran S. da Silva
GECCO2
2010 An evolutionary algorithm to the Density Control, Coverage and Routing Multi-Period Problem in Wireless Sensor Networks
abstract
Wireless Sensor Networks (WSNs) are composed of autonomous and resource-constrained (power, sensing, radios, and processors) nodes. These networks are conceived to have a large number of nodes working on monitoring phenomena. A major challenge for theses networks is to provide solutions that maximize quality of service (QoS) requirements, such as coverage and data routing, and minimize the energy consumption. This paper presents an evolutionary algorithm (EA) to solve the Density Control, Coverage and Routing Multi-Period Problem (DCCRMP) in WSN. The results are compared to the optimal solutions obtained by an Integer Linear Program model and a GRASP heuristic from the literature. The EA obtains significant improvements in both quality of solutions and computational time.
Iuri Bueno Drumond de Andrade, Tiago O. Januario, Gisele L. Pappa, Geraldo Robson Mateus
IEEE Congress on Evolutionary Computation3
2010 Active Learning Genetic programming for record deduplication
abstract
The great majority of genetic programming (GP) algorithms that deal with the classification problem follow a supervised approach, i.e., they consider that all fitness cases available to evaluate their models are labeled. However, in certain application domains, a lot of human effort is required to label training data, and methods following a semi-supervised approach might be more appropriate. This is because they significantly reduce the time required for data labeling while maintaining acceptable accuracy rates. This paper presents the Active Learning GP (AGP), a semi-supervised GP, and instantiates it for the data deduplication problem. AGP uses an active learning approach in which a committee of multi-attribute functions votes for classifying record pairs as duplicates or not. When the committee majority voting is not enough to predict the class of the data pairs, a user is called to solve the conflict. The method was applied to three datasets and compared to two other deduplication methods. Results show that AGP guarantees the quality of the deduplication while reducing the number of labeled examples needed.
Junio de Freitas, Gisele L. Pappa, Altigran S. da Silva, Marcos André Gonçalves, Edleno Silva de Moura, Adriano Veloso, Alberto H. F. Laender, Moisés G. de Carvalho
IEEE Congress on Evolutionary Computation2
2010 Tuning Genetic Programming parameters with factorial designs
abstract
Parameter setting of Evolutionary Algorithms is a time consuming task with two main approaches: parameter tuning and parameter control. In this work we describe a new methodology for tuning parameters of Genetic Programming algorithms using factorial designs, one-factor designs and multiple linear regression. Our experiments show that factorial designs can be used to determine which parameters have the largest effect on the algorithm's performance. This way, parameter setting efforts can focus on them, largely reducing the parameter search space. Two classical GP problems were studied, with six parameters for the first problem and seven for the second. The results show the maximum tree depth as the parameter with the largest effect on both problems. A one-factor design was performed to fine-tune tree depth on the first problem and a multiple linear regression to fine-tune tree depth and number of generations on the second.
Elisa Boari de Lima, Gisele L. Pappa, Jussara M. Almeida, Marcos André Gonçalves, Wagner Meira Jr.
IEEE Congress on Evolutionary Computation2
2010 Exploiting co-occurrence and information quality metrics to recommend tags in web 2.0 applications
abstract
This work addresses the task of recommending high quality tags by exploiting not only previously assigned tags, but also terms extracted from other textual features (e.g., title and description) associated with the target object.To estimate the quality of a candidate tag recommendation, we use several metrics related to both tag co-occurrence and information quality. We also propose a heuristic function to combine the metrics to produce a final ranking of the recommended tags. We evaluate our heuristic function in various scenarios, for three popular Web 2.0 applications. Our experimental results indicate that our heuristic function significantly outperforms two state-of-the-art tag recommendation algorithms.
Fabiano Muniz Belém, Eder Ferreira Martins, Jussara M. Almeida, Marcos André Gonçalves, Gisele L. Pappa
CIKM5
2010 Adaptive Normalization: A novel data normalization approach for non-stationary time series
abstract
Data normalization is a fundamental preprocessing step for mining and learning from data. However, finding an appropriated method to deal with time series normalization is not a simple task. This is because most of the traditional normalization methods make assumptions that do not hold for most time series. The first assumption is that all time series are stationary, i.e., their statistical properties, such as mean and standard deviation, do not change over time. The second assumption is that the volatility of the time series is considered uniform. None of the methods currently available in the literature address these issues. This paper proposes a new method for normalizing non-stationary heteroscedastic (with non-uniform volatility) time series. The method, named Adaptive Normalization (AN), was tested together with an Artificial Neural Network (ANN) in three forecast problems. The results were compared to other four traditional normalization methods, and showed AN improves ANN accuracy in both short- and long-term predictions.
Eduardo S. Ogasawara, Leonardo C. Martinez, Daniel de Oliveira 0001, Geraldo Zimbrão, Gisele L. Pappa, Marta Mattoso
IJCNN5
2010 Demand-Driven Tag Recommendation
Guilherme Vale Menezes, Jussara M. Almeida, Fabiano Muniz Belém, Marcos André Gonçalves, Anísio Lacerda, Edleno Silva de Moura, Gisele L. Pappa, Adriano Veloso, Nivio Ziviani
ECML/PKDD (2)7
2010 Temporally-aware algorithms for document classification
abstract
Automatic Document Classification (ADC) is still one of the major information retrieval problems. It usually employs a supervised learning strategy, where we first build a classification model using pre-classified documents and then use this model to classify unseen documents. The majority of supervised algorithms consider that all documents provide equally important information. However, in practice, a document may be considered more or less important to build the classification model according to several factors, such as its timeliness, the venue where it was published in, its authors, among others. In this paper, we are particularly concerned with the impact that temporal effects may have on ADC and how to minimize such impact. In order to deal with these effects, we introduce a temporal weighting function (TWF) and propose a methodology to determine it for document collections. We applied the proposed methodology to ACM-DL and Medline and found that the TWF of both follows a lognormal. We then extend three ADC algorithms (namely kNN, Rocchio and Naïve Bayes) to incorporate the TWF. Experiments showed that the temporally-aware classifiers achieved significant gains, outperforming (or at least matching) state-of-the-art algorithms.
Thiago Salles, Leonardo Rocha 0001, Gisele L. Pappa, Fernando Mourão, Wagner Meira Jr., Marcos André Gonçalves
SIGIR3
2009 From an artificial neural network to a stock market day-trading system: A case study on the BM&F BOVESPA
abstract
Predicting trends in the stock market is a subject of major interest for both scholars and financial analysts. The main difficulties of this problem are related to the dynamic, complex, evolutive and chaotic nature of the markets. In order to tackle these problems, this work proposes a day-trading system that “translates” the outputs of an artificial neural network into business decisions, pointing out to the investors the best times to trade and make profits. The ANN forecasts the lowest and highest stock prices of the current trading day. The system was tested with the two main stocks of the BM&FBOVESPA, an important and understudied market. A series of experiments were performed using different data input configurations, and compared with four benchmarks. The results were evaluated using both classical evaluation metrics, such as the ANN generalization error, and more general metrics, such as the annualized return. The ANN showed to be more accurate and give more return to the investor than the four benchmarks. The best results obtained by the ANN had an mean absolute percentage error around 50% smaller than the best benchmark, and doubled the capital of the investor.
Leonardo C. Martinez, Diego N. da Hora, João R. M. Palotti, Wagner Meira Jr., Gisele L. Pappa
IJCNN5
2009 Automatically evolving rule induction algorithms tailored to the prediction of postsynaptic activity in proteins
abstract
It is well-known that no classification algorithm is the best in all application domains. The conventional approach for coping with this problem consists of trying to select the best classification algorithm for the target application domain. We prop
Gisele L. Pappa, Alex Alves Freitas
Intell. Data Anal.1
2009 Evolving rule induction algorithms with multi-objective grammar-based genetic programming
Gisele L. Pappa, Alex Alves Freitas
Knowl. Inf. Syst.1
2006 Automatically Evolving Rule Induction Algorithms
Gisele L. Pappa, Alex Alves Freitas
ECML1