EDBT 2026 Demo / reviewers in the wild / expert
Ana Carolina Lorena
dblp:97/2733
· DBLP profile ↗
55ranked-venue papers
12as first author
19since 2021 · last 2026
0000-0002-6140-571XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 12 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Security and privacy · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring the influence of missing data imputation in group fairness metricsabstractMissing data is a common problem in real-world datasets and can be characterized as the lack of information on one or multiple variables in a dataset. The most frequent technique for handling this issue is imputation, which consists in the replacement of the missing values according to a predefined criterion. Since missing values are often imputed based on the known values in the dataset, existing data issues can be propagated during the imputation process. One such issue is fairness, a concept integral to responsible Artificial Intelligence practices. This work investigates the impact of the imputation process on system fairness by examining how imputation affects the fairness of predictions in Machine Learning models. It provides a comprehensive analysis covering thirteen unfair benchmark datasets with six state-of-the-art imputation strategies under synthetic Missing Not At Random and Missing At Random mechanisms in a multivariate scenario with 10%, 20%, 40%, and 60% of missing rates. Fairness was measured by the following metrics: Statistical Parity, Equalized Odds, Equality of Opportunity, Predictive Equality, Equality of Positive, and Negative Predicted Values. The results demonstrate that the missing mechanism, the classifier choice, and the imputation strategy decisively influence the fairness of the predictions obtained by the Machine Learning models. Arthur Dantas Mangussi, Ricardo Cardoso Pereira, Miriam Seoane Santos, Ana Carolina Lorena, Mykola Pechenizkiy, Pedro H. Abreu |
Artif. Intell. | 4 |
| 2026 | pycol-vis: A Python package for image complexity assessmentabstractDataset complexity poses a significant challenge in classification tasks, especially in real-world applications where a combination of factors such as class overlap, data imbalance, noise, and dimensionality can jeopardize a machine learning algorithm’s performance. While measures to quantify complexity have been proposed and studied in depth for tabular datasets, there is a lack of studies and toolkits focused on measuring complexity in non-structured image data. This limitation hinders our understanding of visual complexity, despite the importance of image data in fields such as healthcare, remote sensing, and autonomous navigation. To address this challenge, we introduce pycol-vis, a novel Python package that helps researchers estimate image complexity. The package implements 17 image complexity measures, specifically designed to capture complexity in real-world scenarios. This toolkit is essential for researchers dealing with complex classification problems in the vision domain, providing tools to assess difficulty and overlap in real-world image data. Diogo Apóstolo, Miriam Seoane Santos, Ana Carolina Lorena, Nathalie Japkowicz, Pedro H. Abreu |
Neurocomputing | 3 |
| 2026 | Filtering Instances and Rejecting Predictions to Obtain Reliable Models in Healthcare
Maria Gabriela Valeriano, David Kohan Marzagão, Alfredo Montelongo, Carlos Roberto Veiga Kiffer, Natan Katz, Ana Carolina Lorena |
Mach. Learn. | 6 |
| 2025 | Optimizing view generation in classification datasetsabstractThe divide-and-conquer strategy is a common approach for solving Computer Science problems, where one divides a problem into multiple potentially simpler subproblems whose results are combined. This decomposition approach can also help solve classification problems in Machine Learning (ML). In this paper, we propose a decomposition strategy to generate multiple views from a dataset in an optimized way. The objective is to divide the input features into subsets and build ML models for each. When a new instance has to be classified, the view for which the complexity in classifying the instance is lower is chosen to be used in prediction. We show experimentally how this approach can benefit the nearest neighbor classifier, increasing classification accuracy for complex problems while overcoming the limitations of this method in handling directly high-dimensional data. Victor Castro Nacif De Faria, Rafael M. O. Cruz, Robert Sabourin, Ana Carolina Lorena |
IJCNN | 4 |
| 2025 | Beyond Filtering: Leveraging Instance Hardness for Data-Centric Machine Learning in HealthcareabstractMachine learning models often struggle to generalize to real-world healthcare settings due to data imperfections such as noise, biases, and label inconsistencies. A promising approach to improving model robustness is instance hardness analysis, which quantifies the difficulty of classifying individual data points in a dataset. While previous studies have explored filtering hard instances to enhance predictive performance, our findings indicate that this strategy can lead to overfitting and reduced generalization on external datasets. This work systematically evaluates the effects of filtering hard instances across multiple healthcare-related datasets, revealing the nuanced relationship between data complexity and model performance. Beyond filtering, we propose alternative data-centric strategies, including expert-guided instance selection, explainability-driven insights, and fairness-aware approaches, to ensure that machine learning models remain both accurate and equitable. Our results highlight the need for a holistic view of data quality, moving beyond naive filtering approaches to leverage instance hardness as a tool for dataset refinement and model improvement. Maria Gabriela Valeriano, Carlos Roberto Veiga Kiffer, Ana Carolina Lorena |
IJCNN | 3 |
| 2025 | Studying the robustness of data imputation methodologies against adversarial attacksabstractCybersecurity attacks, such as poisoning and evasion, can intentionally introduce false or misleading information in different forms into data, potentially leading to catastrophic consequences for critical infrastructures, like water supply or energy power plants. While numerous studies have investigated the impact of these attacks on model-based prediction approaches, they often overlook the impurities present in the data used to train these models. One of those forms is missing data, the absence of values in one or more features. This issue is typically addressed by imputing missing values with plausible estimates, which directly impacts the performance of the classifier. The goal of this work is to promote a Data-centric AI approach by investigating how different types of cybersecurity attacks impact the imputation process. To this end, we conducted experiments using four popular evasion and poisoning attacks strategies across 29 real-world datasets, including the NSL-KDD and Edge-IIoT datasets, which were used as case study. For the adversarial attack strategies, we employed the Fast Gradient Sign Method, Carlini & Wagner, Project Gradient Descent, and Poison Attack against Support Vector Machine algorithm. Also, four state-of-the-art imputation strategies were tested under Missing Not At Random, Missing Completely at Random, and Missing At Random mechanisms using three missing rates (5%, 20%, 40%). We assessed imputation quality using MAE, while data distribution shifts were analyzed with the Kolmogorov–Smirnov and Chi-square tests. Furthermore, we measured classification performance by training an XGBoost classifier on the imputed datasets, using F1-score, Accuracy, and AUC. To deepen our analysis, we also incorporated six complexity metrics to characterize how adversarial attacks and imputation strategies impact dataset complexity. Our findings demonstrate that adversarial attacks significantly impact the imputation process. In terms of imputation assessment in what concerns to quality error, the scenario that enrolees imputation with Project Gradient Descent attack proved to be more robust in comparison to other adversarial methods. Regarding data distribution error, results from the Kolmogorov–Smirnov test indicate that in the context of numerical features, all imputation strategies differ from the baseline (without missing data) however for the categorical context Chi-Squared test proved no difference between imputation and the baseline. Arthur Dantas Mangussi, Ricardo Cardoso Pereira, Ana Carolina Lorena, Miriam Seoane Santos, Pedro H. Abreu |
Comput. Secur. | 3 |
| 2025 | Pycol: A Python package for dataset complexity measuresabstractClass overlap presents a significant challenge to machine learning algorithms, especially when class imbalance is present. These factors contribute substantially to the complexity of classification tasks, particularly in real-world scenarios. As a result, measuring overlap is crucial, yet it remains difficult to quantify due to its intricate nature, since it can manifest and be measured in multiple ways. To help mitigate this, recent research has conceptualized a new taxonomy of class overlap measures, divided into multiple families, which allows researchers to obtain a more complete overview of the complexity of the datasets. In line with recent research, we introduce a new Python package for class overlap measurement named pycol. This package implements 29 overlap measures, divided into four overlap families specifically designed to capture class overlap in imbalanced real-world scenarios. This makes pycol an essential tool for researchers dealing with complex classification problems, providing robust solutions to quantify the joint-effect of class overlap and class imbalance effectively. Diogo Apóstolo, Miriam Seoane Santos, Ana Carolina Lorena, Pedro H. Abreu |
Neurocomputing | 3 |
| 2025 | mdatagen: A python library for the artificial generation of missing data
Arthur Dantas Mangussi, Miriam Seoane Santos, Filipe Loyola Lopes, Ricardo Cardoso Pereira, Ana Carolina Lorena, Pedro H. Abreu |
Neurocomputing | 5 |
| 2025 | Novel applications of item response theory for analysing data set complexity and benchmark selection
João Luiz Junho Pereira, Alfredo Antonio Alencar Exposito de Queiroz, Telmo de Menezes e Silva Filho, Ana Carolina Lorena, Rafael Gomes Mantovani, Gisele L. Pappa, Ricardo B. C. Prudêncio |
Mach. Learn. | 4 |
| 2024 | Measuring Latent Traits of Instance Hardness and Classifier Ability using Boltzmann MachinesabstractTraditional Machine Learning (ML) approaches often emphasize evaluating models using global metrics over a dataset, frequently overlooking the nuances of learning data. Analyzing how hard it is to classify each instance, also known as instance hardness, furnishes such information, offering insights into reasons behind particular misclassifications. This paper introduces an unsupervised Deep Boltzmann Machine model integrated with an interpretability module that provides various latent traits related to instance hardness and classifier predictive performance. Such knowledge can facilitate in-depth analyses of the learning dataset's instances and the predictive power or ability of ML algorithms. Herein, we illustrate our approach by assessing five datasets with over 230 learning algorithms. Eduardo Vargas Ferreira, Ricardo B. C. Prudêncio, Ana Carolina Lorena |
IJCNN | 3 |
| 2024 | Assessor Models for Explaining Instance Hardness in Classification ProblemsabstractUnderstanding the difficulty of individual instances in a classification problem is important to define the limits of learning performance in the problem. Previous works are devoted to measuring Instance Hardness (IH), while solutions for explaining IH are still not deeply investigated. In this paper, we rely on using assessor models and eXplanaible AI (XAI) techniques to predict and explain IH. Many XAI techniques have been developed in the literature to explain the predictions of Machine Learning (ML) models. In our work, we are focused on explaining the difficulty of instances. Given a classification dataset, we trained and evaluated a pool of diverse ML models to measure the IH of each instance. Then, we trained an assessor model to predict the IH based on the instances’ features. Once the assessor is built, its predictions (i.e., the expected IH) can be explained using XAI techniques. In our experiments, we produced Partial Dependence Plots (PDP) to inspect the marginal effect of specific features on the IH predicted by the assessor. From the PDPs, we could check how IH is distributed along the instances’ features in a problem, and more specifically, we could visualize areas of high expected predictive difficulty. Ricardo B. C. Prudêncio, Ana Carolina Lorena, Telmo de Menezes e Silva Filho, Patrícia Drapal, Maria Gabriela Valeriano |
IJCNN | 2 |
| 2024 | Optimal selection of benchmarking datasets for unbiased machine learning algorithm evaluation
João Luiz Junho Pereira, Kate Smith-Miles, Mario A. Muñoz, Ana Carolina Lorena |
Data Min. Knowl. Discov. | 4 |
| 2024 | Golden lichtenberg algorithm: a fibonacci sequence approach applied to feature selection
João Luiz Junho Pereira, Matheus Brendon Francisco, Benedict Jun Ma, Guilherme Ferreira Gomes, Ana Carolina Lorena |
Neural Comput. Appl. | 5 |
| 2023 | Complexity-Driven Sampling for Bagging
Carmen Lancho, Marcílio Carlos Pereira de Souto, Ana Carolina Lorena, Isaac Martín de Diego |
IDEAL | 3 |
| 2022 | Let the data speak: analysing data from multiple health centers of the São Paulo metropolitan area for COVID-19 clinical deterioration predictionabstractWith the spread of different COVID-19 variants in the Brazilian territory, the national health system has been facing a constant overload. Using data from five different health centers located in the Sao Paulo metropolitan area, this work seeks to identify key common factors associated with the prognosis of COVID-19 severity. The proxies for severity considered are hospitalization time, death and use of mechanical ventilation. The induced models predicted ob-jective short-term COVID-19 clinical deterioration outcomes with AUC, sensitivity and specificity up to 0.880, 0.824 and 0.833, respectively. Parameters such as C-reactive protein and percentage of neutrophils have shown most influence on the predictions. Given the nature of the lab tests highlighted, we note that innate inflammatory status in admission can play a significant role in patient outcome. Maria Gabriela Valeriano, Carlos Roberto Veiga Kiffer, Giane Higino, Paloma Zanão, Dulce A. Barbosa, Patrícia A. Moreira, Paulo Caleb J. L. Santos, Renato Grinbaum, Ana Carolina Lorena |
CCGRID | 9 |
| 2022 | Relating instance hardness to classification performance in a dataset: a visual approachabstractMachine Learning studies often involve a series of computational experiments in which the predictive performance of multiple models are compared across one or more datasets. The results obtained are usually summarized through average statistics, either in numeric tables or simple plots. Such approaches fail to reveal interesting subtleties about algorithmic performance, including which observations an algorithm may find easy or hard to classify, and also which observations within a dataset may present unique challenges. Recently, a methodology known as Instance Space Analysis was proposed for visualizing algorithm performance across different datasets. This methodology relates predictive performance to estimated instance hardness measures extracted from the datasets. However, the analysis considered an instance as being an entire classification dataset and the algorithm performance was reported for each dataset as an average error across all observations in the dataset. In this paper, we developed a more fine-grained analysis by adapting the ISA methodology. The adapted version of ISA allows the analysis of an individual classification dataset by a 2-D hardness embedding, which provides a visualization of the data according to the difficulty level of its individual observations. This allows deeper analyses of the relationships between instance hardness and predictive performance of classifiers. We also provide an open-access Python package named PyHard, which encapsulates the adapted ISA and provides an interactive visualization interface. We illustrate through case studies how our tool can provide insights about data quality and algorithm performance in the presence of challenges such as noisy and biased data. Pedro Yuri Arbs Paiva, Camila Castro Moreno, Kate Smith-Miles, Maria Gabriela Valeriano, Ana Carolina Lorena |
Mach. Learn. | 5 |
| 2021 | Evaluating Data Characterization Measures for Clustering Problems in Meta-learning
Luiz Henrique dos Santos Fernandes, Marcílio Carlos Pereira de Souto, Ana Carolina Lorena |
ICONIP (1) | 3 |
| 2021 | Assessing the data complexity of imbalanced datasets
Victor H. Barella, Luís Paulo F. Garcia, Marcílio Carlos Pereira de Souto, Ana Carolina Lorena, André C. P. L. F. de Carvalho |
Inf. Sci. | 4 |
| 2021 | An Instance Space Analysis of Regression ProblemsabstractThe quest for greater insights into algorithm strengths and weaknesses, as revealed when studying algorithm performance on large collections of test problems, is supported by interactive visual analytics tools. A recent advance is Instance Space Analysis, which presents a visualization of the space occupied by the test datasets, and the performance of algorithms across the instance space. The strengths and weaknesses of algorithms can be visually assessed, and the adequacy of the test datasets can be scrutinized through visual analytics. This article presents the first Instance Space Analysis of regression problems in Machine Learning, considering the performance of 14 popular algorithms on 4,855 test datasets from a variety of sources. The two-dimensional instance space is defined by measurable characteristics of regression problems, selected from over 26 candidate features. It enables the similarities and differences between test instances to be visualized, along with the predictive performance of regression algorithms across the entire instance space. The purpose of creating this framework for visual analysis of an instance space is twofold: one may assess the capability and suitability of various regression techniques; meanwhile the bias, diversity, and level of difficulty of the regression problems popularly used by the community can be visually revealed. This article shows the applicability of the created regression instance space to provide insights into the strengths and weaknesses of regression algorithms, and the opportunities to diversify the benchmark test instances to support greater insights. Mario A. Muñoz, Matheus R. Leal, Kate Smith-Miles, Ana Carolina Lorena, Gisele L. Pappa, Rômulo Madureira Rodrigues |
ACM Trans. Knowl. Discov. Data | 5 |
| 2020 | Monitoring Night Skies with Deep Learning
Yuri Galindo, Marcelo De Cicco, Marcos G. Quiles, Ana Carolina Lorena |
ICONIP (4) | 4 |
| 2020 | Boosting meta-learning with simulated data complexity measuresabstractMeta-Learning has been largely used over the last years to support the recommendation of the most suitable machine learning algorithm(s) and hyperparameters for new datasets. Traditionally, a meta-base is created containing meta-features extracted from several datasets along with the performance of a pool of machine learning algorithms when applied to these datasets. The meta-features must describe essential aspects of the dataset and distinguish different problems and solutions. However, if one wants the use of Meta-Learning to be computationally efficient, the extraction of the meta-feature values should also show a low computational cost, considering a trade-off between the time spent to run all the algorithms and the time required to extract the meta-features. One class of measures with successful results in the characterization of classification datasets is concerned with estimating the underlying complexity of the classification problem. These data complexity measures take into account the overlap between classes imposed by the feature values, the separability of the classes and distribution of the instances within the classes. However, the extraction of these measures from datasets usually presents a high computational cost. In this paper, we propose an empirical approach designed to decrease the computational cost of computing the data complexity measures, while still keeping their descriptive ability. The proposal consists of a novel Meta-Learning system able to predict the values of the data complexity measures for a dataset by using simpler meta-features as input. In an extensive set of experiments, we show that the predictive performance achieved by Meta-Learning systems which use the predicted data complexity measures is similar to the performance obtained using the original data complexity measures, but the computational cost involved in their computation is significantly reduced. Luís Paulo F. Garcia, Adriano Rivolli, Edesio Alcobaça, Ana Carolina Lorena, André C. P. L. F. de Carvalho |
Intell. Data Anal. | 4 |
| 2019 | Data complexity measures in feature selectionabstractFeature selection (FS) is a pre-processing step often mandatory in data analysis by Machine Learning techniques. Its objective is to reduce data dimensionality by identifying and maintaining only the relevant features from a dataset. In this work we evaluate the use of complexity measures of classification problems in FS. These descriptors allow estimating the intrinsic difficulty of a classification problem by regarding on characteristics of the dataset available for learning. We propose a combined univariate-multivariate FS technique which employs two complexity measures: Fisher's maximum discriminant ratio and sum of intra-extra class distances. The results reveal that the complexity measures are indeed suitable for estimating feature importance in classification datasets. Large reductions in the numbers of features were obtained, while preserving, in general, the predictive accuracy of two strong classification techniques: Support Vector Machines and Random Forests. Lucas Chesini Okimoto, Ana Carolina Lorena |
IJCNN | 2 |
| 2019 | New label noise injection methods for the evaluation of noise filtersabstractNoise is often present in real datasets used for training Machine Learning classifiers. Their disruptive effects in the learning process may include: increasing the complexity of the induced models, a higher processing time and a reduced predictive power in the classification of new examples. Therefore, treating noisy data in a preprocessing step is crucial for improving data quality and to reduce their harmful effects in the learning process. There are various filters using different concepts for identifying noisy examples in a dataset. Their ability in noise preprocessing is usually assessed in the identification of artificial noise injected into one or more datasets. This is performed to overcome the limitation that only a domain expert can guarantee whether a real example is indeed noisy. The most frequently used label noise injection method is the noise at random method, in which a percentage of the training examples have their labels randomly exchanged. This is carried out regardless of the characteristics and example space positions of the selected examples. This paper proposes two novel methods to inject label noise in classification datasets. These methods, based on complexity measures, can produce more challenging and realistic noisy datasets by the disturbance of the labels of critical examples situated close to the decision borders and can improve the noise filtering evaluation. An extensive experimental evaluation of different noise filters is performed using public datasets with imputed label noise and the influence of the noise injection methods are compared in both data preprocessing and classification steps. Luís Paulo F. Garcia, Jens Lehmann 0001, André C. P. L. F. de Carvalho, Ana Carolina Lorena |
Knowl. Based Syst. | 4 |
| 2018 | Automatic Design of Evolutionary Algorithms Based on Entropy TriggersabstractThe field of automatic algorithm design has received increasing attention in recent years. From a multitude of available algorithms, a researcher can effectively design a new one customized to his/her own problem. For this, hyper-heuristics techniques have proven to be useful. Their main objective is to search in the space of heuristics rather than in the problem solution space. The present paper proposes a hyper-heuristic for the automatic design of evolutionary algorithms supported by the use of an entropy metric. This metric is used as a trigger mechanism for switching between the algorithms components, aiding the formation of the new hybrid algorithm. Guilherme Ribeiroda Silva, Márcio P. Basgalupp, Ana Carolina Lorena |
CEC | 3 |
| 2018 | Classifier Recommendation Using Data Complexity MeasuresabstractApplication of machine learning to new and unfamiliar domains calls for increasing automation in choosing a learning algorithm suitable for the data arising from each domain. Meta-learning could address this need since it has been largely used in the last years to support the recommendation of the most suitable algorithms for a new dataset. The use of complexity measures could increase the systematic comprehension over the meta-models and also allow to differentiate the performance of a set of techniques taking into account the overlap between classes imposed by feature values, the separability and distribution of the data points. In this paper we compare the effectiveness of several standard regression models in predicting the accuracies of classifiers for classification problems from the OpenML repository. We show that the models can predict the classifiers' accuracies with low mean-squared-error and identify the best classifier for a problem that results in statistically significant improvements over a randomly chosen classifier or a fixed classifier believed to be good on average. Luís Paulo F. Garcia, Ana Carolina Lorena, Marcílio Carlos Pereira de Souto, Tin Kam Ho |
ICPR | 2 |
| 2018 | Data Complexity Measures for Imbalanced Classification TasksabstractIn imbalanced classification tasks, the training datasets may show class overlapping and classes of low density. In these scenarios, the predictions for the minority class are impaired. Although assessing the imbalance level of a training set is straightforward, it is hard to measure other aspects that may affect the predictive performance of classification algorithms in imbalanced tasks. This paper presents a set of measures designed to understand the difficulty of imbalanced classification tasks by regarding on each class individually. They are adapted from popular data complexity measures for classification problems, which are shown to perform poorly in imbalanced scenarios. Experiments on synthetic datasets with different levels of imbalance, class overlapping and density of the classes show that the proposed adaptations can better explain the difficulty of imbalanced classification tasks. Victor H. Barella, Luís Paulo F. Garcia, Marcílio Carlos Pereira de Souto, Ana Carolina Lorena, André C. P. L. F. de Carvalho |
IJCNN | 4 |
| 2018 | Using Complexity Measures to Evolve Synthetic Classification DatasetsabstractMachine Learning studies usually involve a large volume of experimental work. For instance, any new technique or solution to a classification problem has to be evaluated concerning the predictive performance achieved in many datasets. In order to evaluate the robustness of the algorithm face to different class distributions, it would be interesting to choose a set of datasets that spans different levels of classification difficulty. In this paper, we present a method to generate synthetic classification datasets with varying complexity levels. The idea is to greedly exchange the labeling of a set of synthetically generated points in order to reach a given level of classification complexity, which is assessed by measures that estimate the difficulty of a classification problem based on the geometrical distribution of the data. Vinícius Veloso de Melo, Ana Carolina Lorena |
IJCNN | 2 |
| 2018 | Data complexity meta-features for regression problems
Ana Carolina Lorena, Aron I. Maciel, Péricles B. C. Miranda, Ivan G. Costa, Ricardo B. C. Prudêncio |
Mach. Learn. | 1 |
| 2017 | Score normalization applied to adaptive biometric systemsabstractBiometric authentication systems have certain limitations. Recent studies have shown that biometric features may change over time, which can entail a decrease in recognition performance of the biometric system. An adaptive biometric system addresses this problem by adapting the biometric reference/template over time, thereby tracking the changes automatically. However, the use of these systems usually requires the adoption of a high threshold value to avoid the inclusion of impostor patterns into the genuine biometric reference. In this study, we hypothesize that score normalization procedures, which have been used to improve the recognition performance of biometric systems through a better refinement of their decision, can also improve the overall performance of adaptive systems. With such a normalization, a better threshold choice could also be made, which would then increase the number of genuine samples used for adaptation. To the best of our knowledge, this is the first investigation towards the use of score normalization to enhance adaptive biometric systems dealing with the change of user features over time. Through a systematic experimental design tested on two behavioral biometric traits, the obtained results indeed support our conjecture. Moreover, the experimental results show that the performance gain brought by adaptation can have a higher overall impact than score normalization alone. (C) 2017 Elsevier Ltd. All rights reserved. Paulo Henrique Pisani, Norman Poh, André C. P. L. F. de Carvalho, Ana Carolina Lorena |
Comput. Secur. | 4 |
| 2017 | Adaptive algorithms applied to accelerometer biometrics in a data stream contextabstractThe use of smartphone devices has increased over the last years, as illustrated by the growth in smartphone sales. These devices are currently used for several services, such as bank account access, social networks and storage of personal information Paulo Henrique Pisani, Ana Carolina Lorena, André C. P. L. F. de Carvalho |
Intell. Data Anal. | 2 |
| 2016 | Measuring the complexity of regression problemsabstractMany works have attempted to characterize the complexity of classification problems by measures extracted from their learning datasets. These indexes provide indicatives of the inherent difficulty in solving a given classification problem. Although regression problems are equally frequent, there is a lack of studies in Machine Learning dedicated to understanding their complexity. This paper proposes some measures aimed to characterize the complexity of regression problems. They are experimentally evaluated on a set of synthetic datasets with different complexities. The results show that various measures and their combinations are able to distinguish simple linear problems from more complex variants. Aron I. Maciel, Ivan G. Costa, Ana Carolina Lorena |
IJCNN | 3 |
| 2016 | Enhanced template update: Application to keystroke dynamicsabstractInternational audience Paulo Henrique Pisani, Romain Giot, André C. P. L. F. de Carvalho, Ana Carolina Lorena |
Comput. Secur. | 4 |
| 2016 | Ensembles of label noise filters: a ranking approach
Luís Paulo F. Garcia, Ana Carolina Lorena, Stan Matwin, André C. P. L. F. de Carvalho |
Data Min. Knowl. Discov. | 2 |
| 2016 | Noise detection in the meta-learning level
Luís Paulo F. Garcia, André C. P. L. F. de Carvalho, Ana Carolina Lorena |
Neurocomputing | 3 |
| 2015 | Using Growing Neural Gas in Prototype Generation for Nearest Neighbor Classifiers
Jussara Dias, Marcos G. Quiles, Ana Carolina Lorena |
ICONIP (2) | 3 |
| 2015 | On Measuring the Complexity of Classification Problems
Ana Carolina Lorena, Marcílio Carlos Pereira de Souto |
ICONIP (1) | 1 |
| 2015 | Adaptive approaches for keystroke dynamicsabstractEnhanced authentication mechanisms are currently needed in several situations. Mainly due to the widespread use of the Internet, data exposure became a source of growing concern. Commonly used login and password credentials may not provide enough security in this scenario, as they may be easily stolen or guessed in some cases. The use of biometrics is a prominent alternative for user authentication, such as by the use of keystroke dynamics. This biometric technology allows the recognition of users by their typing rhythm, which can be performed using data provided by a common keyboard. However, recent work has shown that typing rhythm changes over time. As a result, a static biometric model can become outdated, decreasing the predictive performance of the system. In light of this fact, there is a need for new techniques able to dynamically adapt user models over time. This paper evaluates, in a data stream context, algorithms proposed in the literature for user authentication based on keystroke dynamics. Modifications to these algorithms are also proposed and evaluated. A study of the behaviour of the algorithms over time under several aspects is also performed. According to our experiments, adaptive methods can improve predictive performance of user recognition by keystroke dynamics. Paulo Henrique Pisani, Ana Carolina Lorena, André C. P. L. F. de Carvalho |
IJCNN | 2 |
| 2015 | Effect of label noise in the complexity of classification problems
Luís Paulo F. Garcia, André C. P. L. F. de Carvalho, Ana Carolina Lorena |
Neurocomputing | 3 |
| 2015 | Using the One-vs-One decomposition to improve the performance of class noise filters via an aggregation strategy in multi-class classification problemsabstractNoise filters are preprocessing techniques designed to improve data quality in classification tasks by detecting and eliminating examples that contain errors or noise. However, filtering can also remove correct examples and examples containing valuable information, which could be useful for learning. This fact usually implies a margin of improvement on the noise detection accuracy for almost any noise filter. This paper proposes a scheme to improve the performance of noise filters in multi-class classification problems, based on decomposing the dataset into multiple binary subproblems. Decomposition strategies have proven to be successful in improving classification performance in multi-class problems by generating simpler binary subproblems. Similarly, we adapt the principles of the One-vs-One decomposition strategy to noise filtering, making the noise identification process simpler. In order to integrate the filtering results achieved in the binary subproblems, our proposal uses a soft voting approach considering a reliability level based on the aggregation of the noise degree prediction calculated for each binary classifier. The experimental results show that the One-vs-One decomposition strategy usually increases the performance of the noise filters studied, which can detect more accurately the noisy examples. Luís Paulo F. Garcia, José A. Sáez, Julián Luengo, Ana Carolina Lorena, André C. P. L. F. de Carvalho, Francisco Herrera |
Knowl. Based Syst. | 4 |
| 2014 | Advances in intelligent systems
Teresa Bernarda Ludermir, Cleber Zanchettin, Ana Carolina Lorena |
Neurocomputing | 3 |
| 2012 | Evolutionary neural networks applied to keystroke dynamics: Genetic and immune basedabstractThe evolution in the use of digital identities has brought several advancements. However, this evolution has also contributed to the rise of the identity theft. An alternative to curb identity theft is by the identification of anomalous user behavior on the computer, what is known as behavioral intrusion detection. Among the features to be extracted from the user behavior, this paper focuses on keystroke dynamics, which analysis the user typing rhythm. This work uses a neural network to recognize users by keystroke dynamics and draws a comparison among several training algorithms: single backpropagation, three approaches based on genetic algorithms and three approaches based on immune algorithms. Paulo Henrique Pisani, Ana Carolina Lorena |
IEEE Congress on Evolutionary Computation | 2 |
| 2012 | Analysis of complexity indices for classification problems: Cancer gene expression data
Ana Carolina Lorena, Ivan G. Costa, Newton Spolaôr, Marcílio Carlos Pereira de Souto |
Neurocomputing | 1 |
| 2011 | Multi-objective Genetic Algorithm Evaluation in Feature Selection
Newton Spolaôr, Ana Carolina Lorena, Huei Diana Lee |
EMO | 2 |
| 2011 | Comparing machine learning classifiers in potential distribution modelling
Ana Carolina Lorena, Luís F. O. Jacintho, Marinez Ferreira de Siqueira, Renato De Giovanni, Lúcia G. Lohmann, André C. P. L. F. de Carvalho, Missae Yamamoto |
Expert Syst. Appl. | 1 |
| 2010 | Complexity measures of supervised classifications tasks: A case study for cancer gene expression dataabstractMachine Learning algorithms have been widely used for gene expression data classification, despite the fact that these data have often intrinsic limitations, such as high dimensionality and a small number of examples. Few studies try to characterize to which extent these aspects can influence the performance of the classification models induced. In this paper we compute different measures characterizing the complexity of gene expression data sets for cancer diagnosis. We then investigate how these measures relate to the classification performances achieved by support vector machines, a popular Machine Learning technique usually employed in the analysis of gene expression data. The results obtained indicate that some of the complexity indices utilized are indeed successful in explaining the difficulty involved in the classification of cancer gene expression data. Marcílio Carlos Pereira de Souto, Ana Carolina Lorena, Newton Spolaôr, Ivan G. Costa |
IJCNN | 2 |
| 2010 | Building binary-tree-based multiclass classifiers using separability measures
Ana Carolina Lorena, André C. P. L. F. de Carvalho |
Neurocomputing | 1 |
| 2009 | Evaluation Functions for the Evolutionary Design of Multiclass Support Vector MachinesabstractSupport Vector Machines were originally proposed to solve two-class classification problems. When they are applied to multiclass classification problems, usually a decomposition approach is followed, in which the original multiclass problem is decomposed into multiple binary sub-problems, whose solutions are afterwards combined. There are several strategies to decompose the multiclass problem. Genetic Algorithms (GAs) can be used to optimize the decomposition according to the performance obtained in the overall multiclass problem solution. This paper presents a study on possible evaluation functions that can be used by the GA in order to evaluate a given decomposition. Ana Carolina Lorena, André C. P. L. F. de Carvalho |
Int. J. Comput. Intell. Appl. | 1 |
| 2008 | On the Complexity of Gene Expression Classification Data SetsabstractOne of the main kinds of computational tasks regarding gene expression data is the construction of classifiers (models), often via some machine learning (ML) technique and given data sets, to automatically discriminate expression patterns from cancer (tumor) and normal tissues or from subtypes of cancers. A very distinctive characteristic of these data sets is its high dimensionality and the fewer number of data items. Such a characteristic makes the induction of accurate ML models difficult (e.g., it could lead to model overfitting). In this context, we present an empirical study on the complexity of the classification task of gene expression data sets, related to cancer, used for classification purposes. In order to do so, we measure the complexity of the ML models used to perform the tumors' classification. The results indicate that most of these data sets can be effectively discriminated by a simple linear function. Ana Carolina Lorena, Ivan G. Costa, Marcílio Carlos Pereira de Souto |
HIS | 1 |
| 2008 | Ensembles of Pre-processing Techniques for Noise Detection in Gene Expression Data
Giampaolo L. Libralon, André C. P. L. F. de Carvalho, Ana Carolina Lorena |
ICONIP (1) | 3 |
| 2008 | Potential Distribution Modelling Using Machine Learning
Ana Carolina Lorena, Marinez Ferreira de Siqueira, Renato De Giovanni, André C. P. L. F. de Carvalho, Ronaldo C. Prati |
IEA/AIE | 1 |
| 2008 | Evolutionary tuning of SVM parameter values in multiclass problems
Ana Carolina Lorena, André C. P. L. F. de Carvalho |
Neurocomputing | 1 |
| 2007 | Comparing Several Evaluation Functions in the Evolutionary Design of Multiclass Support Vector MachinesabstractSupport Vector Machines were originally designed to solve two-class classification problems. When they are applied to multiclass classification problems, the original problem is usually decomposed into multiple binary sub- problems. Afterwards, individual classifiers are induced to solve each of these binary problems. To obtain the final multiclass prediction, the outputs of these binary classifiers generated are combined. Genetic Algorithms can be used to optimize the combination of binary classifiers, defining the decomposition according to the performance obtained in the multiclass problem solution. This paper investigates several evaluation functions that can be used in order to evaluate the performance of the decompositions evolved by genetic algorithms. Ana Carolina Lorena, André C. P. L. F. de Carvalho |
HIS | 1 |
| 2005 | Support Vector Machines Applied to White Blood Cell RecognitionabstractA clinical decision support system known as Leuko has been developed for leukemia diagnosis using a naive Bayes classifier. The system is able to recognize six types of white blood cells (WBC), including a malignancy. This paper investigates the use of support vector machines (SVMs) classifiers to recognize WBC for future leukemia diagnosis. Since SVMs are originally designed for the solution of two class problems, several strategies for their extension to this multiclass task are investigated and compared. The experimental results evidence the potential of SVMs to leukemia diagnosis and indicate that a hierarchical tree-based multiclass strategy can be better suited to a future update of the Leuko system. Daniela Ushizima, Ana Carolina Lorena, André C. P. L. F. de Carvalho |
HIS | 2 |
| 2005 | Minimum Spanning Trees in Hierarchical Multiclass Support Vector Machines Generation
Ana Carolina Lorena, André C. P. L. F. de Carvalho |
IEA/AIE | 1 |
| 2003 | Human Splice Site Identification with Multiclass Support Vector Machines and Bagging
Ana Carolina Lorena, André C. P. L. F. de Carvalho |
ICANN | 1 |