Miriam Seoane Santos

dblp:173/3027 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-5912-963XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring the influence of missing data imputation in group fairness metrics
abstract
Missing data is a common problem in real-world datasets and can be characterized as the lack of information on one or multiple variables in a dataset. The most frequent technique for handling this issue is imputation, which consists in the replacement of the missing values according to a predefined criterion. Since missing values are often imputed based on the known values in the dataset, existing data issues can be propagated during the imputation process. One such issue is fairness, a concept integral to responsible Artificial Intelligence practices. This work investigates the impact of the imputation process on system fairness by examining how imputation affects the fairness of predictions in Machine Learning models. It provides a comprehensive analysis covering thirteen unfair benchmark datasets with six state-of-the-art imputation strategies under synthetic Missing Not At Random and Missing At Random mechanisms in a multivariate scenario with 10%, 20%, 40%, and 60% of missing rates. Fairness was measured by the following metrics: Statistical Parity, Equalized Odds, Equality of Opportunity, Predictive Equality, Equality of Positive, and Negative Predicted Values. The results demonstrate that the missing mechanism, the classifier choice, and the imputation strategy decisively influence the fairness of the predictions obtained by the Machine Learning models.
Arthur Dantas Mangussi, Ricardo Cardoso Pereira, Miriam Seoane Santos, Ana Carolina Lorena, Mykola Pechenizkiy, Pedro H. Abreu
Artif. Intell.3
2026 pycol-vis: A Python package for image complexity assessment
abstract
Dataset complexity poses a significant challenge in classification tasks, especially in real-world applications where a combination of factors such as class overlap, data imbalance, noise, and dimensionality can jeopardize a machine learning algorithm’s performance. While measures to quantify complexity have been proposed and studied in depth for tabular datasets, there is a lack of studies and toolkits focused on measuring complexity in non-structured image data. This limitation hinders our understanding of visual complexity, despite the importance of image data in fields such as healthcare, remote sensing, and autonomous navigation. To address this challenge, we introduce pycol-vis, a novel Python package that helps researchers estimate image complexity. The package implements 17 image complexity measures, specifically designed to capture complexity in real-world scenarios. This toolkit is essential for researchers dealing with complex classification problems in the vision domain, providing tools to assess difficulty and overlap in real-world image data.
Diogo Apóstolo, Miriam Seoane Santos, Ana Carolina Lorena, Nathalie Japkowicz, Pedro H. Abreu
Neurocomputing2
2026 Fairness in machine learning pipelines: Guided interventions with the Fairforge tool
abstract
With the growing adoption of machine learning systems in high-stakes domains such as healthcare, finance, and public administration, ensuring that these systems behave responsibly has become an urgent concern. Initiatives such as the EU AI Act and the broader movement toward responsible AI have highlighted fairness as a key challenge in the development and deployment of such technologies. Although existing tools support model optimization through hyperparameter tuning and algorithm selection, they often neglect the broader pipeline, overlooking how factors like data bias and model evaluation practices contribute to fairness. This paper presents Fairforge, a tool designed to support responsible ML development by guiding users through essential stages of the pipeline. These include data preprocessing with bias-awareness, fairness-informed model training, and postprocessing correction techniques. Fairforge also provides integrated interfaces for evaluating both performance and fairness metrics. The tool aims to make state-of-the-art fairness techniques accessible to users without deep expertise in the field. To validate the effectiveness of Fairforge, we conducted a series of usability tests involving users with diverse levels of technical backgrounds. The results demonstrate that the tool helps promote fairness in model development, even among non-expert practitioners.
Emanuel Roque, Miriam Seoane Santos, Penousal Machado, Pedro H. Abreu
Neurocomputing2
2025 Studying the robustness of data imputation methodologies against adversarial attacks
abstract
Cybersecurity attacks, such as poisoning and evasion, can intentionally introduce false or misleading information in different forms into data, potentially leading to catastrophic consequences for critical infrastructures, like water supply or energy power plants. While numerous studies have investigated the impact of these attacks on model-based prediction approaches, they often overlook the impurities present in the data used to train these models. One of those forms is missing data, the absence of values in one or more features. This issue is typically addressed by imputing missing values with plausible estimates, which directly impacts the performance of the classifier. The goal of this work is to promote a Data-centric AI approach by investigating how different types of cybersecurity attacks impact the imputation process. To this end, we conducted experiments using four popular evasion and poisoning attacks strategies across 29 real-world datasets, including the NSL-KDD and Edge-IIoT datasets, which were used as case study. For the adversarial attack strategies, we employed the Fast Gradient Sign Method, Carlini & Wagner, Project Gradient Descent, and Poison Attack against Support Vector Machine algorithm. Also, four state-of-the-art imputation strategies were tested under Missing Not At Random, Missing Completely at Random, and Missing At Random mechanisms using three missing rates (5%, 20%, 40%). We assessed imputation quality using MAE, while data distribution shifts were analyzed with the Kolmogorov–Smirnov and Chi-square tests. Furthermore, we measured classification performance by training an XGBoost classifier on the imputed datasets, using F1-score, Accuracy, and AUC. To deepen our analysis, we also incorporated six complexity metrics to characterize how adversarial attacks and imputation strategies impact dataset complexity. Our findings demonstrate that adversarial attacks significantly impact the imputation process. In terms of imputation assessment in what concerns to quality error, the scenario that enrolees imputation with Project Gradient Descent attack proved to be more robust in comparison to other adversarial methods. Regarding data distribution error, results from the Kolmogorov–Smirnov test indicate that in the context of numerical features, all imputation strategies differ from the baseline (without missing data) however for the categorical context Chi-Squared test proved no difference between imputation and the baseline.
Arthur Dantas Mangussi, Ricardo Cardoso Pereira, Ana Carolina Lorena, Miriam Seoane Santos, Pedro H. Abreu
Comput. Secur.4
2025 The Role of Deep Learning in Medical Image Inpainting: A Systematic Review
abstract
Image inpainting is a crucial technique in computer vision, particularly for reconstructing corrupted images. In medical imaging, it addresses issues from instrumental errors, artifacts, or human factors. The development of deep learning techniques has revolutionized image inpainting, allowing for the generation of high-level semantic information to ensure structural and textural consistency in restored images. This article presents a comprehensive review of 53 studies on deep image inpainting in medical imaging, analyzing its evolution, impact, and limitations. The findings highlight the significance of deep image inpainting in artifact removal and enhancing the performance of multi-task approaches by localizing and inpainting regions of interest. Furthermore, the study identifies magnetic resonance imaging and computed tomography as the predominant modalities and highlights generative adversarial networks and U-Net as preferred architectures. Future research directions include the development of blind inpainting techniques, the exploration of techniques suitable for 3D/4D images, multiple artifacts, and multi-task applications, and the improvement of architectures.
Joana Cristo Santos, Hugo Tomás Pereira Alexandre, Miriam Seoane Santos, Pedro H. Abreu
ACM Trans. Comput. Heal.3
2025 Pycol: A Python package for dataset complexity measures
abstract
Class overlap presents a significant challenge to machine learning algorithms, especially when class imbalance is present. These factors contribute substantially to the complexity of classification tasks, particularly in real-world scenarios. As a result, measuring overlap is crucial, yet it remains difficult to quantify due to its intricate nature, since it can manifest and be measured in multiple ways. To help mitigate this, recent research has conceptualized a new taxonomy of class overlap measures, divided into multiple families, which allows researchers to obtain a more complete overview of the complexity of the datasets. In line with recent research, we introduce a new Python package for class overlap measurement named pycol. This package implements 29 overlap measures, divided into four overlap families specifically designed to capture class overlap in imbalanced real-world scenarios. This makes pycol an essential tool for researchers dealing with complex classification problems, providing robust solutions to quantify the joint-effect of class overlap and class imbalance effectively.
Diogo Apóstolo, Miriam Seoane Santos, Ana Carolina Lorena, Pedro H. Abreu
Neurocomputing2
2025 mdatagen: A python library for the artificial generation of missing data
Arthur Dantas Mangussi, Miriam Seoane Santos, Filipe Loyola Lopes, Ricardo Cardoso Pereira, Ana Carolina Lorena, Pedro H. Abreu
Neurocomputing2
2024 Reconstruction of Mammography Projections using Image-to-Image Translation Techniques
abstract
Mammography imaging is the gold standard for breast cancer detection and involves capturing two projections: mediolateral oblique and craniocaudal projections.The implementation of an approach that allows the acquisition of only one projection and reconstructs the other could mitigate patient burden, minimize radiation exposure, and reduce costs.Image-to-image translation has showcased the ability to generate realistic synthetic images in different medical imaging modalities which make these techniques a great candidate for the novel application in mammography.This study aims to compare five image-to-image translation approaches to assess the feasibility of reconstructing a mammography projection from its counterpart.The results indicate that ResViT shows the best overall performance in translating between both projections.
Joana Cristo Santos, Miriam Seoane Santos, Pedro H. Abreu
ESANN2
2024 An Interpretable Human-in-the-Loop Process to Improve Medical Image Classification
Joana Cristo Santos, Miriam Seoane Santos, Pedro H. Abreu
IDA (1)2
2023 ydata-profiling: Accelerating data-centric AI with high-quality data
Fabiana Clemente, Gonçalo Martins Ribeiro, Alexandre Quemy, Miriam Seoane Santos, Ricardo Cardoso Pereira, Alex Barros
Neurocomputing4
2022 The identification of cancer lesions in mammography images with missing pixels: analysis of morphology
abstract
The quality of mammography images is essential for the diagnosis of breast cancer and image imputation has become a popular technique to overcome noise, artifacts, and missing data to aid in the diagnosis of diseases. In this paper, we assess the performance of six imputation methodologies for the reconstruction of missing pixels in different morphologies in mammography images. The images included in this study are collected from four public datasets (CBIS-DDSM, Mini-MIAS, INbreast, and CSAW) and the imputation results are evaluated through the mean absolute error (MAE) and structural similarity index measure (SSIM). This study goes beyond the traditional evaluation of imputation algorithms, analyzing imputation quality, morphology preservation and classification performance. The effects of imputation on the morphology of cancer lesions are of utmost importance since it lays the foundation for physicians to interpret and analyze the imputation results. The results show that DIP is the most promising methodology for higher missing pixel rates, morphology preservation, and classifying malignant and benign images.
Joana Cristo Santos, Pedro H. Abreu, Miriam Seoane Santos
DSAA3
2022 The impact of heterogeneous distance functions on missing data imputation and classification performance
Miriam Seoane Santos, Pedro H. Abreu, Alberto Fernández 0001, Julián Luengo, João A. M. Santos
Eng. Appl. Artif. Intell.1
2020 Assessing the Impact of Distance Functions on K-Nearest Neighbours Imputation of Biomedical Datasets
Miriam Seoane Santos, Pedro H. Abreu, Szymon Wilk, João A. M. Santos
AIME1
2020 Reviewing Autoencoders for Missing Data Imputation: Technical Trends, Applications and Outcomes
abstract
Missing data is a problem often found in real-world datasets and it can degrade the performance of most machine learning models. Several deep learning techniques have been used to address this issue, and one of them is the Autoencoder and its Denoising and Variational variants. These models are able to learn a representation of the data with missing values and generate plausible new ones to replace them. This study surveys the use of Autoencoders for the imputation of tabular data and considers 26 works published between 2014 and 2020. The analysis is mainly focused on discussing patterns and recommendations for the architecture, hyperparameters and training settings of the network, while providing a detailed discussion of the results obtained by Autoencoders when compared to other state-of-the-art methods, and of the data contexts where they have been applied. The conclusions include a set of recommendations for the technical settings of the network, and show that Denoising Autoencoders outperform their competitors, particularly the often used statistical methods.
Ricardo Cardoso Pereira, Miriam Seoane Santos, Pedro Pereira Rodrigues, Pedro H. Abreu
J. Artif. Intell. Res.2
2020 How distance metrics influence missing data imputation with k-nearest neighbours
Miriam Seoane Santos, Pedro H. Abreu, Szymon Wilk, João A. M. Santos
Pattern Recognit. Lett.1
2018 Missing Data Imputation via Denoising Autoencoders: The Untold Story
Adriana Fonseca Costa, Miriam Seoane Santos, Jastin Pompeu Soares, Pedro H. Abreu
IDA2
2018 Analysing the Footprint of Classifiers in Overlapped and Imbalanced Contexts
Marta Mercier, Miriam Seoane Santos, Pedro H. Abreu, Carlos Soares, Jastin Pompeu Soares, João A. M. Santos
IDA2
2018 Exploring the Effects of Data Distribution in Missing Data Imputation
Jastin Pompeu Soares, Miriam Seoane Santos, Pedro H. Abreu, Helder Araújo, João A. M. Santos
IDA2
2017 Influence of Data Distribution in Missing Data Imputation
Miriam Seoane Santos, Jastin Pompeu Soares, Pedro H. Abreu, Helder Araújo, João A. M. Santos
AIME1
2015 A new cluster-based oversampling method for improving survival prediction of hepatocellular carcinoma patients
Miriam Seoane Santos, Pedro H. Abreu, Pedro J. García-Laencina, Adélia Simão, Armando Carvalho
J. Biomed. Informatics1