EDBT 2026 Demo / reviewers in the wild / expert
Pedro H. Abreu
dblp:66/6482 · also Pedro Henriques Abreu
· DBLP profile ↗
53ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0002-9278-8194ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Security and privacy · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploring the influence of missing data imputation in group fairness metricsabstractMissing data is a common problem in real-world datasets and can be characterized as the lack of information on one or multiple variables in a dataset. The most frequent technique for handling this issue is imputation, which consists in the replacement of the missing values according to a predefined criterion. Since missing values are often imputed based on the known values in the dataset, existing data issues can be propagated during the imputation process. One such issue is fairness, a concept integral to responsible Artificial Intelligence practices. This work investigates the impact of the imputation process on system fairness by examining how imputation affects the fairness of predictions in Machine Learning models. It provides a comprehensive analysis covering thirteen unfair benchmark datasets with six state-of-the-art imputation strategies under synthetic Missing Not At Random and Missing At Random mechanisms in a multivariate scenario with 10%, 20%, 40%, and 60% of missing rates. Fairness was measured by the following metrics: Statistical Parity, Equalized Odds, Equality of Opportunity, Predictive Equality, Equality of Positive, and Negative Predicted Values. The results demonstrate that the missing mechanism, the classifier choice, and the imputation strategy decisively influence the fairness of the predictions obtained by the Machine Learning models. Arthur Dantas Mangussi, Ricardo Cardoso Pereira, Miriam Seoane Santos, Ana Carolina Lorena, Mykola Pechenizkiy, Pedro H. Abreu |
Artif. Intell. | 6 |
| 2026 | pycol-vis: A Python package for image complexity assessmentabstractDataset complexity poses a significant challenge in classification tasks, especially in real-world applications where a combination of factors such as class overlap, data imbalance, noise, and dimensionality can jeopardize a machine learning algorithm’s performance. While measures to quantify complexity have been proposed and studied in depth for tabular datasets, there is a lack of studies and toolkits focused on measuring complexity in non-structured image data. This limitation hinders our understanding of visual complexity, despite the importance of image data in fields such as healthcare, remote sensing, and autonomous navigation. To address this challenge, we introduce pycol-vis, a novel Python package that helps researchers estimate image complexity. The package implements 17 image complexity measures, specifically designed to capture complexity in real-world scenarios. This toolkit is essential for researchers dealing with complex classification problems in the vision domain, providing tools to assess difficulty and overlap in real-world image data. Diogo Apóstolo, Miriam Seoane Santos, Ana Carolina Lorena, Nathalie Japkowicz, Pedro H. Abreu |
Neurocomputing | 5 |
| 2026 | mlcpl: A python package for deep multi-label image classification with partial-labels on PyTorch
Chak Fong Chong, Xu Yang 0010, Yapeng Wang 0001, Pedro H. Abreu |
Neurocomputing | 4 |
| 2026 | Fairness in machine learning pipelines: Guided interventions with the Fairforge toolabstractWith the growing adoption of machine learning systems in high-stakes domains such as healthcare, finance, and public administration, ensuring that these systems behave responsibly has become an urgent concern. Initiatives such as the EU AI Act and the broader movement toward responsible AI have highlighted fairness as a key challenge in the development and deployment of such technologies. Although existing tools support model optimization through hyperparameter tuning and algorithm selection, they often neglect the broader pipeline, overlooking how factors like data bias and model evaluation practices contribute to fairness. This paper presents Fairforge, a tool designed to support responsible ML development by guiding users through essential stages of the pipeline. These include data preprocessing with bias-awareness, fairness-informed model training, and postprocessing correction techniques. Fairforge also provides integrated interfaces for evaluating both performance and fairness metrics. The tool aims to make state-of-the-art fairness techniques accessible to users without deep expertise in the field. To validate the effectiveness of Fairforge, we conducted a series of usability tests involving users with diverse levels of technical backgrounds. The results demonstrate that the tool helps promote fairness in model development, even among non-expert practitioners. Emanuel Roque, Miriam Seoane Santos, Penousal Machado, Pedro H. Abreu |
Neurocomputing | 4 |
| 2026 | LogicMix: Sample mixing data augmentation for multi-label image classification with partial labels
Chak Fong Chong, Jielong Guo, Xu Yang 0010, Wei Ke 0001, Pedro H. Abreu, Yapeng Wang 0001, Sio Kei Im |
Pattern Recognit. | 5 |
| 2026 | Unveiling Group-Specific Distributed Concept Drift: A Fairness Imperative in Federated LearningabstractIn the evolving field of machine learning, ensuring group fairness has become a critical concern, prompting the development of algorithms designed to mitigate bias in decision-making processes. Group fairness refers to the principle that a model's decisions should be equitable across different groups defined by sensitive attributes such as gender or race, ensuring that individuals from privileged groups and unprivileged groups are treated fairly and receive similar outcomes. However, achieving fairness in the presence of group-specific concept drift remains an unexplored frontier, and our research represents pioneering efforts in this regard. Group-specific concept drift refers to situations where one group experiences concept drift over time, while another does not, leading to a decrease in fairness even if accuracy (ACC) remains fairly stable. Within the framework of federated learning (FL), where clients collaboratively train models, its distributed nature further amplifies these challenges since each client can experience group-specific concept drift independently while still sharing the same underlying concept, creating a complex and dynamic environment for maintaining fairness. The most significant contribution of our research is the formalization and introduction of the problem of group-specific concept drift and its distributed counterpart, shedding light on its critical importance in the field of fairness. In addition, leveraging insights from prior research, we adapt an existing distributed concept drift adaptation algorithm to tackle group-specific distributed concept drift, which uses a multimodel approach, a local group-specific drift detection mechanism, and continuous clustering of models over time. The findings from our experiments highlight the importance of addressing group-specific concept drift and its distributed counterpart to advance fairness in machine learning. Teresa Salazar, João Gama 0001, Helder Araújo, Pedro H. Abreu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Studying the robustness of data imputation methodologies against adversarial attacksabstractCybersecurity attacks, such as poisoning and evasion, can intentionally introduce false or misleading information in different forms into data, potentially leading to catastrophic consequences for critical infrastructures, like water supply or energy power plants. While numerous studies have investigated the impact of these attacks on model-based prediction approaches, they often overlook the impurities present in the data used to train these models. One of those forms is missing data, the absence of values in one or more features. This issue is typically addressed by imputing missing values with plausible estimates, which directly impacts the performance of the classifier. The goal of this work is to promote a Data-centric AI approach by investigating how different types of cybersecurity attacks impact the imputation process. To this end, we conducted experiments using four popular evasion and poisoning attacks strategies across 29 real-world datasets, including the NSL-KDD and Edge-IIoT datasets, which were used as case study. For the adversarial attack strategies, we employed the Fast Gradient Sign Method, Carlini & Wagner, Project Gradient Descent, and Poison Attack against Support Vector Machine algorithm. Also, four state-of-the-art imputation strategies were tested under Missing Not At Random, Missing Completely at Random, and Missing At Random mechanisms using three missing rates (5%, 20%, 40%). We assessed imputation quality using MAE, while data distribution shifts were analyzed with the Kolmogorov–Smirnov and Chi-square tests. Furthermore, we measured classification performance by training an XGBoost classifier on the imputed datasets, using F1-score, Accuracy, and AUC. To deepen our analysis, we also incorporated six complexity metrics to characterize how adversarial attacks and imputation strategies impact dataset complexity. Our findings demonstrate that adversarial attacks significantly impact the imputation process. In terms of imputation assessment in what concerns to quality error, the scenario that enrolees imputation with Project Gradient Descent attack proved to be more robust in comparison to other adversarial methods. Regarding data distribution error, results from the Kolmogorov–Smirnov test indicate that in the context of numerical features, all imputation strategies differ from the baseline (without missing data) however for the categorical context Chi-Squared test proved no difference between imputation and the baseline. Arthur Dantas Mangussi, Ricardo Cardoso Pereira, Ana Carolina Lorena, Miriam Seoane Santos, Pedro H. Abreu |
Comput. Secur. | 5 |
| 2025 | Reparameterization convolutional neural networks for handling imbalanced datasets in solar panel fault classification
Jielong Guo, Chak Fong Chong, Pedro H. Abreu, Chao Mao, Jiaxuan Li 0003, Chan-Tong Lam, Benjamin K. Ng |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | The Role of Deep Learning in Medical Image Inpainting: A Systematic ReviewabstractImage inpainting is a crucial technique in computer vision, particularly for reconstructing corrupted images. In medical imaging, it addresses issues from instrumental errors, artifacts, or human factors. The development of deep learning techniques has revolutionized image inpainting, allowing for the generation of high-level semantic information to ensure structural and textural consistency in restored images. This article presents a comprehensive review of 53 studies on deep image inpainting in medical imaging, analyzing its evolution, impact, and limitations. The findings highlight the significance of deep image inpainting in artifact removal and enhancing the performance of multi-task approaches by localizing and inpainting regions of interest. Furthermore, the study identifies magnetic resonance imaging and computed tomography as the predominant modalities and highlights generative adversarial networks and U-Net as preferred architectures. Future research directions include the development of blind inpainting techniques, the exploration of techniques suitable for 3D/4D images, multiple artifacts, and multi-task applications, and the improvement of architectures. Joana Cristo Santos, Hugo Tomás Pereira Alexandre, Miriam Seoane Santos, Pedro H. Abreu |
ACM Trans. Comput. Heal. | 4 |
| 2025 | Pycol: A Python package for dataset complexity measuresabstractClass overlap presents a significant challenge to machine learning algorithms, especially when class imbalance is present. These factors contribute substantially to the complexity of classification tasks, particularly in real-world scenarios. As a result, measuring overlap is crucial, yet it remains difficult to quantify due to its intricate nature, since it can manifest and be measured in multiple ways. To help mitigate this, recent research has conceptualized a new taxonomy of class overlap measures, divided into multiple families, which allows researchers to obtain a more complete overview of the complexity of the datasets. In line with recent research, we introduce a new Python package for class overlap measurement named pycol. This package implements 29 overlap measures, divided into four overlap families specifically designed to capture class overlap in imbalanced real-world scenarios. This makes pycol an essential tool for researchers dealing with complex classification problems, providing robust solutions to quantify the joint-effect of class overlap and class imbalance effectively. Diogo Apóstolo, Miriam Seoane Santos, Ana Carolina Lorena, Pedro H. Abreu |
Neurocomputing | 4 |
| 2025 | Category-wise Fine-Tuning: Resisting incorrect pseudo-labels in multi-label image classification with partial labels
Chak Fong Chong, Xinyi Fang, Jielong Guo, Pedro H. Abreu, Yapeng Wang 0001, Xu Yang 0010, Wei Ke 0001, Sio Kei Im |
Neurocomputing | 4 |
| 2025 | mdatagen: A python library for the artificial generation of missing data
Arthur Dantas Mangussi, Miriam Seoane Santos, Filipe Loyola Lopes, Ricardo Cardoso Pereira, Ana Carolina Lorena, Pedro H. Abreu |
Neurocomputing | 6 |
| 2024 | Reconstruction of Mammography Projections using Image-to-Image Translation TechniquesabstractMammography imaging is the gold standard for breast cancer detection and involves capturing two projections: mediolateral oblique and craniocaudal projections.The implementation of an approach that allows the acquisition of only one projection and reconstructs the other could mitigate patient burden, minimize radiation exposure, and reduce costs.Image-to-image translation has showcased the ability to generate realistic synthetic images in different medical imaging modalities which make these techniques a great candidate for the novel application in mammography.This study aims to compare five image-to-image translation approaches to assess the feasibility of reconstructing a mammography projection from its counterpart.The results indicate that ResViT shows the best overall performance in translating between both projections. Joana Cristo Santos, Miriam Seoane Santos, Pedro H. Abreu |
ESANN | 3 |
| 2024 | An Interpretable Human-in-the-Loop Process to Improve Medical Image Classification
Joana Cristo Santos, Miriam Seoane Santos, Pedro H. Abreu |
IDA (1) | 3 |
| 2024 | Data-Centric Federated Learning for Anomaly Detection in Smart Grids and Other Industrial Control SystemsabstractEnergy smart grids and other modern industrial control systems networks impose considerable security management challenges due to several factors: their broad geographic dispersion and capillarity, the constrained nature of many of the devices and network links that integrate them, and the fact that they are often fragmented across multiple domains, owned and managed by different entities which often have nonaligned or even competing interests. Due to this scenario, we propose to improve federated learning-based anomaly detection for smart grids and other industrial control networks, using a federated data-centric methodology that attends to the balance and causality of the data, improving the representation of the different classes of anomalies of the ingested data, which directly impact the classifier's performance. The proposed approach shows up to 33% performance improvements in terms of F1-score for attack classification, compared to the baseline federated approach (not attending to class imbalance and causality) on a broad range of industrial control systems traffic datasets. Dylan Perdigão, Tiago Cruz 0001, Paulo Simões 0001, Pedro H. Abreu |
NOMS | 4 |
| 2024 | Imputation of data Missing Not at Random: Artificial generation and benchmark analysisabstractExperimental assessment of different missing data imputation methods often compute error rates between the original values and the estimated ones. This experimental setup relies on complete datasets that are injected with missing values. The injection process is straightforward for the Missing Completely At Random and Missing At Random mechanisms; however, the Missing Not At Random mechanism poses a major challenge, since the available artificial generation strategies are limited. Furthermore, the studies focused on this latter mechanism tend to disregard a comprehensive baseline of state-of-the-art imputation methods. In this work, both challenges are addressed: four new Missing Not At Random generation strategies are introduced and a benchmark study is conducted to compare six imputation methods in an experimental setup that covers 10 datasets and five missingness levels (10% to 80%). The overall findings are that, for most missing rates and datasets, the best imputation method to deal with Missing Not At Random values is the Multiple Imputation by Chained Equations, whereas for higher missingness rates autoencoders show promising results. Ricardo Cardoso Pereira, Pedro H. Abreu, Pedro Pereira Rodrigues, Mário A. T. Figueiredo |
Expert Syst. Appl. | 2 |
| 2024 | Call for Papers: Data Generation in Healthcare Environments
Ricardo Cardoso Pereira, Pedro Pereira Rodrigues, Irina S. Moreira, Pedro H. Abreu |
J. Biomed. Informatics | 4 |
| 2023 | Evaluating the faithfulness of saliency maps in explaining deep learning models using realistic perturbations
José Pereira Amorim, Pedro H. Abreu, João A. M. Santos, Marc Cortes, Victor Vila |
Inf. Process. Manag. | 2 |
| 2022 | The identification of cancer lesions in mammography images with missing pixels: analysis of morphologyabstractThe quality of mammography images is essential for the diagnosis of breast cancer and image imputation has become a popular technique to overcome noise, artifacts, and missing data to aid in the diagnosis of diseases. In this paper, we assess the performance of six imputation methodologies for the reconstruction of missing pixels in different morphologies in mammography images. The images included in this study are collected from four public datasets (CBIS-DDSM, Mini-MIAS, INbreast, and CSAW) and the imputation results are evaluated through the mean absolute error (MAE) and structural similarity index measure (SSIM). This study goes beyond the traditional evaluation of imputation algorithms, analyzing imputation quality, morphology preservation and classification performance. The effects of imputation on the morphology of cancer lesions are of utmost importance since it lays the foundation for physicians to interpret and analyze the imputation results. The results show that DIP is the most promising methodology for higher missing pixel rates, morphology preservation, and classifying malignant and benign images. Joana Cristo Santos, Pedro H. Abreu, Miriam Seoane Santos |
DSAA | 2 |
| 2022 | The impact of heterogeneous distance functions on missing data imputation and classification performance
Miriam Seoane Santos, Pedro H. Abreu, Alberto Fernández 0001, Julián Luengo, João A. M. Santos |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | Partial Multiple Imputation With Variational Autoencoders: Tackling Not at Randomness in Healthcare DataabstractMissing data can pose severe consequences in critical contexts, such as clinical research based on routinely collected healthcare data. This issue is usually handled with imputation strategies, but these tend to produce poor and biased results under the Missing Not At Random (MNAR) mechanism. A recent trend that has been showing promising results for MNAR is the use of generative models, particularly Variational Autoencoders. However, they have a limitation: the imputed values are the result of a single sample, which can be biased. To tackle it, an extension to the Variational Autoencoder that uses a partial multiple imputation procedure is introduced in this work. The proposed method was compared to 8 state-of-the-art imputation strategies, in an experimental setup with 34 datasets from the medical context, injected with the MNAR mechanism (10% to 80% rates). The results were evaluated through the Mean Absolute Error, with the new method being the overall best in 71% of the datasets, significantly outperforming the remaining ones, particularly for high missing rates. Finally, a case study of a classification task with heart failure data was also conducted, where this method induced improvements in 50% of the classifiers. Ricardo Cardoso Pereira, Pedro H. Abreu, Pedro Pereira Rodrigues |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Assessing the Impact of Distance Functions on K-Nearest Neighbours Imputation of Biomedical Datasets
Miriam Seoane Santos, Pedro H. Abreu, Szymon Wilk, João A. M. Santos |
AIME | 2 |
| 2020 | Missing Image Data Imputation using Variational Autoencoders with Weighted Loss
Ricardo Cardoso Pereira, Joana Cristo Santos, José Pereira Amorim, Pedro Pereira Rodrigues, Pedro H. Abreu |
ESANN | 5 |
| 2020 | Interpretability vs. Complexity: The Friction in Deep Neural NetworksabstractSaliency maps have been used as one possibility to interpret deep neural networks. This method estimates the relevance of each pixel in the image classification, with higher values representing pixels which contribute positively to classification. The goal of this study is to understand how the complexity of the network affects the interpretabilty of the saliency maps in classification tasks. To achieve that, we investigate how changes in the regularization affects the saliency maps produced, and their fidelity to the overall classification process of the network.The experimental setup consists in the calculation of the fidelity of five saliency map methods that were compare, applying them to models trained on the CIFAR-10 dataset, using different levels of weight decay on some or all the layers. Achieved results show that models with lower regularization are statistically (significance of 5%) more interpretable than the other models. Also, regularization applied only to the higher convolutional layers or fully-connected layers produce saliency maps with more fidelity. José Pereira Amorim, Pedro H. Abreu, Mauricio Reyes 0001, João A. M. Santos |
IJCNN | 2 |
| 2020 | VAE-BRIDGE: Variational Autoencoder Filter for Bayesian Ridge Imputation of Missing DataabstractThe missing data issue is often found in real-world datasets and it is usually handled with imputation strategies that replace the missing values with new data. Recently, generative models such as Variational Autoencoders have been applied for this imputation task. However, they were always used to perform the entire imputation, which has presented limited results when comparing to other state-of-the-art methods. In this work, a new approach called Variational Autoencoder Filter for Bayesian Ridge Imputation is introduced. It uses a Variational Autoencoder at the beginning of the imputation pipeline to filter the instances that are later fitted to a Bayesian ridge regression used to predict the new values. The approach was compared to four state-of-the-art imputation methods using 10 datasets from the healthcare context covering clinical trials, all injected with missing values under different rates. The proposed approach significantly outperformed the remaining methods in all settings, achieving an overall improvement between 26% and 67%. Ricardo Cardoso Pereira, Pedro H. Abreu, Pedro Pereira Rodrigues |
IJCNN | 2 |
| 2020 | Reviewing Autoencoders for Missing Data Imputation: Technical Trends, Applications and OutcomesabstractMissing data is a problem often found in real-world datasets and it can degrade the performance of most machine learning models. Several deep learning techniques have been used to address this issue, and one of them is the Autoencoder and its Denoising and Variational variants. These models are able to learn a representation of the data with missing values and generate plausible new ones to replace them. This study surveys the use of Autoencoders for the imputation of tabular data and considers 26 works published between 2014 and 2020. The analysis is mainly focused on discussing patterns and recommendations for the architecture, hyperparameters and training settings of the network, while providing a detailed discussion of the results obtained by Autoencoders when compared to other state-of-the-art methods, and of the data contexts where they have been applied. The conclusions include a set of recommendations for the technical settings of the network, and show that Denoising Autoencoders outperform their competitors, particularly the often used statistical methods. Ricardo Cardoso Pereira, Miriam Seoane Santos, Pedro Pereira Rodrigues, Pedro H. Abreu |
J. Artif. Intell. Res. | 4 |
| 2020 | How distance metrics influence missing data imputation with k-nearest neighbours
Miriam Seoane Santos, Pedro H. Abreu, Szymon Wilk, João A. M. Santos |
Pattern Recognit. Lett. | 2 |
| 2020 | Guest Editorial: Information Fusion for Medical Data: Early, Late, and Deep Fusion Methods for Multimodal DataabstractThe papers in this special section examine important current topics on multimodal data fusion in the medical context. All clinical data, including genomic and proteomic, play a role in the diagnosis and in particular in the treatment planning and follow-up. This is true for all types of data analyses whether in classification, regression, retrieval, clustering, or other. The interaction between several types of information is not always well understood. Experienced clinicians automatically and even unconsciously add multiple sources of information into their decision process, but machine learning tools often concentrate on single information sources. This special issue presents five examples where several data sources are fused. The papers give several examples of fusion techniques and also the results obtained in quite different application scenarios. Inês Domingues, Henning Müller, Andrés Ortiz 0001, Belur V. Dasarathy, Pedro H. Abreu, Vince D. Calhoun |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | Autonomous agents and multi-agent systems applied in healthcare
Sara Montagna, Daniel Castro Silva, Pedro H. Abreu, Márcia Ito, Michael Schumacher 0001, Eloisa Vargiu |
Artif. Intell. Medicine | 3 |
| 2019 | Multiple-Choice Questions in Programming Courses: Can We Use Them and Are Students Motivated by Them?abstractLow performance of nontechnical engineering students in programming courses is a problem that remains unsolved. Over the years, many authors have tried to identify the multiple causes for that failure, but there is unanimity on the fact that motivation is a key factor for the acquisition of knowledge by students. To better understand motivation, a new evaluation strategy has been adopted in a second programming course of a nontechnical degree, consisting of 91 students. The goals of the study were to identify if those students felt more motivated to answer multiple-choice questions in comparison to development questions, and what type of question better allows for testing student knowledge acquisition. Possibilities around the motivational qualities of multiple-choice questions in programming courses will be discussed in light of the results. In conclusion, it seems clear that student performance varies according to the type of question. Our study points out that multiple-choice questions can be seen as a motivational factor for engineering students and it might also be a good way to test acquired programming concepts. Therefore, this type of question could be further explored in the evaluation points. Pedro H. Abreu, Daniel Castro Silva, Anabela Jesus Gomes |
ACM Trans. Comput. Educ. | 1 |
| 2018 | Denial of Service Attacks: Detecting the Frailties of Machine Learning Algorithms in the Classification Process
Ivo Frazão, Pedro H. Abreu, Tiago Cruz 0001, Helder Araújo, Paulo Simões 0001 |
CRITIS | 2 |
| 2018 | Interpreting deep learning models for ordinal problems
José Pereira Amorim, Inês Domingues, Pedro H. Abreu, João A. M. Santos |
ESANN | 3 |
| 2018 | Bi-Rads Classification of Breast Cancer: A New Pre-Processing Pipeline for Deep Models TrainingabstractOne of the main difficulties in the use of deep learning strategies in medical contexts is the training set size. While these methods need large annotated training sets, these datasets are costly to obtain in medical contexts and suffer from intra and inter-subject variability. In the present work, two new pre-processing techniques are introduced to improve a deep classifier performance. First, data augmentation based on co-registration is suggested. Then, multi-scale enhancement based on Difference of Gaussians is proposed. Results are accessed in a public mammogram database, the InBreast, in the context of an ordinal problem, the BI-RADS classification. Moreover, a pre-trained Convolutional Neural Network with the AlexNet architecture was used as a base classifier. The multi-class classification experiments show that the proposed pipeline with the Difference of Gaussians and the data augmentation technique outperforms using the original dataset only and using the original dataset augmented by mirroring the images. Inês Domingues, Pedro H. Abreu, João A. M. Santos |
ICIP | 2 |
| 2018 | Missing Data Imputation via Denoising Autoencoders: The Untold Story
Adriana Fonseca Costa, Miriam Seoane Santos, Jastin Pompeu Soares, Pedro H. Abreu |
IDA | 4 |
| 2018 | Analysing the Footprint of Classifiers in Overlapped and Imbalanced Contexts
Marta Mercier, Miriam Seoane Santos, Pedro H. Abreu, Carlos Soares, Jastin Pompeu Soares, João A. M. Santos |
IDA | 3 |
| 2018 | Exploring the Effects of Data Distribution in Missing Data Imputation
Jastin Pompeu Soares, Miriam Seoane Santos, Pedro H. Abreu, Helder Araújo, João A. M. Santos |
IDA | 3 |
| 2018 | Evaluation of Oversampling Data Balancing Techniques in the Context of Ordinal ClassificationabstractData imbalance is characterized by a discrepancy in the number of examples per class of a dataset. This phenomenon is known to deteriorate the performance of classifiers, since they are less able to learn the characteristics of the less represented classes. For most imbalanced datasets, the application of sampling techniques improves the classifier's performance. For small datasets, oversampling has been shown to be the most appropriate strategy since it augments the original set of samples. Although several oversampling strategies have been proposed and tested over the years, the work has mostly focused on binary or multi-class tasks. Motivated by medical applications, where there is often an order associated with the classes (increasing likelihood of malignancy, for instance), the present work tests some existing oversampling techniques in ordinal contexts. Moreover, four new oversampling techniques are proposed. Experiments were made both on private and public datasets. Private datasets concern the assessment of response to treatment on oncologic diseases. The 15 public datasets were chosen since they are widely used in the literature. Results show that data balance techniques improve classification results on ordinal imbalanced datasets, even when these techniques are not specifically designed for ordinal problems. With our pipeline, better or equal to published results were obtained for 10 out of the 15 public datasets with improvements upon a decrease of 0.43 on MMAE. Inês Domingues, José Pereira Amorim, Pedro H. Abreu, Hugo Duarte, João A. M. Santos |
IJCNN | 3 |
| 2018 | Registration of CT with PET: A Comparison of Intensity-Based Approaches
Gisèle Pereira, Inês Domingues, Pedro Martins 0003, Pedro H. Abreu, Hugo Duarte, João A. M. Santos |
IWCIA | 4 |
| 2017 | Influence of Data Distribution in Missing Data Imputation
Miriam Seoane Santos, Jastin Pompeu Soares, Pedro H. Abreu, Helder Araújo, João A. M. Santos |
AIME | 3 |
| 2016 | Types of assessing student-programming knowledgeabstractHigh failure and dropout rates are common in higher education institutions with introductory programming courses. Some researchers advocate that sometimes teachers don't use correct methods of assessment and that many students pass in programming without knowing how to program. In this paper authors describe the assessment methodology applied to a first year, first semester, Biomedical Engineering programming course (2015/2016). Students' programming skills were tested by playing a game in the first class, then they were assessed with three tests and a final exam, each with topics the authors considered fundamental for the students to master. A correlation analyses between the different types of tests and exam questions is done, to evaluate the most suitable, for assessing programming knowledge, showing that it is possible to use different question types as a pedagogical strategy, to assess student difficulty levels and programming skills, that help students acquire abstract, reasoning and algorithm thinking in an acceptable level. Also, it is shown that different forms of questions are equivalent to assess equal knowledge and that it is possible to predict the ability of a student to program at an early stage. Anabela Jesus Gomes, Fernanda Brito Correia, Pedro H. Abreu |
FIE | 3 |
| 2016 | TweeProfiles3: Visualization of Spatio-Temporal Patterns on Twitter
André Maia, Tiago Cunha 0001, Carlos Soares, Pedro H. Abreu |
WorldCIST (1) | 4 |
| 2015 | MoCaS: Mobile Carpooling System
Álvaro Ribeiro, Daniel Castro Silva, Pedro H. Abreu |
WorldCIST (1) | 3 |
| 2015 | A new cluster-based oversampling method for improving survival prediction of hepatocellular carcinoma patients
Miriam Seoane Santos, Pedro H. Abreu, Pedro J. García-Laencina, Adélia Simão, Armando Carvalho |
J. Biomed. Informatics | 2 |
| 2014 | Augmented Reality Mobile Tourism Application
Flávio Pereira, Daniel Castro Silva, Pedro H. Abreu, António Pinho |
WorldCIST (2) | 3 |
| 2014 | MusE Central: A Data Aggregation System for Music Events
Delfim Simões, Pedro H. Abreu, Daniel Castro Silva |
WorldCIST (2) | 2 |
| 2014 | Strategy planner: Graphical definition of soccer set-plays
João Cravo, Pedro H. Abreu, Luís Paulo Reis, Nuno Lau, Luís Mota |
Data Knowl. Eng. | 3 |
| 2014 | An Inverted Ant Colony Optimization approach to traffic
José Capela Dias, Penousal Machado, Daniel Castro Silva, Pedro H. Abreu |
Eng. Appl. Artif. Intell. | 4 |
| 2014 | Using model-based collaborative filtering techniques to recommend the expected best strategy to defeat a simulated soccer opponentabstractHow to improve the performance of a simulated soccer team using final game statistics? This is the question this research aims to answer using model-based collaborative techniques and a robotic team – FC Portugal – as a case study. After developing a Pedro H. Abreu, Daniel Castro Silva, João Portela, João Mendes-Moreira 0001, Luís Paulo Reis |
Intell. Data Anal. | 1 |
| 2014 | Development of a flexible language for mission description for multi-robot missions
Daniel Castro Silva, Pedro H. Abreu, Luís Paulo Reis, Eugénio Oliveira |
Inf. Sci. | 2 |
| 2013 | An automatic approach to extract goal plans from soccer simulated matches
Pedro H. Abreu, Nuno Lau, Luís Paulo Reis |
Soft Comput. | 2 |
| 2012 | Automatic extraction of goal-scoring behaviors from soccer matchesabstractIn a soccer match, a cooperative behavior emerges from the combined execution of simple actions by players. A cooperative behavior can be planned if players are previously committed to its execution prior to its start or unplanned otherwise. The ability to reproduce some of these behaviors can be useful to help a team achieve better performances. This work presents an approach to identify and extract cooperative behaviors that start from set-pieces and lead to a goal while ball possession is kept. The representation of these behaviors is abstracted using a set-play definition language to promote their reusability. A set of game log files generated with the FC Portugal team and collected from the RoboCup 2010 2D simulated soccer competition were analyzed. The results achieved showed that 25% of the total goals scored originated from set-pieces which attests to the importance of performing this analysis. Several guidelines for the definition of future set-plays were also inferred. In the future, these behaviors shall be tested to infer which are capable of neutralizing an opponent's team strategy and maximize the creation of goal opportunities. Pedro H. Abreu, Nuno Lau, Luís Paulo Reis |
IROS | 2 |
| 2012 | Performance analysis in soccer: a Cartesian coordinates based approach using RoboCup data
Pedro H. Abreu, Daniel Castro Silva, Luís Paulo Reis, Júlio Garganta |
Soft Comput. | 1 |
| 2010 | Human vs. Robotic Soccer: How Far Are They? A Statistical Comparison
Pedro H. Abreu, Israel Costa, Daniel Castelão, Luís Paulo Reis, Júlio Garganta |
RoboCup | 1 |