VLDB 2026 Research / reviewers in the wild / expert
Tânia Pereira 0001
dblp:15/9502
· DBLP profile ↗
18ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0003-1681-2436ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Post-Hoc to Integrated Calibration: Bilevel Training with Doubly Kernelized ECE
João D. Nunes, Felipe Coutinho, Inês Machado, Diana Montezuma, Domingos Oliveira, Tânia Pereira 0001, Jaime S. Cardoso 0001 |
ICPR (15) | 6 |
| 2025 | MYCN-Amplified Neuroblastoma Detection Radiomics Vs. Trainable FeaturesabstractNeuroblastoma (NB) is the most common extracranial tumor in pediatric cases. The MYCN oncogene amplification (MNA) is knowingly correlated with a poor prognosis, making detecting this biomarker crucial for treatment selection and survival prediction. The current clinical protocol for MNA detection includes invasive procedures, such as biopsy. The proposed work aims to develop non-invasive techniques for predicting MNA in patients with diagnosed NB, using AI-based models and Computerized Tomography (CT) scans. Machine learning methods that use the imaging features extracted from the tumor on the CT slices were developed and compared with deep learning (DL) models. Additionally, agnostic explainable methods for imaging were applied to create explanations about the relevant information used by the DL models in the prediction. The results show a better performance for the DL approach, which achieved an AUC of$0.94 \pm 0.04$. The similarity in the explanations produced by the models trained with different data splits showed that feature extraction remains somewhat invariant to shifts in training data, which is relevant given the small amount of data available. Learning models were shown to have predictive potential that, with further improvements, can be integrated into predictive, explainable, and, thus, trustworthy systems to aid clinicians in the decision-making process. Mafalda Malafaia, Francisco Silva 0002, Diogo Costa Carvalho, Sílvia Costa Dias, Helena Torrão, Hélder P. Oliveira, Tânia Pereira 0001 |
BIBE | 8 |
| 2025 | Incrementally Learning to Segment the Lungs: Similarities and Differences Across InstitutionsabstractSegmentation of the lungs in Computed Tomography (CT) is very challenging due to changes in lung shape, size, and parenchyma pattern, as well as differences in imaging acquisition protocols. As a consequence, these models may not be robust and may decrease their performance when deployed in a clinical setting. The Continual Learning paradigm holds great promise since learning models continually acquire incoming knowledge, having the ability to adapt to changing environments. In this work, experience replay with random sampling of past data was implemented, using the original CT images and the corresponding ground-truths. Data from four different institutions were used to develop the experiments, and the models were evaluated on a cross-cohort dataset. Using raw data, the goal was to study how the datasets and their imaging patterns were related and what impact the training datasets have on one another. The catastrophic forgetting effect diminished for almost all datasets. For two of the in-domain test datasets there was forward and backward transfer, results that could be linked to a possible similarity between them. A mean DSC of 0.94 was obtained across all datasets. The results showed how the similarity or disparity between data from different institutions can influence the performance of learning models. Joana Sousa, Hélder P. Oliveira, Tânia Pereira 0001 |
BIBE | 3 |
| 2025 | From Pixels to Pathways: AI-Based Approaches for Multimodal Lung Cancer ClassificationabstractLung cancer remains the leading cause of cancer related deaths globally, responsible for approximately 1.8 million deaths each year. A key contributor to this high mortality rate is the late-stage diagnosis of the disease, underscoring the urgent need for effective early detection strategies. Low-dose computed tomography (CT) has shown great value in early screening, particularly when paired with clinical information. Clinical data, while valuable, lacks spatial and morphological insights essential for comprehensive evaluation. Combining both modalities offers a more holistic approach for lung cancer classification. This study presents AI-based methods for lung cancer classification using unimodal approaches - structured clinical data and chest CT imaging - alongside a novel multimodal deep learning framework that integrates both data types to classify lung nodules as malignant or benign. For the clinical modality, machine learning models including logistic regression, random forests, LightGBM, XGBoost, and multilayer perceptrons were evaluated with extensive hyperparameter tuning. In the imaging modality, ResNet18 and ResNet34 convolutional neural networks were used, with and without data augmentation. The study explored both intermediate and late fusion strategies to combine modality-specific representations. Results show that multimodal models consistently outperformed their unimodal counterparts, achieving a best-case area under the ROC curve (AUC) of 0.9138, with an accuracy of 0.8424 and an F1-score of 0.8422. These findings highlight the complementary strengths of imaging and clinical data and support the growing potential of multimodal deep learning in improving diagnostic accuracy in lung cancer classification. Sofia Gonçalves, Joana Sousa, Margarida Gouveia, Maria Amaro, Hélder P. Oliveira, Tânia Pereira 0001 |
BIBM | 6 |
| 2025 | Toward Generalizable Radiomics Models for EGFR Mutation Prediction: A Multi-Dataset EvaluationabstractEpidermal Growth Factor Receptor (EGFR) is one of the most frequently mutated genes in lung cancer. Its mutation status characterization is crucial for personalized treatment in Non-Small-Cell Lung Cancer (NSCLC). Biopsy is the gold standard for characterizing the EGFR mutation status. However, it is an invasive time-consuming method and is often burdensome or even impractical for some patients. Therefore, it is of utmost importance to identify alternative non-invasive methods for classifying this mutation. Computed Tomography (CT) images represent a non-invasive, safer and faster method to directly characterize lung cancer. This study developed a comprehensive radiomic approach for EGFR mutation classification using CT images, in which two preprocessing strategies were compared and five machine learning algorithms were evaluated across different datasets. We analyzed two independent datasets individually and combined, implementing lung containing nodule versus bounding box around nodule preprocessing approaches. Radiomic features were extracted using PyRadiomics and selected through Principal Component Analysis (PCA)$(65-95 \%$variance thresholds) and pairwise correlation filtering. The results demonstrated that the lung with nodule strategy achieved better and more consistent performance compared to the bounding box around the nodule method. The best performance$(\text{AUC}=0.780)$was achieved using Random Forest with correlation filtering. The results suggest that radiomics may be a potential support tool for EGFR classification when biopsy is not feasible or recommended. This would enable safer and more efficient personalized treatment. Nevertheless, the results underscore the need for larger, diverse datasets to improve model robustness for characterizing such complex and variable information before clinical integration. Madalena Pereira, Tânia Mendes, Venceslau Hespanhol, Hélder P. Oliveira, Tânia Pereira 0001 |
BIBM | 5 |
| 2025 | From CT Scans to 3D Printed Models: A Pipeline for Mandible Surgical PlanningabstractAccurate surgical planning is critical in mandibular reconstruction to restore the oncology patient's function and aesthetics. However, the use of physical three-dimensional (3D) models is often limited by time-consuming manual segmentation procedures or the high cost of commercial solutions. This work addresses the need for an accessible, quick, and low-cost pipeline to obtain a 3D printed model of the segmented mandible from a Computed Tomography (CT) scan. The automatic segmentation stage relied on the two-dimensional U-Net architecture, which was trained and validated with slices across two public datasets (PDDCA, HaN-Seg) and tested with the other two public datasets (TCIA RT, Austrian). The best model achieved an average dice similarity coefficient (DSC) of$0.912 \pm 0.077$across all test sets. The segmentation output was reconstructed into a 3D volume, improved through a post-processing method (with morphological closing, upsample, smoothing, and mesh reduction), and 3D printed through fused deposition modelling. The assessment of a stomatologist confirmed overall high anatomical fidelity to the CT and clinical utility, even though further improvements in important fine anatomical elements were suggested. This solution contributes to a promising alternative to producing 3D personalised mandibles for surgical planning, reducing time and manual effort while improving the quality and accessibility. Future work may explore the use of 3D DL architectures and a broader evaluation of the 3D mandible models. Ana Saraiva, Margarida Gouveia, Catarina Lopes, Jorge Marinho, Tânia Pereira 0001, Joaquim Mendes |
BIBM | 5 |
| 2025 | Swin Transformer Applied to Breast MRI Super-Resolution in a Cross-Cohort DatasetabstractAdvancements in the care for patients with breast cancer have demanded the development of biomechanical breast models for the planning and risk mitigation of such invasive surgical procedures. However, these approaches require large amounts of high-quality magnetic resonance imaging (MRI) training data that is of difficult acquisition and availability. Although this can be solved using synthetic data, generating high resolution images comes at the price of very high computational constraints and tipically low performances. On the other hand, producing lower resolution samples yields better results and efficiency but falls short of meeting health professional standards. Therefore, this work aims to validate a joint approach between lower resolution generative models and the proposed superresolution architecture, titled Shifted Window Image Restoration (SWinIR), which was used to achieve a$4 x$increase in image size of breast cancer patient MRI samples. Results prove to be promising and to further expand upon the super-resolution state-of-the-art, achieving good maximum peak signal-to-noise ratio of 41.36 and structural similarity index values of 0.962 and thus beating traditional methods and other machine learning architectures. Henrique Sousa, Tânia Pereira 0001, Eva Batista, Pedro Gouveia, Hélder P. Oliveira |
CBMS | 3 |
| 2025 | Beyond Accuracy: The Role of Calibration in Computational PathologyabstractDeep learning in computational pathology (CPath) has rapidly advanced in recent years. Research has primarily focused on enhancing accuracy and interpretability across various histology image analysis tasks, from tile-level to slide-level foundation models and novel multiple instance learning (MIL) strategies. However, it is equally important for models to provide well-calibrated confidence estimates. Due to factors such as dataset bias, overfitting, and limited training data, existing models tend to be overly confident on test sets. Promising solutions to address this issue include temperature scaling, a post-hoc method that adjusts logits using a single scalar value. However, the role of calibration in CPath is yet to be clarified. In this study, we evaluate temperature scaling and linear temperature scaling for CPath tasks, analyzing their impact on recalibration in both in-domain and out-of-domain distributions. The results show the limitations of current probability calibration techniques and motivate future work. João D. Nunes, Diana Montezuma, Domingos Oliveira, Tânia Pereira 0001, Inti Zlobec, Jaime S. Cardoso 0001 |
IJCNN | 4 |
| 2025 | A survey on cell nuclei instance segmentation and classification: Leveraging context and attention
João D. Nunes, Diana Montezuma, Domingos Oliveira, Tânia Pereira 0001, Jaime S. Cardoso 0001 |
Medical Image Anal. | 4 |
| 2024 | CNN-based Methods for Survival Prediction using CT images for Lung Cancer PatientsabstractLung Cancer (LC) is still among the top main causes of death worldwide, and it is the leading death number among other cancers. Several AI-based methods have been developed for the early detection of LC, trying to use Computed Tomography (CT) images to identify the initial signs of the disease. The survival prediction could help the clinicians to adequate the treatment plan and all the proceedings, by the identification of the most severe cases that need more attention. In this study, several deep learning models were compared to predict the survival of LC patients using CT images. The best performing model, a CNN with 3 layers, achieved an AUC value of 0.80, a Precision value of 0.56 and a Recall of 0.64. The obtained results showed that CT images carry information that can be used to assess the survival of LC. Maria Amaro, Hélder P. Oliveira, Tânia Pereira 0001 |
CBMS | 3 |
| 2024 | Exploring the differences between Multi-task and Single-task with the use of Explainable AI for lung nodule classificationabstractCurrently, lung cancer is one of the deadliest diseases that affects millions of people globally. However, Artificial Intelligence is being increasingly integrated with healthcare practices, with the goal to aid in the early diagnosis of lung cancer. Although such methods have shown very promising results, they still lack transparency to the user, which consequently could make their generalised adoption a challenging task. Therefore, in this work we explore the use of post-hoc explainable methods, to better understand the inner-workings of an already established multitasking framework that executes the segmentation and the classification task of lung nodules simultaneously. The idea behind such study is to understand how a multitasking approach impacts the model’s performance in the lung nodule classification task when compared to single-task models. Our results show that the multitasking approach works as an attention mechanism by aiding the model to learn more meaningful features. Furthermore, the multitasking framework was able to achieve a better performance in regard to the explainability metric, with an increase of 7% when compared to our baseline, and also during the classification and segmentation task, with an increase of 4.84% and 15.03% for each task respectively, when also compared to the studied baselines. Luís Fernandes, Tânia Pereira 0001, Hélder P. Oliveira |
CBMS | 2 |
| 2024 | Deep Learning Models to Predict Brain Cancer Grade Through MRI AnalysisabstractThe early and accurate detection and the grading characterization of brain cancer will generate a positive impact on the treatment plan of those patients. AI-based models can help analyze the Magnetic Resonance Imaging (MRI) to make an initial assessment of the tumor grading. The objective of this work was to develop an AI-based model to classify the grading of the tumor using the MRI. Two regions of interest were explored, with several levels of complexity for the neural network architecture, and two strategies to deal with imbalanced data. The best results were obtained for the most complex architecture (Resnet50) with a combination of weighted random sampler and data augmentation achieving a balanced accuracy of 62.26%. This work confirmed that complex problems required a more dense neural network and strategies to deal with the imbalanced data. Pedro Vale, Jennifer Boer, Hélder P. Oliveira, Tânia Pereira 0001 |
CBMS | 4 |
| 2023 | Patch-based CNN Models for Bone Marrow Edema Detection Using MRIabstractBone marrow edema (BME) or bone marrow lesion is the term attributed to an observed signal change within the bone marrow in magnetic resonance imaging (MRI). BME can be originated from multiple mechanisms, with pain being the main symptom. The presence of BME is an unspecific but sensitive sign with a wide differential diagnosis, that may act as a guide that leads to a systematic and correct interpretation of the magnetic resonance examination. An automatic approach for BME detection and quantification aims to reduce the overload of clinicians, decreasing human error and accelerating the time to the correct diagnosis. In this work, the bone region on the MRI slice was split into several patches and a CNN-based model was trained to detect BME in each patch from the MRI slice. The learning model developed achieved an AUC of 0.853 ± 0.056, showing that the CNN-based model can be used to detect BME in the MRI and confirming the patch strategy implemented to deal with the small data size and allowing the neural network to learn the specific information related with the classification task by reducing the region of the image to be considered. A learning model that can help clinicians with BME identification will decrease the time and the error for the diagnosis, and represent the first step for a more objective assessment of the BME. André Gomes, Tânia Pereira 0001, Francisco Silva 0002, Pedro Franco, Diogo Costa Carvalho, Sílvia Costa Dias, Hélder P. Oliveira |
BIBM | 2 |
| 2023 | AI-based Models to Predict the Heart Rate Using PPG and Accelerometer Signals During Physical ExerciseabstractPPG signal is a valuable resource for continuous heart rate monitoring; however, this signal suffers from artifact movements, which is particularly relevant during physical exercise and makes this biomedical signal difficult to use for heart rate detection during those activities. The purpose of this study was to develop learning models to determine heart rate using data from wearables (PPG and acceleration signals) and dealing with noise during physical exercise. Learning models based on CNNs and LSTMs were developed to predict the heart rate. The PPG signal was combined with data from accelerometers trying to overcome the noise movement on the PPG signal. Two datasets were used on this work: the 2015 IEEE Signal Processing Cup (SPC) dataset was used for training and testing, and another dataset was used for validation of the learning model (PPG-DaLiA dataset). The predictions obtained by the learning model represented a mean average error of 7.033±5.376 bpm for the SCP dataset, while a mean average error of 9.520±8.443 bpm for the validation set. The use of acceleration data increases the performance of the learning models on the prediction of the heart rate, showing the benefits of using this source of data to overcome the noise movement problem on the PPG signal. The combination of PPG signal with acceleration data could allow the learning models to use more information regarding the motion artifacts that affect the PPG and improve performance on the physiological event detections, which will largely spread the use of wearables on the healthcare applications for continuous monitor the physiological state allowing early and accurate detection of pathological events. Hélder P. Oliveira, Xiao Hu 0002, Tânia Pereira 0001 |
BIBM | 4 |
| 2023 | A Machine Learning Approach for Predicting Microsatellite Instability using RNA-seqabstractMicrosatellite Instability (MSI) is an important biomarker in cancer patients, showing a defective DNA mismatch repair system. Its detection allows the use of immunotherapy to treat cancer, an approach that is revolutionizing cancer treatment. MSI is especially relevant for three types of cancer: Colon Adenocarcinoma (COAD), Stomach Adenocarcinoma (STAD), and Uterus corpus endometrial cancer (UCEC). In this work, learning algorithms were employed to predict MSI using RNA-seq data from The Cancer Genome Atlas (TCGA) database, with a focus on the selection of the most informative genomic features. The Multi-Layer Perceptron (MLP) obtained the best score (AUC = 98.44%), showing that it is possible to exploit information from RNA-seq data to find relevant relationships with the instability levels of microsatellites (MS). The accurate prediction of MSI with transcription data from cancer patients will help with the correct determination of MSI status and adequate prescription of immunotherapy, creating more precise and personalized patient care. At the genetic level, the study revealed a high expression of genes related to cell regulation functions, and a low expression of genes responsible for Mismatch Repair functions, in patients with high instability. Miguel Simões, Tânia Pereira 0001, Francisco Silva 0002, José Machado 0001, Hélder P. Oliveira |
BIBM | 2 |
| 2023 | Lung CT image synthesis using GANs
José Mendes, Tânia Pereira 0001, Francisco Silva 0002, Julieta Frade, Joana Morgado, Cláudia Freitas, Eduardo Negrão, Beatriz Flor De Lima, Miguel Correia Da Silva, António J. Madureira, José Luís Costa, Venceslau Hespanhol, António Cunha, Hélder P. Oliveira |
Expert Syst. Appl. | 2 |
| 2021 | Stacking Approach for Lung Cancer EGFR Mutation Status Prediction from CT ScansabstractDue to the huge mortality rate of lung cancer, there is a strong need for developing solutions that help with the early diagnosis and the definition of the most appropriate treatment. In the particular case of target therapy, effective genotyping of the tumor is fundamental since this treatment uses targeted drugs that can induce death in cancer cells. The biopsy is the traditional method to assess the genotype information but it is extremely invasive and painful. Medical imaging is a valuable alternative to biopsies, considering the potential to extract imaging features correlated with specific genomic alterations. Regarding the limitations of single model approaches for gene mutation status predictions, ensemble strategies might bring valuable benefits by combining the strengths and weaknesses of the aggregated methods. This preliminary work aims to provide further advances in the radiogenomics field by studying the use of ensemble methods to predict the Epidermal Growth Factor Receptor (EGFR) mutation status in lung cancer. The best result obtained for the proposed ensemble approach was an AUC of 0.706 (± 0.122). However, the ensemble did not outperform the single models with AUC values of 0.712 (± 0.119) for Logistic Regression, 0.711 (± 0.119) for Support Vector Machine and 0.712 (± 0.120) for Elastic Net. The high correlation found on the decisions of each single model might be a plausible explanation for this behavior, which caused the ensemble to misclassify the same examples as the single models. Alexandra Ventura, Tânia Pereira 0001, Francisco Silva 0002, Cláudia Freitas, António Cunha, Hélder P. Oliveira |
BIBM | 2 |
| 2020 | A Supervised Approach to Robust Photoplethysmography Quality AssessmentabstractEarly detection of Atrial Fibrillation (AFib) is crucial to prevent stroke recurrence. New tools for monitoring cardiac rhythm are important for risk stratification and stroke prevention. As many of new approaches to long-term AFib detection are now based on photoplethysmogram (PPG) recordings from wearable devices, ensuring high PPG signal-to-noise ratios is a fundamental requirement for a robust detection of AFib episodes. Traditionally, signal quality assessment is often based on the evaluation of similarity between pulses to derive signal quality indices. There are limitations to using this approach for accurate assessment of PPG quality in the presence of arrhythmia, as in the case of AFib, mainly due to substantial changes in pulse morphology. In this paper, we first tested the performance of algorithms selected from a body of studies on PPG quality assessment using a dataset of PPG recordings from patients with AFib. We then propose machine learning approaches for PPG quality assessment in 30-s segments of PPG recording from 13 stroke patients admitted to the University of California San Francisco (UCSF) neuro intensive care unit and another dataset of 3764 patients from one of the five UCSF general intensive care units. We used data acquired from two systems, fingertip PPG (fPPG) from a bedside monitor system, and radial PPG (rPPG) measured using a wearable commercial wristband. We compared various supervised machine learning techniques including k-nearest neighbors, decisions trees, and a two-class support vector machine (SVM). SVM provided the best performance. fPPG signals were used to build the model and achieved 0.9477 accuracy when tested on the data from the fPPG exclusive to the test set, and 0.9589 accuracy when tested on the rPPG data. Tânia Pereira 0001, Kais Gadhoumi, Mitchell Ma, Xiuyun Liu, Rene Colorado, Kevin J. Keenan, Karl Meisel, Xiao Hu 0002 |
IEEE J. Biomed. Health Informatics | 1 |