Helena Aidos

dblp:18/8001 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
7since 2021 · last 2024
0000-0001-6827-4217ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-authorArtificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2024 iDPP@CLEF 2024: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi
ECIR (6)1
2024 Deep Temporal Consensus Clustering for Patient Stratification in Amyotrophic Lateral Sclerosis
abstract
Amyotrophic Lateral Sclerosis (ALS) is a fast-acting neurodegenerative disease, characterized by loss of muscle movement and heterogeneity in disease evolution.This poses a challenge in predicting the best time for therapy administration.Here, we propose Deep Temporal Consensus Clustering (DTCC), a stratification method to uncover patient groups with similar disease progression.Using only the initial 6-month follow-up period, DTCC uncovered five clusters that were evaluated in terms of disease evolution and time-to-event.For three critical events (non-invasive ventilation, gastrostomy and death) the attained groups show distinct 10year progressions, validating the approach.
Miguel Pego Roque, Andreia S. Martins, Marta Gromicho, Mamede de Carvalho, Sara C. Madeira, Pedro Tomás, Helena Aidos
ESANN7
2024 Biclustering data analysis: a comprehensive survey
abstract
Biclustering, the simultaneous clustering of rows and columns of a data matrix, has proved its effectiveness in bioinformatics due to its capacity to produce local instead of global models, evolving from a key technique used in gene expression data analysis into one of the most used approaches for pattern discovery and identification of biological modules, used in both descriptive and predictive learning tasks. This survey presents a comprehensive overview of biclustering. It proposes an updated taxonomy for its fundamental components (bicluster, biclustering solution, biclustering algorithms, and evaluation measures) and applications. We unify scattered concepts in the literature with new definitions to accommodate the diversity of data types (such as tabular, network, and time series data) and the specificities of biological and biomedical data domains. We further propose a pipeline for biclustering data analysis and discuss practical aspects of incorporating biclustering in real-world applications. We highlight prominent application domains, particularly in bioinformatics, and identify typical biclusters to illustrate the analysis output. Moreover, we discuss important aspects to consider when choosing, applying, and evaluating a biclustering algorithm. We also relate biclustering with other data mining tasks (clustering, pattern mining, classification, triclustering, N-way clustering, and graph mining). Thus, it provides theoretical and practical guidance on biclustering data analysis, demonstrating its potential to uncover actionable insights from complex datasets.
Eduardo N. Castanho, Helena Aidos, Sara C. Madeira
Briefings Bioinform.2
2023 iDPP@CLEF 2023: The Intelligent Disease Progression Prediction Challenge
Helena Aidos, Roberto Bergamaschi, Paola Cavalla, Adriano Chiò, Arianna Dagliati, Barbara Di Camillo, Mamede de Carvalho, Nicola Ferro 0001, Piero Fariselli, Jose Manuel García Dominguez, Sara C. Madeira, Eleonora Tavazzi
ECIR (3)1
2023 Artificial intelligence and statistical methods for stratification and prediction of progression in amyotrophic lateral sclerosis: A systematic review
abstract
BACKGROUND: Amyotrophic Lateral Sclerosis (ALS) is a fatal neurodegenerative disorder characterised by the progressive loss of motor neurons in the brain and spinal cord. The fact that ALS's disease course is highly heterogeneous, and its determinants not fully known, combined with ALS's relatively low prevalence, renders the successful application of artificial intelligence (AI) techniques particularly arduous. OBJECTIVE: This systematic review aims at identifying areas of agreement and unanswered questions regarding two notable applications of AI in ALS, namely the automatic, data-driven stratification of patients according to their phenotype, and the prediction of ALS progression. Differently from previous works, this review is focused on the methodological landscape of AI in ALS. METHODS: We conducted a systematic search of the Scopus and PubMed databases, looking for studies on data-driven stratification methods based on unsupervised techniques resulting in (A) automatic group discovery or (B) a transformation of the feature space allowing patient subgroups to be identified; and for studies on internally or externally validated methods for the prediction of ALS progression. We described the selected studies according to the following characteristics, when applicable: variables used, methodology, splitting criteria and number of groups, prediction outcomes, validation schemes, and metrics. RESULTS: Of the starting 1604 unique reports (2837 combined hits between Scopus and PubMed), 239 were selected for thorough screening, leading to the inclusion of 15 studies on patient stratification, 28 on prediction of ALS progression, and 6 on both stratification and prediction. In terms of variables used, most stratification and prediction studies included demographics and features derived from the ALSFRS or ALSFRS-R scores, which were also the main prediction targets. The most represented stratification methods were K-means, and hierarchical and expectation-maximisation clustering; while random forests, logistic regression, the Cox proportional hazard model, and various flavours of deep learning were the most widely used prediction methods. Predictive model validation was, albeit unexpectedly, quite rarely performed in absolute terms (leading to the exclusion of 78 eligible studies), with the overwhelming majority of included studies resorting to internal validation only. CONCLUSION: This systematic review highlighted a general agreement in terms of input variable selection for both stratification and prediction of ALS progression, and in terms of prediction targets. A striking lack of validated models emerged, as well as a general difficulty in reproducing many published studies, mainly due to the absence of the corresponding parameter lists. While deep learning seems promising for prediction applications, its superiority with respect to traditional methods has not been established; there is, instead, ample room for its application in the subfield of patient stratification. Finally, an open question remains on the role of new environmental and behavioural variables collected via novel, real-time sensors.
Erica Tavazzi, Enrico Longato, Martina Vettoretti, Helena Aidos, Isotta Trescato, Chiara Roversi, Andreia S. Martins, Eduardo N. Castanho, Ruben Branco, Diogo F. Soares, Alessandro Guazzo, Giovanni Birolo, Daniele Pala, Pietro Bosoni, Adriano Chiò, Umberto Manera, Mamede de Carvalho, Bruno Miranda, Marta Gromicho, Inês Alves, Riccardo Bellazzi, Arianna Dagliati, Piero Fariselli, Sara C. Madeira, Barbara Di Camillo
Artif. Intell. Medicine4
2022 Biclustering fMRI time series: a comparative study
abstract
BACKGROUND: The effectiveness of biclustering, simultaneous clustering of rows and columns in a data matrix, was shown in gene expression data analysis. Several researchers recognize its potentialities in other research areas. Nevertheless, the last two decades have witnessed the development of a significant number of biclustering algorithms targeting gene expression data analysis and a lack of consistent studies exploring the capacities of biclustering outside this traditional application domain. RESULTS: This work evaluates the potential use of biclustering in fMRI time series data, targeting the Region × Time dimensions by comparing seven state-in-the-art biclustering and three traditional clustering algorithms on artificial and real data. It further proposes a methodology for biclustering evaluation beyond gene expression data analysis. The results discuss the use of different search strategies in both artificial and real fMRI time series showed the superiority of exhaustive biclustering approaches, obtaining the most homogeneous biclusters. However, their high computational costs are a challenge, and further work is needed for the efficient use of biclustering in fMRI data analysis. CONCLUSIONS: This work pinpoints avenues for the use of biclustering in spatio-temporal data analysis, in particular neurosciences applications. The proposed evaluation methodology showed evidence of the effectiveness of biclustering in finding local patterns in fMRI time series data. Further work is needed regarding scalability to promote the application in real scenarios.
Eduardo N. Castanho, Helena Aidos, Sara C. Madeira
BMC Bioinform.2
2022 Exploiting second-order dissimilarity representations for hierarchical clustering and visualization
Helena Aidos
Data Min. Knowl. Discov.1
2017 ECG-based Biometrics using a Deep Autoencoder for Feature Learning - An Empirical Study on Transferability
Afonso Eduardo, Helena Aidos, Ana Fred
ICPRAM2
2017 Discrimination of Alzheimer's Disease using longitudinal information
Helena Aidos, Ana Fred
Data Min. Knowl. Discov.1
2016 Efficient Evidence Accumulation Clustering for Large Datasets
abstract
The unprecedented collection and storage of data in electronic format has given rise to an interest in automated analysis for generation of knowledge and new insights. Cluster analysis is a good candidate since it makes as few assumptions about the data as possible. A vast body of work on clustering methods exist, yet, typically, no single method is able to respond to the specificities of all kinds of data. Evidence Accumulation Clustering (EAC) is a robust state of the art ensemble algorithm that has shown good results. However, this robustness comes with higher computational cost. Currently, its application is slow or restricted to small datasets. The objective of the present work is to scale EAC, allowing its applicability to big datasets, with technology available at a typical workstation. Three approaches for different parts of EAC are presented: a parallel GPU K-Means implementation, a novel strategy to build a sparse CSR matrix specialized to EAC and Single-Link based on Minimum Spanning Trees using an external memory sorting algorithm. Combining these approaches, the application of EAC to much larger datasets than before was accomplished.
Helena Aidos, Ana Fred
ICPRAM2
2015 Semi-Supervised Consensus Clustering for ECG Pathology Classification
Helena Aidos, André Lourenço, Diana Batista, Samuel Rota Bulò, Ana Fred
ECML/PKDD (3)1
2014 Learning Similarities by Accumulating Evidence in a Probabilistic Way
Helena Aidos, Ana Fred
CIARP1
2014 Identifying regions of interest for discriminating Alzheimer's disease from mild cognitive impairment
abstract
Alzheimer's disease (AD) is one of the most common types of dementia that affects elderly people, with no known cure. Early diagnosis of this disease is very important to improve patients' life quality and slow down the disease progression. Over the years, researchers have been proposing several techniques to analyze brain images, like FDG-PET, to automatically find changes in the brain activity. This paper compares regions of voxels identified by an expert with regions of voxels found automatically, in terms of corresponding classification accuracies based on three well-known classifiers. The automatic identification of regions is made by segmenting FDG-PET images, and extracting features that represent each of those regions. Experimental results show that the regions found automatically are very discriminative, outperforming results with expert's defined regions.
Helena Aidos, João M. M. Duarte, Ana Fred
ICIP1
2014 Feature Extraction in Pet Images for the Diagnosis of Alzheimer's Disease
abstract
Alzheimer’s disease accounts for an estimated 60% to 80% of cases of dementia. Recently, several computer-aided diagnosis systems have been developed, based on extracting information from FDG-PET scans. 3-dimensional FDG-PET images, under a voxel-as-feature approach, lead to high-dimensional features spaces which results in system performance problems. We propose a new approach for feature extraction of 3-dimensional images to improve the performance of a diagnosis system, and compare it with Gaussian pyramid technique. To evaluate the performance of our approach we applied it to a data base obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI). Experimental results have shown that the proposed approach is a good option for image feature extraction, outperforming the Gaussian pyramid technique.
João M. M. Duarte, Helena Aidos, Ana Fred
ICPRAM2
2013 Evidence Accumulation Approach applied to EEG Analysis
Helena Aidos, Carlos Carreiras, Hugo Silva 0001, Ana Fred
ICPRAM1
2013 The Area under the ROC Curve as a Criterion for Clustering Evaluation
Helena Aidos, Robert P. W. Duin, Ana Fred
ICPRAM1
2013 Morphological ECG Analysis for Attention Detection
abstract
The electroencephalogram (EEG) signal, acquired on the scalp, has been extensively used to understand cognitive function, and in particular attention. However, this type of signal has several drawbacks in a context of Physiological Computing, being susceptible to noise and requiring the use of impractical head-mounted apparatuses, which impacts normal human-computer interaction. For these reasons, the electrocardiogram (ECG) has been proposed as an alternative source to assess emotion, which is also continuously available, and related with the psychophysiological state of the subject. In this paper we present a study focused on the morphological analysis of the ECG signal acquired from subjects performing a task demanding high levels of attention. The analysis is made using various unsupervised learning techniques, which are validated against evidence found in a previous study by our team, where EEG signals collected for the same task exhibit distinct patterns as the subjects progress in the task.
Carlos Carreiras, André Lourenço, Helena Aidos, Hugo Silva 0001, Ana Fred
IJCCI3
2012 Classification using High Order Dissimilarities in Non-euclidean Spaces
Helena Aidos, Ana Fred, Robert P. W. Duin
ICPRAM (1)1
2012 Statistical modeling of dissimilarity increments for d-dimensional data: Application in partitional clustering
Helena Aidos, Ana Fred
Pattern Recognit.1
2010 An information retrieval perspective on visualization of gene expression data with ontological annotation
abstract
High-dimensional data are often visualized by dimensionality reduction methods whose goals are not directly related to visualization. We use a recent formalization of visualization as information retrieval and apply that formalism to data with structured annotations: we analyze gene expression data with annotations from the Gene Ontology (GO). We show that using the GO information in visualization yields better retrieval with respect to known ontological relationships and allows discovery of data properties not explained by the ontology.
Jaakko Peltonen, Helena Aidos, Nils Gehlenborg, Alvis Brazma, Samuel Kaski
ICASSP2
2010 Information Retrieval Perspective to Nonlinear Dimensionality Reduction for Data Visualization
Jarkko Venna, Jaakko Peltonen, Kristian Nybo, Helena Aidos, Samuel Kaski
J. Mach. Learn. Res.4
2009 Supervised nonlinear dimensionality reduction by Neighbor Retrieval
abstract
Many recent works have combined two machine learning topics, learning of supervised distance metrics and manifold embedding methods, into supervised nonlinear dimensionality reduction methods. We show that a combination of an early metric learning method and a recent unsupervised dimensionality reduction method empirically outperforms previous methods. In our method, the Riemannian distance metric measures local change of class distributions, and the dimensionality reduction method makes a rigorous tradeoff between precision and recall in retrieving similar data points based on the reduced-dimensional display. The resulting supervised visualizations are good for finding (sets of) similar data samples that have similar class distributions.
Jaakko Peltonen, Helena Aidos, Samuel Kaski
ICASSP2