Alexandra M. Carvalho

dblp:68/5595 · DBLP profile ↗
← Back
13ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0001-6607-7711ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-authorTheory of computation · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Causality in Categorical Data Using Geometric Complexity
Alexandra M. Carvalho, Diogo Cruz, Paulo Mateus, Bruno Mera
IDEAL (1)1
2023 Causal Graph Discovery for Explainable Insights on Marine Biotoxin Shellfish Contamination
Filipe Ferraz, Marta B. Lopes, Susana Rodrigues, Pedro Reis Costa, Susana Vinga, Alexandra M. Carvalho
IDEAL7
2023 Using Markov chains and temporal alignment to identify clinical patterns in Dementia
abstract
In the healthcare sector, resorting to big data and advanced analytics is a great advantage when dealing with complex groups of patients in terms of comorbidities, representing a significant step towards personalized targeting. In this work, we focus on understanding key features and clinical pathways of patients with multimorbidity suffering from Dementia. This disease can result from many heterogeneous factors, potentially becoming more prevalent as the population ages. We present a set of methods that allow us to identify medical appointment patterns within a cohort of 1924 patients followed from January 2007 to August 2021 in Hospital da Luz (Lisbon), and to stratify patients into subgroups that exhibit similar patterns of interaction. With Markov Chains, we are able to identify the most prevailing medical appointments attended by Dementia patients, as well as recurring transitions between these. To perform patient stratification, we applied AliClu, a temporal sequence alignment algorithm for clustering longitudinal clinical data, which allowed us to successfully identify patient subgroups with similar medical appointment activity. A feature analysis per cluster obtained allows the identification of distinct patterns and characteristics. This pipeline provides a tool to identify prevailing clinical pathways of medical appointments within the dataset, as well as the most common transitions between medical specialities within Dementia patients. This methodology, alongside demographic and clinical data, has the potential to provide early signalling of the most likely clinical pathways and serve as a support tool for health providers in deciding the best course of treatment, considering a patient as a whole.
Luísa Marote Costa, João Pedro Colaço, Alexandra M. Carvalho, Susana Vinga, Andreia Sofia Teixeira
J. Biomed. Informatics3
2022 Model Complexity in Statistical Manifolds: The Role of Curvature
abstract
Model complexity plays an essential role in its selection, namely, by choosing a model that fits the data and is also succinct. Two-part codes and the minimum description length have been successful in delivering procedures to single out the best models, avoiding overfitting. In this work, we pursue this approach and complement it by performing further assumptions in the parameter space. Concretely, we assume that the parameter space is a smooth manifold, and by using tools of Riemannian geometry, we derive a sharper expression than the standard one given by the stochastic complexity, where the scalar curvature of the Fisher information metric plays a dominant role. Furthermore, we compute a sharper approximation to the capacity for exponential families and apply our results to derive optimal dimensional reduction in the context of principal component analysis.
Bruno Mera, Paulo Mateus, Alexandra M. Carvalho
IEEE Trans. Inf. Theory3
2021 Learning dynamic Bayesian networks from time-dependent and time-independent data: Unraveling disease progression in Amyotrophic Lateral Sclerosis
abstract
Amyotrophic lateral sclerosis (ALS) is a neurodegenerative disease causing patients to quickly lose motor neurons. The disease is characterized by a fast functional impairment and ventilatory decline, leading most patients to die from respiratory failure. To estimate when patients should get ventilatory support, it is helpful to adequately profile the disease progression. For this purpose, we use dynamic Bayesian networks (DBNs), a machine learning model, that graphically represents the conditional dependencies among variables. However, the standard DBN framework only includes dynamic (time-dependent) variables, while most ALS datasets have dynamic and static (time-independent) observations. Therefore, we propose the sdtDBN framework, which learns optimal DBNs with static and dynamic variables. Besides learning DBNs from data, with polynomial-time complexity in the number of variables, the proposed framework enables the user to insert prior knowledge and to make inference in the learned DBNs. We use sdtDBNs to study the progression of 1214 patients from a Portuguese ALS dataset. First, we predict the values of every functional indicator in the patients' consultations, achieving results competitive with state-of-the-art studies. Then, we determine the influence of each variable in patients' decline before and after getting ventilatory support. This insightful information can lead clinicians to pay particular attention to specific variables when evaluating the patients, thus improving prognosis. The case study with ALS shows that sdtDBNs are a promising predictive and descriptive tool, which can also be applied to assess the progression of other diseases, given time-dependent and time-independent clinical observations.
Tiago Leão, Sara C. Madeira, Marta Gromicho, Mamede de Carvalho, Alexandra M. Carvalho
J. Biomed. Informatics5
2015 Polynomial-time algorithm for learning optimal tree-augmented dynamic Bayesian networks
José L. Monteiro, Susana Vinga, Alexandra M. Carvalho
UAI3
2014 Hybrid learning of Bayesian multinets for binary classification
Alexandra M. Carvalho, Pedro Adão, Paulo Mateus
Pattern Recognit.1
2011 Discriminative Learning of Bayesian Networks via Factorized Conditional Log-Likelihood
Alexandra M. Carvalho, Teemu Roos, Arlindo L. Oliveira, Petri Myllymäki
J. Mach. Learn. Res.1
2007 Learning bayesian networks consistent with the optimal branching
abstract
We introduce a polynomial-time algorithm to learn Bayesian networks whose structure is restricted to nodes with in-degree at most k and to edges consistent with the optimal branching, that we call consistent k-graphs (CkG). The optimal branching is used as an heuristic for a primary causality order between network variables, which is subsequently refined, according to a certain score, into an optimal CkG Bayesian network. This approach augments the search space exponentially, in the number of nodes, relatively to trees, yet keeping a polynomial-time bound. The proposed algorithm can be applied to scores that decompose over the network structure, such as the well known LL, MDL, AIC, BIC, K2, BD, BDe, BDeu and MIT scores. We tested the proposed algorithm in a classification task. We show that the induced classifier always score better than or the same as the Naive Bayes and Tree Augmented Naive Bayes classifiers. Experiments on the UCI repository show that, in many cases, the improved scores translate into increased classification accuracy.
Alexandra M. Carvalho, Arlindo L. Oliveira
ICMLA1
2006 RISOTTO: Fast Extraction of Motifs with Mismatches
Nadia Pisanti, Alexandra M. Carvalho, Laurent Marsan, Marie-France Sagot
LATIN2
2006 An Efficient Algorithm for the Identification of Structured Motifs in DNA Promoter Sequences
abstract
We propose a new algorithm for identifying cis-regulatory modules in genomic sequences. The proposed algorithm, named RISO, uses a new data structure, called box-link, to store the information about conserved regions that occur in a well-ordered and regularly spaced manner in the data set sequences. This type of conserved regions, called structured motifs, is extremely relevant in the research of gene regulatory mechanisms since it can effectively represent promoter models. The complexity analysis shows a time and space gain over the best known exact algorithms that is exponential in the spacings between binding sites. A full implementation of the algorithm was developed and made available online. Experimental results show that the algorithm is much faster than existing ones, sometimes by more than four orders of magnitude. The application of the method to biological data sets shows its ability to extract relevant consensi.
Alexandra M. Carvalho, Ana T. Freitas, Arlindo L. Oliveira, Marie-France Sagot
IEEE ACM Trans. Comput. Biol. Bioinform.1
2005 A highly scalable algorithm for the extraction of CIS-regulatory regions
Alexandra M. Carvalho, Ana T. Freitas, Arlindo L. Oliveira, Marie-France Sagot
APBC1
2004 Efficient Extraction of Structured Motifs Using Box-Links
Alexandra M. Carvalho, Ana T. Freitas, Arlindo L. Oliveira, Marie-France Sagot
SPIRE1