VLDB 2026 Research / reviewers in the wild / expert
Rosa Aghdam
dblp:132/6618
· DBLP profile ↗
6ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-9045-9592ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HighDimMixedModels.jl: Robust high-dimensional mixed-effects models across omics dataabstractHigh-dimensional mixed-effects models are an increasingly important form of regression in which the number of covariates rivals or exceeds the number of samples, which are collected in groups or clusters. The penalized likelihood approach to fitting these models relies on a coordinate descent algorithm that lacks guarantees of convergence to a global optimum. Here, we empirically study the behavior of this algorithm on simulated and real examples of three types of data that are common in modern biology: transcriptome, genome-wide association, and microbiome data. Our simulations provide new insights into the algorithm's behavior in these settings, and, comparing the performance of two popular penalties, we demonstrate that the smoothly clipped absolute deviation (SCAD) penalty consistently outperforms the least absolute shrinkage and selection operator (LASSO) penalty in terms of both variable selection and estimation accuracy across omics data. To empower researchers in biology and other fields to fit models with the SCAD penalty, we implement the algorithm in a Julia package, HighDimMixedModels.jl. Evan Gorstein, Rosa Aghdam, Claudia R. Solís-Lemus |
PLoS Comput. Biol. | 2 |
| 2024 | Human limits in machine learning: prediction of potato yield and disease using soil microbiome dataabstractBACKGROUND: The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. RESULTS: Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can't classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. CONCLUSIONS: Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management. Rosa Aghdam, Xudong Tang, Shan Shan, Richard Lankau, Claudia R. Solís-Lemus |
BMC Bioinform. | 1 |
| 2024 | AFITbin: a metagenomic contig binning method using aggregate l-mer frequency based on initial and terminal nucleotidesabstractBACKGROUND: Using next-generation sequencing technologies, scientists can sequence complex microbial communities directly from the environment. Significant insights into the structure, diversity, and ecology of microbial communities have resulted from the study of metagenomics. The assembly of reads into longer contigs, which are then binned into groups of contigs that correspond to different species in the metagenomic sample, is a crucial step in the analysis of metagenomics. It is necessary to organize these contigs into operational taxonomic units (OTUs) for further taxonomic profiling and functional analysis. For binning, which is synonymous with the clustering of OTUs, the tetra-nucleotide frequency (TNF) is typically utilized as a compositional feature for each OTU. RESULTS: In this paper, we present AFIT, a new l-mer statistic vector for each contig, and AFITBin, a novel method for metagenomic binning based on AFIT and a matrix factorization method. To evaluate the performance of the AFIT vector, the t-SNE algorithm is used to compare species clustering based on AFIT and TNF information. In addition, the efficacy of AFITBin is demonstrated on both simulated and real datasets in comparison to state-of-the-art binning methods such as MetaBAT 2, MaxBin 2.0, CONCOT, MetaCon, SolidBin, BusyBee Web, and MetaBinner. To further analyze the performance of the purposed AFIT vector, we compare the barcodes of the AFIT vector and the TNF vector. CONCLUSION: The results demonstrate that AFITBin shows superior performance in taxonomic identification compared to existing methods, leveraging the AFIT vector for improved results in metagenomic binning. This approach holds promise for advancing the analysis of metagenomic data, providing more reliable insights into microbial community composition and function. AVAILABILITY: A python package is available at: https://github.com/SayehSobhani/AFITBin . Amin Darabi, Sayeh Sobhani, Rosa Aghdam, Changiz Eslahchi |
BMC Bioinform. | 3 |
| 2023 | wpLogicNet: logic gate and structure inference in gene regulatory networksabstractMOTIVATION: The gene regulatory process resembles a logic system in which a target gene is regulated by a logic gate among its regulators. While various computational techniques are developed for a gene regulatory network (GRN) reconstruction, the study of logical relationships has received little attention. Here, we propose a novel tool called wpLogicNet that simultaneously infers both the directed GRN structures and logic gates among genes or transcription factors (TFs) that regulate their target genes, based on continuous steady-state gene expressions. RESULTS: wpLogicNet proposes a framework to infer the logic gates among any number of regulators, with a low time-complexity. This distinguishes wpLogicNet from the existing logic-based models that are limited to inferring the gate between two genes or TFs. Our method applies a Bayesian mixture model to estimate the likelihood of the target gene profile and to infer the logic gate a posteriori. Furthermore, in structure-aware mode, wpLogicNet reconstructs the logic gates in TF-gene or gene-gene interaction networks with known structures. The predicted logic gates are validated on simulated datasets of TF-gene interaction networks from Escherichia coli. For the directed-edge inference, the method is validated on datasets from E.coli and DREAM project. The results show that compared to other well-known methods, wpLogicNet is more precise in reconstructing the network and logical relationships among genes. AVAILABILITY AND IMPLEMENTATION: The datasets and R package of wpLogicNet are available in the github repository, https://github.com/CompBioIPM/wpLogicNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Seyed Amir Malekpour, Maryam Shahdoust, Rosa Aghdam, Mehdi Sadeghi |
Bioinform. | 3 |
| 2021 | A neural network-based method for polypharmacy side effects predictionabstractBACKGROUND: Polypharmacy is a type of treatment that involves the concurrent use of multiple medications. Drugs may interact when they are used simultaneously. So, understanding and mitigating polypharmacy side effects are critical for patient safety and health. Since the known polypharmacy side effects are rare and they are not detected in clinical trials, computational methods are developed to model polypharmacy side effects. RESULTS: We propose a neural network-based method for polypharmacy side effects prediction (NNPS) by using novel feature vectors based on mono side effects, and drug-protein interaction information. The proposed method is fast and efficient which allows the investigation of large numbers of polypharmacy side effects. Our novelty is defining new feature vectors for drugs and combining them with a neural network architecture to apply for the context of polypharmacy side effects prediction. We compare NNPS on a benchmark dataset to predict 964 polypharmacy side effects against 5 well-established methods and show that NNPS achieves better results than the results of all 5 methods in terms of accuracy, complexity, and running time speed. NNPS outperforms about 9.2% in Area Under the Receiver-Operating Characteristic, 12.8% in Area Under the Precision-Recall Curve, 8.6% in F-score, 10.3% in Accuracy, and 18.7% in Matthews Correlation Coefficient with 5-fold cross-validation against the best algorithm among other well-established methods (Decagon method). Also, the running time of the Decagon method which is 15 days for one fold of cross-validation is reduced to 8 h by the NNPS method. CONCLUSIONS: The performance of NNPS is benchmarked against 5 well-known methods, Decagon, Concatenated drug features, Deep Walk, DEDICOM, and RESCAL, for 964 polypharmacy side effects. We adopt the 5-fold cross-validation for 50 iterations and use the average of the results to assess the performance of the NNPS method. The evaluation of the NNPS against five well-known methods, in terms of accuracy, complexity, and running time speed shows the performance of the presented method for an essential and challenging problem in pharmacology. Datasets and code for NNPS algorithm are freely accessible at https://github.com/raziyehmasumshah/NNPS . Raziyeh Masumshah, Rosa Aghdam, Changiz Eslahchi |
BMC Bioinform. | 2 |
| 2019 | Some node ordering methods for the K2 algorithmabstractAbstract Inferring Bayesian network structure from data is a challenging issue, and many researchers have been working on this problem. The K2 is a well‐known order‐dependent algorithm to learn Bayesian network. The result of the algorithm is not robust since it achieves different network structure if node orderings are permuted. Consequently, choosing suitable sequential node ordering for the input of the K2 algorithm is a challenging task. In this work, some deterministic methods for selecting a suitable sequential node ordering are introduced. The effectiveness of these methods benchmarked through the Asia, Alarm, Car, and Insurance networks. The results indicate that the methods based on the concept of mutual information and entropy are suitable for finding a sequential node ordering and considerably improves the precision of network inference. The source code and selected data sets are available on http://profiles.bs.ipm.ir/softwares/ordering/ . Rosa Aghdam, Vahid Rezaei Tabar, Hamid Pezeshk |
Comput. Intell. | 1 |