VLDB 2026 Research / reviewers in the wild / expert
Claudia R. Solís-Lemus
dblp:181/4322
· DBLP profile ↗
6ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-9789-8915ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SNaQ.jl: Improved scalability for level-1 phylogenetic network inferenceabstractMOTIVATION: Phylogenetic networks represent complex biological scenarios that are overlooked in trees, such as hybridization and horizontal gene transfer. Although numerous methods have been developed for phylogenetic network inference, their scalability is severely limited by the computational demands of likelihood optimization and the vastness of network space. Composite (or pseudo-) likelihood approaches like SNaQ have improved computational tractability for network inference, but they remain inadequate for datasets of sizes routinely handled by tree inference methods. RESULTS: Here, we introduce SNaQ.jl, a new standalone Julia package with the composite likelihood inference originally implemented within PhyloNetworks.jl as well as new scalability features that enhance computational efficiency through (i) parallelization of quartet likelihood calculations during composite likelihood computation, (ii) weighted random selection of quartets, and (iii) probabilistic decision-making during network search. Through a simulation study and empirical data analysis, we show that this new version of SNaQ.jl (version 1.1) improves average runtimes by up to 499% on average with no change in function parameters or method accuracy. AVAILABILITY AND IMPLEMENTATION: SNaQ.jl is a new open source Julia package available at https://github.com/JuliaPhylo/SNaQ.jl. Nathan Kolbow, Sungsik Kong, Tyler Chafin, Joshua Justison, Cécile Ané, Claudia R. Solís-Lemus |
Bioinform. | 6 |
| 2025 | HighDimMixedModels.jl: Robust high-dimensional mixed-effects models across omics dataabstractHigh-dimensional mixed-effects models are an increasingly important form of regression in which the number of covariates rivals or exceeds the number of samples, which are collected in groups or clusters. The penalized likelihood approach to fitting these models relies on a coordinate descent algorithm that lacks guarantees of convergence to a global optimum. Here, we empirically study the behavior of this algorithm on simulated and real examples of three types of data that are common in modern biology: transcriptome, genome-wide association, and microbiome data. Our simulations provide new insights into the algorithm's behavior in these settings, and, comparing the performance of two popular penalties, we demonstrate that the smoothly clipped absolute deviation (SCAD) penalty consistently outperforms the least absolute shrinkage and selection operator (LASSO) penalty in terms of both variable selection and estimation accuracy across omics data. To empower researchers in biology and other fields to fit models with the SCAD penalty, we implement the algorithm in a Julia package, HighDimMixedModels.jl. Evan Gorstein, Rosa Aghdam, Claudia R. Solís-Lemus |
PLoS Comput. Biol. | 3 |
| 2024 | Human limits in machine learning: prediction of potato yield and disease using soil microbiome dataabstractBACKGROUND: The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. RESULTS: Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can't classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. CONCLUSIONS: Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management. Rosa Aghdam, Xudong Tang, Shan Shan, Richard Lankau, Claudia R. Solís-Lemus |
BMC Bioinform. | 5 |
| 2023 | Machine learning identification of Pseudomonas aeruginosa strains from colony image dataabstractWhen grown on agar surfaces, microbes can produce distinct multicellular spatial structures called colonies, which contain characteristic sizes, shapes, edges, textures, and degrees of opacity and color. For over one hundred years, researchers have used these morphology cues to classify bacteria and guide more targeted treatment of pathogens. Advances in genome sequencing technology have revolutionized our ability to classify bacterial isolates and while genomic methods are in the ascendancy, morphological characterization of bacterial species has made a resurgence due to increased computing capacities and widespread application of machine learning tools. In this paper, we revisit the topic of colony morphotype on the within-species scale and apply concepts from image processing, computer vision, and deep learning to a dataset of 69 environmental and clinical Pseudomonas aeruginosa strains. We find that colony morphology and complexity under common laboratory conditions is a robust, repeatable phenotype on the level of individual strains, and therefore forms a potential basis for strain classification. We then use a deep convolutional neural network approach with a combination of data augmentation and transfer learning to overcome the typical data starvation problem in biological applications of deep learning. Using a train/validation/test split, our results achieve an average validation accuracy of 92.9% and an average test accuracy of 90.7% for the classification of individual strains. These results indicate that bacterial strains have characteristic visual 'fingerprints' that can serve as the basis of classification on a sub-species level. Our work illustrates the potential of image-based classification of bacterial pathogens and highlights the potential to use similar approaches to predict medically relevant strain characteristics like antibiotic resistance and virulence from colony data. Jennifer B. Rattray, Ryan J. Lowhorn, Ryan Walden, Pedro Márquez-Zacarías, Evgeniya Molotkova, Gabriel Perron, Claudia R. Solís-Lemus, Daniel Pimentel Alarcon, Sam P. Brown |
PLoS Comput. Biol. | 7 |
| 2022 | Towards a robust out-of-the-box neural network model for genomic dataabstractBACKGROUND: The accurate prediction of biological features from genomic data is paramount for precision medicine and sustainable agriculture. For decades, neural network models have been widely popular in fields like computer vision, astrophysics and targeted marketing given their prediction accuracy and their robust performance under big data settings. Yet neural network models have not made a successful transition into the medical and biological world due to the ubiquitous characteristics of biological data such as modest sample sizes, sparsity, and extreme heterogeneity. RESULTS: Here, we investigate the robustness, generalization potential and prediction accuracy of widely used convolutional neural network and natural language processing models with a variety of heterogeneous genomic datasets. Mainly, recurrent neural network models outperform convolutional neural network models in terms of prediction accuracy, overfitting and transferability across the datasets under study. CONCLUSIONS: While the perspective of a robust out-of-the-box neural network model is out of reach, we identify certain model characteristics that translate well across datasets and could serve as a baseline model for translational researchers. Songyang Cheng, Claudia R. Solís-Lemus |
BMC Bioinform. | 3 |
| 2017 | Adversarial principal component analysisabstractThis paper studies the following question: where should an adversary place an outlier of a given magnitude in order to maximize the error of the subspace estimated by PCA? We give the exact location of this worst possible outlier, and the exact expression of the maximum possible error. Equivalently, we determine the information-theoretic bounds on how much an outlier can tilt a subspace in its direction. This in turn provides universal (worst-case) error bounds for PCA under arbitrary noisy settings. Our results also have several implications on adaptive PCA, online PCA, and rank-one updates. We illustrate our results with a subspace tracking experiment. Daniel L. Pimentel-Alarcón, Aritra Biswas, Claudia R. Solís-Lemus |
ISIT | 3 |