Adrián Díaz

dblp:311/0323 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
YearPublicationVenuePosition
2025 Combining evolution and protein language models for an interpretable cancer driver mutation prediction with D2Deep
abstract
The mutations driving cancer are being increasingly exposed through tumor-specific genomic data. However, differentiating between cancer-causing driver mutations and random passenger mutations remains challenging. State-of-the-art homology-based predictors contain built-in biases and are often ill-suited to the intricacies of cancer biology. Protein language models have successfully addressed various biological problems but have not yet been tested on the challenging task of cancer driver mutation prediction at a large scale. Additionally, they often fail to offer result interpretation, hindering their effective use in clinical settings. The AI-based D2Deep method we introduce here addresses these challenges by combining two powerful elements: (i) a nonspecialized protein language model that captures the makeup of all protein sequences and (ii) protein-specific evolutionary information that encompasses functional requirements for a particular protein. D2Deep relies exclusively on sequence information, outperforms state-of-the-art predictors, and captures intricate epistatic changes throughout the protein caused by mutations. These epistatic changes correlate with known mutations in the clinical setting and can be used for the interpretation of results. The model is trained on a balanced, somatic training set and so effectively mitigates biases related to hotspot mutations compared to state-of-the-art techniques. The versatility of D2Deep is illustrated by its performance on non-cancer mutation prediction, where most variants still lack known consequences. D2Deep predictions and confidence scores are available via https://tumorscope.be/d2deep to help with clinical interpretation and mutation prioritization.
Konstantina Tzavella, Adrián Díaz, Catharina Olsen, Wim F. Vranken
Briefings Bioinform.2
2024 Large-scale structure-informed multiple sequence alignment of proteins with SIMSApiper
abstract
SUMMARY: SIMSApiper is a Nextflow pipeline that creates reliable, structure-informed MSAs of thousands of protein sequences faster than standard structure-based alignment methods. Structural information can be provided by the user or collected by the pipeline from online resources. Parallelization with sequence identity-based subsets can be activated to significantly speed up the alignment process. Finally, the number of gaps in the final alignment can be reduced by leveraging the position of conserved secondary structure elements. AVAILABILITY AND IMPLEMENTATION: The pipeline is implemented using Nextflow, Python3, and Bash. It is publicly available on github.com/Bio2Byte/simsapiper.
Charlotte Crauwels, Sophie-Luise Heidig, Adrián Díaz, Wim F. Vranken
Bioinform.3
2024 bio2Byte Tools deployment as a Python package and Galaxy tool to predict protein biophysical properties
abstract
SUMMARY: We introduce a unified Python package for the prediction of protein biophysical properties, streamlining previous tools developed by the Bio2Byte research group. This suite facilitates comprehensive assessments of protein characteristics, incorporating predictors for backbone and sidechain dynamics, local secondary structure propensities, early folding, long disorder, beta-sheet aggregation, and fused in sarcoma (FUS)-like phase separation. Our package significantly eases the integration and execution of these tools, enhancing accessibility for both computational and experimental researchers. AVAILABILITY AND IMPLEMENTATION: The suite is available on the Python Package Index (PyPI): https://pypi.org/project/b2bTools/ and Bioconda: https://bioconda.github.io/recipes/b2btools/README.html for Linux and macOS systems, with Docker images hosted on Biocontainers: https://quay.io/repository/biocontainers/b2btools?tab=tags&tag=latest and Docker Hub: https://hub.docker.com/u/bio2byte. Online deployments are available on Galaxy Europe: https://usegalaxy.eu/root?tool_id=b2btools_single_sequence and our online server: https://bio2byte.be/b2btools/. The source code can be found at https://bitbucket.org/bio2byte/b2btools_releases.
Jose Gavaldá-García, Adrián Díaz, Wim F. Vranken
Bioinform.2
2023 The ACPYPE web server for small-molecule MD topology generation
abstract
MOTIVATION: The generation of parameter files for molecular dynamics (MD) simulations of small molecules that are suitable for force fields commonly applied to proteins and nucleic acids is often challenging. The ACPYPE software and website aid the generation of such parameter files. RESULTS: ACPYPE uses OpenBabel and ANTECHAMBER to generate MD input files in Gromacs, AMBER, CHARMM, and CNS formats. It can now take a SMILES string as input, in addition to the original PDB or mol2 coordinate files, with GAFF2 support and GLYCAM force field conversion added. It can be installed locally via Anaconda, PyPI, and Docker distributions, while the web server at https://bio2byte.be/acpype/ was updated with an API, and provides visualization of results for uploaded molecules as well as a pre-generated set of 3738 drug molecules. AVAILABILITY AND IMPLEMENTATION: The web application is freely available at https://www.bio2byte.be/acpype/ and the open-source code can be found at https://github.com/alanwilter/acpype.
Luciano Porto Kagami, Alan Wilter, Adrián Díaz, Wim F. Vranken
Bioinform.3
2022 Learning a Battery of COVID-19 Mortality Prediction Models by Multi-objective Optimization
Mario Martínez-García, Susana García-Gutierrez, Rubén Armañanzas, Adrián Díaz, Iñaki Inza, José Antonio Lozano 0001
AIME4
2021 Derivation of a Cost-Sensitive COVID-19 Mortality Risk Indicator Using a Multistart Framework
abstract
The overall global death rate for COVID-19 patients has escalated to 2.13% after more than a year of worldwide spread. Despite strong research on the infection pathogenesis, the molecular mechanisms involved in a fatal course are still poorly understood. Machine learning constitutes a perfect tool to develop algorithms for predicting a patient’s hospitalization outcome at triage. This paper presents a probabilistic model, referred to as a mortality risk indicator, able to assess the risk of a fatal outcome for new patients. The derivation of the model was done over a database of 2,547 patients from the first COVID-19 wave in Spain. Model learning was tackled through a five multistart configuration that guaranteed good generalization power and low variance error estimators. The training algorithm made use of a class weighting correction to account for the mortality class imbalance and two regularization learners, logistic and lasso regressors. Outcome probabilities were adjusted to obtain cost-sensitive predictions by minimizing the type II error. Our mortality indicator returns both a binary outcome and a three-stage mortality risk level. The estimated AUC across multistarts reaches an average of 0.907. At the optimal cutoff for the binary outcome, the model attains an average sensitivity of 0.898, with a 0.745 specificity. An independent set of 121 patients later released from the same consortium attained perfect sensitivity (1), with a 0.759 specificity when predicted by our model. Best performance for the indicator is achieved when the prediction’s time horizon is within two weeks since admission to hospital. In addition to a strong predictive performance, the set of selected features highlights the relevance of several underrated molecules in COVID-19 research, such as blood eosinophils, bilirubin, and urea levels.
Rubén Armañanzas, Adrián Díaz, Mario Martínez-García, Santiago Mazuelas
BIBM2