Alejandro Santos-Díaz

dblp:199/0189 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-5235-7325ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021
YearPublicationVenuePosition
2025 Multi-Scale Genomic Signatures and Machine Learning for Enhanced Prediction of Antimicrobial Resistance
abstract
Antimicrobial resistance (AMR) is a growing global health challenge that necessitates accurate computational methods for predicting resistant phenotypes from genomic data. In this study, we evaluate the performance of traditional machine learning (ML) models-Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), and kNearest Neighbors (k-NN)—alongside deep learning (DL) approaches, including Multilayer Perceptron (MLP), Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN). We also introduce TabTransformer as an additional model for AMR classification. Our models leverage k-mer-based features, resistance gene presence/absence, and SNP data across three feature sets. Our results show that LR, RF, and SVM consistently achieved an over 90% accuracy, while k -NN underperformed (77.78 %) due to sensitivity to high-dimensional feature spaces. Among deep learning models, RNN demonstrated superior performance in high-dimensional settings (97.22 % accuracy on 1024 features), while CNN and MLP experienced performance declines. TabTransformer also exhibited robust accuracy across different feature sets, outperforming CNN and MLP. Increasing the number of selected features beyond 256 did not significantly improve performance, emphasizing the importance of feature selection. Future work will explore nucleotide models and further integration of biological metadata to enhance model interpretability and prediction accuracy.
Axel Alejandro Ramos García, Cuauhtémoc Licona Cassani, Alejandro Santos-Díaz
CBMS3
2025 Comparison of ConvNeXt and Vision-Language Models for Breast Density Assessment in Screening Mammography
abstract
Mammographic breast density classification is essential for cancer risk assessment but remains challenging due to subjective interpretation and inter-observer variability. This study compares multimodal and CNN-based methods for automated classification using the BI-RADS system, evaluating BioMedCLIP and ConvNeXt across three learning scenarios: zero-shot classification, linear probing with textual descriptions, and fine-tuning with numerical labels. Results show that zero-shot classification achieved modest performance, while the fine-tuned ConvNeXt model outperformed the BioMedCLIP linear probe. Although linear probing demonstrated potential with pretrained embeddings, it was less effective than full fine-tuning. These findings suggest that despite the promise of multimodal learning, CNN-based models with end-to-end fine-tuning provide stronger performance for specialized medical imaging. The study underscores the need for more detailed textual representations and domain-specific adaptations in future radiology applications.
Yusdivia Molina-Román, David Gómez-Ortiz, Ernestina Menasalvas Ruiz, José G. Tamez-Peña, Alejandro Santos-Díaz
CBMS5
2025 FuNTB: a functional network clustering tool for the analysis of genome-wide genetic variants in Mycobacterium tuberculosis
abstract
MOTIVATION: Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mtb), still claims around 1.25 million lives each year. The growing threat of drug resistance-often driven by single‑nucleotide polymorphisms (SNPs) in Mtb genomes underscores the need for high‑quality genomic data and powerful bioinformatics tools. We present FuNTB, a python‑based pipeline that detects non‑synonymous SNPs in Mtb and builds functional network clusters to reveal genotype-phenotype relationships. RESULTS: FuNTB profiles non‑synonymous SNPs at the gene level across user‑defined phenotypes, pinpointing both shared and unique mutations. It ingests annotated Variant Call Format (VCF) files or MTBseq outputs and merges them with clinical metadata to produce network‑XML files compatible with Cytoscape and Gephi. When applied to the CRyPTIC Mtb collection, FuNTB rapidly recovered established resistance genes and surfaced novel candidates, validating its utility for mapping genotype-phenotype associations. AVAILABILITY AND IMPLEMENTATION: FuNTB is implemented in Python 3.8+ and is freely available under the MIT license at https://doi.org/10.5281/zenodo.15399917.
Axel Alejandro Ramos García, Paulina M. Mejía-Ponce, Nelly Sélem-Mojica, Alejandro Santos-Díaz, Emmanuel Martínez 0001, Cuauhtémoc Licona Cassani
Bioinform.4
2024 Named Entity Recognition in Mammography Radiology Reports using a Multilingual Transfer Learning Approach
abstract
This study explores a multilingual transfer learning strategy for Named Entity Recognition (NER) in mammography radiology reports, aiming to improve breast cancer diagnosis. By utilizing a dataset from TecSalud, which includes mammograms and Electronic Health Records (EHRs) over ten years, this study seeks to address the linguistic barriers in medical documentation through advanced Natural Language Processing (NLP) models. Our approach involves meticulously labeling twenty-four distinct entities within the predominantly Spanish dataset, covering a range of diagnostic features and interpretive findings, highlighting the challenge of linguistic diversity in medical records and the potential of NLP to bridge this gap.The results demonstrate that fine-tuning on the last layer offers a balanced approach between simplicity and accuracy, avoiding overfitting and achieving state-of-art results.
Esteban Ricardo Salazar Cabrera, Alejandro Santos-Díaz, Ernestina Menasalvas Ruiz, José G. Tamez-Peña, Víctor Robles
CBMS2
2024 Characterization of hippocampal local field potentials using fractal dimension analysis
abstract
Alzheimer’s disease (AD) is one of the most common neurodegenerative disorders, affecting more than 50 million people worldwide. Current detection methods are often inefficient and inaccessible or invasive and inadequate for the early stages, as pathological biomarkers have not been developed. This study introduces a novel method to characterize hippocampal local field potentials (LFPs) to aid in the early detection of AD. We used fractal dimension (FD) analysis to process LFP recordings of the hippocampus region of animal models, recorded in both basal and kainic acid-induced active states. The LFP signals were classified using time-series clustering to identify the existence of more than one trend. The FD values of different states of an LFP in an AD model were analyzed. The results show that the transition phase triggers fluctuations in the FD value despite the basal- and active-state values remaining consistent. The findings of this analysis provide a promising basis for the development of a digital biomarker for conditions that disrupt normal brain behavior, such as AD. This approach could potentially improve our understanding of these diseases and contribute to more effective diagnostic tools.
Pedro Flores-Ortiz, Luis Montesinos, Arturo G. Isla, Alejandro Santos-Díaz, Luis Enrique Arroyo-Garcia
CBMS4