EDBT 2026 Demo / reviewers in the wild / expert
Rubén Armañanzas
dblp:78/1152
· DBLP profile ↗
19ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0003-4049-0000ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bidirectional Floating Feature Selection Guided by Uncertainty Quantification
Marcos Lopez-De-Castro, José González-Gomariz, Alberto García-Galindo, Kewei Ni, Farnoosh Abbas Aghababazadeh, Benjamin Haibe-Kains, Rubén Armañanzas |
AIME (2) | 7 |
| 2026 | Conformal Recursive Feature EliminationabstractIn this work, we introduce a novel feature selection method that leverages the conformal prediction framework. This new method, named Conformal Recursive Feature Elimination (CRFE), recursively identifies and removes features that increase the non-conformity of a dataset measured under additively decomposable non-conformity functions, re-estimating feature relevance after each elimination step analogously to classical Recursive Feature Elimination (RFE). We also introduce a geometrically motivated stopping criterion for recursive feature selectors. CRFE outputs are compared with the classical RFE algorithm and the results of other state-of-the-art feature selectors across several benchmark and real-world datasets. Our experiments show that CRFE generally improves prediction-set efficiency and subset stability relative to RFE, while maintaining empirical coverage close to the nominal level. Finally, we show how the automatic stopping criterion effectively halted the recursive selection of features, selecting subsets of effective and non-redundant features. • A novel feature selector named CRFE is presented. • CRFE leverages conformal prediction to recursively remove features. • CRFE improves standard RFE in terms of performance and stability. • A data-driven stopping criterion is presented and validated. Marcos Lopez-De-Castro, Alberto García-Galindo, Rubén Armañanzas |
Pattern Recognit. | 3 |
| 2025 | Conformal inference for reliable single cell RNA-seq annotationabstractMOTIVATION: Despite the inherent complexity associated to automatic cell type assignments, most supervised learning models overlook rigorous uncertainty quantification on the annotations. Although some existing pipelines incorporate rejection options under predefined circumstances, they usually rely on arbitrary assumptions and do not provide statistical guarantees. In this work, we propose a methodology based on the conformal prediction framework to provide reliable single-cell annotations. Conformal prediction provides statistical guarantees on the outcome predictions without making any assumption about the underlying distribution of the data. Our methodological proposal leverages conformal inference to address two critical challenges in single-cell RNA sequencing annotations: (i) detect out-of-distribution cell types in the query data; and, (ii) perform reliable uncertainty quantification of the cell annotations through well-calibrated prediction sets. RESULTS: We evaluated the anomaly detector and the uncertainty-aware annotator in 10 batched experiments derived from various tissues. Specifically, we studied three different annotation taxonomies (standard, classwise, and cluster) alongside three different non-conformity measures. The results showed that our anomaly detector effectively identified previously unseen cell types, producing well-calibrated prediction sets. This rigorous annotation helped maintain coverage probabilities at the expected significance level. Finally, we illustrate how the integration of conformal prediction outputs enhanced further downstream analyses. AVAILABILITY AND IMPLEMENTATION: The automatic scRNA-seq annotator is available at https://github.com/digital-medicine-research-group-UNAV/conformalized_single_cell_annotator and https://doi.org/10.5281/zenodo.15870599. Marcos Lopez-De-Castro, Alberto García-Galindo, José González-Gomariz, Rubén Armañanzas |
Bioinform. | 4 |
| 2025 | Fair prediction sets through multi-objective hyperparameter optimizationabstractAbstract The widespread implementation of machine learning in safety-critical domains has raised ethical concerns regarding algorithmic discrimination. In such settings, the integration of fairness-aware algorithms with uncertainty quantification tools enables the development of reliable and safe decision-making. In this paper, we introduce a novel methodology that combines conformal prediction, offering rigorous prediction sets, with multi-objective optimization via evolutionary learning. The proposed meta-algorithm optimizes the hyperparameter configuration of classifiers to produce confidence predictors that balance efficiency and equalized coverage guarantees, addressing fairness concerns related to sensitive attributes. We empirically evaluate our methodology with four real-world problems and demonstrate its efficacy in exploring this trade-off and producing a repertoire of Pareto optimal conformal predictors. In this way, our contribution offers different modeling alternatives from which to choose depending on the policy adopted by stakeholders, thus illustrating its capability to enhance equitable decision-making. Alberto García-Galindo, Marcos Lopez-De-Castro, Rubén Armañanzas |
Mach. Learn. | 3 |
| 2025 | RNACOREX - RNA coregulatory network explorer and classifierabstractMicro-RNAs (miRNA) and their relationship with messenger RNAs (mRNA) have been widely associated with disease development and progression. Post-transcriptional coregulatory networks are sets of miRNA-mRNA interactions that regulate specific genetic behaviors through their combined activity. However, identifying reliable sets of such interactions associated with specific diseases remains challenging, partly due to the high rate of false positives and the lack of user-friendly tools developed for this purpose. In this work, we introduce a new Python package called RNACOREX (RNA CORegulatory network EXplorer and classifier). RNACOREX is a new, easy-to-use tool that allows researchers to find disease associated post-transcriptional coregulatory networks and use them to classify new unseen observations of miRNA and mRNA quantifications. RNACOREX combines structural information from curated databases with expression data analysis, using conditional mutual information to infer reliable sets of miRNA-mRNA interactions. These sets are then used to build probabilistic models based on Conditional Linear Gaussian (CLG) classifiers, which allow both prediction on new samples and validation of the inferred networks. To demonstrate its capabilities, we tested RNACOREX in 13 different databases from the The Cancer Genome Atlas Program, generating the associated post-transcriptional coregulatory networks and extracting classification performance metrics for each tumor type. Specifically, we used RNACOREX to classify patients according to their survival time in each cancer type, highlighting miRNA-mRNA interactions that consistently appeared across different cancer types. The results show that RNACOREX achieves competitive predictive performance compared to widely used classification algorithms, while offering the added benefit of interpretability through its graph-based modeling framework. Aitor Oviedo-Madrid, José González-Gomariz, Rubén Armañanzas |
PLoS Comput. Biol. | 3 |
| 2023 | Network Community Detection in Connectomics Data using Graph TheoryabstractThis work presents a comprehensive descriptive analysis of a complex network derived from individual neuronal reconstructions from the vinegar fly (Drosophila melanogaster) using graph theory. The primary aim was to establish an informed criterion for selecting the most suitable community detection algorithm based on the network’s topology. We implemented four state-of-the-art community detection algorithms, namely Louvain, Greedy Modularity, Label Propagation, and Infomap, and analyzed the intrinsic architecture and organization of the outcomes. In addition, to offer a more robust and consensual perspective on the results, we introduced a novel Rough Clustering method that allows a refined interpretation of communities within the overall network structure. Leandro González-Montesino, Darian H. Grass-Boada, Rubén Armañanzas |
BIBM | 3 |
| 2023 | Electronic Near Point of Convergence: A Tool to Aid the Assessment of ConcussionabstractNear Point of Convergence (NPC) is a widely used test when assessing oculomotor dysfunction in patients with mild Traumatic Brain Injury (mTBI) or concussion. A new electronic version of the Near Point Convergence test (eNPC) was developed to automatically measure NPC distance, addressing the need to improve the administration and precision of NPC testing. Age-adjusted normative values were achieved by using goodness-of-fit and generalized linear regression models. Target subjects included males and females aged 13 to 50 years. Normative values were computed as a function of age for records from 230 healthy volunteers. eNPC values from an independent cohort of 237 patients clinically deemed to have concussion and 245 matched controls were transformed into z-scores relative to age-expected normal eNPC values and then compared to demonstrate performance in a clinical population. Mean differences between matched controls and concussed patients showed a highly significant difference for eNPC z-scores between groups. The maximum Odds-Ratio (OR) value for z-scores was reached in the cutoff interval 0.75 to 1 with a mean of 2.614 and a median of 2.627. Usability for the eNPC test was high with a rate of obtaining measurements between 92% and 98%. Using z-scores helped reduce variance of the measurement and accounted for the effect of age. Significant separation between concussed patients and matched controls validated the use of the eNPC mathematical model. Further, high OR values using z-scores demonstrated potential improved clinical utility of the eNPC in the prediction of oculomotor dysfunction. Leslie S. Prichep, Saloni Kanakia, Jeffrey J. Bazarian, Tracey Covassin, R. J. Elbin, Gillian A. Hotz, Jeb F. Struder, Rubén Armañanzas |
BIBM | 9 |
| 2022 | Learning a Battery of COVID-19 Mortality Prediction Models by Multi-objective Optimization
Mario Martínez-García, Susana García-Gutierrez, Rubén Armañanzas, Adrián Díaz, Iñaki Inza, José Antonio Lozano 0001 |
AIME | 3 |
| 2021 | Derivation of a Cost-Sensitive COVID-19 Mortality Risk Indicator Using a Multistart FrameworkabstractThe overall global death rate for COVID-19 patients has escalated to 2.13% after more than a year of worldwide spread. Despite strong research on the infection pathogenesis, the molecular mechanisms involved in a fatal course are still poorly understood. Machine learning constitutes a perfect tool to develop algorithms for predicting a patient’s hospitalization outcome at triage. This paper presents a probabilistic model, referred to as a mortality risk indicator, able to assess the risk of a fatal outcome for new patients. The derivation of the model was done over a database of 2,547 patients from the first COVID-19 wave in Spain. Model learning was tackled through a five multistart configuration that guaranteed good generalization power and low variance error estimators. The training algorithm made use of a class weighting correction to account for the mortality class imbalance and two regularization learners, logistic and lasso regressors. Outcome probabilities were adjusted to obtain cost-sensitive predictions by minimizing the type II error. Our mortality indicator returns both a binary outcome and a three-stage mortality risk level. The estimated AUC across multistarts reaches an average of 0.907. At the optimal cutoff for the binary outcome, the model attains an average sensitivity of 0.898, with a 0.745 specificity. An independent set of 121 patients later released from the same consortium attained perfect sensitivity (1), with a 0.759 specificity when predicted by our model. Best performance for the indicator is achieved when the prediction’s time horizon is within two weeks since admission to hospital. In addition to a strong predictive performance, the set of selected features highlights the relevance of several underrated molecules in COVID-19 research, such as blood eosinophils, bilirubin, and urea levels. Rubén Armañanzas, Adrián Díaz, Mario Martínez-García, Santiago Mazuelas |
BIBM | 1 |
| 2019 | PaperBot: open-source web-based search and metadata organization of scientific literatureabstractBACKGROUND: The biomedical literature is expanding at ever-increasing rates, and it has become extremely challenging for researchers to keep abreast of new data and discoveries even in their own domains of expertise. We introduce PaperBot, a configurable, modular, open-source crawler to automatically find and efficiently index peer-reviewed publications based on periodic full-text searches across publisher web portals. RESULTS: PaperBot may operate stand-alone or it can be easily integrated with other software platforms and knowledge bases. Without user interactions, PaperBot retrieves and stores the bibliographic information (full reference, corresponding email contact, and full-text keyword hits) based on pre-set search logic from a wide range of sources including Elsevier, Wiley, Springer, PubMed/PubMedCentral, Nature, and Google Scholar. Although different publishing sites require different search configurations, the common interface of PaperBot unifies the process from the user perspective. Once saved, all information becomes web accessible allowing efficient triage of articles based on their actual relevance and seamless annotation of suitable metadata content. The platform allows the agile reconfiguration of all key details, such as the selection of search portals, keywords, and metadata dimensions. The tool also provides a one-click option for adding articles manually via digital object identifier or PubMed ID. The microservice architecture of PaperBot implements these capabilities as a loosely coupled collection of distinct modules devised to work separately, as a whole, or to be integrated with or replaced by additional software. All metadata is stored in a schema-less NoSQL database designed to scale efficiently in clusters by minimizing the impedance mismatch between relational model and in-memory data structures. CONCLUSIONS: As a testbed, we deployed PaperBot to help identify and manage peer-reviewed articles pertaining to digital reconstructions of neuronal morphology in support of the NeuroMorpho.Org data repository. PaperBot enabled the custom definition of both general and neuroscience-specific metadata dimensions, such as animal species, brain region, neuron type, and digital tracing system. Since deployment, PaperBot helped NeuroMorpho.Org more than quintuple the yearly volume of processed information while maintaining a stable personnel workforce. Patricia Maraver, Rubén Armañanzas, Todd A. Gillette, Giorgio A. Ascoli |
BMC Bioinform. | 2 |
| 2017 | Ensemble graphs to reveal post-transcriptional regulatory networks in Alzheimer's diseaseabstractIntegration of multiple datasets grants in-silico investigations with higher statistical and reasoning power to elucidate secondary discoveries hidden to the initial data producers. Here we introduce a novel method for the network analysis of messenger RNA regulation. Post-translational regulation of gene activity by microRNA molecules is investigated, combining expression data and sequence binding predictions. A set of sounding machine learning techniques allows the integration of these structural and functional results, conveying them into an ensemble graph of regulations. Ensemble graphs are embedded in Bayesian network classifiers following an ascending order of complexity and finally evaluated by their goodness-of-fit and classification performances. The new proposal is put to the test in the integration of four Alzheimer's disease datasets, reaching optimal values of 94.39% ± 2.34 accuracy and 0.9794 ± 0.01 for the area under the ROC curve. Detected regulations within the optimal network structure match the state-of-the-art literature. Additional dependences suggest previously unreported regulations in Alzheimer's disease research. Rubén Armañanzas |
BIBM | 1 |
| 2017 | Voxel-Based Diagnosis of Alzheimer's Disease Using Classifier EnsemblesabstractFunctional magnetic resonance imaging (fMRI) is one of the most promising noninvasive techniques for early Alzheimer's disease (AD) diagnosis. In this paper, we explore the application of different machine learning techniques to the classification of fMRI data for this purpose. The functional images were first preprocessed using the statistical parametric mapping toolbox to output individual maps of statistically activated voxels. A fast filter was applied afterwards to select voxels commonly activated across demented and nondemented groups. Four feature ranking selection techniques were embedded into a wrapper scheme using an inner-outer loop for the selection of relevant voxels. The wrapper approach was guided by the performance of six pattern recognition models, three of which were ensemble classifiers based on stochastic searches. Final classification performance was assessed from the nested internal and external cross-validation loops taking several voxel sets ordered by importance. Numerical performance was evaluated using statistical tests, and the best combination of voxel selection and classification reached a 97.14% average accuracy. Results repeatedly pointed out Brodmann regions with distinct activation patterns between demented and nondemented profiles, indicating that the machine learning analysis described is a powerful method to detect differences in several brain regions between both groups. Rubén Armañanzas, Martina Iglesias, Dinora A. Morales, Lidia Alonso-Nanclares |
IEEE J. Biomed. Health Informatics | 1 |
| 2016 | Genetic algorithms and Gaussian Bayesian networks to uncover the predictive core set of bibliometric indicesabstractThe diversity of bibliometric indices today poses the challenge of exploiting the relationships among them. Our research uncovers the best core set of relevant indices for predicting other bibliometric indices. An added difficulty is to select the role of each variable, that is, which bibliometric indices are predictive variables and which are response variables. This results in a novel multioutput regression problem where the role of each variable (predictor or response) is unknown beforehand. We use Gaussian Bayesian networks to solve the this problem and discover multivariate relationships among bibliometric indices. These networks are learnt by a genetic algorithm that looks for the optimal models that best predict bibliometric data. Results show that the optimal induced Gaussian Bayesian networks corroborate previous relationships between several indices, but also suggest new, previously unreported interactions. An extended analysis of the best model illustrates that a set of 12 bibliometric indices can be accurately predicted using only a smaller predictive core subset composed of citations, g‐index, q2‐index, and hr‐index. This research is performed using bibliometric data on Spanish full professors associated with the computer science area. Alfonso Ibáñez, Rubén Armañanzas, Concha Bielza, Pedro Larrañaga |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2013 | Unveiling relevant non-motor Parkinson's disease severity symptoms using a machine learning approach
Rubén Armañanzas, Concha Bielza, Kallol Ray Chaudhuri, Pablo Martínez-Martín, Pedro Larrañaga |
Artif. Intell. Medicine | 1 |
| 2013 | Comparison of metaheuristic strategies for peakbin selection in proteomic mass spectrometry data
Miguel García-Torres, Rubén Armañanzas, Concha Bielza, Pedro Larrañaga |
Inf. Sci. | 2 |
| 2011 | Peakbin Selection in Mass Spectrometry Data Using a Consensus Approach with Estimation of Distribution AlgorithmsabstractProgress is continuously being made in the quest for stable biomarkers linked to complex diseases. Mass spectrometers are one of the devices for tackling this problem. The data profiles they produce are noisy and unstable. In these profiles, biomarkers are detected as signal regions (peaks), where control and disease samples behave differently. Mass spectrometry (MS) data generally contain a limited number of samples described by a high number of features. In this work, we present a novel class of evolutionary algorithms, estimation of distribution algorithms (EDA), as an efficient peak selector in this MS domain. There is a trade-of f between the reliability of the detected biomarkers and the low number of samples for analysis. For this reason, we introduce a consensus approach, built upon the classical EDA scheme, that improves stability and robustness of the final set of relevant peaks. An entire data workflow is designed to yield unbiased results. Four publicly available MS data sets (two MALDI-TOF and another two SELDI-TOF) are analyzed. The results are compared to the original works, and a new plot (peak frequential plot) for graphically inspecting the relevant peaks is introduced. A complete online supplementary page, which can be found at http://www.sc.ehu.es/ccwbayes/members/ruben/ms, includes extended info and results, in addition to Matlab scripts and references. Rubén Armañanzas, Yvan Saeys, Iñaki Inza, Miguel García-Torres, Concha Bielza, Yves Van de Peer, Pedro Larrañaga |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2009 | Microarray Analysis of Autoimmune Diseases by Machine Learning ProceduresabstractMicroarray-based global gene expression profiling, with the use of sophisticated statistical algorithms is providing new insights into the pathogenesis of autoimmune diseases. We have applied a novel statistical technique for gene selection based on machine learning approaches to analyze microarray expression data gathered from patients with systemic lupus erythematosus (SLE) and primary antiphospholipid syndrome (PAPS), two autoimmune diseases of unknown genetic origin that share many common features. The methodology included a combination of three data discretization policies, a consensus gene selection method, and a multivariate correlation measurement. A set of 150 genes was found to discriminate SLE and PAPS patients from healthy individuals. Statistical validations demonstrate the relevance of this gene set from an univariate and multivariate perspective. Moreover, functional characterization of these genes identified an interferon-regulated gene signature, consistent with previous reports. It also revealed the existence of other regulatory pathways, including those regulated by PTEN, TNF, and BCL-2, which are altered in SLE and PAPS. Remarkably, a significant number of these genes carry E2F binding motifs in their promoters, projecting a role for E2F in the regulation of autoimmunity. Rubén Armañanzas, Borja Calvo, Iñaki Inza, Marcos López-Hoyos, Víctor Martínez-Taboada, Eduardo Ucar, Irantzu Bernales, Asier Fullaondo, Pedro Larrañaga, Ana M. Zubiaga |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2006 | Machine learning in bioinformaticsabstractThis article reviews machine learning methods for bioinformatics. It presents modelling methods, such as supervised classification, clustering and probabilistic graphical models for knowledge discovery, as well as deterministic and stochastic heuristics for optimization. Applications in genomics, proteomics, systems biology, evolution and text mining are also shown. Pedro Larrañaga, Borja Calvo, Roberto Santana 0001, Concha Bielza, Josu Galdiano, Iñaki Inza, José Antonio Lozano 0001, Rubén Armañanzas, Guzmán Santafé, Aritz Pérez Martínez, Víctor Robles |
Briefings Bioinform. | 8 |
| 2005 | A multiobjective approach to the portfolio optimization problemabstractThe portfolio optimization problem uses mathematical approaches to model stock exchange investments. Its aim is to find an optimal set of assets to invest on, as well as the optimal investments for each asset. In the present work, the problem is treated as a multi-objective optimization problem. Three well-known optimization techniques greedy search, simulated annealing and ant colony optimization are adapted to this multi-objective context. Pareto fronts for five stock indexes are collected, showing the different behaviors of the algorithms adapted. Finally, the results are discussed. Rubén Armañanzas, José Antonio Lozano 0001 |
Congress on Evolutionary Computation | 1 |