Angel Rubio

dblp:31/275 · also Ángel Rubio · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-authorSystems, architecture and hardware · 4 · 2 first-author
YearPublicationVenuePosition
2026 Foundation models and deep learning for cancer drug response prediction: a framework for data, metrics, and validation
abstract
The emergence of large-scale omics data and foundational models has renewed efforts in the field of cancer drug response prediction (DRP). Despite recent progress, challenges such as limited tumor heterogeneity in standard cell lines and inconsistencies in experimental protocols across studies persist. However, these challenges also open significant opportunities for innovation. The complex nature of drug responses, influenced by variations in new patients and new drugs, presents a critical area for advancing validation approaches that traditional machine learning approaches often overlook. This review provides a comprehensive overview of the current state of DRP using advanced machine-learning models, discussing data sources, model designs, and evaluation methods. We introduce a unified framework for testing these models with a focus on clinically relevant metrics. By evaluating a range of foundational and deep-learning models within this framework, we identify performance gaps and propose concrete strategies to advance these computational models for reliable use in personalized cancer treatment, thereby unlocking their full clinical potential.
Katyna Sada del Real, Vinay S. Swamy, Josefina Arcagni, Raul Rabadan, Angel Rubio
Briefings Bioinform.6
2023 Precision oncology: a review to assess interpretability in several explainable methods
abstract
Great efforts have been made to develop precision medicine-based treatments using machine learning. In this field, where the goal is to provide the optimal treatment for each patient based on his/her medical history and genomic characteristics, it is not sufficient to make excellent predictions. The challenge is to understand and trust the model's decisions while also being able to easily implement it. However, one of the issues with machine learning algorithms-particularly deep learning-is their lack of interpretability. This review compares six different machine learning methods to provide guidance for defining interpretability by focusing on accuracy, multi-omics capability, explainability and implementability. Our selection of algorithms includes tree-, regression- and kernel-based methods, which we selected for their ease of interpretation for the clinician. We also included two novel explainable methods in the comparison. No significant differences in accuracy were observed when comparing the methods, but an improvement was observed when using gene expression instead of mutational status as input for these methods. We concentrated on the current intriguing challenge: model comprehension and ease of use. Our comparison suggests that the tree-based methods are the most interpretable of those tested.
Marian Gimeno, Katyna Sada del Real, Angel Rubio
Briefings Bioinform.3
2022 Rediscover: an R package to identify mutually exclusive mutations
abstract
MOTIVATION: Discover is an algorithm developed to identify mutually exclusive genomic events. Its main contribution is a statistical analysis based on the Poisson-Binomial (PB) distribution to take into account the mutation rate of genes and samples. Discover is very effective for identifying mutually exclusive mutations at the expense of speed in large datasets: the PB is computationally costly to estimate, and checking all the potential mutually exclusive alterations requires millions of tests. RESULTS: We have implemented a new version of the package called Rediscover that implements exact and approximate computations of the PB. Rediscover exact implementation is slightly faster than Discover for large and medium-sized datasets. The approximation is 100-1000 times faster for them making it possible to get results in less than a minute with a standard desktop. The memory footprint is also smaller in Rediscover. The new package is available at CRAN and provides some functions to integrate its usage with other R packages such as maftools and TCGAbiolinks. AVAILABILITY AND IMPLEMENTATION: Rediscover is available at CRAN (https://cran.r-project.org/web/packages/Rediscover/index.html). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Juan A. Ferrer-Bonsoms, Laura Jareno, Angel Rubio
Bioinform.3
2022 On the identifiability of the isoform deconvolution problem: application to select the proper fragment length in an RNA-seq library
abstract
MOTIVATION: Isoform deconvolution is an NP-hard problem. The accuracy of the proposed solutions is far from perfect. At present, it is not known if gene structure and isoform concentration can be uniquely inferred given paired-end reads, and there is no objective method to select the fragment length to improve the number of identifiable genes. Different pieces of evidence suggest that the optimal fragment length is gene-dependent, stressing the need for a method that selects the fragment length according to a reasonable trade-off across all the genes in the whole genome. RESULTS: A gene is considered to be identifiable if it is possible to get both the structure and concentration of its transcripts univocally. Here, we present a method to state the identifiability of this deconvolution problem. Assuming a given transcriptome and that the coverage is sufficient to interrogate all junction reads of the transcripts, this method states whether or not a gene is identifiable given the read length and fragment length distribution. Applying this method using different read and fragment length combinations, the optimal average fragment length for the human transcriptome is around 400-600 nt for coding genes and 150-200 nt for long non-coding RNAs. The optimal read length is the largest one that fits in the fragment length. It is also discussed the potential profit of combining several libraries to reconstruct the transcriptome. Combining two libraries of very different fragment lengths results in a significant improvement in gene identifiability. AVAILABILITY AND IMPLEMENTATION: Code is available in GitHub (https://github.com/JFerrer-B/transcriptome-identifiability). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Juan A. Ferrer-Bonsoms, Xabier Morales, Pegah T. Afshar, Wing Hung Wong, Angel Rubio
Bioinform.5
2022 BOSO: A novel feature selection algorithm for linear regression with high-dimensional data
abstract
With the frenetic growth of high-dimensional datasets in different biomedical domains, there is an urgent need to develop predictive methods able to deal with this complexity. Feature selection is a relevant strategy in machine learning to address this challenge. We introduce a novel feature selection algorithm for linear regression called BOSO (Bilevel Optimization Selector Operator). We conducted a benchmark of BOSO with key algorithms in the literature, finding a superior accuracy for feature selection in high-dimensional datasets. Proof-of-concept of BOSO for predicting drug sensitivity in cancer is presented. A detailed analysis is carried out for methotrexate, a well-studied drug targeting cancer metabolism.
Luis Vitores Valcarcel, Edurne San José-Enériz, Xabier Cendoya, Angel Rubio, Xabier Agirre, Felipe Prósper, Francisco J. Planes
PLoS Comput. Biol.4
2019 Upstream analysis of alternative splicing: a review of computational approaches to predict context-dependent splicing factors
abstract
Alternative splicing (AS) has shown to play a pivotal role in the development of diseases, including cancer. Specifically, all the hallmarks of cancer (angiogenesis, cell immortality, avoiding immune system response, etc.) are found to have a counterpart in aberrant splicing of key genes. Identifying the context-specific regulators of splicing provides valuable information to find new biomarkers, as well as to define alternative therapeutic strategies. The computational models to identify these regulators are not trivial and require three conceptual steps: the detection of AS events, the identification of splicing factors that potentially regulate these events and the contextualization of these pieces of information for a specific experiment. In this work, we review the different algorithmic methodologies developed for each of these tasks. Main weaknesses and strengths of the different steps of the pipeline are discussed. Finally, a case study is detailed to help the reader be aware of the potential and limitations of this computational approach.
Fernando Carazo, Juan P. Romero, Angel Rubio
Briefings Bioinform.3
2015 Advances in network-based metabolic pathway analysis and gene expression data integration
abstract
With the emergence of metabolic networks, novel mathematical pathway concepts were introduced in the past decade, aiming to go beyond canonical maps. However, the use of network-based pathways to interpret 'omics' data has been limited owing to the fact that their computation has, until very recently, been infeasible in large (genome-scale) metabolic networks. In this review article, we describe the progress made in the past few years in the field of network-based metabolic pathway analysis. In particular, we review in detail novel optimization techniques to compute elementary flux modes, an important pathway concept in this field. In addition, we summarize approaches for the integration of metabolic pathways with gene expression data, discussing recent advances using network-based pathway concepts.
Alberto Rezola, Jon Pey, Luis Tobalina, Angel Rubio, John E. Beasley, Francisco J. Planes
Briefings Bioinform.4
2014 Recent Memory and Performance Improvements in Octopus Code
Joseba Alberdi-Rodriguez, Micael J. T. Oliveira, Pablo García-Risueño, Fernando Nogueira, Javier Muguerza, Agustin Arruabarrena, Angel Rubio
ICCSA (4)7
2014 Channel and feature selection for a surface electromyographic pattern recognition task
Iker Mesa, Angel Rubio, Imanol Tubia, Joaquín de No
Expert Syst. Appl.2
2013 Joint analysis of miRNA and mRNA expression data
abstract
miRNAs are small RNA molecules ('22 nt) that interact with their target mRNAs inhibiting translation or/and cleavaging the target mRNA. This interaction is guided by sequence complentarity and results in the reduction of mRNA and/or protein levels. miRNAs are involved in key biological processes and different diseases. Therefore, deciphering miRNA targets is crucial for diagnostics and therapeutics. However, miRNA regulatory mechanisms are complex and there is still no high-throughput and low-cost miRNA target screening technique. In recent years, several computational methods based on sequence complementarity of the miRNA and the mRNAs have been developed. However, the predicted interactions using these computational methods are inconsistent and the expected false positive rates are still large. Recently, it has been proposed to use the expression values of miRNAs and mRNAs (and/or proteins) to refine the results of sequence-based putative targets for a particular experiment. These methods have shown to be effective identifying the most prominent interactions from the databases of putative targets. Here, we review these methods that combine both expression and sequence-based putative targets to predict miRNA targets.
Ander Muniategui, Jon Pey, Francisco J. Planes, Angel Rubio
Briefings Bioinform.4
2013 Selection of human tissue-specific elementary flux modes using gene expression data
abstract
MOTIVATION: The analysis of high-throughput molecular data in the context of metabolic pathways is essential to uncover their underlying functional structure. Among different metabolic pathway concepts in systems biology, elementary flux modes (EFMs) hold a predominant place, as they naturally capture the complexity and plasticity of cellular metabolism and go beyond predefined metabolic maps. However, their use to interpret high-throughput data has been limited so far, mainly because their computation in genome-scale metabolic networks has been unfeasible. To face this issue, different optimization-based techniques have been recently introduced and their application to human metabolism is promising. RESULTS: In this article, we exploit and generalize the K-shortest EFM algorithm to determine a subset of EFMs in a human genome-scale metabolic network. This subset of EFMs involves a wide number of reported human metabolic pathways, as well as potential novel routes, and constitutes a valuable database where high-throughput data can be mapped and contextualized from a metabolic perspective. To illustrate this, we took expression data of 10 healthy human tissues from a previous study and predicted their characteristic EFMs based on enrichment analysis. We used a multivariate hypergeometric test and showed that it leads to more biologically meaningful results than standard hypergeometric. Finally, a biological discussion on the characteristic EFMs obtained in liver is conducted, finding a high level of agreement when compared with the literature.
Alberto Rezola, Jon Pey, Luis F. de Figueiredo, Adam Podhorski, Stefan Schuster, Angel Rubio, Francisco J. Planes
Bioinform.6
2012 CalMaTe: a method and software to improve allele-specific copy number of SNP arrays for downstream segmentation
abstract
SUMMARY: CalMaTe calibrates preprocessed allele-specific copy number estimates (ASCNs) from DNA microarrays by controlling for single-nucleotide polymorphism-specific allelic crosstalk. The resulting ASCNs are on average more accurate, which increases the power of segmentation methods for detecting changes between copy number states in tumor studies including copy neutral loss of heterozygosity. CalMaTe applies to any ASCNs regardless of preprocessing method and microarray technology, e.g. Affymetrix and Illumina. AVAILABILITY: The method is available on CRAN (http://cran.r-project.org/) in the open-source R package calmate, which also includes an add-on to the Aroma Project framework (http://www.aroma-project.org/).
Maria Ortiz-Estevez, Ander Aramburu, Henrik Bengtsson, Pierre Neuvial, Angel Rubio
Bioinform.5
2010 ACNE: a summarization method to estimate allele-specific copy numbers for Affymetrix SNP arrays
abstract
MOTIVATION: Current algorithms for estimating DNA copy numbers (CNs) borrow concepts from gene expression analysis methods. However, single nucleotide polymorphism (SNP) arrays have special characteristics that, if taken into account, can improve the overall performance. For example, cross hybridization between alleles occurs in SNP probe pairs. In addition, most of the current CN methods are focused on total CNs, while it has been shown that allele-specific CNs are of paramount importance for some studies. Therefore, we have developed a summarization method that estimates high-quality allele-specific CNs. RESULTS: The proposed method estimates the allele-specific DNA CNs for all Affymetrix SNP arrays dealing directly with the cross hybridization between probes within SNP probesets. This algorithm outperforms (or at least it performs as well as) other state-of-the-art algorithms for computing DNA CNs. It better discerns an aberration from a normal state and it also gives more precise allele-specific CNs. AVAILABILITY: The method is available in the open-source R package ACNE, which also includes an add on to the aroma.affymetrix framework (http://www.aroma-project.org/).
Maria Ortiz-Estevez, Henrik Bengtsson, Angel Rubio
Bioinform.3
2010 Improvements to previous algorithms to predict gene structure and isoform concentrations using Affymetrix Exon arrays
abstract
BACKGROUND: Exon arrays provide a way to measure the expression of different isoforms of genes in an organism. Most of the procedures to deal with these arrays are focused on gene expression or on exon expression. Although the only biological analytes that can be properly assigned a concentration are transcripts, there are very few algorithms that focus on them. The reason is that previously developed summarization methods do not work well if applied to transcripts. In addition, gene structure prediction, i.e., the correspondence between probes and novel isoforms, is a field which is still unexplored. RESULTS: We have modified and adapted a previous algorithm to take advantage of the special characteristics of the Affymetrix exon arrays. The structure and concentration of transcripts -some of them possibly unknown- in microarray experiments were predicted using this algorithm. Simulations showed that the suggested modifications improved both specificity (SP) and sensitivity (ST) of the predictions. The algorithm was also applied to different real datasets showing its effectiveness and the concordance with PCR validated results. CONCLUSIONS: The proposed algorithm shows a substantial improvement in the performance over the previous version. This improvement is mainly due to the exploitation of the redundancy of the Affymetrix exon arrays. An R-Package of SPACE with the updated algorithms have been developed and is freely available.
Miguel Ángel Antón, Ander Aramburu, Angel Rubio
BMC Bioinform.3
2009 Computing the shortest elementary flux modes in genome-scale metabolic networks
abstract
MOTIVATION: Elementary flux modes (EFMs) represent a key concept to analyze metabolic networks from a pathway-oriented perspective. In spite of considerable work in this field, the computation of the full set of elementary flux modes in large-scale metabolic networks still constitutes a challenging issue due to its underlying combinatorial complexity. RESULTS: In this article, we illustrate that the full set of EFMs can be enumerated in increasing order of number of reactions via integer linear programming. In this light, we present a novel procedure to efficiently determine the K-shortest EFMs in large-scale metabolic networks. Our method was applied to find the K-shortest EFMs that produce lysine in the genome-scale metabolic networks of Escherichia coli and Corynebacterium glutamicum. A detailed analysis of the biological significance of the K-shortest EFMs was conducted, finding that glucose catabolism, ammonium assimilation, lysine anabolism and cofactor balancing were correctly predicted. The work presented here represents an important step forward in the analysis and computation of EFMs for large-scale metabolic networks, where traditional methods fail for networks of even moderate size. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Luis F. de Figueiredo, Adam Podhorski, Angel Rubio, Christoph Kaleta, John E. Beasley, Stefan Schuster, Francisco J. Planes
Bioinform.3
2005 Correlation between Gene Expression and GO Semantic Similarity
abstract
This research analyzes some aspects of the relationship between gene expression, gene function, and gene annotation. Many recent studies are implicitly based on the assumption that gene products that are biologically and functionally related would maintain this similarity both in their expression profiles as well as in their Gene Ontology (GO) annotation. We analyze how accurate this assumption proves to be using real publicly available data. We also aim to validate a measure of semantic similarity for GO annotation. We use the Pearson correlation coefficient and its absolute value as a measure of similarity between expression profiles of gene products. We explore a number of semantic similarity measures (Resnik, Jiang, and Lin) and compute the similarity between gene products annotated using the GO. Finally, we compute correlation coefficients to compare gene expression similarity against GO semantic similarity. Our results suggest that the Resnik similarity measure outperforms the others and seems better suited for use in Gene Ontology. We also deduce that there seems to be correlation between semantic similarity in the GO annotation and gene expression for the three GO ontologies. We show that this correlation is negligible up to a certain semantic similarity value; then, for higher similarity values, the relationship trend becomes almost linear. These results can be used to augment the knowledge provided by clustering algorithms and in the development of bioinformatic tools for finding and characterizing gene products.
José L. Sevilla Campo, Victor Segura, Adam Podhorski, Elizabeth Guruceaga, José M. Mato, Luis A. Martínez-Cruz, Fernando J. Corrales, Angel Rubio
IEEE ACM Trans. Comput. Biol. Bioinform.8
2003 GARBAN: genomic analysis and rapid biological annotation of cDNA microarray and proteomic data
abstract
SUMMARY: Genomic Analysis and Rapid Biological ANnotation (GARBAN) is a new tool that provides an integrated framework to analyze simultaneously and compare multiple data sets derived from microarray or proteomic experiments. It carries out automated classifications of genes or proteins according to the criteria of the Gene Ontology Consortium at a level of depth defined by the user. Additionally, it performs clustering analysis of all sets based on functional categories or on differential expression levels. GARBAN also provides graphical representations of the biological pathways in which all the genes/proteins participate. AVAILABILITY: http://garban.tecnun.es.
Luis A. Martínez-Cruz, Angel Rubio, María L. Martínez-Chantar, Alberto Labarga, Isabel Barrio, Adam Podhorski, Victor Segura, José L. Sevilla Campo, Matías A. Avila, José M. Mato
Bioinform.2
2002 Involving the operator in a singularity avoidance strategy for a redundant slave manipulator in a teleoperated application
abstract
This paper proposes a control scheme to apply in a redundant slave robot in a teleoperated application. Singularities are avoided through a damped least-squares formulation of the inverse kinematics problem. A direct approach for selecting the damping factor is proposed. The redundant degree of freedom (DOF) is used to keep the manipulator in the most manipulable position and within the joint limits. In addition, all the redundancy resolution is performed whilst feeding back information to the operator who can help to improve the performance and robustness of the whole system.
Jaime Rubí, Angel Rubio, Alejo Avello
IROS2
2002 An intuitive force feed-back to avoid singularity proximity and workspace boundaries in bilateral controlled systems based on virtual springs
abstract
Kinematic and dynamic performance of robotic manipulators is poor when they are near singularities and workspace edges. This problem becomes more complex in teleoperated systems where there are two robots whose kinematics can be different. In these cases, neither should approach singularities or workspace boundaries. This work proposes a means of preventing the robots approaching places of risk. The algorithm computes a force proportional to the distance to singularity and/or workspace boundary. This force stops the robot moving towards that direction. It should be stressed that this strategy works in a very intuitive way in bilateral controlled systems, since the operator feels as if there are virtual springs which serve to prevent him entering forbidden regions. Finally, the algorithm has been tested successfully in a telerobotic system that consists of a Stewart platform and a 6-DOF open-chain manipulator.
Emilio Sanchez, Angel Rubio, Alejo Avello
IROS2
2000 On the Use of Virtual Springs to Avoid Singularities and Workspace Boundaries in Force-Feedback Teleoperation
abstract
A force-force controller of a master-slave force-feedback teleoperation system is proposed. This scheme is especially suitable for non-backdriveable manipulators. Forces and torques exerted by the operator and the environment and measured by a six axis force sensor are used to compute in real time a desired trajectory which is tracked by a PD controller. This system has been applied to a master-slave system, in which the master robot is a Stewart platform. The use of virtual springs prevents the platform from running into singularities and from getting close to workspace boundaries. Experimental results the satisfactory performance of control algorithm with virtual springs.
Angel Rubio, Alejo Avello, Julián Flórez 0001
ICRA1
1999 Adaptive Impedance Modification of a Master-Slave Manipulator
abstract
Master-slave teleoperators have to cope with tasks so different such as unconstrained motion and hard contact tasks. If the controller is tuned on one of them, the other will not perform well and may, eventually, lead to an unstable system. In this paper, an adaptive modification of the desired robot behaviour in the orthogonal direction to the contact task is proposed. This strategy makes the system stable in hard contact tasks while easy to control in unconstrained motion. Experimental results obtained with a 2-degree-of-freedom manipulator prove the effectiveness of the method.
Angel Rubio, Alejo Avello, Julián Flórez 0001
ICRA1