Alexandre Perera-Lluna

dblp:53/1787 · also Alexandre Perera · DBLP profile ↗
← Back
29ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0001-6427-851XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 8 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 singIST: An integrative method for comparative single-cell transcriptomics between disease models and humans
abstract
MOTIVATION: Disease models are fundamental tools in drug discovery and early-stage drug development, but they only approximate human disease, and selecting a suitable model is challenging. Quantitative computational methods exist to assess molecular resemblance to human conditions, but approaching that work at single-cell resolution, and doing so in an explainable and generalizable way, remain very limited. RESULTS: We present singIST, a computational method for comparative single-cell transcriptomics analysis between disease models and human conditions. singIST provides explainable quantitative measures on disease model similarity to the human reference at the pathway, cell type and gene levels. These measures jointly account for gene orthology, cell type presence in the model, cell type and gene importance in the human condition, and gene level fold changes in the model, within a unifying framework that controls for the intrinsic complexities of single-cell data. We first test singIST in three well-characterized murine models against moderate-to-severe Atopic Dermatitis, showing that it recapitulates established biology while generating new hypotheses. We then apply it to Hidradenitis Suppurativa, comparing in vivo human lesions with ex vivo skin explants with and without CD3/CD28 stimulation, and show that stimulation selectively improves pathways that already recapitulate the human signal. Finally, we perform simulation studies that: (i) unit-test the implementation and behaviour of the algorithm under controlled scenarios and (ii) compare singIST against a naïve baseline based on overlapping differentially expressed genes.
Aitor Moruno-Cuenca, Sergio Picart-Armada, Rachael Bogle, Jennifer Fox, Lam C. Tsoi, Johann Eli Gudjonsson, Alexandre Perera-Lluna, Francesc Fernández-Albert
PLoS Comput. Biol.7
2025 Functional Enrichment of Lipidomics Data to Characterize the Metabolism of Diabetes Mellitus and Subclinical Carotid Atherosclerosis
abstract
Pathway enrichment analysis aims to uncover global functional patterns in omics data, including metabolomics. However, its application to lipidomics, the comprehensive study of lipids in biological systems, is hindered by the limited availability of lipid-specific pathways in current databases. Tools such as LINEX2and LION/web offer solutions for functional enrichment of lipids. Nevertheless, more information can still be extracted from large databases like LIPID MAPS. In this work, we present a novel functional enrichment algorithm based on a hierarchical graph constructed through unsupervised clustering of molecular embeddings. These embeddings were generated by applying a SMILES (Simplified Molecular Input Line Entry System) language model to all lipids in LIPID MAPS. We apply this algorithm to five comparisons involving type 1 and type 2 diabetes, as well as subclinical carotid atherosclerosis in type 2 diabetes. Additionally, we perform enrichment using LINEX2on the same datasets and compare the outcomes to highlight complementary insights.
Maria Barranco-Altirriba, Josch K. Pauling, Dídac Mauricio, Alexandre Perera-Lluna
BIBE4
2025 Cross Cohort Integration of Whole Blood RNA-Seq Samples: Normalization Units Comparison and Stable Housekeeping Gene Selection
abstract
The comparability of whole-blood RNA-seq data across heterogeneous human cohorts remains a critical challenge for transcriptomic research. In this study, we analyzed baseline fasting whole blood RNA-seq profiles from three distinct cohorts (non-elite athletes, individuals with familial venous thrombosis, and participants in an obesity study) to quantify non-biological variance and evaluate normalization strategies for improving cross-dataset integration. Two common approaches, transcripts per million (TPM) and counts per million (CPM), were compared in combination with batch effect correction. Principal Component Analysis (PCA) showed that TPM normalized data failed to achieve meaningful alignment across cohorts even after batch correction. In contrast, CPM normalization combined with batch effect correction more effectively reduced unwanted technical variation and improved cross-cohort alignment. Further, we assessed housekeeping gene stability and identified SDHA, GUSB, and TBP as the most suitable housekeeping genes for whole blood RNA-seq, providing reliable internal references for normalization and dataset comparability. Collectively, our findings underscore the importance of selecting appropriate normalization strategies and leveraging stable housekeeping genes as critical for improving cross-cohort comparability in whole blood transcriptomic studies.
Pol Ezquerra-Condeminas, Alexandre Perera-Lluna, José Manuel Soria
BIBE2
2025 Characterizing Temperature-Related Risks in Patients with Chronic Conditions in the Vallès Occidental, Spain
abstract
This study assesses the relationship between daily maximum temperature and healthcare interactions (admissions and emergency visits) among patients with cardiovascular, respiratory, and diabetic conditions in Vallès Occidental, Spain (2015–2022). Using distributed lag non-linear models (DLNM), we quantify temperature-related risk while accounting for lagged effects and seasonality. Results show significant associations especially with higher temperatures. These findings underscore the need for improved, disease-specific heat-health warning systems and more granular consideration of comorbidities in climatehealth vulnerability assessments.
Blanca Alaejos Pardo, Alexandre Perera-Lluna, Jordi Fonollosa, David Dalmau, Marc Prohom
BIBE2
2025 A deep attention-based encoder for the prediction of type 2 diabetes longitudinal outcomes from routinely collected health care data
abstract
Recent evidence indicates that Type 2 Diabetes Mellitus (T2DM) is a complex and highly heterogeneous disease involving various pathophysiological and genetic pathways, which presents clinicians with challenges in disease management. While deep learning models have made significant progress in helping practitioners manage T2DM treatments, several important limitations persist. In this paper we propose DARE, a model based on the transformer encoder, designed for analyzing longitudinal heterogeneous diabetes data. The model can be easily fine-tuned for various clinical prediction tasks, enabling a computational approach to assist clinicians in the management of the disease. We trained DARE using data from over 200,000 diabetic subjects from the primary healthcare SIDIAP database, which includes diagnosis and drug codes, along with various clinical and analytical measurements. After an unsupervised pre-training phase, we fine-tuned the model for predicting three specific clinical outcomes: i) occurrence of comorbidity, ii) achievement of target glycemic control (defined as glycated hemoglobin < 7 % ) and iii) changes in glucose-lowering treatment. In cross-validation, the embedding vectors generated by DARE outperformed those from baseline models (comorbidities prediction task A U C = 0 . 88 , treatment prediction task A U C = 0 . 91 , HbA1c target prediction task A U C = 0 . 82 ). Our findings suggest that attention-based encoders improve results with respect to different deep learning and classical baseline models when used to predict different clinical relevant outcomes from T2DM longitudinal data. • Deep learning shows promise in managing T2DM and predicting its progression • We developed DARE, a transformer model for analyzing T2DM longitudinal records • DARE was trained on data from 200K diabetic patients spanning a 5-year period • DARE forecasts treatment changes, HbA1c targets, and diabetes comorbidities.
Enrico Manzini, Bogdan Vlacho, Josep Franch-Nadal, Joan Escudero, Ana Génova, Elisenda Reixach, Erich Andrés, Israel Pizarro, Dídac Mauricio, Alexandre Perera-Lluna
Expert Syst. Appl.10
2023 IoT Gas Sensors Array for Unobtrusive Tracking of Cooking Activity
abstract
Recent research on remote tracking environments has strengthened smart home IoT ecosystems by the integration of multiple sensing tools that capture not only contextual data in a private setting, but also information about its residents. This shift paves the way for remote health industries, as information traditionally out of reach is available 24/7. Gas sensing, moving away from privacy-invasive tracking paradigms, emerges within this context, inspiring the monitoring of activities of daily living (ADLs) that could facilitate the remote healthcare supervision of the elderly. In this paper, we present how a gas sensing array based on low-cost commercial metal oxide (MOX) gas sensors has been assembled for the development of Principal Component Analysis (PCA) model which detects cooking activity within a household. Our resulting unobtrusive tracking system requiring no user input and posing no privacy concerns, suitable to other ADL use cases, highlights how AI-equipped IoT cloud infrastructures and accurate gas sensors are called to revolutionise remote healthcare.
Zouhair Haddi, Joshua Llano, Miquel Alfaras, Daniel Marin, Alexandre Perera-Lluna, Xavier Llauradó, Narcís Avellena, Jordi Fonollosa, Eduard Llobet Valero
CoDIT5
2022 Mapping layperson medical terminology into the Human Phenotype Ontology using neural machine translation models
abstract
In the medical domain there exists a terminological gap between patients and caregivers and the healthcare professionals. This gap may hinder the success of the communication between healthcare consumers and professionals in the field, with negative emotional and clinical consequences. In this work, we build a machine learning-based tool for the automatic translation between the terminology used by laypeople and that of the Human Phenotype Ontology (HPO). HPO is a structured vocabulary of phenotypic abnormalities found in human disease. Our method uses a vector space to represent an HPO-specific embedding as the output space for a neural network model trained on vector representations of layperson versions and other textual descriptors of medical terms. We explored different output embeddings coupled to different neural network architectures for the machine translation stage. We compute a similarity measure to evaluate the ability of the model to assign an HPO term to a layperson input. The best-performing models resulted with a similarity higher than 0.7 for more than 80% of the terms, with a median between 0.98 and 1. The translator model is made available in a web application at this link: https://hpotranslator.b2slab.upc.edu.
Enrico Manzini, Jon Garrido-Aguirre, Jordi Fonollosa, Alexandre Perera-Lluna
Expert Syst. Appl.4
2022 Longitudinal deep learning clustering of Type 2 Diabetes Mellitus trajectories using routinely collected health records
abstract
Type 2 diabetes mellitus (T2DM) is a highly heterogeneous chronic disease with different pathophysiological and genetic characteristics affecting its progression, associated complications and response to therapies. The advances in deep learning (DL) techniques and the availability of a large amount of healthcare data allow us to investigate T2DM characteristics and evolution with a completely new approach, studying common disease trajectories rather than cross sectional values. We used an Kernelized-AutoEncoder algorithm to map 5 years of data of 11,028 subjects diagnosed with T2DM in a latent space that embedded similarities and differences between patients in terms of the evolution of the disease. Once we obtained the latent space, we used classical clustering algorithms to create longitudinal clusters representing different evolutions of the diabetic disease. Our unsupervised DL clustering algorithm suggested seven different longitudinal clusters. Different mean ages were observed among the clusters (ranging from 65.3±11.6 to 72.8±9.4). Subjects in clusters B (Hypercholesteraemic) and E (Hypertensive) had shorter diabetes duration (9.2±3.9 and 9.5±3.9 years respectively). Subjects in Cluster G (Metabolic) had the poorest glycaemic control (mean glycated hemoglobin 7.99±1.42%), while cluster E had the best one (mean glycated hemoglobin 7.04±1.11%). Obesity was observed mainly in clusters A (Neuropathic), C (Multiple Complications), F (Retinopathy) and G. A dashboard is available at dm2.b2slab.upc.edu to visualize the different trajectories corresponding to the 7 clusters.
Enrico Manzini, Bogdan Vlacho, Josep Franch-Nadal, Joan Escudero, Ana Génova, Elisenda Reixach, Erik Andrés, Israel Pizarro, José-Luis Portero, Dídac Mauricio, Alexandre Perera-Lluna
J. Biomed. Informatics11
2021 MultiPaths: a Python framework for analyzing multi-layer biological networks using diffusion algorithms
abstract
SUMMARY: High-throughput screening yields vast amounts of biological data which can be highly challenging to interpret. In response, knowledge-driven approaches emerged as possible solutions to analyze large datasets by leveraging prior knowledge of biomolecular interactions represented in the form of biological networks. Nonetheless, given their size and complexity, their manual investigation quickly becomes impractical. Thus, computational approaches, such as diffusion algorithms, are often employed to interpret and contextualize the results of high-throughput experiments. Here, we present MultiPaths, a framework consisting of two independent Python packages for network analysis. While the first package, DiffuPy, comprises numerous commonly used diffusion algorithms applicable to any generic network, the second, DiffuPath, enables the application of these algorithms on multi-layer biological networks. To facilitate its usability, the framework includes a command line interface, reproducible examples and documentation. To demonstrate the framework, we conducted several diffusion experiments on three independent multi-omics datasets over disparate networks generated from pathway databases, thus, highlighting the ability of multi-layer networks to integrate multiple modalities. Finally, the results of these experiments demonstrate how the generation of harmonized networks from disparate databases can improve predictive performance with respect to individual resources. AVAILABILITY AND IMPLEMENTATION: DiffuPy and DiffuPath are publicly available under the Apache License 2.0 at https://github.com/multipaths. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Josep Marín-Llaó, Sarah Mubeen, Alexandre Perera-Lluna, Martin Hofmann-Apitius, Sergio Picart-Armada, Daniel Domingo-Fernández
Bioinform.3
2021 The effect of statistical normalization on network propagation scores
abstract
MOTIVATION: Network diffusion and label propagation are fundamental tools in computational biology, with applications like gene-disease association, protein function prediction and module discovery. More recently, several publications have introduced a permutation analysis after the propagation process, due to concerns that network topology can bias diffusion scores. This opens the question of the statistical properties and the presence of bias of such diffusion processes in each of its applications. In this work, we characterized some common null models behind the permutation analysis and the statistical properties of the diffusion scores. We benchmarked seven diffusion scores on three case studies: synthetic signals on a yeast interactome, simulated differential gene expression on a protein-protein interaction network and prospective gene set prediction on another interaction network. For clarity, all the datasets were based on binary labels, but we also present theoretical results for quantitative labels. RESULTS: Diffusion scores starting from binary labels were affected by the label codification and exhibited a problem-dependent topological bias that could be removed by the statistical normalization. Parametric and non-parametric normalization addressed both points by being codification-independent and by equalizing the bias. We identified and quantified two sources of bias-mean value and variance-that yielded performance differences when normalizing the scores. We provided closed formulae for both and showed how the null covariance is related to the spectral properties of the graph. Despite none of the proposed scores systematically outperformed the others, normalization was preferred when the sought positive labels were not aligned with the bias. We conclude that the decision on bias removal should be problem and data-driven, i.e. based on a quantitative analysis of the bias and its relation to the positive entities. AVAILABILITY: The code is publicly available at https://github.com/b2slab/diffuBench and the data underlying this article are available at https://github.com/b2slab/retroData. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sergio Picart-Armada, Wesley K. Thompson, Alfonso Buil, Alexandre Perera-Lluna
Bioinform.4
2019 Multiview: a software package for multiview pattern recognition methods
abstract
SUMMARY: Multiview datasets are the norm in bioinformatics, often under the label multi-omics. Multiview data are gathered from several experiments, measurements or feature sets available for the same subjects. Recent studies in pattern recognition have shown the advantage of using multiview methods of clustering and dimensionality reduction; however, none of these methods are readily available to the extent of our knowledge. Multiview extensions of four well-known pattern recognition methods are proposed here. Three multiview dimensionality reduction methods: multiview t-distributed stochastic neighbour embedding, multiview multidimensional scaling and multiview minimum curvilinearity embedding, as well as a multiview spectral clustering method. Often they produce better results than their single-view counterparts, tested here on four multiview datasets. AVAILABILITY AND IMPLEMENTATION: R package at the B2SLab site: http://b2slab.upc.edu/software-and-tutorials/ and Python package: https://pypi.python.org/pypi/multiview. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Samir Kanaan-Izquierdo, Andrey Ziyatdinov, Maria Araceli Burgueño, Alexandre Perera-Lluna
Bioinform.4
2019 Benchmarking network propagation methods for disease gene identification
abstract
In-silico identification of potential target genes for disease is an essential aspect of drug target discovery. Recent studies suggest that successful targets can be found through by leveraging genetic, genomic and protein interaction information. Here, we systematically tested the ability of 12 varied algorithms, based on network propagation, to identify genes that have been targeted by any drug, on gene-disease data from 22 common non-cancerous diseases in OpenTargets. We considered two biological networks, six performance metrics and compared two types of input gene-disease association scores. The impact of the design factors in performance was quantified through additive explanatory models. Standard cross-validation led to over-optimistic performance estimates due to the presence of protein complexes. In order to obtain realistic estimates, we introduced two novel protein complex-aware cross-validation schemes. When seeding biological networks with known drug targets, machine learning and diffusion-based methods found around 2-4 true targets within the top 20 suggestions. Seeding the networks with genes associated to disease by genetics decreased performance below 1 true hit on average. The use of a larger network, although noisier, improved overall performance. We conclude that diffusion-based prioritisers and machine learning applied to diffusion-based features are suited for drug discovery in practice and improve over simpler neighbour-voting methods. We also demonstrate the large impact of choosing an adequate validation strategy and the definition of seed disease genes.
Sergio Picart-Armada, Steven J. Barrett, David R. Willé, Alexandre Perera-Lluna, Alex Gutteridge, Benoit H. Dessailly
PLoS Comput. Biol.4
2018 diffuStats: an R package to compute diffusion-based scores on biological networks
abstract
Summary: Label propagation and diffusion over biological networks are a common mathematical formalism in computational biology for giving context to molecular entities and prioritizing novel candidates in the area of study. There are several choices in conceiving the diffusion process-involving the graph kernel, the score definitions and the presence of a posterior statistical normalization-which have an impact on the results. This manuscript describes diffuStats, an R package that provides a collection of graph kernels and diffusion scores, as well as a parallel permutation analysis for the normalized scores, that eases the computation of the scores and their benchmarking for an optimal choice. Availability and implementation: The R package diffuStats is publicly available in Bioconductor, https://bioconductor.org, under the GPL-3 license. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Sergio Picart-Armada, Wesley K. Thompson, Alfonso Buil, Alexandre Perera-Lluna
Bioinform.4
2018 FELLA: an R package to enrich metabolomics data
abstract
BACKGROUND: Pathway enrichment techniques are useful for understanding experimental metabolomics data. Their purpose is to give context to the affected metabolites in terms of the prior knowledge contained in metabolic pathways. However, the interpretation of a prioritized pathway list is still challenging, as pathways show overlap and cross talk effects. RESULTS: We introduce FELLA, an R package to perform a network-based enrichment of a list of affected metabolites. FELLA builds a hierarchical representation of an organism biochemistry from the Kyoto Encyclopedia of Genes and Genomes (KEGG), containing pathways, modules, enzymes, reactions and metabolites. In addition to providing a list of pathways, FELLA reports intermediate entities (modules, enzymes, reactions) that link the input metabolites to them. This sheds light on pathway cross talk and potential enzymes or metabolites as targets for the condition under study. FELLA has been applied to six public datasets -three from Homo sapiens, two from Danio rerio and one from Mus musculus- and has reproduced findings from the original studies and from independent literature. CONCLUSIONS: The R package FELLA offers an innovative enrichment concept starting from a list of metabolites, based on a knowledge graph representation of the KEGG database that focuses on interpretability. Besides reporting a list of pathways, FELLA suggests intermediate entities that are of interest per se. Its usefulness has been shown at several molecular levels on six public datasets, including human and animal models. The user can run the enrichment analysis through a simple interactive graphical interface or programmatically. FELLA is publicly available in Bioconductor under the GPL-3 license.
Sergio Picart-Armada, Francesc Fernández-Albert, Maria Vinaixa, Oscar Yanes, Alexandre Perera-Lluna
BMC Bioinform.5
2018 Multiview and multifeature spectral clustering using common eigenvectors
Samir Kanaan-Izquierdo, Andrey Ziyatdinov, Alexandre Perera-Lluna
Pattern Recognit. Lett.3
2016 solarius: an R interface to SOLAR for variance component analysis in pedigrees
abstract
UNLABELLED: : The open source environment R is one of the most widely used software for statistical computing. It provides a variety of applications including statistical genetics. Most of the powerful tools for quantitative genetic analyses are stand-alone free programs developed by researchers in academia. SOLAR is one of the standard software programs to perform linkage and association mappings of the quantitative trait loci (QTLs) in pedigrees of arbitrary size and complexity. solarius allows the user to exploit the variance component methods implemented in SOLAR. It automates such routine operations as formatting pedigree and phenotype data. It parses also the model output and contains summary and plotting functions for exploration of the results. In addition, solarius enables parallel computing of the linkage and association analyses that makes the calculation of genome-wide scans more efficient. AVAILABILITY AND IMPLEMENTATION: solarius is available on CRAN and on GitHub https://github.com/ugcd/solarius CONTACT: : [email protected].
Andrey Ziyatdinov, Helena Brunel, Angel Martinez-Perez, Alfonso Buil, Alexandre Perera-Lluna, José Manuel Soria
Bioinform.5
2015 Sequence information gain based motif analysis
abstract
BACKGROUND: The detection of regulatory regions in candidate sequences is essential for the understanding of the regulation of a particular gene and the mechanisms involved. This paper proposes a novel methodology based on information theoretic metrics for finding regulatory sequences in promoter regions. RESULTS: This methodology (SIGMA) has been tested on genomic sequence data for Homo sapiens and Mus musculus. SIGMA has been compared with different publicly available alternatives for motif detection, such as MEME/MAST, Biostrings (Bioconductor package), MotifRegressor, and previous work such Qresiduals projections or information theoretic based detectors. Comparative results, in the form of Receiver Operating Characteristic curves, show how, in 70% of the studied Transcription Factor Binding Sites, the SIGMA detector has a better performance and behaves more robustly than the methods compared, while having a similar computational time. The performance of SIGMA can be explained by its parametric simplicity in the modelling of the non-linear co-variability in the binding motif positions. CONCLUSIONS: Sequence Information Gain based Motif Analysis is a generalisation of a non-linear model of the cis-regulatory sequences detection based on Information Theory. This generalisation allows us to detect transcription factor binding sites with maximum performance disregarding the covariability observed in the positions of the training set of sequences. SIGMA is freely available to the public at http://b2slab.upc.edu.
Joan Maynou, Erola Pairo, Santiago Marco, Alexandre Perera-Lluna
BMC Bioinform.4
2014 An R package to analyse LC/MS metabolomic data: MAIT (Metabolite Automatic Identification Toolkit)
abstract
UNLABELLED: Current tools for liquid chromatography and mass spectrometry for metabolomic data cover a limited number of processing steps, whereas online tools are hard to use in a programmable fashion. This article introduces the Metabolite Automatic Identification Toolkit (MAIT) package, which makes it possible for users to perform metabolomic end-to-end liquid chromatography and mass spectrometry data analysis. MAIT is focused on improving the peak annotation stage and provides essential tools to validate statistical analysis results. MAIT generates output files with the statistical results, peak annotation and metabolite identification. AVAILABILITY AND IMPLEMENTATION: http://b2slab.upc.edu/software-and-downloads/metabolite-automatic-identification-toolkit/.
Francesc Fernández-Albert, Rafael Llorach, Cristina Andres-Lacueva, Alexandre Perera-Lluna
Bioinform.4
2014 Intensity drift removal in LC/MS metabolomics by common variance compensation
abstract
UNLABELLED: Liquid chromatography coupled to mass spectrometry (LC/MS) has become widely used in Metabolomics. Several artefacts have been identified during the acquisition step in large LC/MS metabolomics experiments, including ion suppression, carryover or changes in the sensitivity and intensity. Several sources have been pointed out as responsible for these effects. In this context, the drift effects of the peak intensity is one of the most frequent and may even constitute the main source of variance in the data, resulting in misleading statistical results when the samples are analysed. In this article, we propose the introduction of a methodology based on a common variance analysis before the data normalization to address this issue. This methodology was tested and compared with four other methods by calculating the Dunn and Silhouette indices of the quality control classes. The results showed that our proposed methodology performed better than any of the other four methods. As far as we know, this is the first time that this kind of approach has been applied in the metabolomics context. AVAILABILITY AND IMPLEMENTATION: The source code of the methods is available as the R package intCor at http://b2slab.upc.edu/software-and-downloads/intensity-drift-correction/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Francesc Fernández-Albert, Rafael Llorach, Mar Garcia-Aloy, Andrey Ziyatdinov, Cristina Andres-Lacueva, Alexandre Perera-Lluna
Bioinform.6
2013 Predictability of gene ontology slim-terms from primary structure information in Embryophyta plant proteins
abstract
BACKGROUND: Proteins are the key elements on the path from genetic information to the development of life. The roles played by the different proteins are difficult to uncover experimentally as this process involves complex procedures such as genetic modifications, injection of fluorescent proteins, gene knock-out methods and others. The knowledge learned from each protein is usually annotated in databases through different methods such as the proposed by The Gene Ontology (GO) consortium. Different methods have been proposed in order to predict GO terms from primary structure information, but very few are available for large-scale functional annotation of plants, and reported success rates are much less than the reported by other non-plant predictors. This paper explores the predictability of GO annotations on proteins belonging to the Embryophyta group from a set of features extracted solely from their primary amino acid sequence. RESULTS: High predictability of several GO terms was found for Molecular Function and Cellular Component. As expected, a lower degree of predictability was found on Biological Process ontology annotations, although a few biological processes were easily predicted. Proteins related to transport and transcription were particularly well predicted from primary structure information. The most discriminant features for prediction were those related to electric charges of the amino-acid sequence and hydropathicity derived features. CONCLUSIONS: An analysis of GO-slim terms predictability in plants was carried out, in order to determine single categories or groups of functions that are most related with primary structure information. For each highly predictable GO term, the responsible features of such successfulness were identified and discussed. In addition to most published studies, focused on few categories or single ontologies, results in this paper comprise a complete landscape of GO predictability from primary structure encompassing 75 GO terms at molecular, cellular and phenotypical level. Thus, it provides a valuable guide for researchers interested on further advances in protein function prediction on Embryophyta plants.
Jorge Alberto Jaramillo-Garzón, Joan-Josep Gallardo-Chacón, Germán Castellanos-Domínguez, Alexandre Perera-Lluna
BMC Bioinform.4
2012 A subspace method for the detection of transcription factor binding sites
abstract
MOTIVATION: The identification of the sites at which transcription factors (TFs) bind to Deoxyribonucleic acid (DNA) is an important problem in molecular biology. Many computational methods have been developed for motif finding, most of them based on position-specific scoring matrices (PSSMs) which assume the independence of positions within a binding site. However, some experimental and computational studies demonstrate that interdependences within the positions exist. RESULTS: In this article, we introduce a novel motif finding method which constructs a subspace based on the covariance of numerical DNA sequences. When a candidate sequence is projected into the modeled subspace, a threshold in the Q-residuals confidence allows us to predict whether this sequence is a binding site. Using the TRANSFAC and JASPAR databases, we compared our Q-residuals detector with existing PSSM methods. In most of the studied TF binding sites, the Q-residuals detector performs significantly better and faster than MATCH and MAST. As compared with Motifscan, a method which takes into account interdependences, the performance of the Q-residuals detector is better when the number of available sequences is small.
Erola Pairo, Joan Maynou, Santiago Marco, Alexandre Perera-Lluna
Bioinform.4
2010 Fault detection, identification, and reconstruction of faulty chemical gas sensors under drift conditions, using Principal Component Analysis and Multiscale-PCA
abstract
Statistical methods like Principal Components Analysis (PCA) or Partial Least Squares (PLS) and multiscale approaches, have been reported to be very useful in the task of fault diagnosis of malfunctioning sensors for several types of faults. In this work, we compare the performance of PCA and Multiscale-PCA on a fault based on a change of sensor sensitivity. This type of fault affects chemical gas sensors and it is one of the effects of the sensor poisoning. These two methods will be applied on a dataset composed by the signals of 17 conductive polymer gas sensors, measuring three analytes at several concentration levels during 10 months. Therefore, additionally to performance's comparison, both method's stability along the time will be tested. The comparison between both techniques will be made regarding three aspects; detection, identification of the faulty sensors and correction of faulty sensors response.
Marta Padilla, Alexandre Perera-Lluna, Ivan Montoliu, A. Chaudry, Krishna C. Persaud, Santiago Marco
IJCNN2
2010 MISS: a non-linear methodology based on mutual information for genetic association studies in both population and sib-pairs analysis
abstract
MOTIVATION: Finding association between genetic variants and phenotypes related to disease has become an important vehicle for the study of complex disorders. In this context, multi-loci genetic association might unravel additional information when compared with single loci search. The main goal of this work is to propose a non-linear methodology based on information theory for finding combinatorial association between multi-SNPs and a given phenotype. RESULTS: The proposed methodology, called MISS (mutual information statistical significance), has been integrated jointly with a feature selection algorithm and has been tested on a synthetic dataset with a controlled phenotype and in the particular case of the F7 gene. The MISS methodology has been contrasted with a multiple linear regression (MLR) method used for genetic association in both, a population-based study and a sib-pairs analysis and with the maximum entropy conditional probability modelling (MECPM) method, which searches for predictive multi-locus interactions. Several sets of SNPs within the F7 gene region have been found to show a significant correlation with the FVII levels in blood. The proposed multi-site approach unveils combinations of SNPs that explain more significant information of the phenotype than their individual polymorphisms. MISS is able to find more correlations between SNPs and the phenotype than MLR and MECPM. Most of the marked SNPs appear in the literature as functional variants with real effect on the protein FVII levels in blood. AVAILABILITY: The code is available at http://sisbio.recerca.upc.edu/R/MISS_0.2.tar.gz
Helena Brunel, Joan-Josep Gallardo-Chacón, Alfonso Buil, Montserrat Vallverdú, José Manuel Soria, Pere Caminal, Alexandre Perera-Lluna
Bioinform.7
2010 Computational detection of transcription factor binding sites through differential Rényi entropy
abstract
Regulatory sequence detection is a critical facet for understanding the cell mechanisms in order to coordinate the response to stimuli. Protein synthesis involves the binding of a transcription factor to specific sequences in a process related to the gene expression initiation. A characteristic of this binding process is that the same factor binds with different sequences placed along all genome. Thus, any computational approach shows many difficulties related with this variability observed from the binding sequences. This paper proposes the detection of transcription factor binding sites based on a parametric uncertainty measurement (Rényi entropy). This detection algorithm evaluates the variation on the total Rényi entropy of a set of sequences when a candidate sequence is assumed to be a true binding site belonging to the set. The efficiency of the method is measured in form of receiver operating characteristic (ROC) curves on different transcription factors fromSaccharomyces cerevisiaeorganism. The results are compared with other known motif detection algorithms such as Motif Discovery scan (MDscan) and multiple expectation–maximization (EM) for motif elicitation (MEME).
Joan Maynou, Joan-Josep Gallardo-Chacón, Montserrat Vallverdú, Pere Caminal, Alexandre Perera-Lluna
IEEE Trans. Inf. Theory5
2009 Dimensionality Reduction Oriented Toward the Feature Visualization for Ischemia Detection
abstract
An effective data representation methodology on high-dimension feature spaces is presented, which allows a better interpretation of subjacent physiological phenomena (namely, cardiac behavior related to cardiovascular diseases), and is based on search criteria over a feature set resulting in an increase in the detection capability of ischemic pathologies, but also connecting these features with the physiologic representation of the ECG. The proposed dimension reduction scheme consists of three levels: projection, interpretation, and visualization. First, a hybrid algorithm is described that projects the multidimensional data to a lower dimension space, gathering the features that contribute similarly in the meaning of the covariance reconstruction in order to find information of clinical relevance over the initial training space. Next, an algorithm of variable selection is provided that further reduces the dimension, taking into account only the variables that offer greater class separability, and finally, the selected feature set is projected to a 2-D space in order to verify the performance of the suggested dimension reduction algorithm in terms of the discrimination capability for ischemia detection. The ECG recordings used in this study are from the European ST-T database and from the Universidad Nacional de Colombia database. In both cases, over 99% feature reduction was obtained, and classification precision was over 99% using a five-nearest-neighbor classifier (5-NN).
Edilson Delgado-Trejos, Alexandre Perera-Lluna, Montserrat Vallverdú, Pere Caminal, Germán Castellanos-Domínguez
IEEE Trans. Inf. Technol. Biomed.2
2008 Floating Feature Selection for multiloci association of quantitative traits in sib-pairs analysis
abstract
Finding association between genotypic differences and disease traits has become one of the main objectives in current genetic research. It has been published that some of the underlying factors in the dynamics of the coagulation process have a genetic compound, showing significant hereditability. This is the case of the Factor VII. In this work, we propose a method for selecting sets of Single Nucleotide Polymorphisms (SNPs) of the F7 gene that are significantly related with the phenotype (Factor VII levels). The methodology is applied to the sib pairs from the GAIT project sample. The method consists of an adapted Sequential Floating Feature Selection (SFFS) algorithm. This algorithm is applied with two relevance criteria, one linear and one non linear. The SNPs sets found with linear models are included in the sets found with non linear techniques. The results fit in with previous results in clinical area.
Helena Brunel, Alexandre Perera-Lluna, Alfonso Buil, Maria Sabater Lleal, Juan Carlos Souto, J. Fontcuberta, Montserrat Vallverdú, José Manuel Soria, Pere Caminal
BIBE2
2008 Use of Gene Ontology semantic information in protein interaction data visualization
abstract
The Gene Ontology project is an effort to structure knowledge on biological products and processes by adding semantic information to them. This is done in a systematic way so that this additional information can be automatically processed. In this contribution a protein-protein interaction visualization algorithm is proposed, which combines protein interaction data with Gene Ontology semantic information. The information is integrated using a semantic distance measure defined in ontologies or taxonomies. Multidimensional scaling is applied to this measure and the output complements protein interaction data in building an interaction visualization map.
Raimon Massanet Vila, Pere Caminal, Alexandre Perera-Lluna
BIBE3
2008 Detection of transcription factor binding sites using Rényi entropy
abstract
During the process of protein synthesis, transcription of DNA to messenger RNA starts with the binding of the transcription factors to the promoter. One of the issues on the prediction of transcription factor binding is that sequences corresponding to the binding present variability. In this manuscript a method for the detection of binding site is proposed, based on a parametric uncertainty measurement (Renyi entropy). This measurement is done through an estimation of the probability for each nucleotide avoiding any numerical representation of the nucleotides. We obtain values of the efficiency of the method as receiver operating characteristic curves found on ABF1 and ROX1 binding sites in chromosome I and XVI of the organism Saccharomyces cerevisiae.
Joan Maynou, Montserrat Vallverdú, Francesc Claria, Alexandre Perera-Lluna, Pere Caminal
BIBE4
2004 Sensor-based machine olfaction with a neurodynamics model of the olfactory bulb
abstract
We propose a biologically inspired model of olfactory processing for chemosensor arrays. The model captures three functions in the early olfactory pathway: chemotopic convergence of receptor neurons onto the olfactory bulb, center on-off surround lateral interactions, and adaptation to sustained stimuli. The projection of ORNs onto glomerular units is simulated with a self-organizing model of chemotopic convergence, which leads to odor specific spatial patterning. This information serves as an input to a network of mitral cells with center on-off surround lateral inhibition, which enhances the initial contrast among odors and decouples odor identity from intensity. Finally, slow adaptation of mitral cells adds a temporal dimension to the spatial patterns that further enhances odor discrimination. The model is validated using experimental data from an array of temperature-modulated metal-oxide sensors.
Baranidharan Raman, Agustin Gutierrez-Galvez, Alexandre Perera-Lluna, Ricardo Gutierrez-Osuna
IROS3