Jesper Tegnér

dblp:24/4407 · also Jesper N. Tegnér · DBLP profile ↗
← Back
36ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-9568-5588ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 5 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Algorithmic Complexity Underpins When Path Information Improves Graph Neural Networks Performance on Molecular Graphs
Abdulrahman Ibraheem, Narsis A. Kiani, Jesper Tegnér
ICCSA (1)3
2026 Decoding extremophiles: insights from bioinformatics, machine learning, and data-driven approaches
abstract
Life thrives in Earth's most inhospitable environments, from boiling hydrothermal vents to hypersaline lakes and frozen polar deserts, thanks to the remarkable adaptations of extremophilic microorganisms. The study of these organisms has rapidly evolved from early cultivation-based discoveries to a data-rich discipline powered by advanced omics technologies. This review comprehensively outlines the current landscape and future directions in extremophile research, emphasizing the pivotal role of bioinformatics, machine learning (ML), and data-driven approaches. We begin by charting the evolution of methodologies, from innovative in situ cultivation techniques and robust biomolecule extraction protocols to modern multi-omics workflows (metagenomics, transcriptomics, proteomics, and metabolomics) that decode the genetic and functional basis of extremophiles. We then catalogue essential bioinformatics resources and specialized databases critical for annotating extremophile genomes and uncovering their unique adaptive strategies, including protein stabilization and syntrophic metabolic relationships. Finally, we explore the transformative potential of artificial intelligence (AI) and ML in overcoming fundamental challenges in the field. These include predicting the functions of uncharacterized "hypothetical" proteins, identifying novel extremozymes, modeling complex genotype-phenotype relationships, and guiding the targeted engineering of industrially relevant strains. By synthesizing insights across these domains, this review highlights how integrating computational biology and AI is poised to unlock the full biotechnological potential of extremophiles and redefine the boundaries of life itself.
Maria N. Chasapi, Nicholas Kontis, Robert Lehmann 0002, Ruqaiya Tasneem, Niketan S. Patel, Sumeer Ahmad Khan, Xabier Martinez de Morentin, Iro N. Chasapi, Eleni Aplakidou, Alexandros Galaras, Lila Aldakheel, Minjing Su, Fotis A. Baltoumas, Kasthuri Venkateswaran, Vincenzo Lagani, David Gomez-Cabrero, Jesper Tegnér, Georgios A. Pavlopoulos, Alexandre Soares Rosado
Briefings Bioinform.17
2025 Interpretable Causal Representation Learning for Biological Data in the Pathway Space
abstract
Predicting the impact of genomic and drug perturbations in cellular function is crucial for understanding gene functions and drug effects, ultimately leading to improved therapies. To this end, Causal Representation Learning (CRL) constitutes one of the most promising approaches, as it aims to identify the latent factors that causally govern biological systems, thus facilitating the prediction of the effect of unseen perturbations. Yet, current CRL methods fail in reconciling their principled latent representations with known biological processes, leading to models that are not interpretable. To address this major issue, in this work we present SENA-discrepancy-VAE, a model based on the recently proposed CRL method discrepancy-VAE, that produces representations where each latent factor can be interpreted as the (linear) combination of the activity of a (learned) set of biological processes. To this extent, we present an encoder, SENA-$\delta$, that efficiently compute and map biological processes' activity levels to the latent causal factors. We show that SENA-discrepancy-VAE achieves predictive performances on unseen combinations of interventions that are comparable with its original, non-interpretable counterpart, while inferring causal latent factors that are biologically meaningful.
Jesus de la Fuente Cedeño, Robert Lehmann 0002, Carlos Ruiz-Arenas, Jan Voges, Irene Marín-Goñi, Xabier Martinez de Morentin, David Gomez-Cabrero, Idoia Ochoa, Jesper Tegnér, Vincenzo Lagani, Mikel Hernaez
ICLR9
2025 GeneSetCluster 2.0: a comprehensive toolset for summarizing and integrating gene-sets analysis
abstract
BACKGROUND: Gene-Set Analysis (GSA) is commonly used to analyze high-throughput experiments. However, GSA cannot readily disentangle clusters or pathways due to redundancies in upstream knowledge bases, which hinders comprehensive exploration and interpretation of biological findings. To address this challenge, we developed GeneSetCluster, an R package designed to summarize and integrate GSA results. Over time, we and users as well identified limitations in the original version, such as difficulties in managing redundancies across multiple gene-sets, large computational times, and its lack of accessibility for users without programming expertise. RESULTS: We present GeneSetCluster 2.0, a comprehensive upgrade that delivers methodological, computational, interpretative, and user-experience enhancements. Methodologically, GeneSetCluster 2.0 introduces a novel approach to address duplicated gene-sets and implements a seriation-based clustering algorithm that reorders results, aiding pattern identification. Computationally, the package is optimized for parallel processing, significantly reducing execution time. GeneSetCluster 2.0 enhances cluster annotations by associating clusters with relevant tissues and biological processes to improve biological interpretation, particularly for human and mouse data. To broaden accessibility, we have developed a user-friendly web application enabling non-programmers to use it. This version also ensures seamless integration between the R package, catering to users with programming expertise, and the web application for broader audiences. We evaluated the updates in a single-cell RNA public dataset. CONCLUSION: GeneSetCluster 2.0 offers substantial improvements over its predecessor. Furthermore, by bridging the gap between bioinformaticians and clinicians in multidisciplinary teams, GeneSetCluster 2.0 facilitates collaborative research. The R package and web application, along with detailed installation and usage guides, are available on GitHub ( https://github.com/TranslationalBioinformaticsUnit/GeneSetCluster2.0 ), and the web application can be accessed at https://translationalbio.shinyapps.io/genesetcluster/ .
Asier Ortega-Legarreta, Alberto Maillo, Daniel Mouzo, Ana Rosa López-Pérez, Lara Kular, Majid Pahlevan Kakhki, Maja Jagodic, Jesper Tegnér, Vincenzo Lagani, Ewoud Ewing, David Gomez-Cabrero
BMC Bioinform.8
2025 Minimal algorithmic information loss methods for dimension reduction, feature selection and network sparsification
abstract
We present a novel, domain-agnostic, model-independent, unsupervised, and universally applicable Machine Learning approach for dimensionality reduction based on the principles of algorithmic complexity. Specifically, but without loss of generality, we focus on addressing the challenge of reducing certain dimensionality aspects, such as the number of edges in a network, while retaining essential features of interest. These features include preserving crucial network properties like degree distribution, clustering coefficient, edge betweenness, and degree and eigenvector centralities but can also go beyond edges to nodes and weights for network pruning and trimming. Our approach outperforms classical statistical Machine Learning techniques and state-of-the-art dimensionality reduction algorithms by preserving a greater number of data features that statistical algorithms would miss, particularly nonlinear patterns stemming from deterministic recursive processes that may look statistically random but are not. Moreover, previous approaches heavily rely on a priori feature selection, which requires constant supervision. Our findings demonstrate the effectiveness of the algorithms in overcoming some of these limitations while maintaining a time-efficient computational profile. Our approach not only matches, but also exceeds, the performance of established and state-of-the-art dimensionality reduction algorithms. We extend the applicability of our method to lossy compression tasks involving images and any multi-dimensional data. This highlights the versatility and broad utility of the approach in multiple domains.
Hector Zenil, Narsis A. Kiani, Alyssa M. Adams, Felipe S. Abrahão, Antonio Rueda-Toicen, Allan A. Zea, Luan Ozelim, Jesper Tegnér
Inf. Sci.8
2024 ClustAll: An R package for patient stratification in complex diseases
abstract
In the era of precision medicine, it is necessary to understand heterogeneity among patients with complex diseases to improve personalized prevention and management strategies. Here, we introduce ClustAll, a Bioconductor package designed for unsupervised patient stratification using clinical data. ClustAll is based on the previously validated methodology ClustAll, a clustering framework that effectively handles intricacies in clinical data, including mixed data types, missing values, and collinearity. Additionally, ClustAll stands out in its ability to identify multiple patient stratifications within the same population while ensuring their robustness. The updated implementation of ClustAll features S4 classes, parallel computing for enhanced computational efficiency, and user-friendly tools for exploring and comparing stratifications against clinical phenotypes. The performance of ClustAll has been validated using two public clinical datasets, confirming its effectiveness in patient stratification and highlighting its potential impact on clinical management. In summary, ClustAll is a powerful tool for patient stratification in personalized medicine.
Asier Ortega-Legarreta, Sara Palomino-Echeverria, Estefania Huergo, Vincenzo Lagani, Narsis A. Kiani, Pierre-Emmanuel Rautou, Nuria Planell-Picola, Jesper Tegnér, David Gomez-Cabrero
PLoS Comput. Biol.8
2023 Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
abstract
Srijith Radhakrishnan, Chao-Han Yang, Sumeer Khan, Rohit Kumar, Narsis Kiani, David Gomez-Cabrero, Jesper Tegnér. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Srijith Radhakrishnan, Chao-Han Huck Yang, Sumeer Ahmad Khan, Narsis A. Kiani, David Gomez-Cabrero, Jesper Tegnér
EMNLP7
2023 A Parameter-Efficient Learning Approach to Arabic Dialect Identification with Pre-Trained General-Purpose Speech Model
abstract
In this work, we explore Parameter-Efficient-Learning (PEL) techniques to repurpose a General-Purpose-Speech (GSM) model for Arabic dialect identification (ADI). Specifically, we investigate different setups to incorporate trainable features into a multi-layer encoder-decoder GSM formulation under frozen pre-trained settings. Our architecture includes residual adapter and model reprogramming (input-prompting). We design a token-level label mapping to condition the GSM for Arabic Dialect Identification (ADI). We achieve new state-of-the-art accuracy on the ADI-17 dataset by vanilla fine-tuning. We further reduce the training budgets with the PEL method, which performs within 1.86% accuracy to fine-tuning using only 2.5% of (extra) network trainable parameters.
Srijith Radhakrishnan, Chao-Han Huck Yang, Sumeer Ahmad Khan, Narsis A. Kiani, David Gomez-Cabrero, Jesper Tegnér
INTERSPEECH6
2022 Non-local Attention Improves Description Generation for Retinal Images
abstract
Automatically generating medical reports from retinal images is a difficult task in which an algorithm must generate semantically coherent descriptions for a given retinal image. Existing methods mainly rely on the input image to generate descriptions. However, many abstract medical concepts or descriptions cannot be generated based on image information only. In this work, we integrate additional information to help solve this task; we observe that early in the diagnosis process, ophthalmologists have usually written down a small set of keywords denoting important information. These keywords are then subsequently used to aid the later creation of medical reports for a patient. Since these keywords commonly exist and are useful for generating medical reports, we incorporate them into automatic report generation. Since we have two types of inputs expert-defined unordered keywords and images - effectively fusing features from these different modalities is challenging. To that end, we propose a new keyword-driven medical report generation method based on a non-local attention-based multi-modal feature fusion approach, TransFuser, which is capable of fusing features from different types of inputs based on such attention. Our experiments show the proposed method successfully captures the mutual information of keywords and image content. We further show our proposed keyword-driven generation model reinforced by the TransFuser is superior to baselines under the popular text evaluation metrics BLEU, CIDEr, and ROUGE. Trans-Fuser Github: https://github.com/Jhhuangkay/Non-local-Attention-ImprovesDescription-Generation-for-Retinal-Images.
Jia-Hong Huang, Ting-Wei Wu, Chao-Han Huck Yang, Zenglin Shi, I-Hung Lin, Jesper Tegnér, Marcel Worring
WACV6
2021 DeepOpht: Medical Report Generation for Retinal Images via Deep Models and Visual Explanation
abstract
In this work, we propose an AI-based method that intends to improve the conventional retinal disease treatment procedure and help ophthalmologists increase diagnosis efficiency and accuracy. The proposed method is composed of a deep neural networks-based (DNN-based) module, including a retinal disease identifier and clinical description generator, and a DNN visual explanation module. To train and validate the effectiveness of our DNN-based module, we propose a large-scale retinal disease image dataset. Also, as ground truth, we provide a retinal image dataset manually labeled by ophthalmologists to qualitatively show the proposed AI-based method is effective. With our experimental results, we show that the proposed method is quantitatively and qualitatively effective. Our method is capable of creating meaningful retinal image descriptions and visual explanations that are clinically relevant.https://github.com/Jhhuangkay/DeepOpht-Medical-Report-Generation-for-Retinal-Images-via-Deep-Models-and-Visual-Explanation.
Jia-Hong Huang, Chao-Han Huck Yang, Fangyu Liu 0001, Meng Tian 0002, Yi-Chieh Liu, Ting-Wei Wu, I-Hung Lin, Hiromasa Morikawa, Hernghua Chang, Jesper Tegnér, Marcel Worring
WACV11
2021 DeepViral: prediction of novel virus-host interactions from protein sequences and infectious disease phenotypes
abstract
MOTIVATION: Infectious diseases caused by novel viruses have become a major public health concern. Rapid identification of virus-host interactions can reveal mechanistic insights into infectious diseases and shed light on potential treatments. Current computational prediction methods for novel viruses are based mainly on protein sequences. However, it is not clear to what extent other important features, such as the symptoms caused by the viruses, could contribute to a predictor. Disease phenotypes (i.e. signs and symptoms) are readily accessible from clinical diagnosis and we hypothesize that they may act as a potential proxy and an additional source of information for the underlying molecular interactions between the pathogens and hosts. RESULTS: We developed DeepViral, a deep learning based method that predicts protein-protein interactions (PPI) between humans and viruses. Motivated by the potential utility of infectious disease phenotypes, we first embedded human proteins and viruses in a shared space using their associated phenotypes and functions, supported by formalized background knowledge from biomedical ontologies. By jointly learning from protein sequences and phenotype features, DeepViral significantly improves over existing sequence-based methods for intra- and inter-species PPI prediction. AVAILABILITY AND IMPLEMENTATION: Code and datasets for reproduction and customization are available at https://github.com/bio-ontology-research-group/DeepViral. Prediction results for 14 virus families are available at https://doi.org/10.5281/zenodo.4429824. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wang Liu-Wei, Senay Kafkas, Jun Chen 0021, Nicholas J. Dimonaco, Jesper Tegnér, Robert Hoehndorf
Bioinform.5
2020 Evolving Neural Networks through a Reverse Encoding Tree
abstract
NeuroEvolution is one of the most competitive evolutionary learning strategies for designing novel neural networks for use in specific tasks, such as logic circuit design and digital gaming. However, the application of benchmark methods such as the NeuroEvolution of Augmenting Topologies (NEAT) remains a challenge, in terms of their computational cost and search time inefficiency. This paper advances a method which incorporates a type of topological edge coding, named Reverse Encoding Tree (RET), for evolving scalable neural networks efficiently. Using RET, two types of approaches - NEAT with Binary search encoding (Bi-NEAT) and NEAT with Golden-Section search encoding (GS-NEAT) - have been designed to solve problems in benchmark continuous learning environments such as logic gates, Cartpole, and Lunar Lander, and tested against classical NEAT and FS-NEAT as baselines. Additionally, we conduct a robustness test to evaluate the resilience of the proposed NEAT approaches. The results show that the two proposed approaches deliver improved performance, characterized by (1) a higher accumulated reward within a finite number of time steps; (2) using fewer episodes to solve problems in targeted environments, and (3) maintaining adaptive robustness under noisy perturbations, which outperform the baselines in all tested cases. Our analysis also demonstrates that RET expends potential future research directions in dynamic environments. Code is available from https://github.com/HaolingZHANG/ReverseEncodingTree.
Haoling Zhang, Chao-Han Huck Yang, Hector Zenil, Narsis A. Kiani, Jesper Tegnér
CEC6
2020 Interpretable Self-Attention Temporal Reasoning for Driving Behavior Understanding
abstract
Performing driving behaviors based on causal reasoning is essential to ensure driving safety. In this work, we investigated how state-of-the-art 3D Convolutional Neural Networks (CNNs) perform on classifying driving behaviors based on causal reasoning. We proposed a perturbation-based visual explanation method to inspect the models' performance visually. By examining the video attention saliency, we found that existing models could not precisely capture the causes (e.g., traffic light) of the specific action (e.g., stopping). Therefore, the Temporal Reasoning Block (TRB) was proposed and introduced to the models. With the TRB models, we achieved the accuracy of $\mathbf{86.3\%}$, which outperform the state-of-the-art 3D CNNs from previous works. The attention saliency also demonstrated that TRB helped models focus on the causes more precisely. With both numerical and visual evaluations, we concluded that our proposed TRB models were able to provide accurate driving behavior prediction by learning the causal reasoning of the behaviors.
Yi-Chieh Liu, Yung-An Hsieh, Min-Hung Chen, Chao-Han Huck Yang, Jesper Tegnér, Yichang James Tsai
ICASSP5
2019 Building gene regulatory networks from scATAC-seq and scRNA-seq using Linked Self Organizing Maps
abstract
Rapid advances in single-cell assays have outpaced methods for analysis of those data types. Different single-cell assays show extensive variation in sensitivity and signal to noise levels. In particular, scATAC-seq generates extremely sparse and noisy datasets. Existing methods developed to analyze this data require cells amenable to pseudo-time analysis or require datasets with drastically different cell-types. We describe a novel approach using self-organizing maps (SOM) to link scATAC-seq regions with scRNA-seq genes that overcomes these challenges and can generate draft regulatory networks. Our SOMatic package generates chromatin and gene expression SOMs separately and combines them using a linking function. We applied SOMatic on a mouse pre-B cell differentiation time-course using controlled Ikaros over-expression to recover gene ontology enrichments, identify motifs in genomic regions showing similar single-cell profiles, and generate a gene regulatory network that both recovers known interactions and predicts new Ikaros targets during the differentiation process. The ability of linked SOMs to detect emergent properties from multiple types of highly-dimensional genomic data with very different signal properties opens new avenues for integrative analysis of heterogeneous data.
Camden Jansen, Ricardo N. Ramirez, Nicole C. El-Ali, David Gomez-Cabrero, Jesper Tegnér, Matthias Merkenschlager, Ana Conesa
PLoS Comput. Biol.5
2017 HiDi: an efficient reverse engineering schema for large-scale dynamic regulatory network reconstruction using adaptive differentiation
abstract
MOTIVATION: The use of differential equations (ODE) is one of the most promising approaches to network inference. The success of ODE-based approaches has, however, been limited, due to the difficulty in estimating parameters and by their lack of scalability. Here, we introduce a novel method and pipeline to reverse engineer gene regulatory networks from gene expression of time series and perturbation data based upon an improvement on the calculation scheme of the derivatives and a pre-filtration step to reduce the number of possible links. The method introduces a linear differential equation model with adaptive numerical differentiation that is scalable to extremely large regulatory networks. RESULTS: We demonstrate the ability of this method to outperform current state-of-the-art methods applied to experimental and synthetic data using test data from the DREAM4 and DREAM5 challenges. Our method displays greater accuracy and scalability. We benchmark the performance of the pipeline with respect to dataset size and levels of noise. We show that the computation time is linear over various network sizes. AVAILABILITY AND IMPLEMENTATION: The Matlab code of the HiDi implementation is available at: www.complexitycalculator.com/HiDiScript.zip. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yue Deng 0008, Hector Zenil, Jesper Tegnér, Narsis A. Kiani
Bioinform.3
2017 Dynamics and heterogeneity of brain damage in multiple sclerosis
abstract
Multiple Sclerosis (MS) is an autoimmune disease driving inflammatory and degenerative processes that damage the central nervous system (CNS). However, it is not well understood how these events interact and evolve to evoke such a highly dynamic and heterogeneous disease. We established a hypothesis whereby the variability in the course of MS is driven by the very same pathogenic mechanisms responsible for the disease, the autoimmune attack on the CNS that leads to chronic inflammation, neuroaxonal degeneration and remyelination. We propose that each of these processes acts more or less severely and at different times in each of the clinical subgroups. To test this hypothesis, we developed a mathematical model that was constrained by experimental data (the expanded disability status scale [EDSS] time series) obtained from a retrospective longitudinal cohort of 66 MS patients with a long-term follow-up (up to 20 years). Moreover, we validated this model in a second prospective cohort of 120 MS patients with a three-year follow-up, for which EDSS data and brain volume time series were available. The clinical heterogeneity in the datasets was reduced by grouping the EDSS time series using an unsupervised clustering analysis. We found that by adjusting certain parameters, albeit within their biological range, the mathematical model reproduced the different disease courses, supporting the dynamic CNS damage hypothesis to explain MS heterogeneity. Our analysis suggests that the irreversible axon degeneration produced in the early stages of progressive MS is mainly due to the higher rate of myelinated axon degeneration, coupled to the lower capacity for remyelination. However, and in agreement with recent pathological studies, degeneration of chronically demyelinated axons is not a key feature that distinguishes this phenotype. Moreover, the model reveals that lower rates of axon degeneration and more rapid remyelination make relapsing MS more resilient than the progressive subtype. Therefore, our results support the hypothesis of a common pathogenesis for the different MS subtypes, even in the presence of genetic and environmental heterogeneity. Hence, MS can be considered as a single disease in which specific dynamics can provoke a variety of clinical outcomes in different patient groups. These results have important implications for the design of therapeutic interventions for MS at different stages of the disease.
Ekaterina Kotelnikova, Narsis A. Kiani, Elena Abad, Elena H. Martinez-Lapiscina, Magí Andorrà, Irati Zubizarreta, Irene Pulido-Valdeolivas, Inna Pertsovskaya, Leonidas G. Alexopoulos, Tomas Olsson, Roland Martin, Friedemann Paul, Jesper Tegnér, Jordi García-Ojalvo, Pablo Villoslada
PLoS Comput. Biol.13
2016 Normalization of circulating microRNA expression data obtained by quantitative real-time RT-PCR
abstract
The high-throughput analysis of microRNAs (miRNAs) circulating within the blood of healthy and diseased individuals is an active area of biomarker research. Whereas quantitative real-time reverse transcription polymerase chain reaction (qPCR)-based methods are widely used, it is yet unresolved how the data should be normalized. Here, we show that a combination of different algorithms results in the identification of candidate reference miRNAs that can be exploited as normalizers, in both discovery and validation phases. Using the methodology considered here, we identify normalizers that are able to reduce nonbiological variation in the data and we present several case studies, to illustrate the relevance in the context of physiological or pathological scenarios. In conclusion, the discovery of stable reference miRNAs from high-throughput studies allows appropriate normalization of focused qPCR assays.
Francesco Marabita, Paola de Candia, Anna Torri, Jesper Tegnér, Sergio Abrignani, Riccardo L. Rossi
Briefings Bioinform.4
2016 From comorbidities of chronic obstructive pulmonary disease to identification of shared molecular mechanisms by data integration
abstract
BACKGROUND: Deep mining of healthcare data has provided maps of comorbidity relationships between diseases. In parallel, integrative multi-omics investigations have generated high-resolution molecular maps of putative relevance for understanding disease initiation and progression. Yet, it is unclear how to advance an observation of comorbidity relations (one disease to others) to a molecular understanding of the driver processes and associated biomarkers. RESULTS: Since Chronic Obstructive Pulmonary disease (COPD) has emerged as a central hub in temporal comorbidity networks, we developed a systematic integrative data-driven framework to identify shared disease-associated genes and pathways, as a proxy for the underlying generative mechanisms inducing comorbidity. We integrated records from approximately 13 M patients from the Medicare database with disease-gene maps that we derived from several resources including a semantic-derived knowledge-base. Using rank-based statistics we not only recovered known comorbidities but also discovered a novel association between COPD and digestive diseases. Furthermore, our analysis provides the first set of COPD co-morbidity candidate biomarkers, including IL15, TNF and JUP, and characterizes their association to aging and life-style conditions, such as smoking and physical activity. CONCLUSIONS: The developed framework provides novel insights in COPD and especially COPD co-morbidity associated mechanisms. The methodology could be used to discover and decipher the molecular underpinning of other comorbidity relationships and furthermore, allow the identification of candidate co-morbidity biomarkers.
David Gomez-Cabrero, Jörg Menche, Claudia Vargas, Isaac Cano, Dieter Maier, Albert-László Barabási, Jesper Tegnér, Josep Roca
BMC Bioinform.7
2013 Algorithmic complexity of motifs clusters superfamilies of networks
abstract
Representing biological systems as networks has proved to be very powerful. For example, local graph analysis of substructures such as subgraph over representation (or motifs) has elucidated different sub-types of networks. Here we report that using numerical approximations of Kolmogorov complexity, by means of algorithmic probability, clusters different classes of networks. For this, we numerically estimate the algorithmic probability of the sub-matrices from the adjacency matrix of the original network (hence including motifs). We conclude that algorithmic information theory is a powerful tool supplementing other network analysis techniques.
Hector Zenil, Narsis A. Kiani, Jesper Tegnér
BIBM3
2013 A beta-mixture quantile normalization method for correcting probe design bias in Illumina Infinium 450 k DNA methylation data
abstract
MOTIVATION: The Illumina Infinium 450 k DNA Methylation Beadchip is a prime candidate technology for Epigenome-Wide Association Studies (EWAS). However, a difficulty associated with these beadarrays is that probes come in two different designs, characterized by widely different DNA methylation distributions and dynamic range, which may bias downstream analyses. A key statistical issue is therefore how best to adjust for the two different probe designs. RESULTS: Here we propose a novel model-based intra-array normalization strategy for 450 k data, called BMIQ (Beta MIxture Quantile dilation), to adjust the beta-values of type2 design probes into a statistical distribution characteristic of type1 probes. The strategy involves application of a three-state beta-mixture model to assign probes to methylation states, subsequent transformation of probabilities into quantiles and finally a methylation-dependent dilation transformation to preserve the monotonicity and continuity of the data. We validate our method on cell-line data, fresh frozen and paraffin-embedded tumour tissue samples and demonstrate that BMIQ compares favourably with two competing methods. Specifically, we show that BMIQ improves the robustness of the normalization procedure, reduces the technical variation and bias of type2 probe values and successfully eliminates the type1 enrichment bias caused by the lower dynamic range of type2 probes. BMIQ will be useful as a preprocessing step for any study using the Illumina Infinium 450 k platform. AVAILABILITY: BMIQ is freely available from http://code.google.com/p/bmiq/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Andrew E. Teschendorff, Francesco Marabita, Matthias Lechner, Thomas E. Bartlett, Jesper Tegnér, David Gomez-Cabrero, Stephan Beck 0002
Bioinform.5
2009 On reliable discovery of molecular signatures
abstract
BACKGROUND: Molecular signatures are sets of genes, proteins, genetic variants or other variables that can be used as markers for a particular phenotype. Reliable signature discovery methods could yield valuable insight into cell biology and mechanisms of human disease. However, it is currently not clear how to control error rates such as the false discovery rate (FDR) in signature discovery. Moreover, signatures for cancer gene expression have been shown to be unstable, that is, difficult to replicate in independent studies, casting doubts on their reliability. RESULTS: We demonstrate that with modern prediction methods, signatures that yield accurate predictions may still have a high FDR. Further, we show that even signatures with low FDR may fail to replicate in independent studies due to limited statistical power. Thus, neither stability nor predictive accuracy are relevant when FDR control is the primary goal. We therefore develop a general statistical hypothesis testing framework that for the first time provides FDR control for signature discovery. Our method is demonstrated to be correct in simulation studies. When applied to five cancer data sets, the method was able to discover molecular signatures with 5% FDR in three cases, while two data sets yielded no significant findings. CONCLUSION: Our approach enables reliable discovery of molecular signatures from genome-wide data with current sample sizes. The statistical framework developed herein is potentially applicable to a wide range of prediction problems in bioinformatics.
Roland Nilsson, Johan Björkegren, Jesper Tegnér
BMC Bioinform.3
2009 An Algorithm for Reading Dependencies from the Minimal Undirected Independence Map of a Graphoid that Satisfies Weak Transitivity
José M. Peña 0001, Roland Nilsson, Johan Björkegren, Jesper Tegnér
J. Mach. Learn. Res.4
2008 Networks in Biology - From Identification, Analysis to Interpretation
Jesper Tegnér
GD1
2008 Electrotonic Signals along Intracellular Membranes May Interconnect Dendritic Spines and Nucleus
abstract
Synapses on dendritic spines of pyramidal neurons show a remarkable ability to induce phosphorylation of transcription factors at the nuclear level with a short latency, incompatible with a diffusion process from the dendritic spines to the nucleus. To account for these findings, we formulated a novel extension of the classical cable theory by considering the fact that the endoplasmic reticulum (ER) is an effective charge separator, forming an intrinsic compartment that extends from the spine to the nuclear membrane. We use realistic parameters to show that an electrotonic signal may be transmitted along the ER from the dendritic spines to the nucleus. We found that this type of signal transduction can additionally account for the remarkable ability of the cell nucleus to differentiate between depolarizing synaptic signals that originate from the dendritic spines and back-propagating action potentials. This study considers a novel computational role for dendritic spines, and sheds new light on how spines and ER may jointly create an additional level of processing within the single neuron.
Isaac Shemer, Björn Brinne, Jesper Tegnér, Sten Grillner
PLoS Comput. Biol.3
2007 Detecting multivariate differentially expressed genes
abstract
BACKGROUND: Gene expression is governed by complex networks, and differences in expression patterns between distinct biological conditions may therefore be complex and multivariate in nature. Yet, current statistical methods for detecting differential expression merely consider the univariate difference in expression level of each gene in isolation, thus potentially neglecting many genes of biological importance. RESULTS: We have developed a novel algorithm for detecting multivariate expression patterns, named Recursive Independence Test (RIT). This algorithm generalizes differential expression testing to more complex expression patterns, while still including genes found by the univariate approach. We prove that RIT is consistent and controls error rates for small sample sizes. Simulation studies confirm that RIT offers more power than univariate differential expression analysis when multivariate effects are present. We apply RIT to gene expression data sets from diabetes and cancer studies, revealing several putative disease genes that were not detected by univariate differential expression analysis. CONCLUSION: The proposed RIT algorithm increases the power of gene expression analysis by considering multivariate effects while retaining error rate control, and may be useful when conventional differential expression tests yield few findings.
Roland Nilsson, José M. Peña 0001, Johan Björkegren, Jesper Tegnér
BMC Bioinform.4
2007 Towards scalable and data efficient learning of Markov boundaries
José M. Peña 0001, Roland Nilsson, Johan Björkegren, Jesper Tegnér
Int. J. Approx. Reason.4
2007 Consistent Feature Selection for Pattern Recognition in Polynomial Time
Roland Nilsson, José M. Peña 0001, Johan Björkegren, Jesper Tegnér
J. Mach. Learn. Res.4
2006 Evaluating Feature Selection for SVMs in High Dimensions
Roland Nilsson, José M. Peña 0001, Johan Björkegren, Jesper Tegnér
ECML4
2006 Identifying the Relevant Nodes Without Learning the Model
José M. Peña 0001, Roland Nilsson, Johan Björkegren, Jesper Tegnér
UAI4
2006 Detection of compound mode of action by computational integration of whole-genome measurements and genetic perturbations
abstract
BACKGROUND: A key problem of drug development is to decide which compounds to evaluate further in expensive clinical trials (Phase I- III). This decision is primarily based on the primary targets and mechanisms of action of the chemical compounds under consideration. Whole-genome expression measurements have shown to be useful for this process but current approaches suffer from requiring either a large number of mutant experiments or a detailed understanding of the regulatory networks. RESULTS: We have designed an algorithm, CutTree that when applied to whole-genome expression datasets identifies the primary affected genes (PAGs) of a chemical compound by separating them from downstream, indirectly affected genes. Unlike previous methods requiring whole-genome deletion libraries or a complete map of gene network architecture, CutTree identifies PAGs from a limited set of experimental perturbations without requiring any prior information about the underlying pathways. The principle for CutTree is to iteratively filter out PAGs from other recurrently active genes (RAGs) that are not PAGs. The in silico validation predicted that CutTree should be able to identify 3-4 out of 5 known PAGs (approximately 70%). In accordance, when we applied CutTree to whole-genome expression profiles from 17 genetic perturbations in the presence of galactose in Yeast, CutTree identified four out of five known primary galactose targets (80%). Using an exhaustive search strategy to detect these PAGs would not have been feasible (>1012 combinations). CONCLUSION: In combination with genetic perturbation techniques like short interfering RNA (siRNA) followed by whole-genome expression measurements, CutTree sets the stage for compound target identification in less well-characterized but more disease-relevant mammalian cell systems.
Kristofer Hallén, Johan Björkegren, Jesper Tegnér
BMC Bioinform.3
2005 Scalable, Efficient and Correct Learning of Markov Boundaries Under the Faithfulness Assumption
José M. Peña 0001, Johan Björkegren, Jesper Tegnér
ECSQARU3
2005 Learning dynamic Bayesian network models via cross-validation
José M. Peña 0001, Johan Björkegren, Jesper Tegnér
Pattern Recognit. Lett.3
2002 An adaptive spike-timing-dependent plasticity rule
Jesper Tegnér, Ádám Kepecs
Neurocomputing1
2001 Why Neuronal Dynamics Should Control Synaptic Learning Rules
abstract
Hebbian learning rules are generally formulated as static rules. Un(cid:173) der changing condition (e.g. neuromodulation, input statistics) most rules are sensitive to parameters. In particular, recent work has focused on two different formulations of spike-timing-dependent plasticity rules. Additive STDP [1] is remarkably versatile but also very fragile, whereas multiplicative STDP [2, 3] is more ro(cid:173) bust but lacks attractive features such as synaptic competition and rate stabilization. Here we address the problem of robustness in the additive STDP rule. We derive an adaptive control scheme, where the learning function is under fast dynamic control by post(cid:173) synaptic activity to stabilize learning under a variety of conditions. Such a control scheme can be implemented using known biophysical mechanisms of synapses. We show that this adaptive rule makes the addit ive STDP more robust. Finally, we give an example how meta plasticity of the adaptive rule can be used to guide STDP into different type of learning regimes.
Jesper Tegnér, Ádám Kepecs
NIPS1
1999 Control of burst proportion and frequency range by drive-dependent modulation of adaptation
Jeanette Kotaleski, Jesper Tegnér, Sten Grillner, Anders Lansner
Neurocomputing2
1999 The synaptic NMDA component desynchronizes neural bursters
Jesper Tegnér, Jeanette Kotaleski
Neurocomputing1