Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Djork-Arné Clevert

dblp:03/1617 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0003-4191-2156ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Generative modeling · 29% Representation and self-supervised learning · 19% Question answering and dialogue systems · 18%
Interdisciplinary, comprehensive, and emerging computing
10 papers
Bioinformatics and computational biology · 93% Medical and health informatics · 7%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 30 heaviest of 41, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
biomedical question answering
0.912025
KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA · ICLR 2025
Natural language and speech › Question answering and dialogue systems
knowledge-intensive question answering
0.912025
KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA · ICLR 2025
Knowledge graphs
knowledge graph reasoning
0.912025
KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA · ICLR 2025
Machine learning › Generative modeling › molecular generation
3d molecule generation
0.812024
Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule Generation · ICLR 2024
Machine learning › Generative modeling
diffusion model
0.812024
Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule Generation · ICLR 2024
Machine learning › Generative modeling › diffusion model › geometric diffusion model
equivariant diffusion model
0.812024
Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule Generation · ICLR 2024
Bioinformatics and computational biology
epigenomics
0.812024
A fast machine learning dataloader for epigenetic tracks from BigWig files · Bioinform. 2024
Bioinformatics and computational biology › genomics
machine learning for genomics
0.812024
A fast machine learning dataloader for epigenetic tracks from BigWig files · Bioinform. 2024
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning
0.612022
Unsupervised Learning of Group Invariant and Equivariant Representations · NeurIPS 2022
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning
0.612022
Unsupervised Learning of Group Invariant and Equivariant Representations · NeurIPS 2022
Machine learning › Learning paradigms
unsupervised learning
0.612022
Unsupervised Learning of Group Invariant and Equivariant Representations · NeurIPS 2022
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning
0.522017
Rectified factor networks for biclustering of omics data · Bioinform. 2017
Rectified Factor Networks · NIPS 2015
Machine learning › Graph learning
graph generation
0.512021
Permutation-Invariant Variational Autoencoder for Graph-Level Representation Learning · NeurIPS 2021
Machine learning › Trustworthy machine learning › interpretability
graph neural network explanation
0.512021
Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity · ICML 2021
Machine learning › Graph learning
graph representation learning
0.512021
Permutation-Invariant Variational Autoencoder for Graph-Level Representation Learning · NeurIPS 2021
Machine learning › Trustworthy machine learning
interpretability
0.512021
Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity · ICML 2021
Machine learning › Graph learning › molecular representation learning › molecular graph learning
molecular graph neural network
0.512021
Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity · ICML 2021
Machine learning › Generative modeling
variational autoencoder
0.512021
Permutation-Invariant Variational Autoencoder for Graph-Level Representation Learning · NeurIPS 2021
Medical and health informatics › clinical data analysis
phenotyping
0.512021
Self-supervised feature extraction from image time series in plant phenotyping using triplet networks · Bioinform. 2021
Bioinformatics and computational biology › molecular property prediction
pka prediction
0.512021
pKPDB: a protein data bank extension database of pKa and pI theoretical values · Bioinform. 2021
Bioinformatics and computational biology › plant biology
plant phenotyping
0.512021
Self-supervised feature extraction from image time series in plant phenotyping using triplet networks · Bioinform. 2021
Bioinformatics and computational biology › molecular property prediction
protein property prediction
0.512021
pKPDB: a protein data bank extension database of pKa and pI theoretical values · Bioinform. 2021
Bioinformatics and computational biology
gene expression analysis
0.532017
Rectified factor networks for biclustering of omics data · Bioinform. 2017
FABIA: factor analysis for bicluster acquisition · Bioinform. 2010
I/NI-calls for the exclusion of non-informative genes: a highly effective filtering tool for microarray data · Bioinform. 2007
Bioinformatics and computational biology
drug discovery
0.412020
grünifai: interactive multiparameter optimization of molecules in a continuous vector space · Bioinform. 2020
Bioinformatics and computational biology › drug discovery
molecular optimization
0.412020
grünifai: interactive multiparameter optimization of molecules in a continuous vector space · Bioinform. 2020
Bioinformatics and computational biology › gene expression analysis
biclustering
0.422017
Rectified factor networks for biclustering of omics data · Bioinform. 2017
FABIA: factor analysis for bicluster acquisition · Bioinform. 2010
Bioinformatics and computational biology › genome editing
CRISPR guide RNA design
0.412019
PAVOOC: designing CRISPR sgRNAs using 3D protein structures and functional domain annotations · Bioinform. 2019
Bioinformatics and computational biology
genome editing
0.412019
PAVOOC: designing CRISPR sgRNAs using 3D protein structures and functional domain annotations · Bioinform. 2019
Bioinformatics and computational biology › molecular informatics
molecular design
0.212024
Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule Generation · ICLR 2024
Machine learning › Representation and self-supervised learning
pre-training
0.212015
Rectified Factor Networks · NIPS 2015

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 1.7large language model · 1.7time-dependent loss weighting · 1.5in silico modeling · 0.9parallel interval processing · 0.8e(3)-equivariant graph neural networks · 0.8e(3)-equivariant graph neural network · 0.8GPU decompression · 0.8group theory · 0.6encoder-decoder framework · 0.6posterior regularization · 0.5alternating minimization · 0.5triplet network · 0.5transfer learning · 0.5self-supervised learning · 0.5poisson-boltzmann calculation · 0.5monte carlo calculation · 0.5gini regularization · 0.5
YearPublicationVenuePosition
2025 KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA
abstract
Biomedical reasoning integrates structured, codified knowledge with tacit, experience-driven insights. Depending on the context, quantity, and nature of available evidence, researchers and clinicians use diverse strategies, including rule-based, prototype-based, and case-based reasoning. Effective medical AI models must handle this complexity while ensuring reliability and adaptability. We introduce KGARevion, a knowledge graph-based agent that answers knowledge-intensive questions. Upon receiving a query, KGARevion generates relevant triplets by leveraging the latent knowledge embedded in a large language model. It then verifies these triplets against a grounded knowledge graph, filtering out errors and retaining only accurate, contextually relevant information for the final answer. This multi-step process strengthens reasoning, adapts to different models of medical inference, and outperforms retrieval-augmented generation-based approaches that lack effective verification mechanisms. Evaluations on medical QA benchmarks show that KGARevion improves accuracy by over 5.2% over 15 models in handling complex medical queries. To further assess its effectiveness, we curated three new medical QA datasets with varying levels of semantic complexity, where KGARevion improved accuracy by 10.4%. The agent integrates with different LLMs and biomedical knowledge graphs for broad applicability across knowledge-intensive tasks. We evaluated KGARevion on AfriMed-QA, a newly introduced dataset focused on African healthcare, demonstrating its strong zero-shot generalization to underrepresented medical contexts.
Xiao-Rui Su 0001, Yibo Wang 0001, Shanghua Gao, Xiaolong Liu 0012, Valentina Giunchiglia, Djork-Arné Clevert, Marinka Zitnik
ICLR6
2025 Diffusion Generative Modeling on Lie Group Representations
abstract
We introduce a novel class of score-based diffusion processes that operate directly in the representation space of Lie groups. Leveraging the framework of Generalized Score Matching, we derive a class of Langevin dynamics that decomposes as a direct sum of Lie algebra representations, enabling the modeling of any target distribution on any (non-Abelian) Lie group. Standard score-matching emerges as a special case of our framework when the Lie group is the translation group. We prove that our generalized generative processes arise as solutions to a new class of paired stochastic differential equations (SDEs), introduced here for the first time. We validate our approach through experiments on diverse data types, demonstrating its effectiveness in real-world applications such as $\text{SO}(3)$-guided molecular conformer generation and modeling ligand-specific global $\text{SE}(3)$ transformations for molecular docking, showing improvement in comparison to Riemannian diffusion on the group itself. We show that an appropriate choice of Lie group enhances learning efficiency by reducing the effective dimensionality of the trajectory space and enables the modeling of transitions between complex data distributions.
Marco Bertolini, Djork-Arné Clevert
NeurIPS3
2024 Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule Generation
abstract
Deep generative diffusion models are a promising avenue for 3D de novo molecular design in materials science and drug discovery. However, their utility is still limited by suboptimal performance on large molecular structures and limited training data. To address this gap, we explore the design space of E(3)-equivariant diffusion models, focusing on previously unexplored areas. Our extensive comparative analysis evaluates the interplay between continuous and discrete state spaces. From this investigation, we present the EQGAT-diff model, which consistently outperforms established models for the QM9 and GEOM-Drugs datasets. Significantly, EQGAT-diff takes continuous atom positions, while chemical elements and bond types are categorical and uses time-dependent loss weighting, substantially increasing training convergence, the quality of generated samples, and inference time. We also showcase that including chemically motivated additional features like hybridization states in the diffusion process enhances the validity of generated molecules. To further strengthen the applicability of diffusion models to limited training data, we investigate the transferability of EQGAT-diff trained on the large PubChem3D dataset with implicit hydrogen atoms to target different data distributions. Fine-tuning EQGAT-diff for just a few iterations shows an efficient distribution shift, further improving performance throughout data sets. Finally, we test our model on the Crossdocked data set for structure-based de novo ligand generation, underlining the importance of our findings showing state-of-the-art performance on Vina docking scores.
Tuan Le, Julian Cremer, Frank Noé, Djork-Arné Clevert, Kristof Schütt
ICLR4
2024 A fast machine learning dataloader for epigenetic tracks from BigWig files
abstract
SUMMARY: We created bigwig-loader, a data-loader for epigenetic profiles from BigWig files that decompresses and processes information for multiple intervals from multiple BigWig files in parallel. This is an access pattern needed to create training batches for typical machine learning models on epigenetics data. Using a new codec, the decompression can be done on a graphical processing unit (GPU) making it fast enough to create the training batches during training, mitigating the need for saving preprocessed training examples to disk. AVAILABILITY AND IMPLEMENTATION: The bigwig-loader installation instructions and source code can be accessed at https://github.com/pfizer-opensource/bigwig-loader.
Joren Sebastian Retel, Andreas Poehlmann, Josh Chiou, Andreas Steffen, Djork-Arné Clevert
Bioinform.5
2023 Explaining, Evaluating and Enhancing Neural Networks' Learned Representations
abstract
Most efforts in interpretability in deep learning have focused on (1) extracting explanations of a specific downstream task in relation to the input features and (2) imposing constraints on the model, often at the expense of predictive performance. New advances in (unsupervised) representation learning and transfer learning, however, raise the need for an explanatory framework for networks that are trained without a specific downstream task. We address these challenges by showing how explainability can be an aid, rather than an obstacle, towards better and more efficient representations. Specifically, we propose a natural aggregation method generalizing attribution maps between any two (convolutional) layers of a neural network. Additionally, we employ such attributions to define two novel scores for evaluating the informativeness and the disentanglement of latent embeddings. Extensive experiments show that the proposed scores do correlate with the desired properties. We also confirm and extend previously known results concerning the independence of some common saliency strategies from the model parameters. Finally, we show that adopting our proposed scores as constraints during the training of a representation learning task improves the downstream performance of the model.
Marco Bertolini, Djork-Arné Clevert, Floriane Montanari
ICANN (5)2
2022 Unsupervised Learning of Group Invariant and Equivariant Representations
abstract
Equivariant neural networks, whose hidden features transform according to representations of a group $G$ acting on the data, exhibit training efficiency and an improved generalisation performance. In this work, we extend group invariant and equivariant representation learning to the field of unsupervised deep learning. We propose a general learning strategy based on an encoder-decoder framework in which the latent representation is separated in an invariant term and an equivariant group action component. The key idea is that the network learns to encode and decode data to and from a group-invariant representation by additionally learning to predict the appropriate group action to align input and output pose to solve the reconstruction task. We derive the necessary conditions on the equivariant encoder, and we present a construction valid for any $G$, both discrete and continuous. We describe explicitly our construction for rotations, translations and permutations. We test the validity and the robustness of our approach in a variety of experiments with diverse data types employing different network architectures.
Robin Winter, Marco Bertolini, Tuan Le, Frank Noé, Djork-Arné Clevert
NeurIPS5
2021 Parameterized Hypercomplex Graph Neural Networks for Graph Classification
Tuan Le, Marco Bertolini, Frank Noé, Djork-Arné Clevert
ICANN (3)4
2021 Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity
abstract
Rationalizing which parts of a molecule drive the predictions of a molecular graph convolutional neural network (GCNN) can be difficult. To help, we propose two simple regularization techniques to apply during the training of GCNNs: Batch Representation Orthonormalization (BRO) and Gini regularization. BRO, inspired by molecular orbital theory, encourages graph convolution operations to generate orthonormal node embeddings. Gini regularization is applied to the weights of the output layer and constrains the number of dimensions the model can use to make predictions. We show that Gini and BRO regularization can improve the accuracy of state-of-the-art GCNN attribution methods on artificial benchmark datasets. In a real-world setting, we demonstrate that medicinal chemists significantly prefer explanations extracted from regularized models. While we only study these regularizers in the context of GCNNs, both can be applied to other types of neural networks.
Ryan Henderson, Djork-Arné Clevert, Floriane Montanari
ICML2
2021 Permutation-Invariant Variational Autoencoder for Graph-Level Representation Learning
abstract
Recently, there has been great success in applying deep neural networks on graph structured data. Most work, however, focuses on either node- or graph-level supervised learning, such as node, link or graph classification or node-level unsupervised learning (e.g. node clustering). Despite its wide range of possible applications, graph-level unsupervised learning has not received much attention yet. This might be mainly attributed to the high representation complexity of graphs, which can be represented by $n!$ equivalent adjacency matrices, where $n$ is the number of nodes.In this work we address this issue by proposing a permutation-invariant variational autoencoder for graph structured data. Our proposed model indirectly learns to match the node ordering of input and output graph, without imposing a particular node ordering or performing expensive graph matching. We demonstrate the effectiveness of our proposed model for graph reconstruction, generation and interpolation and evaluate the expressive power of extracted representations for downstream graph-level classification and regression.
Robin Winter, Frank Noé, Djork-Arné Clevert
NeurIPS3
2021 pKPDB: a protein data bank extension database of pKa and pI theoretical values
abstract
SUMMARY: pKa values of ionizable residues and isoelectric points of proteins provide valuable local and global insights about their structure and function. These properties can be estimated with reasonably good accuracy using Poisson-Boltzmann and Monte Carlo calculations at a considerable computational cost (from some minutes to several hours). pKPDB is a database of over 12 M theoretical pKa values calculated over 120k protein structures deposited in the Protein Data Bank. By providing precomputed pKa and pI values, users can retrieve results instantaneously for their protein(s) of interest while also saving countless hours and resources that would be spent on repeated calculations. Furthermore, there is an ever-growing imbalance between experimental pKa and pI values and the number of resolved structures. This database will complement the experimental and computational data already available and can also provide crucial information regarding buried residues that are under-represented in experimental measurements. AVAILABILITY AND IMPLEMENTATION: Gzipped csv files containing p Ka and isoelectric point values can be downloaded from https://pypka.org/pKPDB. To query a single PDB code please use the PypKa free server at https://pypka.org. The pKPDB source code can be found at https://github.com/mms-fcul/pKPDB. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pedro B. P. S. Reis, Djork-Arné Clevert, Miguel Machuqueiro
Bioinform.2
2021 Self-supervised feature extraction from image time series in plant phenotyping using triplet networks
abstract
MOTIVATION: Image-based profiling combines high-throughput screening with multiparametric feature analysis to capture the effect of perturbations on biological systems. This technology has attracted increasing interest in the field of plant phenotyping, promising to accelerate the discovery of novel herbicides. However, the extraction of meaningful features from unlabeled plant images remains a big challenge. RESULTS: We describe a novel data-driven approach to find feature representations from plant time-series images in a self-supervised manner by using time as a proxy for image similarity. In the spirit of transfer learning, we first apply an ImageNet-pretrained architecture as a base feature extractor. Then, we extend this architecture with a triplet network to refine and reduce the dimensionality of extracted features by ranking relative similarities between consecutive and non-consecutive time points. Without using any labels, we produce compact, organized representations of plant phenotypes and demonstrate their superior applicability to clustering, image retrieval and classification tasks. Besides time, our approach could be applied using other surrogate measures of phenotype similarity, thus providing a versatile method of general interest to the phenotypic profiling community. AVAILABILITY AND IMPLEMENTATION: Source code is provided in https://github.com/bayer-science-for-a-better-life/plant-triplet-net. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Paula A. Marin Zapata, Sina Roth, Dirk Schmutzler, Erica Manesso, Djork-Arné Clevert
Bioinform.6
2020 grünifai: interactive multiparameter optimization of molecules in a continuous vector space
abstract
SUMMARY: Optimizing small molecules in a drug discovery project is a notoriously difficult task as multiple molecular properties have to be considered and balanced at the same time. In this work, we present our novel interactive in silico compound optimization platform termed grünifai to support the ideation of the next generation of compounds under the constraints of a multiparameter objective. grünifai integrates adjustable in silico models, a continuous representation of the chemical space, a scalable particle swarm optimization algorithm and the possibility to actively steer the compound optimization through providing feedback on generated intermediate structures. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available under an MIT license and are openly available on GitHub (https://github.com/jrwnter/gruenifai). The backend, including the optimization method and distribution on multiple GPU nodes is written in Python 3. The frontend is written in ReactJS.
Robin Winter, Joren Sebastian Retel, Frank Noé, Djork-Arné Clevert, Andreas Steffen
Bioinform.4
2019 PAVOOC: designing CRISPR sgRNAs using 3D protein structures and functional domain annotations
abstract
SUMMARY: Single-guide RNAs (sgRNAs) targeting the same gene can significantly vary in terms of efficacy and specificity. PAVOOC (Prediction And Visualization of On- and Off-targets for CRISPR) is a web-based CRISPR sgRNA design tool that employs state of the art machine learning models to prioritize most effective candidate sgRNAs. In contrast to other tools, it maps sgRNAs to functional domains and protein structures and visualizes cut sites on corresponding protein crystal structures. Furthermore, PAVOOC supports homology-directed repair template generation for genome editing experiments and the visualization of the mutated amino acids in 3D. AVAILABILITY AND IMPLEMENTATION: PAVOOC is available under https://pavooc.me and accessible using modern browsers (Chrome/Chromium recommended). The source code is hosted at github.com/moritzschaefer/pavooc under the MIT License. The backend, including data processing steps, and the frontend are implemented in Python 3 and ReactJS, respectively. All components run in a simple Docker environment. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Moritz Schaefer, Djork-Arné Clevert, Bertram Weiss, Andreas Steffen
Bioinform.2
2017 Rectified factor networks for biclustering of omics data
abstract
MOTIVATION: Biclustering has become a major tool for analyzing large datasets given as matrix of samples times features and has been successfully applied in life sciences and e-commerce for drug design and recommender systems, respectively. actor nalysis for cluster cquisition (FABIA), one of the most successful biclustering methods, is a generative model that represents each bicluster by two sparse membership vectors: one for the samples and one for the features. However, FABIA is restricted to about 20 code units because of the high computational complexity of computing the posterior. Furthermore, code units are sometimes insufficiently decorrelated and sample membership is difficult to determine. We propose to use the recently introduced unsupervised Deep Learning approach Rectified Factor Networks (RFNs) to overcome the drawbacks of existing biclustering methods. RFNs efficiently construct very sparse, non-linear, high-dimensional representations of the input via their posterior means. RFN learning is a generalized alternating minimization algorithm based on the posterior regularization method which enforces non-negative and normalized posterior means. Each code unit represents a bicluster, where samples for which the code unit is active belong to the bicluster and features that have activating weights to the code unit belong to the bicluster. RESULTS: On 400 benchmark datasets and on three gene expression datasets with known clusters, RFN outperformed 13 other biclustering methods including FABIA. On data of the 1000 Genomes Project, RFN could identify DNA segments which indicate, that interbreeding with other hominins starting already before ancestors of modern humans left Africa. AVAILABILITY AND IMPLEMENTATION: https://github.com/bioinf-jku/librfn. CONTACT: [email protected] or [email protected].
Djork-Arné Clevert, Thomas Unterthiner, Gundula Povysil, Sepp Hochreiter
Bioinform.1
2015 Rectified Factor Networks
abstract
We propose rectified factor networks (RFNs) to efficiently construct very sparse, non-linear, high-dimensional representations of the input. RFN models identify rare and small events, have a low interference between code units, have a small reconstruction error, and explain the data covariance structure. RFN learning is a generalized alternating minimization algorithm derived from the posterior regularization method which enforces non-negative and normalized posterior means. We proof convergence and correctness of the RFN learning algorithm.On benchmarks, RFNs are compared to other unsupervised methods like autoencoders, RBMs, factor analysis, ICA, and PCA. In contrast to previous sparse coding methods, RFNs yield sparser codes, capture the data's covariance structure more precisely, and have a significantly smaller reconstruction error. We test RFNs as pretraining technique of deep networks on different vision datasets, where RFNs were superior to RBMs and autoencoders. On gene expression data from two pharmaceutical drug discovery studies, RFNs detected small and rare gene modules that revealed highly relevant new biological insights which were so far missed by other unsupervised methods.RFN package for GPU/CPU is available at http://www.bioinf.jku.at/software/rfn.
Djork-Arné Clevert, Thomas Unterthiner, Sepp Hochreiter
NIPS1
2010 FABIA: factor analysis for bicluster acquisition
abstract
MOTIVATION: Biclustering of transcriptomic data groups genes and samples simultaneously. It is emerging as a standard tool for extracting knowledge from gene expression measurements. We propose a novel generative approach for biclustering called 'FABIA: Factor Analysis for Bicluster Acquisition'. FABIA is based on a multiplicative model, which accounts for linear dependencies between gene expression and conditions, and also captures heavy-tailed distributions as observed in real-world transcriptomic data. The generative framework allows to utilize well-founded model selection methods and to apply Bayesian techniques. RESULTS: On 100 simulated datasets with known true, artificially implanted biclusters, FABIA clearly outperformed all 11 competitors. On these datasets, FABIA was able to separate spurious biclusters from true biclusters by ranking biclusters according to their information content. FABIA was tested on three microarray datasets with known subclusters, where it was two times the best and once the second best method among the compared biclustering approaches. AVAILABILITY: FABIA is available as an R package on Bioconductor (http://www.bioconductor.org). All datasets, results and software are available at http://www.bioinf.jku.at/software/fabia/fabia.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sepp Hochreiter, Ulrich Bodenhofer, Martin Heusel, Andreas Mitterecker, Adetayo Kasim, Tatsiana Khamiakova, Suzy Van Sanden, Dan Lin 0004, Willem Talloen, Luc Bijnens, Hinrich W. H. Göhlmann, Ziv Shkedy, Djork-Arné Clevert
Bioinform.14
2007 I/NI-calls for the exclusion of non-informative genes: a highly effective filtering tool for microarray data
abstract
MOTIVATION: DNA microarray technology typically generates many measurements of which only a relatively small subset is informative for the interpretation of the experiment. To avoid false positive results, it is therefore critical to select the informative genes from the large noisy data before the actual analysis. Most currently available filtering techniques are supervised and therefore suffer from a potential risk of overfitting. The unsupervised filtering techniques, on the other hand, are either not very efficient or too stringent as they may mix up signal with noise. We propose to use the multiple probes measuring the same target mRNA as repeated measures to quantify the signal-to-noise ratio of that specific probe set. A Bayesian factor analysis with specifically chosen prior settings, which models this probe level information, is providing an objective feature filtering technique, named informative/non-informative calls (I/NI calls). RESULTS: Based on 30 real-life data sets (including various human, rat, mice and Arabidopsis studies) and a spiked-in data set, it is shown that I/NI calls is highly effective, with exclusion rates ranging from 70% to 99%. Consequently, it offers a critical solution to the curse of high-dimensionality in the analysis of microarray data. AVAILABILITY: This filtering approach is publicly available as a function implemented in the R package FARMS (www.bioinf.jku.at/software/farms/farms.html).
Willem Talloen, Djork-Arné Clevert, Sepp Hochreiter, Dhammika Amaratunga, Luc Bijnens, Stefan Kass, Hinrich W. H. Göhlmann
Bioinform.2
2006 A new summarization method for affymetrix probe level data
abstract
MOTIVATION: We propose a new model-based technique for summarizing high-density oligonucleotide array data at probe level for Affymetrix GeneChips. The new summarization method is based on a factor analysis model for which a Bayesian maximum a posteriori method optimizes the model parameters under the assumption of Gaussian measurement noise. Thereafter, the RNA concentration is estimated from the model. In contrast to previous methods our new method called 'Factor Analysis for Robust Microarray Summarization (FARMS)' supplies both P-values indicating interesting information and signal intensity values. RESULTS: We compare FARMS on Affymetrix's spike-in and Gene Logic's dilution data to established algorithms like Affymetrix Microarray Suite (MAS) 5.0, Model Based Expression Index (MBEI), Robust Multi-array Average (RMA). Further, we compared FARMS with 43 other methods via the 'Affycomp II' competition. The experimental results show that FARMS with default parameters outperforms previous methods if both sensitivity and specificity are simultaneously considered by the area under the receiver operating curve (AUC). We measured two quantities through the AUC: correctly detected expression changes versus wrongly detected (fold change) and correctly detected significantly different expressed genes in two sets of arrays versus wrongly detected (P-value). Furthermore FARMS is computationally less expensive then RMA, MAS and MBEI. AVAILABILITY: The FARMS R package is available from http://www.bioinf.jku.at/software/farms/farms.html. SUPPLEMENTARY INFORMATION: http://www.bioinf.jku.at/publications/papers/farms/supplementary.ps
Sepp Hochreiter, Djork-Arné Clevert, Klaus Obermayer
Bioinform.2