EDBT 2026 Demo / reviewers in the wild / expert
Djork-Arné Clevert
dblp:03/1617
· DBLP profile ↗
18ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0003-4191-2156ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Generative modeling · 29% Representation and self-supervised learning · 19% Question answering and dialogue systems · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Bioinformatics and computational biology · 93% Medical and health informatics · 7% | |
| Databases, data mining, and information retrieval
1 paper |
Knowledge graphs · 100% |
Topics — the 30 heaviest of 41, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › domain-specific question answering
biomedical question answering |
0.9 | 1 | 2025 | KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA · ICLR 2025 |
Natural language and speech › Question answering and dialogue systems
knowledge-intensive question answering |
0.9 | 1 | 2025 | KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA · ICLR 2025 |
Knowledge graphs
knowledge graph reasoning |
0.9 | 1 | 2025 | KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA · ICLR 2025 |
Machine learning › Generative modeling › molecular generation
3d molecule generation |
0.8 | 1 | 2024 | Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule Generation · ICLR 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule Generation · ICLR 2024 |
Machine learning › Generative modeling › diffusion model › geometric diffusion model
equivariant diffusion model |
0.8 | 1 | 2024 | Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule Generation · ICLR 2024 |
Bioinformatics and computational biology
epigenomics |
0.8 | 1 | 2024 | A fast machine learning dataloader for epigenetic tracks from BigWig files · Bioinform. 2024 |
Bioinformatics and computational biology › genomics
machine learning for genomics |
0.8 | 1 | 2024 | A fast machine learning dataloader for epigenetic tracks from BigWig files · Bioinform. 2024 |
Machine learning › Representation and self-supervised learning › equivariance
equivariant representation learning |
0.6 | 1 | 2022 | Unsupervised Learning of Group Invariant and Equivariant Representations · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning |
0.6 | 1 | 2022 | Unsupervised Learning of Group Invariant and Equivariant Representations · NeurIPS 2022 |
Machine learning › Learning paradigms
unsupervised learning |
0.6 | 1 | 2022 | Unsupervised Learning of Group Invariant and Equivariant Representations · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › representation learning
unsupervised representation learning |
0.5 | 2 | 2017 | Rectified factor networks for biclustering of omics data · Bioinform. 2017 Rectified Factor Networks · NIPS 2015 |
Machine learning › Graph learning
graph generation |
0.5 | 1 | 2021 | Permutation-Invariant Variational Autoencoder for Graph-Level Representation Learning · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › interpretability
graph neural network explanation |
0.5 | 1 | 2021 | Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity · ICML 2021 |
Machine learning › Graph learning
graph representation learning |
0.5 | 1 | 2021 | Permutation-Invariant Variational Autoencoder for Graph-Level Representation Learning · NeurIPS 2021 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity · ICML 2021 |
Machine learning › Graph learning › molecular representation learning › molecular graph learning
molecular graph neural network |
0.5 | 1 | 2021 | Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced Sparsity · ICML 2021 |
Machine learning › Generative modeling
variational autoencoder |
0.5 | 1 | 2021 | Permutation-Invariant Variational Autoencoder for Graph-Level Representation Learning · NeurIPS 2021 |
Medical and health informatics › clinical data analysis
phenotyping |
0.5 | 1 | 2021 | Self-supervised feature extraction from image time series in plant phenotyping using triplet networks · Bioinform. 2021 |
Bioinformatics and computational biology › molecular property prediction
pka prediction |
0.5 | 1 | 2021 | pKPDB: a protein data bank extension database of pKa and pI theoretical values · Bioinform. 2021 |
Bioinformatics and computational biology › plant biology
plant phenotyping |
0.5 | 1 | 2021 | Self-supervised feature extraction from image time series in plant phenotyping using triplet networks · Bioinform. 2021 |
Bioinformatics and computational biology › molecular property prediction
protein property prediction |
0.5 | 1 | 2021 | pKPDB: a protein data bank extension database of pKa and pI theoretical values · Bioinform. 2021 |
Bioinformatics and computational biology
gene expression analysis |
0.5 | 3 | 2017 | Rectified factor networks for biclustering of omics data · Bioinform. 2017 FABIA: factor analysis for bicluster acquisition · Bioinform. 2010 I/NI-calls for the exclusion of non-informative genes: a highly effective filtering tool for microarray data · Bioinform. 2007 |
Bioinformatics and computational biology
drug discovery |
0.4 | 1 | 2020 | grünifai: interactive multiparameter optimization of molecules in a continuous vector space · Bioinform. 2020 |
Bioinformatics and computational biology › drug discovery
molecular optimization |
0.4 | 1 | 2020 | grünifai: interactive multiparameter optimization of molecules in a continuous vector space · Bioinform. 2020 |
Bioinformatics and computational biology › gene expression analysis
biclustering |
0.4 | 2 | 2017 | Rectified factor networks for biclustering of omics data · Bioinform. 2017 FABIA: factor analysis for bicluster acquisition · Bioinform. 2010 |
Bioinformatics and computational biology › genome editing
CRISPR guide RNA design |
0.4 | 1 | 2019 | PAVOOC: designing CRISPR sgRNAs using 3D protein structures and functional domain annotations · Bioinform. 2019 |
Bioinformatics and computational biology
genome editing |
0.4 | 1 | 2019 | PAVOOC: designing CRISPR sgRNAs using 3D protein structures and functional domain annotations · Bioinform. 2019 |
Bioinformatics and computational biology › molecular informatics
molecular design |
0.2 | 1 | 2024 | Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule Generation · ICLR 2024 |
Machine learning › Representation and self-supervised learning
pre-training |
0.2 | 1 | 2015 | Rectified Factor Networks · NIPS 2015 |
Methods — techniques the papers use, named apart from their topics
retrieval-augmented generation · 1.7large language model · 1.7time-dependent loss weighting · 1.5in silico modeling · 0.9parallel interval processing · 0.8e(3)-equivariant graph neural networks · 0.8e(3)-equivariant graph neural network · 0.8GPU decompression · 0.8group theory · 0.6encoder-decoder framework · 0.6posterior regularization · 0.5alternating minimization · 0.5triplet network · 0.5transfer learning · 0.5self-supervised learning · 0.5poisson-boltzmann calculation · 0.5monte carlo calculation · 0.5gini regularization · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KGARevion: An AI Agent for Knowledge-Intensive Biomedical QAabstractBiomedical reasoning integrates structured, codified knowledge with tacit, experience-driven insights. Depending on the context, quantity, and nature of available evidence, researchers and clinicians use diverse strategies, including rule-based, prototype-based, and case-based reasoning. Effective medical AI models must handle this complexity while ensuring reliability and adaptability. We introduce KGARevion, a knowledge graph-based agent that answers knowledge-intensive questions. Upon receiving a query, KGARevion generates relevant triplets by leveraging the latent knowledge embedded in a large language model. It then verifies these triplets against a grounded knowledge graph, filtering out errors and retaining only accurate, contextually relevant information for the final answer. This multi-step process strengthens reasoning, adapts to different models of medical inference, and outperforms retrieval-augmented generation-based approaches that lack effective verification mechanisms. Evaluations on medical QA benchmarks show that KGARevion improves accuracy by over 5.2% over 15 models in handling complex medical queries. To further assess its effectiveness, we curated three new medical QA datasets with varying levels of semantic complexity, where KGARevion improved accuracy by 10.4%. The agent integrates with different LLMs and biomedical knowledge graphs for broad applicability across knowledge-intensive tasks. We evaluated KGARevion on AfriMed-QA, a newly introduced dataset focused on African healthcare, demonstrating its strong zero-shot generalization to underrepresented medical contexts. Xiao-Rui Su 0001, Yibo Wang 0001, Shanghua Gao, Xiaolong Liu 0012, Valentina Giunchiglia, Djork-Arné Clevert, Marinka Zitnik |
ICLR | 6 |
| 2025 | Diffusion Generative Modeling on Lie Group RepresentationsabstractWe introduce a novel class of score-based diffusion processes that operate directly in the representation space of Lie groups.
Leveraging the framework of Generalized Score Matching, we derive a class of Langevin dynamics that decomposes as a direct sum of Lie algebra representations,
enabling the modeling of any target distribution on any (non-Abelian) Lie group.
Standard score-matching emerges as a special case of our framework when the Lie group is the translation group.
We prove that our generalized generative processes arise as solutions to a new class of paired stochastic differential equations (SDEs), introduced here for the first time.
We validate our approach through experiments on diverse data types, demonstrating its effectiveness in real-world applications such as $\text{SO}(3)$-guided molecular conformer generation and modeling ligand-specific global $\text{SE}(3)$ transformations for molecular docking, showing improvement in comparison to Riemannian diffusion on the group itself.
We show that an appropriate choice of Lie group enhances learning efficiency by reducing the effective dimensionality of the trajectory space and enables the modeling of transitions between complex data distributions. Marco Bertolini, Djork-Arné Clevert |
NeurIPS | 3 |
| 2024 | Navigating the Design Space of Equivariant Diffusion-Based Generative Models for De Novo 3D Molecule GenerationabstractDeep generative diffusion models are a promising avenue for 3D de novo molecular design in materials science and drug discovery.
However, their utility is still limited by suboptimal performance on large molecular structures and limited training data.
To address this gap, we explore the design space of E(3)-equivariant diffusion models, focusing on previously unexplored areas.
Our extensive comparative analysis evaluates the interplay between continuous and discrete state spaces.
From this investigation, we present the EQGAT-diff model, which consistently outperforms established models for the QM9 and GEOM-Drugs datasets.
Significantly, EQGAT-diff takes continuous atom positions, while chemical elements and bond types are categorical and uses time-dependent loss weighting, substantially increasing training convergence, the quality of generated samples, and inference time. We also showcase that including chemically motivated additional features like hybridization states in the diffusion process enhances the validity of generated molecules.
To further strengthen the applicability of diffusion models to limited training data, we investigate the transferability of EQGAT-diff trained on the large PubChem3D dataset with implicit hydrogen atoms to target different data distributions. Fine-tuning EQGAT-diff for just a few iterations shows an efficient distribution shift, further improving performance throughout data sets.
Finally, we test our model on the Crossdocked data set for structure-based de novo ligand generation, underlining the importance of our findings showing state-of-the-art performance on Vina docking scores. Tuan Le, Julian Cremer, Frank Noé, Djork-Arné Clevert, Kristof Schütt |
ICLR | 4 |
| 2024 | A fast machine learning dataloader for epigenetic tracks from BigWig filesabstractSUMMARY: We created bigwig-loader, a data-loader for epigenetic profiles from BigWig files that decompresses and processes information for multiple intervals from multiple BigWig files in parallel. This is an access pattern needed to create training batches for typical machine learning models on epigenetics data. Using a new codec, the decompression can be done on a graphical processing unit (GPU) making it fast enough to create the training batches during training, mitigating the need for saving preprocessed training examples to disk. AVAILABILITY AND IMPLEMENTATION: The bigwig-loader installation instructions and source code can be accessed at https://github.com/pfizer-opensource/bigwig-loader. Joren Sebastian Retel, Andreas Poehlmann, Josh Chiou, Andreas Steffen, Djork-Arné Clevert |
Bioinform. | 5 |
| 2023 | Explaining, Evaluating and Enhancing Neural Networks' Learned RepresentationsabstractMost efforts in interpretability in deep learning have focused on (1) extracting explanations of a specific downstream task in relation to the input features and (2) imposing constraints on the model, often at the expense of predictive performance. New advances in (unsupervised) representation learning and transfer learning, however, raise the need for an explanatory framework for networks that are trained without a specific downstream task. We address these challenges by showing how explainability can be an aid, rather than an obstacle, towards better and more efficient representations. Specifically, we propose a natural aggregation method generalizing attribution maps between any two (convolutional) layers of a neural network. Additionally, we employ such attributions to define two novel scores for evaluating the informativeness and the disentanglement of latent embeddings. Extensive experiments show that the proposed scores do correlate with the desired properties. We also confirm and extend previously known results concerning the independence of some common saliency strategies from the model parameters. Finally, we show that adopting our proposed scores as constraints during the training of a representation learning task improves the downstream performance of the model. Marco Bertolini, Djork-Arné Clevert, Floriane Montanari |
ICANN (5) | 2 |
| 2022 | Unsupervised Learning of Group Invariant and Equivariant RepresentationsabstractEquivariant neural networks, whose hidden features transform according to representations of a group $G$ acting on the data, exhibit training efficiency and an improved generalisation performance. In this work, we extend group invariant and equivariant representation learning to the field of unsupervised deep learning. We propose a general learning strategy based on an encoder-decoder framework in which the latent representation is separated in an invariant term and an equivariant group action component. The key idea is that the network learns to encode and decode data to and from a group-invariant representation by additionally learning to predict the appropriate group action to align input and output pose to solve the reconstruction task. We derive the necessary conditions on the equivariant encoder, and we present a construction valid for any $G$, both discrete and continuous. We describe explicitly our construction for rotations, translations and permutations. We test the validity and the robustness of our approach in a variety of experiments with diverse data types employing different network architectures. Robin Winter, Marco Bertolini, Tuan Le, Frank Noé, Djork-Arné Clevert |
NeurIPS | 5 |
| 2021 | Parameterized Hypercomplex Graph Neural Networks for Graph Classification
Tuan Le, Marco Bertolini, Frank Noé, Djork-Arné Clevert |
ICANN (3) | 4 |
| 2021 | Improving Molecular Graph Neural Network Explainability with Orthonormalization and Induced SparsityabstractRationalizing which parts of a molecule drive the predictions of a molecular graph convolutional neural network (GCNN) can be difficult. To help, we propose two simple regularization techniques to apply during the training of GCNNs: Batch Representation Orthonormalization (BRO) and Gini regularization. BRO, inspired by molecular orbital theory, encourages graph convolution operations to generate orthonormal node embeddings. Gini regularization is applied to the weights of the output layer and constrains the number of dimensions the model can use to make predictions. We show that Gini and BRO regularization can improve the accuracy of state-of-the-art GCNN attribution methods on artificial benchmark datasets. In a real-world setting, we demonstrate that medicinal chemists significantly prefer explanations extracted from regularized models. While we only study these regularizers in the context of GCNNs, both can be applied to other types of neural networks. Ryan Henderson, Djork-Arné Clevert, Floriane Montanari |
ICML | 2 |
| 2021 | Permutation-Invariant Variational Autoencoder for Graph-Level Representation LearningabstractRecently, there has been great success in applying deep neural networks on graph structured data. Most work, however, focuses on either node- or graph-level supervised learning, such as node, link or graph classification or node-level unsupervised learning (e.g. node clustering). Despite its wide range of possible applications, graph-level unsupervised learning has not received much attention yet. This might be mainly attributed to the high representation complexity of graphs, which can be represented by $n!$ equivalent adjacency matrices, where $n$ is the number of nodes.In this work we address this issue by proposing a permutation-invariant variational autoencoder for graph structured data. Our proposed model indirectly learns to match the node ordering of input and output graph, without imposing a particular node ordering or performing expensive graph matching. We demonstrate the effectiveness of our proposed model for graph reconstruction, generation and interpolation and evaluate the expressive power of extracted representations for downstream graph-level classification and regression. Robin Winter, Frank Noé, Djork-Arné Clevert |
NeurIPS | 3 |
| 2021 | pKPDB: a protein data bank extension database of pKa and pI theoretical valuesabstractSUMMARY: pKa values of ionizable residues and isoelectric points of proteins provide valuable local and global insights about their structure and function. These properties can be estimated with reasonably good accuracy using Poisson-Boltzmann and Monte Carlo calculations at a considerable computational cost (from some minutes to several hours). pKPDB is a database of over 12 M theoretical pKa values calculated over 120k protein structures deposited in the Protein Data Bank. By providing precomputed pKa and pI values, users can retrieve results instantaneously for their protein(s) of interest while also saving countless hours and resources that would be spent on repeated calculations. Furthermore, there is an ever-growing imbalance between experimental pKa and pI values and the number of resolved structures. This database will complement the experimental and computational data already available and can also provide crucial information regarding buried residues that are under-represented in experimental measurements. AVAILABILITY AND IMPLEMENTATION: Gzipped csv files containing p Ka and isoelectric point values can be downloaded from https://pypka.org/pKPDB. To query a single PDB code please use the PypKa free server at https://pypka.org. The pKPDB source code can be found at https://github.com/mms-fcul/pKPDB. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pedro B. P. S. Reis, Djork-Arné Clevert, Miguel Machuqueiro |
Bioinform. | 2 |
| 2021 | Self-supervised feature extraction from image time series in plant phenotyping using triplet networksabstractMOTIVATION: Image-based profiling combines high-throughput screening with multiparametric feature analysis to capture the effect of perturbations on biological systems. This technology has attracted increasing interest in the field of plant phenotyping, promising to accelerate the discovery of novel herbicides. However, the extraction of meaningful features from unlabeled plant images remains a big challenge. RESULTS: We describe a novel data-driven approach to find feature representations from plant time-series images in a self-supervised manner by using time as a proxy for image similarity. In the spirit of transfer learning, we first apply an ImageNet-pretrained architecture as a base feature extractor. Then, we extend this architecture with a triplet network to refine and reduce the dimensionality of extracted features by ranking relative similarities between consecutive and non-consecutive time points. Without using any labels, we produce compact, organized representations of plant phenotypes and demonstrate their superior applicability to clustering, image retrieval and classification tasks. Besides time, our approach could be applied using other surrogate measures of phenotype similarity, thus providing a versatile method of general interest to the phenotypic profiling community. AVAILABILITY AND IMPLEMENTATION: Source code is provided in https://github.com/bayer-science-for-a-better-life/plant-triplet-net. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Paula A. Marin Zapata, Sina Roth, Dirk Schmutzler, Erica Manesso, Djork-Arné Clevert |
Bioinform. | 6 |
| 2020 | grünifai: interactive multiparameter optimization of molecules in a continuous vector spaceabstractSUMMARY: Optimizing small molecules in a drug discovery project is a notoriously difficult task as multiple molecular properties have to be considered and balanced at the same time. In this work, we present our novel interactive in silico compound optimization platform termed grünifai to support the ideation of the next generation of compounds under the constraints of a multiparameter objective. grünifai integrates adjustable in silico models, a continuous representation of the chemical space, a scalable particle swarm optimization algorithm and the possibility to actively steer the compound optimization through providing feedback on generated intermediate structures. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available under an MIT license and are openly available on GitHub (https://github.com/jrwnter/gruenifai). The backend, including the optimization method and distribution on multiple GPU nodes is written in Python 3. The frontend is written in ReactJS. Robin Winter, Joren Sebastian Retel, Frank Noé, Djork-Arné Clevert, Andreas Steffen |
Bioinform. | 4 |
| 2019 | PAVOOC: designing CRISPR sgRNAs using 3D protein structures and functional domain annotationsabstractSUMMARY: Single-guide RNAs (sgRNAs) targeting the same gene can significantly vary in terms of efficacy and specificity. PAVOOC (Prediction And Visualization of On- and Off-targets for CRISPR) is a web-based CRISPR sgRNA design tool that employs state of the art machine learning models to prioritize most effective candidate sgRNAs. In contrast to other tools, it maps sgRNAs to functional domains and protein structures and visualizes cut sites on corresponding protein crystal structures. Furthermore, PAVOOC supports homology-directed repair template generation for genome editing experiments and the visualization of the mutated amino acids in 3D. AVAILABILITY AND IMPLEMENTATION: PAVOOC is available under https://pavooc.me and accessible using modern browsers (Chrome/Chromium recommended). The source code is hosted at github.com/moritzschaefer/pavooc under the MIT License. The backend, including data processing steps, and the frontend are implemented in Python 3 and ReactJS, respectively. All components run in a simple Docker environment. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Moritz Schaefer, Djork-Arné Clevert, Bertram Weiss, Andreas Steffen |
Bioinform. | 2 |
| 2017 | Rectified factor networks for biclustering of omics dataabstractMOTIVATION: Biclustering has become a major tool for analyzing large datasets given as matrix of samples times features and has been successfully applied in life sciences and e-commerce for drug design and recommender systems, respectively. actor nalysis for cluster cquisition (FABIA), one of the most successful biclustering methods, is a generative model that represents each bicluster by two sparse membership vectors: one for the samples and one for the features. However, FABIA is restricted to about 20 code units because of the high computational complexity of computing the posterior. Furthermore, code units are sometimes insufficiently decorrelated and sample membership is difficult to determine. We propose to use the recently introduced unsupervised Deep Learning approach Rectified Factor Networks (RFNs) to overcome the drawbacks of existing biclustering methods. RFNs efficiently construct very sparse, non-linear, high-dimensional representations of the input via their posterior means. RFN learning is a generalized alternating minimization algorithm based on the posterior regularization method which enforces non-negative and normalized posterior means. Each code unit represents a bicluster, where samples for which the code unit is active belong to the bicluster and features that have activating weights to the code unit belong to the bicluster. RESULTS: On 400 benchmark datasets and on three gene expression datasets with known clusters, RFN outperformed 13 other biclustering methods including FABIA. On data of the 1000 Genomes Project, RFN could identify DNA segments which indicate, that interbreeding with other hominins starting already before ancestors of modern humans left Africa. AVAILABILITY AND IMPLEMENTATION: https://github.com/bioinf-jku/librfn. CONTACT: [email protected] or [email protected]. Djork-Arné Clevert, Thomas Unterthiner, Gundula Povysil, Sepp Hochreiter |
Bioinform. | 1 |
| 2015 | Rectified Factor NetworksabstractWe propose rectified factor networks (RFNs) to efficiently construct very sparse, non-linear, high-dimensional representations of the input. RFN models identify rare and small events, have a low interference between code units, have a small reconstruction error, and explain the data covariance structure. RFN learning is a generalized alternating minimization algorithm derived from the posterior regularization method which enforces non-negative and normalized posterior means. We proof convergence and correctness of the RFN learning algorithm.On benchmarks, RFNs are compared to other unsupervised methods like autoencoders, RBMs, factor analysis, ICA, and PCA. In contrast to previous sparse coding methods, RFNs yield sparser codes, capture the data's covariance structure more precisely, and have a significantly smaller reconstruction error. We test RFNs as pretraining technique of deep networks on different vision datasets, where RFNs were superior to RBMs and autoencoders. On gene expression data from two pharmaceutical drug discovery studies, RFNs detected small and rare gene modules that revealed highly relevant new biological insights which were so far missed by other unsupervised methods.RFN package for GPU/CPU is available at http://www.bioinf.jku.at/software/rfn. Djork-Arné Clevert, Thomas Unterthiner, Sepp Hochreiter |
NIPS | 1 |
| 2010 | FABIA: factor analysis for bicluster acquisitionabstractMOTIVATION: Biclustering of transcriptomic data groups genes and samples simultaneously. It is emerging as a standard tool for extracting knowledge from gene expression measurements. We propose a novel generative approach for biclustering called 'FABIA: Factor Analysis for Bicluster Acquisition'. FABIA is based on a multiplicative model, which accounts for linear dependencies between gene expression and conditions, and also captures heavy-tailed distributions as observed in real-world transcriptomic data. The generative framework allows to utilize well-founded model selection methods and to apply Bayesian techniques. RESULTS: On 100 simulated datasets with known true, artificially implanted biclusters, FABIA clearly outperformed all 11 competitors. On these datasets, FABIA was able to separate spurious biclusters from true biclusters by ranking biclusters according to their information content. FABIA was tested on three microarray datasets with known subclusters, where it was two times the best and once the second best method among the compared biclustering approaches. AVAILABILITY: FABIA is available as an R package on Bioconductor (http://www.bioconductor.org). All datasets, results and software are available at http://www.bioinf.jku.at/software/fabia/fabia.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Sepp Hochreiter, Ulrich Bodenhofer, Martin Heusel, Andreas Mitterecker, Adetayo Kasim, Tatsiana Khamiakova, Suzy Van Sanden, Dan Lin 0004, Willem Talloen, Luc Bijnens, Hinrich W. H. Göhlmann, Ziv Shkedy, Djork-Arné Clevert |
Bioinform. | 14 |
| 2007 | I/NI-calls for the exclusion of non-informative genes: a highly effective filtering tool for microarray dataabstractMOTIVATION: DNA microarray technology typically generates many measurements of which only a relatively small subset is informative for the interpretation of the experiment. To avoid false positive results, it is therefore critical to select the informative genes from the large noisy data before the actual analysis. Most currently available filtering techniques are supervised and therefore suffer from a potential risk of overfitting. The unsupervised filtering techniques, on the other hand, are either not very efficient or too stringent as they may mix up signal with noise. We propose to use the multiple probes measuring the same target mRNA as repeated measures to quantify the signal-to-noise ratio of that specific probe set. A Bayesian factor analysis with specifically chosen prior settings, which models this probe level information, is providing an objective feature filtering technique, named informative/non-informative calls (I/NI calls). RESULTS: Based on 30 real-life data sets (including various human, rat, mice and Arabidopsis studies) and a spiked-in data set, it is shown that I/NI calls is highly effective, with exclusion rates ranging from 70% to 99%. Consequently, it offers a critical solution to the curse of high-dimensionality in the analysis of microarray data. AVAILABILITY: This filtering approach is publicly available as a function implemented in the R package FARMS (www.bioinf.jku.at/software/farms/farms.html). Willem Talloen, Djork-Arné Clevert, Sepp Hochreiter, Dhammika Amaratunga, Luc Bijnens, Stefan Kass, Hinrich W. H. Göhlmann |
Bioinform. | 2 |
| 2006 | A new summarization method for affymetrix probe level dataabstractMOTIVATION: We propose a new model-based technique for summarizing high-density oligonucleotide array data at probe level for Affymetrix GeneChips. The new summarization method is based on a factor analysis model for which a Bayesian maximum a posteriori method optimizes the model parameters under the assumption of Gaussian measurement noise. Thereafter, the RNA concentration is estimated from the model. In contrast to previous methods our new method called 'Factor Analysis for Robust Microarray Summarization (FARMS)' supplies both P-values indicating interesting information and signal intensity values. RESULTS: We compare FARMS on Affymetrix's spike-in and Gene Logic's dilution data to established algorithms like Affymetrix Microarray Suite (MAS) 5.0, Model Based Expression Index (MBEI), Robust Multi-array Average (RMA). Further, we compared FARMS with 43 other methods via the 'Affycomp II' competition. The experimental results show that FARMS with default parameters outperforms previous methods if both sensitivity and specificity are simultaneously considered by the area under the receiver operating curve (AUC). We measured two quantities through the AUC: correctly detected expression changes versus wrongly detected (fold change) and correctly detected significantly different expressed genes in two sets of arrays versus wrongly detected (P-value). Furthermore FARMS is computationally less expensive then RMA, MAS and MBEI. AVAILABILITY: The FARMS R package is available from http://www.bioinf.jku.at/software/farms/farms.html. SUPPLEMENTARY INFORMATION: http://www.bioinf.jku.at/publications/papers/farms/supplementary.ps Sepp Hochreiter, Djork-Arné Clevert, Klaus Obermayer |
Bioinform. | 2 |