VLDB 2026 Research / reviewers in the wild / expert
Michael Habeck
dblp:67/5618
· DBLP profile ↗
17ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-2188-5667ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 1 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Probabilistic and Bayesian machine learning · 88% Knowledge representation and reasoning · 12% | |
| Interdisciplinary, comprehensive, and emerging computing
8 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 19 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo |
2.4 | 4 | 2025 | Geodesic Slice Sampling on the Sphere · J. Mach. Learn. Res. 2025 Parallel Affine Transformation Tuning of Markov Chain Monte Carlo · ICML 2024 Gibbsian Polar Slice Sampling · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
slice sampling |
1.5 | 2 | 2025 | Geodesic Slice Sampling on the Sphere · J. Mach. Learn. Res. 2025 Gibbsian Polar Slice Sampling · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
adaptive MCMC |
0.8 | 1 | 2024 | Parallel Affine Transformation Tuning of Markov Chain Monte Carlo · ICML 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
symbolic regression |
0.8 | 1 | 2024 | Scaling Up Unbiased Search-based Symbolic Regression · IJCAI 2024 |
Bioinformatics and computational biology
structural biology |
0.3 | 4 | 2010 | A New Algorithm for Improving the Resolution of Cryo-EM Density Maps · RECOMB 2010 ISD: a software package for Bayesian NMR structure calculation · Bioinform. 2008 ARIA2: Automated NOE assignment and data integration in NMR structure calculation · Bioinform. 2007 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › importance sampling
annealed importance sampling |
0.3 | 1 | 2017 | Model evidence from nonequilibrium simulations · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian model selection
marginal likelihood estimation |
0.3 | 1 | 2017 | Model evidence from nonequilibrium simulations · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning
directional statistics |
0.3 | 1 | 2025 | Geodesic Slice Sampling on the Sphere · J. Mach. Learn. Res. 2025 |
Bioinformatics and computational biology
protein structure analysis |
0.2 | 1 | 2016 | A probabilistic model for detecting rigid domains in protein structures · Bioinform. 2016 |
Bioinformatics and computational biology › structural biology
NMR structure calculation |
0.2 | 3 | 2008 | ISD: a software package for Bayesian NMR structure calculation · Bioinform. 2008 ARIA2: Automated NOE assignment and data integration in NMR structure calculation · Bioinform. 2007 ARIA: automated NOE assignment and NMR structure calculation · Bioinform. 2003 |
Bioinformatics and computational biology › structural bioinformatics
molecular structure analysis |
0.1 | 1 | 2012 | CSB: a Python framework for structural bioinformatics · Bioinform. 2012 |
Bioinformatics and computational biology
structural bioinformatics |
0.1 | 1 | 2012 | CSB: a Python framework for structural bioinformatics · Bioinform. 2012 |
Bioinformatics and computational biology
protein structure prediction |
0.1 | 1 | 2011 | HHfrag: HMM-based fragment detection using HHpred · Bioinform. 2011 |
Bioinformatics and computational biology › protein structure analysis › protein flexibility
conformational diversity |
0.1 | 1 | 2008 | Mixture models for protein structure ensembles · Bioinform. 2008 |
Bioinformatics and computational biology › structural bioinformatics
protein structure |
0.1 | 1 | 2008 | Mixture models for protein structure ensembles · Bioinform. 2008 |
Bioinformatics and computational biology › structural bioinformatics › protein conformational analysis
protein structure ensemble analysis |
0.1 | 1 | 2008 | Mixture models for protein structure ensembles · Bioinform. 2008 |
Bioinformatics and computational biology › protein dynamics
conformational change |
0.1 | 1 | 2016 | A probabilistic model for detecting rigid domains in protein structures · Bioinform. 2016 |
Bioinformatics and computational biology › structural bioinformatics
protein structure determination |
0.0 | 1 | 2003 | ARIA: automated NOE assignment and NMR structure calculation · Bioinform. 2003 |
Bioinformatics and computational biology › genomics
structural genomics |
0.0 | 1 | 2003 | ARIA: automated NOE assignment and NMR structure calculation · Bioinform. 2003 |
Methods — techniques the papers use, named apart from their topics
markov chain monte carlo · 1.8unbiased search · 1.5genetic programming · 1.5uniform ergodicity · 0.9shrinkage · 0.9geodesic slice sampling · 0.9gibbsian polar slice sampling · 0.8bayesian inference · 0.3reverse AIS · 0.3forward-backward simulation · 0.3annealed importance sampling · 0.3gibbs sampling · 0.2profile-profile comparison · 0.1hidden markov model · 0.1ambiguous restraints iterative assignment · 0.1gaussian mixture model · 0.1expectation-maximization · 0.1CCPN data model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Geodesic Slice Sampling on the SphereabstractProbability measures on the sphere form an important class of statistical models and are used, for example, in modeling directional data or shapes. Due to their widespread use, but also as an algorithmic building block, efficient sampling of distributions on the sphere is highly desirable. We propose a shrinkage based and an idealized geodesic slice sampling Markov chain, designed to generate approximate samples from distributions on the sphere. In particular, the shrinkage-based version of the algorithm can be implemented such that it runs efficiently and has no tuning parameters. We verify reversibility and prove that under weak regularity conditions geodesic slice sampling is uniformly ergodic. Numerical experiments show that the proposed slice samplers achieve excellent mixing on challenging targets including distributions arising in rigid-registration problems and mixtures of von Mises-Fisher distributions. In these settings our approach outperforms standard samplers such as random-walk Metropolis-Hastings and Hamiltonian Monte Carlo. Michael Habeck, Mareike Hasenpflug, Shantanu Kodgirwar, Daniel Rudolf |
J. Mach. Learn. Res. | 1 |
| 2024 | Parallel Affine Transformation Tuning of Markov Chain Monte CarloabstractThe performance of Markov chain Monte Carlo samplers strongly depends on the properties of the target distribution such as its covariance structure, the location of its probability mass and its tail behavior. We explore the use of bijective affine transformations of the sample space to improve the properties of the target distribution and thereby the performance of samplers running in the transformed space. In particular, we propose a flexible and user-friendly scheme for adaptively learning the affine transformation during sampling. Moreover, the combination of our scheme with Gibbsian polar slice sampling is shown to produce samples of high quality at comparatively low computational cost in several settings based on real-world data. Philip Schär, Michael Habeck, Daniel Rudolf |
ICML | 2 |
| 2024 | Scaling Up Unbiased Search-based Symbolic Regression
Paul Kahlmeyer, Joachim Giesen, Michael Habeck, Henrik Voigt |
IJCAI | 3 |
| 2023 | Gibbsian Polar Slice SamplingabstractPolar slice sampling (Roberts & Rosenthal, 2002) is a Markov chain approach for approximate sampling of distributions that is difficult, if not impossible, to implement efficiently, but behaves provably well with respect to the dimension. By updating the directional and radial components of chain iterates separately, we obtain a family of samplers that mimic polar slice sampling, and yet can be implemented efficiently. Numerical experiments in a variety of settings indicate that our proposed algorithm outperforms the two most closely related approaches, elliptical slice sampling (Murray et al., 2010) and hit-and-run uniform slice sampling (MacKay, 2003). We prove the well-definedness and convergence of our methods under suitable assumptions on the target distribution. Philip Schär, Michael Habeck, Daniel Rudolf |
ICML | 2 |
| 2021 | A graph-based algorithm for detecting rigid domains in protein structuresabstractBACKGROUND: Conformational transitions are implicated in the biological function of many proteins. Structural changes in proteins can be described approximately as the relative movement of rigid domains against each other. Despite previous efforts, there is a need to develop new domain segmentation algorithms that are capable of analysing the entire structure database efficiently and do not require the choice of protein-dependent tuning parameters such as the number of rigid domains. RESULTS: We develop a graph-based method for detecting rigid domains in proteins. Structural information from multiple conformational states is represented by a graph whose nodes correspond to amino acids. Graph clustering algorithms allow us to reduce the graph and run the Viterbi algorithm on the associated line graph to obtain a segmentation of the input structures into rigid domains. In contrast to many alternative methods, our approach does not require knowledge about the number of rigid domains. Moreover, we identified default values for the algorithmic parameters that are suitable for a large number of conformational ensembles. We test our algorithm on examples from the DynDom database and illustrate our method on various challenging systems whose structural transitions have been studied extensively. CONCLUSIONS: The results strongly suggest that our graph-based algorithm forms a novel framework to characterize structural transitions in proteins via detecting their rigid domains. The web server is available at http://azifi.tz.agrar.uni-goettingen.de/webservice/ . Truong Khanh Linh Dang, Thach Nguyen, Michael Habeck, Mehmet Gültas, Stephan Waack |
BMC Bioinform. | 3 |
| 2017 | Model evidence from nonequilibrium simulationsabstractThe marginal likelihood, or model evidence, is a key quantity in Bayesian parameter estimation and model comparison. For many probabilistic models, computation of the marginal likelihood is challenging, because it involves a sum or integral over an enormous parameter space. Markov chain Monte Carlo (MCMC) is a powerful approach to compute marginal likelihoods. Various MCMC algorithms and evidence estimators have been proposed in the literature. Here we discuss the use of nonequilibrium techniques for estimating the marginal likelihood. Nonequilibrium estimators build on recent developments in statistical physics and are known as annealed importance sampling (AIS) and reverse AIS in probabilistic machine learning. We introduce estimators for the model evidence that combine forward and backward simulations and show for various challenging models that the evidence estimators outperform forward and reverse AIS. Michael Habeck |
NIPS | 1 |
| 2016 | A probabilistic model for detecting rigid domains in protein structuresabstractMOTIVATION: Large-scale conformational changes in proteins are implicated in many important biological functions. These structural transitions can often be rationalized in terms of relative movements of rigid domains. There is a need for objective and automated methods that identify rigid domains in sets of protein structures showing alternative conformational states. RESULTS: We present a probabilistic model for detecting rigid-body movements in protein structures. Our model aims to approximate alternative conformational states by a few structural parts that are rigidly transformed under the action of a rotation and a translation. By using Bayesian inference and Markov chain Monte Carlo sampling, we estimate all parameters of the model, including a segmentation of the protein into rigid domains, the structures of the domains themselves, and the rigid transformations that generate the observed structures. We find that our Gibbs sampling algorithm can also estimate the optimal number of rigid domains with high efficiency and accuracy. We assess the power of our method on several thousand entries of the DynDom database and discuss applications to various complex biomolecular systems. AVAILABILITY AND IMPLEMENTATION: The Python source code for protein ensemble analysis is available at: https://github.com/thachnguyen/motion_detection CONTACT: : [email protected]. Thach Nguyen, Michael Habeck |
Bioinform. | 2 |
| 2016 | Inferential Structure Determination of Chromosomes from Single-Cell Hi-C DataabstractChromosome conformation capture (3C) techniques have revealed many fascinating insights into the spatial organization of genomes. 3C methods typically provide information about chromosomal contacts in a large population of cells, which makes it difficult to draw conclusions about the three-dimensional organization of genomes in individual cells. Recently it became possible to study single cells with Hi-C, a genome-wide 3C variant, demonstrating a high cell-to-cell variability of genome organization. In principle, restraint-based modeling should allow us to infer the 3D structure of chromosomes from single-cell contact data, but suffers from the sparsity and low resolution of chromosomal contacts. To address these challenges, we adapt the Bayesian Inferential Structure Determination (ISD) framework, originally developed for NMR structure determination of proteins, to infer statistical ensembles of chromosome structures from single-cell data. Using ISD, we are able to compute structural error bars and estimate model parameters, thereby eliminating potential bias imposed by ad hoc parameter choices. We apply and compare different models for representing the chromatin fiber and for incorporating singe-cell contact information. Finally, we extend our approach to the analysis of diploid chromosome data. Simeon Carstens, Michael Nilges, Michael Habeck |
PLoS Comput. Biol. | 3 |
| 2012 | A blind deconvolution approach for pseudo CT prediction from MR image pairsabstractPredicting a CT image or a map of the linear attenuation coefficients from the information provided by magnetic resonance imaging (MRI) is a challenging task. This problem is of significant importance for combined positron emission tomography (PET)/MRI scanners, as quantitative PET image reconstruction requires an attenuation map. In PET/CT this attenuation map is derived from the CT scan or from a rotating source, however, current PET/MR systems can not directly measure attenuation images - and indeed it is desirable to save the patient from the additional radiation exposure. Recent approaches tackle this problem by using MR sequences with ultra-short echo times (UTE). At the price of lower effective image resolution, the UTE image yields signal from bone and therefore provides valuable information for calculating the attenuation map. We propose a novel approach to this problem based on nonnegative blind deconvolution and present the first method that explicitly models the image degradation of the UTE image. Incorporating prior knowledge such as smoothness and a novel orthogonality constraint alleviates the deconvolution process. Due to its probabilistic formulation our approach allows hyperparameter estimation and is therefore parameter-free. Michael Hirsch 0001, Matthias Hofmann, Frederic Mantlik, Bernd J. Pichler, Bernhard Schölkopf, Michael Habeck |
ICIP | 6 |
| 2012 | CSB: a Python framework for structural bioinformaticsabstractSUMMARY: Computational Structural Biology Toolbox (CSB) is a cross-platform Python class library for reading, storing and analyzing biomolecular structures with rich support for statistical analyses. CSB is designed for reusability and extensibility and comes with a clean, well-documented API following good object-oriented engineering practice. AVAILABILITY: Stable release packages are available for download from the Python Package Index (PyPI) as well as from the project's website http://csb.codeplex.com. CONTACTS: [email protected] or [email protected] Ivan Kalev, Martin Mechelke, Klaus O. Kopec, Thomas Holder, Simeon Carstens, Michael Habeck |
Bioinform. | 6 |
| 2011 | HHfrag: HMM-based fragment detection using HHpredabstractAbstract Motivation: Over the last decade, both static and dynamic fragment libraries for protein structure prediction have been introduced. The former are built from clusters in either sequence or structure space and aim to extract a universal structural alphabet. The latter are tailored for a particular query protein sequence and aim to provide local structural templates that need to be assembled in order to build the full-length structure. Results: Here, we introduce HHfrag, a dynamic HMM-based fragment search method built on the profile–profile comparison tool HHpred. We show that HHfrag provides advantages over existing fragment assignment methods in that it: (i) improves the precision of the fragments at the expense of a minor loss in sequence coverage; (ii) detects fragments of variable length (6–21 amino acid residues); (iii) allows for gapped fragments and (iv) does not assign fragments to regions where there is no clear sequence conservation. We illustrate the usefulness of fragments detected by HHfrag on targets from most recent CASP. Availability: A web server for running HHfrag is available at http://toolkit.tuebingen.mpg.de/hhfrag. The source code is available at http://www.eb.tuebingen.mpg.de/departments/1-protein-evolution/michael-habeck/HHfrag.tar.gz Contact: [email protected] Supplementary Information: Supplementary data are available at Bioinformatics online. Ivan Kalev, Michael Habeck |
Bioinform. | 2 |
| 2010 | A New Algorithm for Improving the Resolution of Cryo-EM Density Maps
Michael Hirsch 0001, Bernhard Schölkopf, Michael Habeck |
RECOMB | 3 |
| 2010 | Robust probabilistic superposition and comparison of protein structuresabstractBACKGROUND: Protein structure comparison is a central issue in structural bioinformatics. The standard dissimilarity measure for protein structures is the root mean square deviation (RMSD) of representative atom positions such as alpha-carbons. To evaluate the RMSD the structures under comparison must be superimposed optimally so as to minimize the RMSD. How to evaluate optimal fits becomes a matter of debate, if the structures contain regions which differ largely--a situation encountered in NMR ensembles and proteins undergoing large-scale conformational transitions. RESULTS: We present a probabilistic method for robust superposition and comparison of protein structures. Our method aims to identify the largest structurally invariant core. To do so, we model non-rigid displacements in protein structures with outlier-tolerant probability distributions. These distributions exhibit heavier tails than the Gaussian distribution underlying standard RMSD minimization and thus accommodate highly divergent structural regions. The drawback is that under a heavy-tailed model analytical expressions for the optimal superposition no longer exist. To circumvent this problem we work with a scale mixture representation, which implies a weighted RMSD. We develop two iterative procedures, an Expectation Maximization algorithm and a Gibbs sampler, to estimate the local weights, the optimal superposition, and the parameters of the heavy-tailed distribution. Applications demonstrate that heavy-tailed models capture differences between structures undergoing substantial conformational changes and can be used to assess the precision of NMR structures. By comparing Bayes factors we can automatically choose the most adequate model. Therefore our method is parameter-free. CONCLUSIONS: Heavy-tailed distributions are well-suited to describe large-scale conformational differences in protein structures. A scale mixture representation facilitates the fitting of these distributions and enables outlier-tolerant superposition. Martin Mechelke, Michael Habeck |
BMC Bioinform. | 2 |
| 2008 | Mixture models for protein structure ensemblesabstractMOTIVATION: Protein structure ensembles provide important insight into the dynamics and function of a protein and contain information that is not captured with a single static structure. However, it is not clear a priori to what extent the variability within an ensemble is caused by internal structural changes. Additional variability results from overall translations and rotations of the molecule. And most experimental data do not provide information to relate the structures to a common reference frame. To report meaningful values of intrinsic dynamics, structural precision, conformational entropy, etc., it is therefore important to disentangle local from global conformational heterogeneity. RESULTS: We consider the task of disentangling local from global heterogeneity as an inference problem. We use probabilistic methods to infer from the protein ensemble missing information on reference frames and stable conformational sub-states. To this end, we model a protein ensemble as a mixture of Gaussian probability distributions of either entire conformations or structural segments. We learn these models from a protein ensemble using the expectation-maximization algorithm. Our first model can be used to find multiple conformers in a structure ensemble. The second model partitions the protein chain into locally stable structural segments or core elements and less structured regions typically found in loops. Both models are simple to implement and contain only a single free parameter: the number of conformers or structural segments. Our models can be used to analyse experimental ensembles, molecular dynamics trajectories and conformational change in proteins. AVAILABILITY: The Python source code for protein ensemble analysis is available from the authors upon request. Michael Hirsch 0001, Michael Habeck |
Bioinform. | 2 |
| 2008 | ISD: a software package for Bayesian NMR structure calculationabstractUNLABELLED: The conventional approach to calculating biomolecular structures from nuclear magnetic resonance (NMR) data is often viewed as subjective due to its dependence on rules of thumb for deriving geometric constraints and suitable values for theory parameters from noisy experimental data. As a result, it can be difficult to judge the precision of an NMR structure in an objective manner. The inferential structure determination (ISD) framework, which has been introduced recently, addresses this problem by using Bayesian inference to derive a probability distribution that represents both the unknown structure and its uncertainty. It also determines additional unknowns, such as theory parameters, that normally need to be chosen empirically. Here we give an overview of the ISD software package, which implements this methodology. AVAILABILITY: http://www.bioc.cam.ac.uk/isd Wolfgang Rieping, Michael Nilges, Michael Habeck |
Bioinform. | 3 |
| 2007 | ARIA2: Automated NOE assignment and data integration in NMR structure calculationabstractUNLABELLED: Modern structural genomics projects demand for integrated methods for the interpretation and storage of nuclear magnetic resonance (NMR) data. Here we present version 2.1 of our program ARIA (Ambiguous Restraints for Iterative Assignment) for automated assignment of nuclear Overhauser enhancement (NOE) data and NMR structure calculation. We report on recent developments, most notably a graphical user interface, and the incorporation of the object-oriented data model of the Collaborative Computing Project for NMR (CCPN). The CCPN data model defines a storage model for NMR data, which greatly facilitates the transfer of data between different NMR software packages. AVAILABILITY: A distribution with the source code of ARIA 2.1 is freely available at http://www.pasteur.fr/recherche/unites/Binfs/aria2. Wolfgang Rieping, Michael Habeck, Benjamin Bardiaux, Aymeric Bernard, Therese E. Malliavin, Michael Nilges |
Bioinform. | 2 |
| 2003 | ARIA: automated NOE assignment and NMR structure calculationabstractMOTIVATION: In the light of several ongoing structural genomics projects, faster and more reliable methods for structure calculation from NMR data are in great demand. The major bottleneck in the determination of solution NMR structures is the assignment of NOE peaks (nuclear Overhauser effect). Due to the high complexity of the assignment problem, most NOEs cannot be directly converted into unambiguous inter-proton distance restraints. RESULTS: We present version 1.2 of our program ARIA (Ambiguous Restraints for Iterative Assignment) for automated assignment of NOE data and NMR structure calculation. We summarize recent progress in correcting for spin diffusion with a relaxation matrix approach, representing non-bonded interactions in the force field and refining final structures in explicit solvent. We also discuss book-keeping, data exchange with spectra assignment programs and deposition of the analysed experimental data to the databases. AVAILABILITY: ARIA 1.2 is available from: http://www.pasteur.fr/recherche/unites/Binfs/aria/. SUPPLEMENTARY INFORMATION: XML DTDs (for chemical shifts and NOE crosspeaks), Python scripts for the conversion of various NMR data formats and the results of example calculations using data from the S. cerevisiae HRDC domain are available from: http://www.pasteur.fr/recherche/unites/Binfs/aria/ Jens P. Linge, Michael Habeck, Wolfgang Rieping, Michael Nilges |
Bioinform. | 2 |