Hannes Stärk

dblp:300/4627 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Generative modeling · 73% Graph learning · 8% Representation and self-supervised learning · 8%
Interdisciplinary, comprehensive, and emerging computing
9 papers
Bioinformatics and computational biology · 79% Computational science and engineering · 21%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.332025
Think while You Generate: Discrete Diffusion with Planned Denoising · ICLR 2025
Generative Modeling of Molecular Dynamics Trajectories · NeurIPS 2024
DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking · ICLR 2023
Machine learning › Generative modeling
flow matching
2.332024
ET-Flow: Equivariant Flow-Matching for Molecular Conformer Generation · NeurIPS 2024
Dirichlet Flow Matching with Applications to DNA Sequence Design · ICML 2024
Harmonic Self-Conditioned Flow Matching for joint Multi-Ligand Docking and Binding Site Design · ICML 2024
Machine learning › Generative modeling › diffusion model
discrete diffusion model
1.622025
Think while You Generate: Discrete Diffusion with Planned Denoising · ICLR 2025
Dirichlet Flow Matching with Applications to DNA Sequence Design · ICML 2024
Bioinformatics and computational biology
protein structure prediction
1.012026
Protein FID: improved evaluation of protein structure generative models · Bioinform. 2026
Machine learning › Generative modeling
3d generative model
0.912025
ProtComposer: Compositional Protein Structure Generation with 3D Ellipsoids · ICLR 2025
Machine learning › Generative modeling › diffusion model › discrete diffusion model
diffusion language model
0.912025
Think while You Generate: Discrete Diffusion with Planned Denoising · ICLR 2025
Machine learning › Generative modeling › protein design
protein structure generation
0.912025
ProtComposer: Compositional Protein Structure Generation with 3D Ellipsoids · ICLR 2025
Bioinformatics and computational biology › protein design
protein structure generation
0.912025
ProtComposer: Compositional Protein Structure Generation with 3D Ellipsoids · ICLR 2025
Machine learning › Generative modeling › flow matching
equivariant flow matching
0.812024
ET-Flow: Equivariant Flow-Matching for Molecular Conformer Generation · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
molecular conformation generation
0.812024
ET-Flow: Equivariant Flow-Matching for Molecular Conformer Generation · NeurIPS 2024
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation
0.812024
Dirichlet Flow Matching with Applications to DNA Sequence Design · ICML 2024
Computational science and engineering
computational chemistry
0.812024
Generative Modeling of Molecular Dynamics Trajectories · NeurIPS 2024
Bioinformatics and computational biology › drug discovery
computational drug discovery
0.812024
ET-Flow: Equivariant Flow-Matching for Molecular Conformer Generation · NeurIPS 2024
Bioinformatics and computational biology › molecular informatics › molecular modeling
molecular conformation generation
0.812024
ET-Flow: Equivariant Flow-Matching for Molecular Conformer Generation · NeurIPS 2024
Computational science and engineering › computational chemistry › molecular simulation
molecular dynamics
0.812024
Generative Modeling of Molecular Dynamics Trajectories · NeurIPS 2024
Bioinformatics and computational biology
protein design
0.812024
Harmonic Self-Conditioned Flow Matching for joint Multi-Ligand Docking and Binding Site Design · ICML 2024
Machine learning › Generative modeling › diffusion model › geometric diffusion model
equivariant diffusion model
0.712023
DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking · ICLR 2023
Bioinformatics and computational biology › molecular informatics › molecular modeling
molecular docking
0.712023
DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking · ICLR 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
3D Infomax improves GNNs for Molecular Property Prediction · ICML 2022
Computer vision › 3D vision
geometric deep learning
0.612022
EquiBind: Geometric Deep Learning for Drug Binding Structure Prediction · ICML 2022
Machine learning › Graph learning
graph neural network
0.612022
3D Infomax improves GNNs for Molecular Property Prediction · ICML 2022
Machine learning › Graph learning › molecular representation learning › molecular graph learning
molecular property prediction
0.612022
3D Infomax improves GNNs for Molecular Property Prediction · ICML 2022
Machine learning › Representation and self-supervised learning
mutual information maximization
0.612022
3D Infomax improves GNNs for Molecular Property Prediction · ICML 2022
Bioinformatics and computational biology
drug discovery
0.612022
EquiBind: Geometric Deep Learning for Drug Binding Structure Prediction · ICML 2022
Natural language and speech › Information extraction and text analysis › distributional semantics
distributional similarity
0.312026
Protein FID: improved evaluation of protein structure generative models · Bioinform. 2026
Geometric modeling and processing
shape representation
0.312025
ProtComposer: Compositional Protein Structure Generation with 3D Ellipsoids · ICLR 2025
Machine learning › Generative modeling › diffusion model
conditional generation
0.212024
Generative Modeling of Molecular Dynamics Trajectories · NeurIPS 2024
Bioinformatics and computational biology › synthetic biology
DNA sequence design
0.212024
Dirichlet Flow Matching with Applications to DNA Sequence Design · ICML 2024
Bioinformatics and computational biology › molecular informatics
molecular representation learning
0.212022
3D Infomax improves GNNs for Molecular Property Prediction · ICML 2022
Geometric modeling and processing › geometric deep learning
equivariant neural networks
0.212022
EquiBind: Geometric Deep Learning for Drug Binding Structure Prediction · ICML 2022

Methods — techniques the papers use, named apart from their topics

diffusion model · 3.9fine-tuning · 2.9statistical modeling · 2.6optimal transport distance · 2.0foldseek clustering · 2.0equivariant transformer · 1.5distillation · 1.5classifier-free guidance · 1.5planner-denoiser framework · 0.9mask diffusion · 0.9harmonic prior · 0.8flow matching · 0.8SE(3)-equivariant network · 0.6
YearPublicationVenuePosition
2026 Protein FID: improved evaluation of protein structure generative models
abstract
MOTIVATION: Protein structure generative models have seen a recent surge of interest, but meaningfully evaluating them computationally is an active area of research. While current metrics have driven useful progress, they do not capture how well models sample the design space represented by the training data. We argue for a protein Frechet Inception Distance (FID) metric to supplement current evaluations with a measure of distributional similarity in a semantically meaningful latent space. RESULTS: Our FID behaves desirably under protein structure perturbations and correctly recapitulates similarities between protein samples: it correlates with optimal transport distances and recovers FoldSeek clusters and the CATH hierarchy. Evaluating current protein structure generative models with FID shows that they fall short of modeling the distribution of PDB proteins. AVAILABILITY: Code is available at: https://github.com/ffaltings/protfid.
Felix Faltings, Hannes Stärk, Tommi S. Jaakkola, Regina Barzilay
Bioinform.2
2025 Think while You Generate: Discrete Diffusion with Planned Denoising
abstract
Discrete diffusion has achieved state-of-the-art performance, outperforming or approaching autoregressive models on standard benchmarks. In this work, we introduce *Discrete Diffusion with Planned Denoising* (DDPD), a novel framework that separates the generation process into two models: a planner and a denoiser. At inference time, the planner selects which positions to denoise next by identifying the most corrupted positions in need of denoising, including both initially corrupted and those requiring additional refinement. This plan-and-denoise approach enables more efficient reconstruction during generation by iteratively identifying and denoising corruptions in the optimal order. DDPD outperforms traditional denoiser-only mask diffusion methods, achieving superior results on language modeling benchmarks such as *text8*, *OpenWebText*, and token-based generation on *ImageNet 256 × 256*. Notably, in language modeling, DDPD significantly reduces the performance gap between diffusion-based and autoregressive methods in terms of generative perplexity. Code is available at [github.com/liusulin/DDPD](https://github.com/liusulin/DDPD).
Sulin Liu, Juno Nam, Hannes Stärk, Tommi S. Jaakkola, Rafael Gómez-Bombarelli
ICLR4
2025 ProtComposer: Compositional Protein Structure Generation with 3D Ellipsoids
abstract
We develop ProtComposer to generate protein structures conditioned on spatial protein layouts that are specified via a set of 3D ellipsoids capturing substructure shapes and semantics. At inference time, we condition on ellipsoids that are hand-constructed, extracted from existing proteins, or from a statistical model, with each option unlocking new capabilities. Hand-specifying ellipsoids enables users to control the location, size, orientation, secondary structure, and approximate shape of protein substructures. Conditioning on ellipsoids of existing proteins enables redesigning their substructure's connectivity or editing substructure properties. By conditioning on novel and diverse ellipsoid layouts from a simple statistical model, we improve protein generation with expanded Pareto frontiers between designability, novelty, and diversity. Further, this enables sampling designable proteins with a helix-fraction that matches PDB proteins, unlike existing generative models that commonly oversample conceptually simple helix bundles. Code is available at https://github.com/NVlabs/protcomposer.
Hannes Stärk, Bowen Jing 0002, Tomas Geffner, Jason Yim, Tommi S. Jaakkola, Arash Vahdat, Karsten Kreis
ICLR1
2024 Harmonic Self-Conditioned Flow Matching for joint Multi-Ligand Docking and Binding Site Design
abstract
A significant amount of protein function requires binding small molecules, including enzymatic catalysis. As such, designing binding pockets for small molecules has several impactful applications ranging from drug synthesis to energy storage. Towards this goal, we first develop HarmonicFlow, an improved generative process over 3D protein-ligand binding structures based on our self-conditioned flow matching objective. FlowSite extends this flow model to jointly generate a protein pocket’s discrete residue types and the molecule’s binding 3D structure. We show that HarmonicFlow improves upon state-of-the-art generative processes for docking in simplicity, generality, and average sample quality in pocket-level docking. Enabled by this structure modeling, FlowSite designs binding sites substantially better than baseline approaches.
Hannes Stärk, Bowen Jing 0002, Regina Barzilay, Tommi S. Jaakkola
ICML1
2024 Dirichlet Flow Matching with Applications to DNA Sequence Design
abstract
Discrete diffusion or flow models could enable faster and more controllable sequence generation than autoregressive models. We show that naive linear flow matching on the simplex is insufficient toward this goal since it suffers from discontinuities in the training target and further pathologies. To overcome this, we develop Dirichlet flow matching on the simplex based on mixtures of Dirichlet distributions as probability paths. In this framework, we derive a connection between the mixtures' scores and the flow's vector field that allows for classifier and classifier-free guidance. Further, we provide distilled Dirichlet flow matching, which enables one-step sequence generation with minimal performance hits, resulting in $O(L)$ speedups compared to autoregressive models. On complex DNA sequence generation tasks, we demonstrate superior performance compared to all baselines in distributional metrics and in achieving desired design targets for generated sequences. Finally, we show that our classifier-free guidance approach improves unconditional generation and is effective for generating DNA that satisfies design targets.
Hannes Stärk, Bowen Jing 0002, Chenyu Wang 0003, Gabriele Corso, Bonnie Berger, Regina Barzilay, Tommi S. Jaakkola
ICML1
2024 ET-Flow: Equivariant Flow-Matching for Molecular Conformer Generation
abstract
Predicting low-energy molecular conformations given a molecular graph is an important but challenging task in computational drug discovery. Existing state- of-the-art approaches either resort to large scale transformer-based models that diffuse over conformer fields, or use computationally expensive methods to gen- erate initial structures and diffuse over torsion angles. In this work, we introduce Equivariant Transformer Flow (ET-Flow). We showcase that a well-designed flow matching approach with equivariance and harmonic prior alleviates the need for complex internal geometry calculations and large architectures, contrary to the prevailing methods in the field. Our approach results in a straightforward and scalable method that directly operates on all-atom coordinates with minimal assumptions. With the advantages of equivariance and flow matching, ET-Flow significantly increases the precision and physical validity of the generated con- formers, while being a lighter model and faster at inference. Code is available https://github.com/shenoynikhil/ETFlow.
Majdi Hassan, Nikhil Shenoy, Jungyoon Lee, Hannes Stärk, Stephan Thaler, Dominique Beaini
NeurIPS4
2024 Generative Modeling of Molecular Dynamics Trajectories
abstract
Molecular dynamics (MD) is a powerful technique for studying microscopic phenomena, but its computational cost has driven significant interest in the development of deep learning-based surrogate models. We introduce generative modeling of molecular trajectories as a paradigm for learning flexible multi-task surrogate models of MD from data. By conditioning on appropriately chosen frames of the trajectory, we show such generative models can be adapted to diverse tasks such as forward simulation, transition path sampling, and trajectory upsampling. By alternatively conditioning on part of the molecular system and inpainting the rest, we also demonstrate the first steps towards dynamics-conditioned molecular design. We validate the full set of these capabilities on tetrapeptide simulations and show preliminary results on scaling to protein monomers. Altogether, our work illustrates how generative modeling can unlock value from MD data towards diverse downstream tasks that are not straightforward to address with existing methods or even MD itself. Code is available at https://github.com/bjing2016/mdgen.
Bowen Jing 0002, Hannes Stärk, Tommi S. Jaakkola, Bonnie Berger
NeurIPS2
2023 DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking
Gabriele Corso, Hannes Stärk, Bowen Jing 0002, Regina Barzilay, Tommi S. Jaakkola
ICLR2
2022 3D Infomax improves GNNs for Molecular Property Prediction
abstract
Molecular property prediction is one of the fastest-growing applications of deep learning with critical real-world impacts. Although the 3D molecular graph structure is necessary for models to achieve strong performance on many tasks, it is infeasible to obtain 3D structures at the scale required by many real-world applications. To tackle this issue, we propose to use existing 3D molecular datasets to pre-train a model to reason about the geometry of molecules given only their 2D molecular graphs. Our method, called 3D Infomax, maximizes the mutual information between learned 3D summary vectors and the representations of a graph neural network (GNN). During fine-tuning on molecules with unknown geometry, the GNN is still able to produce implicit 3D information and uses it for downstream tasks. We show that 3D Infomax provides significant improvements for a wide range of properties, including a 22% average MAE reduction on QM9 quantum mechanical properties. Moreover, the learned representations can be effectively transferred between datasets in different molecular spaces.
Hannes Stärk, Dominique Beaini, Gabriele Corso, Prudencio Tossou, Christian Dallago, Stephan Günnemann, Pietro Liò
ICML1
2022 EquiBind: Geometric Deep Learning for Drug Binding Structure Prediction
abstract
Predicting how a drug-like molecule binds to a specific protein target is a core problem in drug discovery. An extremely fast computational binding method would enable key applications such as fast virtual screening or drug engineering. Existing methods are computationally expensive as they rely on heavy candidate sampling coupled with scoring, ranking, and fine-tuning steps. We challenge this paradigm with EquiBind, an SE(3)-equivariant geometric deep learning model performing direct-shot prediction of both i) the receptor binding location (blind docking) and ii) the ligand’s bound pose and orientation. EquiBind achieves significant speed-ups and better quality compared to traditional and recent baselines. Further, we show extra improvements when coupling it with existing fine-tuning techniques at the cost of increased running time. Finally, we propose a novel and fast fine-tuning model that adjusts torsion angles of a ligand’s rotatable bonds based on closed form global minima of the von Mises angular distance to a given input atomic point cloud, avoiding previous expensive differential evolution strategies for energy minimization.
Hannes Stärk, Octavian-Eugen Ganea, Lagnajit Pattanaik, Regina Barzilay, Tommi S. Jaakkola
ICML1