VLDB 2026 Research / reviewers in the wild / expert
Kieran Didi
dblp:336/6909
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Generative modeling · 75% Graph learning · 15% Representation and self-supervised learning · 8% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% |
Topics — the 17 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
protein design |
1.6 | 2 | 2025 | Proteina: Scaling Flow-based Protein Structure Generative Models · ICLR 2025 Dynamics-Informed Protein Design with Structure Conditioning · ICLR 2024 |
Machine learning › Generative modeling
diffusion model |
1.5 | 2 | 2024 | DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform · NeurIPS 2024 Dynamics-Informed Protein Design with Structure Conditioning · ICLR 2024 |
Machine learning › Generative modeling › diffusion model › controllable generation
compositional generation |
0.9 | 1 | 2025 | Compositional Flows for 3D Molecule and Synthesis Pathway Co-design · ICML 2025 |
Machine learning › Generative modeling
flow matching |
0.9 | 1 | 2025 | Compositional Flows for 3D Molecule and Synthesis Pathway Co-design · ICML 2025 |
Machine learning › Generative modeling
normalizing flow |
0.9 | 1 | 2025 | Proteina: Scaling Flow-based Protein Structure Generative Models · ICLR 2025 |
Machine learning › Generative modeling › protein design
protein structure generation |
0.9 | 1 | 2025 | Proteina: Scaling Flow-based Protein Structure Generative Models · ICLR 2025 |
Bioinformatics and computational biology › drug discovery
drug design |
0.9 | 1 | 2025 | Compositional Flows for 3D Molecule and Synthesis Pathway Co-design · ICML 2025 |
Bioinformatics and computational biology › molecular informatics › cheminformatics › molecule generation
synthesizable molecule design |
0.9 | 1 | 2025 | Compositional Flows for 3D Molecule and Synthesis Pathway Co-design · ICML 2025 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.8 | 1 | 2024 | Dynamics-Informed Protein Design with Structure Conditioning · ICLR 2024 |
Machine learning › Generative modeling › diffusion model
conditional sampling |
0.8 | 1 | 2024 | DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform · NeurIPS 2024 |
Machine learning › Graph learning › graph neural network › geometric graph neural network
equivariant graph neural network |
0.8 | 1 | 2024 | Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024 |
Machine learning › Graph learning
graph neural network |
0.8 | 1 | 2024 | Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024 |
Machine learning › Generative modeling › diffusion model
inverse problem solving |
0.8 | 1 | 2024 | DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning
pre-training |
0.8 | 1 | 2024 | Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024 |
Bioinformatics and computational biology › structural bioinformatics
protein structure |
0.8 | 1 | 2024 | Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024 |
Bioinformatics and computational biology › structural bioinformatics › protein structure representation
protein structure representation learning |
0.8 | 1 | 2024 | Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024 |
Machine learning › Generative modeling
generative flow networks |
0.3 | 1 | 2025 | Compositional Flows for 3D Molecule and Synthesis Pathway Co-design · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
flow matching · 3.5transformer · 1.7reward-guided sampling · 1.7generative flow networks · 1.7classifier-free guidance · 1.7LoRA · 1.7pre-training · 1.5normal mode analysis · 1.5geometric graph neural networks · 1.5classifier guidance · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Proteina: Scaling Flow-based Protein Structure Generative ModelsabstractRecently, diffusion- and flow-based generative models of protein structures have emerged as a powerful tool for de novo protein design. Here, we develop *Proteina*, a new large-scale flow-based protein backbone generator that utilizes hierarchical fold class labels for conditioning and relies on a tailored scalable transformer architecture with up to $5\times$ as many parameters as previous models. To meaningfully quantify performance, we introduce a new set of metrics that directly measure the distributional similarity of generated proteins with reference sets, complementing existing metrics. We further explore scaling training data to millions of synthetic protein structures and explore improved training and sampling recipes adapted to protein backbone generation. This includes fine-tuning strategies like LoRA for protein backbones, new guidance methods like classifier-free guidance and autoguidance for protein backbones, and new adjusted training objectives. Proteina achieves state-of-the-art performance on de novo protein backbone design and produces diverse and designable proteins at unprecedented length, up to 800 residues. The hierarchical conditioning offers novel control, enabling high-level secondary-structure guidance as well as low-level fold-specific generation. Tomas Geffner, Kieran Didi, Zuobai Zhang, Danny Reidenbach, Zhonglin Cao, Jason Yim, Mario Geiger, Christian Dallago, Emine Küçükbenli, Arash Vahdat, Karsten Kreis |
ICLR | 2 |
| 2025 | Compositional Flows for 3D Molecule and Synthesis Pathway Co-designabstractMany generative applications, such as synthesis-based 3D molecular design, involve constructing compositional objects with continuous features.
Here, we introduce Compositional Generative Flows (CGFlow), a novel framework that extends flow matching to generate objects in compositional steps while modeling continuous states.
Our key insight is that modeling compositional state transitions can be formulated as a straightforward extension of the flow matching interpolation process.
We further build upon the theoretical foundations of generative flow networks (GFlowNets), enabling reward-guided sampling of compositional structures.
We apply CGFlow to synthesizable drug design by jointly designing the molecule's synthetic pathway with its 3D binding pose.
Our approach achieves state-of-the-art binding affinity and synthesizability on all 15 targets from the LIT-PCBA benchmark, and 4.2x improvement in sampling efficiency compared to 2D synthesis-based baseline.
To our best knowledge, our method is also the first to achieve state of-art-performance in both Vina Dock (-9.42) and AiZynth success rate (36.1\%) on the CrossDocked2020 benchmark. Tony Shen, Seonghwan Seo, Ross Irwin, Kieran Didi, Simon Olsson, Woo Youn Kim, Martin Ester |
ICML | 4 |
| 2024 | Evaluating Representation Learning on the Protein Structure UniverseabstractWe introduce ProteinWorkshop, a comprehensive benchmark suite for representation learning on protein structures with Geometric Graph Neural Networks. We consider large-scale pre-training and downstream tasks on both experimental and predicted structures to enable the systematic evaluation of the quality of the learned structural representation and their usefulness in capturing functional relationships for downstream tasks. We find that: (1) large-scale pretraining on AlphaFold structures and auxiliary tasks consistently improve the performance of both rotation-invariant and equivariant GNNs, and (2) more expressive equivariant GNNs benefit from pretraining to a greater extent compared to invariant models.
We aim to establish a common ground for the machine learning and computational biology communities to rigorously compare and advance protein structure representation learning. Our open-source codebase reduces the barrier to entry for working with large protein structure datasets by providing: (1) storage-efficient dataloaders for large-scale structural databases including AlphaFoldDB and ESM Atlas, as well as (2) utilities for constructing new tasks from the entire PDB. ProteinWorkshop is available at: github.com/a-r-j/ProteinWorkshop. Arian Rokkum Jamasb, Alex Morehead, Chaitanya K. Joshi, Zuobai Zhang, Kieran Didi, Simon V. Mathis, Charles Harris, Jian Tang 0005, Jianlin Cheng, Pietro Liò, Tom L. Blundell |
ICLR | 5 |
| 2024 | Dynamics-Informed Protein Design with Structure ConditioningabstractCurrent protein generative models are able to design novel backbones with desired shapes or functional motifs. However, despite the importance of a protein’s dynamical properties for its function, conditioning on dynamical properties remains elusive. We present a new approach to protein generative modeling by leveraging Normal Mode Analysis that enables us to capture dynamical properties too. We introduce a method for conditioning the diffusion probabilistic models on protein dynamics, specifically on the lowest non-trivial normal mode of oscillation. Our method, similar to the classifier guidance conditioning, formulates the sampling process as being driven by conditional and unconditional terms. However, unlike previous works, we approximate the conditional term with a simple analytical function rather than an external neural network, thus making the eigenvector calculations approachable. We present the corresponding SDE theory as a formal justification of our approach. We extend our framework to conditioning on structure and dynamics at the same time, enabling scaffolding of the dynamical motifs. We demonstrate the empirical effectiveness of our method by turning the open-source unconditional protein diffusion model Genie into the conditional model with no retraining. Generated proteins exhibit the desired dynamical and structural properties while still being biologically plausible. Our work represents a first step towards incorporating dynamical behaviour in protein design and may open the door to designing more flexible and functional proteins in the future. Urszula Julia Komorowska, Simon V. Mathis, Kieran Didi, Francisco Vargas 0001, Pietro Liò, Mateja Jamnik |
ICLR | 3 |
| 2024 | DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transformabstractGenerative modelling paradigms based on denoising diffusion processes have emerged as a leading candidate for conditional sampling in inverse problems.
In many real-world applications, we often have access to large, expensively trained unconditional diffusion models, which we aim to exploit for improving conditional sampling.
Most recent approaches are motivated heuristically and lack a unifying framework, obscuring connections between them. Further, they often suffer from issues such as being very sensitive to hyperparameters, being expensive to train or needing access to weights hidden behind a closed API. In this work, we unify conditional training and sampling using the mathematically well-understood Doob's h-transform. This new perspective allows us to unify many existing methods under a common umbrella. Under this framework, we propose DEFT (Doob's h-transform Efficient FineTuning), a new approach for conditional generation that simply fine-tunes a very small network to quickly learn the conditional $h$-transform, while keeping the larger unconditional network unchanged. DEFT is much faster than existing baselines while achieving state-of-the-art performance across a variety of linear and non-linear benchmarks. On image reconstruction tasks, we achieve speedups of up to 1.6$\times$, while having the best perceptual quality on natural images and reconstruction performance on medical images. Further, we also provide initial experiments on protein motif scaffolding and outperform reconstruction guidance methods. Alexander Denker, Francisco Vargas 0001, Shreyas Padhy, Kieran Didi, Simon V. Mathis, Riccardo Barbano, Vincent Dutordoir, Emile Mathieu, Urszula Julia Komorowska, Pietro Liò |
NeurIPS | 4 |