Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gabriele Corso

dblp:262/6499 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
11since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Generative modeling · 60% Graph learning · 22% Representation and self-supervised learning · 8%
Interdisciplinary, comprehensive, and emerging computing
8 papers
Bioinformatics and computational biology · 88% Computational science and engineering · 12%

Topics — the 30 heaviest of 34, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
4.162024
DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete Latents · ICML 2024
Particle Guidance: non-I.I.D. Diverse Sampling with Diffusion Models · ICLR 2024
Deep Confident Steps to New Pockets: Strategies for Docking Generalization · ICLR 2024
Bioinformatics and computational biology › molecular informatics › molecular modeling
molecular docking
2.332025
Composing Unbalanced Flows for Flexible Docking and Relaxation · ICLR 2025
Deep Confident Steps to New Pockets: Strategies for Docking Generalization · ICLR 2024
DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking · ICLR 2023
Machine learning › Generative modeling
flow matching
1.622025
Composing Unbalanced Flows for Flexible Docking and Relaxation · ICLR 2025
Dirichlet Flow Matching with Applications to DNA Sequence Design · ICML 2024
Machine learning › Graph learning
graph neural network
1.532022
3D Infomax improves GNNs for Molecular Property Prediction · ICML 2022
Directional Graph Networks · ICML 2021
Principal Neighbourhood Aggregation for Graph Nets · NeurIPS 2020
Bioinformatics and computational biology › molecular informatics › molecular modeling › molecular docking
flexible docking
0.912025
Composing Unbalanced Flows for Flexible Docking and Relaxation · ICLR 2025
Bioinformatics and computational biology › structural bioinformatics
protein structure
0.912025
Composing Unbalanced Flows for Flexible Docking and Relaxation · ICLR 2025
Bioinformatics and computational biology › molecular informatics › molecular modeling
molecular conformation generation
0.822024
Torsional Diffusion for Molecular Conformer Generation · NeurIPS 2022
Particle Guidance: non-I.I.D. Diverse Sampling with Diffusion Models · ICLR 2024
Machine learning › Generative modeling › diffusion model
discrete diffusion model
0.812024
Dirichlet Flow Matching with Applications to DNA Sequence Design · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.812024
DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete Latents · ICML 2024
Machine learning › Generative modeling › generative model evaluation
sample diversity
0.812024
Particle Guidance: non-I.I.D. Diverse Sampling with Diffusion Models · ICLR 2024
Machine learning › Deep learning architectures and training › sequence modeling
sequence generation
0.812024
Dirichlet Flow Matching with Applications to DNA Sequence Design · ICML 2024
Bioinformatics and computational biology › molecular informatics › molecular modeling › molecular docking
blind docking
0.812024
Deep Confident Steps to New Pockets: Strategies for Docking Generalization · ICLR 2024
Bioinformatics and computational biology › structural bioinformatics
protein-ligand binding
0.812024
Deep Confident Steps to New Pockets: Strategies for Docking Generalization · ICLR 2024
Machine learning › Generative modeling › diffusion model › geometric diffusion model
equivariant diffusion model
0.712023
DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking · ICLR 2023
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
3D Infomax improves GNNs for Molecular Property Prediction · ICML 2022
Machine learning › Generative modeling › diffusion model
molecular conformation generation
0.612022
Torsional Diffusion for Molecular Conformer Generation · NeurIPS 2022
Machine learning › Graph learning › molecular representation learning › molecular graph learning
molecular property prediction
0.612022
3D Infomax improves GNNs for Molecular Property Prediction · ICML 2022
Machine learning › Representation and self-supervised learning
mutual information maximization
0.612022
3D Infomax improves GNNs for Molecular Property Prediction · ICML 2022
Machine learning › Generative modeling › diffusion model › efficient diffusion model
subspace diffusion model
0.612022
Subspace Diffusion Generative Models · ECCV (23) 2022
Computational science and engineering
computational chemistry
0.612022
Torsional Diffusion for Molecular Conformer Generation · NeurIPS 2022
Machine learning › Graph learning › graph representation learning › structural encoding
distance encoding
0.512021
Neural Distance Embeddings for Biological Sequences · NeurIPS 2021
Computational science and engineering › graph learning
hyperbolic embedding
0.512021
Neural Distance Embeddings for Biological Sequences · NeurIPS 2021
Bioinformatics and computational biology
sequence analysis
0.512021
Neural Distance Embeddings for Biological Sequences · NeurIPS 2021
Bioinformatics and computational biology › sequence analysis › sequence feature extraction
sequence embedding
0.512021
Neural Distance Embeddings for Biological Sequences · NeurIPS 2021
Machine learning › Graph learning › graph neural network
aggregation function
0.412020
Principal Neighbourhood Aggregation for Graph Nets · NeurIPS 2020
Machine learning › Graph learning › graph neural network
expressive power
0.412020
Principal Neighbourhood Aggregation for Graph Nets · NeurIPS 2020
Machine learning › Generative modeling
autoregressive model
0.212024
DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete Latents · ICML 2024
Machine learning › Generative modeling › image generation
conditional image generation
0.212024
Particle Guidance: non-I.I.D. Diverse Sampling with Diffusion Models · ICLR 2024
Bioinformatics and computational biology › synthetic biology
DNA sequence design
0.212024
Dirichlet Flow Matching with Applications to DNA Sequence Design · ICML 2024
Computer vision › 3D vision
molecular structure
0.212022
Torsional Diffusion for Molecular Conformer Generation · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.8optimal transport · 1.7flow matching · 1.7synthetic data · 1.5scaling laws · 1.5joint-particle potential · 1.5distillation · 1.5diffusion sampling · 1.5confidence model · 1.5classifier-free guidance · 1.5
YearPublicationVenuePosition
2025 Composing Unbalanced Flows for Flexible Docking and Relaxation
abstract
Diffusion models have emerged as a successful approach for molecular docking, but they often cannot model protein flexibility or generate nonphysical poses. We argue that both these challenges can be tackled by framing the problem as a transport between distributions. Still, existing paradigms lack the flexibility to define effective maps between such complex distributions. To address this limitation, we propose Unbalanced Flow Matching, a generalization of Flow Matching (FM) that allows trading off sample efficiency with approximation accuracy and enables more accurate transport. Empirically, we apply Unbalanced FM on flexible docking and structure relaxation, demonstrating our ability to model protein flexibility and generate energetically favorable poses. On the PDBBind docking benchmark, our method FlexDock improves the docking performance while increasing the proportion of energetically favorable poses from 30% to 73%.
Gabriele Corso, Vignesh Ram Somnath, Noah Getz, Regina Barzilay, Tommi S. Jaakkola, Andreas Krause 0001
ICLR1
2024 Deep Confident Steps to New Pockets: Strategies for Docking Generalization
abstract
Accurate blind docking has the potential to lead to new biological breakthroughs, but for this promise to be realized, docking methods must generalize well across the proteome. Existing benchmarks, however, fail to rigorously assess generalizability. Therefore, we develop DockGen, a new benchmark based on the ligand-binding domains of proteins, and we show that existing machine learning-based docking models have very weak generalization abilities. We carefully analyze the scaling laws of ML-based docking and show that, by scaling data and model size, as well as integrating synthetic data strategies, we are able to significantly increase the generalization capacity and set new state-of-the-art performance across benchmarks. Further, we propose Confidence Bootstrapping, a new training paradigm that solely relies on the interaction between diffusion and confidence models and exploits the multi-resolution generation process of diffusion models. We demonstrate that Confidence Bootstrapping significantly improves the ability of ML-based docking methods to dock to unseen protein classes, edging closer to accurate and generalizable blind docking methods.
Gabriele Corso, Arthur Deng, Nicholas Polizzi, Regina Barzilay, Tommi S. Jaakkola
ICLR1
2024 Particle Guidance: non-I.I.D. Diverse Sampling with Diffusion Models
abstract
In light of the widespread success of generative models, a significant amount of research has gone into speeding up their sampling time. However, generative models are often sampled multiple times to obtain a diverse set incurring a cost that is orthogonal to sampling time. We tackle the question of how to improve diversity and sample efficiency by moving beyond the common assumption of independent samples. We propose particle guidance, an extension of diffusion-based generative sampling where a joint-particle time-evolving potential enforces diversity. We analyze theoretically the joint distribution that particle guidance generates, how to learn a potential that achieves optimal diversity, and the connections with methods in other disciplines. Empirically, we test the framework both in the setting of conditional image generation, where we are able to increase diversity without affecting quality, and molecular conformer generation, where we reduce the state-of-the-art median error by 13% on average.
Gabriele Corso, Valentin De Bortoli, Regina Barzilay, Tommi S. Jaakkola
ICLR1
2024 Dirichlet Flow Matching with Applications to DNA Sequence Design
abstract
Discrete diffusion or flow models could enable faster and more controllable sequence generation than autoregressive models. We show that naive linear flow matching on the simplex is insufficient toward this goal since it suffers from discontinuities in the training target and further pathologies. To overcome this, we develop Dirichlet flow matching on the simplex based on mixtures of Dirichlet distributions as probability paths. In this framework, we derive a connection between the mixtures' scores and the flow's vector field that allows for classifier and classifier-free guidance. Further, we provide distilled Dirichlet flow matching, which enables one-step sequence generation with minimal performance hits, resulting in $O(L)$ speedups compared to autoregressive models. On complex DNA sequence generation tasks, we demonstrate superior performance compared to all baselines in distributional metrics and in achieving desired design targets for generated sequences. Finally, we show that our classifier-free guidance approach improves unconditional generation and is effective for generating DNA that satisfies design targets.
Hannes Stärk, Bowen Jing 0002, Chenyu Wang 0003, Gabriele Corso, Bonnie Berger, Regina Barzilay, Tommi S. Jaakkola
ICML4
2024 DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete Latents
abstract
Diffusion models (DMs) have revolutionized generative learning. They utilize a diffusion process to encode data into a simple Gaussian distribution. However, encoding a complex, potentially multimodal data distribution into a single *continuous* Gaussian distribution arguably represents an unnecessarily challenging learning problem. We propose ***Dis**crete-**Co**ntinuous Latent Variable **Diff**usion Models (DisCo-Diff)* to simplify this task by introducing complementary *discrete* latent variables. We augment DMs with learnable discrete latents, inferred with an encoder, and train DM and encoder end-to-end. DisCo-Diff does not rely on pre-trained networks, making the framework universally applicable. The discrete latents significantly simplify learning the DM's complex noise-to-data mapping by reducing the curvature of the DM's generative ODE. An additional autoregressive transformer models the distribution of the discrete latents, a simple step because DisCo-Diff requires only few discrete variables with small codebooks. We validate DisCo-Diff on toy data, several image synthesis tasks as well as molecular docking, and find that introducing discrete latents consistently improves model performance. For example, DisCo-Diff achieves state-of-the-art FID scores on class-conditioned ImageNet-64/128 datasets with ODE sampler.
Gabriele Corso, Tommi S. Jaakkola, Arash Vahdat, Karsten Kreis
ICML2
2023 DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking
Gabriele Corso, Hannes Stärk, Bowen Jing 0002, Regina Barzilay, Tommi S. Jaakkola
ICLR1
2022 Subspace Diffusion Generative Models
Bowen Jing 0002, Gabriele Corso, Renato Berlinghieri, Tommi S. Jaakkola
ECCV (23)2
2022 3D Infomax improves GNNs for Molecular Property Prediction
abstract
Molecular property prediction is one of the fastest-growing applications of deep learning with critical real-world impacts. Although the 3D molecular graph structure is necessary for models to achieve strong performance on many tasks, it is infeasible to obtain 3D structures at the scale required by many real-world applications. To tackle this issue, we propose to use existing 3D molecular datasets to pre-train a model to reason about the geometry of molecules given only their 2D molecular graphs. Our method, called 3D Infomax, maximizes the mutual information between learned 3D summary vectors and the representations of a graph neural network (GNN). During fine-tuning on molecules with unknown geometry, the GNN is still able to produce implicit 3D information and uses it for downstream tasks. We show that 3D Infomax provides significant improvements for a wide range of properties, including a 22% average MAE reduction on QM9 quantum mechanical properties. Moreover, the learned representations can be effectively transferred between datasets in different molecular spaces.
Hannes Stärk, Dominique Beaini, Gabriele Corso, Prudencio Tossou, Christian Dallago, Stephan Günnemann, Pietro Liò
ICML3
2022 Torsional Diffusion for Molecular Conformer Generation
abstract
Molecular conformer generation is a fundamental task in computational chemistry. Several machine learning approaches have been developed, but none have outperformed state-of-the-art cheminformatics methods. We propose torsional diffusion, a novel diffusion framework that operates on the space of torsion angles via a diffusion process on the hypertorus and an extrinsic-to-intrinsic score model. On a standard benchmark of drug-like molecules, torsional diffusion generates superior conformer ensembles compared to machine learning and cheminformatics methods in terms of both RMSD and chemical properties, and is orders of magnitude faster than previous diffusion-based models. Moreover, our model provides exact likelihoods, which we employ to build the first generalizable Boltzmann generator. Code is available at https://github.com/gcorso/torsional-diffusion.
Bowen Jing 0002, Gabriele Corso, Jeffrey Chang, Regina Barzilay, Tommi S. Jaakkola
NeurIPS2
2021 Directional Graph Networks
Dominique Beaini, Saro Passaro, Vincent Létourneau, William L. Hamilton, Gabriele Corso, Pietro Liò
ICML5
2021 Neural Distance Embeddings for Biological Sequences
abstract
The development of data-dependent heuristics and representations for biological sequences that reflect their evolutionary distance is critical for large-scale biological research. However, popular machine learning approaches, based on continuous Euclidean spaces, have struggled with the discrete combinatorial formulation of the edit distance that models evolution and the hierarchical relationship that characterises real-world datasets. We present Neural Distance Embeddings (NeuroSEED), a general framework to embed sequences in geometric vector spaces, and illustrate the effectiveness of the hyperbolic space that captures the hierarchical structure and provides an average 38% reduction in embedding RMSE against the best competing geometry. The capacity of the framework and the significance of these improvements are then demonstrated devising supervised and unsupervised NeuroSEED approaches to multiple core tasks in bioinformatics. Benchmarked with common baselines, the proposed approaches display significant accuracy and/or runtime improvements on real-world datasets. As an example for hierarchical clustering, the proposed pretrained and from-scratch methods match the quality of competing baselines with 30x and 15x runtime reduction, respectively.
Gabriele Corso, Rex Ying, Michal Pándy, Petar Velickovic, Jure Leskovec, Pietro Liò
NeurIPS1
2020 Principal Neighbourhood Aggregation for Graph Nets
abstract
Graph Neural Networks (GNNs) have been shown to be effective models for different predictive tasks on graph-structured data. Recent work on their expressive power has focused on isomorphism tasks and countable feature spaces. We extend this theoretical framework to include continuous features---which occur regularly in real-world input domains and within the hidden layers of GNNs---and we demonstrate the requirement for multiple aggregation functions in this context. Accordingly, we propose Principal Neighbourhood Aggregation (PNA), a novel architecture combining multiple aggregators with degree-scalers (which generalize the sum aggregator). Finally, we compare the capacity of different models to capture and exploit the graph structure via a novel benchmark containing multiple tasks taken from classical graph theory, alongside existing benchmarks from real-world domains, all of which demonstrate the strength of our model. With this work we hope to steer some of the GNN research towards new aggregation methods which we believe are essential in the search for powerful and robust models.
Gabriele Corso, Luca Cavalleri, Dominique Beaini, Pietro Liò, Petar Velickovic
NeurIPS1