Simon V. Mathis

dblp:338/5638 · also Simon Valentin Mathis · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 50% Graph learning · 37% Representation and self-supervised learning · 10%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.522024
DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform · NeurIPS 2024
Dynamics-Informed Protein Design with Structure Conditioning · ICLR 2024
Bioinformatics and computational biology
geometric deep learning
0.912025
gRNAde: Geometric Deep Learning for 3D RNA inverse design · ICLR 2025
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA design
0.912025
gRNAde: Geometric Deep Learning for 3D RNA inverse design · ICLR 2025
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.812024
Dynamics-Informed Protein Design with Structure Conditioning · ICLR 2024
Machine learning › Generative modeling › diffusion model
conditional sampling
0.812024
DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform · NeurIPS 2024
Machine learning › Graph learning › graph neural network › geometric graph neural network
equivariant graph neural network
0.812024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Machine learning › Graph learning
graph neural network
0.812024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Machine learning › Generative modeling › diffusion model
inverse problem solving
0.812024
DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform · NeurIPS 2024
Machine learning › Representation and self-supervised learning
pre-training
0.812024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Bioinformatics and computational biology
protein design
0.812024
Dynamics-Informed Protein Design with Structure Conditioning · ICLR 2024
Bioinformatics and computational biology › structural bioinformatics
protein structure
0.812024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Bioinformatics and computational biology › structural bioinformatics › protein structure representation
protein structure representation learning
0.812024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Machine learning › Graph learning › graph neural network
expressive power
0.712023
On the Expressive Power of Geometric Graph Neural Networks · ICML 2023
Machine learning › Graph learning › graph neural network
geometric graph neural network
0.712023
On the Expressive Power of Geometric Graph Neural Networks · ICML 2023
Graph algorithms and graph theory
graph isomorphism
0.712023
On the Expressive Power of Geometric Graph Neural Networks · ICML 2023
Graph algorithms and graph theory › graph isomorphism
weisfeiler-leman algorithm
0.712023
On the Expressive Power of Geometric Graph Neural Networks · ICML 2023

Methods — techniques the papers use, named apart from their topics

multi-state GNN · 1.7graph neural network · 1.7autoregressive decoding · 1.7pre-training · 1.5normal mode analysis · 1.5geometric graph neural networks · 1.5classifier guidance · 1.5weisfeiler-leman test · 1.3universal approximation · 1.3stochastic differential equations · 0.8stochastic differential equation · 0.8fine-tuning · 0.8doob's h-transform · 0.8
YearPublicationVenuePosition
2025 gRNAde: Geometric Deep Learning for 3D RNA inverse design
abstract
Computational RNA design tasks are often posed as inverse problems, where sequences are designed based on adopting a single desired secondary structure without considering 3D conformational diversity. We introduce gRNAde, a geometric RNA design pipeline operating on 3D RNA backbones to design sequences that explicitly account for structure and dynamics. gRNAde uses a multi-state Graph Neural Network and autoregressive decoding to generates candidate RNA sequences conditioned on one or more 3D backbone structures where the identities of the bases are unknown. On a single-state fixed backbone re-design benchmark of 14 RNA structures from the PDB identified by Das et al. (2010), gRNAde obtains higher native sequence recovery rates (56% on average) compared to Rosetta (45% on average), taking under a second to produce designs compared to the reported hours for Rosetta. We further demonstrate the utility of gRNAde on a new benchmark of multi-state design for structurally flexible RNAs, as well as zero-shot ranking of mutational fitness landscapes in a retrospective analysis of a recent ribozyme. Experimental wet lab validation on 10 different structured RNA backbones finds that gRNAde has a success rate of 50% at designing pseudoknotted RNA structures, a significant advance over 35% for Rosetta. Open source code and tutorials are available at: github.com/chaitjo/geometric-rna-design
Chaitanya K. Joshi, Arian Rokkum Jamasb, Ramón Viñas 0001, Charles Harris, Simon V. Mathis, Alex Morehead, Rishabh Anand, Pietro Liò
ICLR5
2024 Evaluating Representation Learning on the Protein Structure Universe
abstract
We introduce ProteinWorkshop, a comprehensive benchmark suite for representation learning on protein structures with Geometric Graph Neural Networks. We consider large-scale pre-training and downstream tasks on both experimental and predicted structures to enable the systematic evaluation of the quality of the learned structural representation and their usefulness in capturing functional relationships for downstream tasks. We find that: (1) large-scale pretraining on AlphaFold structures and auxiliary tasks consistently improve the performance of both rotation-invariant and equivariant GNNs, and (2) more expressive equivariant GNNs benefit from pretraining to a greater extent compared to invariant models. We aim to establish a common ground for the machine learning and computational biology communities to rigorously compare and advance protein structure representation learning. Our open-source codebase reduces the barrier to entry for working with large protein structure datasets by providing: (1) storage-efficient dataloaders for large-scale structural databases including AlphaFoldDB and ESM Atlas, as well as (2) utilities for constructing new tasks from the entire PDB. ProteinWorkshop is available at: github.com/a-r-j/ProteinWorkshop.
Arian Rokkum Jamasb, Alex Morehead, Chaitanya K. Joshi, Zuobai Zhang, Kieran Didi, Simon V. Mathis, Charles Harris, Jian Tang 0005, Jianlin Cheng, Pietro Liò, Tom L. Blundell
ICLR6
2024 Dynamics-Informed Protein Design with Structure Conditioning
abstract
Current protein generative models are able to design novel backbones with desired shapes or functional motifs. However, despite the importance of a protein’s dynamical properties for its function, conditioning on dynamical properties remains elusive. We present a new approach to protein generative modeling by leveraging Normal Mode Analysis that enables us to capture dynamical properties too. We introduce a method for conditioning the diffusion probabilistic models on protein dynamics, specifically on the lowest non-trivial normal mode of oscillation. Our method, similar to the classifier guidance conditioning, formulates the sampling process as being driven by conditional and unconditional terms. However, unlike previous works, we approximate the conditional term with a simple analytical function rather than an external neural network, thus making the eigenvector calculations approachable. We present the corresponding SDE theory as a formal justification of our approach. We extend our framework to conditioning on structure and dynamics at the same time, enabling scaffolding of the dynamical motifs. We demonstrate the empirical effectiveness of our method by turning the open-source unconditional protein diffusion model Genie into the conditional model with no retraining. Generated proteins exhibit the desired dynamical and structural properties while still being biologically plausible. Our work represents a first step towards incorporating dynamical behaviour in protein design and may open the door to designing more flexible and functional proteins in the future.
Urszula Julia Komorowska, Simon V. Mathis, Kieran Didi, Francisco Vargas 0001, Pietro Liò, Mateja Jamnik
ICLR2
2024 DEFT: Efficient Fine-tuning of Diffusion Models by Learning the Generalised $h$-transform
abstract
Generative modelling paradigms based on denoising diffusion processes have emerged as a leading candidate for conditional sampling in inverse problems. In many real-world applications, we often have access to large, expensively trained unconditional diffusion models, which we aim to exploit for improving conditional sampling. Most recent approaches are motivated heuristically and lack a unifying framework, obscuring connections between them. Further, they often suffer from issues such as being very sensitive to hyperparameters, being expensive to train or needing access to weights hidden behind a closed API. In this work, we unify conditional training and sampling using the mathematically well-understood Doob's h-transform. This new perspective allows us to unify many existing methods under a common umbrella. Under this framework, we propose DEFT (Doob's h-transform Efficient FineTuning), a new approach for conditional generation that simply fine-tunes a very small network to quickly learn the conditional $h$-transform, while keeping the larger unconditional network unchanged. DEFT is much faster than existing baselines while achieving state-of-the-art performance across a variety of linear and non-linear benchmarks. On image reconstruction tasks, we achieve speedups of up to 1.6$\times$, while having the best perceptual quality on natural images and reconstruction performance on medical images. Further, we also provide initial experiments on protein motif scaffolding and outperform reconstruction guidance methods.
Alexander Denker, Francisco Vargas 0001, Shreyas Padhy, Kieran Didi, Simon V. Mathis, Riccardo Barbano, Vincent Dutordoir, Emile Mathieu, Urszula Julia Komorowska, Pietro Liò
NeurIPS5
2023 On the Expressive Power of Geometric Graph Neural Networks
abstract
The expressive power of Graph Neural Networks (GNNs) has been studied extensively through the Weisfeiler-Leman (WL) graph isomorphism test. However, standard GNNs and the WL framework are inapplicable for geometric graphs embedded in Euclidean space, such as biomolecules, materials, and other physical systems. In this work, we propose a geometric version of the WL test (GWL) for discriminating geometric graphs while respecting the underlying physical symmetries: permutations, rotation, reflection, and translation. We use GWL to characterise the expressive power of geometric GNNs that are invariant or equivariant to physical symmetries in terms of distinguishing geometric graphs. GWL unpacks how key design choices influence geometric GNN expressivity: (1) Invariant layers have limited expressivity as they cannot distinguish one-hop identical geometric graphs; (2) Equivariant layers distinguish a larger class of graphs by propagating geometric information beyond local neighbourhoods; (3) Higher order tensors and scalarisation enable maximally powerful geometric GNNs; and (4) GWL's discrimination-based perspective is equivalent to universal approximation. Synthetic experiments supplementing our results are available at https://github.com/chaitjo/geometric-gnn-dojo
Chaitanya K. Joshi, Cristian Bodnar, Simon V. Mathis, Taco Cohen, Pietro Liò
ICML3