Chaitanya K. Joshi

dblp:202/2132 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0003-4722-1815ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Graph learning · 60% Generative modeling · 15% Representation and self-supervised learning · 13%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 79% Computational science and engineering · 21%
Theoretical computer science
1 paper
Graph algorithms and graph theory · 100%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Graph learning
graph neural network
1.422024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Benchmarking Graph Neural Networks · J. Mach. Learn. Res. 2023
Machine learning › Generative modeling
diffusion model
0.912025
All-atom Diffusion Transformers: Unified generative modelling of molecules and materials · ICML 2025
Computational science and engineering
computational chemistry
0.912025
All-atom Diffusion Transformers: Unified generative modelling of molecules and materials · ICML 2025
Bioinformatics and computational biology
geometric deep learning
0.912025
gRNAde: Geometric Deep Learning for 3D RNA inverse design · ICLR 2025
Bioinformatics and computational biology › RNA biology › RNA analysis › RNA bioinformatics
RNA design
0.912025
gRNAde: Geometric Deep Learning for 3D RNA inverse design · ICLR 2025
Machine learning › Graph learning › graph neural network › geometric graph neural network
equivariant graph neural network
0.812024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Machine learning › Representation and self-supervised learning
pre-training
0.812024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Bioinformatics and computational biology › structural bioinformatics
protein structure
0.812024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Bioinformatics and computational biology › structural bioinformatics › protein structure representation
protein structure representation learning
0.812024
Evaluating Representation Learning on the Protein Structure Universe · ICLR 2024
Machine learning › Graph learning › graph neural network
expressive power
0.712023
On the Expressive Power of Geometric Graph Neural Networks · ICML 2023
Machine learning › Graph learning › graph neural network
geometric graph neural network
0.712023
On the Expressive Power of Geometric Graph Neural Networks · ICML 2023
Machine learning › Deep learning architectures and training
positional encoding
0.712023
Benchmarking Graph Neural Networks · J. Mach. Learn. Res. 2023
Graph algorithms and graph theory
graph isomorphism
0.712023
On the Expressive Power of Geometric Graph Neural Networks · ICML 2023
Graph algorithms and graph theory › graph isomorphism
weisfeiler-leman algorithm
0.712023
On the Expressive Power of Geometric Graph Neural Networks · ICML 2023
Empirical software engineering
reproducibility
0.212023
Benchmarking Graph Neural Networks · J. Mach. Learn. Res. 2023

Methods — techniques the papers use, named apart from their topics

transformer · 1.7multi-state GNN · 1.7latent diffusion · 1.7graph neural network · 1.7autoregressive decoding · 1.7autoencoder · 1.7pre-training · 1.5geometric graph neural networks · 1.5weisfeiler-leman test · 1.3universal approximation · 1.3model comparison · 1.3benchmarking · 1.3
YearPublicationVenuePosition
2025 gRNAde: Geometric Deep Learning for 3D RNA inverse design
abstract
Computational RNA design tasks are often posed as inverse problems, where sequences are designed based on adopting a single desired secondary structure without considering 3D conformational diversity. We introduce gRNAde, a geometric RNA design pipeline operating on 3D RNA backbones to design sequences that explicitly account for structure and dynamics. gRNAde uses a multi-state Graph Neural Network and autoregressive decoding to generates candidate RNA sequences conditioned on one or more 3D backbone structures where the identities of the bases are unknown. On a single-state fixed backbone re-design benchmark of 14 RNA structures from the PDB identified by Das et al. (2010), gRNAde obtains higher native sequence recovery rates (56% on average) compared to Rosetta (45% on average), taking under a second to produce designs compared to the reported hours for Rosetta. We further demonstrate the utility of gRNAde on a new benchmark of multi-state design for structurally flexible RNAs, as well as zero-shot ranking of mutational fitness landscapes in a retrospective analysis of a recent ribozyme. Experimental wet lab validation on 10 different structured RNA backbones finds that gRNAde has a success rate of 50% at designing pseudoknotted RNA structures, a significant advance over 35% for Rosetta. Open source code and tutorials are available at: github.com/chaitjo/geometric-rna-design
Chaitanya K. Joshi, Arian Rokkum Jamasb, Ramón Viñas 0001, Charles Harris, Simon V. Mathis, Alex Morehead, Rishabh Anand, Pietro Liò
ICLR1
2025 All-atom Diffusion Transformers: Unified generative modelling of molecules and materials
abstract
Diffusion models are the standard toolkit for generative modelling of 3D atomic systems. However, for different types of atomic systems -- such as molecules and materials -- the generative processes are usually highly specific to the target system despite the underlying physics being the same. We introduce the All-atom Diffusion Transformer (ADiT), a unified latent diffusion framework for jointly generating both periodic materials and non-periodic molecular systems using the same model: (1) An autoencoder maps a unified, all-atom representations of molecules and materials to a shared latent embedding space; and (2) A diffusion model is trained to generate new latent embeddings that the autoencoder can decode to sample new molecules or materials. Experiments on MP20, QM9 and GEOM-DRUGS datasets demonstrate that jointly trained ADiT generates realistic and valid molecules as well as materials, obtaining state-of-the-art results on par with molecule and crystal-specific models. ADiT uses standard Transformers with minimal inductive biases for both the autoencoder and diffusion model, resulting in significant speedups during training and inference compared to equivariant diffusion models. Scaling ADiT up to half a billion parameters predictably improves performance, representing a step towards broadly generalizable foundation models for generative chemistry. Open source code: https://github.com/facebookresearch/all-atom-diffusion-transformer
Chaitanya K. Joshi, Xiang Fu 0005, Yi-Lun Liao, Vahe Gharakhanyan, Benjamin Kurt Miller, Anuroop Sriram, Zachary W. Ulissi
ICML1
2024 Evaluating Representation Learning on the Protein Structure Universe
abstract
We introduce ProteinWorkshop, a comprehensive benchmark suite for representation learning on protein structures with Geometric Graph Neural Networks. We consider large-scale pre-training and downstream tasks on both experimental and predicted structures to enable the systematic evaluation of the quality of the learned structural representation and their usefulness in capturing functional relationships for downstream tasks. We find that: (1) large-scale pretraining on AlphaFold structures and auxiliary tasks consistently improve the performance of both rotation-invariant and equivariant GNNs, and (2) more expressive equivariant GNNs benefit from pretraining to a greater extent compared to invariant models. We aim to establish a common ground for the machine learning and computational biology communities to rigorously compare and advance protein structure representation learning. Our open-source codebase reduces the barrier to entry for working with large protein structure datasets by providing: (1) storage-efficient dataloaders for large-scale structural databases including AlphaFoldDB and ESM Atlas, as well as (2) utilities for constructing new tasks from the entire PDB. ProteinWorkshop is available at: github.com/a-r-j/ProteinWorkshop.
Arian Rokkum Jamasb, Alex Morehead, Chaitanya K. Joshi, Zuobai Zhang, Kieran Didi, Simon V. Mathis, Charles Harris, Jian Tang 0005, Jianlin Cheng, Pietro Liò, Tom L. Blundell
ICLR3
2024 On Representation Knowledge Distillation for Graph Neural Networks
abstract
Knowledge distillation (KD) is a learning paradigm for boosting resource-efficient graph neural networks (GNNs) using more expressive yet cumbersome teacher models. Past work on distillation for GNNs proposed the local structure preserving (LSP) loss, which matches local structural relationships defined over edges across the student and teacher's node embeddings. This article studies whether preserving the global topology of how the teacher embeds graph data can be a more effective distillation objective for GNNs, as real-world graphs often contain latent interactions and noisy edges. We propose graph contrastive representation distillation (G-CRD), which uses contrastive learning to implicitly preserve global topology by aligning the student node embeddings to those of the teacher in a shared representation space. Additionally, we introduce an expanded set of benchmarks on large-scale real-world datasets where the performance gap between teacher and student GNNs is non-negligible. Experiments across four datasets and 14 heterogeneous GNN architectures show that G-CRD consistently boosts the performance and robustness of lightweight GNNs, outperforming LSP (and a global structure preserving (GSP) variant of LSP) as well as baselines from 2-D computer vision. An analysis of the representational similarity among teacher and student embedding spaces reveals that G-CRD balances preserving local and global relationships, while structure preserving approaches are best at preserving one or the other.
Chaitanya K. Joshi, Fayao Liu, Xu Xun, Jie Lin 0001, Chuan-Sheng Foo
IEEE Trans. Neural Networks Learn. Syst.1
2023 On the Expressive Power of Geometric Graph Neural Networks
abstract
The expressive power of Graph Neural Networks (GNNs) has been studied extensively through the Weisfeiler-Leman (WL) graph isomorphism test. However, standard GNNs and the WL framework are inapplicable for geometric graphs embedded in Euclidean space, such as biomolecules, materials, and other physical systems. In this work, we propose a geometric version of the WL test (GWL) for discriminating geometric graphs while respecting the underlying physical symmetries: permutations, rotation, reflection, and translation. We use GWL to characterise the expressive power of geometric GNNs that are invariant or equivariant to physical symmetries in terms of distinguishing geometric graphs. GWL unpacks how key design choices influence geometric GNN expressivity: (1) Invariant layers have limited expressivity as they cannot distinguish one-hop identical geometric graphs; (2) Equivariant layers distinguish a larger class of graphs by propagating geometric information beyond local neighbourhoods; (3) Higher order tensors and scalarisation enable maximally powerful geometric GNNs; and (4) GWL's discrimination-based perspective is equivalent to universal approximation. Synthetic experiments supplementing our results are available at https://github.com/chaitjo/geometric-gnn-dojo
Chaitanya K. Joshi, Cristian Bodnar, Simon V. Mathis, Taco Cohen, Pietro Liò
ICML1
2023 Benchmarking Graph Neural Networks
abstract
In the last few years, graph neural networks (GNNs) have become the standard toolkit for analyzing and learning from data on graphs. This emerging field has witnessed an extensive growth of promising techniques that have been applied with success to computer science, mathematics, biology, physics and chemistry. But for any successful field to become mainstream and reliable, benchmarks must be developed to quantify progress. This led us in March 2020 to release a benchmark framework that i) comprises of a diverse collection of mathematical and real-world graphs, ii) enables fair model comparison with the same parameter budget to identify key architectures, iii) has an open-source, easy-to use and reproducible code infrastructure, and iv) is flexible for researchers to experiment with new theoretical ideas. As of December 2022, the GitHub repository has reached 2,000 stars and 380 forks, which demonstrates the utility of the proposed open-source framework through the wide usage by the GNN community. In this paper, we present an updated version of our benchmark with a concise presentation of the aforementioned framework characteristics, an additional medium-sized molecular dataset AQSOL, similar to the popular ZINC, but with a real-world measured chemical target, and discuss how this framework can be leveraged to explore new GNN designs and insights. As a proof of value of our benchmark, we study the case of graph positional encoding (PE) in GNNs, which was introduced with this benchmark and has since spurred interest of exploring more powerful PE for Transformers and GNNs in a robust experimental setting.
Vijay Prakash Dwivedi, Chaitanya K. Joshi, Anh Tuan Luu, Thomas Laurent 0001, Yoshua Bengio, Xavier Bresson
J. Mach. Learn. Res.2
2022 Point Discriminative Learning for Data-efficient 3D Point Cloud Analysis
abstract
3D point cloud analysis has drawn a lot of research attention due to its wide applications. However, collecting massive labelled 3D point cloud data is both time-consuming and labor-intensive. This calls for data-efficient learning methods. In this work we propose PointDisc, a point discriminative learning method to leverage self-supervisions for data-efficient 3D point cloud classification and segmentation. PointDisc imposes a novel point discrimination loss on the middle and global level features produced by the backbone network. This point discrimination loss enforces learned features to be consistent with points belonging to the corresponding local shape region and inconsistent with randomly sampled noisy points. We conduct extensive experiments on 3D object classification, 3D semantic and part segmentation, showing the benefits of PointDisc for data-efficient learning. Detailed analysis demonstrate that PointDisc learns unsupervised features that well capture local and global geometry.
Fayao Liu, Guosheng Lin, Chuan-Sheng Foo, Chaitanya K. Joshi, Jie Lin 0001
3DV4
2022 Multigraph Transformer for Free-Hand Sketch Recognition
abstract
Learning meaningful representations of free-hand sketches remains a challenging task given the signal sparsity and the high-level abstraction of sketches. Existing techniques have focused on exploiting either the static nature of sketches with convolutional neural networks (CNNs) or the temporal sequential property with recurrent neural networks (RNNs). In this work, we propose a new representation of sketches as multiple sparsely connected graphs. We design a novel graph neural network (GNN), the multigraph transformer (MGT), for learning representations of sketches from multiple graphs, which simultaneously capture global and local geometric stroke structures as well as temporal information. We report extensive numerical experiments on a sketch recognition task to demonstrate the performance of the proposed approach. Particularly, MGT applied on 414k sketches from Google QuickDraw: 1) achieves a small recognition gap to the CNN-based performance upper bound (72.80% versus 74.22%) and infers faster than the CNN competitors and 2) outperforms all RNN-based models by a significant margin. To the best of our knowledge, this is the first work proposing to represent sketches as graphs and apply GNNs for sketch recognition. Code and trained models are available at https://github.com/PengBoXiangShang/multigraph_transformer.
Peng Xu 0005, Chaitanya K. Joshi, Xavier Bresson
IEEE Trans. Neural Networks Learn. Syst.2
2021 Learning TSP Requires Rethinking Generalization
abstract
End-to-end training of neural network solvers for combinatorial optimization problems such as the Travelling Salesman Problem is intractable and inefficient beyond a few hundreds of nodes. While state-of-the-art Machine Learning approaches perform closely to classical solvers when trained on trivially small sizes, they are unable to generalize the learnt policy to larger instances of practical scales. Towards leveraging transfer learning to solve large-scale TSPs, this paper identifies inductive biases, model architectures and learning algorithms that promote generalization to instances larger than those seen in training. Our controlled experiments provide the first principled investigation into such zero-shot generalization, revealing that extrapolating beyond training data requires rethinking the neural combinatorial optimization pipeline, from network layers and learning paradigms to evaluation protocols.
Chaitanya K. Joshi, Quentin Cappart, Louis-Martin Rousseau, Thomas Laurent 0001
CP1