Kishalay Das

dblp:258/3218 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 79% Graph learning · 15% Language models and text generation · 6%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.722025
LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation · NeurIPS 2025
Periodic Materials Generation using Text-Guided Joint Diffusion Model · ICLR 2025
Computational science and engineering
materials science
1.122025
LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation · NeurIPS 2025
CrysGNN: Distilling Pre-trained Knowledge to Enhance Property Prediction for Crystalline Materials · AAAI 2023
Machine learning › Generative modeling › diffusion model
periodic material generation
0.912025
Periodic Materials Generation using Text-Guided Joint Diffusion Model · ICLR 2025
Machine learning › Generative modeling › diffusion model › conditional diffusion model
text-guided diffusion model
0.912025
Periodic Materials Generation using Text-Guided Joint Diffusion Model · ICLR 2025
Computational science and engineering › materials science › materials discovery
crystal structure generation
0.912025
LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation · NeurIPS 2025
Machine learning › Graph learning
graph neural network
0.712023
CrysGNN: Distilling Pre-trained Knowledge to Enhance Property Prediction for Crystalline Materials · AAAI 2023
Natural language and speech › Language models and text generation
large language model
0.312025
LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation · NeurIPS 2025
Computational science and engineering › materials informatics
crystal property prediction
0.212023
CrysGNN: Distilling Pre-trained Knowledge to Enhance Property Prediction for Crystalline Materials · AAAI 2023

Methods — techniques the papers use, named apart from their topics

large language model · 1.7fine-tuning · 1.7equivariant diffusion · 1.7pre-training · 1.3knowledge distillation · 1.3joint diffusion · 0.9equivariant graph neural network · 0.9
YearPublicationVenuePosition
2025 Periodic Materials Generation using Text-Guided Joint Diffusion Model
abstract
Equivariant diffusion models have emerged as the prevailing approach for generat- ing novel crystal materials due to their ability to leverage the physical symmetries of periodic material structures. However, current models do not effectively learn the joint distribution of atom types, fractional coordinates, and lattice structure of the crystal material in a cohesive end-to-end diffusion framework. Also, none of these models work under realistic setups, where users specify the desired characteristics that the generated structures must match. In this work, we introduce TGDMat, a novel text-guided diffusion model designed for 3D periodic material generation. Our approach integrates global structural knowledge through textual descriptions at each denoising step while jointly generating atom coordinates, types, and lattice structure using a periodic-E(3)-equivariant graph neural network (GNN). Extensive experiments using popular datasets on benchmark tasks reveal that TGDMat out- performs existing baseline methods by a good margin. Notably, for the structure prediction task, with just one generated sample, TGDMat outperforms all baseline models, highlighting the importance of text-guided diffusion. Further, in the genera- tion task, TGDMat surpasses all baselines and their text-fusion variants, showcasing the effectiveness of the joint diffusion paradigm. Additionally, incorporating textual knowledge reduces overall training and sampling computational overhead while enhancing generative performance when utilizing real-world textual prompts from experts. Code is available at https://github.com/kdmsit/TGDMat
Kishalay Das, Subhojyoti Khastagir, Pawan Goyal 0002, Seung-Cheol Lee, Satadeep Bhattacharjee, Niloy Ganguly
ICLR1
2025 LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation
abstract
Recent advances in generative modeling have shown significant promise in designing novel periodic crystal structures. Existing approaches typically rely on either large language models (LLMs) or equivariant denoising models, each with complementary strengths: LLMs excel at handling discrete atomic types but often struggle with continuous features such as atomic positions and lattice parameters, while denoising models are effective at modeling continuous variables but encounter difficulties in generating accurate atomic compositions. To bridge this gap, we propose CrysLLMGen, a hybrid framework that integrates an LLM with a diffusion model to leverage their complementary strengths for crystal material generation. During sampling, CrysLLMGen first employs a fine-tuned LLM to produce an intermediate representation of atom types, atomic coordinates, and lattice structure. While retaining the predicted atom types, it passes the atomic coordinates and lattice structure to a pre-trained equivariant diffusion model for refinement. Our framework outperforms state-of-the-art generative models across several benchmark tasks and datasets. Specifically, CrysLLMGen not only achieves a balanced performance in terms of structural and compositional validity but also generates more stable and novel materials compared to LLM-based and denoising-based models Furthermore, CrysLLMGen exhibits strong conditional generation capabilities, effectively producing materials that satisfy user-defined constraints. Code is available at \url{https://github.com/kdmsit/crysllmgen}
Subhojyoti Khastagir, Kishalay Das, Pawan Goyal 0002, Seung-Cheol Lee, Satadeep Bhattacharjee, Niloy Ganguly
NeurIPS2
2023 CrysGNN: Distilling Pre-trained Knowledge to Enhance Property Prediction for Crystalline Materials
abstract
In recent years, graph neural network (GNN) based approaches have emerged as a powerful technique to encode complex topological structure of crystal materials in an enriched repre- sentation space. These models are often supervised in nature and using the property-specific training data, learn relation- ship between crystal structure and different properties like formation energy, bandgap, bulk modulus, etc. Most of these methods require a huge amount of property-tagged data to train the system which may not be available for different prop- erties. However, there is an availability of a huge amount of crystal data with its chemical composition and structural bonds. To leverage these untapped data, this paper presents CrysGNN, a new pre-trained GNN framework for crystalline materials, which captures both node and graph level structural information of crystal graphs using a huge amount of unla- belled material data. Further, we extract distilled knowledge from CrysGNN and inject into different state of the art prop- erty predictors to enhance their property prediction accuracy. We conduct extensive experiments to show that with distilled knowledge from the pre-trained model, all the SOTA algo- rithms are able to outperform their own vanilla version with good margins. We also observe that the distillation process provides significant improvement over the conventional ap- proach of finetuning the pre-trained model. We will release the pre-trained model along with the large dataset of 800K crys- tal graph which we carefully curated; so that the pre-trained model can be plugged into any existing and upcoming models to enhance their prediction accuracy.
Kishalay Das, Bidisha Samanta, Pawan Goyal 0002, Seung-Cheol Lee, Satadeep Bhattacharjee, Niloy Ganguly
AAAI1
2023 CrysMMNet: Multimodal Representation for Crystal Property Prediction
abstract
Machine Learning models have emerged as a powerful tool for fast and accurate prediction of different crystalline properties. Exiting state-of-the-art models rely on a single modality of crystal data i.e crystal graph structure, where they construct multi-graph by establishing edges between nearby atoms in 3D space and apply GNN to learn materials representation. Thereby, they encode local chemical semantics around the atoms successfully but fail to capture important global periodic structural information like space group number, crystal symmetry, rotational information etc, which influence different crystal properties. In this work, we leverage textual descriptions of materials to model global structural information into graph structure and learn a more robust and enriched representation of crystalline materials. To this effect, we first curate a textual dataset for crystalline material databases containing descriptions of each material. Further, we propose CrysMMNet, a simple multi-modal framework, which fuses both structural and textual representation together to generate a joint multimodal representation of crystalline materials. We conduct extensive experiments on two benchmark datasets across ten different properties to show that CrysMMNet outperforms existing state-of-the-art baseline methods with a good margin. We also observe that fusing the textual representation with crystal graph structure provides consistent improvement for all the SOTA GNN models compared to their own vanilla versions. We have shared the textual dataset, that we have curated for both the benchmark material databases, with the community for future use..
Kishalay Das, Pawan Goyal 0002, Seung-Cheol Lee, Satadeep Bhattacharjee, Niloy Ganguly
UAI1
2020 Hypergraph Attention Isomorphism Network by Learning Line Graph Expansion
abstract
Graph neural networks (GNNs) are able to achieve state-of-the-art performance for node representation and classification in a network. But, most of the existing GNNs can be applied to simple graphs, where an edge connects only a pair of nodes. Studies have shown that hypergraphs are effective to model real-world relationships which are of higher order in nature. Recently, graph neural networks are proposed for hypergraphs, but they implicitly use clique or star expansions to convert the hypergraph to a simple graph, or use computationally expensive hypergraph Laplacian.In this work, we propose a novel hypergraph neural network for semi-supervised hypernode classification, which operates directly on the hypergraphs with varying hyperedge sizes. Within each layer, it indirectly works on the line graph of the given hypergraph, without actually forming the line graph explicitly. Moreover, it also employs a self-attention mechanism to learn the weights of those edge relationships. Experimentally, HAIN is able to improve the state-of-the-art hypernode classification performance on all the datasets we use. We make the source code available to ease the reproducibility of the results.
Sambaran Bandyopadhyay, Kishalay Das, M. Narasimha Murty
IEEE BigData2