Seung-Cheol Lee

dblp:76/8410 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-8995-7463ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 79% Graph learning · 15% Language models and text generation · 6%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Computational science and engineering · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.722025
LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation · NeurIPS 2025
Periodic Materials Generation using Text-Guided Joint Diffusion Model · ICLR 2025
Computational science and engineering
materials science
1.122025
LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation · NeurIPS 2025
CrysGNN: Distilling Pre-trained Knowledge to Enhance Property Prediction for Crystalline Materials · AAAI 2023
Machine learning › Generative modeling › diffusion model
periodic material generation
0.912025
Periodic Materials Generation using Text-Guided Joint Diffusion Model · ICLR 2025
Machine learning › Generative modeling › diffusion model › conditional diffusion model
text-guided diffusion model
0.912025
Periodic Materials Generation using Text-Guided Joint Diffusion Model · ICLR 2025
Computational science and engineering › materials science › materials discovery
crystal structure generation
0.912025
LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation · NeurIPS 2025
Machine learning › Graph learning
graph neural network
0.712023
CrysGNN: Distilling Pre-trained Knowledge to Enhance Property Prediction for Crystalline Materials · AAAI 2023
Natural language and speech › Language models and text generation
large language model
0.312025
LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation · NeurIPS 2025
Computational science and engineering › materials informatics
crystal property prediction
0.212023
CrysGNN: Distilling Pre-trained Knowledge to Enhance Property Prediction for Crystalline Materials · AAAI 2023

Methods — techniques the papers use, named apart from their topics

large language model · 1.7fine-tuning · 1.7equivariant diffusion · 1.7pre-training · 1.3knowledge distillation · 1.3joint diffusion · 0.9equivariant graph neural network · 0.9
YearPublicationVenuePosition
2025 Periodic Materials Generation using Text-Guided Joint Diffusion Model
abstract
Equivariant diffusion models have emerged as the prevailing approach for generat- ing novel crystal materials due to their ability to leverage the physical symmetries of periodic material structures. However, current models do not effectively learn the joint distribution of atom types, fractional coordinates, and lattice structure of the crystal material in a cohesive end-to-end diffusion framework. Also, none of these models work under realistic setups, where users specify the desired characteristics that the generated structures must match. In this work, we introduce TGDMat, a novel text-guided diffusion model designed for 3D periodic material generation. Our approach integrates global structural knowledge through textual descriptions at each denoising step while jointly generating atom coordinates, types, and lattice structure using a periodic-E(3)-equivariant graph neural network (GNN). Extensive experiments using popular datasets on benchmark tasks reveal that TGDMat out- performs existing baseline methods by a good margin. Notably, for the structure prediction task, with just one generated sample, TGDMat outperforms all baseline models, highlighting the importance of text-guided diffusion. Further, in the genera- tion task, TGDMat surpasses all baselines and their text-fusion variants, showcasing the effectiveness of the joint diffusion paradigm. Additionally, incorporating textual knowledge reduces overall training and sampling computational overhead while enhancing generative performance when utilizing real-world textual prompts from experts. Code is available at https://github.com/kdmsit/TGDMat
Kishalay Das, Subhojyoti Khastagir, Pawan Goyal 0002, Seung-Cheol Lee, Satadeep Bhattacharjee, Niloy Ganguly
ICLR4
2025 LLM Meets Diffusion: A Hybrid Framework for Crystal Material Generation
abstract
Recent advances in generative modeling have shown significant promise in designing novel periodic crystal structures. Existing approaches typically rely on either large language models (LLMs) or equivariant denoising models, each with complementary strengths: LLMs excel at handling discrete atomic types but often struggle with continuous features such as atomic positions and lattice parameters, while denoising models are effective at modeling continuous variables but encounter difficulties in generating accurate atomic compositions. To bridge this gap, we propose CrysLLMGen, a hybrid framework that integrates an LLM with a diffusion model to leverage their complementary strengths for crystal material generation. During sampling, CrysLLMGen first employs a fine-tuned LLM to produce an intermediate representation of atom types, atomic coordinates, and lattice structure. While retaining the predicted atom types, it passes the atomic coordinates and lattice structure to a pre-trained equivariant diffusion model for refinement. Our framework outperforms state-of-the-art generative models across several benchmark tasks and datasets. Specifically, CrysLLMGen not only achieves a balanced performance in terms of structural and compositional validity but also generates more stable and novel materials compared to LLM-based and denoising-based models Furthermore, CrysLLMGen exhibits strong conditional generation capabilities, effectively producing materials that satisfy user-defined constraints. Code is available at \url{https://github.com/kdmsit/crysllmgen}
Subhojyoti Khastagir, Kishalay Das, Pawan Goyal 0002, Seung-Cheol Lee, Satadeep Bhattacharjee, Niloy Ganguly
NeurIPS4
2024 Determining the best feature combination through text and probabilistic feature analysis for GPT-2-based mobile app review detection
abstract
Abstract Mobile apps, used by many people worldwide, have become an essential part of life. Before using a mobile app, users judge the reliability of apps according to their reviews. Therefore, app reviews are essential components of management for companies. Unfortunately, some fake reviewers write negative reviews for competing apps. Moreover, artificial intelligence (AI)-based macro bot programs that generate app reviews have emerged and can create large numbers of reviews with malicious purposes in a short time. One notable AI technology that can generate such reviews is Generative Pre-trained Transformer-2 (GPT-2). The reviews generated by GPT-2 use human-like grammar; therefore, it is difficult to detect them with only text mining techniques, which use tools like part-of-speech (POS) tagging and sentiment scores. Thus, probability-based sampling techniques in GPT-2 must be used. In this study, we identified features to detect reviews generated by GPT-2 and determined the optimal feature combination for improving detection performance. To achieve this, based on the analysis results, we built a training dataset to find the best feature combination for detecting the generated reviews. Various machine learning models were then trained and evaluated using this dataset. As a result, the model that used both text mining and probability-based sampling techniques detected generated reviews more effectively than the model that used only text mining techniques. This model achieved a top classification accuracy of 90% and a macro F1 of 0.90. We expect the results of this study to help app developers maintain a more stable mobile app ecosystem. Graphical abstract
Seung-Cheol Lee, Dong-Gun Lee, Yeong-Seok Seo
Appl. Intell.1
2023 CrysGNN: Distilling Pre-trained Knowledge to Enhance Property Prediction for Crystalline Materials
abstract
In recent years, graph neural network (GNN) based approaches have emerged as a powerful technique to encode complex topological structure of crystal materials in an enriched repre- sentation space. These models are often supervised in nature and using the property-specific training data, learn relation- ship between crystal structure and different properties like formation energy, bandgap, bulk modulus, etc. Most of these methods require a huge amount of property-tagged data to train the system which may not be available for different prop- erties. However, there is an availability of a huge amount of crystal data with its chemical composition and structural bonds. To leverage these untapped data, this paper presents CrysGNN, a new pre-trained GNN framework for crystalline materials, which captures both node and graph level structural information of crystal graphs using a huge amount of unla- belled material data. Further, we extract distilled knowledge from CrysGNN and inject into different state of the art prop- erty predictors to enhance their property prediction accuracy. We conduct extensive experiments to show that with distilled knowledge from the pre-trained model, all the SOTA algo- rithms are able to outperform their own vanilla version with good margins. We also observe that the distillation process provides significant improvement over the conventional ap- proach of finetuning the pre-trained model. We will release the pre-trained model along with the large dataset of 800K crys- tal graph which we carefully curated; so that the pre-trained model can be plugged into any existing and upcoming models to enhance their prediction accuracy.
Kishalay Das, Bidisha Samanta, Pawan Goyal 0002, Seung-Cheol Lee, Satadeep Bhattacharjee, Niloy Ganguly
AAAI4
2023 CrysMMNet: Multimodal Representation for Crystal Property Prediction
abstract
Machine Learning models have emerged as a powerful tool for fast and accurate prediction of different crystalline properties. Exiting state-of-the-art models rely on a single modality of crystal data i.e crystal graph structure, where they construct multi-graph by establishing edges between nearby atoms in 3D space and apply GNN to learn materials representation. Thereby, they encode local chemical semantics around the atoms successfully but fail to capture important global periodic structural information like space group number, crystal symmetry, rotational information etc, which influence different crystal properties. In this work, we leverage textual descriptions of materials to model global structural information into graph structure and learn a more robust and enriched representation of crystalline materials. To this effect, we first curate a textual dataset for crystalline material databases containing descriptions of each material. Further, we propose CrysMMNet, a simple multi-modal framework, which fuses both structural and textual representation together to generate a joint multimodal representation of crystalline materials. We conduct extensive experiments on two benchmark datasets across ten different properties to show that CrysMMNet outperforms existing state-of-the-art baseline methods with a good margin. We also observe that fusing the textual representation with crystal graph structure provides consistent improvement for all the SOTA GNN models compared to their own vanilla versions. We have shared the textual dataset, that we have curated for both the benchmark material databases, with the community for future use..
Kishalay Das, Pawan Goyal 0002, Seung-Cheol Lee, Satadeep Bhattacharjee, Niloy Ganguly
UAI3
2022 Using Sentence-level Classification Helps Entity Extraction from Material Science Literature
abstract
In the last few years, several attempts have been made on extracting information from material science research domain. Material Science research articles are a rich source of information about various entities related to material science such as names of the materials used for experiments, the computational software used along with its parameters, the method used in the experiments, etc. But the distribution of these entities is not uniform across different sections of research articles. Most of the sentences in the research articles do not contain any entity. In this work, we first use a sentence-level classifier to identify sentences containing at least one entity mention. Next, we apply the information extraction models only on the filtered sentences, to extract various entities of interest. Our experiments for named entity recognition in the material science research articles show that this additional sentence-level classification step helps to improve the F1 score by more than 4%.
Ankan Mullick, Shubhraneel Pal, Tapas Nayak, Seung-Cheol Lee, Satadeep Bhattacharjee, Pawan Goyal 0002
LREC4
2020 Query-Based Video Synopsis for Intelligent Traffic Monitoring Applications
abstract
Synopsis of a long-duration video has many applications in intelligent transportation systems. It can help to monitor traffic with lesser manpower. However, generating meaningful synopsis of a long-duration video recording can be challenging. Often summarized outputs include redundant contents or activities that may not be helpful to the observer. Moving object trajectories are possible sources of information that can be used to generate the synopsis of long-duration videos. The synopsis generation faces challenges due to object tracking, grouping of the trajectories with respect to activity type, object category, and contextual information, and generating smooth synopsis according to a query. In this paper, we propose a method to generate meaningful and smooth synopsis of long-duration videos according to the users' query. We have tracked moving objects and adopted deep learning to classify the objects into known categories (e.g., car, bike, and pedestrians). We then identify regions in the surveillance scene with the help of unsupervised clustering. Each tube (spatiotemporal object trajectory) is represented by the source and the destination. In the final stage, we take a query from the user and generate the synopsis video by smoothly blending the appropriate tubes over the background frame through energy minimization. The proposed method has been evaluated on two publicly available datasets and our own surveillance datasets. We have compared the method with popular state-of-the-art techniques. The experiments reveal that the proposed method is superior to the existing techniques and it produces visually seamless video synopsis.
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Renuka Patnaik, Seung-Cheol Lee, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim
IEEE Trans. Intell. Transp. Syst.5