Ke Liu 0012

dblp:32/2948-12 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0001-9698-817XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 A Denoising Pre-training Framework for Accelerating Novel Material Discovery
abstract
Crystal materials play an important role in the development of society. The discovery of new materials is critical to achieving sustainable development goals (SDGs), such as climate change mitigation, affordable and clean energy, and fostering innovation in industry and infrastructure. Recent advances in deep learning for crystal property prediction have accelerated material discovery, but these methods typically rely on labeled data, which is often limited and varies across different properties. This limitation hinders the full utilization of the vast amount of unlabeled data in materials science. To overcome this challenge, we introduce an unsupervised Denoising Pre-training Framework (DPF) tailored for crystal structures. DPF trains a model to reconstruct the original crystal structure by recovering the masked atom types, perturbed atom positions, and perturbed crystal lattices. Through pre-training, models learn the intrinsic features of crystal structures and capture the key features influencing crystal properties. We pre-train models on a dataset of 380,743 unlabeled crystal structures and fine-tune them on downstream property prediction tasks. Extensive experiments demonstrate the effectiveness of our framework, showing its potential to significantly advance material science and contribute to the development of society by accelerating the discovery of materials crucial for sustainable technologies.
Shuaike Shen, Ke Liu 0012, Muzhi Zhu, Hao Chen 0041
AAAI2
2025 A Universal Periodicity Injection Module for Crystal Property Prediction
Yichao Fu, Ke Liu 0012, Shangde Gao, Te Qiao
ICIC (26)2
2025 Physics Aware Neural Networks for Unsupervised Binding Energy Prediction
abstract
Developing models for protein-ligand interactions holds substantial significance for drug discovery. Supervised methods often failed due to the lack of labeled data for predicting the protein-ligand binding energy, like antibodies. Therefore, unsupervised approaches are urged to make full use of the unlabeled data. To tackle the problem, we propose an efficient, unsupervised protein-ligand binding energy prediction model via the conservation of energy (CEBind), which follows the physical laws. Specifically, given a protein-ligand complex, we randomly sample forces for each atom in the ligand. Then these forces are applied rigidly to the ligand to perturb its position, following the law of rigid body dynamics. Finally, CEBind predicts the energy of both the unperturbed complex and the perturbed complex. The energy gap between two complexes equals the work of the outer forces, following the law of conservation of energy. Extensive experiments are conducted on the unsupervised protein-ligand binding energy prediction benchmarks, comparing them with previous works. Empirical results and theoretic analysis demonstrate that CEBind is more efficient and outperforms previous unsupervised models on benchmarks.
Ke Liu 0012, Hao Cheng 0012, Chunhua Shen
ICML1
2025 Mat-Instructions: A Large-Scale Inorganic Material Instruction Dataset for Large Language Models
abstract
Recent advancements in large language models (LLMs) have revolutionized research discovery across various scientific disciplines, including materials science. The discovery of novel materials, particularly crystal materials, is essential for achieving sustainable development goals (SDGs), as they drive breakthroughs in climate change mitigation, clean and affordable energy, and the promotion of industrial innovation. However, unlocking the full potential of LLMs in materials research remains challenging due to the lack of high-quality, diverse, and instruction-based datasets. Such datasets are crucial for guiding these models in understanding and predicting the structure, property, and function of materials across various tasks. To address this limitation, we introduce Mat-Instruction, a large-scale inorganic material instruction dataset, specifically designed to unlock the potential of LLMs in materials science. Extensive experiments on fine-tuning LLaMA with our Mat-Instruction dataset demonstrate its effectiveness in advancing progress for materials science. The code and dataset are available at https://github.com/zjuKeLiu/Mat-Instructions
Ke Liu 0012, Shangde Gao, Yichao Fu, Xiaoliang Wu 0003, Shuo Tong, Ajitha Rajan
IJCAI1
2025 Matrix Factorization with Dynamic Multi-view Clustering for Recommender System
abstract
Matrix factorization (MF), a cornerstone of recommender systems, decomposes user-item interaction matrices into latent representations. Traditional MF approaches, however, employ a two-stage, non-end-to-end paradigm, sequentially performing recommendation and clustering, resulting in prohibitive computational costs for large-scale applications like e-commerce and IoT, where billions of users interact with trillions of items. To address this, we propose Matrix Factorization with Dynamic Multi-view Clustering (MFDMC), a unified framework that balances efficient end-to-end training with comprehensive utilization of web-scale data and enhances interpretability. MFDMC leverages dynamic multi-view clustering to learn user and item representations, adaptively pruning poorly formed clusters. Each entity's representation is modeled as a weighted projection of robust clusters, capturing its diverse roles across views. This design maximizes representation space utilization, improves interpretability, and ensures resilience for downstream tasks. Extensive experiments demonstrate MFDMC's superior performance in recommender systems and other representation learning domains, such as computer vision, highlighting its scalability and versatility.
Shangde Gao, Ke Liu 0012, Yichao Fu, Jian Wu 0001
IJCNN2
2025 Uncertainty-Aware Multi-expert Knowledge Distillation for Imbalanced Disease Grading
Shuo Tong, Shangde Gao, Ke Liu 0012, Haochao Ying, Jian Wu 0001
MICCAI (13)3
2025 Towards Generalizable Retina Vessel Segmentation with Deformable Graph Priors
abstract
Retinal vessel segmentation is critical for medical diagnosis, yet existing models often struggle to generalize across domains due to appearance variability, limited annotations, and complex vascular morphology. We propose GraphSeg, a variational Bayesian framework that integrates anatomical graph priors with structure-aware image decomposition to enhance cross-domain segmentation. GraphSeg factorizes retinal images into structure-preserved and structure-degraded components, enabling domain-invariant representation. A deformable graph prior, derived from a statistical retinal atlas, is incorporated via a differentiable alignment and guided by an unsupervised energy function. Experiments on three public benchmarks (CHASE, DRIVE, HRF) show that GraphSeg consistently outperforms existing methods under domain shifts. These results highlight the importance of jointly modeling anatomical topology and image structure for robust generalizable vessel segmentation.
Ke Liu 0012, Shangde Gao, Yichao Fu, Shangqi Gao
NeurIPS1
2024 Floating Anchor Diffusion Model for Multi-motif Scaffolding
abstract
Motif scaffolding seeks to design scaffold structures for constructing proteins with functions derived from the desired motif, which is crucial for the design of vaccines and enzymes. Previous works approach the problem by inpainting or conditional generation. Both of them can only scaffold motifs with fixed positions, and the conditional generation cannot guarantee the presence of motifs. However, prior knowledge of the relative motif positions in a protein is not readily available, and constructing a protein with multiple functions in one protein is more general and significant because of the synergies between functions. We propose a Floating Anchor Diffusion (FADiff) model. FADiff allows motifs to float rigidly and independently in the process of diffusion, which guarantees the presence of motifs and automates the motif position design. Our experiments demonstrate the efficacy of FADiff with high success rates and designable novel scaffolds. To the best of our knowledge, FADiff is the first work to tackle the challenge of scaffolding multiple motifs without relying on the expertise of relative motif positions in the protein. Code is available at https://github.com/aim-uofa/FADiff.
Ke Liu 0012, Weian Mao, Shuaike Shen, Xiaoran Jiao, Hao Cheng 0012, Chunhua Shen
ICML1
2024 Collaborative knowledge amalgamation: Preserving discriminability and transferability in unsupervised learning
Shangde Gao, Yichao Fu, Ke Liu 0012, Wei Gao 0001, Jian Wu 0001, Yuqiang Han
Inf. Sci.3
2024 A periodicity aware transformer for crystal property prediction
Ke Liu 0012, Kaifan Yang, Shangde Gao
Neural Comput. Appl.1
2023 Contrastive Knowledge Amalgamation for Unsupervised Image Classification
Shangde Gao, Yichao Fu, Ke Liu 0012, Yuqiang Han
ICANN (2)3
2023 PCVAE: A Physics-informed Neural Network for Determining the Symmetry and Geometry of Crystals
abstract
The symmetry and geometry of a crystal fundamentally determine its various physical and chemical properties. However, crystal structure prediction, including space group determination and crystal structure optimization, remains an ongoing challenge because traditional DFT-based approaches are time- and computational-intensive even for one specific set of material, not to mention structure prediction of massive materials. This paper determines the geometric structure of massive crystals solely from chemical formulae from scratch. In addition, due to various phases or changing environmental conditions, different pressure and temperature for example, there could be multiple crystal structures corresponding to one chemical formula (MS4OF), which has been overlooked or poorly addressed in previous research. Hereby, we propose a Physics-informed Conditional Variational Auto Encoder (PCVAE) to encode possible symmetry and geometry distribution as well as various phases of a crystal with Gaussian distributions. PCVAE achieves a new state-of-the-art in crystal structure prediction. Extensive experiments demonstrate the strong predictive power of PCVAE. The code and datasets are available at https://github.com/zjuKeLiu/PCVAE.
Ke Liu 0012, Shangde Gao, Kaifan Yang, Yuqiang Han
IJCNN1
2023 Empowering General-purpose User Representation with Full-life Cycle Behavior Modeling
Bei Yang, Ke Liu 0012, Renjun Xu, Qinghui Sun
KDD3
2023 E(2)-Equivariant Vision Transformer
abstract
Vision Transformer (ViT) has achieved remarkable performance in computer vision. However, positional encoding in ViT makes it substantially difficult to learn the intrinsic equivariance in data. Ini- tial attempts have been made on designing equiv- ariant ViT but are proved defective in some cases in this paper. To address this issue, we design a Group Equivariant Vision Transformer (GE-ViT) via a novel, effective positional encoding opera- tor. We prove that GE-ViT meets all the theoreti- cal requirements of an equivariant neural network. Comprehensive experiments are conducted on standard benchmark datasets, demonstrating that GE-ViT significantly outperforms non-equivariant self-attention networks. The code is available at https://github.com/ZJUCDSYangKaifan/GEVit.
Renjun Xu, Kaifan Yang, Ke Liu 0012, Fengxiang He
UAI3
2022 S2SNet: A Pretrained Neural Network for Superconductivity Discovery
abstract
Superconductivity allows electrical current to flow without any energy loss, and thus making solids superconducting is a grand goal of physics, material science, and electrical engineering. More than 16 Nobel Laureates have been awarded for their contribution in superconductivity research. Superconductors are valuable for sustainable development goals (SDGs), such as climate change mitigation, affordable and clean energy, industry, innovation and infrastructure, and so on. However, a unified physics theory explaining all superconductivity mechanism is still unknown. It is believed that superconductivity is microscopically due to not only molecular compositions but also the geometric crystal structure. Hence a new dataset, S2S, containing both crystal structures and superconducting critical temperature, is built upon SuperCon and Material Project. Based on this new dataset, we propose a novel model, S2SNet, which utilizes the attention mechanism for superconductivity prediction. To overcome the shortage of data, S2SNet is pre-trained on the whole Material Project dataset with Masked-Language Modeling (MLM). S2SNet makes a new state-of-the-art, with out-of-sample accuracy of 92% and Area Under Curve (AUC) of 0.92. To the best of our knowledge, S2SNet is the first work to predict superconductivity with only information of crystal structures. This work is beneficial to superconductivity discovery and further SDGs. The code and datasets are available at https://github.com/supercond/S2SNet
Ke Liu 0012, Kaifan Yang, Jiahong Zhang, Renjun Xu
IJCAI1
2022 Learning Interest-oriented Universal User Representation via Self-supervision
abstract
User representation is essential for providing high-quality commercial services in industry. In our business scenarios, we face the challenge of learning universal (general-purpose) user representation. The universal representation is expected to be informative, and can handle various types of real-world applications without fine-tuning (e.g., applicable for both user profiling and the recall process in advertising). It shows great advantages compared to the solution of training a specific model for each downstream application. Specifically, we attempt to improve universal user representation from two points of views. First, a contrastive self-supervised learning paradigm is presented to guide the representation model training. It provides a unified framework that allows for long-term or short-term interest representation learning in a data-driven manner. Moreover, a novel multi-interest extraction module is presented. The module introduces an interest dictionary to capture principal interests of the given user, and then generate his/her interest-oriented representations via behavior aggregation. Experimental results demonstrate the effectiveness and applicability of the learned user representations. Such an industrial solution has now been deployed in various real-world tasks.
Qinghui Sun, Renjun Xu, Ke Liu 0012, Bei Yang
ACM Multimedia5