Yinjun Jia

dblp:333/7450 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
7 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
7 papers
Generative modeling · 50% Representation and self-supervised learning · 20% Graph learning · 12%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › structural bioinformatics
protein structure
1.722025
AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation · NeurIPS 2025
CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design · NeurIPS 2025
Bioinformatics and computational biology
drug discovery
1.522025
AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation · NeurIPS 2025
DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening · NeurIPS 2023
Bioinformatics and computational biology › drug discovery
virtual screening
1.522025
AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation · NeurIPS 2025
DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening · NeurIPS 2023
Machine learning › Representation and self-supervised learning
contrastive learning
1.022024
Self-supervised Pocket Pretraining via Protein Fragment-Surroundings Alignment · ICLR 2024
DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening · NeurIPS 2023
Machine learning › Generative modeling › generative model › continuous-time generative model
bayesian flow network
0.912025
Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent Space · NeurIPS 2025
Machine learning › Generative modeling
molecular generation
0.912025
Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent Space · NeurIPS 2025
Machine learning › Generative modeling
variational autoencoder
0.912025
Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent Space · NeurIPS 2025
Bioinformatics and computational biology › gene regulation
binding site prediction
0.912025
AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation · NeurIPS 2025
Bioinformatics and computational biology › drug discovery
bioactivity prediction
0.912025
Redefining the task of Bioactivity Prediction · ICLR 2025
Bioinformatics and computational biology
molecular property prediction
0.912025
Redefining the task of Bioactivity Prediction · ICLR 2025
Bioinformatics and computational biology
protein design
0.912025
CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.812024
Full-Atom Peptide Design with Geometric Latent Diffusion · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
latent diffusion model
0.812024
Full-Atom Peptide Design with Geometric Latent Diffusion · NeurIPS 2024
Machine learning › Representation and self-supervised learning › structured representation
molecular representation
0.812024
Self-supervised Pocket Pretraining via Protein Fragment-Surroundings Alignment · ICLR 2024
Machine learning › Graph learning
molecular representation learning
0.812024
Protein-ligand binding representation learning from fine-grained interactions · ICLR 2024
Bioinformatics and computational biology › protein design
peptide design
0.812024
Full-Atom Peptide Design with Geometric Latent Diffusion · NeurIPS 2024
Bioinformatics and computational biology › protein analysis
protein-ligand interaction
0.812024
Protein-ligand binding representation learning from fine-grained interactions · ICLR 2024
Computer vision › Vision and language › vision-language pretraining
contrastive vision-language pretraining
0.712023
DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening · NeurIPS 2023
Machine learning › Graph learning
protein structure prediction
0.312025
CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design · NeurIPS 2025
Machine learning › Generative modeling › molecular generation
molecular structure generation
0.212024
Full-Atom Peptide Design with Geometric Latent Diffusion · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

contrastive learning · 3.7structure filtering · 1.7machine learning · 1.7docking · 1.7dataset curation · 1.7variational autoencoder · 1.6pairwise distance prediction · 1.5masked modeling · 1.5e(3)-equivariant neural network · 0.9cross-attention adapter · 0.9bayesian flow network · 0.9transformer · 0.8fragment-surroundings alignment · 0.8
YearPublicationVenuePosition
2025 Redefining the task of Bioactivity Prediction
abstract
Small molecules are vital to modern medicine, and accurately predicting their bioactivity against protein targets is crucial for therapeutic discovery and development. However, current machine learning models often rely on spurious features, leading to biased outcomes. Notably, a simple pocket-only baseline can achieve results comparable to, and sometimes better than, more complex models that incorporate both the protein pockets and the small molecules. Our analysis reveals that this phenomenon arises from insufficient training data and an improper evaluation process, which is typically conducted at the pocket level rather than the small molecule level. To address these issues, we redefine the bioactivity prediction task by introducing the SIU dataset-a million-scale Structural small molecule-protein Interaction dataset for Unbiased bioactivity prediction task, which is 50 times larger than the widely used PDBbind. The bioactivity labels in SIU are derived from wet experiments and organized by label types, ensuring greater accuracy and comparability. The complexes in SIU are constructed using a majority vote from three commonly used docking software programs, enhancing their reliability. Additionally, the structure of SIU allows for multiple small molecules to be associated with each protein pocket, enabling the redefinition of evaluation metrics like Pearson and Spearman correlations across different small molecules targeting the same protein pocket. Experimental results demonstrate that this new task provides a more challenging and meaningful benchmark for training and evaluating bioactivity prediction models, ultimately offering a more robust assessment of model performance.
Yanwen Huang, Yinjun Jia, Hongbo Ma, Wei-Ying Ma, Ya-Qin Zhang, Yanyan Lan
ICLR3
2025 Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent Space
abstract
Medicinal chemists often optimize drugs considering their 3D structures and designing structurally distinct molecules that retain key features, such as shapes, pharmacophores, or chemical properties. Previous deep learning approaches address this through supervised tasks like molecule inpainting or property-guided optimization. In this work, we propose a flexible zero-shot molecule manipulation method by navigating in a shared latent space of 3D molecules. We introduce a Variational AutoEncoder (VAE) for 3D molecules, named MolFLAE, which learns a fixed-dimensional, E(3)-equivariant latent space independent of atom counts. MolFLAE encodes 3D molecules using an E(3)-equivariant neural network into fixed number of latent nodes, distinguished by learned embeddings. The latent space is regularized, and molecular structures are reconstructed via a Bayesian Flow Network (BFN) conditioned on the encoder’s latent output. MolFLAE achieves competitive performance on standard unconditional 3D molecule generation benchmarks. Moreover, the latent space of MolFLAE enables zero-shot molecule manipulation, including atom number editing, structure reconstruction, and coordinated latent interpolation for both structure and properties. We further demonstrate our approach on a drug optimization task for the human glucocorticoid receptor, generating molecules with improved hydrophilicity while preserving key interactions, under computational evaluations. These results highlight the flexibility, robustness, and real-world utility of our method, opening new avenues for molecule editing and optimization.
Yinjun Jia, Zitong Tian, Wei-Ying Ma, Yanyan Lan
NeurIPS2
2025 CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design
abstract
Cyclic peptides exhibit better binding affinity and proteolytic stability compared to their linear counterparts. However, the development of cyclic peptide design models is hindered by the scarcity of data. To address this, we introduce **CPSea**(**C**yclic **P**eptide **Sea**), a dataset of 2.71 million cyclic peptide-receptor complexes, curated through systematic mining of the AlphaFold Database (AFDB). Our pipeline extracts compact domains from AFDB, identifies cyclization sites using the $\beta$-carbon (C$_\beta$) distance thresholds, and applies multi-stage filtering to ensure structure fidelity and binding compatibility. Compared with experimental data of cyclic peptides, CPSea shows similar distributions in metrics on structure fidelity and wet-lab compatibility. To our knowledge, CPSea is the largest cyclic peptide-receptor dataset to date, enabling end-to-end model training for the first time. The dataset also showcases the feasibility of simulating inter-chain interactions using intra-chain interactions, expanding available resources for machine-learning models on protein-protein interactions. The dataset and relevant scripts are accessible on GitHub ([https://github.com/YZY010418/CPSea](https://github.com/YZY010418/CPSea)).
Ziyi Yang 0011, Hanyuan Xie, Yinjun Jia, Xiangzhe Kong, Jiqing Zheng, Ziting Zhang, Yang Liu 0003, Lei Liu 0049, Yanyan Lan
NeurIPS3
2025 AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation
abstract
Virtual screening (VS) is a critical component of modern drug discovery, yet most existing methods—whether physics-based or deep learning-based—are developed around *holo* protein structures with known ligand-bound pockets. Consequently, their performance degrades significantly on *apo* or predicted structures such as those from AlphaFold2, which are more representative of real-world early-stage drug discovery, where pocket information is often missing. In this paper, we introduce an alignment-and-aggregation framework to enable accurate virtual screening under structural uncertainty. Our method comprises two core components: (1) a tri-modal contrastive learning module that aligns representations of the ligand, the *holo* pocket, and cavities detected from structures, thereby enhancing robustness to pocket localization error; and (2) a cross-attention based adapter for dynamically aggregating candidate binding sites, enabling the model to learn from activity data even without precise pocket annotations. We evaluated our method on a newly curated benchmark of *apo* structures, where it significantly outperforms state-of-the-art methods in blind apo setting, improving the early enrichment factor (EF1\%) from 11.75 to 37.19. Notably, it also maintains strong performance on *holo* structures. These results demonstrate the promise of our approach in advancing first-in-class drug discovery, particularly in scenarios lacking experimentally resolved protein-ligand complexes. Our implementation is publicly available at [https://github.com/Wiley-Z/AANet](https://github.com/Wiley-Z/AANet).
Wenyu Zhu, Yinjun Jia, Haichuan Tan, Ya-Qin Zhang, Wei-Ying Ma, Yanyan Lan
NeurIPS4
2024 Protein-ligand binding representation learning from fine-grained interactions
abstract
The binding between proteins and ligands plays a crucial role in the realm of drug discovery. Previous deep learning approaches have shown promising results over traditional computationally intensive methods, but resulting in poor generalization due to limited supervised data. In this paper, we propose to learn protein-ligand binding representation in a self-supervised learning manner. Different from existing pre-training approaches which treat proteins and ligands individually, we emphasize to discern the intricate binding patterns from fine-grained interactions. Specifically, this self-supervised learning problem is formulated as a prediction of the conclusive binding complex structure given a pocket and ligand with a Transformer based interaction module, which naturally emulates the binding process. To ensure the representation of rich binding information, we introduce two pre-training tasks, i.e. atomic pairwise distance map prediction and mask ligand reconstruction, which comprehensively model the fine-grained interactions from both structure and feature space. Extensive experiments have demonstrated the superiority of our method across various binding tasks, including protein-ligand affinity prediction, virtual screening and protein-ligand docking.
Shikun Feng, Yinjun Jia, Wei-Ying Ma, Yanyan Lan
ICLR3
2024 Self-supervised Pocket Pretraining via Protein Fragment-Surroundings Alignment
abstract
Pocket representations play a vital role in various biomedical applications, such as druggability estimation, ligand affinity prediction, and de novo drug design. While existing geometric features and pretrained representations have demonstrated promising results, they usually treat pockets independent of ligands, neglecting the fundamental interactions between them. However, the limited pocket-ligand complex structures available in the PDB database (less than 100 thousand non-redundant pairs) hampers large-scale pretraining endeavors for interaction modeling. To address this constraint, we propose a novel pocket pretraining approach that leverages knowledge from high-resolution atomic protein structures, assisted by highly effective pretrained small molecule representations. By segmenting protein structures into drug-like fragments and their corresponding pockets, we obtain a reasonable simulation of ligand-receptor interactions, resulting in the generation of over 5 million complexes. Subsequently, the pocket encoder is trained in a contrastive manner to align with the representation of pseudo-ligand furnished by some pretrained small molecule encoders. Our method, named ProFSA, achieves state-of-the-art performance across various tasks, including pocket druggability prediction, pocket matching, and ligand binding affinity prediction. Notably, ProFSA surpasses other pretraining methods by a substantial margin. Moreover, our work opens up a new avenue for mitigating the scarcity of protein-ligand complex data through the utilization of high-quality and diverse protein structure databases.
Yinjun Jia, Yuanle Mo, Yuyan Ni, Wei-Ying Ma, Zhiming Ma, Yanyan Lan
ICLR2
2024 Full-Atom Peptide Design with Geometric Latent Diffusion
abstract
Peptide design plays a pivotal role in therapeutics, allowing brand new possibility to leverage target binding sites that are previously undruggable. Most existing methods are either inefficient or only concerned with the target-agnostic design of 1D sequences. In this paper, we propose a generative model for full-atom Peptide design with Geometric LAtent Diffusion (PepGLAD) given the binding site. We first establish a benchmark consisting of both 1D sequences and 3D structures from Protein Data Bank (PDB) and literature for systematic evaluation. We then identify two major challenges of leveraging current diffusion-based models for peptide design: the full-atom geometry and the variable binding geometry. To tackle the first challenge, PepGLAD derives a variational autoencoder that first encodes full-atom residues of variable size into fixed-dimensional latent representations, and then decodes back to the residue space after conducting the diffusion process in the latent space. For the second issue, PepGLAD explores a receptor-specific affine transformation to convert the 3D coordinates into a shared standard space, enabling better generalization ability across different binding shapes. Experimental Results show that our method not only improves diversity and binding affinity significantly in the task of sequence-structure co-design, but also excels at recovering reference structures for binding conformation generation.
Xiangzhe Kong, Yinjun Jia, Wenbing Huang 0001, Yang Liu 0005
NeurIPS2
2023 DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening
Bo Qiang, Haichuan Tan, Yinjun Jia, Minsi Ren, Minsi Lu, Wei-Ying Ma, Yanyan Lan
NeurIPS4