VLDB 2026 Research / reviewers in the wild / expert
Yinjun Jia
dblp:333/7450
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
7 papers |
Generative modeling · 50% Representation and self-supervised learning · 20% Graph learning · 12% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › structural bioinformatics
protein structure |
1.7 | 2 | 2025 | AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation · NeurIPS 2025 CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design · NeurIPS 2025 |
Bioinformatics and computational biology
drug discovery |
1.5 | 2 | 2025 | AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation · NeurIPS 2025 DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening · NeurIPS 2023 |
Bioinformatics and computational biology › drug discovery
virtual screening |
1.5 | 2 | 2025 | AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation · NeurIPS 2025 DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening · NeurIPS 2023 |
Machine learning › Representation and self-supervised learning
contrastive learning |
1.0 | 2 | 2024 | Self-supervised Pocket Pretraining via Protein Fragment-Surroundings Alignment · ICLR 2024 DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening · NeurIPS 2023 |
Machine learning › Generative modeling › generative model › continuous-time generative model
bayesian flow network |
0.9 | 1 | 2025 | Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent Space · NeurIPS 2025 |
Machine learning › Generative modeling
molecular generation |
0.9 | 1 | 2025 | Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent Space · NeurIPS 2025 |
Machine learning › Generative modeling
variational autoencoder |
0.9 | 1 | 2025 | Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent Space · NeurIPS 2025 |
Bioinformatics and computational biology › gene regulation
binding site prediction |
0.9 | 1 | 2025 | AANet: Virtual Screening under Structural Uncertainty via Alignment and Aggregation · NeurIPS 2025 |
Bioinformatics and computational biology › drug discovery
bioactivity prediction |
0.9 | 1 | 2025 | Redefining the task of Bioactivity Prediction · ICLR 2025 |
Bioinformatics and computational biology
molecular property prediction |
0.9 | 1 | 2025 | Redefining the task of Bioactivity Prediction · ICLR 2025 |
Bioinformatics and computational biology
protein design |
0.9 | 1 | 2025 | CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design · NeurIPS 2025 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Full-Atom Peptide Design with Geometric Latent Diffusion · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.8 | 1 | 2024 | Full-Atom Peptide Design with Geometric Latent Diffusion · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › structured representation
molecular representation |
0.8 | 1 | 2024 | Self-supervised Pocket Pretraining via Protein Fragment-Surroundings Alignment · ICLR 2024 |
Machine learning › Graph learning
molecular representation learning |
0.8 | 1 | 2024 | Protein-ligand binding representation learning from fine-grained interactions · ICLR 2024 |
Bioinformatics and computational biology › protein design
peptide design |
0.8 | 1 | 2024 | Full-Atom Peptide Design with Geometric Latent Diffusion · NeurIPS 2024 |
Bioinformatics and computational biology › protein analysis
protein-ligand interaction |
0.8 | 1 | 2024 | Protein-ligand binding representation learning from fine-grained interactions · ICLR 2024 |
Computer vision › Vision and language › vision-language pretraining
contrastive vision-language pretraining |
0.7 | 1 | 2023 | DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening · NeurIPS 2023 |
Machine learning › Graph learning
protein structure prediction |
0.3 | 1 | 2025 | CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide design · NeurIPS 2025 |
Machine learning › Generative modeling › molecular generation
molecular structure generation |
0.2 | 1 | 2024 | Full-Atom Peptide Design with Geometric Latent Diffusion · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 3.7structure filtering · 1.7machine learning · 1.7docking · 1.7dataset curation · 1.7variational autoencoder · 1.6pairwise distance prediction · 1.5masked modeling · 1.5e(3)-equivariant neural network · 0.9cross-attention adapter · 0.9bayesian flow network · 0.9transformer · 0.8fragment-surroundings alignment · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Redefining the task of Bioactivity PredictionabstractSmall molecules are vital to modern medicine, and accurately predicting their bioactivity against protein targets is crucial for therapeutic discovery and development. However, current machine learning models often rely on spurious features, leading to biased outcomes. Notably, a simple pocket-only baseline can achieve results comparable to, and sometimes better than, more complex models that incorporate both the protein pockets and the small molecules. Our analysis reveals that this phenomenon arises from insufficient training data and an improper evaluation process, which is typically conducted at the pocket level rather than the small molecule level. To address these issues, we redefine the bioactivity prediction task by introducing the SIU dataset-a million-scale Structural small molecule-protein Interaction dataset for Unbiased bioactivity prediction task, which is 50 times larger than the widely used PDBbind. The bioactivity labels in SIU are derived from wet experiments and organized by label types, ensuring greater accuracy and comparability. The complexes in SIU are constructed using a majority vote from three commonly used docking software programs, enhancing their reliability. Additionally, the structure of SIU allows for multiple small molecules to be associated with each protein pocket, enabling the redefinition of evaluation metrics like Pearson and Spearman correlations across different small molecules targeting the same protein pocket. Experimental results demonstrate that this new task provides a more challenging and meaningful benchmark for training and evaluating bioactivity prediction models, ultimately offering a more robust assessment of model performance. Yanwen Huang, Yinjun Jia, Hongbo Ma, Wei-Ying Ma, Ya-Qin Zhang, Yanyan Lan |
ICLR | 3 |
| 2025 | Manipulating 3D Molecules in a Fixed-Dimensional E(3)-Equivariant Latent SpaceabstractMedicinal chemists often optimize drugs considering their 3D structures and designing structurally distinct molecules that retain key features, such as shapes, pharmacophores, or chemical properties. Previous deep learning approaches address this through supervised tasks like molecule inpainting or property-guided optimization. In this work, we propose a flexible zero-shot molecule manipulation method by navigating in a shared latent space of 3D molecules. We introduce a Variational AutoEncoder (VAE) for 3D molecules, named MolFLAE, which learns a fixed-dimensional, E(3)-equivariant latent space independent of atom counts. MolFLAE encodes 3D molecules using an E(3)-equivariant neural network into fixed number of latent nodes, distinguished by learned embeddings. The latent space is regularized, and molecular structures are reconstructed via a Bayesian Flow Network (BFN) conditioned on the encoder’s latent output. MolFLAE achieves competitive performance on standard unconditional 3D molecule generation benchmarks. Moreover, the latent space of MolFLAE enables zero-shot molecule manipulation, including atom number editing, structure reconstruction, and coordinated latent interpolation for both structure and properties. We further demonstrate our approach on a drug optimization task for the human glucocorticoid receptor, generating molecules with improved hydrophilicity while preserving key interactions, under computational evaluations. These results highlight the flexibility, robustness, and real-world utility of our method, opening new avenues for molecule editing and optimization. Yinjun Jia, Zitong Tian, Wei-Ying Ma, Yanyan Lan |
NeurIPS | 2 |
| 2025 | CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide designabstractCyclic peptides exhibit better binding affinity and proteolytic stability compared to their linear counterparts. However, the development of cyclic peptide design models is hindered by the scarcity of data. To address this, we introduce **CPSea**(**C**yclic **P**eptide **Sea**), a dataset of 2.71 million cyclic peptide-receptor complexes, curated through systematic mining of the AlphaFold Database (AFDB). Our pipeline extracts compact domains from AFDB, identifies cyclization sites using the $\beta$-carbon (C$_\beta$) distance thresholds, and applies multi-stage filtering to ensure structure fidelity and binding compatibility. Compared with experimental data of cyclic peptides, CPSea shows similar distributions in metrics on structure fidelity and wet-lab compatibility. To our knowledge, CPSea is the largest cyclic peptide-receptor dataset to date, enabling end-to-end model training for the first time. The dataset also showcases the feasibility of simulating inter-chain interactions using intra-chain interactions, expanding available resources for machine-learning models on protein-protein interactions. The dataset and relevant scripts are accessible on GitHub ([https://github.com/YZY010418/CPSea](https://github.com/YZY010418/CPSea)). Ziyi Yang 0011, Hanyuan Xie, Yinjun Jia, Xiangzhe Kong, Jiqing Zheng, Ziting Zhang, Yang Liu 0003, Lei Liu 0049, Yanyan Lan |
NeurIPS | 3 |
| 2025 | AANet: Virtual Screening under Structural Uncertainty via Alignment and AggregationabstractVirtual screening (VS) is a critical component of modern drug discovery, yet most existing methods—whether physics-based or deep learning-based—are developed around *holo* protein structures with known ligand-bound pockets. Consequently, their performance degrades significantly on *apo* or predicted structures such as those from AlphaFold2, which are more representative of real-world early-stage drug discovery, where pocket information is often missing. In this paper, we introduce an alignment-and-aggregation framework to enable accurate virtual screening under structural uncertainty. Our method comprises two core components: (1) a tri-modal contrastive learning module that aligns representations of the ligand, the *holo* pocket, and cavities detected from structures, thereby enhancing robustness to pocket localization error; and (2) a cross-attention based adapter for dynamically aggregating candidate binding sites, enabling the model to learn from activity data even without precise pocket annotations. We evaluated our method on a newly curated benchmark of *apo* structures, where it significantly outperforms state-of-the-art methods in blind apo setting, improving the early enrichment factor (EF1\%) from 11.75 to 37.19. Notably, it also maintains strong performance on *holo* structures. These results demonstrate the promise of our approach in advancing first-in-class drug discovery, particularly in scenarios lacking experimentally resolved protein-ligand complexes. Our implementation is publicly available at [https://github.com/Wiley-Z/AANet](https://github.com/Wiley-Z/AANet). Wenyu Zhu, Yinjun Jia, Haichuan Tan, Ya-Qin Zhang, Wei-Ying Ma, Yanyan Lan |
NeurIPS | 4 |
| 2024 | Protein-ligand binding representation learning from fine-grained interactionsabstractThe binding between proteins and ligands plays a crucial role in the realm of drug discovery. Previous deep learning approaches have shown promising results over traditional computationally intensive methods, but resulting in poor generalization due to limited supervised data. In this paper, we propose to learn protein-ligand binding representation in a self-supervised learning manner. Different from existing pre-training approaches which treat proteins and ligands individually, we emphasize to discern the intricate binding patterns from fine-grained interactions. Specifically, this self-supervised learning problem is formulated as a prediction of the conclusive binding complex structure given a pocket and ligand with a Transformer based interaction module, which naturally emulates the binding process. To ensure the representation of rich binding information, we introduce two pre-training tasks, i.e. atomic pairwise distance map prediction and mask ligand reconstruction, which comprehensively model the fine-grained interactions from both structure and feature space. Extensive experiments have demonstrated the superiority of our method across various binding tasks, including protein-ligand affinity prediction, virtual screening and protein-ligand docking. Shikun Feng, Yinjun Jia, Wei-Ying Ma, Yanyan Lan |
ICLR | 3 |
| 2024 | Self-supervised Pocket Pretraining via Protein Fragment-Surroundings AlignmentabstractPocket representations play a vital role in various biomedical applications, such as druggability estimation, ligand affinity prediction, and de novo drug design. While existing geometric features and pretrained representations have demonstrated promising results, they usually treat pockets independent of ligands, neglecting the fundamental interactions between them. However, the limited pocket-ligand complex structures available in the PDB database (less than 100 thousand non-redundant pairs) hampers large-scale pretraining endeavors for interaction modeling. To address this constraint, we propose a novel pocket pretraining approach that leverages knowledge from high-resolution atomic protein structures, assisted by highly effective pretrained small molecule representations. By segmenting protein structures into drug-like fragments and their corresponding pockets, we obtain a reasonable simulation of ligand-receptor interactions, resulting in the generation of over 5 million complexes. Subsequently, the pocket encoder is trained in a contrastive manner to align with the representation of pseudo-ligand furnished by some pretrained small molecule encoders. Our method, named ProFSA, achieves state-of-the-art performance across various tasks, including pocket druggability prediction, pocket matching, and ligand binding affinity prediction. Notably, ProFSA surpasses other pretraining methods by a substantial margin. Moreover, our work opens up a new avenue for mitigating the scarcity of protein-ligand complex data through the utilization of high-quality and diverse protein structure databases. Yinjun Jia, Yuanle Mo, Yuyan Ni, Wei-Ying Ma, Zhiming Ma, Yanyan Lan |
ICLR | 2 |
| 2024 | Full-Atom Peptide Design with Geometric Latent DiffusionabstractPeptide design plays a pivotal role in therapeutics, allowing brand new possibility to leverage target binding sites that are previously undruggable. Most existing methods are either inefficient or only concerned with the target-agnostic design of 1D sequences. In this paper, we propose a generative model for full-atom Peptide design with Geometric LAtent Diffusion (PepGLAD) given the binding site. We first establish a benchmark consisting of both 1D sequences and 3D structures from Protein Data Bank (PDB) and literature for systematic evaluation. We then identify two major challenges of leveraging current diffusion-based models for peptide design: the full-atom geometry and the variable binding geometry. To tackle the first challenge, PepGLAD derives a variational autoencoder that first encodes full-atom residues of variable size into fixed-dimensional latent representations, and then decodes back to the residue space after conducting the diffusion process in the latent space. For the second issue, PepGLAD explores a receptor-specific affine transformation to convert the 3D coordinates into a shared standard space, enabling better generalization ability across different binding shapes. Experimental Results show that our method not only improves diversity and binding affinity significantly in the task of sequence-structure co-design, but also excels at recovering reference structures for binding conformation generation. Xiangzhe Kong, Yinjun Jia, Wenbing Huang 0001, Yang Liu 0005 |
NeurIPS | 2 |
| 2023 | DrugCLIP: Contrasive Protein-Molecule Representation Learning for Virtual Screening
Bo Qiang, Haichuan Tan, Yinjun Jia, Minsi Ren, Minsi Lu, Wei-Ying Ma, Yanyan Lan |
NeurIPS | 4 |