Shaorong Chen

dblp:284/3994 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
4 papers
Graph learning · 27% Generative modeling · 18% Deep learning architectures and training · 15%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › proteomics › peptide sequencing
de novo peptide sequencing
5.162026
Regressor-guided Diffusion Model for De Novo Peptide Sequencing with Explicit Mass Control · AAAI 2026
A Comprehensive and Systematic Review for Deep Learning-Based De Novo Peptide Sequencing · IJCAI 2025
Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo · ICLR 2025
Bioinformatics and computational biology
proteomics
5.162026
Regressor-guided Diffusion Model for De Novo Peptide Sequencing with Explicit Mass Control · AAAI 2026
A Comprehensive and Systematic Review for Deep Learning-Based De Novo Peptide Sequencing · IJCAI 2025
Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo · ICLR 2025
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
1.222026
ReNovo: Retrieval-Based \emph{De Novo} Mass Spectrometry Peptide Sequencing · ICLR 2025
Regressor-guided Diffusion Model for De Novo Peptide Sequencing with Explicit Mass Control · AAAI 2026
Machine learning › Generative modeling
diffusion model
1.012026
Regressor-guided Diffusion Model for De Novo Peptide Sequencing with Explicit Mass Control · AAAI 2026
Bioinformatics and computational biology › sequence analysis
database search
0.912025
Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo · ICLR 2025
Machine learning › Learning theory › information theory
conditional mutual information
0.812024
AdaNovo: Towards Robust \emph{De Novo} Peptide Sequencing in Proteomics against Data Biases · NeurIPS 2024
Machine learning › Representation and self-supervised learning › contrastive learning
graph contrastive learning
0.812024
DiscoGNN: A Sample-Efficient Framework for Self-Supervised Graph Representation Learning · ICDE 2024
Machine learning › Graph learning
graph representation learning
0.812024
DiscoGNN: A Sample-Efficient Framework for Self-Supervised Graph Representation Learning · ICDE 2024
Machine learning › Graph learning
graph self-supervised learning
0.812024
DiscoGNN: A Sample-Efficient Framework for Self-Supervised Graph Representation Learning · ICDE 2024
Machine learning › Trustworthy machine learning › robustness
robust learning
0.812024
AdaNovo: Towards Robust \emph{De Novo} Peptide Sequencing in Proteomics against Data Biases · NeurIPS 2024
Performance modeling and evaluation
benchmarking
0.812024
NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024
Performance modeling and evaluation › benchmarking › machine learning benchmarking
deep learning benchmarks
0.812024
NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

deep learning · 3.9tandem mass spectrometry · 3.3gradient-based guidance · 2.0diffusion model · 2.0deep neural network · 1.7conditional mutual information · 1.5retrieval-based inference · 0.9datastore · 0.9masking and correction · 0.8contrastive learning · 0.8
YearPublicationVenuePosition
2026 Regressor-guided Diffusion Model for De Novo Peptide Sequencing with Explicit Mass Control
abstract
The discovery of novel proteins relies on sensitive protein identification, for which de novo peptide sequencing (DNPS) from mass spectra is a crucial approach. While deep learning has advanced DNPS, existing models inadequately enforce the fundamental mass consistency constraint—that a predicted peptide's mass must match the experimental measured precursor mass. Previous DNPS methods often treat this critical information as a simple input feature or use it in post-processing, leading to numerous implausible predictions that do not adhere to this fundamental physical property. To address this limitation, we introduce DiffuNovo, a novel regressor-guided diffusion model for de novo peptide sequencing that provides explicit peptide-level mass control. Our approach integrates the mass constraint at two critical stages: during training, a novel peptide-level mass loss guides model optimization, while at inference, regressor-based guidance from gradient-based updates in the latent space steers the generation to compel the predicted peptide adheres to the mass constraint. Comprehensive evaluations on established benchmarks demonstrate that DiffuNovo surpasses state-of-the-art methods in DNPS accuracy. Additionally, as the first DNPS model to employ a diffusion model as its core backbone, DiffuNovo leverages the powerful controllability of diffusion architecture and achieves a significant reduction in mass error, thereby producing much more physically plausible peptides. These innovations represent a substantial advancement toward robust and broadly applicable DNPS. The source code is available in the supplementary material.
Shaorong Chen, Jun Xia 0001
AAAI1
2025 ReNovo: Retrieval-Based \emph{De Novo} Mass Spectrometry Peptide Sequencing
abstract
Proteomics is the large-scale study of proteins. Tandem mass spectrometry, as the only high-throughput technique for protein sequence identification, plays a pivotal role in proteomics research. One of the long-standing challenges in this field is peptide identification, which entails determining the specific peptide (sequence of amino acids) that corresponds to each observed mass spectrum. The conventional approach involves database searching, wherein the observed mass spectrum is scored against a pre-constructed peptide database. However, the reliance on pre-existing databases limits applicability in scenarios where the peptide is absent from existing databases. Such circumstances necessitate \emph{de novo} peptide sequencing, which derives peptide sequence solely from input mass spectrum, independent of any peptide database. Despite ongoing advancements in \emph{de novo} peptide sequencing, its performance still has considerable room for improvement, which limits its application in large-scale experiments. In this study, we introduce a novel \textbf{Re}trieval-based \emph{De \textbf{Novo}} peptide sequencing methodology, termed \textbf{ReNovo}, which draws inspiration from database search methods. Specifically, by constructing a datastore from training data, ReNovo can retrieve information from the datastore during the inference stage to conduct retrieval-based inference, thereby achieving improved performance. This innovative approach enables ReNovo to effectively combine the strengths of both methods: utilizing the assistance of the datastore while also being capable of predicting novel peptides that are not present in pre-existing databases. A series of experiments have confirmed that ReNovo outperforms state-of-the-art models across multiple widely-used datasets, incurring only minor storage and time consumption, representing a significant advancement in proteomics. Supplementary materials include the code.
Shaorong Chen, Jun Xia 0001, Lecheng Zhang, Zhangyang Gao, Bozhen Hu, Cheng Tan 0012, Wenjie Du 0003, Stan Z. Li
ICLR1
2025 Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo
Jun Xia 0001, Sizhe Liu, Shaorong Chen, Hongxin Xiang, Zicheng Liu 0006, Yue Liu 0008, Stan Z. Li
ICLR4
2025 A Comprehensive and Systematic Review for Deep Learning-Based De Novo Peptide Sequencing
abstract
Tandem mass spectrometry (MS/MS) has revolutionized the field of proteomics, enabling the high-throughput identification of proteins. However, one of the central challenges in mass spectrometry-based proteomics remains peptide identification, especially in the absence of a comprehensive peptide database. While traditional database search methods compare observed mass spectra to pre-existing protein databases, they are limited by the availability and completeness of these databases. \emph{De novo} peptide sequencing, which derives peptide sequences directly from mass spectra, has emerged as a crucial approach in such cases. In recent years, deep learning has made significant strides in this domain. These methods train deep neural networks for translating mass spectra into peptide sequences without relying on any pre-constructed databases. Despite significant progress, this field still lacks a comprehensive and systematic review. In this paper, we provide the first review of deep learning-based \emph{de novo} peptide sequencing techniques from the perspectives of data types, model architectures, decoding strategies, applications and evaluation metrics. We also identify key challenges and highlight promising avenues for future research, providing a valuable resource for the AI and scientific communities.
Jun Xia 0001, Shaorong Chen, Tianze Ling, Stan Z. Li
IJCAI3
2024 DiscoGNN: A Sample-Efficient Framework for Self-Supervised Graph Representation Learning
abstract
Self-supervised graph representation learning has received increasing research interest recently, with generative and contrastive modeling being two dominant ways. Typically, generative learning first masks parts of each graph and then recovers the masked parts based on the encoding results of the corrupted graph. However, these methods only mask fixed parts of each graph and fail to train on all the nodes and edges, which hinders them from getting the most out of each graph. As a remedy, we propose a novel self-supervised strategy, dubbed DetCor, where we first randomly replace some nodes and edges with alternative ones and then pre-train GNNs to detect and correct the replaced ones from all the nodes and edges. Additionally, for graph-level learning, the vanilla contrastive framework cannot reflect the distinction between the in-batch negatives. To alleviate this issue, we propose RankGCL, which enables the contrastive framework to capture the similarity ranking information between graphs and shows special superiority in graph similarity-based practical tasks. DetCor and RankGCL together constitute a unified self-supervised framework, DiscoGNN, which matches or outperforms state-of-the-art strategies on multiple datasets from various domains. Also, DiscoGNN is a sample-efficient framework that can achieve better performance than competitive methods with much less pre-training data. We release the codes at: https://github.com/junxia97/DiscoGNN-ICDE.
Jun Xia 0001, Shaorong Chen, Yue Liu 0008, Zhangyang Gao, Jiangbin Zheng 0002, Xihong Yang, Stan Z. Li
ICDE2
2024 AdaNovo: Towards Robust \emph{De Novo} Peptide Sequencing in Proteomics against Data Biases
abstract
Tandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the high-throughput analysis of protein composition in biological tissues. Despite the development of several deep learning methods for predicting amino acid sequences (peptides) responsible for generating the observed mass spectra, training data biases hinder further advancements of \emph{de novo} peptide sequencing. Firstly, prior methods struggle to identify amino acids with Post-Translational Modifications (PTMs) due to their lower frequency in training data compared to canonical amino acids, further resulting in unsatisfactory peptide sequencing performance. Secondly, various noise and missing peaks in mass spectra reduce the reliability of training data (Peptide-Spectrum Matches, PSMs). To address these challenges, we propose AdaNovo, a novel and domain knowledge-inspired framework that calculates Conditional Mutual Information (CMI) between the mass spectra and amino acids or peptides, using CMI for robust training against above biases. Extensive experiments indicate that AdaNovo outperforms previous competitors on the widely-used 9-species benchmark, meanwhile yielding 3.6\% - 9.4\% improvements in PTMs identification. The supplements contain the code.
Jun Xia 0001, Shaorong Chen, Xiaojun Shan, Wenjie Du 0003, Zhangyang Gao, Cheng Tan 0012, Bozhen Hu, Jiangbin Zheng 0002, Stan Z. Li
NeurIPS2
2024 NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics
abstract
Tandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the analysis of protein composition in biological tissues. Many deep learning methods have been developed for \emph{de novo} peptide sequencing task, i.e., predicting the peptide sequence for the observed mass spectrum. However, two key challenges seriously hinder the further research of this important task. Firstly, since there is no consensus for the evaluation datasets, the empirical results in different research papers are often not comparable, leading to unfair comparison. Secondly, the current methods are usually limited to amino acid-level or peptide-level precision and recall metrics. In this work, we present the first unified benchmark NovoBench for \emph{de novo} peptide sequencing, which comprises diverse mass spectrum data, integrated models, and comprehensive evaluation metrics. Recent impressive methods, including DeepNovo, PointNovo, Casanovo, InstaNovo, AdaNovo and $\pi$-HelixNovo are integrated into our framework. In addition to amino acid-level and peptide-level precision and recall, we also evaluate the models' performance in terms of identifying post-tranlational modifications (PTMs), efficiency and robustness to peptide length, noise peaks and missing fragment ratio, which are important influencing factors while seldom be considered. Leveraging this benchmark, we conduct a large-scale study of current methods, report many insightful findings that open up new possibilities for future development. The benchmark is open-sourced to facilitate future research and application. The code is available at \url{https://github.com/Westlake-OmicsAI/NovoBench}.
Shaorong Chen, Jun Xia 0001, Sizhe Liu, Tianze Ling, Wenjie Du 0003, Yue Liu 0008, Jianwei Yin, Stan Z. Li
NeurIPS2