VLDB 2026 Research / reviewers in the wild / expert
Shaofeng Lin
dblp:197/4236
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-1177-5480ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencingabstractSUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077). Jiayin Dai, Jiongming Ma, Kunqi Chen, Jia Meng 0001, Daniel J. Rigden, Shaofeng Lin, Qingru Xu |
Bioinform. | 9 |
| 2025 | Optimizing Secure Data Transmission in GPU-Accelerated Collaborative Computing via Metadata ReconstructionabstractAs the use of GPUs in collaborative computing environments increases, the demand for memory capacity continues to grow. The introduction of Compute Express Link (CXL) technology enables GPUs to access extended memory but also introduces new security challenges. Expressly, data transmission between secure memory regions incurs overhead due to ciphertext conversion, driven by independent security meta data for each memory region, which results in additional memory access overhead and degrades performance. Existing solutions alleviate this issue by sharing security metadata, thus reducing overhead from additional memory accesses. However, transmitting security metadata still consumes significant bandwidth, limiting overall system performance. To address this limitation, we propose a data transmission optimization method based on secure metadata reconstruction tailored to the unique characteristics of GPU multi-memory architectures in collaborative computing. Our method minimizes the amount of security metadata transmitted and reconstructs complete metadata at the receiver's end. This approach significantly reduces bandwidth usage and enhances system performance. Experimental results demonstrate that our method outperforms the Salus approach, achieving an 11.0% improvement in IPC (Instructions Per Cycle) performance. Shaofeng Lin, Qiming Zhou, Yeping He, Hengtai Ma |
CSCWD | 1 |
| 2025 | Gradient-Threshold-Based Integrity Tree Optimization for Secure GPU MemoryabstractWith the increasing adoption of GPUs in neural networks, memory security concerns have become more pronounced. As a widely used memory integrity protection mechanism, the Integrity Tree effectively defends against replay attacks. However, the access and processing overhead associated with Integrity Trees has emerged as a significant performance bottleneck. Existing optimization strategies primarily rely on runtime prediction of frequently accessed regions to adjust the Integrity Tree structure, thereby reducing node access and processing costs. Nevertheless, due to the spatial accumulation and temporal locality characteristics of GPU memory access, these methods suffer from limitations in prediction accuracy and responsiveness: the influence of historical data can lead to inaccurate predictions, while the dynamic evolution of frequently accessed regions is challenging to detect in real-time, ultimately impacting the optimization effectiveness of the Integrity Tree. To address these challenges, this paper proposes an Integrity Tree optimization method based on a gradient-threshold mechanism to improve the access efficiency of secure GPU memory. The proposed approach introduces a gradient-based access weighting mechanism to mitigate the influence of historical data, enabling predictions to capture current memory access patterns more accurately. Additionally, a threshold-based real-time detection strategy is employed to continuously monitor changes in frequently accessed regions and adapt the Integrity Tree structure accordingly. Experimental results demonstrate that the proposed method significantly reduces the node traversal overhead associated with Integrity Tree. Shaofeng Lin, Yeping He, Qiming Zhou, Hengtai Ma |
IJCNN | 1 |
| 2025 | Towards Few-shot LLM-based Vulnerability Reproduction and Verification for Industrial Web ApplicationsabstractThe inherent vulnerability of Web system poses a severe security threat to modern industrial systems. Nevertheless, automatically obtaining accurate and reproducible vulnerability intelligence in a time-acceptable manner remains a challenge. To address this issue, we propose VulRV, an end-to-end vulnerability reproduction and verification system for industrial Web applications. VulRV employs large language models (LLMs) with few-shot Chain-of-Thought (CoT) prompting to generate the environment configuration. And then VulRV utilizes tool-calling method to establish corresponding vulnerability reproduction and validation environment. To evaluate the effectiveness of the proposed method, we collected reports of 22 vulnerabilities of mainstream industrial Web application spanning four different categories. Experimental results demonstrate that our method can effectively extract information from the reports to construct Docker containers-based environments to automatically reproduce and validate all these 22 vulnerabilities. Thereby the results show VulRV are capable of supporting subsequent red-team testing and risk assessment of industrial Web applications. Lin Ni, Shaofeng Lin |
INDIN | 3 |
| 2025 | Deciphering the MHC immunopeptidome of human cancers with Ligand.MHC atlasabstractA fundamental principle of immunotherapy is that T cells are capable of detecting tumor epitopes presented on cancer cell surfaces. Immunopeptidomic strategies empowered by liquid chromatography-tandem mass spectrometry have transformed tumor epitopes identification and provided novel insights into tumor immunology. It enables in-depth profiling of major histocompatibility complex (MHC) presented ligands, thereby offering valuable perspectives on the molecular dialog among tumor and T cells. Here, we developed an immune-ligand identification and analysis pipeline from large-scale immunopeptidomics data. Through an extensive collection and processing of 5821 immunopeptidomic samples, which amounted to 305.7 million MS2 spectra, we identified 24 380 595 peptide-spectrum matches from these samples and further detected a total of 1 017 731 unique MHC immune ligands. These ligands were deconvolved and classified to specific HLA alleles. In total, we detected 582 852 HLA-I peptides and 434 879 HLA-II peptides that can bind to 292 HLA alleles, thereby greatly expanding the cancer immunopeptidome. Additionally, we identified and annotated 372 720 tumor-associated post-translational modification (PTM) peptides, revealing the comprehensive landscape of PTM antigens. All ligands and annotations were aggregated into Ligand.MHC Atlas, a comprehensive repository dedicated to tumor-derived HLA-presented ligands across 26 major human cancers (54 subtypes). Overall, our study uniquely integrates batch-effect correction, leverages the optimized software with novel deconvolution approach for immunopeptidomics analysis and ligand identification, and provides a public web portal with a comprehensive HLA ligand repository. Ligand.MHC Atlas functions as an invaluable resource, offering crucial understandings into immunology investigations. It will accelerate the advancement of cancer vaccines and immunotherapies. Ligand.MHC Atlas is available at http://modinfor.com/Ligand.MHC-Atlas/. Zhi Ran, Meilin Mu, Shaofeng Lin, Lan Kuang, Kunqi Chen, Shengbao Suo, Hao-Dong Xu |
Briefings Bioinform. | 3 |
| 2024 | Ensuring Data Integrity and Freshness in GPU-CXL Transfers With Tamper-Resistant MetadataabstractGPUs have become critical accelerators in applications such as scientific computing and deep learning. The demand for more robust security measures has significantly increased as their usage expands across a broader range of applications. To address this need, Trusted Execution Environments (TEEs) have been integrated into GPU systems to safeguard applications and data. However, with technologies like Compute Express Link expanding the memory capacity of GPUs, cross-memory data transfers now pose new security challenges. Existing methods attempt to optimize the performance overhead associated with these transfers by sharing security metadata. However, this shared metadata is unreliable, raising concerns about data integrity and freshness and making systems vulnerable to replay attacks. In response to these issues, this paper introduces the trusted shared security metadata concept. Leveraging the tamper-resistant of MoveCtr, the paper develops a trusted shared security metadata scheme and a secure cross-memory integrity tree, ensuring the integrity and freshness of transferred data and protecting against replay attacks. Experimental results demonstrate that this method provides strong security guarantees while maintaining performance overhead within an acceptable range. Shaofeng Lin, Yeping He, Qiming Zhou, Hengtai Ma |
HPCC | 1 |
| 2024 | MetaDegron: multimodal feature-integrated protein language model for predicting E3 ligase targeted degronsabstractProtein degradation through the ubiquitin proteasome system at the spatial and temporal regulation is essential for many cellular processes. E3 ligases and degradation signals (degrons), the sequences they recognize in the target proteins, are key parts of the ubiquitin-mediated proteolysis, and their interactions determine the degradation specificity and maintain cellular homeostasis. To date, only a limited number of targeted degron instances have been identified, and their properties are not yet fully characterized. To tackle on this challenge, here we develop a novel deep-learning framework, namely MetaDegron, for predicting E3 ligase targeted degron by integrating the protein language model and comprehensive featurization strategies. Through extensive evaluations using benchmark datasets and comparison with existing method, such as Degpred, we demonstrate the superior performance of MetaDegron. Among functional features, MetaDegron allows batch prediction of targeted degrons of 21 E3 ligases, and provides functional annotations and visualization of multiple degron-related structural and physicochemical features. MetaDegron is freely available at http://modinfor.com/MetaDegron/. We anticipate that MetaDegron will serve as a useful tool for the clinical and translational community to elucidate the mechanisms of regulation of protein homeostasis, cancer research, and drug development. Mengqiu Zheng, Shaofeng Lin, Kunqi Chen, Ruifeng Hu 0002, Zhongming Zhao, Haodong Xu |
Briefings Bioinform. | 2 |
| 2022 | GPS-Uber: a hybrid-learning framework for prediction of general and E3-specific lysine ubiquitination sitesabstractAs an important post-translational modification, lysine ubiquitination participates in numerous biological processes and is involved in human diseases, whereas the site specificity of ubiquitination is mainly decided by ubiquitin-protein ligases (E3s). Although numerous ubiquitination predictors have been developed, computational prediction of E3-specific ubiquitination sites is still a great challenge. Here, we carefully reviewed the existing tools for the prediction of general ubiquitination sites. Also, we developed a tool named GPS-Uber for the prediction of general and E3-specific ubiquitination sites. From the literature, we manually collected 1311 experimentally identified site-specific E3-substrate relations, which were classified into different clusters based on corresponding E3s at different levels. To predict general ubiquitination sites, we integrated 10 types of sequence and structure features, as well as three types of algorithms including penalized logistic regression, deep neural network and convolutional neural network. Compared with other existing tools, the general model in GPS-Uber exhibited a highly competitive accuracy, with an area under curve values of 0.7649. Then, transfer learning was adopted for each E3 cluster to construct E3-specific models, and in total 112 individual E3-specific predictors were implemented. Using GPS-Uber, we conducted a systematic prediction of human cancer-associated ubiquitination events, which could be helpful for further experimental consideration. GPS-Uber will be regularly updated, and its online service is free for academic research at http://gpsuber.biocuckoo.cn/. Xiaodan Tan, Dachao Tang, Yujie Gou, Wanshan Ning, Shaofeng Lin, Weizhi Zhang 0002, Yu Xue 0001 |
Briefings Bioinform. | 7 |
| 2021 | EPSD: a well-annotated data resource of protein phosphorylation sites in eukaryotesabstractAs an important post-translational modification (PTM), protein phosphorylation is involved in the regulation of almost all of biological processes in eukaryotes. Due to the rapid progress in mass spectrometry-based phosphoproteomics, a large number of phosphorylation sites (p-sites) have been characterized but remain to be curated. Here, we briefly summarized the current progresses in the development of data resources for the collection, curation, integration and annotation of p-sites in eukaryotic proteins. Also, we designed the eukaryotic phosphorylation site database (EPSD), which contained 1 616 804 experimentally identified p-sites in 209 326 phosphoproteins from 68 eukaryotic species. In EPSD, we not only collected 1 451 629 newly identified p-sites from high-throughput (HTP) phosphoproteomic studies, but also integrated known p-sites from 13 additional databases. Moreover, we carefully annotated the phosphoproteins and p-sites of eight model organisms by integrating the knowledge from 100 additional resources that covered 15 aspects, including phosphorylation regulator, genetic variation and mutation, functional annotation, structural annotation, physicochemical property, functional domain, disease-associated information, protein-protein interaction, drug-target relation, orthologous information, biological pathway, transcriptional regulator, mRNA expression, protein expression/proteomics and subcellular localization. We anticipate that the EPSD can serve as a useful resource for further analysis of eukaryotic phosphorylation. With a data volume of 14.1 GB, EPSD is free for all users at http://epsd.biocuckoo.cn/. Shaofeng Lin, Jiaqi Zhou 0003, Chen Ruan, Yiran Tu, Lan Yao, Yu Xue 0001 |
Briefings Bioinform. | 1 |