VLDB 2026 Research / reviewers in the wild / expert
Yi Shi 0007
dblp:00/3680-7
· DBLP profile ↗
8ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-5279-3239ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Theory of computation · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive Context Length Optimization with Low-Frequency Truncation for Multi-Agent Reinforcement LearningabstractRecently, deep multi-agent reinforcement learning (MARL) has demonstrated promising performance for solving challenging tasks, such as long-term dependencies and non-Markovian environments. Its success is partly attributed to conditioning policies on large fixed context length. However, such large fixed context lengths may lead to limited exploration efficiency and redundant information. In this paper, we propose a novel MARL framework to obtain adaptive and effective contextual information. Specifically, we design a central agent that dynamically optimizes context length via temporal gradient analysis, enhancing exploration to facilitate convergence to global optima in MARL. Furthermore, to enhance the adaptive optimization capability of the context length, we present an efficient input representation for the central agent, which effectively filters redundant information. By leveraging a Fourier-based low-frequency truncation method, we extract global temporal trends across decentralized agents, providing an effective and efficient representation of the MARL environment. Extensive experiments demonstrate that the proposed method achieves state-of-the-art (SOTA) performance on long-term dependency tasks, including PettingZoo, MiniGrid, Google Research Football (GRF), and StarCraft Multi-Agent Challenge v2 (SMACv2). Wenchang Duan, Yaoliang Yu, Jiwan He, Yi Shi 0007 |
NeurIPS | 4 |
| 2025 | PSTP: accurate residue-level phase separation prediction using protein conformational and language model embeddingsabstractPhase separation (PS) is essential in cellular processes and disease mechanisms, highlighting the need for predictive algorithms to analyze uncharacterized sequences and accelerate experimental validation. Current high-accuracy methods often rely on extensive annotations or handcrafted features, limiting their generalizability to sequences lacking such annotations and making it difficult to identify key protein regions involved in PS. We introduce Phase Separation's Transfer-learning Prediction (PSTP), which combines conformational embeddings with large language model embeddings, enabling state-of-the-art PS predictions from protein sequences alone. PSTP performs well across various prediction scenarios and shows potential for predicting novel-designed artificial proteins. Additionally, PSTP provides residue-level predictions that are highly correlated with experimentally validated PS regions. By analyzing 160 000+ variants, PSTP characterizes the strong link between the incidence of pathogenic variants and residue-level PS propensities in unconserved intrinsically disordered regions, offering insights into underexplored mutation effects. PSTP's sliding-window optimization reduces its memory usage to a few hundred megabytes, facilitating rapid execution on typical CPUs and GPUs. Offered via both a web server and an installable Python package, PSTP provides a versatile tool for decoding protein PS behavior and supporting disease-focused research. Mofan Feng, Liangjie Liu, Zhuo-Ning Xian, Xiaoxi Wei, Wenqian Yan, Yi Shi 0007, Guang He |
Briefings Bioinform. | 8 |
| 2022 | 3D genome assisted protein-protein interaction prediction
Zehua Guo 0004, Liangjie Liu, Mofan Feng, Runqiu Chi, Xianbin Su, Luming Meng, Dan Cao, Guang He, Yi Shi 0007 |
Future Gener. Comput. Syst. | 16 |
| 2022 | A simple linear time algorithm to solve the MIST problem on interval graphs
Peng Li 0065, Jianhui Shang, Yi Shi 0007 |
Theor. Comput. Sci. | 3 |
| 2021 | Improving Protein-protein Interaction Prediction by Incorporating 3D Genome Information
Zehua Guo 0004, Liangjie Liu, Xianbin Su, Mofan Feng, Runqiu Chi, Luming Meng, Guang He, Yi Shi 0007 |
ISBRA | 11 |
| 2021 | The longest cycle problem is polynomial on interval graphs
Jianhui Shang, Peng Li 0065, Yi Shi 0007 |
Theor. Comput. Sci. | 3 |
| 2020 | DeepAntigen: a novel method for neoantigen prioritization via 3D genome and deep sparse learningabstractMOTIVATION: The mutations of cancers can encode the seeds of their own destruction, in the form of T-cell recognizable immunogenic peptides, also known as neoantigens. It is computationally challenging, however, to accurately prioritize the potential neoantigen candidates according to their ability of activating the T-cell immunoresponse, especially when the somatic mutations are abundant. Although a few neoantigen prioritization methods have been proposed to address this issue, advanced machine learning model that is specifically designed to tackle this problem is still lacking. Moreover, none of the existing methods considers the original DNA loci of the neoantigens in the perspective of 3D genome which may provide key information for inferring neoantigens' immunogenicity. RESULTS: In this study, we discovered that DNA loci of the immunopositive and immunonegative MHC-I neoantigens have distinct spatial distribution patterns across the genome. We therefore used the 3D genome information along with an ensemble pMHC-I coding strategy, and developed a group feature selection-based deep sparse neural network model (DNN-GFS) that is optimized for neoantigen prioritization. DNN-GFS demonstrated increased neoantigen prioritization power comparing to existing sequence-based approaches. We also developed a webserver named deepAntigen (http://yishi.sjtu.edu.cn/deepAntigen) that implements the DNN-GFS as well as other machine learning methods. We believe that this work provides a new perspective toward more accurate neoantigen prediction which eventually contribute to personalized cancer immunotherapy. AVAILABILITY AND IMPLEMENTATION: Data and implementation are available on webserver: http://yishi.sjtu.edu.cn/deepAntigen. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi Shi 0007, Zehua Guo 0004, Xianbin Su, Luming Meng, Minhua Zheng, Xueyin Shang, Wangqiu Cheng, Yaoliang Yu, Yujia Cai, Chaoyi Zhang, Tom Weidong Cai, Guang He, Zeguang Han |
Bioinform. | 1 |
| 2016 | DeepGene: an advanced cancer type classifier based on deep learning and somatic point mutationsabstractBACKGROUND: With the developments of DNA sequencing technology, large amounts of sequencing data have become available in recent years and provide unprecedented opportunities for advanced association studies between somatic point mutations and cancer types/subtypes, which may contribute to more accurate somatic point mutation based cancer classification (SMCC). However in existing SMCC methods, issues like high data sparsity, small volume of sample size, and the application of simple linear classifiers, are major obstacles in improving the classification performance. RESULTS: To address the obstacles in existing SMCC studies, we propose DeepGene, an advanced deep neural network (DNN) based classifier, that consists of three steps: firstly, the clustered gene filtering (CGF) concentrates the gene data by mutation occurrence frequency, filtering out the majority of irrelevant genes; secondly, the indexed sparsity reduction (ISR) converts the gene data into indexes of its non-zero elements, thereby significantly suppressing the impact of data sparsity; finally, the data after CGF and ISR is fed into a DNN classifier, which extracts high-level features for accurate classification. Experimental results on our curated TCGA-DeepGene dataset, which is a reformulated subset of the TCGA dataset containing 12 selected types of cancer, show that CGF, ISR and DNN all contribute in improving the overall classification performance. We further compare DeepGene with three widely adopted classifiers and demonstrate that DeepGene has at least 24% performance improvement in terms of testing accuracy. CONCLUSIONS: Based on deep learning and somatic point mutation data, we devise DeepGene, an advanced cancer type classifier, which addresses the obstacles in existing SMCC studies. Experiments indicate that DeepGene outperforms three widely adopted existing classifiers, which is mainly attributed to its deep learning module that is able to extract the high level features between combinatorial somatic point mutations and cancer types. Yuchen Yuan, Yi Shi 0007, ChangYang Li, Jinman Kim, Tom Weidong Cai, Zeguang Han, David Dagan Feng |
BMC Bioinform. | 2 |