VLDB 2026 Research / reviewers in the wild / expert
Sijie Yao
dblp:238/1978
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-4228-0832ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
biomarker discovery |
0.6 | 1 | 2022 | Efficient gradient boosting for prognostic biomarker discovery · Bioinform. 2022 |
Bioinformatics and computational biology › biomarker discovery
prognostic biomarker identification |
0.6 | 1 | 2022 | Efficient gradient boosting for prognostic biomarker discovery · Bioinform. 2022 |
Methods — techniques the papers use, named apart from their topics
light gradient boosting · 0.6gradient boosting decision tree · 0.6extreme gradient boosting · 0.6cox regression · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Shared-weight graph framework for comprehensive protein stability prediction across diverse mutation typesabstractResearch on protein stability changes is vital for understanding disease mechanisms and optimizing industrial enzymes. Protein thermal stability can be modified by variants leading to changes in ΔΔG values between wild-type and mutant proteins. Despite advances, most models focus on single-point mutations, overlooking multipoint and indel mutations. Typically, the single-point mutation is expected to have a relatively limited impact on the function of a protein, necessitating more drastic modifications to meet new challenges. Current methods for multipoint mutations yield poor results, and no method exists for any length of indel mutations. To address this, we introduce UniMutStab, a shared-graph convolutional network leveraging protein language models and residue interaction networks to access any type of mutation. An embedded edge weight module enhances the integration of residue node features and interactions, improving prediction accuracy. Trained on the "Mega-scale" dataset with ~780 000 mutations, UniMutStab surpasses existing methods in predicting protein stability changes. It is a purely sequence-based approach to predict arbitrary mutation types, demonstrating robust generalization across multiple tasks and potentially contributing significantly to protein engineering, personalized therapeutics, and diagnostic methodologies. Sijie Yao, Long Fan |
Briefings Bioinform. | 2 |
| 2022 | Efficient gradient boosting for prognostic biomarker discoveryabstractMOTIVATION: A gradient boosting decision tree (GBDT) is a powerful ensemble machine-learning method that has the potential to accelerate biomarker discovery from high-dimensional molecular data. Recent algorithmic advances, such as extreme gradient boosting (XGB) and light gradient boosting (LGB), have rendered the GBDT training more efficient, scalable and accurate. However, these modern techniques have not yet been widely adopted in discovering biomarkers for censored survival outcomes, which are key clinical outcomes or endpoints in cancer studies. RESULTS: In this paper, we present a new R package 'Xsurv' as an integrated solution that applies two modern GBDT training frameworks namely, XGB and LGB, for the modeling of right-censored survival outcomes. Based on our simulations, we benchmark the new approaches against traditional methods including the stepwise Cox regression model and the original gradient boosting function implemented in the package 'gbm'. We also demonstrate the application of Xsurv in analyzing a melanoma methylation dataset. Together, these results suggest that Xsurv is a useful and computationally viable tool for screening a large number of prognostic candidate biomarkers, which may facilitate future translational and clinical research. AVAILABILITY AND IMPLEMENTATION: 'Xsurv' is freely available as an R package at: https://github.com/topycyao/Xsurv. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Kaiqiao Li, Sijie Yao, Biwei Cao, Denise Kalos, Pei Fen Kuan, Ruoqing Zhu |
Bioinform. | 2 |
| 2021 | Recognition and counting of wheat mites in wheat fields by a three-step deep learning method
Peng Chen 0001, Weilu Li, Sijie Yao, Chun Ma, Jun Zhang 0011, Bing Wang 0004, Chun-Hou Zheng 0001, Chengjun Xie |
Neurocomputing | 3 |