VLDB 2026 Research / reviewers in the wild / expert
Sizhe Liu
dblp:81/11466
· DBLP profile ↗
11ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Embedded and real-time systems · 66% Performance modeling and evaluation · 27% Electronic design automation · 8% | |
| Databases, data mining, and information retrieval
1 paper |
Machine learning and data management · 100% |
Topics — the 16 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › proteomics › peptide sequencing
de novo peptide sequencing |
1.6 | 2 | 2025 | Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo · ICLR 2025 NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024 |
Bioinformatics and computational biology
proteomics |
1.6 | 2 | 2025 | Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo · ICLR 2025 NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024 |
Bioinformatics and computational biology › sequence analysis
database search |
0.9 | 1 | 2025 | Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo · ICLR 2025 |
Bioinformatics and computational biology
drug discovery |
0.9 | 1 | 2025 | SP-DTI: subpocket-informed transformer for drug-target interaction prediction · Bioinform. 2025 |
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction |
0.9 | 1 | 2025 | SP-DTI: subpocket-informed transformer for drug-target interaction prediction · Bioinform. 2025 |
Embedded and real-time systems › real-time scheduling › deadline scheduling
EDF scheduling |
0.9 | 1 | 2025 | Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025 |
Embedded and real-time systems › real-time scheduling
preemptive scheduling |
0.9 | 1 | 2025 | Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025 |
Embedded and real-time systems › real-time scheduling
priority scheduling |
0.9 | 1 | 2025 | Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025 |
Embedded and real-time systems
real-time scheduling |
0.9 | 1 | 2025 | Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025 |
Bioinformatics and computational biology › machine learning for biology
molecular relational learning |
0.8 | 1 | 2024 | FlexMol: A Flexible Toolkit for Benchmarking Molecular Relational Learning · NeurIPS 2024 |
Performance modeling and evaluation
benchmarking |
0.8 | 1 | 2024 | NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024 |
Performance modeling and evaluation › benchmarking › machine learning benchmarking
deep learning benchmarks |
0.8 | 1 | 2024 | NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024 |
Embedded and real-time systems › cyber-physical system platforms
robot operating system |
0.3 | 1 | 2025 | Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025 |
Electronic design automation › hardware verification and test
diagnosis |
0.1 | 1 | 2012 | Test-data volume optimization for diagnosis · DAC 2012 |
Electronic design automation
hardware verification and test |
0.1 | 1 | 2012 | Test-data volume optimization for diagnosis · DAC 2012 |
Electronic design automation › hardware verification and test
test data volume reduction |
0.1 | 1 | 2012 | Test-data volume optimization for diagnosis · DAC 2012 |
Methods — techniques the papers use, named apart from their topics
deep learning · 2.4tandem mass spectrometry · 1.5protein encoders · 1.5drug encoders · 1.5transformer · 0.9pre-trained language model · 0.9graph neural network · 0.9interaction layers · 0.8interaction layer · 0.8statistical learning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semisupervised Cross-Domain Capacity Prediction for Batteries via Granular Modeling and Confidence Aware PseudolabelingabstractIn practical applications, the degradation behavior of lithium-ion batteries exhibits significant differences due to variations in operating conditions. Meanwhile, the scarcity of labeled data poses considerable challenges for capacity prediction in terms of both accuracy and generalization. To address these issues, this article proposes a cross-domain semisupervised capacity prediction framework that integrates multigranularity feature modeling with a confidence controlled pseudolabel selection mechanism. Specifically, the proposed method enhances the model’s ability to capture the granularity of nonlinear degradation trends in battery capacity, thereby improving prediction accuracy and stability. In addition, a pseudolabel learning strategy based on confidence filtering and stagewise regulation is designed to dynamically guide high-quality pseudolabels in the target domain into training, effectively reducing the risk of noisy label propagation. Experiments conducted on eight tasks across two heterogeneous battery datasets demonstrate R$^{2}$improvements of 1.3%–8.7% and Mean Absolute Error (MAE) reductions of 38%–80%, validating the practical potential of the proposed method under complex degradation scenarios. Sizhe Liu, Dezhi Xu, Chao Shen 0001, Yujian Ye, Chengxi Zhang, Yan Wang 0049 |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo
Jun Xia 0001, Sizhe Liu, Shaorong Chen, Hongxin Xiang, Zicheng Liu 0006, Yue Liu 0008, Stan Z. Li |
ICLR | 2 |
| 2025 | Large Language Models Are Cross-Lingual Knowledge-Free ReasonersabstractPeng Hu, Sizhe Liu, Changjiang Gao, Xin Huang, Xue Han, Junlan Feng, Chao Deng, Shujian Huang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Sizhe Liu, Changjiang Gao, Xue Han 0018, Junlan Feng, Chao Deng 0002, Shujian Huang |
NAACL (Long Papers) | 2 |
| 2025 | Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2abstractThe Robot Operating System 2 (ROS) is heavily used in autonomous systems due to its large ecosystem and modular design. However, ROS remains problematic for real-time applications despite prior efforts to improve its real-time capabilities. Many of these problems are rooted in the implementation of the ROS executor, which does not support preemption or userspecified priorities. These properties are fundamental to real-time scheduling and desirable for general systems to reduce latency. ROS variants in prior work have separately supported the preemption and prioritization of callbacks but place restrictions on the application, preventing their adoption in a real-world workload. This paper addresses these deficiencies with a novel executor framework,$\operatorname{ROS}^{\text{RT}}$, which is compatible with any type of ROS application while supporting preemptive, priority-driven scheduling. Additionally, to support flexible EDF scheduling in ROS${ }^{\text {RT }}$, a custom EDF scheduler implementation is proposed using the new Linux scheduling class SCHED_EXT. ROS${ }^{\text {RT }}$s flexibility and real-time compatibility do not come at the cost of overheads: it achieves a significant decrease in publisher-tosubscriber overhead from the native ROS executor. Finally, this paper concludes with a case study performed on the Autoware Reference System, which simulates the execution of the LiDAR module in an autonomous driving application, demonstrating the ability of$\operatorname{ROS}^{\text{RT}}$at real-time scheduling in real-world scenarios. Sizhe Liu, Rohan Wagle, Shareef Ahmed, Zelin Tong, James H. Anderson |
RTSS | 1 |
| 2025 | SP-DTI: subpocket-informed transformer for drug-target interaction predictionabstractMOTIVATION: Drug-target interaction (DTI) prediction is crucial for drug discovery, significantly reducing costs and time in experimental searches across vast drug compound spaces. While deep learning has advanced DTI prediction accuracy, challenges remain: (i) existing methods often lack generalizability, with performance dropping significantly on unseen proteins and cross-domain settings; and (ii) current molecular relational learning often overlooks subpocket-level interactions, which are vital for a detailed understanding of binding sites. RESULTS: We introduce SP-DTI, a subpocket-informed transformer model designed to address these challenges through: (i) detailed subpocket analysis using the Cavity Identification and Analysis Routine for interaction modeling at both global and local levels, and (ii) integration of pre-trained language models into graph neural networks to encode drugs and proteins, enhancing generalizability to unlabeled data. Benchmark evaluations show that SP-DTI consistently outperforms state-of-the-art models, achieving an area under the receiver operating characteristic curve of 0.873 in unseen protein settings, an 11% improvement over the best baseline. AVAILABILITY AND IMPLEMENTATION: The model scripts are available at https://github.com/Steven51516/SP-DTI. Sizhe Liu, Haofeng Xu, Jun Xia 0001, Stan Z. Li |
Bioinform. | 1 |
| 2025 | Intelligent fault diagnosis for unbalanced battery data using adversarial domain expansion and enhanced stochastic configuration networks
Sizhe Liu, Dezhi Xu, Yujian Ye, Tinglong Pan |
Inf. Sci. | 1 |
| 2024 | Autonomy Today: Many Delay-Prone Black Boxes
Sizhe Liu, Rohan Wagle, James H. Anderson, Ming Yang 0036, Yunhua Li |
ECRTS | 1 |
| 2024 | FlexMol: A Flexible Toolkit for Benchmarking Molecular Relational LearningabstractMolecular relational learning (MRL) is crucial for understanding the interaction behaviors between molecular pairs, a critical aspect of drug discovery and development. However, the large feasible model space of MRL poses significant challenges to benchmarking, and existing MRL frameworks face limitations in flexibility and scope. To address these challenges, avoid repetitive coding efforts, and ensure fair comparison of models, we introduce FlexMol, a comprehensive toolkit designed to facilitate the construction and evaluation of diverse model architectures across various datasets and performance metrics. FlexMol offers a robust suite of preset model components, including 16 drug encoders, 13 protein sequence encoders, 9 protein structure encoders, and 7 interaction layers. With its easy-to-use API and flexibility, FlexMol supports the dynamic construction of over 70, 000 distinct combinations of model architectures. Additionally, we provide detailed benchmark results and code examples to demonstrate FlexMol’s effectiveness in simplifying and standardizing MRL model development and comparison. FlexMol is open-sourced and available at https://github.com/Steven51516/FlexMol. Sizhe Liu, Jun Xia 0001, Lecheng Zhang, Yue Liu 0008, Wenjie Du 0003, Zhangyang Gao, Bozhen Hu, Cheng Tan 0012, Hongxin Xiang, Stan Z. Li |
NeurIPS | 1 |
| 2024 | NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in ProteomicsabstractTandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the analysis of protein composition in biological tissues. Many deep learning methods have been developed for \emph{de novo} peptide sequencing task, i.e., predicting the peptide sequence for the observed mass spectrum. However, two key challenges seriously hinder the further research of this important task. Firstly, since there is no consensus for the evaluation datasets, the empirical results in different research papers are often not comparable, leading to unfair comparison. Secondly, the current methods are usually limited to amino acid-level or peptide-level precision and recall metrics. In this work, we present the first unified benchmark NovoBench for \emph{de novo} peptide sequencing, which comprises diverse mass spectrum data, integrated models, and comprehensive evaluation metrics. Recent impressive methods, including DeepNovo, PointNovo, Casanovo, InstaNovo, AdaNovo and $\pi$-HelixNovo are integrated into our framework. In addition to amino acid-level and peptide-level precision and recall, we also evaluate the models' performance in terms of identifying post-tranlational modifications (PTMs), efficiency and robustness to peptide length, noise peaks and missing fragment ratio, which are important influencing factors while seldom be considered. Leveraging this benchmark, we conduct a large-scale study of current methods, report many insightful findings that open up new possibilities for future development. The benchmark is open-sourced to facilitate future research and application. The code is available at \url{https://github.com/Westlake-OmicsAI/NovoBench}. Shaorong Chen, Jun Xia 0001, Sizhe Liu, Tianze Ling, Wenjie Du 0003, Yue Liu 0008, Jianwei Yin, Stan Z. Li |
NeurIPS | 4 |
| 2024 | Adaptive fusion transfer learning-based digital multitwin-assised intelligent fault diagnosis
Sizhe Liu, Yongsheng Qi, Liqiang Liu |
Knowl. Based Syst. | 1 |
| 2012 | Test-data volume optimization for diagnosisabstractTest data collection for a failing integrated circuit (IC) can be very expensive and time consuming. Many companies now collect a fix amount of test data regardless of the failure characteristics. As a result, limited data collection could lead to inaccurate diagnosis, while an excessive amount increases the cost not only in terms of unnecessary test data collection but also increased cost for test execution and data-storage. In this work, the objective is to develop a method for predicting the precise amount of test data necessary to produce an accurate diagnosis. By analyzing the failing outputs of an IC during its actual test, the developed method dynamically determines which failing test pattern to terminate testing, producing an amount of test data that is sufficient for an accurate diagnosis analysis. The method leverages several statistical learning techniques, and is evaluated using actual data from a population of failing chips and five standard benchmarks. Experiments demonstrate that test-data collection can be reduced by > 30% (as compared to collecting the full-failure response) while at the same time ensuring >90% diagnosis accuracy. Prematurely terminating test-data collection at fixed levels (e.g., 100 failing bits) is also shown to negatively impact diagnosis accuracy. Osei Poku, Xiaochun Yu, Sizhe Liu, Ibrahima Komara, R. D. (Shawn) Blanton |
DAC | 4 |