Sizhe Liu

dblp:81/11466 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Embedded and real-time systems · 66% Performance modeling and evaluation · 27% Electronic design automation · 8%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 16 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › proteomics › peptide sequencing
de novo peptide sequencing
1.622025
Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo · ICLR 2025
NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024
Bioinformatics and computational biology
proteomics
1.622025
Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo · ICLR 2025
NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024
Bioinformatics and computational biology › sequence analysis
database search
0.912025
Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo · ICLR 2025
Bioinformatics and computational biology
drug discovery
0.912025
SP-DTI: subpocket-informed transformer for drug-target interaction prediction · Bioinform. 2025
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction
0.912025
SP-DTI: subpocket-informed transformer for drug-target interaction prediction · Bioinform. 2025
Embedded and real-time systems › real-time scheduling › deadline scheduling
EDF scheduling
0.912025
Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025
Embedded and real-time systems › real-time scheduling
preemptive scheduling
0.912025
Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025
Embedded and real-time systems › real-time scheduling
priority scheduling
0.912025
Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025
Embedded and real-time systems
real-time scheduling
0.912025
Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025
Bioinformatics and computational biology › machine learning for biology
molecular relational learning
0.812024
FlexMol: A Flexible Toolkit for Benchmarking Molecular Relational Learning · NeurIPS 2024
Performance modeling and evaluation
benchmarking
0.812024
NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024
Performance modeling and evaluation › benchmarking › machine learning benchmarking
deep learning benchmarks
0.812024
NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics · NeurIPS 2024
Embedded and real-time systems › cyber-physical system platforms
robot operating system
0.312025
Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2 · RTSS 2025
Electronic design automation › hardware verification and test
diagnosis
0.112012
Test-data volume optimization for diagnosis · DAC 2012
Electronic design automation
hardware verification and test
0.112012
Test-data volume optimization for diagnosis · DAC 2012
Electronic design automation › hardware verification and test
test data volume reduction
0.112012
Test-data volume optimization for diagnosis · DAC 2012

Methods — techniques the papers use, named apart from their topics

deep learning · 2.4tandem mass spectrometry · 1.5protein encoders · 1.5drug encoders · 1.5transformer · 0.9pre-trained language model · 0.9graph neural network · 0.9interaction layers · 0.8interaction layer · 0.8statistical learning · 0.1
YearPublicationVenuePosition
2026 Semisupervised Cross-Domain Capacity Prediction for Batteries via Granular Modeling and Confidence Aware Pseudolabeling
abstract
In practical applications, the degradation behavior of lithium-ion batteries exhibits significant differences due to variations in operating conditions. Meanwhile, the scarcity of labeled data poses considerable challenges for capacity prediction in terms of both accuracy and generalization. To address these issues, this article proposes a cross-domain semisupervised capacity prediction framework that integrates multigranularity feature modeling with a confidence controlled pseudolabel selection mechanism. Specifically, the proposed method enhances the model’s ability to capture the granularity of nonlinear degradation trends in battery capacity, thereby improving prediction accuracy and stability. In addition, a pseudolabel learning strategy based on confidence filtering and stagewise regulation is designed to dynamically guide high-quality pseudolabels in the target domain into training, effectively reducing the risk of noisy label propagation. Experiments conducted on eight tasks across two heterogeneous battery datasets demonstrate R$^{2}$improvements of 1.3%–8.7% and Mean Absolute Error (MAE) reductions of 38%–80%, validating the practical potential of the proposed method under complex degradation scenarios.
Sizhe Liu, Dezhi Xu, Chao Shen 0001, Yujian Ye, Chengxi Zhang, Yan Wang 0049
IEEE Trans. Ind. Informatics1
2025 Bridging the Gap between Database Search and De Novo Peptide Sequencing with SearchNovo
Jun Xia 0001, Sizhe Liu, Shaorong Chen, Hongxin Xiang, Zicheng Liu 0006, Yue Liu 0008, Stan Z. Li
ICLR2
2025 Large Language Models Are Cross-Lingual Knowledge-Free Reasoners
abstract
Peng Hu, Sizhe Liu, Changjiang Gao, Xin Huang, Xue Han, Junlan Feng, Chao Deng, Shujian Huang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sizhe Liu, Changjiang Gao, Xue Han 0018, Junlan Feng, Chao Deng 0002, Shujian Huang
NAACL (Long Papers)2
2025 Ros ${ }^{\text {RT }}$: Enabling Flexible Scheduling in Ros 2
abstract
The Robot Operating System 2 (ROS) is heavily used in autonomous systems due to its large ecosystem and modular design. However, ROS remains problematic for real-time applications despite prior efforts to improve its real-time capabilities. Many of these problems are rooted in the implementation of the ROS executor, which does not support preemption or userspecified priorities. These properties are fundamental to real-time scheduling and desirable for general systems to reduce latency. ROS variants in prior work have separately supported the preemption and prioritization of callbacks but place restrictions on the application, preventing their adoption in a real-world workload. This paper addresses these deficiencies with a novel executor framework,$\operatorname{ROS}^{\text{RT}}$, which is compatible with any type of ROS application while supporting preemptive, priority-driven scheduling. Additionally, to support flexible EDF scheduling in ROS${ }^{\text {RT }}$, a custom EDF scheduler implementation is proposed using the new Linux scheduling class SCHED_EXT. ROS${ }^{\text {RT }}$s flexibility and real-time compatibility do not come at the cost of overheads: it achieves a significant decrease in publisher-tosubscriber overhead from the native ROS executor. Finally, this paper concludes with a case study performed on the Autoware Reference System, which simulates the execution of the LiDAR module in an autonomous driving application, demonstrating the ability of$\operatorname{ROS}^{\text{RT}}$at real-time scheduling in real-world scenarios.
Sizhe Liu, Rohan Wagle, Shareef Ahmed, Zelin Tong, James H. Anderson
RTSS1
2025 SP-DTI: subpocket-informed transformer for drug-target interaction prediction
abstract
MOTIVATION: Drug-target interaction (DTI) prediction is crucial for drug discovery, significantly reducing costs and time in experimental searches across vast drug compound spaces. While deep learning has advanced DTI prediction accuracy, challenges remain: (i) existing methods often lack generalizability, with performance dropping significantly on unseen proteins and cross-domain settings; and (ii) current molecular relational learning often overlooks subpocket-level interactions, which are vital for a detailed understanding of binding sites. RESULTS: We introduce SP-DTI, a subpocket-informed transformer model designed to address these challenges through: (i) detailed subpocket analysis using the Cavity Identification and Analysis Routine for interaction modeling at both global and local levels, and (ii) integration of pre-trained language models into graph neural networks to encode drugs and proteins, enhancing generalizability to unlabeled data. Benchmark evaluations show that SP-DTI consistently outperforms state-of-the-art models, achieving an area under the receiver operating characteristic curve of 0.873 in unseen protein settings, an 11% improvement over the best baseline. AVAILABILITY AND IMPLEMENTATION: The model scripts are available at https://github.com/Steven51516/SP-DTI.
Sizhe Liu, Haofeng Xu, Jun Xia 0001, Stan Z. Li
Bioinform.1
2025 Intelligent fault diagnosis for unbalanced battery data using adversarial domain expansion and enhanced stochastic configuration networks
Sizhe Liu, Dezhi Xu, Yujian Ye, Tinglong Pan
Inf. Sci.1
2024 Autonomy Today: Many Delay-Prone Black Boxes
Sizhe Liu, Rohan Wagle, James H. Anderson, Ming Yang 0036, Yunhua Li
ECRTS1
2024 FlexMol: A Flexible Toolkit for Benchmarking Molecular Relational Learning
abstract
Molecular relational learning (MRL) is crucial for understanding the interaction behaviors between molecular pairs, a critical aspect of drug discovery and development. However, the large feasible model space of MRL poses significant challenges to benchmarking, and existing MRL frameworks face limitations in flexibility and scope. To address these challenges, avoid repetitive coding efforts, and ensure fair comparison of models, we introduce FlexMol, a comprehensive toolkit designed to facilitate the construction and evaluation of diverse model architectures across various datasets and performance metrics. FlexMol offers a robust suite of preset model components, including 16 drug encoders, 13 protein sequence encoders, 9 protein structure encoders, and 7 interaction layers. With its easy-to-use API and flexibility, FlexMol supports the dynamic construction of over 70, 000 distinct combinations of model architectures. Additionally, we provide detailed benchmark results and code examples to demonstrate FlexMol’s effectiveness in simplifying and standardizing MRL model development and comparison. FlexMol is open-sourced and available at https://github.com/Steven51516/FlexMol.
Sizhe Liu, Jun Xia 0001, Lecheng Zhang, Yue Liu 0008, Wenjie Du 0003, Zhangyang Gao, Bozhen Hu, Cheng Tan 0012, Hongxin Xiang, Stan Z. Li
NeurIPS1
2024 NovoBench: Benchmarking Deep Learning-based \emph{De Novo} Sequencing Methods in Proteomics
abstract
Tandem mass spectrometry has played a pivotal role in advancing proteomics, enabling the analysis of protein composition in biological tissues. Many deep learning methods have been developed for \emph{de novo} peptide sequencing task, i.e., predicting the peptide sequence for the observed mass spectrum. However, two key challenges seriously hinder the further research of this important task. Firstly, since there is no consensus for the evaluation datasets, the empirical results in different research papers are often not comparable, leading to unfair comparison. Secondly, the current methods are usually limited to amino acid-level or peptide-level precision and recall metrics. In this work, we present the first unified benchmark NovoBench for \emph{de novo} peptide sequencing, which comprises diverse mass spectrum data, integrated models, and comprehensive evaluation metrics. Recent impressive methods, including DeepNovo, PointNovo, Casanovo, InstaNovo, AdaNovo and $\pi$-HelixNovo are integrated into our framework. In addition to amino acid-level and peptide-level precision and recall, we also evaluate the models' performance in terms of identifying post-tranlational modifications (PTMs), efficiency and robustness to peptide length, noise peaks and missing fragment ratio, which are important influencing factors while seldom be considered. Leveraging this benchmark, we conduct a large-scale study of current methods, report many insightful findings that open up new possibilities for future development. The benchmark is open-sourced to facilitate future research and application. The code is available at \url{https://github.com/Westlake-OmicsAI/NovoBench}.
Shaorong Chen, Jun Xia 0001, Sizhe Liu, Tianze Ling, Wenjie Du 0003, Yue Liu 0008, Jianwei Yin, Stan Z. Li
NeurIPS4
2024 Adaptive fusion transfer learning-based digital multitwin-assised intelligent fault diagnosis
Sizhe Liu, Yongsheng Qi, Liqiang Liu
Knowl. Based Syst.1
2012 Test-data volume optimization for diagnosis
abstract
Test data collection for a failing integrated circuit (IC) can be very expensive and time consuming. Many companies now collect a fix amount of test data regardless of the failure characteristics. As a result, limited data collection could lead to inaccurate diagnosis, while an excessive amount increases the cost not only in terms of unnecessary test data collection but also increased cost for test execution and data-storage. In this work, the objective is to develop a method for predicting the precise amount of test data necessary to produce an accurate diagnosis. By analyzing the failing outputs of an IC during its actual test, the developed method dynamically determines which failing test pattern to terminate testing, producing an amount of test data that is sufficient for an accurate diagnosis analysis. The method leverages several statistical learning techniques, and is evaluated using actual data from a population of failing chips and five standard benchmarks. Experiments demonstrate that test-data collection can be reduced by > 30% (as compared to collecting the full-failure response) while at the same time ensuring >90% diagnosis accuracy. Prematurely terminating test-data collection at fixed levels (e.g., 100 failing bits) is also shown to negatively impact diagnosis accuracy.
Osei Poku, Xiaochun Yu, Sizhe Liu, Ibrahima Komara, R. D. (Shawn) Blanton
DAC4