VLDB 2026 Research / reviewers in the wild / expert
Sizhe Qiu
dblp:357/9604
· DBLP profile ↗
5ranked-venue papers
5as first author
5since 2021 · last 2026
0000-0002-1936-1223ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Response to "Addressing flaws in the Seq2Topt dataset for the prediction of enzyme optimal temperature"abstractTo the Editor, We would like to thank the rigorous dataset inspection work conducted by the author of “Addressing flaws in the Seq2Topt dataset for the prediction of enzyme optimal temperature”. The dataset of enzyme optimal temperature (Topt) used to develop Seq2Topt [1] was obtained from https://github.com/jafetgado/tomer, which was curated by Li et al. 2019 [2] and Gado et al. 2020 [3] from the BRENDA database [4]. In “Addressing flaws in the Seq2Topt dataset for the prediction of enzyme optimal temperature,” the author stated that the Topt values of F9FU71, A9U908, J7HET3, P94368, Q5YXQ1, Q83V33, Q50539, Q50538, Q2WCS9 are erroneously annotated in the dataset used by Seq2Topt. Here, we presented all references associated with these data entries in the BRENDA database in Table 1, and the Topt values of F9FU71, P94368, Q5YXQ1, Q83V33, Q50539, and Q50538 were not found in relevant references [5–9]. Dhar et al. 2013 [10] stated that the Topt values of A9U908 (sodA) and J7HET3 (sodB) were 4°C, but we agree with the critique that no enzyme activity measurements at lower temperatures were conducted to confirm that 4°C was the Topt value of A9U908 and J7HET3. For Q2WCS9, De Angelis et al. 2010 [11] reported that 13°C was the measured Topt value, whereas the critique letter did not provide a source for the claimed measured Topt value of 37°C. Selected nine entries from Seq2Topt dataset. Overall, we concur with the critique that the dataset used to develop enzyme Topt predictive models (and machine learning models for other protein properties) should be carefully reviewed and revised to ensure validity, rather than simply extracting data from established databases such as BRENDA. While such curation may be costly, it remains necessary until a high level of validity is automatically guaranteed by these biochemical databases. Accordingly, future updates of Seq2Topt will be based on a more rigorously vetted dataset. The author declares that there is no conflict of interest. None declared. No new data was generated in this study. This response letter refers to the data provided in the Seq2Topt study [1]. Sizhe Qiu, Aidong Yang |
Briefings Bioinform. | 1 |
| 2025 | Seq2Topt: a sequence-based deep learning predictor of enzyme optimal temperatureabstractAn accurate deep learning predictor is needed for enzyme optimal temperature (${T}_{opt}$), which quantitatively describes how temperature affects the enzyme catalytic activity. In comparison with existing models, a new model developed in this study, Seq2Topt, reached a superior accuracy on ${T}_{opt}$ prediction just using protein sequences (RMSE = 12.26°C and R2 = 0.57), and could capture key protein regions for enzyme ${T}_{opt}$ with multi-head attention on residues. Through case studies on thermophilic enzyme selection and predicting enzyme ${T}_{opt}$ shifts caused by point mutations, Seq2Topt was demonstrated as a promising computational tool for enzyme mining and in-silico enzyme design. Additionally, accurate deep learning predictors of enzyme optimal pH (Seq2pHopt, RMSE = 0.88 and R2 = 0.42) and melting temperature (Seq2Tm, RMSE = 7.57 °C and R2 = 0.64) were developed based on the model architecture of Seq2Topt, suggesting that the development of Seq2Topt could potentially give rise to a useful prediction platform of enzymes. Sizhe Qiu, Bozhen Hu, Weiren Xu, Aidong Yang |
Briefings Bioinform. | 1 |
| 2024 | DLTKcat: deep learning-based prediction of temperature-dependent enzyme turnover ratesabstractThe enzyme turnover rate, ${k}_{cat}$, quantifies enzyme kinetics by indicating the maximum efficiency of enzyme catalysis. Despite its importance, ${k}_{cat}$ values remain scarce in databases for most organisms, primarily because of the cost of experimental measurements. To predict ${k}_{cat}$ and account for its strong temperature dependence, DLTKcat was developed in this study and demonstrated superior performance (log10-scale root mean squared error = 0.88, R-squared = 0.66) than previously published models. Through two case studies, DLTKcat showed its ability to predict the effects of protein sequence mutations and temperature changes on ${k}_{cat}$ values. Although its quantitative accuracy is not high enough yet to model the responses of cellular metabolism to temperature changes, DLTKcat has the potential to eventually become a computational tool to describe the temperature dependence of biological systems. Sizhe Qiu, Simiao Zhao, Aidong Yang |
Briefings Bioinform. | 1 |
| 2024 | Inferred regulons are consistent with regulator binding sequences in E. coliabstractThe transcriptional regulatory network (TRN) of E. coli consists of thousands of interactions between regulators and DNA sequences. Regulons are typically determined either from resource-intensive experimental measurement of functional binding sites, or inferred from analysis of high-throughput gene expression datasets. Recently, independent component analysis (ICA) of RNA-seq compendia has shown to be a powerful method for inferring bacterial regulons. However, it remains unclear to what extent regulons predicted by ICA structure have a biochemical basis in promoter sequences. Here, we address this question by developing machine learning models that predict inferred regulon structures in E. coli based on promoter sequence features. Models were constructed successfully (cross-validation AUROC > = 0.8) for 85% (40/47) of ICA-inferred E. coli regulons. We found that: 1) The presence of a high scoring regulator motif in the promoter region was sufficient to specify regulatory activity in 40% (19/47) of the regulons, 2) Additional features, such as DNA shape and extended motifs that can account for regulator multimeric binding, helped to specify regulon structure for the remaining 60% of regulons (28/47); 3) investigating regulons where initial machine learning models failed revealed new regulator-specific sequence features that improved model accuracy. Finally, we found that strong regulatory binding sequences underlie both the genes shared between ICA-inferred and experimental regulons as well as genes in the E. coli core pan-regulon of Fur. This work demonstrates that the structure of ICA-inferred regulons largely can be understood through the strength of regulator binding sites in promoter regions, reinforcing the utility of top-down inference for regulon discovery. Sizhe Qiu, Xinlong Wan, Yueshan Liang, Cameron R. Lamoureux, Amir Akbari, Bernhard O. Palsson, Daniel C. Zielinski |
PLoS Comput. Biol. | 1 |
| 2023 | Flux balance analysis-based metabolic modeling of microbial secondary metabolism: Current status and outlookabstractIn microorganisms, different from primary metabolism for cellular growth, secondary metabolism is for ecological interactions and stress responses and an important source of natural products widely used in various areas such as pharmaceutics and food additives. With advancements of sequencing technologies and bioinformatics tools, a large number of biosynthetic gene clusters of secondary metabolites have been discovered from microbial genomes. However, due to challenges from the difficulty of genome-scale pathway reconstruction and the limitation of conventional flux balance analysis (FBA) on secondary metabolism, the quantitative modeling of secondary metabolism is poorly established, in contrast to that of primary metabolism. This review first discusses current efforts on the reconstruction of secondary metabolic pathways in genome-scale metabolic models (GSMMs), as well as related FBA-based modeling techniques. Additionally, potential extensions of FBA are suggested to improve the prediction accuracy of secondary metabolite production. As this review posits, biosynthetic pathway reconstruction for various secondary metabolites will become automated and a modeling framework capturing secondary metabolism onset will enhance the predictive power. Expectedly, an improved FBA-based modeling workflow will facilitate quantitative study of secondary metabolism and in silico design of engineering strategies for natural product production. Sizhe Qiu, Aidong Yang, Hong Zeng 0003 |
PLoS Comput. Biol. | 1 |