VLDB 2026 Research / reviewers in the wild / expert
Huiying Zhao
dblp:36/8323
· DBLP profile ↗
25ranked-venue papers
2as first author
20since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Locating the Root Cause of Poor Coverage in Mobile Communication Networks Based on Spatio-temporal Graph Message PropagationabstractPoor coverage quality is a common cause of poor wireless communication network quality, which seriously affects the user experience in mobile communication. Currently, the front line mainly adopts a manual trial-and-error method, which has problems such as low efficiency and high human cost. How to use artificial intelligence algorithms to quickly and accurately identify and solve the problem of poor coverage quality based on existing data is one of the important research directions in the field of wireless networks. The data of wireless networks is essentially spatio-temporal data, but most of the existing methods are based on time-domain and space-domain data for analysis and modeling, and the information mining in the spatio domain is not sufficient. In the spatio domain, the distribution of base stations is not uniform in Euclidean space, which increases the difficulty of spatio-temporal modeling. In view of the natural advantages of graph mining technology for modeling and processing unstructured data, this paper proposes a model named Spatio-Temporal Graph Message Propagation (STGMP) based on graph technology. This method uses spatio-temporal graphs to represent the historical states of related service cells, proposes a processing layer that combines the time and spatio domains, and maps the actual problem to a multi-classification task, thereby achieving the identification of the causes of poor coverage quality. This paper also conducts experiments on real data sets, and the results show that the proposed method STGMP is very effective. Zhipu Xie, Bin Yang 0038, Jinchao Huang 0001, Huiying Zhao, Lexi Xu, Ruiqi Liu 0002 |
IWCMC | 4 |
| 2024 | Cross-Layer Alarm Association Rules Discovery of Cloud-Network based on Knowledge GraphabstractThe fragmented architecture, cloud-based infrastructure, and functionally virtualized network elements within the 5 G core network have significantly surged the volume and diversity of alarms generated on cloud network service platforms that it supports. Given the inherently cross-layered nature of failure scenarios on these platforms, identifying the root causes presents a significant challenge. Alarm association rule mining has become an effective means to address the problems of alarm correlation and root cause localization. In this paper, an explainable alarm association rule mining approach based on knowledge graph, referred to as ARK-G, is proposed. Initially, a cloud-network cross-layer alarm association knowledge graph (CA2KG) is constructed. Subsequently, the knowledge embedding based graph convolutional network is employed to perform knowledge graph embedding on CA2KG. This embedding is then utilized to enhance the RNNLogic algorithm, thereby facilitating cross-layer alarm association rule mining with interpretable paths. Finally, a weighted rule tree is derived from a subset of CA2KG and the generated explainable rules, enabling the deduction of the root alarm. Experimental results demonstrate that the proposed ARK-G approach for association rule mining yields a higher hit rate compared to the baseline model, which provides valuable assistance in the faults analysis of 5 G cloud-network platforms. Huiying Zhao, Hongwu Li, Bin Wu 0001, Ruiqi Liu 0002, Lexi Xu, Bingming Huang, Zhipu Xie, Xinzhou Cheng |
IWCMC | 1 |
| 2024 | An uncertainty-based interpretable deep learning framework for predicting breast cancer outcomeabstractBACKGROUND: Predicting outcome of breast cancer is important for selecting appropriate treatments and prolonging the survival periods of patients. Recently, different deep learning-based methods have been carefully designed for cancer outcome prediction. However, the application of these methods is still challenged by interpretability. In this study, we proposed a novel multitask deep neural network called UISNet to predict the outcome of breast cancer. The UISNet is able to interpret the importance of features for the prediction model via an uncertainty-based integrated gradients algorithm. UISNet improved the prediction by introducing prior biological pathway knowledge and utilizing patient heterogeneity information. RESULTS: The model was tested in seven public datasets of breast cancer, and showed better performance (average C-index = 0.691) than the state-of-the-art methods (average C-index = 0.650, ranged from 0.619 to 0.677). Importantly, the UISNet identified 20 genes as associated with breast cancer, among which 11 have been proven to be associated with breast cancer by previous studies, and others are novel findings of this study. CONCLUSIONS: Our proposed method is accurate and robust in predicting breast cancer outcomes, and it is an effective way to identify breast cancer-associated genes. The method codes are available at: https://github.com/chh171/UISNet . Siyin Lin, Junqi Lin, Minfan He, Yuedong Yang, Yongzhong OuYang, Huiying Zhao |
BMC Bioinform. | 7 |
| 2023 | Accurately Identifying Muscle-Invasive Bladder Cancer from MRI via Weakly Supervised LearningabstractBladder cancer (BCa) is one of the most common malignancies in the world, which can be categorized into muscleinvasive (MIBC) and non-muscle-invasive (NMIBC). These two types of BCa must be treated differently, and thus it is essential to correctly distinguish MIBC and NMIBC patients preoperatively for adopting different treatment methods accordingly. Currently, the two types can be distinguished through MRI images by radiologists, but manual inspection is time and labor-consuming. Existing machine learning based methods attempt to free radiologists from manual inspection. However, they fail to take full advantage of image features and always require extra laborious refined manual labeling in addition to the classification labels. In this study, we propose a Tumor Staging and Localization Network (TSLNet) to perform preoperative non-invasive assessment of muscle invasion of BCa, which can automatically distinguish MIBC patients from NMIBC patients based on MRI T2-weighted images of BCa. The model adopts the weakly supervised learning method. Specifically, self-produced guidance is used as pixellevel segmentation pseudo labels for auxiliary supervision to extract basic features, and location-recognition based fine-grained image classification technology and inexact consistency labels are used for auxiliary supervision to extract fine-grained features. Moreover, the model can visualize the critical regions of the lesions, which can provide practical reference and a basis for clinicians’ clinical diagnosis. Experimental results show that the model achieves high AUC, accuracy, specificity, sensitivity, and F1-score, which is comparable to experienced clinicians. Fudan Zheng, Yuedong Yang, Tianxin Lin, Shaoxu Wu, Yutong Lu, Zhiguang Chen 0001, Huiying Zhao |
BIBM | 9 |
| 2023 | Temporal Semantic Attention Network for Aspect-Based Sentiment Analysis
Bin Yang 0038, Xinyang Tong, Huiying Zhao, Zhipu Xie |
DEXA (2) | 5 |
| 2023 | Multi-Granularity Cross-Attention Network for Visual Question AnsweringabstractVisual Question Answering (VQA) is a recent hot topic that involves multimedia analysis, computer vision (CV), natural language processing (NLP), and even a broad perspective of artificial intelligence, which is challenging and has obtained increasing attention. VQA needs a complete understanding of the spatial relationship, textual clues, as well as the common sense for an actual image. However, most existing approaches simply embed and concatenate the features of questions and images to predict answers. Treating all embeddings equally without consideration of relation consistency hinders the model performance. In this paper, we propose an explicit Multi-Granularity Cross-Attention network (MGCAN) that mutually learns the multi-modal branches. MGCAN jointly matches word-level representation with whole image, and patch-level representation with the whole question that infers the high-order vision-semantic relationship. Experiments conducted on VQA datasets demonstrate that the proposed MGCAN outperforms previous baselines. The cross-attention mechanism explicitly exploits the relevant visual and textual clues that lead to superior prediction. Xinzhou Cheng, Huiying Zhao, Zhipu Xie, Lexi Xu |
TrustCom | 5 |
| 2023 | Accurately identifying nucleic-acid-binding sites through geometric graph learning on language model predicted structuresabstractThe interactions between nucleic acids and proteins are important in diverse biological processes. The high-quality prediction of nucleic-acid-binding sites continues to pose a significant challenge. Presently, the predictive efficacy of sequence-based methods is constrained by their exclusive consideration of sequence context information, whereas structure-based methods are unsuitable for proteins lacking known tertiary structures. Though protein structures predicted by AlphaFold2 could be used, the extensive computing requirement of AlphaFold2 hinders its use for genome-wide applications. Based on the recent breakthrough of ESMFold for fast prediction of protein structures, we have developed GLMSite, which accurately identifies DNA- and RNA-binding sites using geometric graph learning on ESMFold predicted structures. Here, the predicted protein structures are employed to construct protein structural graph with residues as nodes and spatially neighboring residue pairs for edges. The node representations are further enhanced through the pre-trained language model ProtTrans. The network was trained using a geometric vector perceptron, and the geometric embeddings were subsequently fed into a common network to acquire common binding characteristics. Finally, these characteristics were input into two fully connected layers to predict binding sites with DNA and RNA, respectively. Through comprehensive tests on DNA/RNA benchmark datasets, GLMSite was shown to surpass the latest sequence-based methods and be comparable with structure-based methods. Moreover, the prediction was shown useful for inferring nucleic-acid-binding proteins, demonstrating its potential for protein function discovery. The datasets, codes, and trained models are available at https://github.com/biomed-AI/nucleic-acid-binding. Yidong Song, Qianmu Yuan, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 3 |
| 2023 | Fast and accurate protein function prediction from sequence through pretrained language model and homology-based label diffusionabstractProtein function prediction is an essential task in bioinformatics which benefits disease mechanism elucidation and drug target discovery. Due to the explosive growth of proteins in sequence databases and the diversity of their functions, it remains challenging to fast and accurately predict protein functions from sequences alone. Although many methods have integrated protein structures, biological networks or literature information to improve performance, these extra features are often unavailable for most proteins. Here, we propose SPROF-GO, a Sequence-based alignment-free PROtein Function predictor, which leverages a pretrained language model to efficiently extract informative sequence embeddings and employs self-attention pooling to focus on important residues. The prediction is further advanced by exploiting the homology information and accounting for the overlapping communities of proteins with related functions through the label diffusion algorithm. SPROF-GO was shown to surpass state-of-the-art sequence-based and even network-based approaches by more than 14.5, 27.3 and 10.1% in area under the precision-recall curve on the three sub-ontology test sets, respectively. Our method was also demonstrated to generalize well on non-homologous proteins and unseen species. Finally, visualization based on the attention mechanism indicated that SPROF-GO is able to capture sequence domains useful for function prediction. The datasets, source codes and trained models of SPROF-GO are available at https://github.com/biomed-AI/SPROF-GO. The SPROF-GO web server is freely available at http://bio-web1.nscc-gz.cn/app/sprof-go. Qianmu Yuan, Jiancong Xie, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 4 |
| 2022 | Identifying Patient Subgroups with Different Mortality Risks in ICU Based on Dynamically Acquired Clinical Data
Guilan Kong, Shuai Jin, Huiying Zhao |
AMIA | 4 |
| 2022 | Accurately Identifying Coronary Atherosclerotic Heart Disease through Merged Beats of ElectrocardiogramabstractCoronary Atherosclerotic Heart Disease (CAHD) is one kind of severe heart disease that is the dominating cause of death from non-communicable diseases worldwide. CAHD can be early detected through pre-symptomatic health check-ups, and the electrocardiogram (ECG) is common for non-invasive health check diagnoses. Traditionally, ECG signals are utilized to extract clinical features that are then input into machine learning methods for training and prediction. While these extracted features are interpretable, they are difficult to break through known features. On the other hand, ECG can be directly input to deep learning techniques, but such methods are usually limited by small sample sizes. Here, we propose to merge multiple beats of raw signal into one beat, which greatly reduces the complexity while maintaining the raw information. Moreover, we have constructed the largest benchmark dataset for 1113 CAHD patients of 12-lead ECG signals from the UK Biobank database and used the data to train a deep learning model. The results indicated that merged beat signals could achieve the best performance corresponding to an AUC of 0.71 and accuracy of 0.7, which is 4% higher than models using the raw signals and 6% higher than those using the clinical features. Further intuitive interpretation revealed that ST waves in lead II and V3 are the most closely associated with CAHD, consistent with clinical observations. Xinfeng Wang, Mengling Qi, Chengzhi Dong, Yuedong Yang, Huiying Zhao |
BIBM | 6 |
| 2022 | Genetic and phenotypic relationships between coronary atherosclerotic heart disease and electrocardiographic traitsabstractObservational studies have revealed that Coronary Atherosclerotic Heart Disease (CAHD) is associated with abnormal electrocardiogram (ECG) traits. However, it remains unclear whether there are genetic correlations between ECG and CAHD. Here, we explored genetic correlations and putative causal relationships between CAHD and ECG by performing Mendelian randomization (MR) and Polygenic risk score (PRS) analyses on the summary statistics from a large-scale genome-wide association study (GWAS) for CAHD (FinnGen: Ncase 23363, Ncontrol 187840) and ECG traits (UK Biobank: Ncase=1137, Ncontrol=40823). Results showed a causal genetic relationship between CAHD and six ECG traits in the lead V6. These ECG traits combining with age and gender have predicted CAHD risk with an AUC of 0.76. Further summary data-based Mendelian randomization (SMR) analysis identified 11 risk genes associated with the causality between CAHD and ECG. Thus, the revealed putative causal effects of CAHD on ECG traits provide genetic evidence to support the importance of monitoring CAHD risk through the ECG. Xinfeng Wang, Xuehao Xiu, Mengling Qi, Yuedong Yang, Huiying Zhao |
BIBM | 6 |
| 2022 | Capturing large genomic contexts for accurately predicting enhancer-promoter interactionsabstractEnhancer-promoter interaction (EPI) is a key mechanism underlying gene regulation. EPI prediction has always been a challenging task because enhancers could regulate promoters of distant target genes. Although many machine learning models have been developed, they leverage only the features in enhancers and promoters, or simply add the average genomic signals in the regions between enhancers and promoters, without utilizing detailed features between or outside enhancers and promoters. Due to a lack of large-scale features, existing methods could achieve only moderate performance, especially for predicting EPIs in different cell types. Here, we present a Transformer-based model, TransEPI, for EPI prediction by capturing large genomic contexts. TransEPI was developed based on EPI datasets derived from Hi-C or ChIA-PET data in six cell lines. To avoid over-fitting, we evaluated the TransEPI model by testing it on independent test datasets where the cell line and chromosome are different from the training data. TransEPI not only achieved consistent performance across the cross-validation and test datasets from different cell types but also outperformed the state-of-the-art machine learning and deep learning models. In addition, we found that the improved performance of TransEPI was attributed to the integration of large genomic contexts. Lastly, TransEPI was extended to study the non-coding mutations associated with brain disorders or neural diseases, and we found that TransEPI was also useful for predicting the target genes of non-coding mutations. Ken Chen 0006, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 2 |
| 2022 | AlphaFold2-aware protein-DNA binding site prediction using graph transformerabstractProtein-DNA interactions play crucial roles in the biological systems, and identifying protein-DNA binding sites is the first step for mechanistic understanding of various biological activities (such as transcription and repair) and designing novel drugs. How to accurately identify DNA-binding residues from only protein sequence remains a challenging task. Currently, most existing sequence-based methods only consider contextual features of the sequential neighbors, which are limited to capture spatial information. Based on the recent breakthrough in protein structure prediction by AlphaFold2, we propose an accurate predictor, GraphSite, for identifying DNA-binding residues based on the structural models predicted by AlphaFold2. Here, we convert the binding site prediction problem into a graph node classification task and employ a transformer-based variant model to take the protein structural information into account. By leveraging predicted protein structures and graph transformer, GraphSite substantially improves over the latest sequence-based and structure-based methods. The algorithm is further confirmed on the independent test set of 181 proteins, where GraphSite surpasses the state-of-the-art structure-based method by 16.4% in area under the precision-recall curve and 11.2% in Matthews correlation coefficient, respectively. We provide the datasets, the predicted structures and the source codes along with the pre-trained models of GraphSite at https://github.com/biomed-AI/GraphSite. The GraphSite web server is freely available at https://biomed.nscc-gz.cn/apps/GraphSite. Qianmu Yuan, Jiahua Rao, Shuangjia Zheng, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 5 |
| 2022 | Alignment-free metal ion-binding site prediction from protein sequence through pretrained language model and multi-task learningabstractMore than one-third of the proteins contain metal ions in the Protein Data Bank. Correct identification of metal ion-binding residues is important for understanding protein functions and designing novel drugs. Due to the small size and high versatility of metal ions, it remains challenging to computationally predict their binding sites from protein sequence. Existing sequence-based methods are of low accuracy due to the lack of structural information, and time-consuming owing to the usage of multi-sequence alignment. Here, we propose LMetalSite, an alignment-free sequence-based predictor for binding sites of the four most frequently seen metal ions in BioLiP (Zn2+, Ca2+, Mg2+ and Mn2+). LMetalSite leverages the pretrained language model to rapidly generate informative sequence representations and employs transformer to capture long-range dependencies. Multi-task learning is adopted to compensate for the scarcity of training data and capture the intrinsic similarities between different metal ions. LMetalSite was shown to surpass state-of-the-art structure-based methods by more than 19.7, 14.4, 36.8 and 12.6% in area under the precision recall on the four independent tests, respectively. Further analyses indicated that the self-attention modules are effective to learn the structural contexts of residues from protein sequence. We provide the data sets, source codes and trained models of LMetalSite at https://github.com/biomed-AI/LMetalSite. Qianmu Yuan, Yu Wang 0008, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 4 |
| 2022 | A coarse-refine segmentation network for COVID-19 CT imagesabstractThe rapid spread of the novel coronavirus disease 2019 (COVID-19) causes a significant impact on public health. It is critical to diagnose COVID-19 patients so that they can receive reasonable treatments quickly. The doctors can obtain a precise estimate of the infection's progression and decide more effective treatment options by segmenting the CT images of COVID-19 patients. However, it is challenging to segment infected regions in CT slices because the infected regions are multi-scale, and the boundary is not clear due to the low contrast between the infected area and the normal area. In this paper, a coarse-refine segmentation network is proposed to address these challenges. The coarse-refine architecture and hybrid loss is used to guide the model to predict the delicate structures with clear boundaries to address the problem of unclear boundaries. The atrous spatial pyramid pooling module in the network is added to improve the performance in detecting infected regions with different scales. Experimental results show that the model in the segmentation of COVID-19 CT images outperforms other familiar medical segmentation models, enabling the doctor to get a more accurate estimate on the progression of the infection and thus can provide more reasonable treatment options. Ziwang Huang, Xiang Zhang 0012, Huiying Zhao, Yutian Chong, Hejun Wu, Yuedong Yang, Jun Shen 0008, Yunfei Zha |
IET Image Process. | 6 |
| 2022 | To Improve Prediction of Binding Residues With DNA, RNA, Carbohydrate, and Peptide Via Multi-Task Deep Neural NetworksabstractMOTIVATION: The interactions of proteins with DNA, RNA, peptide, and carbohydrate play key roles in various biological processes. The studies of uncharacterized protein-molecules interactions could be aided by accurate predictions of residues that bind with partner molecules. However, the existing methods for predicting binding residues on proteins remain of relatively low accuracies due to the limited number of complex structures in databases. As different types of molecules partially share chemical mechanisms, the predictions for each molecular type should benefit from the binding information with other molecule types. RESULTS: In this study, we employed a multiple task deep learning strategy to develop a new sequence-based method for simultaneously predicting binding residues/sites with multiple important molecule types named MTDsite. By combining four training sets for DNA, RNA, peptide, and carbohydrate-binding proteins, our method yielded accurate and robust predictions with AUC values of 0.852, 0836, 0.758, and 0.776 on their respective independent test sets, which are 0.52 to 6.6% better than other state-of-the-art methods. To my best knowledge, this is the first method using multi-task framework to predict multiple molecular binding sites simultaneously. Shuangjia Zheng, Huiying Zhao, Zhangming Niu, Yutong Lu, Yi Pan 0001, Yuedong Yang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2021 | DGAT-onco: A differential analysis method to detect oncogenes by integrating functional information of mutationsabstractIt is a common strategy to predict oncogenes by differential analysis between somatic mutations and background mutations. Most previous methods only utilize mutations in the cancer population to model its background mutation, which have an obvious bias. A recent method, DiffMut, improves this issue by conducting differential mutational analysis with both mutations in the cancer population and the natural population. However, it assumes the impacts of all mutations are equal, neglecting their functional difference. Thus, we developed a method, DGAT-onco that integrated the functional impacts of mutations to the differential mutational analysis framework of DiffMut. We performed DGAT-onco analysis with 33 cancer types from the Cancer Genome Atlas (TCGA) dataset. Its reliability was further evaluated on an independent test set including 22 cancers from other sources (TS22). Using oncogenes from the Cancer Gene Census (CGC) as the gold standard, our method achieves higher classification performance in oncogene discovery than five alternative methods (i.e., DiffMut, WITER, OncodriveCLUSTL, OncodriveFML, and MutSigCV) with an average AUPRC of 0.197 and 0.187 in TCGA and TS22 respectively. The source code and supplementary materials of DGAT-onco are available at https://github.com/zhanghaoyang0/DGAT-onco. Junkang Wei, Zifeng Liu, Yutian Chong, Yutong Lu, Huiying Zhao, Yuedong Yang |
BIBM | 7 |
| 2021 | scAdapt: virtual adversarial domain adaptation network for single cell RNA-seq data classification across platforms and speciesabstractIn single cell analyses, cell types are conventionally identified based on expressions of known marker genes, whose identifications are time-consuming and irreproducible. To solve this issue, many supervised approaches have been developed to identify cell types based on the rapid accumulation of public datasets. However, these approaches are sensitive to batch effects or biological variations since the data distributions are different in cross-platforms or species predictions. In this study, we developed scAdapt, a virtual adversarial domain adaptation network, to transfer cell labels between datasets with batch effects. scAdapt used both the labeled source and unlabeled target data to train an enhanced classifier and aligned the labeled source centroids and pseudo-labeled target centroids to generate a joint embedding. The scAdapt was demonstrated to outperform existing methods for classification in simulated, cross-platforms, cross-species, spatial transcriptomic and COVID-19 immune datasets. Further quantitative evaluations and visualizations for the aligned embeddings confirm the superiority in cell mixing and the ability to preserve discriminative cluster structure present in the original datasets. Yuansong Zeng, Huiying Zhao, Yuedong Yang |
Briefings Bioinform. | 4 |
| 2021 | Structure-aware protein-protein interaction site prediction using deep graph convolutional networkabstractMOTIVATION: Protein-protein interactions (PPI) play crucial roles in many biological processes, and identifying PPI sites is an important step for mechanistic understanding of diseases and design of novel drugs. Since experimental approaches for PPI site identification are expensive and time-consuming, many computational methods have been developed as screening tools. However, these methods are mostly based on neighbored features in sequence, and thus limited to capture spatial information. RESULTS: We propose a deep graph-based framework deep Graph convolutional network for Protein-Protein-Interacting Site prediction (GraphPPIS) for PPI site prediction, where the PPI site prediction problem was converted into a graph node classification task and solved by deep learning using the initial residual and identity mapping techniques. We showed that a deeper architecture (up to eight layers) allows significant performance improvement over other sequence-based and structure-based methods by more than 12.5% and 10.5% on AUPRC and MCC, respectively. Further analyses indicated that the predicted interacting sites by GraphPPIS are more spatially clustered and closer to the native ones even when false-positive predictions are made. The results highlight the importance of capturing spatially neighboring residues for interacting site prediction. AVAILABILITY AND IMPLEMENTATION: The datasets, the pre-computed features, and the source codes along with the pre-trained models of GraphPPIS are available at https://github.com/biomed-AI/GraphPPIS. The GraphPPIS web server is freely available at https://biomed.nscc-gz.cn/apps/GraphPPIS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qianmu Yuan, Huiying Zhao, Yaoqi Zhou, Yuedong Yang |
Bioinform. | 3 |
| 2021 | Deep Learning Enables Accurate Diagnosis of Novel Coronavirus (COVID-19) With CT ImagesabstractA novel coronavirus (COVID-19) recently emerged as an acute respiratory syndrome, and has caused a pneumonia outbreak world-widely. As the COVID-19 continues to spread rapidly across the world, computed tomography (CT) has become essentially important for fast diagnoses. Thus, it is urgent to develop an accurate computer-aided method to assist clinicians to identify COVID-19-infected patients by CT images. Here, we have collected chest CT scans of 88 patients diagnosed with COVID-19 from hospitals of two provinces in China, 100 patients infected with bacteria pneumonia, and 86 healthy persons for comparison and modeling. Based on the data, a deep learning-based CT diagnosis system was developed to identify patients with COVID-19. The experimental results showed that our model could accurately discriminate the COVID-19 patients from the bacteria pneumonia patients with an AUC of 0.95, recall (sensitivity) of 0.96, and precision of 0.79. When integrating three types of CT images, our model achieved a recall of 0.93 with precision of 0.86 for discriminating COVID-19 patients from others. Moreover, our model could extract main lesion features, especially the ground-glass opacity (GGO), which are visually helpful for assisted diagnoses by doctors. An online server is available for online diagnoses with CT images by our server (http://biomed.nscc-gz.cn/model.php). Source codes and datasets are available at our GitHub (https://github.com/SY575/COVID19-CT). Shuangjia Zheng, Xiang Zhang 0012, Ziwang Huang, Huiying Zhao, Yutian Chong, Jun Shen 0008, Yunfei Zha, Yuedong Yang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 9 |
| 2020 | Deep Learning Based Prediction Towards Designing A Smart Building Assistant SystemabstractNowadays, smart building infrastructures are equipped with hundreds of sensors to monitor building environments and provide smart solutions for occupant comfortability and energy efficiency. Ideally, an automated system can predict and adjust the physical features (e.g., lighting, air quality, temperature, and so on) in a person’s office based on his/her personalized preferences and activities. However, since the data is from one person, there may not be sufficient data for machine learning model training, and the data’s quality may be low (e.g., with noises). Then, it is a challenge to conduct accurate predictions to provide personalized environment adjustment. To handle this problem, in this paper, we propose a smart building assistance system consisting of different sensor data analysis approaches and a deep neural network (DNN)-based prediction model to make a more accurate prediction despite low-quality sensor data. First, we collected a year-long smart building dataset from four different data sources (i.e., sensors, calendar, weather, and survey). Second, we perform different feature engineering approaches (i.e., concretization, one-hot encoding, and multiple feature combination) on the data as inputs for the prediction models. Third, we identify a support vector regression-based prediction model and propose a hybrid DNN model consisting of several recurrent neural network blocks and a feed-forward DNN block to predict different preferred physical features considering different activities of a person (e.g., meeting, lunch, research activities). Finally, we conduct experimental studies to evaluate the performance of the proposed prediction models compared to other existing machine learning models in terms of accuracy. Our predicted preferred physical features match the occupant’s preferred ranges of different physical features during a specific activity. We also open-sourced our code on GitHub. Ankur Sarker, Fan Yao 0002, Haiying Shen, Huiying Zhao, Haroon R. Lone, Bradford Campbell, Mitchel Rosen |
MASS | 4 |
| 2020 | Accurate prediction of genome-wide RNA secondary structure profile based on extreme gradient boostingabstractMOTIVATION: RNA secondary structure plays a vital role in fundamental cellular processes, and identification of RNA secondary structure is a key step to understand RNA functions. Recently, a few experimental methods were developed to profile genome-wide RNA secondary structure, i.e. the pairing probability of each nucleotide, through high-throughput sequencing techniques. However, these high-throughput methods have low precision and cannot cover all nucleotides due to limited sequencing coverage. RESULTS: Here, we have developed a new method for the prediction of genome-wide RNA secondary structure profile from RNA sequence based on the extreme gradient boosting technique. The method achieves predictions with areas under the receiver operating characteristic curve (AUC) >0.9 on three different datasets, and AUC of 0.888 by another independent test on the recently released Zika virus data. These AUCs are consistently >5% greater than those by the CROSS method recently developed based on a shallow neural network. Further analysis on the 1000 Genome Project data showed that our predicted unpaired probabilities are highly correlated (>0.8) with the minor allele frequencies at synonymous, non-synonymous mutations, and mutations in untranslated regions, which were higher than those generated by RNAplfold. Moreover, the prediction over all human mRNA indicated a consistent result with previous observation that there is a periodic distribution of unpaired probability on codons. The accurate predictions by our method indicate that such model trained on genome-wide experimental data might be an alternative for analytical methods. AVAILABILITY AND IMPLEMENTATION: The GRASP is available for academic use at https://github.com/sysu-yanglab/GRASP. SUPPLEMENTARY INFORMATION: Supplementary data are available online. Yaobin Ke, Jiahua Rao, Huiying Zhao, Yutong Lu, Nong Xiao 0001, Yuedong Yang |
Bioinform. | 3 |
| 2017 | Equivalent modeling and simulation for PV system on dynamic clustering equivalent strategyabstractAs the large-scale integration of photovoltaic (PV) power station, a higher requirement is put forward by power grid analysis on the accuracy of PV station model. For the model of photovoltaic system is complex, and it requires solving large-scale model equation, which is not conductive for simulation and big data analysis when a large-scale PV power station connected to grid. This paper focuses on the same type of photovoltaic power generation, assuming that the photovoltaic power plant is composed of multiple photovoltaic power generation units connected by the collector line, the photovoltaic power plant can be modeled to consider the operation of similar power generation unit grouped together, so that reduce the simulation scale. This paper points out that the dynamic clustering equivalence strategy can be classified into two cases. During the fault and fault-over, the transient condition is required to be grouped when the active power ramp recovery control module is operating, and the another single-unit equivalent model can be performed when the module is ignoring or not running into a non-fault condition, which is the second condition. Simulation results show that the proposed method has good adaptability for irradiance disturbance, different fault time. Hang Meng, Xiaohui Ye, Xinli Song, Zhida Su, Wenzhuo Liu, Lingtong Luo, Huiying Zhao |
IECON | 8 |
| 2011 | Improving protein fold recognition and template-based modeling by employing probabilistic-based matching between predicted one-dimensional structural properties of query and corresponding native properties of templatesabstractMOTIVATION: In recent years, development of a single-method fold-recognition server lags behind consensus and multiple template techniques. However, a good consensus prediction relies on the accuracy of individual methods. This article reports our efforts to further improve a single-method fold recognition technique called SPARKS by changing the alignment scoring function and incorporating the SPINE-X techniques that make improved prediction of secondary structure, backbone torsion angle and solvent accessible surface area. RESULTS: The new method called SPARKS-X was tested with the SALIGN benchmark for alignment accuracy, Lindahl and SCOP benchmarks for fold recognition, and CASP 9 blind test for structure prediction. The method is compared to several state-of-the-art techniques such as HHPRED and BoostThreader. Results show that SPARKS-X is one of the best single-method fold recognition techniques. We further note that incorporating multiple templates and refinement in model building will likely further improve SPARKS-X. AVAILABILITY: The method is available as a SPARKS-X server at http://sparks.informatics.iupui.edu/ Yuedong Yang, Eshel Faraggi, Huiying Zhao, Yaoqi Zhou |
Bioinform. | 3 |
| 2010 | Structure-based prediction of DNA-binding proteins by structural alignment and a volume-fraction corrected DFIRE-based energy functionabstractMOTIVATION: Template-based prediction of DNA binding proteins requires not only structural similarity between target and template structures but also prediction of binding affinity between the target and DNA to ensure binding. Here, we propose to predict protein-DNA binding affinity by introducing a new volume-fraction correction to a statistical energy function based on a distance-scaled, finite, ideal-gas reference (DFIRE) state. RESULTS: We showed that this energy function together with the structural alignment program TM-align achieves the Matthews correlation coefficient (MCC) of 0.76 with an accuracy of 98%, a precision of 93% and a sensitivity of 64%, for predicting DNA binding proteins in a benchmark of 179 DNA binding proteins and 3797 non-binding proteins. The MCC value is substantially higher than the best MCC value of 0.69 given by previous methods. Application of this method to 2235 structural genomics targets uncovered 37 as DNA binding proteins, 27 (73%) of which are putatively DNA binding and only 1 protein whose annotated functions do not contain DNA binding, while the remaining proteins have unknown function. The method provides a highly accurate and sensitive technique for structure-based prediction of DNA binding proteins. AVAILABILITY: The method is implemented as a part of the Structure-based function-Prediction On-line Tools (SPOT) package available at http://sparks.informatics.iupui.edu/spot Huiying Zhao, Yuedong Yang, Yaoqi Zhou |
Bioinform. | 1 |