VLDB 2026 Research / reviewers in the wild / expert
Fengfeng Zhou
dblp:84/933
· DBLP profile ↗
38ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0002-8108-6007ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HpMiX: A Disease ceRNA biomarker prediction framework driven by graph topology-constrained Mixup and hypergraph residual enhancement
Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Fengfeng Zhou |
Neural Networks | 6 |
| 2026 | PepHarmony: a multi-view contrastive learning framework for integrated sequence and structure-based peptide representation
Ruochi Zhang, Chang Liu 0082, Huaping Li, Yuqian Wu, Fengfeng Zhou, Xin Gao 0001 |
Neural Networks | 10 |
| 2026 | DeepSelective: Interpretable prognosis prediction via feature selection and compression in EHR data
Ruochi Zhang, Xiaoyang Wang 0009, Qiong Zhou, Ziqi Deng, Yueying Wang, Yusi Fan, Jiale Zhang 0002, Lan Huang 0002, Chang Liu 0082, Fengfeng Zhou |
Pattern Recognit. | 13 |
| 2026 | A Dynamic Multi-Scale Hypergraph Learning Framework Driven by Features and Structures for ceRNA-Disease Association PredictionabstractCompetitive endogenous RNA (ceRNA) networks are pivotal for uncovering disease molecular mechanisms. Graph representation learning is a cornerstone for modeling biological regulatory networks and predicting disease-related biomarkers. However, current methods face challenges: traditional graph neural network (GNN) rely on low-order graph structures, which struggle to capture high-order molecular interactions, resulting in topological information loss; shallow GNN fail to model long-range dependencies, while deep architectures suffer from over-smoothing, limiting complex regulatory expression; static embeddings overlook dynamic molecular interactions, reducing biomarker accuracy. These limitations highlight the need for advanced graph learning frameworks. To address these challenges, we propose DMHLF, a Dynamic Multi-scale Hypergraph Learning Framework for predicting disease-associated ceRNA biomarkers. The framework first integrates multiple regulatory relationships among miRNAs, lncRNAs, circRNAs, mRNAs, and diseases to construct disease-specific ceRNA regulatory networks, capturing local and global regulatory patterns through multi-Hop hyperedges. Subsequently, we devise a Hypergraph-Weighted Dynamic Random Walk (HEDRW) method to dynamically extract node meta-embeddings that encode high-order regulatory information. Concurrently, we extend Eigen-GNN spectral analysis to hypergraph structures, incorporating a residual-enhanced hypergraph neural network to preserve the global topological properties of shallow hypergraphs. Finally, a cross-scale attention mechanism aligns and fuses multi-scale features to generate high-quality node embeddings for disease-ceRNA association prediction. Experiments on diverse datasets demonstrate that DMHLF significantly outperforms existing methods. Case study further validates the framework's efficacy in identifying disease-related ceRNA biomarkers, providing a reliable predictive tool for biomedical research. Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Fengfeng Zhou, Yu-Qing Li |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Reliable Multimodal Cancer Survival Prediction With Confidence-Aware Risk ModelingabstractMultimodal survival methods that integrate histology whole-slide images and transcriptomic profiles hold significant promise for understanding patient prognostication and guiding personalized treatment strategies. However, existing approaches primarily focus on improving predictive performance through multimodal information fusion, often neglecting the reliability estimation of the prediction results and the inherent alignment noise across modalities. Thus, we propose ReCaSP, a novel and reliable cancer survival prediction framework that effectively integrates histology and transcriptomics data via multimodal alignment and fusion, providing the auxiliary confidence levels for survival predictions through a confidence-aware risk modeling mechanism. Specifically, our approach incorporates a fine-grained risk classifier that models risk labels jointly over both multiple time intervals and censorship status, utilizing evidential deep learning to yield fine-grained risk predictions accompanied by confidence scores. Additionally, to mitigate the inherent noise in multimodal data alignment, we introduce a cross-attention alignment module that effectively aligns histology data with transcriptomics data prior to multimodal fusion, thereby facilitating cross-modal interaction learning. Extensive experiments on five datasets demonstrate that ReCaSP significantly outperforms state-of-the-art methods, achieving a 4.58% improvement in the overall C-Index. Xuping Xie, Qixing Yang, Lan Huang 0002, Fengfeng Zhou, Yan Wang 0028 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | PepLand: a large-scale pre-trained peptide representation model for a comprehensive landscape of both canonical and non-canonical amino acidsabstractThe recent interest in peptides incorporating non-canonical amino acids has surged within the scientific community, driven by their enhanced stability and resistance to proteolytic degradation. These so-called non-canonical peptides offer significant potential for modifying biological, pharmacological, and physiochemical characteristics in both native and synthetic contexts. Despite their advantages, there remains a notable gap in the availability of an efficient pre-trained model capable of effectively capturing feature representations from such intricate peptide sequences. This study herein introduces PepLand, a novel pre-training framework designed for the comprehensive representation and analysis of peptides, encompassing both canonical and non-canonical amino acids. PepLand leverages a general-purpose multi-view heterogeneous graph neural network to unveil the subtle structural representations of peptides. Our empirical evaluations demonstrate PepLand's proficiency in a range of peptide property prediction tasks, including cell penetrability, solubility, and protein-peptide binding affinity. These rigorous assessments affirm PepLand's superior capability in discerning critical representations of peptides with both canonical and non-canonical amino acids, and provide a robust foundation for transformative advances in peptide-focused pharmaceutical research. We have made the entire source code and datasets available at http://www.healthinformaticslab.org/supp/resources.php or https://github.com/zhangruochi/PepLand. Ruochi Zhang, Chang Liu 0082, Yuting Xiu, Ningning Chen, Yu Wang 0225, Yan Wang 0028, Xin Gao 0001, Fengfeng Zhou |
Briefings Bioinform. | 11 |
| 2025 | TaiChiNet: PCA-based Ying-Yang dilution of inter- and intra-BERT layers to represent anti-coronavirus peptides
Shiying Ding, Yusi Fan, Yannan Sun, Gongyou Zhang, Ruochi Zhang, Lan Huang 0002, Fengfeng Zhou |
Expert Syst. Appl. | 10 |
| 2025 | SeqFReD: A reinforcement-learning framework for sequence feature representation and dimension reduction
Xuechen Mu, Haotian Zhang 0018, Yusi Fan, Fengfeng Zhou |
Knowl. Based Syst. | 8 |
| 2024 | MolFeSCue: enhancing molecular property prediction in data-limited and imbalanced contexts using few-shot and contrastive learningabstractMOTIVATION: Predicting molecular properties is a pivotal task in various scientific domains, including drug discovery, material science, and computational chemistry. This problem is often hindered by the lack of annotated data and imbalanced class distributions, which pose significant challenges in developing accurate and robust predictive models. RESULTS: This study tackles these issues by employing pretrained molecular models within a few-shot learning framework. A novel dynamic contrastive loss function is utilized to further improve model performance in the situation of class imbalance. The proposed MolFeSCue framework not only facilitates rapid generalization from minimal samples, but also employs a contrastive loss function to extract meaningful molecular representations from imbalanced datasets. Extensive evaluations and comparisons of MolFeSCue and state-of-the-art algorithms have been conducted on multiple benchmark datasets, and the experimental data demonstrate our algorithm's effectiveness in molecular representations and its broad applicability across various pretrained models. Our findings underscore MolFeSCues potential to accelerate advancements in drug discovery. AVAILABILITY AND IMPLEMENTATION: We have made all the source code utilized in this study publicly accessible via GitHub at http://www.healthinformaticslab.org/supp/ or https://github.com/zhangruochi/MolFeSCue. The code (MolFeSCue-v1-00) is also available as the supplementary file of this paper. Ruochi Zhang, Chang Liu 0082, Yan Wang 0028, Lan Huang 0002, Fengfeng Zhou |
Bioinform. | 8 |
| 2024 | FairCare: Adversarial training of a heterogeneous graph neural network with attention mechanism to learn fair representations of electronic health records
Yan Wang 0028, Ruochi Zhang, Qiong Zhou, Shengde Zhang, Yusi Fan, Lan Huang 0002, Fengfeng Zhou |
Inf. Process. Manag. | 9 |
| 2024 | A probabilistic knowledge graph for target identificationabstractEarly identification of safe and efficacious disease targets is crucial to alleviating the tremendous cost of drug discovery projects. However, existing experimental methods for identifying new targets are generally labor-intensive and failure-prone. On the other hand, computational approaches, especially machine learning-based frameworks, have shown remarkable application potential in drug discovery. In this work, we propose Progeni, a novel machine learning-based framework for target identification. In addition to fully exploiting the known heterogeneous biological networks from various sources, Progeni integrates literature evidence about the relations between biological entities to construct a probabilistic knowledge graph. Graph neural networks are then employed in Progeni to learn the feature embeddings of biological entities to facilitate the identification of biologically relevant target candidates. A comprehensive evaluation of Progeni demonstrated its superior predictive power over the baseline methods on the target identification task. In addition, our extensive tests showed that Progeni exhibited high robustness to the negative effect of exposure bias, a common phenomenon in recommendation systems, and effectively identified new targets that can be strongly supported by the literature. Moreover, our wet lab experiments successfully validated the biological significance of the top target candidates predicted by Progeni for melanoma and colorectal cancer. All these results suggested that Progeni can identify biologically effective targets and thus provide a powerful and useful tool for advancing the drug discovery process. Chang Liu 0082, Kaimin Xiao, Cuinan Yu, Yipin Lei, Kangbo Lyu, Tingzhong Tian, Dan Zhao 0004, Fengfeng Zhou, Haidong Tang, Jianyang Zeng 0001 |
PLoS Comput. Biol. | 8 |
| 2023 | Orchestrating information across tissues via a novel multitask GAT framework to improve quantitative gene regulation relation modeling for survival analysisabstractSurvival analysis is critical to cancer prognosis estimation. High-throughput technologies facilitate the increase in the dimension of genic features, but the number of clinical samples in cohorts is relatively small due to various reasons, including difficulties in participant recruitment and high data-generation costs. Transcriptome is one of the most abundantly available OMIC (referring to the high-throughput data, including genomic, transcriptomic, proteomic and epigenomic) data types. This study introduced a multitask graph attention network (GAT) framework DQSurv for the survival analysis task. We first used a large dataset of healthy tissue samples to pretrain the GAT-based HealthModel for the quantitative measurement of the gene regulatory relations. The multitask survival analysis framework DQSurv used the idea of transfer learning to initiate the GAT model with the pretrained HealthModel and further fine-tuned this model using two tasks i.e. the main task of survival analysis and the auxiliary task of gene expression prediction. This refined GAT was denoted as DiseaseModel. We fused the original transcriptomic features with the difference vector between the latent features encoded by the HealthModel and DiseaseModel for the final task of survival analysis. The proposed DQSurv model stably outperformed the existing models for the survival analysis of 10 benchmark cancer types and an independent dataset. The ablation study also supported the necessity of the main modules. We released the codes and the pretrained HealthModel to facilitate the feature encodings and survival analysis of transcriptome-based future studies, especially on small datasets. The model and the code are available at http://www.healthinformaticslab.org/supp/. Meiyu Duan, Yueying Wang, Gongyou Zhang, Haotian Zhang 0018, Lan Huang 0002, Ruochi Zhang, Fengfeng Zhou |
Briefings Bioinform. | 10 |
| 2023 | GSRNet, an adversarial training-based deep framework with multi-scale CNN and BiGRU for predicting genomic signals and regions
Gancheng Zhu, Yusi Fan, Fei Li 0039, Annebella Tsz Ho Choi, Zhikang Tan, Yiruo Cheng, Changfan Luo, Gongyou Zhang, Zhaomin Yao, Lan Huang 0002, Fengfeng Zhou |
Expert Syst. Appl. | 15 |
| 2023 | INS-GNN: Improving graph imbalance learning with self-supervision
Xin Juan, Fengfeng Zhou, Wentao Wang 0006, Wei Jin 0009, Jiliang Tang, Xin Wang 0035 |
Inf. Sci. | 2 |
| 2023 | EvaGoNet: An integrated network of variational autoencoder and Wasserstein generative adversarial network with gradient penalty for binary classification tasks
Changfan Luo, Yongkang Shao, Jianzheng Hu, Meiyu Duan, Lan Huang 0002, Fengfeng Zhou |
Inf. Sci. | 10 |
| 2023 | FMGNN: A Method to Predict Compound-Protein Interaction With Pharmacophore Features and Physicochemical Properties of Amino AcidsabstractIdentifying interactions between compounds and proteins is an essential task in drug discovery. To recommend compounds as new drug candidates, applying the computational approaches has a lower cost than conducting the wet-lab experiments. Machine learning-based methods, especially deep learning-based methods, have advantages in learning complex feature interactions between compounds and proteins. However, deep learning models will over-generalize and lead to the problem of predicting less relevant compound-protein pairs when the compound-protein feature interactions are high-dimensional sparse. This problem can be overcome by learning both low-order and high-order feature interactions. In this paper, we propose a novel hybrid model with Factorization Machines and Graph Neural Network called FMGNN to extract the low-order and high-order features, respectively. Then, we design a compound-protein interactions (CPIs) prediction method with pharmacophore features of compound and physicochemical properties of amino acids. The pharmacophore features can ensure that the prediction results much more fit the expectation of biological experiment and the physicochemical properties of amino acids are loaded into the embedding layer to improve the convergence speed and accuracy of protein feature learning. The experimental results on several datasets, especially on an imbalanced large-scale dataset, showed that our proposed method outperforms other existing methods for CPI prediction. The western blot experiment results on wogonin and its candidate target proteins also showed that our proposed method is effective and accurate for finding target proteins. The computer program of implementing the model FMGNN is available at https://github.com/tcygxu2021/FMGNN. Chunyan Tang, Fengfeng Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | HLAB: learning the BiLSTM features from the ProtBert-encoded proteins for the class I HLA-peptide binding predictionabstractHuman Leukocyte Antigen (HLA) is a type of molecule residing on the surfaces of most human cells and exerts an essential role in the immune system responding to the invasive items. The T cell antigen receptors may recognize the HLA-peptide complexes on the surfaces of cancer cells and destroy these cancer cells through toxic T lymphocytes. The computational determination of HLA-binding peptides will facilitate the rapid development of cancer immunotherapies. This study hypothesized that the natural language processing-encoded peptide features may be further enriched by another deep neural network. The hypothesis was tested with the Bi-directional Long Short-Term Memory-extracted features from the pretrained Protein Bidirectional Encoder Representations from Transformers-encoded features of the class I HLA (HLA-I)-binding peptides. The experimental data showed that our proposed HLAB feature engineering algorithm outperformed the existing ones in detecting the HLA-I-binding peptides. The extensive evaluation data show that the proposed HLAB algorithm outperforms all the seven existing studies on predicting the peptides binding to the HLA-A*01:01 allele in AUC and achieves the best average AUC values on the six out of the seven k-mers (k=8,9,...,14, respectively represent the prediction task of a polypeptide consisting of k amino acids) except for the 9-mer prediction tasks. The source code and the fine-tuned feature extraction models are available at http://www.healthinformaticslab.org/supp/resources.php. Gancheng Zhu, Fei Li 0039, Lan Huang 0002, Meiyu Duan, Fengfeng Zhou |
Briefings Bioinform. | 7 |
| 2022 | Superpixel-Level Global and Local Similarity Graph-Based Clustering for Large Hyperspectral ImagesabstractDue to the scarcity of labeled samples, clustering in hyperspectral images (HSIs) has a great potential and application value. However, current clustering methods are mainly pixel-level techniques that neglect the large spectral variability of a scene and suffer from massive time and memory consumption when dealing with large HSIs. In this article, we propose a superpixel-level global and local similarity graph-based clustering (SGLSC) algorithm that can classify ground objects exploiting spectral and spatial dimensions with reasonable time and memory consumption on large HSIs. The proposed SGLSC exploits the superpixel concept, which is treated as a homogeneous entity, into the clustering process. For modeling the essential structure of HSIs, a similarity graph combing the global and local information is constructed and inserted into the spectral clustering to partition the superpixel-level graph structure. The proposed method was tested on three benchmark HSIs’ datasets and compared with some advanced literature algorithms. Experiments demonstrate that it can obtain promising results. Haishi Zhao, Fengfeng Zhou, Lorenzo Bruzzone, Renchu Guan, Chen Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | A comprehensive comparison of residue-level methylation levels with the regression-based gene-level methylation estimations by ReGearabstractMOTIVATION: DNA methylation is a biological process impacting the gene functions without changing the underlying DNA sequence. The DNA methylation machinery usually attaches methyl groups to some specific cytosine residues, which modify the chromatin architectures. Such modifications in the promoter regions will inactivate some tumor-suppressor genes. DNA methylation within the coding region may significantly reduce the transcription elongation efficiency. The gene function may be tuned through some cytosines are methylated. METHODS: This study hypothesizes that the overall methylation level across a gene may have a better association with the sample labels like diseases than the methylations of individual cytosines. The gene methylation level is formulated as a regression model using the methylation levels of all the cytosines within this gene. A comprehensive evaluation of various feature selection algorithms and classification algorithms is carried out between the gene-level and residue-level methylation levels. RESULTS: A comprehensive evaluation was conducted to compare the gene and cytosine methylation levels for their associations with the sample labels and classification performances. The unsupervised clustering was also improved using the gene methylation levels. Some genes demonstrated statistically significant associations with the class label, even when no residue-level methylation features have statistically significant associations with the class label. So in summary, the trained gene methylation levels improved various methylome-based machine learning models. Both methodology development of regression algorithms and experimental validation of the gene-level methylation biomarkers are worth of further investigations in the future studies. The source code, example data files and manual are available at http://www.healthinformaticslab.org/supp/. Jinpu Cai, Yuyang Xu, Shiying Ding, Yuewei Sun, Jingyi Lyu, Meiyu Duan, Shuai Liu 0010, Lan Huang 0002, Fengfeng Zhou |
Briefings Bioinform. | 10 |
| 2021 | Application of Bayesian phylogenetic inference modelling for evolutionary genetic analysis and dynamic changes in 2019-nCoVabstractThe novel coronavirus (2019-nCoV) has recently caused a large-scale outbreak of viral pneumonia both in China and worldwide. In this study, we obtained the entire genome sequence of 777 new coronavirus strains as of 29 February 2020 from a public gene bank. Bioinformatics analysis of these strains indicated that the mutation rate of these new coronaviruses is not high at present, similar to the mutation rate of the severe acute respiratory syndrome (SARS) virus. The similarities of 2019-nCoV and SARS virus suggested that the S and ORF6 proteins shared a low similarity, while the E protein shared the higher similarity. The 2019-nCoV sequence has similar potential phosphorylation sites and glycosylation sites on the surface protein and the ORF1ab polyprotein as the SARS virus; however, there are differences in potential modification sites between the Chinese strain and some American strains. At the same time, we proposed two possible recombination sites for 2019-nCoV. Based on the results of the skyline, we speculate that the activity of the gene population of 2019-nCoV may be before the end of 2019. As the scope of the 2019-nCoV infection further expands, it may produce different adaptive evolutions due to different environments. Finally, evolutionary genetic analysis can be a useful resource for studying the spread and virulence of 2019-nCoV, which are essential aspects of preventive and precise medicine. Tong Shao, Wenfang Wang, Meiyu Duan, Zhuoyuan Xin, Baoyue Liu, Fengfeng Zhou |
Briefings Bioinform. | 7 |
| 2021 | A dynamic recursive feature elimination framework (dRFE) to further refine a set of OMIC biomarkersabstractMOTIVATION: A feature selection algorithm may select the subset of features with the best associations with the class labels. The recursive feature elimination (RFE) is a heuristic feature screening framework and has been widely used to select the biological OMIC biomarkers. This study proposed a dynamic recursive feature elimination (dRFE) framework with more flexible feature elimination operations. The proposed dRFE was comprehensively compared with 11 existing feature selection algorithms and five classifiers on the eight difficult transcriptome datasets from a previous study, the ten newly collected transcriptome datasets and the five methylome datasets. RESULTS: The experimental data suggested that the regular RFE framework did not perform well, and dRFE outperformed the existing feature selection algorithms in most cases. The dRFE-detected features achieved Acc = 1.0000 for the two methylome datasets GSE53045 and GSE66695. The best prediction accuracies of the dRFE-detected features were 0.9259, 0.9424 and 0.8601 for the other three methylome datasets GSE74845, GSE103186 and GSE80970, respectively. Four transcriptome datasets received Acc = 1.0000 using the dRFE-detected features, and the prediction accuracies for the other six newly collected transcriptome datasets were between 0.6301 and 0.9917. AVAILABILITY AND IMPLEMENTATION: The experiments in this study are implemented and tested using the programming language Python version 3.7.6. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lan Huang 0002, Fengfeng Zhou |
Bioinform. | 3 |
| 2021 | Finding branched pathways in metabolic network via atom group trackingabstractFinding non-standard or new metabolic pathways has important applications in metabolic engineering, synthetic biology and the analysis and reconstruction of metabolic networks. Branched metabolic pathways dominate in metabolic networks and depict a more comprehensive picture of metabolism compared to linear pathways. Although progress has been developed to find branched metabolic pathways, few efforts have been made in identifying branched metabolic pathways via atom group tracking. In this paper, we present a pathfinding method called BPFinder for finding branched metabolic pathways by atom group tracking, which aims to guide the synthetic design of metabolic pathways. BPFinder enumerates linear metabolic pathways by tracking the movements of atom groups in metabolic network and merges the linear atom group conserving pathways into branched pathways. Two merging rules based on the structure of conserved atom groups are proposed to accurately merge the branched compounds of linear pathways to identify branched pathways. Furthermore, the integrated information of compound similarity, thermodynamic feasibility and conserved atom groups is also used to rank the pathfinding results for feasible branched pathways. Experimental results show that BPFinder is more capable of recovering known branched metabolic pathways as compared to other existing methods, and is able to return biologically relevant branched pathways and discover alternative branched pathways of biochemical interest. The online server of BPFinder is available at http://114.215.129.245:8080/atomic/. The program, source code and data can be downloaded from https://github.com/hyr0771/BPFinder. Yusi Xie, Fengfeng Zhou |
PLoS Comput. Biol. | 4 |
| 2021 | Spectral-Spatial Genetic Algorithm-Based Unsupervised Band Selection for Hyperspectral Image ClassificationabstractBand selection (BS) can mitigate the “curse of dimensionality” problem and improve the performance of hyperspectral image (HSI) classification. Genetic algorithms (GAs) have been applied to the task of hyperspectral BS showing significant advantages compared with other literature methods. However, the traditional GAs-based methods often select sets of bands having residual redundancy due to the large search space related to hyperspectral BS and the limitation of premature convergence in GAs. Moreover, existing GAs-based methods often are supervised, and that needs a large number of labeled samples to compute the fitness value for assessing the quality of selected bands. In this article, an unsupervised BS approach based on an improved GA is proposed. A fitness function based on the fisher score combined with superpixel is designed for evaluating the discriminability of band subsets considering both spectral and spatial information. Then, modified genetic operations are constructed to restrain the search space and reduce the redundancy of selected bands. The performance of the proposed spectral-spatial GA-based BS method is evaluated on three HSIs. The experimental results demonstrate that the proposed method is superior to the traditional GA-based method and seven state-of-the-art unsupervised methods. Haishi Zhao, Lorenzo Bruzzone, Renchu Guan, Fengfeng Zhou, Chen Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2020 | Heuristic Black-Box Adversarial Attacks on Video Recognition ModelsabstractWe study the problem of attacking video recognition models in the black-box setting, where the model information is unknown and the adversary can only make queries to detect the predicted top-1 class and its probability. Compared with the black-box attack on images, attacking videos is more challenging as the computation cost for searching the adversarial perturbations on a video is much higher due to its high dimensionality. To overcome this challenge, we propose a heuristic black-box attack model that generates adversarial perturbations only on the selected frames and regions. More specifically, a heuristic-based algorithm is proposed to measure the importance of each frame in the video towards generating the adversarial examples. Based on the frames' importance, the proposed algorithm heuristically searches a subset of frames where the generated adversarial example has strong adversarial attack ability while keeps the perturbations lower than the given bound. Besides, to further boost the attack efficiency, we propose to generate the perturbations only on the salient regions of the selected frames. In this way, the generated perturbations are sparse in both temporal and spatial domains. Experimental results of attacking two mainstream video recognition methods on the UCF-101 dataset and the HMDB-51 dataset demonstrate that the proposed heuristic black-box adversarial attack method can significantly reduce the computation cost and lead to more than 28% reduction in query numbers for the untargeted attack on both datasets. Zhipeng Wei 0001, Jingjing Chen 0001, Linxi Jiang, Tat-Seng Chua, Fengfeng Zhou, Yu-Gang Jiang 0001 |
AAAI | 6 |
| 2020 | Feature selection may improve deep neural networks for the bioinformatics problemsabstractMOTIVATION: Deep neural network (DNN) algorithms were utilized in predicting various biomedical phenotypes recently, and demonstrated very good prediction performances without selecting features. This study proposed a hypothesis that the DNN models may be further improved by feature selection algorithms. RESULTS: A comprehensive comparative study was carried out by evaluating 11 feature selection algorithms on three conventional DNN algorithms, i.e. convolution neural network (CNN), deep belief network (DBN) and recurrent neural network (RNN), and three recent DNNs, i.e. MobilenetV2, ShufflenetV2 and Squeezenet. Five binary classification methylomic datasets were chosen to calculate the prediction performances of CNN/DBN/RNN models using feature selected by the 11 feature selection algorithms. Seventeen binary classification transcriptome and two multi-class transcriptome datasets were also utilized to evaluate how the hypothesis may generalize to different data types. The experimental data supported our hypothesis that feature selection algorithms may improve DNN models, and the DBN models using features selected by SVM-RFE usually achieved the best prediction accuracies on the five methylomic datasets. AVAILABILITY AND IMPLEMENTATION: All the algorithms were implemented and tested under the programming environment Python version 3.6.6. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shuainan Li, Xiaoyue Feng, Xin Feng 0004, Yexian Zhang, Meiyu Duan, Lan Huang 0002, Fengfeng Zhou |
Bioinform. | 12 |
| 2020 | sefOri: selecting the best-engineered sequence features to predict DNA replication originsabstractMOTIVATION: Cell divisions start from replicating the double-stranded DNA, and the DNA replication process needs to be precisely regulated both spatially and temporally. The DNA is replicated starting from the DNA replication origins. A few successful prediction models were generated based on the assumption that the DNA replication origin regions have sequence level features like physicochemical properties significantly different from the other DNA regions. RESULTS: This study proposed a feature selection procedure to further refine the classification model of the DNA replication origins. The experimental data demonstrated that as large as 26% improvement in the prediction accuracy may be achieved on the yeast Saccharomyces cerevisiae. Moreover, the prediction accuracies of the DNA replication origins were improved for all the four yeast genomes investigated in this study. AVAILABILITY AND IMPLEMENTATION: The software sefOri version 1.0 was available at http://www.healthinformaticslab.org/supp/resources.php. An online server was also provided for the convenience of the users, and its web link may be found in the above-mentioned web page. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Chenwei Lou, Ruoyao Shi, Wenyang Zhou, Yubo Wang 0006, Lan Huang 0002, Xin Feng 0004, Fengfeng Zhou |
Bioinform. | 10 |
| 2020 | Combining High Speed ELM Learning with a Deep Convolutional Neural Network Feature Encoding for Predicting Protein-RNA InteractionsabstractEmerging evidence has shown that RNA plays a crucial role in many cellular processes, and their biological functions are primarily achieved by binding with a variety of proteins. High-throughput biological experiments provide a lot of valuable information for the initial identification of RNA-protein interactions (RPIs), but with the increasing complexity of RPIs networks, this method gradually falls into expensive and time-consuming situations. Therefore, there is an urgent need for high speed and reliable methods to predict RNA-protein interactions. In this study, we propose a computational method for predicting the RNA-protein interactions using sequence information. The deep learning convolution neural network (CNN) algorithm is utilized to mine the hidden high-level discriminative features from the RNA and protein sequences and feed it into the extreme learning machine (ELM) classifier. The experimental results with 5-fold cross-validation indicate that the proposed method achieves superior performance on benchmark datasets (RPI1807, RPI2241, and RPI369) with the accuracy of 98.83, 90.83, and 85.63 percent, respectively. We further evaluate the performance of the proposed model by comparing it with the state-of-the-art SVM classifier and other existing methods on the same benchmark data set. In addition, we predicted the independent NPInter v2.0 data set using the model trained on RPI369. The experimental results show that our model can serve as a useful tool for predicting RNA-protein interactions. Lei Wang 0121, Zhu-Hong You, De-Shuang Huang, Fengfeng Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | DietLens-Eout: Large Scale Restaurant Food Photo RecognitionabstractRestaurant dishes represent a significant portion of food that people consume in their daily life. While people are becoming health-conscious in their food intake, convenient restaurant food tracking becomes an essential task in wellness and fitness applications. Given the huge number of dishes (food categories) involved, it becomes extremely challenging for traditional food photo classification to be feasible in both algorithm design and training data availability. In this work, we present a demo that runs on restaurant dish images in a city of millions of residents and tens of thousand restaurants. We propose a rank-loss based convolutional neural network to optimize the image features representation. Context information such as GPS location of the recognition request is also used to further improve the performance. Our experimental results are highly promising. We have shown in our demo that the proposed algorithm is near ready to be deployed in real-world applications. Zhipeng Wei 0001, Jingjing Chen 0001, Zhaoyan Ming, Chong-Wah Ngo, Tat-Seng Chua, Fengfeng Zhou |
ICMR | 6 |
| 2018 | pyHIVE, a health-related image visualization and engineering system using PythonabstractBACKGROUND: Imaging is one of the major biomedical technologies to investigate the status of a living object. But the biomedical image based data mining problem requires extensive knowledge across multiple disciplinaries, e.g. biology, mathematics and computer science, etc. RESULTS: pyHIVE (a Health-related Image Visualization and Engineering system using Python) was implemented as an image processing system, providing five widely used image feature engineering algorithms. A standard binary classification pipeline was also provided to help researchers build data models immediately after the data is collected. pyHIVE may calculate five widely-used image feature engineering algorithms efficiently using multiple computing cores, and also featured the modules of Principal Component Analysis (PCA) based preprocessing and normalization. CONCLUSIONS: The demonstrative example shows that the image features generated by pyHIVE achieved very good classification performances based on the gastrointestinal endoscopic images. This system pyHIVE and the demonstrative example are freely available and maintained at http://www.healthinformaticslab.org/supp/resources.php . Ruochi Zhang, Ruixue Zhao, Xin Feng 0004, Fengfeng Zhou |
BMC Bioinform. | 7 |
| 2017 | hMuLab: A Biomedical Hybrid MUlti-LABel Classifier Based on Multiple Linear RegressionabstractMany biomedical classification problems are multi-label by nature, e.g., a gene involved in a variety of functions and a patient with multiple diseases. The majority of existing classification algorithms assumes each sample with only one class label, and the multi-label classification problem remains to be a challenge for biomedical researchers. This study proposes a novel multi-label learning algorithm, hMuLab, by integrating both feature-based and neighbor-based similarity scores. The multiple linear regression modeling techniques make hMuLab capable of producing multiple label assignments for a query sample. The comparison results over six commonly-used multi-label performance measurements suggest that hMuLab performs accurately and stably for the biomedical datasets, and may serve as a complement to the existing literature. Ruiquan Ge, Manli Zhou, Fengfeng Zhou |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2016 | McTwo: a two-step feature selection algorithm based on maximal information coefficientabstractBACKGROUND: High-throughput bio-OMIC technologies are producing high-dimension data from bio-samples at an ever increasing rate, whereas the training sample number in a traditional experiment remains small due to various difficulties. This "large p, small n" paradigm in the area of biomedical "big data" may be at least partly solved by feature selection algorithms, which select only features significantly associated with phenotypes. Feature selection is an NP-hard problem. Due to the exponentially increased time requirement for finding the globally optimal solution, all the existing feature selection algorithms employ heuristic rules to find locally optimal solutions, and their solutions achieve different performances on different datasets. RESULTS: This work describes a feature selection algorithm based on a recently published correlation measurement, Maximal Information Coefficient (MIC). The proposed algorithm, McTwo, aims to select features associated with phenotypes, independently of each other, and achieving high classification performance of the nearest neighbor algorithm. Based on the comparative study of 17 datasets, McTwo performs about as well as or better than existing algorithms, with significantly reduced numbers of selected features. The features selected by McTwo also appear to have particular biomedical relevance to the phenotypes from the literature. CONCLUSION: McTwo selects a feature subset with very good classification performance, as well as a small feature number. So McTwo may represent a complementary feature selection algorithm for the high-dimensional biomedical datasets. Ruiquan Ge, Manli Zhou, Youxi Luo, Qinghan Meng, Guoqin Mai, Dongli Ma, Fengfeng Zhou |
BMC Bioinform. | 8 |
| 2014 | WinHAP2: an extremely fast haplotype phasing program for long genotype sequencesabstractBACKGROUND: The haplotype phasing problem tries to screen for phenotype associated genomic variations from millions of candidate data. Most of the current computer programs handle this problem with high requirements of computing power and memory. By replacing the computation-intensive step of constructing the maximum spanning tree with a heuristics of estimated initial haplotype, we released the WinHAP algorithm version 1.0, which outperforms the other algorithms in terms of both running speed and overall accuracy. RESULTS: This work further speeds up the WinHAP algorithm to version 2.0 (WinHAP2) by utilizing the divide-and-conquer strategy and the OpenMP parallel computing mode. WinHAP2 can phase 500 genotypes with 1,000,000 SNPs using just 12.8 MB in memory and 2.5 hours on a personal computer, whereas the other programs require unacceptable memory or running times. The parallel running mode further improves WinHAP2's running speed with several orders of magnitudes, compared with the other programs, including Beagle, SHAPEIT2 and 2SNP. CONCLUSIONS: WinHAP2 is an extremely fast haplotype phasing program which can handle a large-scale genotyping study with any number of SNPs in the current literature and at least in the near future. Weihua Pan, Fengfeng Zhou |
BMC Bioinform. | 4 |
| 2010 | cBar: a computer program to distinguish plasmid-derived from chromosome-derived sequence fragments in metagenomics dataabstractSUMMARY: Huge amount of metagenomic sequence data have been produced as a result of the rapidly increasing efforts worldwide in studying microbial communities as a whole. Most, if not all, sequenced metagenomes are complex mixtures of chromosomal and plasmid sequence fragments from multiple organisms, possibly from different kingdoms. Computational methods for prediction of genomic elements such as genes are significantly different for chromosomes and plasmids, hence raising the need for separation of chromosomal from plasmid sequences in a metagenome. We present a program for classification of a metagenome set into chromosomal and plasmid sequences, based on their distinguishing pentamer frequencies. On a large training set consisting of all the sequenced prokaryotic chromosomes and plasmids, the program achieves approximately 92% in classification accuracy. On a large set of simulated metagenomes with sequence lengths ranging from 300 bp to 100 kbp, the program has classification accuracy from 64.45% to 88.75%. On a large independent test set, the program achieves 88.29% classification accuracy. AVAILABILITY: The program has been implemented as a standalone prediction program, cBar, which is available at http://csbl.bmb.uga.edu/~ffzhou/cBar. Fengfeng Zhou, Ying Xu 0001 |
Bioinform. | 1 |
| 2009 | De novo computational prediction of non-coding RNA genes in prokaryotic genomesabstractMOTIVATION: The computational identification of non-coding RNA (ncRNA) genes represents one of the most important and challenging problems in computational biology. Existing methods for ncRNA gene prediction rely mostly on homology information, thus limiting their applications to ncRNA genes with known homologues. RESULTS: We present a novel de novo prediction algorithm for ncRNA genes using features derived from the sequences and structures of known ncRNA genes in comparison to decoys. Using these features, we have trained a neural network-based classifier and have applied it to Escherichia coli and Sulfolobus solfataricus for genome-wide prediction of ncRNAs. Our method has an average prediction sensitivity and specificity of 68% and 70%, respectively, for identifying windows with potential for ncRNA genes in E.coli. By combining windows of different sizes and using positional filtering strategies, we predicted 601 candidate ncRNAs and recovered 41% of known ncRNAs in E.coli. We experimentally investigated six novel candidates using Northern blot analysis and found expression of three candidates: one represents a potential new ncRNA, one is associated with stable mRNA decay intermediates and one is a case of either a potential riboswitch or transcription attenuator involved in the regulation of cell division. In general, our approach enables the identification of both cis- and trans-acting ncRNAs in partially or completely sequenced microbial genomes without requiring homology or structural conservation. AVAILABILITY: The source code and results are available at http://csbl.bmb.uga.edu/publications/materials/tran/. Thao T. Tran, Fengfeng Zhou, Sarah Marshburn, Mark Stead, Sidney R. Kushner, Ying Xu 0001 |
Bioinform. | 2 |
| 2008 | Barcodes for genomes and applicationsabstractBACKGROUND: Each genome has a stable distribution of the combined frequency for each k-mer and its reverse complement measured in sequence fragments as short as 1000 bps across the whole genome, for 1<k<6. The collection of these k-mer frequency distributions is unique to each genome and termed the genome's barcode. RESULTS: We found that for each genome, the majority of its short sequence fragments have highly similar barcodes while sequence fragments with different barcodes typically correspond to genes that are horizontally transferred or highly expressed. This observation has led to new and more effective ways for addressing two challenging problems: metagenome binning problem and identification of horizontally transferred genes. Our barcode-based metagenome binning algorithm substantially improves the state of the art in terms of both binning accuracies and the scope of applicability. Other attractive properties of genomes barcodes include (a) the barcodes have different and identifiable characteristics for different classes of genomes like prokaryotes, eukaryotes, mitochondria and plastids, and (b) barcodes similarities are generally proportional to the genomes' phylogenetic closeness. CONCLUSION: These and other properties of genomes barcodes make them a new and effective tool for studying numerous genome and metagenome analysis problems. Fengfeng Zhou, Victor Olman, Ying Xu 0001 |
BMC Bioinform. | 1 |
| 2006 | CSS-Palm: palmitoylation site prediction with a clustering and scoring strategy (CSS)abstractUNLABELLED: Palmitoylation is an important post-translational lipid modification of proteins. Unlike prenylation and myristoylation, palmitoylation is a reversible covalent modification, allowing for dynamic regulation of multiple complex cellular systems. However, in vivo or in vitro identification of palmitoylation sites is usually time-consuming and labor-intensive. So in silico predictions could help to narrow down the possible palmitoylation sites, which can be used to guide further experimental design. Previous studies suggested that there is no unique canonical motif for palmitoylation sites, so we hypothesize that the bona fide pattern might be compromised by heterogeneity of multiple structural determinants with different features. Based on this hypothesis, we partition the known palmitoylation sites into three clusters and score the similarity between the query peptide and the training ones based on BLOSUM62 matrix. We have implemented a computer program for palmitoylation site prediction, Clustering and Scoring Strategy for Palmitoylation Sites Prediction (CSS-Palm) system, and found that the program's prediction performance is encouraging with highly positive Jack-Knife validation results (sensitivity 82.16% and specificity 83.17% for cut-off score 2.6). Our analyses indicate that CSS-Palm could provide a powerful and effective tool to studies of palmitoylation sites. AVAILABILITY: CSS-Palm is implemented in PHP/PERL+MySQL and can be freely accessed at http://bioinformatics.lcd-ustc.org/css_palm/ CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bionformatics online. Fengfeng Zhou, Yu Xue 0001, Xuebiao Yao, Ying Xu 0001 |
Bioinform. | 1 |
| 2005 | No-wait scheduling in single-hop multi-channel LANs
Fengfeng Zhou, Yinlong Xu 0001, Guoliang Chen 0001 |
Inf. Process. Lett. | 1 |
| 2003 | Minimizing ADMs on WDM Directed Fiber Trees
Fengfeng Zhou, Guoliang Chen 0001, Yinlong Xu 0001 |
J. Comput. Sci. Technol. | 1 |