VLDB 2026 Research / reviewers in the wild / expert
Minfeng Xiao
dblp:290/5780
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-0507-7352ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PBIP: a deep learning framework for predicting phage-bacterium interactions at the strain levelabstractPhage therapy has received great attention as a promising antimicrobial treatment, and its core technique, namely predicting phage-bacterium interactions (PBIs), is crucial for understanding infection mechanisms and optimizing therapeutic strategies. However, existing computational methods mainly focus on the species or higher taxonomic levels, and usually neglect the potential of deep embedding representations, limiting their ability to capture complex biological patterns inherent in sequences. This hinders the discovery of rich sequence features, and restricts the clinical application of phage therapy. To address these limitations, we propose a novel deep learning framework (called PBIP) for strain-level PBI prediction. In PBIP, we first identify strain-level interactions through biological infection experiments and sequencing of Klebsiella pneumoniae isolated from the clinical environment of Xiangya Hospital. Then, we utilize a pretrained unified representation model to convert protein sequences of phages and bacteria into deep embeddings. Next, we apply the synthetic minority oversampling technique to generate positive interactions in the embedding space to address the data imbalance issue. Subsequently, we design a deep neural network that uses a convolutional neural network to extract local features, a bi-directional gated recurrent unit to capture global features, and an attention module to highlight significant features. Finally, a fully connected layer integrates this information for PBI prediction. Experimental results show the superiority of PBIP over the state-of-the-art methods in predicting PBIs. The code and datasets are available at https://github.com/a1678019300/PBIP. Lijia Ma, Gufeng Liu, Yuan Bai, Qiuzhen Lin, Jianqiang Li 0001, Minfeng Xiao |
Briefings Bioinform. | 7 |
| 2025 | BERTPVP: Identifying and Classifying Phage Virion Proteins Using Bidirectional Encoder Representations-Based TransformersabstractPhage virion proteins (PVPs), which form the structural components of phages, are crucial for maintaining phage structures and infecting host bacteria. Identifying PVPs can lead to the development of novel therapeutic agents to combat bacterial infections, attracting great research attention in recent years. However, most of the existing methods for PVP identification heavily depend on the effectiveness of feature extraction and have no ability to precisely classify specific classes. In this article, we propose a bidirectional encoder representations-based Transformer model called BERTPVP for the identification and classification of PVPs. BERTPVP uses a stack of transformer encoders to effectively capture contextual information from the entire protein sequence through the multi-head self-attention mechanism. We firstly pre-train the model using the masked language modeling task to learn contextual information from phage protein sequences, which aids in understanding the significance and relevance of individual components in relation to the entire sequence. Subsequently, the pre-trained model is fine-tuned for PVP identification and classification tasks. Our experimental results show the superiority of the proposed BERTPVP over the state-of-the-art methods in identifying and classifying PVPs. Moreover, the ablation study demonstrates the necessity of the pre-training and fine-tuning components of BERTPVP in accelerating convergence and improving prediction performance. Lijia Ma, Wenxiang Zhou, Yuan Bai, Minfeng Xiao, Jianqiang Li 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2023 | PhaGAA: an integrated web server platform for phage genome annotation and analysisabstractMOTIVATION: Phage genome annotation plays a key role in the design of phage therapy. To date, there have been various genome annotation tools for phages, but most of these tools focus on mono-functional annotation and have complex operational processes. Accordingly, comprehensive and user-friendly platforms for phage genome annotation are needed. RESULTS: Here, we propose PhaGAA, an online integrated platform for phage genome annotation and analysis. By incorporating several annotation tools, PhaGAA is constructed to annotate the prophage genome at DNA and protein levels and provide the analytical results. Furthermore, PhaGAA could mine and annotate phage genomes from bacterial genome or metagenome. In summary, PhaGAA will be a useful resource for experimental biologists and help advance the phage synthetic biology in basic and application research. AVAILABILITY AND IMPLEMENTATION: PhaGAA is freely available at http://phage.xialab.info/. Qingrui Liu, Jiliang Xu, Junyin Zhang, Minfeng Xiao, Yannan Bin, Junfeng Xia |
Bioinform. | 7 |
| 2023 | Identifying Phage Sequences From Metagenomic Data Using Deep Neural Network With Word Embedding and Attention MechanismabstractPhages are the functional viruses that infect bacteria and they play important roles in microbial communities and ecosystems. Phage research has attracted great attention due to the wide applications of phage therapy in treating bacterial infection in recent years. Metagenomics sequencing technique can sequence microbial communities directly from an environmental sample. Identifying phage sequences from metagenomic data is a vital step in the downstream of phage analysis. However, the existing methods for phage identification suffer from some limitations in the utilization of the phage feature for prediction, and therefore their prediction performance still need to be improved further. In this article, we propose a novel deep neural network (called MetaPhaPred) for identifying phages from metagenomic data. In MetaPhaPred, we first use a word embedding technique to encode the metagenomic sequences into word vectors, extracting the latent feature vectors of DNA words. Then, we design a deep neural network with a convolutional neural network (CNN) to capture the feature maps in sequences, and with a bi-directional long short-term memory network (Bi-LSTM) to capture the long-term dependencies between features from both forward and backward directions. The feature map consists of a set of feature patterns, each of which is the weighted feature extracted by a convolution filter with convolution kernels in the CNN slide along the input feature vectors. Next, an attention mechanism is used to enhance contributions of important features. Experimental results on both simulated and real metagenomic data with different lengths demonstrate the superiority of the proposed MetaPhaPred over the state-of-the-art methods in identifying phage sequences. Lijia Ma, Wenwei Deng, Yuan Bai, Zhanwei Du, Minfeng Xiao, Lin Wang 0012, Jianqiang Li 0001, Asoke K. Nandi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | DPProm: A Two-Layer Predictor for Identifying Promoters and Their Types on Phage Genome Using Deep LearningabstractWith the number of phage genomes increasing, it is urgent to develop new bioinformatics methods for phage genome annotation. Promoter, a DNA region, is important for gene transcriptional regulation. In the era of post-genomics, the availability of data makes it possible to establish computational models for promoter identification with robustness. In this work, we introduce DPProm, a two-layer model composed of DPProm-1L and DPProm-2L, to predict promoters and their types for phages. On the first layer, as a dual-channel deep neural network ensemble method fusing multi-view features (sequence feature and handcrafted feature), the model DPProm-1L is proposed to identify whether a DNA sequence is a promoter or non-promoter. The sequence feature is extracted with convolutional neural network (CNN). And the handcrafted feature is the combination of free energy, GC content, cumulative skew, and Z curve features. On the second layer, DPProm-2L based on CNN is trained to predict the promoters' types (host or phage). For the realization of prediction on the whole genomes, the model DPProm, combines with a novel sequence data processing workflow, which contains sliding window and merging sequences modules. Experimental results show that DPProm outperforms the state-of-the-art methods, and decreases the false positive rate effectively on whole genome prediction. Furthermore, we provide a user-friendly web at http://bioinfo.ahu.edu.cn/DPProm. We expect that DPProm can serve as a useful tool for identification of promoters and their types. Junyin Zhang, Minfeng Xiao, Junfeng Xia, Yannan Bin |
IEEE J. Biomed. Health Informatics | 5 |