Junjie Chen 0004

dblp:04/4498-4 · also JunJie Chen 0004 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-0483-303XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Multi-modal Hierarchical Clustering Network for Cancer Subtype Identification of Multi-omics Data
abstract
Multi-modal clustering is an effective method for integrating multiple omics data in identifying and analyzing cancer diseases, enabling the unsupervised discovery of latent cluster patterns within multi-omics cancer data. However, existing multi-modal clustering methods primarily focus on single-level flat partitions, often overlooking the hierarchical subtype structures of real-world data. To address this issue, we propose a novel Multi-modal Hierarchical Clustering method for cancer Subtype identification of multi-omics data, termed Subtype-MHC. Specifically, Subtype-MHC is a hyperbolic neural network incorporating multiple Poincaré autoencoders, where the latent representations of each omic modality are mapped from the Euclidean space to the hyperbolic space, thus explicitly modeling hierarchical subtype structures during multi-omics integration process. On one hand, we leverage inter-omics complementarity by utilizing an inter-omics Poincaré reconstruction loss to capture modality-specific information within each omic, while a hyperbolic diversity loss is introduced to promote cluster separability on the Poincaré ball. On the other hand, in order to exploit cross-omics consistency, we design a self-weighted cross-omics integration loss to extract the shared hierarchies across all omic modalities. Additionally, a prototype-based weighting strategy is applied on the representation alignment, compressing task-relevant cluster information and mitigating the representation degradation caused by the quality differences of all omic modalities. Extensive experiments on ten multi-omics datasets demonstrate the hierarchical representation ability and clustering effectiveness of Subtype-MHC for cancer subtype identification.
Fangfei Lin, Jie Xu 0044, Yazhou Ren 0001, Junjie Chen 0004, Irwin King, Zenglin Xu
IJCNN4
2025 Directed evolution of antimicrobial peptides using multi-objective zeroth-order optimization
abstract
Antimicrobial peptides (AMPs) emerge as a type of promising therapeutic compounds that exhibit broad spectrum antimicrobial activity with high specificity and good tolerability. Natural AMPs usually need further rational design for improving antimicrobial activity and decreasing toxicity to human cells. Although several algorithms have been developed to optimize AMPs with desired properties, they explored the variations of AMPs in a discrete amino acid sequence space, usually suffering from low efficiency, lack diversity, and local optimum. In this work, we propose a novel directed evolution method, named PepZOO, for optimizing multi-properties of AMPs in a continuous representation space guided by multi-objective zeroth-order optimization. PepZOO projects AMPs from a discrete amino acid sequence space into continuous latent representation space by a variational autoencoder. Subsequently, the latent embeddings of prototype AMPs are taken as start points and iteratively updated according to the guidance of multi-objective zeroth-order optimization. Experimental results demonstrate PepZOO outperforms state-of-the-art methods on improving the multi-properties in terms of antimicrobial function, activity, toxicity, and binding affinity to the targets. Molecular docking and molecular dynamics simulations are further employed to validate the effectiveness of our method. Moreover, PepZOO can reveal important motifs which are required to maintain a particular property during the evolution by aligning the evolutionary sequences. PepZOO provides a novel research paradigm that optimizes AMPs by exploring property change instead of exploring sequence mutations, accelerating the discovery of potential therapeutic peptides.
Xianliang Liu, Yang Zhang 0057, Junjie Chen 0004
Briefings Bioinform.5
2025 INAB: identify nucleic acid binding domain via cross-modal protein language models and multiscale computation
abstract
Protein-nucleic acid interactions play a crucial role in biological processes, including gene regulation and editing. Accurately identifying nucleic acid-binding domains in proteins is essential to unravel these interactions, yet traditional experimental methods like X-ray crystallography remain costly and time-intensive. Computational approaches have thus emerged as indispensable tools to complement wet-lab techniques. Here, we introduce a framework for nucleic acid-binding domain prediction by integrating cross-modal protein language models with a multiscale computational architecture. The proposed method leverages a structurally annotated benchmark dataset, which quantifies binding likelihood through hierarchical, proximity-based labels derived from experimental complexes. Evaluations demonstrate that the approach achieves state-of-the-art performance, providing a new insight into the design of multimodal learning systems in protein-nucleic acid interaction analysis and an open resource to accelerate discoveries in functional genomics and drug design.
Jun Zhang 0078, Junjie Chen 0004, Zexuan Zhu 0001
Briefings Bioinform.3
2025 HiPHD: Hierarchical Classification for Protein Remote Homology Detection by Incorporating Protein Sequential and Structural Information
abstract
Protein remote homology detection is crucial in various biological tasks, such as protein function annotation and structure prediction. Computational methods have been developed to improve efficiency and accuracy of protein homology detection. However, proteins with remote homology usually share similar structures and low sequence identity, resulting in limited performance of sequence/structure alignment-based methods. This study introduces HiPHD, a hierarchical classification framework for protein remote homology detection. HiPHD integrates protein sequential information embedded by protein language models and structural information encoded by graph neural networks, effectively combining spatial and sequential features. Experimental results demonstrate that HiPHD outperforms existing methods in terms of accuracy at all hierarchical levels in both SCOPe and CATH databases. It's anticipated HiPHD will become a valuable tool for protein homology detection and representation learning.
Fuchuan Qu, Yijin Zhao, Jun Zhang 0078, Junjie Chen 0004
IEEE Trans. Comput. Biol. Bioinform.6
2025 ProFun-SOM: Protein Function Prediction for Specific Ontology Based on Multiple Sequence Alignment Reconstruction
abstract
Protein function prediction is crucial for understanding species evolution, including viral mutations. Gene ontology (GO) is a standardized representation framework for describing protein functions with annotated terms. Each ontology is a specific functional category containing multiple child ontologies, and the relationships of parent and child ontologies create a directed acyclic graph. Protein functions are categorized using GO, which divides them into three main groups: cellular component ontology, molecular function ontology, and biological process ontology. Therefore, the GO annotation of protein is a hierarchical multilabel classification problem. This hierarchical relationship introduces complexities such as mixed ontology problem, leading to performance bottlenecks in existing computational methods due to label dependency and data sparsity. To overcome bottleneck issues brought by mixed ontology problem, we propose ProFun-SOM, an innovative multilabel classifier that utilizes multiple sequence alignments (MSAs) to accurately annotate gene ontologies. ProFun-SOM enhances the initial MSAs through a reconstruction process and integrates them into a deep learning architecture. It then predicts annotations within the cellular component, molecular function, biological process, and mixed ontologies. Our evaluation results on three datasets (CAFA3, SwissProt, and NetGO2) demonstrate that ProFun-SOM surpasses state-of-the-art methods. This study confirmed that utilizing MSAs of proteins can effectively overcome the two main bottlenecks issues, label dependency and data sparsity, thereby alleviating the root problem, mixed ontology. A freely accessible web server is available at http://bliulab.net/ ProFun-SOM/.
Jiangyi Shao, Junjie Chen 0004, Bin Liu 0014
IEEE Trans. Neural Networks Learn. Syst.2
2024 Discriminative Forests Improve Generative Diversity for Generative Adversarial Networks
abstract
Improving the diversity of Artificial Intelligence Generated Content (AIGC) is one of the fundamental problems in the theory of generative models such as generative adversarial networks (GANs). Previous studies have demonstrated that the discriminator in GANs should have high capacity and robustness to achieve the diversity of generated data. However, a discriminator with high capacity tends to overfit and guide the generator toward collapsed equilibrium. In this study, we propose a novel discriminative forest GAN, named Forest-GAN, that replaces the discriminator to improve the capacity and robustness for modeling statistics in real-world data distribution. A discriminative forest is composed of multiple independent discriminators built on bootstrapped data. We prove that a discriminative forest has a generalization error bound, which is determined by the strength of individual discriminators and the correlations among them. Hence, a discriminative forest can provide very large capacity without any risk of overfitting, which subsequently improves the generative diversity. With the discriminative forest framework, we significantly improved the performance of AutoGAN with a new record FID of 19.27 from 30.71 on STL10 and improved the performance of StyleGAN2-ADA with a new record FID of 6.87 from 9.22 on LSUN-cat.
Junjie Chen 0004, Qingcai Chen, Hongchang Gao, Wendy Hui Wang, Zenglin Xu, Xinghua Shi
AAAI1
2023 CFAGO: cross-fusion of network and attributes based on attention mechanism for protein function prediction
abstract
MOTIVATION: Protein function annotation is fundamental to understanding biological mechanisms. The abundant genome-scale protein-protein interaction (PPI) networks, together with other protein biological attributes, provide rich information for annotating protein functions. As PPI networks and biological attributes describe protein functions from different perspectives, it is highly challenging to cross-fuse them for protein function prediction. Recently, several methods combine the PPI networks and protein attributes via the graph neural networks (GNNs). However, GNNs may inherit or even magnify the bias caused by noisy edges in PPI networks. Besides, GNNs with stacking of many layers may cause the over-smoothing problem of node representations. RESULTS: We develop a novel protein function prediction method, CFAGO, to integrate single-species PPI networks and protein biological attributes via a multi-head attention mechanism. CFAGO is first pre-trained with an encoder-decoder architecture to capture the universal protein representation of the two sources. It is then fine-tuned to learn more effective protein representations for protein function prediction. Benchmark experiments on human and mouse datasets show CFAGO outperforms state-of-the-art single-species network-based methods by at least 7.59%, 6.90%, 11.68% in terms of m-AUPR, M-AUPR, and Fmax, respectively, demonstrating cross-fusion by multi-head attention mechanism can greatly improve the protein function prediction. We further evaluate the quality of captured protein representations in terms of Davies Bouldin Score, whose results show that cross-fused protein representations by multi-head attention mechanism are at least 2.7% better than that of original and concatenated representations. We believe CFAGO is an effective tool for protein function prediction. AVAILABILITY AND IMPLEMENTATION: The source code of CFAGO and experiments data are available at: http://bliulab.net/CFAGO/.
Zhourun Wu, Mingyue Guo 0001, Xiaopeng Jin, Junjie Chen 0004, Bin Liu 0014
Bioinform.4
2022 ProtRe-CN: Protein Remote Homology Detection by Combining Classification Methods and Network Methods via Learning to Rank
abstract
Protein remote homology detection is one of fundamental research tasks for downstream analysis (i.e., protein structure and function prediction). Many advanced methods are proposed from different views with complementary detection ability, such as the classification method, the network method, and the ranking method. A framework integrating these heterogeneous methods is urgently desired to reduce the false positive rate and predictive bias. We propose a novel ranking method called ProtRe-CN by fusing the classification methods and network methods via Learning to Rank. Experimental results on the benchmark dataset and the independent dataset show that ProtRe-CN outperforms other existing state-of-the-art predictors. ProtRe-CN improves the detective performance via correcting the false positives in the ranking list by combining the heterogeneous methods. The web server of ProtRe-CN can be accessed at http://bliulab.net/ProtRe-CN.
Jiangyi Shao, Junjie Chen 0004, Bin Liu 0014
IEEE ACM Trans. Comput. Biol. Bioinform.2
2021 PAR-GAN: Improving the Generalization of Generative Adversarial Networks Against Membership Inference Attacks
abstract
Recent works have shown that Generative Adversarial Networks (GANs) may generalize poorly and thus are vulnerable to privacy attacks. In this paper, we seek to improve the generalization of GANs from a perspective of privacy protection, specifically in terms of defending against the membership inference attack (MIA) which aims to infer whether a particular sample was used for model training. We design a GAN framework, partition GAN (PAR-GAN), which consists of one generator and multiple discriminators trained over disjoint partitions of the training data. The key idea of PAR-GAN is to reduce the generalization gap by approximating a mixture distribution of all partitions of the training data. Our theoretical analysis shows that PAR-GAN can achieve global optimality just like the original GAN. Our experimental results on simulated data and multiple popular datasets demonstrate that PAR-GAN can improve the generalization of GANs while mitigating information leakage induced by MIA.
Junjie Chen 0004, Wendy Hui Wang, Hongchang Gao, Xinghua Shi
KDD1
2019 Protein Remote Homology Detection and Fold Recognition Based on Sequence-Order Frequency Matrix
abstract
Protein remote homology detection and fold recognition are two critical tasks for the studies of protein structures and functions. Currently, the profile-based methods achieve the state-of-the-art performance in these fields. However, the widely used sequence profiles, like position-specific frequency matrix (PSFM) and position-specific scoring matrix (PSSM), ignore the sequence-order effects along protein sequence. In this study, we have proposed a novel profile, called sequence-order frequency matrix (SOFM), to extract the sequence-order information of neighboring residues from multiple sequence alignment (MSA). Combined with two profile feature extraction approaches, top-n-grams and the Smith-Waterman algorithm, the SOFMs are applied to protein remote homology detection and fold recognition, and two predictors called SOFM-Top and SOFM-SW are proposed. Experimental results show that SOFM contains more information content than other profiles, and these two predictors outperform other state-of-the-art methods. It is anticipated that SOFM will become a very useful profile in the studies of protein structures and functions.
Bin Liu 0014, Junjie Chen 0004, Mingyue Guo 0001, Xiaolong Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 A comprehensive review and comparison of different computational methods for protein remote homology detection
abstract
Protein remote homology detection is one of the most fundamental and central problems for the studies of protein structures and functions, aiming to detect the distantly evolutionary relationships among proteins via computational methods. During the past decades, many computational approaches have been proposed to solve this important task. These methods have made a substantial contribution to protein remote homology detection. Therefore, it is necessary to give a comprehensive review and comparison on these computational methods. In this article, we divide these computational approaches into three categories, including alignment methods, discriminative methods and ranking methods. Their advantages and disadvantages are discussed in a comprehensive perspective, and their performance is compared on widely used benchmark data sets. Finally, some open questions in this field are further explored and discussed.
Junjie Chen 0004, Mingyue Guo 0001, Xiaolong Wang 0001, Bin Liu 0014
Briefings Bioinform.1
2017 SOFM-Top: Protein Remote Homology Detection and Fold Recognition Based on Sequence-Order Frequency Matrix
Junjie Chen 0004, Mingyue Guo 0001, Xiaolong Wang 0001, Bin Liu 0014
ICIC (2)1
2017 ProtDec-LTR2.0: an improved method for protein remote homology detection by combining pseudo protein and supervised Learning to Rank
abstract
SUMMARY: As one of the most important tasks in protein sequence analysis, protein remote homology detection is critical for both basic research and practical applications. Here, we present an effective web server for protein remote homology detection called ProtDec-LTR2.0 by combining ProtDec-Learning to Rank (LTR) and pseudo protein representation. Experimental results showed that the detection performance is obviously improved. The web server provides a user-friendly interface to explore the sequence and structure information of candidate proteins and find their conserved domains by launching a multiple sequence alignment tool. AVAILABILITY AND IMPLEMENTATION: The web server is free and open to all users with no login requirement at http://bioinformatics.hitsz.edu.cn/ProtDec-LTR2.0/. CONTACT: [email protected].
Junjie Chen 0004, Mingyue Guo 0001, Bin Liu 0014
Bioinform.1
2017 Protein remote homology detection based on bidirectional long short-term memory
abstract
BACKGROUND: Protein remote homology detection plays a vital role in studies of protein structures and functions. Almost all of the traditional machine leaning methods require fixed length features to represent the protein sequences. However, it is never an easy task to extract the discriminative features with limited knowledge of proteins. On the other hand, deep learning technique has demonstrated its advantage in automatically learning representations. It is worthwhile to explore the applications of deep learning techniques to the protein remote homology detection. RESULTS: In this study, we employ the Bidirectional Long Short-Term Memory (BLSTM) to learn effective features from pseudo proteins, also propose a predictor called ProDec-BLSTM: it includes input layer, bidirectional LSTM, time distributed dense layer and output layer. This neural network can automatically extract the discriminative features by using bidirectional LSTM and the time distributed dense layer. CONCLUSION: Experimental results on a widely-used benchmark dataset show that ProDec-BLSTM outperforms other related methods in terms of both the mean ROC and mean ROC50 scores. This promising result shows that ProDec-BLSTM is a useful tool for protein remote homology detection. Furthermore, the hidden patterns learnt by ProDec-BLSTM can be interpreted and visualized, and therefore, additional useful information can be obtained.
Junjie Chen 0004, Bin Liu 0014
BMC Bioinform.2
2015 Application of learning to rank to protein remote homology detection
abstract
MOTIVATION: Protein remote homology detection is one of the fundamental problems in computational biology, aiming to find protein sequences in a database of known structures that are evolutionarily related to a given query protein. Some computational methods treat this problem as a ranking problem and achieve the state-of-the-art performance, such as PSI-BLAST, HHblits and ProtEmbed. This raises the possibility to combine these methods to improve the predictive performance. In this regard, we are to propose a new computational method called ProtDec-LTR for protein remote homology detection, which is able to combine various ranking methods in a supervised manner via using the Learning to Rank (LTR) algorithm derived from natural language processing. RESULTS: Experimental results on a widely used benchmark dataset showed that ProtDec-LTR can achieve an ROC1 score of 0.8442 and an ROC50 score of 0.9023 outperforming all the individual predictors and some state-of-the-art methods. These results indicate that it is correct to treat protein remote homology detection as a ranking problem, and predictive performance improvement can be achieved by combining different ranking approaches in a supervised manner via using LTR. AVAILABILITY AND IMPLEMENTATION: For users' convenience, the software tools of three basic ranking predictors and Learning to Rank algorithm were provided at http://bioinformatics.hitsz.edu.cn/ProtDec-LTR/home/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bin Liu 0014, Junjie Chen 0004, Xiaolong Wang 0001
Bioinform.2