Zhixia Teng

dblp:130/1057 · DBLP profile ↗
← Back
17ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-6968-4354ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive multi-view information bottleneck for multi-omics data clustering
abstract
MOTIVATION: Recent advances in single-cell sequencing have transformed precise measurement of gene expression at cellular resolution, enabling unprecedented dissection of cellular heterogeneity and intricate biological processes. The accumulation of multi-omics data offers new avenues for cell clustering-a critical foundation for cell-type identification and downstream analyses. However, substantial challenges persist in simultaneously achieving effective integration of complementary information in multi-omics data and their appropriate weight allocation. RESULTS: Here, we propose an Adaptive Multi-View clustering framework with the Information Bottleneck principle to solve the multi-omics data clustering task (named scAMVIB). The proposed model could learn multi-view omics representations that capture both inter-omics associations and omics-specific patterns, with the adaptive weight allocation. Specifically, multi-view data comprise two components: (i) the integrated omics feature matrix derived from the similarity network fusion strategy and (ii) omics-specific representations from distinct platforms. These inputs are processed through a multi-view information bottleneck clustering framework that leverages cross-view complementarity to enhance representations. View weights are adaptively assigned via maximum entropy regularization, proportional to their information content. The final cell partitions are obtained through sequential iterative optimization. Comprehensive experiments across multiple datasets demonstrate that scAMVIB has strong competitiveness in clustering while maintaining biological interpretability.
Zhen Tian 0004, Xiaojiao Wei, Zhengzheng Lou, Zhixia Teng, Shouli Fu
Briefings Bioinform.4
2026 MFCLDTA: Multi-scale feature contrastive learning for predicting drug-target binding affinity
Zhen Tian 0004, Saisai Zhu, Zhixia Teng, Tao Wang 0082
Expert Syst. Appl.3
2025 Cancer Drug Response Prediction Via Cross-Modal Multilevel Homogeneous and Heterogeneous Feature Learning
abstract
Accurately predicting cancer drug response (CDR) is crucial for personalized cancer therapies and drug repositioning. Efficient CDR prediction requires to integrate multimodal data including sequences, structures, multilevel omics, and diverse biological networks of drugs and cell lines, to capture intricate underlying patterns. So we proposed TransGCDR, a novel method that integrates cross-modal multilevel homogeneous and heterogeneous features to derive highly discriminative embeddings, thereby enhancing CDR prediction. TransGCDR first learns structural representations of drugs using a Transformer from fingerprint substructures and Graph Convolutional Networks for molecular graphs. Meanwhile, it utilizes a Graph Attention Network to extract cell line representations from graphs integrating multilevel omics data, including gene expression, copy number variation, and somatic mutations. Next, it generates homogeneous embeddings for drugs and cell lines from these cross-modal representations while extracting drug- and cell linecentric heterogeneous contextual embeddings from prior CDRs. These embeddings are then integrated using a Residual Attention Graph Convolutional Network to obtain comprehensive representations. Finally, an MLP-based model is trained on the learned embeddings to predict CDRs. Extensive evaluations demonstrate that TransGCDR outperforms state-of-the-art methods in CDR prediction across various metrics. Ablation studies confirm that every module enhances the precise CDR characterization from the multimodal and multilevel data, which is key to the success of TransGCDR. Moreover, case studies highlight the significant clinical potential of TransGCDR to predict CDRs.
Zhixia Teng, Mingxin Yin, Guohua Wang 0001
BIBM1
2025 Cross-RNA transferable sequence representation learning for lncRNA m6A site detection via novel deep domain separation networks
abstract
N6-methyladenosine (m6A) is a key epitranscriptomic marker enriched in long noncoding RNAs (lncRNAs) that is closely involved in complex disease mechanisms. Although accurate detection of m6A sites in lncRNAs is essential for understanding disease mechanisms, the development of effective computational predictors remains challenging due to the limited number of annotated sites. Moreover, most existing predictors are specifically designed for messenger RNAs (mRNAs) based on abundant mRNA-specific knowledge, yet they exhibit limited generalizability to lncRNAs. Given the similarities between mRNAs and lncRNAs, a transferable framework capable of leveraging their shared features is critical for advancing m6A site prediction in lncRNAs. To address this challenge, we propose DSNm6A, a deep learning framework that learns cross-RNA transferable sequence representations for effective lncRNA m6A site detection. To comprehensively capture patterns and signals of m6A sites, lncRNA and mRNA sequences are first encoded from complementary multiple facets, including One-Hot encoding, nucleotide physicochemical properties and cumulative frequency, and position-specific propensity. Based on these sequence encodings, a domain separation network integrating CNN, Bi-LSTM, and BERT modules is then employed to explicitly disentangle domain-invariant features shared between mRNAs and lncRNAs from their domain-specific counterparts. The shared features are finally fed into a fully connected layer for accurate lncRNA m6A sites prediction. Cross-validation and independent test results demonstrate that DSNm6A consistently outperforms existing methods across nearly all performance metrics, attributed to its superior capacity to learn transferable m6A-related features across RNA types. In addition, DSNm6A exhibits strong robustness and generalization across species.
Zhixia Teng, Chunyu Wang 0002, Guohua Wang 0001
Briefings Bioinform.1
2025 Long Noncoding RNA function prediction via multiview cross-contrastive learning combined with multiscale semantic adaptive optimization
abstract
Understanding long noncoding RNA (lncRNA) function is essential for revealing molecular mechanisms and developing effective therapies for complex diseases, as lncRNAs play important regulatory roles in many disease-related biological processes. However, existing lncRNA function predictors struggle to extract discriminative features from multimodal omics data and to model the semantic and topological structure of the gene ontology (GO), which severely limits their ability to achieve biologically meaningful and functionally informative predictions. To address these challenges, we propose a novel framework for lncRNA function prediction, namely MiCLSAO. Firstly, MiCLSAO utilizes multiview cross-contrastive learning with attention mechanisms to extract highly discriminative lncRNA features from diverse omics similarity networks. Secondly, graph convolutional networks are applied to learn initial features of GO terms, while multiscale topological and semantic relationships are incorporated to adaptively refine term representations. Finally, an lncRNA function predictor is developed by dynamically integrating the representations of lncRNAs and GO terms using a Kolmogorov-Arnold network. Extensive experiments demonstrate that MiCLSAO consistently outperforms state-of-the-art methods across multiple metrics, with significant capability to recover known functions and uncover novel ones. Moreover, MiCLSAO demonstrates remarkable practical utility and potential value by providing more functionally informative annotations for lncRNAs.
Zhixia Teng, Qingqi Li, Chunyu Wang 0002, Guohua Wang 0001
Briefings Bioinform.1
2025 Enhancing LncRNA-miRNA interaction prediction with multimodal contrastive representation learning
abstract
Interactions between long non-coding RNAs (lncRNAs) and microRNAs (miRNAs) play an important role in the development of complex human diseases by collaboratively regulating gene transcription and expression. Therefore, identifying lncRNA-miRNA interactions (LMIs) is essential for diagnosing and treating complex human diseases. Because identifying LMIs with wet experiments is time-consuming and labor-intensive, some computational methods have been developed to infer LMIs. However, these approaches excel at utilizing single-modal information but struggle to integrate multimodal data from lncRNAs and miRNAs, which is essential for uncovering complex patterns in LMIs, ultimately limiting their performance. Therefore, this article proposes a novel multimodal contrastive representation learning model (MCRLMI) for LMI predictions. The model fully integrates multi-source similarity information and sequence encodings of lncRNAs and miRNAs. It leverages a graph convolutional network (GCN) and a Transformer to capture local neighborhood structural features and long-distance dependencies, respectively, enabling the collaborative modeling of structural and semantic information. Subsequently, to effectively integrate multimodal characteristics with encoded information, a multichannel attention mechanism and contrastive learning are introduced to fuse the extracted features. Finally, a Kolmogorov-Arnold Network (KAN) is trained with the optimized embeddings to predict LMIs. Extensive experiments show that the proposed MCRLMI consistently outperforms existing methods. Moreover, case studies further validate the potential of MCRLMI to identify novel LMIs in practical applications.
Zhixia Teng, Zhaowen Tian, Murong Zhou, Guohua Wang 0001, Zhen Tian 0004
Briefings Bioinform.1
2025 Hierarchical feature aggregation with mixed attention mechanism for single-cell RNA-seq analysis
Wanning Zhou, Zhixia Teng, Zhen Tian 0004
Expert Syst. Appl.5
2025 Drug repositioning by collaborative learning based on graph convolutional inductive network
Zhixia Teng, Yongliang Li, Zhen Tian 0004, Yingjian Liang, Guohua Wang 0001
Future Gener. Comput. Syst.1
2025 DSANIB: Drug-Target Interaction Predictions With Dual-View Synergistic Attention Network and Information Bottleneck Strategy
abstract
Prediction of drug-target interactions (DTIs) is one of the crucial steps for drug repositioning. Identifying DTIs through bio-experimental manners is always expensive and time-consuming. Recently, deep learning-based approaches have shown promising advancements in DTI prediction, but they face two notable challenges: (i) how to explicitly capture local interactions between drug-target pairs and learn their higher-order substructure embeddings; (ii) How to filter out redundant information to obtain effective embeddings for drugs and targets. Results: In this study, we propose a novel approach, termed DSANIB, to infer potential interactions between drugs and targets. DSANIB comprises two primary components: (1) DSAN component: The Inter-view Attention Network Module explicitly learns the local interactions between drugs and targets, while the Intra-view Attention Network Module aggregates information from local interaction features to obtain their higher-order substructure embeddings. (2) Information Bottleneck (IB) component: DSANIB adopts the IB strategy, which could retain relevant information while minimizing the redundant features to obtain their discriminative representations. Extensive experimental results demonstrate that DSANIB outperforms other SOTA prediction models. In addition, visualization of drug and target embeddings learned through DSANIB could provide interpretable insights for the prediction results.
Zhen Tian 0004, Wanning Zhou, Zhixia Teng, Quan Zou 0001
IEEE J. Biomed. Health Informatics4
2024 MGCNSS: miRNA-disease association prediction with multi-layer graph convolution and distance-based negative sample selection strategy
abstract
Identifying disease-associated microRNAs (miRNAs) could help understand the deep mechanism of diseases, which promotes the development of new medicine. Recently, network-based approaches have been widely proposed for inferring the potential associations between miRNAs and diseases. However, these approaches ignore the importance of different relations in meta-paths when learning the embeddings of miRNAs and diseases. Besides, they pay little attention to screening out reliable negative samples which is crucial for improving the prediction accuracy. In this study, we propose a novel approach named MGCNSS with the multi-layer graph convolution and high-quality negative sample selection strategy. Specifically, MGCNSS first constructs a comprehensive heterogeneous network by integrating miRNA and disease similarity networks coupled with their known association relationships. Then, we employ the multi-layer graph convolution to automatically capture the meta-path relations with different lengths in the heterogeneous network and learn the discriminative representations of miRNAs and diseases. After that, MGCNSS establishes a highly reliable negative sample set from the unlabeled sample set with the negative distance-based sample selection strategy. Finally, we train MGCNSS under an unsupervised learning manner and predict the potential associations between miRNAs and diseases. The experimental results fully demonstrate that MGCNSS outperforms all baseline methods on both balanced and imbalanced datasets. More importantly, we conduct case studies on colon neoplasms and esophageal neoplasms, further confirming the ability of MGCNSS to detect potential candidate miRNAs. The source code is publicly available on GitHub https://github.com/15136943622/MGCNSS/tree/master.
Zhen Tian 0004, Chenguang Han, Lewen Xu, Zhixia Teng
Briefings Bioinform.4
2023 GOGCN: Graph Convolutional Network on Gene Ontology for Functional Similarity Analysis of Genes
abstract
The measurement of gene functional similarity plays a critical role in numerous biological applications, such as gene clustering, the construction of gene similarity networks. However, most existing approaches still rely heavily on traditional computational strategies, which are not guaranteed to achieve satisfactory performance. In this study, we propose a novel computational approach called GOGCN to measure gene functional similarity by modeling the Gene Ontology (GO) through Graph Convolutional Network (GCN). GOGCN is a graph-based approach that performs sufficient representation learning for terms and relations in the GO graph. First, GOGCN employs the GCN-based knowledge graph embedding (KGE) model to learn vector representations (i.e., embeddings) for all entities (i.e., terms). Second, GOGCN calculates the semantic similarity between two terms based on their corresponding vector representations. Finally, GOGCN estimates gene functional similarity by making use of the pair-wise strategy. During the representation learning period, GOGCN promotes semantic interaction between terms through GCN, thereby capturing the rich structural information of the GO graph. Further experimental results on various datasets suggest that GOGCN is superior to the other state-of-the-art approaches, which shows its reliability and effectiveness.
Zhen Tian 0004, Haichuan Fang, Zhixia Teng, Yangdong Ye
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Predicting miRNA-disease associations via learning multimodal networks and fusing mixed neighborhood information
abstract
MOTIVATION: In recent years, a large number of biological experiments have strongly shown that miRNAs play an important role in understanding disease pathogenesis. The discovery of miRNA-disease associations is beneficial for disease diagnosis and treatment. Since inferring these associations through biological experiments is time-consuming and expensive, researchers have sought to identify the associations utilizing computational approaches. Graph Convolutional Networks (GCNs), which exhibit excellent performance in link prediction problems, have been successfully used in miRNA-disease association prediction. However, GCNs only consider 1st-order neighborhood information at one layer but fail to capture information from high-order neighbors to learn miRNA and disease representations through information propagation. Therefore, how to aggregate information from high-order neighborhood effectively in an explicit way is still challenging. RESULTS: To address such a challenge, we propose a novel method called mixed neighborhood information for miRNA-disease association (MINIMDA), which could fuse mixed high-order neighborhood information of miRNAs and diseases in multimodal networks. First, MINIMDA constructs the integrated miRNA similarity network and integrated disease similarity network respectively with their multisource information. Then, the embedding representations of miRNAs and diseases are obtained by fusing mixed high-order neighborhood information from multimodal network which are the integrated miRNA similarity network, integrated disease similarity network and the miRNA-disease association networks. Finally, we concentrate the multimodal embedding representations of miRNAs and diseases and feed them into the multilayer perceptron (MLP) to predict their underlying associations. Extensive experimental results show that MINIMDA is superior to other state-of-the-art methods overall. Moreover, the outstanding performance on case studies for esophageal cancer, colon tumor and lung cancer further demonstrates the effectiveness of MINIMDA. AVAILABILITY AND IMPLEMENTATION: https://github.com/chengxu123/MINIMDA and http://120.79.173.96/.
Zhengzheng Lou, Zhaoxu Cheng, Zhixia Teng, Zhen Tian 0004
Briefings Bioinform.4
2021 ReRF-Pred: predicting amyloidogenic regions of proteins based on their pseudo amino acid composition and tripeptide composition
abstract
BACKGROUND: Amyloids are insoluble fibrillar aggregates that are highly associated with complex human diseases, such as Alzheimer's disease, Parkinson's disease, and type II diabetes. Recently, many studies reported that some specific regions of amino acid sequences may be responsible for the amyloidosis of proteins. It has become very important for elucidating the mechanism of amyloids that identifying the amyloidogenic regions. Accordingly, several computational methods have been put forward to discover amyloidogenic regions. The majority of these methods predicted amyloidogenic regions based on the physicochemical properties of amino acids. In fact, position, order, and correlation of amino acids may also influence the amyloidosis of proteins, which should be also considered in detecting amyloidogenic regions. RESULTS: To address this problem, we proposed a novel machine-learning approach for predicting amyloidogenic regions, called ReRF-Pred. Firstly, the pseudo amino acid composition (PseAAC) was exploited to characterize physicochemical properties and correlation of amino acids. Secondly, tripeptides composition (TPC) was employed to represent the order and position of amino acids. To improve the distinguishability of TPC, all possible tripeptides were analyzed by the binomial distribution method, and only those which have significantly different distribution between positive and negative samples remained. Finally, all samples were characterized by PseAAC and TPC of their amino acid sequence, and a random forest-based amyloidogenic regions predictor was trained on these samples. It was proved by validation experiments that the feature set consisted of PseAAC and TPC is the most distinguishable one for detecting amyloidosis. Meanwhile, random forest is superior to other concerned classifiers on almost all metrics. To validate the effectiveness of our model, ReRF-Pred is compared with a series of gold-standard methods on two datasets: Pep-251 and Reg33. The results suggested our method has the best overall performance and makes significant improvements in discovering amyloidogenic regions. CONCLUSIONS: The advantages of our method are mainly attributed to that PseAAC and TPC can describe the differences between amyloids and other proteins successfully. The ReRF-Pred server can be accessed at http://106.12.83.135:8080/ReRF-Pred/.
Zhixia Teng, Zitong Zhang 0002, Zhen Tian 0004, Yanjuan Li, Guohua Wang 0001
BMC Bioinform.1
2016 Revealing protein functions based on relationships of interacting proteins and GO terms
abstract
numerous computational methods predicted protein function based on the protein-protein interaction (PPI) network. These methods supposed that two proteins share the same function if they interact with each other. However, it is reported by recent studies that the functions of two interacting proteins may be related but different. In this paper, the functional relationship between interacting proteins is studied and a novel method, called as GoDIN, is advanced to annotate functions of the interacting protein in Gene Ontology (GO) context. It is assumed that the functional difference between interacting proteins can be expressed by semantic difference between GO term and its relatives. Thus, the method uses GO term and its relatives to annotate the interacting proteins separately according to their functional roles in the PPI network. The method is validated by a series of experiments and compared with the related method. The experimental results confirm the assumption and suggest that GoDIN is effective on predicting functions of protein.
Zhixia Teng, Maozu Guo 0001, Zhen Tian 0004, Kai Che
BIBM1
2016 SGFSC: speeding the gene functional similarity calculation based on hash tables
abstract
BACKGROUND: In recent years, many measures of gene functional similarity have been proposed and widely used in all kinds of essential research. These methods are mainly divided into two categories: pairwise approaches and group-wise approaches. However, a common problem with these methods is their time consumption, especially when measuring the gene functional similarities of a large number of gene pairs. The problem of computational efficiency for pairwise approaches is even more prominent because they are dependent on the combination of semantic similarity. Therefore, the efficient measurement of gene functional similarity remains a challenging problem. RESULTS: To speed current gene functional similarity calculation methods, a novel two-step computing strategy is proposed: (1) establish a hash table for each method to store essential information obtained from the Gene Ontology (GO) graph and (2) measure gene functional similarity based on the corresponding hash table. There is no need to traverse the GO graph repeatedly for each method with the help of the hash table. The analysis of time complexity shows that the computational efficiency of these methods is significantly improved. We also implement a novel Speeding Gene Functional Similarity Calculation tool, namely SGFSC, which is bundled with seven typical measures using our proposed strategy. Further experiments show the great advantage of SGFSC in measuring gene functional similarity on the whole genomic scale. CONCLUSIONS: The proposed strategy is successful in speeding current gene functional similarity calculation methods. SGFSC is an efficient tool that is freely available at http://nclab.hit.edu.cn/SGFSC . The source code of SGFSC can be downloaded from http://pan.baidu.com/s/1dFFmvpZ .
Zhen Tian 0004, Chunyu Wang 0002, Maozu Guo 0001, Zhixia Teng
BMC Bioinform.5
2014 CPL: Detecting Protein Complexes by Propagating Labels on Protein-Protein Interaction Network
Qiguo Dai, Maozu Guo 0001, Zhixia Teng, Chunyu Wang 0002
J. Comput. Sci. Technol.4
2013 Measuring gene functional similarity based on group-wise comparison of GO terms
abstract
MOTIVATION: Compared with sequence and structure similarity, functional similarity is more informative for understanding the biological roles and functions of genes. Many important applications in computational molecular biology require functional similarity, such as gene clustering, protein function prediction, protein interaction evaluation and disease gene prioritization. Gene Ontology (GO) is now widely used as the basis for measuring gene functional similarity. Some existing methods combined semantic similarity scores of single term pairs to estimate gene functional similarity, whereas others compared terms in groups to measure it. However, these methods may make error-prone judgments about gene functional similarity. It remains a challenge that measuring gene functional similarity reliably. RESULT: We propose a novel method called SORA to measure gene functional similarity in GO context. First of all, SORA computes the information content (IC) of a term making use of semantic specificity and coverage. Second, SORA measures the IC of a term set by means of combining inherited and extended IC of the terms based on the structure of GO. Finally, SORA estimates gene functional similarity using the IC overlap ratio of term sets. SORA is evaluated against five state-of-the-art methods in the file on the public platform for collaborative evaluation of GO-based semantic similarity measure. The carefully comparisons show SORA is superior to other methods in general. Further analysis suggests that it primarily benefits from the structure of GO, which implies expressive information about gene function. SORA offers an effective and reliable way to compare gene function. AVAILABILITY: The web service of SORA is freely available at http://nclab.hit.edu.cn/SORA/
Zhixia Teng, Maozu Guo 0001, Qiguo Dai, Chunyu Wang 0002, Ping Xuan
Bioinform.1