EDBT 2026 Demo / reviewers in the wild / expert
Zhen Tian 0004
dblp:84/8525-4
· DBLP profile ↗
33ranked-venue papers
13as first author
27since 2021 · last 2026
0000-0003-0945-8168ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 11 first-author · 18 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive multi-view information bottleneck for multi-omics data clusteringabstractMOTIVATION: Recent advances in single-cell sequencing have transformed precise measurement of gene expression at cellular resolution, enabling unprecedented dissection of cellular heterogeneity and intricate biological processes. The accumulation of multi-omics data offers new avenues for cell clustering-a critical foundation for cell-type identification and downstream analyses. However, substantial challenges persist in simultaneously achieving effective integration of complementary information in multi-omics data and their appropriate weight allocation. RESULTS: Here, we propose an Adaptive Multi-View clustering framework with the Information Bottleneck principle to solve the multi-omics data clustering task (named scAMVIB). The proposed model could learn multi-view omics representations that capture both inter-omics associations and omics-specific patterns, with the adaptive weight allocation. Specifically, multi-view data comprise two components: (i) the integrated omics feature matrix derived from the similarity network fusion strategy and (ii) omics-specific representations from distinct platforms. These inputs are processed through a multi-view information bottleneck clustering framework that leverages cross-view complementarity to enhance representations. View weights are adaptively assigned via maximum entropy regularization, proportional to their information content. The final cell partitions are obtained through sequential iterative optimization. Comprehensive experiments across multiple datasets demonstrate that scAMVIB has strong competitiveness in clustering while maintaining biological interpretability. Zhen Tian 0004, Xiaojiao Wei, Zhengzheng Lou, Zhixia Teng, Shouli Fu |
Briefings Bioinform. | 1 |
| 2026 | MFCLDTA: Multi-scale feature contrastive learning for predicting drug-target binding affinity
Zhen Tian 0004, Saisai Zhu, Zhixia Teng, Tao Wang 0082 |
Expert Syst. Appl. | 1 |
| 2026 | A Stitch in Time Saves Nine: Progressive Information Bottleneck for Incremental Multiview ClusteringabstractIncremental multiview clustering (IMVC) leverages consistent information between historical and new views to benefit the clustering task. However, existing IMVC approaches ignore the redundant information in individual views, leading to an accumulation of irrelevance. Besides, with the continuous arrivals of new views, the knowledge learned from historical views is often forgotten, which hinders the learning models from achieving long-term dependencies across incremental views. In this study, we propose a novel progressive information bottleneck (PIB), which is capable of removing redundant information in a timely manner and selectively updating historical knowledge based on information gain of new views. Specifically, to facilitate the knowledge transfer from historical views to incoming one, an information-aware knowledge library is built to store the representative samples of historical views. With the emergence of new views, we first devise a matrix-based mutual information (MI) constraint on an encoder to compress redundant information, which facilitates the training of a neural network with analyzable gradients and obtain a compact yet discriminative representation. Then, a dual-selective updating strategy is proposed to preserve historical knowledge in time when it contributes more to the information gain of the knowledge library than the new view. Finally, relevant samples to the new view in knowledge library are migrated to maximize the cross-level consistency between historical and new views. To the best of our knowledge, this is the first work that designs a gradient-analyzable MI measurement for incremental multiview learning and employs information gain to guide the selective update of the knowledge library. Empirical evaluations on six benchmark datasets show that our method outperforms state-of-the-art baseline methods by an average of 8.1%, 7.6%, and 7.8% on clustering accuracy (ACC), normalized mutual information (NMI), and adjusted Rand index (ARI) metrics, respectively. Fengshou Han, Yiqiao Mao, Zhen Tian 0004, Witold Pedrycz, Hui Yu 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2026 | Multi-View Clustering With Cauchy-Schwarz Mutual Information MaximinabstractInformation bottleneck (IB) leverages information theory to guide the learning process of deep multi-view clustering (MVC). It optimizes the trade-off between multi-view compression and preservation by minimizing and maximizing mutual information (MI). Although existing deep MVC based on IB has witnessed great achievements, they usually resort to variational inference to estimate the MI lower bound, which typically introduces estimation errors and results in an unstable lower bound of MI. In this study, we propose a novel Cauchy-Schwarz Mutual Information Maximin (CS-MIM) method, which directly estimates MI with closed-form expressions without requiring variational inference, possessing explicit multi-view information modeling capabilities. Specifically, we first present a non-parametric MI estimation method with Cauchy-Schwarz (CS) divergence, which leverages multi-kernel Gram matrices to capture distributional similarities and avoids the approximation errors introduced by the neural estimators of variational inference. Then, based on the new estimation method, a MI maximin mechanism is devised to parameterize the IB principle with analytical gradients, which facilitates effective compression of multi-view data while preserving the relevant features. Finally, we design a cross-view adaptive attention (CAA) mechanism constrained by the MI based on CS divergence, which further captures the complementarity across views under the guidance of the fused multi-view representation. Extensive evaluation results on 12 public available datasets demonstrate that the CS-MIM remarkably outperforms existing SOTA approaches. Zhen Tian 0004, Enlai Ouyang, Quan Zou 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | Contrastive Variational Graph Symmetric Autoencoder for Full Extrapolation Over Temporal Knowledge GraphsabstractTemporal knowledge graph (TKG) reasoning is potent for discovering plausible facts with observable knowledge. However, evolving TKGs often bring newly unseen entities, relations, and timestamps, which pose significant challenges for current TKG reasoning methods as they struggle to extrapolate facts involving unseen components. We term this realistic yet under-investigated problem as full extrapolation over TKGs, aiming to model all unseen components for reasoning within evolving TKGs. To tackle this, we propose a novel contrastive variational graph symmetric autoencoder (CVGSAE) framework under the meta-learning setting. Initially, we first extract the spatiotemporal connections between relations and integrate a functional time encoding strategy to simultaneously initialize unseen entities, relations, and timestamps. Subsequently, we design an interactive variational graph symmetric autoencoder to encode and then reconstruct entity and relation representations by capturing their interactions to generate contrastive views. Finally, a dual-contrastive learning scheme is imposed on the encoded and reconstructed representations of both entities and relations to enhance their discriminability and task relevance. Extensive experiments on two benchmark datasets demonstrate that CVGSAE yields significant advancements beyond state-of-the-art baselines across all evaluation metrics. The source code is available at:https://github.com/hcfun/CVGSAE Haichuan Fang, Zhen Tian 0004, Yangdong Ye |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | EchoDiffusion: Waveform Conditioned Diffusion Models for Echo-Based Depth EstimationabstractTo extract spatial information, depth estimation using conventional echo-based methods typically employs models with encoder-decoder architectures, such as UNet. However, these methods may face challenges in extracting fine details from echo waveforms and handling multi-scale feature extraction with high precision. To address these challenges, we introduce EchoDiffusion, a framework that incorporates diffusion models conditioned on waveform embeddings for echo-based depth estimation. This framework employs the Multi-Scale Adaptive Latent Feature Network (MALF-Net) to extract multi-scale spatial features and perform adaptive fusion, encoding the echo spectrograms into the latent space. Additionally, we propose the Echo Waveform Detail Embedder (EWDE), which leverages a pre-trained Wav2Vec model to extract detailed spatial information from echo waveforms, using these details as conditional inputs to guide the reverse diffusion process in the latent space. By embedding the echo waveforms into the reverse diffusion process, we can more accurately guide the generation of depth maps. Our extensive evaluations on the Replica and Matterport3D datasets demonstrate that EchoDiffusion establishes new benchmarks for state-of-the-art performance in echo-based depth estimation. Wenjie Zhang 0008, Xiaoheng Jiang, Zhen Tian 0004, Mingliang Xu 0001 |
AAAI | 6 |
| 2025 | Drug-Target Interaction Prediction with Pretrained LLM Feature Learning and Cross-Modal Fusion StrategyabstractAccurate prediction of drug-target interaction (DTI) is crucial for drug discovery. In deep learning (DL)-based approaches, achieving robust and generalizable DTI prediction hinges on informative feature representations and effective modeling strategies. However, existing methods often depend on prior domain knowledge, which may limit their adaptability to unseen samples. Moreover, many approaches fail to exploit the complementary nature of multimodal biological data. To address these limitations, we propose MMF-DTI, a novel multimodal fusion framework that synergizes pretrained large language models (LLMs). These LLMs are employed to extract generalized representations, enhancing the model's capability to generalize to previously unseen molecules. To effectively integrate these modalities, we propose a cross-modal attention module that facilitates interaction between sequence-based and structural features. In comparison with other baselines, MMF-DTI demonstrates a consistent improvement. We further demonstrate the interpretability by aligning attention weights with biophysical interactions, validated on the 1M17 protein. Rongqi Fan, Zhiyuan Dong, Zhen Tian 0004, Pengyuan Li 0014, Yufeng Ling |
BIBM | 3 |
| 2025 | CDIB: Consistency Discovery-guided Information Bottleneck for Multi-modal Knowledge Graph ReasoningabstractMulti-modal knowledge graph reasoning (MKGR) seeks to conjecture plausible facts in MKGs by learning effective representations from various modalities (e.g., structure, text, and image). However, due to holistic redundancy (i.e., each modality carries task-irrelevant redundancy) and modality conflict (i.e., different modalities contain contradictory information), the reasoning performance of current methods is substantially impaired. In this paper, we propose a novel Consistency Discovery-guided Information Bottleneck (CDIB) framework to address the aforementioned challenges. Specifically, a modality compression module is first designed to learn modality-private entity representations of alleviating redundant information. Then, a consistency discovery module is developed to discover cross-modal consistency during multi-modal fusion to learn the comprehensive entity representations. To retain task-relevant information, an information preservation module is devised to further enrich the comprehensive entity representations to be predictive for MKGR. Extensive experiments indicate that CDIB achieves state-of-the-art reasoning ability on two benchmark datasets over current MKGR baselines, and also exhibits promising robustness against noise. Haichuan Fang, Yulin Du, Qiang Guo 0012, Zhen Tian 0004, Yangdong Ye |
ACM Multimedia | 5 |
| 2025 | HSSPPI: hierarchical and spatial-sequential modeling for PPIs predictionabstractMOTIVATION: Protein-protein interactions play a fundamental role in biological systems. Accurate detection of protein-protein interaction sites (PPIs) remains a challenge. And, the methods of PPIs prediction based on biological experiments are expensive. Recently, a lot of computation-based methods have been developed and made great progress. However, current computational methods only focus on one form of protein, using only protein spatial conformation or primary sequence. And, the protein's natural hierarchical structure is ignored. RESULTS: In this study, we propose a novel network architecture, HSSPPI, through hierarchical and spatial-sequential modeling of protein for PPIs prediction. In this network, we represent protein as a hierarchical graph, in which a node in the protein is a residue (residue-level graph) and a node in the residue is an atom (atom-level graph). Moreover, we design a spatial-sequential block for capturing complex interaction relationships from spatial and sequential forms of protein. We evaluate HSSPPI on public benchmark datasets and the predicting results outperform the comparative models. This indicates the effectiveness of hierarchical protein modeling and also illustrates that HSSPPI has a strong feature extraction ability by considering spatial and sequential information simultaneously. AVAILABILITY AND IMPLEMENTATION: The code of HSSPPI is available at https://github.com/biolushuai/Hierarchical-Spatial-Sequential-Modeling-of-Protein. Yuguang Li, Zhen Tian 0004, Xiaofei Nan, Shoutao Zhang, Qinglei Zhou |
Briefings Bioinform. | 2 |
| 2025 | Enhancing LncRNA-miRNA interaction prediction with multimodal contrastive representation learningabstractInteractions between long non-coding RNAs (lncRNAs) and microRNAs (miRNAs) play an important role in the development of complex human diseases by collaboratively regulating gene transcription and expression. Therefore, identifying lncRNA-miRNA interactions (LMIs) is essential for diagnosing and treating complex human diseases. Because identifying LMIs with wet experiments is time-consuming and labor-intensive, some computational methods have been developed to infer LMIs. However, these approaches excel at utilizing single-modal information but struggle to integrate multimodal data from lncRNAs and miRNAs, which is essential for uncovering complex patterns in LMIs, ultimately limiting their performance. Therefore, this article proposes a novel multimodal contrastive representation learning model (MCRLMI) for LMI predictions. The model fully integrates multi-source similarity information and sequence encodings of lncRNAs and miRNAs. It leverages a graph convolutional network (GCN) and a Transformer to capture local neighborhood structural features and long-distance dependencies, respectively, enabling the collaborative modeling of structural and semantic information. Subsequently, to effectively integrate multimodal characteristics with encoded information, a multichannel attention mechanism and contrastive learning are introduced to fuse the extracted features. Finally, a Kolmogorov-Arnold Network (KAN) is trained with the optimized embeddings to predict LMIs. Extensive experiments show that the proposed MCRLMI consistently outperforms existing methods. Moreover, case studies further validate the potential of MCRLMI to identify novel LMIs in practical applications. Zhixia Teng, Zhaowen Tian, Murong Zhou, Guohua Wang 0001, Zhen Tian 0004 |
Briefings Bioinform. | 5 |
| 2025 | MAEST: accurately spatial domain detection in spatial transcriptomics with graph masked autoencoderabstractSpatial transcriptomics (ST) technology provides gene expression profiles with spatial context, offering critical insights into cellular interactions and tissue architecture. A core task in ST is spatial domain identification, which involves detecting coherent regions with similar spatial expression patterns. However, existing methods often fail to fully exploit spatial information, leading to limited representational capacity and suboptimal clustering accuracy. Here, we introduce MAEST, a novel graph neural network model designed to address these limitations in ST data. MAEST leverages graph masked autoencoders to denoise and refine representations while incorporating graph contrastive learning to prevent feature collapse and enhance model robustness. By integrating one-hop and multi-hop representations, MAEST effectively captures both local and global spatial relationships, improving clustering precision. Extensive experiments across diverse datasets, including the human brain, mouse hippocampus, olfactory bulb, brain, and embryo, demonstrate that MAEST outperforms seven state-of-the-art methods in spatial domain identification. Furthermore, MAEST showcases its ability to integrate multi-slice data, identifying joint domains across horizontal tissue sections with high accuracy. These results highlight MAEST's versatility and effectiveness in unraveling the spatial organization of complex tissues. The source code of MAEST can be obtained at https://github.com/clearlove2333/MAEST. Han Shu, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Zhen Tian 0004, Tao Wang 0082 |
Briefings Bioinform. | 9 |
| 2025 | Differentiable graph clustering with structural grouping for single-cell RNA-seq dataabstractMOTIVATION: Clustering cells into subpopulations is one of the most crucial tasks in single-cell RNA sequencing (scRNA-seq) data analysis, which provides support for biological research at cellular level. With the development of graph neural networks, deep graph clustering approaches have achieved excellent performance by modeling the topological relationships between cells. However, existing approaches rely on cell node and its neighbors to obtain the cell feature representation, which ignore the graph cluster structure hidden in scRNA-seq data. Besides, how to bridge the heterogeneous gap between cell node feature and its structural information remains a highly challenging problem. RESULTS: Here, we propose a novel differentiable graph clustering with structural grouping (DGCSG) for scRNA-seq data, which incorporates graph cluster information into deep graph clustering model by designing a differentiable clustering mechanism to learn clustering-friendly representation. Firstly, an interactive module is devised to dynamically transfer node representations learned by autoencoder (AE) to graph attention autoencoder (GATE) in layer-by-layer manner. Then, to characterize graph cluster information, a differentiable clustering mechanism is proposed to transform K-way normalized cuts from a discrete optimization problem into differentiable learning objective through spectral relaxation, which jointly optimizes the GATE by allocating more attention scores to nodes in the same graph cluster. Finally, a decoupled self-supervised optimization is proposed, which guides the representation learning of AE and GATE in the interactive module. Extensive evaluations on 14 scRNA-seq benchmarks verify the superiority of DGCSG compared with state-of-the-art baselines. AVAILABILITY AND IMPLEMENTATION: The code associated with this work is available on GitHub (https://github.com/Xiaoqiang-Yan/DGCSG). Shike Du, Quan Zou 0001, Zhen Tian 0004 |
Bioinform. | 4 |
| 2025 | Hierarchical feature aggregation with mixed attention mechanism for single-cell RNA-seq analysis
Wanning Zhou, Zhixia Teng, Zhen Tian 0004 |
Expert Syst. Appl. | 6 |
| 2025 | Drug repositioning by collaborative learning based on graph convolutional inductive network
Zhixia Teng, Yongliang Li, Zhen Tian 0004, Yingjian Liang, Guohua Wang 0001 |
Future Gener. Comput. Syst. | 3 |
| 2025 | Few-Shot Object Counting with frequency attention and multi-perception head
Gaoxin Ma, Xingquan Zhu 0001, Zhen Tian 0004, Yangdong Ye, Zhenfeng Zhu |
Neurocomputing | 3 |
| 2025 | Decoupled contrastive multi-view clustering with adaptive false negative elimination for cancer subtypingabstractCancer's heterogeneity necessitates precise subtype identification for effective diagnosis and treatment, which can be achieved by integrating multi-omics data to reveal distinct molecular characteristics and enable personalized therapies. Recently, significant efforts have been made through contrastive clustering methods to efficiently identify cancer subtypes. However, existing approaches remain limited in effectively capturing inter- and intra-view relationships in multi-omics data. Additionally, most cancer subtyping methods often rely on random sampling to construct negative pairs, which may inadvertently engender false negatives. To overcome these challenges, we propose a novel end-to-end self-supervised learning model named Decoupled Contrastive Multi-view Clustering with adaptive false negative elimination (DCMC). Specifically, DCMC adopts a multi-view clustering architecture that facilitates intra- and inter-view contrastive learning across distinct embedding spaces, allowing view-specific information to be preserved while maintaining cross-view consistency. We further introduce an adaptive false negative elimination framework to progressively screen potential false negatives. Finally, pseudo-label rectification is applied to enhance the quality of the learned representations and further refine the clustering process. DCMC is evaluated on 10 commonly used cancer datasets against 19 state-of-the-art methods, with experimental results validating its superior performance. In the Liver Hepatocellular Carcinoma case study, differential expression analysis is performed to identify potential biomarkers, while the cancer subtypes identified by DCMC are validated for their responses to specific therapeutic drugs. The datasets and source code for DCMC are available online at https://github.com/LinMengX/DCMC. Mengxiang Lin, Rongqi Fan, Saisai Zhu, Quan Zou 0001, Zhen Tian 0004 |
PLoS Comput. Biol. | 6 |
| 2025 | DSANIB: Drug-Target Interaction Predictions With Dual-View Synergistic Attention Network and Information Bottleneck StrategyabstractPrediction of drug-target interactions (DTIs) is one of the crucial steps for drug repositioning. Identifying DTIs through bio-experimental manners is always expensive and time-consuming. Recently, deep learning-based approaches have shown promising advancements in DTI prediction, but they face two notable challenges: (i) how to explicitly capture local interactions between drug-target pairs and learn their higher-order substructure embeddings; (ii) How to filter out redundant information to obtain effective embeddings for drugs and targets. Results: In this study, we propose a novel approach, termed DSANIB, to infer potential interactions between drugs and targets. DSANIB comprises two primary components: (1) DSAN component: The Inter-view Attention Network Module explicitly learns the local interactions between drugs and targets, while the Intra-view Attention Network Module aggregates information from local interaction features to obtain their higher-order substructure embeddings. (2) Information Bottleneck (IB) component: DSANIB adopts the IB strategy, which could retain relevant information while minimizing the redundant features to obtain their discriminative representations. Extensive experimental results demonstrate that DSANIB outperforms other SOTA prediction models. In addition, visualization of drug and target embeddings learned through DSANIB could provide interpretable insights for the prediction results. Zhen Tian 0004, Wanning Zhou, Zhixia Teng, Quan Zou 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Clustering scRNA-seq data with the cross-view collaborative information fusion strategyabstractSingle-cell RNA sequencing (scRNA-seq) technology has revolutionized biological research by enabling high-throughput, cellular-resolution gene expression profiling. A critical step in scRNA-seq data analysis is cell clustering, which supports downstream analyses. However, the high-dimensional and sparse nature of scRNA-seq data poses significant challenges to existing clustering methods. Furthermore, integrating gene expression information with potential cell structure data remains largely unexplored. Here, we present scCFIB, a novel information bottleneck (IB)-based clustering algorithm that leverages the power of IB for efficient processing of high-dimensional sparse data and incorporates a cross-view fusion strategy to achieve robust cell clustering. scCFIB constructs a multi-feature space by establishing two distinct views from the original features. We then formulate the cell clustering problem as a target loss function within the IB framework, employing a collaborative information fusion strategy. To further optimize scCFIB's performance, we introduce a novel sequential optimization approach through an iterative process. Benchmarking against established methods on diverse scRNA-seq datasets demonstrates that scCFIB achieves superior performance in scRNA-seq data clustering tasks. Availability: the source code is publicly available on GitHub: https://github.com/weixiaojiao/scCFIB. Zhengzheng Lou, Xiaojiao Wei, Yuanhao Hu, Shizhe Hu, Yucong Wu, Zhen Tian 0004 |
Briefings Bioinform. | 6 |
| 2024 | MGCNSS: miRNA-disease association prediction with multi-layer graph convolution and distance-based negative sample selection strategyabstractIdentifying disease-associated microRNAs (miRNAs) could help understand the deep mechanism of diseases, which promotes the development of new medicine. Recently, network-based approaches have been widely proposed for inferring the potential associations between miRNAs and diseases. However, these approaches ignore the importance of different relations in meta-paths when learning the embeddings of miRNAs and diseases. Besides, they pay little attention to screening out reliable negative samples which is crucial for improving the prediction accuracy. In this study, we propose a novel approach named MGCNSS with the multi-layer graph convolution and high-quality negative sample selection strategy. Specifically, MGCNSS first constructs a comprehensive heterogeneous network by integrating miRNA and disease similarity networks coupled with their known association relationships. Then, we employ the multi-layer graph convolution to automatically capture the meta-path relations with different lengths in the heterogeneous network and learn the discriminative representations of miRNAs and diseases. After that, MGCNSS establishes a highly reliable negative sample set from the unlabeled sample set with the negative distance-based sample selection strategy. Finally, we train MGCNSS under an unsupervised learning manner and predict the potential associations between miRNAs and diseases. The experimental results fully demonstrate that MGCNSS outperforms all baseline methods on both balanced and imbalanced datasets. More importantly, we conduct case studies on colon neoplasms and esophageal neoplasms, further confirming the ability of MGCNSS to detect potential candidate miRNAs. The source code is publicly available on GitHub https://github.com/15136943622/MGCNSS/tree/master. Zhen Tian 0004, Chenguang Han, Lewen Xu, Zhixia Teng |
Briefings Bioinform. | 1 |
| 2024 | Drug-target interaction predictions with multi-view similarity network fusion strategy and deep interactive attention mechanismabstractMOTIVATION: Accurately identifying the drug-target interactions (DTIs) is one of the crucial steps in the drug discovery and drug repositioning process. Currently, many computational-based models have already been proposed for DTI prediction and achieved some significant improvement. However, these approaches pay little attention to fuse the multi-view similarity networks related to drugs and targets in an appropriate way. Besides, how to fully incorporate the known interaction relationships to accurately represent drugs and targets is not well investigated. Therefore, there is still a need to improve the accuracy of DTI prediction models. RESULTS: In this study, we propose a novel approach that employs Multi-view similarity network fusion strategy and deep Interactive attention mechanism to predict Drug-Target Interactions (MIDTI). First, MIDTI constructs multi-view similarity networks of drugs and targets with their diverse information and integrates these similarity networks effectively in an unsupervised manner. Then, MIDTI obtains the embeddings of drugs and targets from multi-type networks simultaneously. After that, MIDTI adopts the deep interactive attention mechanism to further learn their discriminative embeddings comprehensively with the known DTI relationships. Finally, we feed the learned representations of drugs and targets to the multilayer perceptron model and predict the underlying interactions. Extensive results indicate that MIDTI significantly outperforms other baseline methods on the DTI prediction task. The results of the ablation experiments also confirm the effectiveness of the attention mechanism in the multi-view similarity network fusion strategy and the deep interactive attention mechanism. AVAILABILITY AND IMPLEMENTATION: https://github.com/XuLew/MIDTI. Lewen Xu, Chenguang Han, Zhen Tian 0004, Quan Zou 0001 |
Bioinform. | 4 |
| 2023 | Predicting microbe-drug associations with structure-enhanced contrastive learning and self-paced negative sampling strategyabstractMOTIVATION: Predicting the associations between human microbes and drugs (MDAs) is one critical step in drug development and precision medicine areas. Since discovering these associations through wet experiments is time-consuming and labor-intensive, computational methods have already been an effective way to tackle this problem. Recently, graph contrastive learning (GCL) approaches have shown great advantages in learning the embeddings of nodes from heterogeneous biological graphs (HBGs). However, most GCL-based approaches don't fully capture the rich structure information in HBGs. Besides, fewer MDA prediction methods could screen out the most informative negative samples for effectively training the classifier. Therefore, it still needs to improve the accuracy of MDA predictions. RESULTS: In this study, we propose a novel approach that employs the Structure-enhanced Contrastive learning and Self-paced negative sampling strategy for Microbe-Drug Association predictions (SCSMDA). Firstly, SCSMDA constructs the similarity networks of microbes and drugs, as well as their different meta-path-induced networks. Then SCSMDA employs the representations of microbes and drugs learned from meta-path-induced networks to enhance their embeddings learned from the similarity networks by the contrastive learning strategy. After that, we adopt the self-paced negative sampling strategy to select the most informative negative samples to train the MLP classifier. Lastly, SCSMDA predicts the potential microbe-drug associations with the trained MLP classifier. The embeddings of microbes and drugs learning from the similarity networks are enhanced with the contrastive learning strategy, which could obtain their discriminative representations. Extensive results on three public datasets indicate that SCSMDA significantly outperforms other baseline methods on the MDA prediction task. Case studies for two common drugs could further demonstrate the effectiveness of SCSMDA in finding novel MDA associations. AVAILABILITY: The source code is publicly available on GitHub https://github.com/Yue-Yuu/SCSMDA-master. Zhen Tian 0004, Haichuan Fang, Weixin Xie, Maozu Guo 0001 |
Briefings Bioinform. | 1 |
| 2023 | Learning knowledge graph embedding with a dual-attention embedding network
Haichuan Fang, Zhen Tian 0004, Yangdong Ye |
Expert Syst. Appl. | 3 |
| 2023 | GOGCN: Graph Convolutional Network on Gene Ontology for Functional Similarity Analysis of GenesabstractThe measurement of gene functional similarity plays a critical role in numerous biological applications, such as gene clustering, the construction of gene similarity networks. However, most existing approaches still rely heavily on traditional computational strategies, which are not guaranteed to achieve satisfactory performance. In this study, we propose a novel computational approach called GOGCN to measure gene functional similarity by modeling the Gene Ontology (GO) through Graph Convolutional Network (GCN). GOGCN is a graph-based approach that performs sufficient representation learning for terms and relations in the GO graph. First, GOGCN employs the GCN-based knowledge graph embedding (KGE) model to learn vector representations (i.e., embeddings) for all entities (i.e., terms). Second, GOGCN calculates the semantic similarity between two terms based on their corresponding vector representations. Finally, GOGCN estimates gene functional similarity by making use of the pair-wise strategy. During the representation learning period, GOGCN promotes semantic interaction between terms through GCN, thereby capturing the rich structural information of the GO graph. Further experimental results on various datasets suggest that GOGCN is superior to the other state-of-the-art approaches, which shows its reliability and effectiveness. Zhen Tian 0004, Haichuan Fang, Zhixia Teng, Yangdong Ye |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | Predicting miRNA-disease associations via learning multimodal networks and fusing mixed neighborhood informationabstractMOTIVATION: In recent years, a large number of biological experiments have strongly shown that miRNAs play an important role in understanding disease pathogenesis. The discovery of miRNA-disease associations is beneficial for disease diagnosis and treatment. Since inferring these associations through biological experiments is time-consuming and expensive, researchers have sought to identify the associations utilizing computational approaches. Graph Convolutional Networks (GCNs), which exhibit excellent performance in link prediction problems, have been successfully used in miRNA-disease association prediction. However, GCNs only consider 1st-order neighborhood information at one layer but fail to capture information from high-order neighbors to learn miRNA and disease representations through information propagation. Therefore, how to aggregate information from high-order neighborhood effectively in an explicit way is still challenging. RESULTS: To address such a challenge, we propose a novel method called mixed neighborhood information for miRNA-disease association (MINIMDA), which could fuse mixed high-order neighborhood information of miRNAs and diseases in multimodal networks. First, MINIMDA constructs the integrated miRNA similarity network and integrated disease similarity network respectively with their multisource information. Then, the embedding representations of miRNAs and diseases are obtained by fusing mixed high-order neighborhood information from multimodal network which are the integrated miRNA similarity network, integrated disease similarity network and the miRNA-disease association networks. Finally, we concentrate the multimodal embedding representations of miRNAs and diseases and feed them into the multilayer perceptron (MLP) to predict their underlying associations. Extensive experimental results show that MINIMDA is superior to other state-of-the-art methods overall. Moreover, the outstanding performance on case studies for esophageal cancer, colon tumor and lung cancer further demonstrates the effectiveness of MINIMDA. AVAILABILITY AND IMPLEMENTATION: https://github.com/chengxu123/MINIMDA and http://120.79.173.96/. Zhengzheng Lou, Zhaoxu Cheng, Zhixia Teng, Zhen Tian 0004 |
Briefings Bioinform. | 6 |
| 2022 | MHADTI: predicting drug-target interactions via multiview heterogeneous information network embedding with hierarchical attention mechanismsabstractMOTIVATION: Discovering the drug-target interactions (DTIs) is a crucial step in drug development such as the identification of drug side effects and drug repositioning. Since identifying DTIs by web-biological experiments is time-consuming and costly, many computational-based approaches have been proposed and have become an efficient manner to infer the potential interactions. Although extensive effort is invested to solve this task, the prediction accuracy still needs to be improved. More especially, heterogeneous network-based approaches do not fully consider the complex structure and rich semantic information in these heterogeneous networks. Therefore, it is still a challenge to predict DTIs efficiently. RESULTS: In this study, we develop a novel method via Multiview heterogeneous information network embedding with Hierarchical Attention mechanisms to discover potential Drug-Target Interactions (MHADTI). Firstly, MHADTI constructs different similarity networks for drugs and targets by utilizing their multisource information. Combined with the known DTI network, three drug-target heterogeneous information networks (HINs) with different views are established. Secondly, MHADTI learns embeddings of drugs and targets from multiview HINs with hierarchical attention mechanisms, which include the node-level, semantic-level and graph-level attentions. Lastly, MHADTI employs the multilayer perceptron to predict DTIs with the learned deep feature representations. The hierarchical attention mechanisms could fully consider the importance of nodes, meta-paths and graphs in learning the feature representations of drugs and targets, which makes their embeddings more comprehensively. Extensive experimental results demonstrate that MHADTI performs better than other SOTA prediction models. Moreover, analysis of prediction results for some interested drugs and targets further indicates that MHADTI has advantages in discovering DTIs. AVAILABILITY AND IMPLEMENTATION: https://github.com/pxystudy/MHADTI. Zhen Tian 0004, Haichuan Fang, Wenjie Zhang 0008, Qiguo Dai, Yangdong Ye |
Briefings Bioinform. | 1 |
| 2022 | A novel gene functional similarity calculation model by utilizing the specificity of terms and relationships in gene ontologyabstractBACKGROUND: Recently, with the foundation and development of gene ontology (GO) resources, numerous works have been proposed to compute functional similarity of genes and achieved series of successes in some research fields. Focusing on the calculation of the information content (IC) of terms is the main idea of these methods, which is essential for measuring functional similarity of genes. However, most approaches have some deficiencies, especially when measuring the IC of both GO terms and their corresponding annotated term sets. To this end, measuring functional similarity of genes accurately is still challenging. RESULTS: In this article, we proposed a novel gene functional similarity calculation method, which especially encapsulates the specificity of terms and edges (STE). The proposed method mainly contains three steps. Firstly, a novel computing model is put forward to compute the IC of terms. This model has the ability to exploit the specific structural information of GO terms. Secondly, the IC of term sets are computed by capturing the genetic structure between the terms contained in the set. Lastly, we measure the gene functional similarity according to the IC overlap ratio of the corresponding annotated genes sets. The proposed method accurately measures the IC of not only GO terms but also the annotated term sets by leveraging the specificity of edges in the GO graph. CONCLUSIONS: We conduct experiments on gene functional classification in biological pathways, gene expression datasets, and protein-protein interaction datasets. Extensive experimental results show the better performances of our proposed STE against several baseline methods. Zhen Tian 0004, Haichuan Fang, Yangdong Ye, Zhenfeng Zhu |
BMC Bioinform. | 1 |
| 2021 | ReRF-Pred: predicting amyloidogenic regions of proteins based on their pseudo amino acid composition and tripeptide compositionabstractBACKGROUND: Amyloids are insoluble fibrillar aggregates that are highly associated with complex human diseases, such as Alzheimer's disease, Parkinson's disease, and type II diabetes. Recently, many studies reported that some specific regions of amino acid sequences may be responsible for the amyloidosis of proteins. It has become very important for elucidating the mechanism of amyloids that identifying the amyloidogenic regions. Accordingly, several computational methods have been put forward to discover amyloidogenic regions. The majority of these methods predicted amyloidogenic regions based on the physicochemical properties of amino acids. In fact, position, order, and correlation of amino acids may also influence the amyloidosis of proteins, which should be also considered in detecting amyloidogenic regions. RESULTS: To address this problem, we proposed a novel machine-learning approach for predicting amyloidogenic regions, called ReRF-Pred. Firstly, the pseudo amino acid composition (PseAAC) was exploited to characterize physicochemical properties and correlation of amino acids. Secondly, tripeptides composition (TPC) was employed to represent the order and position of amino acids. To improve the distinguishability of TPC, all possible tripeptides were analyzed by the binomial distribution method, and only those which have significantly different distribution between positive and negative samples remained. Finally, all samples were characterized by PseAAC and TPC of their amino acid sequence, and a random forest-based amyloidogenic regions predictor was trained on these samples. It was proved by validation experiments that the feature set consisted of PseAAC and TPC is the most distinguishable one for detecting amyloidosis. Meanwhile, random forest is superior to other concerned classifiers on almost all metrics. To validate the effectiveness of our model, ReRF-Pred is compared with a series of gold-standard methods on two datasets: Pep-251 and Reg33. The results suggested our method has the best overall performance and makes significant improvements in discovering amyloidogenic regions. CONCLUSIONS: The advantages of our method are mainly attributed to that PseAAC and TPC can describe the differences between amyloids and other proteins successfully. The ReRF-Pred server can be accessed at http://106.12.83.135:8080/ReRF-Pred/. Zhixia Teng, Zitong Zhang 0002, Zhen Tian 0004, Yanjuan Li, Guohua Wang 0001 |
BMC Bioinform. | 3 |
| 2020 | A Stacked Ensemble Learning Framework with Heterogeneous Feature Combinations for Predicting ncRNA-Protein InteractionabstractThe interaction between ncRNA and protein is a kind of crucial molecular activities in a cell. Developing computational methods to predict ncRNA-protein interactions has attracted increasing attentions in recent years. In this work, a novel stacked ensemble learning framework is presented for predicting ncRNA-protein interaction based on heterogeneous feature combinations, named HFC-RPI. Firstly, the compositional features of k-mer with different orders were extracted from the primary sequence and secondary structure of RNA and protein respectively. Secondly, we trained a set of base learners using a variety of heterogeneous combinations of the extracted features respectively. Thirdly, the prediction results of these base learners were employed to train the stacked learner, which output the final prediction result at the higher layer in HFC-RPI. Moreover, in order to improve the generalization of HFC-RPI, when training the base learners, a cross-validation based method was applied. Extensive experimental results showed that the proposed learning framework HFC-RPI was effective and feasible for predicting the interaction of ncRNA and protein. By comparing with state-of-the-art methods, HFC-RPI was superior to them on most performance evaluation metrics. Qiguo Dai, Zhaowei Wang 0005, Jinmiao Song, Xiaodong Duan, Maozu Guo 0001, Zhen Tian 0004 |
BIBM | 6 |
| 2020 | SWE: a novel method with semantic-weighted edge for measuring gene functional similarityabstractIn recent years, functional similarity has played an independent role in some biological fields such as gene clustering, gene functional prediction, and evaluation for protein-protein interaction. In this premise, some effective methods have already been proposed based on Gene Ontology (GO). Although these mainstream methods achieve the purpose for measuring gene functional similarity, they may have some deficiency when calculating the Information Content (IC) of GO terms. Consequently, measuring the functional similarity accurately is still a meaningful objective of research. In this paper, a novel method called SWE, is proposed for measuring gene functional similarity based on the GO graph. Firstly, an algorithm to measure terms' semantics based on their information in the GO graph is put forward. The information of GO terms mainly contains their depth, ancestors and descendants. Secondly, we calculate the IC of a term set by means of retrieving the inherited relationship between terms in a term set. Finally, the functional similarity between two genes is computed based on the IC overlap ratio of term sets annotating two genes respectively. Results demonstrate that SWE is superior to existing methods in some experiments such as functional classification of genes in a biological pathway, protein-protein interaction and gene expression experiment. Further analysis demonstrates that SWE takes not only the specificity of terms into account, but their information in the GO graph, both of which are shown to be consistent with human perspectives. Zhen Tian 0004, Haichuan Fang, Yangdong Ye, Zhenfeng Zhu |
BIBM | 1 |
| 2017 | Refine gene functional similarity network based on interaction networksabstractBACKGROUND: In recent years, biological interaction networks have become the basis of some essential study and achieved success in many applications. Some typical networks such as protein-protein interaction networks have already been investigated systematically. However, little work has been available for the construction of gene functional similarity networks so far. In this research, we will try to build a high reliable gene functional similarity network to promote its further application. RESULTS: Here, we propose a novel method to construct and refine the gene functional similarity network. It mainly contains three steps. First, we establish an integrated gene functional similarity networks based on different functional similarity calculation methods. Then, we construct a referenced gene-gene association network based on the protein-protein interaction networks. At last, we refine the spurious edges in the integrated gene functional similarity network with the help of the referenced gene-gene association network. Experiment results indicate that the refined gene functional similarity network (RGFSN) exhibits a scale-free, small world and modular architecture, with its degrees fit best to power law distribution. In addition, we conduct protein complex prediction experiment for human based on RGFSN and achieve an outstanding result, which implies it has high reliability and wide application significance. CONCLUSIONS: Our efforts are insightful for constructing and refining gene functional similarity networks, which can be applied to build other high quality biological networks. Zhen Tian 0004, Maozu Guo 0001, Chunyu Wang 0002 |
BMC Bioinform. | 1 |
| 2016 | Revealing protein functions based on relationships of interacting proteins and GO termsabstractnumerous computational methods predicted protein function based on the protein-protein interaction (PPI) network. These methods supposed that two proteins share the same function if they interact with each other. However, it is reported by recent studies that the functions of two interacting proteins may be related but different. In this paper, the functional relationship between interacting proteins is studied and a novel method, called as GoDIN, is advanced to annotate functions of the interacting protein in Gene Ontology (GO) context. It is assumed that the functional difference between interacting proteins can be expressed by semantic difference between GO term and its relatives. Thus, the method uses GO term and its relatives to annotate the interacting proteins separately according to their functional roles in the PPI network. The method is validated by a series of experiments and compared with the related method. The experimental results confirm the assumption and suggest that GoDIN is effective on predicting functions of protein. Zhixia Teng, Maozu Guo 0001, Zhen Tian 0004, Kai Che |
BIBM | 4 |
| 2016 | Constructing an integrated gene similarity network for the identification of disease genesabstractDiscovering novel genes that are involved in human diseases is a challenging task. In recent years, several computational approaches have been proposed to prioritize candidate disease genes. Most of these methods are mainly based on protein-protein interaction (PPI) networks. However, since these PPI networks contain false positives and only cover less half of known human genes, their reliability and coverage are both very low. Therefore, it is highly necessary to fuse multiple genomic data to construct a reliable gene similarity network and then infer disease genes on the whole genomic scale. Here, we proposed a novel method, named RWRB, to infer causal genes of interested disease. First, we construct five individual gene (protein) similarity networks based on multiple genomic data of human genes. Then, an integrated gene similarity network (IGSN) is reconstructed based on similarity network fusion (SNF) method. Finally, we employ the random walk with restart algorithm on the phenotype-gene bilayer network, which combines phenotype similarity network, IGSN as well as the phenotype-gene association network, to prioritize candidate disease genes. We investigate the effectiveness of RWRB through leave-one-out cross-validation methods in inferring phenotype-gene relationships. Results show that RWRB is more accurate than state-of-the-art methods on most evaluation metrics. Further analysis shows that the success of RWRB is benefited from IGSN which has a wider coverage and higher reliability comparing with current PPI networks. Zhen Tian 0004, Maozu Guo 0001, Chunyu Wang 0002, Linlin Xing, Lei Wang 0085, Yin Zhang 0009 |
BIBM | 1 |
| 2016 | SGFSC: speeding the gene functional similarity calculation based on hash tablesabstractBACKGROUND: In recent years, many measures of gene functional similarity have been proposed and widely used in all kinds of essential research. These methods are mainly divided into two categories: pairwise approaches and group-wise approaches. However, a common problem with these methods is their time consumption, especially when measuring the gene functional similarities of a large number of gene pairs. The problem of computational efficiency for pairwise approaches is even more prominent because they are dependent on the combination of semantic similarity. Therefore, the efficient measurement of gene functional similarity remains a challenging problem. RESULTS: To speed current gene functional similarity calculation methods, a novel two-step computing strategy is proposed: (1) establish a hash table for each method to store essential information obtained from the Gene Ontology (GO) graph and (2) measure gene functional similarity based on the corresponding hash table. There is no need to traverse the GO graph repeatedly for each method with the help of the hash table. The analysis of time complexity shows that the computational efficiency of these methods is significantly improved. We also implement a novel Speeding Gene Functional Similarity Calculation tool, namely SGFSC, which is bundled with seven typical measures using our proposed strategy. Further experiments show the great advantage of SGFSC in measuring gene functional similarity on the whole genomic scale. CONCLUSIONS: The proposed strategy is successful in speeding current gene functional similarity calculation methods. SGFSC is an efficient tool that is freely available at http://nclab.hit.edu.cn/SGFSC . The source code of SGFSC can be downloaded from http://pan.baidu.com/s/1dFFmvpZ . Zhen Tian 0004, Chunyu Wang 0002, Maozu Guo 0001, Zhixia Teng |
BMC Bioinform. | 1 |