Haitao Fu

dblp:97/137 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 12 since 2021Systems, architecture and hardware · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Analog Mixed-Signal Circuit Splitting and Defect Simulation Method Based on Signal Flow Diagram
abstract
Comprehensive defect simulation in the integrated circuit design stage is the basic guarantee to improve testability, which is also the basic requirement of current automotive-grade chip design. For example, there are 6 types of basic defect models for a single MOS transistor in a circuit, and the defect set of analog/hybrid integrated circuits containing thousands of devices is extremely large. It is very time-consuming to inject a large number of defects one by one and solve them in a large-scale matrix. In fact, the complete defect simulation often takes months or even years, and there is no good solution at present. In this paper, a circuit splitting method based on signal flow and graph analysis is proposed, which models the control relationship between the device current and voltage nodes of the circuit with a directed dependency graph through signal flow analysis. Then, more weak connections are obtained by identifying and breaking the feedback loops of larger strongly connected components, and the weak connectivity in strongly connected components are split to realize circuit splitting. The acceleration can then be achieved by parallel defect simulation. At the same time, in order to investigate the observability of the internal defects in the previous strongly connected components in the subsequent strongly connected components, the defects were simulated according to the one-way transfer relationship between each strongly connected components in directed acyclic graph. The equivalent compression method of the defects and the signal in the transfer process is proposed, which reduces the number of simulation times of the defects. Finally, this paper takes the benchmark circuit Bandgap, LDO and SARADC circuits as examples, and uses HSPICE simulation to verify the effectiveness of the graph splitting method. Compared with the latest transistor level circuit splitting scheme, the circuit splitting algorithm proposed in this paper can split more and smaller circuit blocks. Using the circuit blocks we have divided to implement the parallel defect simulation scheme and the defect transfer simulation scheme is more effective than the traditional full-circuit and full-defect simulation one by one.
Hanyu Wen, Xinhong Huang, Qihao Zhang, Haitao Fu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 DMFF-EPI: A Dual-Modality Feature Fusion Network with Contrastive Learning for Enhancer-Promoter Interaction Prediction
abstract
Enhancer-Promoter Interactions (EPIs) play a pivotal role in transcriptional gene regulation, and their precise identification is critical for understanding disease mechanisms. Despite significant advances in computational models, achieving robust cross-cell-line generalization and comprehensive feature representation remains an open challenge. To tackle this, we propose DMFF-EPI, a dual-modality feature fusion network enhanced with contrastive learning for enhancer-promoter interaction prediction. Specifically, DMFF-EPI constructs parallel feature extraction modules for enhancer and promoter sequences. One branch combines a convolutional neural network (CNN) with a bidirectional gated recurrent unit (BiGRU) to capture both local motifs and long-range sequential dependencies. The parallel branch employs a multilayer perceptron (MLP) to encode k-mer frequency embeddings as prior structural knowledge. The core innovation lies in a contrastive learning-based regularizer that optimizes the feature space by aligning two complementary representations derived from the same sample-dynamic sequence features and static frequency features. This alignment enforces intra-sample consistency and improves model robustness across cell lines. A joint loss function combining binary cross-entropy for EPI classification and contrastive loss for feature alignment enables more stable training and stronger generalization. We conduct extensive performance evaluations on six publicly available cell lines (GM12878, K562, HeLa-S3, IMR90, HUVEC, and NHEK). The results suggest that DMFF-EPI achieves superior performance compared to existing baseline models in terms of both AUROC and AUPR. Cross-cell-line prediction experiments further evaluate the model's generalization to unseen cellular contexts. The results demonstrate that DMFF-EPI maintains competitive performance even under challenging cross-cell-line scenarios, highlighting its robustness and broad applicability across diverse biological contexts. The source code is publicly available at https://github.com/Wujiahui95/DMFF-EPI.
Amu Erermeri, Zhaoxiang Liu, Haitao Fu
BIBM4
2025 DMCRL: Deep Multi-View Contrastive Representation Learning for Drug-Target Affinity Prediction
abstract
Accurate prediction of drug-target affinity (DTA) is central to virtual screening and lead optimization. However, most existing approaches rely on a single data modality (e.g., molecular graphs or protein sequences), which limits their generalization to new chemical scaffolds and uncharacterized protein families. We present DMCRL, a multi-view contrastive representation learning framework that jointly mitigates view-specific biases and integrates three complementary modalities: (i) a molecular-graph view capturing fine-grained chemical structures, (ii) a protein-sequence view representing target semantics, and (iii) a global drug-target affinity graph encoding relational topology. A contrastive learning objective is introduced to maximize agreement among heterogeneous representations of the same drug-target pair, yielding embeddings that are both view-consistent and view-informative. These unified embeddings are fed into a lightweight MLP predictor for affinity regression (DTA) or interaction classification (DTI) without additional feature engineering. DMCRL achieves state-of-the-art or comparable performance across six benchmark datasets, excelling in regression tasks (MSE, CI,$r_{m}^{2}$) and classification tasks (Precision, Recall, AUC). Ablation and cold-start analyses further confirm the effectiveness of the affinity-graph encoder, contrastive alignment, and the model's strong generalization to unseen drugs and proteins. Overall, aligning molecular, sequence, and relational views through contrastive learning provides an effective paradigm for robust and generalizable DTA prediction. The source code is available at https://github.com/gmjjjj1127/DMCRL.
Haitao Fu, Maojun Gan, Biyong Deng, Shushan Hu
BIBM1
2025 MOIRL: A Multi-Objective Inverse Reinforcement Learning-Driven Model for Therapeutic Peptide Generation
abstract
Peptide drug development faces persistent challenges due to the tension between rational design efficiency and the vast sequence diversity of natural peptides. Traditional rulesbased methods often falter in navigating the high-dimensional combinatorial space of amino acid sequences, limiting their generalizability across diverse functional constraints. To address these limitations, we propose MOIRL, a multi-objective inverse reinforcement learning-driven model for therapeutic peptide generation. This framework synergistically integrates inverse reinforcement learning with a LSTM-Transformer hybrid architecture and a multi-objective reward mechanism, enabling it to jointly model the biological relevance and physicochemical constraints embedded within peptide generation tasks. Empirical evaluations conducted on benchmark datasets for antimicrobial peptide design reveal that MOIRL consistently outperforms state-of-the-art baselines across multiple evaluation dimensions, including structural diversity, predicted bioactivity, and sequence stability. These results indicate that the proposed framework not only improves the fidelity of functional peptide generation but also marks a meaningful advance toward scalable, data-driven approaches for complex molecular design in computational biology. The source code is publicly available at https://github.com/Fingertips-li/MOIRL
Xiaotong Shi, Maojun Gan, Haitao Fu, Shushan Hu
BIBM4
2025 MLPET: a Multi-Level Prompt-Enhanced Transformer for Unified Molecular Property and Drug-Drug Interaction Event Prediction
abstract
Accurate prediction of molecular properties and drug-drug interaction (DDI) events is crucial for drug discovery and pharmacological safety assessment. However, existing approaches typically focus on single-level molecular representation and lack generalizable strategies for multi-task knowledge transfer. In this work, we propose a unified Multi-Level PromptEnhanced Transformer framework (MLPET), jointly modeling both multi-level molecular features and inter-drug relational semantics. To capture multi-level structural knowledge, we design three complementary self-supervised pretraining tasks: (1) atom prediction, which predicts randomly masked atom types to learn local chemical context; (2) bond prediction, which infers the existence of chemical bonds between atom pairs to capture topological dependencies; and (3) distance prediction, which regresses the 3D spatial distance between atoms to encode molecular geometry. In the downstream stage, we introduce a lightweight prompt fusion mechanism that integrates task-specific prompts into a unified vector to guide fine-tuning. This enables flexible and efficient knowledge transfer to multiple tasks, including molecular property prediction and multi-class DDI event classification. Extensive experiments on MoleculeNet, Ryu, and Deng benchmarks demonstrate that MLPET consistently outperforms state-of-the-art baselines, particularly in low-resource and longtail scenarios. Our results highlight the potential of promptguided multi-task pretraining as a generalizable paradigm for molecular representation learning. The source code of MLPET is available at https://github.com/fuhaitao95/MLPET.
Qiqi Zhu, Haitao Fu
BIBM3
2025 DistRMI: a deep distance-aware neural network for explainable RNA loop motif-small molecule interaction prediction
abstract
RNA participates in the occurrence and development of various diseases by regulating gene expression. Owing to its potential to circumvent the limitations of traditional "undruggable" protein targets, it is regarded as a core direction for next-generation precision therapy. Against this backdrop, the accurate and interpretable prediction of RNA-small molecule interactions has become a key link in accelerating the discovery of RNA-targeted drugs. However, existing methods suffer from insufficient prediction accuracy and interpretability, failing to effectively guide lead compound screening or elucidate the mechanism of action. This study presents DistRMI, which integrates Transformers and graph neural networks to capture, respectively, the sequence information of RNA loop motifs and the chemical topological features of small molecules while introducing distance priors between them and leveraging a distance-aware attention mechanism to capture their interaction information. The results show that DistRMI outperforms baseline models, and its performance remains robust even when confronted with unknown RNA loop motifs and small molecules. Visualization of attention weights reveals that bases near the paired bases of RNA loop motifs contribute significantly. Furthermore, retrospective case studies validate the model's reliability. Predicting the binding preferences between RNA loop motifs and small molecules while providing interpretability facilitates an in-depth understanding of RNA-small molecule interactions, promotes in-depth research on RNA and related drugs, and opens up new avenues for disease treatment.
Zhaoxiang Liu, Qiqi Zhu, Qingyan Tian, Yingxiang Deng, Dengguo Wei, Haitao Fu
Briefings Bioinform.7
2025 3A Multi-Classification Division-Aggregation Framework for Fake News Detection
abstract
Nowadays, as human activities are shifting to social media, fake news detection has been a crucial problem. Existing methods ignore the classification difference in online news and cannot take full advantage of multi-classification knowledges. For example, when coping with a post “A mouse is frightened by a cat,” a model that learns “computer” knowledge tends to misunderstand “mouse” and give a fake label, but a model that learns “animal” knowledge tends to give a true label. Therefore, this research proposes a multi-classification division-aggregation framework to detect fake news, namedCKA, which innovatively learns classification knowledges during training stages and aggregates them during prediction stages. It consists of three main components: a news characterizer, an ensemble coordinator, and a truth predictor. The news characterizer is responsible for extracting news features and obtaining news classifications. Cooperating with the news characterizer, the ensemble coordinator generates classification-specifical models for the maximum reservation of classification knowledges during the training stage, where each classification-specifical model maximizes the detection performance of fake news on corresponding news classifications. Further, to aggregate the classification knowledges during the prediction stage, the truth predictor uses the truth discovery technology to aggregate the predictions from different classification-specifical models based on reliability evaluation of classification-specifical models. Extensive experiments prove that our proposedCKAoutperforms state-of-the-art baselines in fake news detection.
Wen Zhang 0008, Haitao Fu, Huan Wang 0005, Zhiguo Gong, Pan Zhou 0001, Di Wang 0015
IEEE Trans. Big Data2
2024 PolyGlotFake: A Novel Multilingual and Multimodal DeepFake Dataset
Yang Hou 0003, Haitao Fu, Chunkai Chen, Zida Li
ICPR (14)2
2024 Multiview representation learning for identification of novel cancer genes and their causative biological mechanisms
abstract
Tumorigenesis arises from the dysfunction of cancer genes, leading to uncontrolled cell proliferation through various mechanisms. Establishing a complete cancer gene catalogue will make precision oncology possible. Although existing methods based on graph neural networks (GNN) are effective in identifying cancer genes, they fall short in effectively integrating data from multiple views and interpreting predictive outcomes. To address these shortcomings, an interpretable representation learning framework IMVRL-GCN is proposed to capture both shared and specific representations from multiview data, offering significant insights into the identification of cancer genes. Experimental results demonstrate that IMVRL-GCN outperforms state-of-the-art cancer gene identification methods and several baselines. Furthermore, IMVRL-GCN is employed to identify a total of 74 high-confidence novel cancer genes, and multiview data analysis highlights the pivotal roles of shared, mutation-specific, and structure-specific representations in discriminating distinctive cancer genes. Exploration of the mechanisms behind their discriminative capabilities suggests that shared representations are strongly associated with gene functions, while mutation-specific and structure-specific representations are linked to mutagenic propensity and functional synergy, respectively. Finally, our in-depth analyses of these candidates suggest potential insights for individualized treatments: afatinib could counteract many mutation-driven risks, and targeting interactions with cancer gene SRC is a reasonable strategy to mitigate interaction-induced risks for NR3C1, RXRA, HNF4A, and SP1.
Jianye Yang 0002, Haitao Fu, Fei-Yang Xue, Menglu Li, Yuyang Wu, Zhanhui Yu, Haohui Luo, Xiaohui Niu
Briefings Bioinform.2
2023 Physical-aware Interconnect Testing and Repairing of Chiplets
abstract
As the interconnect density of chiplets increases rapidly, some physics related defects appeared, such as coupling defects, etc. These defects are hard to detect with ordinary pseudo-random sequence patterns, some special test patterns are needed. Besides, the chip warpage caused by the thinning of 3D chips manufacturing and the uneven stress around TSVs or micro bumps will bring clustered faults of interconnections. For these defects, the repair rate of conventional interconnect redundancy method will be decreased. This paper proposes a physical-aware interconnect testing and repairing method of chiplets, using specific test patterns and clustered faults redundancy circuits to improve the interconnect test coverage and repair rate of chiplets. We also propose automatic repair circuits and the repair data synchronization scheme between multiple dies, so that the calculating and programming of repair data do not need to rely on the ATE programming, and the synchronization of the repair data between multiple dies can be done by hardware circuits automatically, which ensure the interconnection correctly after repairing.
Changming Cui, Tuanhui Xu, Haitao Fu
ETS3
2023 HimGNN: a novel hierarchical molecular graph representation learning framework for property prediction
abstract
Accurate prediction of molecular properties is an important topic in drug discovery. Recent works have developed various representation schemes for molecular structures to capture different chemical information in molecules. The atom and motif can be viewed as hierarchical molecular structures that are widely used for learning molecular representations to predict chemical properties. Previous works have attempted to exploit both atom and motif to address the problem of information loss in single representation learning for various tasks. To further fuse such hierarchical information, the correspondence between learned chemical features from different molecular structures should be considered. Herein, we propose a novel framework for molecular property prediction, called hierarchical molecular graph neural networks (HimGNN). HimGNN learns hierarchical topology representations by applying graph neural networks on atom- and motif-based graphs. In order to boost the representational power of the motif feature, we design a Transformer-based local augmentation module to enrich motif features by introducing heterogeneous atom information in motif representation learning. Besides, we focus on the molecular hierarchical relationship and propose a simple yet effective rescaling module, called contextual self-rescaling, that adaptively recalibrates molecular representations by explicitly modelling interdependencies between atom and motif features. Extensive computational experiments demonstrate that HimGNN can achieve promising performances over state-of-the-art baselines on both classification and regression tasks in molecular property prediction.
Shen Han, Haitao Fu, Yuyang Wu, Ganglan Zhao, Feng Huang 0004, Zhongfei Zhang, Shichao Liu 0002, Wen Zhang 0008
Briefings Bioinform.2
2023 DRLM: A Robust Drug Representation Learning Method and its Applications
abstract
Learning representations from data is a fundamental step for machine learning. High-quality and robust drug representations can broaden the understanding of pharmacology, and improve the modeling of multiple drug-related prediction tasks, which further facilitates drug development. Although there are a number of models developed for drug representation learning from various data sources, few researches extract drug representations from gene expression profiles. Since gene expression profiles of drug-treated cells are widely used in clinical diagnosis and therapy, it is believed that leveraging them to eliminate cell specificity can promote drug representation learning. In this paper, we propose a three-stage deep learning method for drug representation learning, named DRLM, which integrates gene expression profiles of drug-related cells and the therapeutic use information of drugs. Firstly, we construct a stacked autoencoder to learn low-dimensional compact drug representations. Secondly, we utilize an iterative clustering module to reduce the negative effects of cell specificity and noise in gene expression profiles on the low-dimensional drug representations. Thirdly, a therapeutic use discriminator is designed to incorporate therapeutic use information into the drug representations. The visualization analysis of drug representations demonstrates DRLM can reduce cell specificity and integrate therapeutic use information effectively. Extensive experiments on three types of prediction tasks are conducted based on different drug representations, and they show that the drug representations learned by DRLM outperform other representations in terms of most metrics. The ablation analysis also demonstrates DRLM's effectiveness of merging the gene expression profiles with the therapeutic use information.
Haitao Fu, Cecheng Zhao, Xiaohui Niu, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 GraphCDR: a graph neural network method with contrastive learning for cancer drug response prediction
abstract
Predicting the response of a cancer cell line to a therapeutic drug is an important topic in modern oncology that can help personalized treatment for cancers. Although numerous machine learning methods have been developed for cancer drug response (CDR) prediction, integrating diverse information about cancer cell lines, drugs and their known responses still remains a great challenge. In this paper, we propose a graph neural network method with contrastive learning for CDR prediction. GraphCDR constructs a graph neural network based on multi-omics profiles of cancer cell lines, the chemical structure of drugs and known cancer cell line-drug responses for CDR prediction, while a contrastive learning task is presented as a regularizer within a multi-task learning paradigm to enhance the generalization ability. In the computational experiments, GraphCDR outperforms state-of-the-art methods under different experimental configurations, and the ablation study reveals the key components of GraphCDR: biological features, known cancer cell line-drug responses and contrastive learning are important for the high-accuracy CDR prediction. The experimental analyses imply the predictive power of GraphCDR and its potential value in guiding anti-cancer drug selection.
Xuan Liu 0010, Congzhi Song, Feng Huang 0004, Haitao Fu, Wenjie Xiao, Wen Zhang 0008
Briefings Bioinform.4
2022 MVGCN: data integration through multi-view graph convolutional network for predicting links in biomedical bipartite networks
abstract
MOTIVATION: There are various interaction/association bipartite networks in biomolecular systems. Identifying unobserved links in biomedical bipartite networks helps to understand the underlying molecular mechanisms of human complex diseases and thus benefits the diagnosis and treatment of diseases. Although a great number of computational methods have been proposed to predict links in biomedical bipartite networks, most of them heavily depend on features and structures involving the bioentities in one specific bipartite network, which limits the generalization capacity of applying the models to other bipartite networks. Meanwhile, bioentities usually have multiple features, and how to leverage them has also been challenging. RESULTS: In this study, we propose a novel multi-view graph convolution network (MVGCN) framework for link prediction in biomedical bipartite networks. We first construct a multi-view heterogeneous network (MVHN) by combining the similarity networks with the biomedical bipartite network, and then perform a self-supervised learning strategy on the bipartite network to obtain node attributes as initial embeddings. Further, a neighborhood information aggregation (NIA) layer is designed for iteratively updating the embeddings of nodes by aggregating information from inter- and intra-domain neighbors in every view of the MVHN. Next, we combine embeddings of multiple NIA layers in each view, and integrate multiple views to obtain the final node embeddings, which are then fed into a discriminator to predict the existence of links. Extensive experiments show MVGCN performs better than or on par with baseline methods and has the generalization capacity on six benchmark datasets involving three typical tasks. AVAILABILITY AND IMPLEMENTATION: Source code and data can be downloaded from https://github.com/fuhaitao95/MVGCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haitao Fu, Feng Huang 0004, Xuan Liu 0010, Wen Zhang 0008
Bioinform.1
2022 Hierarchical graph representation learning for the prediction of drug-target binding affinity
Zhaoyang Chu, Feng Huang 0004, Haitao Fu, Yuan Quan, Xionghui Zhou, Shichao Liu 0002, Wen Zhang 0008
Inf. Sci.3
2021 A robust drug representation learning model for eliminating cell specificity in gene expression profile and its application
abstract
Learning high-quality drug representations is important for drug development and the understanding of drug action mechanisms. Leveraging the gene expression profile of drug treated cells and eliminating cell specificity can facilitate drug representation learning. In this paper, we propose a four stage deep learning model that aims for drug representation learning based on integrating gene expression profile and the therapeutic use information of drugs, abbreviated as “DGERN”. The stacked autoencoder module is employed for data dimension reduction; the iterative clustering module is used to eliminate cell specificity; the subclass pre-training module and the label classifier module are utilized to integrate the therapeutic use information of drugs into drug representations. Visualization of the drug representations proves that DGERN eliminates cell specificity and integrates the therapeutic use information of drugs effectively. The drug representations learned by DGERN are used in the subsequent and prediction tasks of drug development. In the task of predicting drug-disease associations, DGERN combined with random forest achieves the best performance reaching 0.67 on AUC, exceeding 0.60 of the second-placed one; in the drug-drug interaction prediction task, DGERN combined with random forest gets 0.73 on AUC, which is second in comparison with other drug representations.
Cecheng Zhao, Hui Wang 0065, Haitao Fu, Yingjie Gao 0002, Xiaohui Niu
BIBM4
2021 The Advancement of 1149.10
abstract
Traditional scan test interfaces though general purpose (GP) IOs suffer from bandwidth limitations, contributing to increasing manufacturing test costs. Applying scan test through high speed IOs (i.e. SERDES) is a key interest area for the industry to overcome bandwidth limitations of traditional scan test, but is under the development in the industry today due to the need for compatible technologies across the entire ecosystem of EDA, ATE and IC designers. Guided by IEEE 1149.10, a group of pioneer engineers / users are deploying their robust and complete solutions. In this industry session, engineers from fabless, ATE and EDA companies are sharing their experiences in this area.
Haitao Fu, Edward Seng, Marc Hutner, Jean-François Côté, Geir Eide
ITC-Asia2