Hongjie Wu

dblp:128/5547 · DBLP profile ↗
← Back
73ranked-venue papers
14as first author
44since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 52 · 9 first-author · 27 since 2021Artificial intelligence and machine learning · 15 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VRQA: Context-adaptive view routing for long-document question-answer generation
abstract
Long-document question–answer (QA) generation often employs large-scale generation with strong filtering to reduce off-topic deviations. However, such strategies tend to concentrate generation on a few safe entry points, leading to uneven coverage and similar question expression. Conversely, expanding the scope of questioning may weaken topical consistency and the verifiability of supporting evidence. To address these challenges, we propose VRQA, a long-document QA generation framework with context-adaptive view routing under hierarchical anchor constraints. It selects better question entry points for different chunks and generates QA pairs with more dispersed semantic focuses. VRQA first constructs a semantic representation from chunk-level evidence and key points, then encodes view names and their descriptions into view prototypes, and obtains view-aware contextual features through conditional modulation. Subsequently, it updates the view routing policy online through reinforcement learning, where the reward is constructed using anchor consistency and QA answer quality scores from an evaluator. This enables VRQA to adaptively select the optimal view set and coordinate with submodular selection to achieve a dynamic balance between quality and diversity. To implement this framework, VRQA adopts a three-stage coupled training schedule, consisting of router warm-up, view-conditioned representation learning, and online updating, through which the router is progressively refined by evaluator feedback. Experiments on three vertical domains from the Elsevier OA CC-BY dataset show that VRQA achieves state-of-the-art performance, improving quality by 11.46% on AS-Chk and semantic diversity by 7.96% on VS over the strongest baseline, and thus offering a better trade-off between quality and diversity for long-document QA generation.
Mengting Huang, Hongjie Wu, Fuyuan Hu, Lanhui Liu, Qiming Fu 0001
Expert Syst. Appl.4
2025 Enhancing Diffusion Model Stability for Image Restoration via Gradient Management
abstract
Diffusion models have shown remarkable promise for image restoration by leveraging powerful priors. Prominent methods typically frame the restoration problem within a Bayesian inference framework, which iteratively combines a denoising step with a likelihood guidance step. However, the interactions between these two components in the generation process remain underexplored. In this paper, we analyze the underlying gradient dynamics of these components and identify significant instabilities. Specifically, we demonstrate conflicts between the prior and likelihood gradient directions, alongside temporal fluctuations in the likelihood gradient itself. We show that these instabilities disrupt the generative process and compromise restoration performance. To address these issues, we propose Stabilized Progressive Gradient Diffusion (SPGD), a novel gradient management technique. SPGD integrates two synergistic components: (1) a progressive likelihood warm-up strategy to mitigate gradient conflicts; and (2) adaptive directional momentum (ADM) smoothing to reduce fluctuations in the likelihood gradient. Extensive experiments across diverse restoration tasks demonstrate that SPGD significantly enhances generation stability, leading to state-of-the-art performance in quantitative metrics and visually superior results. Code is available at https://github.com/74587887/SPGD.
Hongjie Wu, Mingqin Zhang, Linchao He, Jizhe Zhou 0001, Jiancheng Lv 0001
ACM Multimedia1
2025 LinFa-Q: Accurate Q-learning with linear function approximation
Zhechao Wang, Qiming Fu 0001, Quan Liu 0004, You Lu 0004, Hongjie Wu, Fuyuan Hu
Neurocomputing6
2025 ICCR-Diff: Identity-preserving and controllable craniofacial reconstruction with diffusion models
Mingqin Zhang, Hongjie Wu, Zhengqing Zang, Jian Wang 0124, Chaoqun Niu, Jiancheng Lv 0001
Knowl. Based Syst.2
2025 TrGPCR: GPCR-Ligand Binding Affinity Prediction Based on Dynamic Deep Transfer Learning
abstract
Predicting G protein-coupled receptor (GPCR) -ligand binding affinity plays a crucial role in drug development. However, determining GPCR-ligand binding affinities is time-consuming and resource-intensive. Although many studies used data-driven methods to predict binding affinity, most of these methods required protein 3D structure, which was often unknown. Moreover, part of these studies only considered the sequence characteristics of the protein, ignoring the secondary structure of the protein. The number of known GPCR for affinity prediction is only a few thousand, which is insufficient for deep learning training. Therefore, this study aimed to propose a deep transfer learning method called TrGPCR, which used dynamic transfer learning to solve the problem of insufficient GPCR data. We used the Binding Database (BindingDB) as the source domain and the GLASS (GPCR-Ligand Association) database as the target domain. We also introduced protein secondary structures, called pockets, as features to predict binding affinities. Compared with DeepDTA, our model improved by 5.2% on RMSE (root mean square error) and 4.5% on MAE (mean squared error).
Yaoyao Lu, Tengsheng Jiang, Qiming Fu 0001, Zhiming Cui 0002, Hongjie Wu
IEEE J. Biomed. Health Informatics6
2024 Diffusion Posterior Proximal Sampling for Image Restoration
abstract
Diffusion models have demonstrated remarkable efficacy in generating high-quality samples. Existing diffusion-based image restoration algorithms exploit pre-trained diffusion models to leverage data priors, yet they still preserve elements inherited from the unconditional generation paradigm. These strategies initiate the denoising process with pure white noise and incorporate random noise at each generative step, leading to over-smoothed results. In this paper, we present a refined paradigm for diffusion-based image restoration. Specifically, we opt for a sample consistent with the measurement identity at each generative step, exploiting the sampling selection as an avenue for output stability and enhancement. The number of candidate samples used for selection is adaptively determined based on the signal-to-noise ratio of the timestep. Additionally, we start the restoration process with an initialization combined with the measurement signal, providing supplementary information to better align the generative process. Extensive experimental results and analyses validate that our proposed method significantly enhances image restoration performance while consuming negligible additional computational resources.
Hongjie Wu, Linchao He, Mingqin Zhang, Dongdong Chen 0004, Kunming Luo, Mengting Luo, Jizhe Zhou 0001, Hu Chen 0002, Jiancheng Lv 0001
ACM Multimedia1
2024 AMDGT: Attention aware multi-modal fusion using a dual graph transformer for drug-disease associations prediction
abstract
Identification of new indications for existing drugs is crucial through the various stages of drug discovery. Computational methods are valuable in establishing meaningful associations between drugs and diseases. However, most methods predict the drug-disease associations based solely on similarity data, neglecting valuable biological and chemical information. These methods often use basic concatenation to integrate information from different modalities, limiting their ability to capture features from a comprehensive and in-depth perspective. Therefore, a novel multimodal framework called AMDGT was proposed to predict new drug associations based on dual-graph transformer modules. By combining similarity data and complex biochemical information, AMDGT understands the multimodal feature fusion of drugs and diseases effectively and comprehensively with an attention-aware modality interaction architecture. Extensive experimental results indicate that AMDGT surpasses state-of-the-art methods in real-world datasets. Moreover, case and molecular docking studies demonstrated that AMDGT is an effective tool for drug repositioning. Our code is available at GitHub: https://github.com/JK-Liu7/AMDGT.
Quan Zou 0001, Hongjie Wu, Prayag Tiwari, Yijie Ding
Knowl. Based Syst.4
2024 AttentionMGT-DTA: A multi-modal drug-target affinity prediction using graph transformer and attention mechanism
abstract
The accurate prediction of drug-target affinity (DTA) is a crucial step in drug discovery and design. Traditional experiments are very expensive and time-consuming. Recently, deep learning methods have achieved notable performance improvements in DTA prediction. However, one challenge for deep learning-based models is appropriate and accurate representations of drugs and targets, especially the lack of effective exploration of target representations. Another challenge is how to comprehensively capture the interaction information between different instances, which is also important for predicting DTA. In this study, we propose AttentionMGT-DTA, a multi-modal attention-based model for DTA prediction. AttentionMGT-DTA represents drugs and targets by a molecular graph and binding pocket graph, respectively. Two attention mechanisms are adopted to integrate and interact information between different protein modalities and drug-target pairs. The experimental results showed that our proposed model outperformed state-of-the-art baselines on two benchmark datasets. In addition, AttentionMGT-DTA also had high interpretability by modeling the interaction strength between drug atoms and protein residues. Our code is available at https://github.com/JK-Liu7/AttentionMGT-DTA.
Hongjie Wu, Tengsheng Jiang, Quan Zou 0001, Shujie Qi, Zhiming Cui 0002, Prayag Tiwari, Yijie Ding
Neural Networks1
2024 Self-supervised Domain Adaptation with Significance-Oriented Masking for Pelvic Organ Prolapse detection
Hongjie Wu, Chenwei Tang, Dongdong Chen 0004, Yueyue Chen, Ling Mei 0002, Jiancheng Lv 0001
Pattern Recognit. Lett.2
2024 MultiModRLBP: A Deep Learning Approach for Multi-Modal RNA-Small Molecule Ligand Binding Sites Prediction
abstract
This study aims to tackle the intricate challenge of predicting RNA-small molecule binding sites to explore the potential value in the field of RNA drug targets. To address this challenge, we propose the MultiModRLBP method, which integrates multi-modal features using deep learning algorithms. These features include 3D structural properties at the nucleotide base level of the RNA molecule, relational graphs based on overall RNA structure, and rich RNA semantic information. In our investigation, we gathered 851 interactions between RNA and small molecule ligand from the RNAglib dataset and RLBind training set. Unlike conventional training sets, this collection broadened its scope by including RNA complexes that have the same RNA sequence but change their respective binding sites due to structural differences or the presence of different ligands. This enhancement enables the MultiModRLBP model to more accurately capture subtle changes at the structural level, ultimately improving its ability to discern nuances among similar RNA conformations. Furthermore, we evaluated MultiModRLBP on two classic test sets, Test18 and Test3, highlighting its performance disparities on small molecules based on metal and non-metal ions. Additionally, we conducted a structural sensitivity analysis on specific complex categories, considering RNA instances with varying degrees of structural changes and whether they share the same ligands. The research results indicate that MultiModRLBP outperforms the current state-of-the-art methods on multiple classic test sets, particularly excelling in predicting binding sites for non-metal ions and instances where the binding sites are widely distributed along the sequence. MultiModRLBP also can be used as a potential tool when the RNA structure is perturbed or the RNA experimental tertiary structure is not available. Most importantly, MultiModRLBP exhibits the capability to distinguish binding characteristics of RNA that are structurally diverse yet exhibit sequence similarity. These advancements hold promise in reducing the costs associated with the development of RNA-targeted drugs.
Lijun Quan, Hongjie Wu, Xuhao Ma, Jingxin Xie, Deng Pan 0006, Taoning Chen, Tingfang Wu, Qiang Lyu
IEEE J. Biomed. Health Informatics4
2024 Euclidean Distance is Not Your Swiss Army Knife
abstract
Graph-based multi-view learning, which has hitherto been used to discover the intrinsic patterns of graph data giving the credit to its convenience of implementation and effectiveness. Note that even though these approaches have been increasingly adopted in multi-view clustering and have generated promising outcomes, they are still faced with the sub-optimal solution. For one thing, multi-view data can be corrupted in the raw feature space. For the other, most existing approaches normally utilize euclidean distance to obtain the similarity between two samples, which can not be the best option for all types of real-world data and leads to inferior results. Therefore, to overcome the aforementioned issues, we integrate multi-metric learning, graph filtering, and subspace learning into a collaborative learning framework for multi-view clustering. Particularly, we prefer to recover a smooth representation of data by graph filtering, which can reserve the geometric structure of the original multi-view data and discard the corruptions simultaneously. Furthermore, instead of using euclidean distance as a Swiss army knife, multiple metrics are utilized to fully exploit the correlation of data based on the smooth representation, hence finally facilitating the downstream clustering task. Extensive experiments on multi-view clustering tasks validate our theoretical findings of ours and prove the improvement of our method over the SOTA approaches.
Yuze Tan, Yixi Liu, Hongjie Wu, Shudong Huang, Zenglin Xu, Ivor W. Tsang, Jiancheng Lv 0001
IEEE Trans. Knowl. Data Eng.3
2023 Metric Multi-View Graph Clustering
abstract
Graph-based methods have hitherto been used to pursue the coherent patterns of data due to its ease of implementation and efficiency. These methods have been increasingly applied in multi-view learning and achieved promising performance in various clustering tasks. However, despite their noticeable empirical success, existing graph-based multi-view clustering methods may still suffer the suboptimal solution considering that multi-view data can be very complicated in raw feature space. Moreover, existing methods usually adopt the similarity metric by an ad hoc approach, which largely simplifies the relationship among real-world data and results in an inaccurate output. To address these issues, we propose to seamlessly integrates metric learning and graph learning for multi-view clustering. Specifically, we employ a useful metric to depict the inherent structure with linearity-aware of affinity graph representation learned based on the self-expressiveness property. Furthermore, instead of directly utilizing the raw features, we prefer to recover a smooth representation such that the geometric structure of the original data can be retained. We model the above concerns into a unified learning framework, and hence complements each learning subtask in a mutual reinforcement manner. The empirical studies corroborate our theoretical findings, and demonstrate that the proposed method is able to boost the multi-view clustering performance.
Yuze Tan, Yixi Liu, Hongjie Wu, Jiancheng Lv 0001, Shudong Huang
AAAI3
2023 Deep Learning-Based Prediction of Drug-Target Binding Affinities by Incorporating Local Structure of Protein
Baozhong Zhu, Tengsheng Jiang, Hongjie Wu
ICIC (3)5
2023 A Transformer-Based Deep Learning Approach with Multi-layer Feature Processing for Accurate Prediction of Protein-DNA Binding Residues
Haipeng Zhao, Baozhong Zhu, Tengsheng Jiang, Hongjie Wu
ICIC (3)5
2023 Drug-Target Interaction Prediction Based on Interpretable Graph Transformer Model
Baozhong Zhu, Tengsheng Jiang, Hongjie Wu
ICIC (3)5
2023 ST-CA YOLOv5: Improved YOLOv5 Based on Swin Transformer and Coordinate Attention for Surface Defect Detection
abstract
Surface defect detection plays a crucial role in industrial equipment to ensure industrial safety. With the development of deep learning, a series of deep learning-based surface defect detection algorithms are proposed and achieved significant success. However, the application of the algorithms in real-world scenarios is restricted by computing power and memory resource, resulting in either deployment issues or performance degradation. In order to balance memory consumption and detection accuracy, we propose a novel method named ST-CA YOLOv5 for surface defect detection. Based on YOLOv5, The Swin Transformer Block (ST) is introduced to design the C3STR module, enhancing the ability to capture long-range semantic information. In the prediction head, we present the CAHead structure by utilizing the lightweight attention module, Coordinate Attention (CA), to fuse feature information. With the benefit of the ST and CA module, the model improves the detection ability for small objects and yields better overall detection performance. Extensive experiments conducted on several real-world datasets demonstrate the effectiveness and superiority of the proposed method, compared with the state-of-the-art methods in terms of detection performance.
Hongjie Wu, Chenwei Tang, Jiancheng Lv 0001
IJCNN2
2023 Preserving Local and Global Information: An Effective Metric-based Subspace Clustering
abstract
Subspace clustering, which recoveries the subspace representation in the form of an affinity graph, has drawn tons of attention due to its effectiveness in various clustering tasks. However, existing subspace clustering methods are usually fed with raw data, which may lead to a suboptimal result since it is difficult to directly and accurately depict the inherent relation between data points. In this paper, we propose a novel subspace clustering method by holistically utilizing the pairwise similarity and graph geometric structure. Our model first constructs an initial subspace representation by means of self-expression, which is able to depict the global structure of data. Then, we use an effective metric to recover an intrinsic matrix with pairwise similarity based on the obtained representation, which further preserves the local structure. Besides, we propose to facilitate the downstream subspace learning task by searching for a smooth representation of the original data, which is obtained by applying a low-pass filter to retain the graph geometric features. By leveraging the subtasks of learning the smooth representation, performing the subspace learning, and recovering the intrinsic similarity matrix in a unified learning framework, each subtask can be alternately boosted. Experiments on several benchmark data sets have been conducted to verify the proposed method.
Yixi Liu, Yuze Tan, Hongjie Wu, Shudong Huang, Yazhou Ren 0001, Jiancheng Lv 0001
ACM Multimedia3
2023 Reinforcement Learning in Few-Shot Scenarios: A Survey
Zhechao Wang, Qiming Fu 0001, You Lu 0004, Hongjie Wu
J. Grid Comput.6
2023 Extracting biomedical relation from cross-sentence text using syntactic dependency graph attention network
Xueyang Zhou, Qiming Fu 0001, Lanhui Liu, You Lu 0004, Hongjie Wu
J. Biomed. Informatics7
2023 Pure graph-guided multi-view subspace clustering
Hongjie Wu, Shudong Huang, Chenwei Tang, Yancheng Zhang, Jiancheng Lv 0001
Pattern Recognit.1
2023 MV-H-RKM: A Multiple View-Based Hypergraph Regularized Restricted Kernel Machine for Predicting DNA-Binding Proteins
abstract
DNA-binding proteins (DBPs) have a significant impact on many life activities, so identification of DBPs is a crucial issue. And it is greatly helpful to understand the mechanism of protein-DNA interactions. In traditional experimental methods, it is significant time-consuming and labor-consuming to identify DBPs. In recent years, many researchers have proposed lots of different DBP identification methods based on machine learning algorithm to overcome shortcomings mentioned above. However, most existing methods cannot get satisfactory results. In this paper, we focus on developing a new predictor of DBPs, called Multi-View Hypergraph Restricted Kernel Machines (MV-H-RKM). In this method, we extract five features from the three views of the proteins. To fuse these features, we couple them by means of the shared hidden vector. Besides, we employ the hypergraph regularization to enforce the structure consistency between original features and the hidden vector. Experimental results show that the accuracy of MV-H-RKM is 84.09% and 85.48% on PDB1075 and PDB186 data set respectively, and demonstrate that our proposed method performs better than other state-of-the-art approaches. The code is publicly available at https://github.com/ShixuanGG/MV-H-RKM.
Yuqing Qian, Tengsheng Jiang, Min Jiang 0009, Yijie Ding, Hongjie Wu
IEEE ACM Trans. Comput. Biol. Bioinform.6
2023 Protein-DNA Binding Residues Prediction Using a Deep Learning Model With Hierarchical Feature Extraction
abstract
Biologically important effects occur when proteins bind to other substances, of which binding to DNA is a crucial one. Therefore, accurate identification of protein-DNA binding residues is important for further understanding of the protein-DNA interaction mechanism. Although wet-lab methods can accurately obtain the location of bound residues, it requires significant human, financial and time costs. There is thus an urgent need to develop efficient computational-based methods. Most current state-of-the-art methods are two-step approaches: the first step uses a sliding window technique to extract residue features; the second step uses each residue as an input to the model for prediction. This has a negative impact on the efficiency of prediction and ease of use. In this study, we propose a sequence-to-sequence (seq2seq) model that can input the entire protein sequence of variable length and use two modules, Transformer Encoder Block and Feature Extracting Block, for hierarchical feature extraction, where Transformer Encoder Block is used to extract global features, and then Feature Extracting Block is used to extract local features to further improve the recognition capability of the model. The comparison results on two benchmark datasets, namely PDNA-543 and PDNA-41, prove the effectiveness of our method in identifying protein-DNA binding residues.
Quan Zou 0001, Hongjie Wu, Yijie Ding
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 Study on Path Planning of Multi-storey Parking Lot Based on Combined Loss Function
Zhongtian Hu, Yuli Wang, Qiming Fu 0001, Weizhong Lu, Hongjie Wu
ICIC (3)7
2022 Drug-Target Interaction Prediction Based on Graph Neural Network and Recommendation System
Chang-an Yuan 0001, Hongjie Wu, Xingming Zhao
ICIC (2)3
2022 Drug-Target Interaction Prediction Based on Transformer
Tengsheng Jiang, Yaoyao Lu, Hongjie Wu
ICIC (2)4
2022 Protein-Ligand Binding Affinity Prediction Based on Deep Learning
Yaoyao Lu, Tengsheng Jiang, Hongjie Wu
ICIC (2)5
2022 Local Feature for Visible-Thermal PReID Based on Transformer
Quanyi Pu, Chang-an Yuan 0001, Hongjie Wu, Xingming Zhao
ICIC (1)3
2022 An Effective Method for Yemeni License Plate Recognition Based on Deep Neural Networks
Hamdan Taleb, Zhipeng Li 0002, Chang-an Yuan 0001, Hongjie Wu, Xingming Zhao, Fahd A. Ghanem
ICIC (3)4
2022 Comprehensive Evaluation of BERT Model for DNA-Language for Prediction of DNA Sequence Binding Specificities in Fine-Tuning Phase
Xianbao Tan, Chang-an Yuan 0001, Hongjie Wu, Xingming Zhao
ICIC (2)3
2022 Using Deep Learning to Predict Transcription Factor Binding Sites Based on Multiple-omics Data
Youhong Xu, Chang-an Yuan 0001, Hongjie Wu, Xingming Zhao
ICIC (1)3
2022 Multi-view Subspace Clustering on Topological Manifold
abstract
Multi-view subspace clustering aims to exploit a common affinity representation by means of self-expression. Plenty of works have been presented to boost the clustering performance, yet seldom considering the topological structure in data, which is crucial for clustering data on manifold. Orthogonal to existing works, in this paper, we argue that it is beneficial to explore the implied data manifold by learning the topological relationship between data points. Our model seamlessly integrates multiple affinity graphs into a consensus one with the topological relevance considered. Meanwhile, we manipulate the consensus graph by a connectivity constraint such that the connected components precisely indicate different clusters. Hence our model is able to directly obtain the final clustering result without reliance on any label discretization strategy as previous methods do. Experimental results on several benchmark datasets illustrate the effectiveness of the proposed model, compared to the state-of-the-art competitors over the clustering performance.
Shudong Huang, Hongjie Wu, Yazhou Ren 0001, Ivor W. Tsang, Zenglin Xu, Wentao Feng, Jiancheng Lv 0001
NeurIPS2
2022 G Protein-Coupled Receptor Interaction Prediction Based on Deep Transfer Learning
abstract
G protein-coupled receptors (GPCRs) account for about 40% to 50% of drug targets. Many human diseases are related to G protein coupled receptors. Accurate prediction of GPCR interaction is not only essential to understand its structural role, but also helps design more effective drugs. At present, the prediction of GPCR interaction mainly uses machine learning methods. Machine learning methods generally require a large number of independent and identically distributed samples to achieve good results. However, the number of available GPCR samples that have been marked is scarce. Transfer learning has a strong advantage in dealing with such small sample problems. Therefore, this paper proposes a transfer learning method based on sample similarity, using XGBoost as a weak classifier and using the TrAdaBoost algorithm based on JS divergence for data weight initialization to transfer samples to construct a data set. After that, the deep neural network based on the attention mechanism is used for model training. The existing GPCR is used for prediction. In short-distance contact prediction, the accuracy of our method is 0.26 higher than similar methods.
Tengsheng Jiang, Yuhui Chen, Zhongtian Hu, Weizhong Lu, Qiming Fu 0001, Yijie Ding, Haiou Li, Hongjie Wu
IEEE ACM Trans. Comput. Biol. Bioinform.9
2021 A Reinforcement Learning-Based Model for Human MicroRNA-Disease Association Prediction
Linqian Cui, You Lu 0004, Qiming Fu 0001, Yijie Ding, Hongjie Wu
ICIC (3)7
2021 DNA-Binding Protein Prediction Based on Deep Learning Feature Fusion
Tengsheng Jiang, Weizhong Lu, Qiming Fu 0001, Haiou Li, Hongjie Wu
ICIC (3)6
2021 Plant Leaf Recognition Network Based on Fine-Grained Visual Classification
Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu
ICIC (1)4
2021 Research on RNA Secondary Structure Prediction Based on MLP
Weizhong Lu, Yu Zhang 0027, Hongjie Wu, Yijie Ding
ICIC (3)4
2021 Membrane Protein Identification via Multiple Kernel Fuzzy SVM
Weizhong Lu, Yuqing Qian, Hongjie Wu, Yijie Ding
ICIC (3)4
2021 Serialized Local Feature Representation Learning for Infrared-Visible Person Re-identification
Si-Zhe Wan, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu
ICIC (1)4
2021 Using Deep Learning to Predict Transcription Factor Binding Sites Combining Raw DNA Sequence, Evolutionary Information and Epigenomic Data
Youhong Xu, Qinghu Zhang, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu
ICIC (3)6
2021 Super-Large Medical Image Storage and Display Technology Based on Concentrated Points of Interest
Yuli Wang, Haiou Li, Weizhong Lu, Hongjie Wu
ICIC (1)5
2021 Automatic Cataract Detection with Multi-Task Learning
abstract
Cataract is one of the most prevalent diseases among the elderly. As the population ages, the incidence of cataracts is on the rise. Early diagnosis and treatment are essential for cataracts. The routine early diagnosis relies on B-scan eye ultrasound images, developing deep learning-based automatic cataract detection makes great sense. However, ultrasound images are complex and contain irrelevant backgrounds, the lens takes up only a small part. Besides, detection networks commonly use only one label as supervision, which leads to low classification accuracy and poor generalization. This paper focuses on making the most of the information in the images, thus proposing a new paradigm for automatic cataract detection. First, an object detection network is included to locate the eyeball area and eliminate the influence of the background. Next, we construct a dataset with multiple labels for each image. We extract the text descriptions of ultrasound images into labels so that each image is tagged with multiple labels. Then we applied the multi-task learning (MTL) methods to cataract detection. The accuracy of classification is significantly improved compared to data with only one label. Last, we propose two gradient-guided auxiliary learning methods to make the auxiliary tasks improve the performance of the main task (cataract detection). The experimental results show that our proposed methods further improve the classification accuracy.
Hongjie Wu, Jiancheng Lv 0001, Jian Wang 0124
IJCNN1
2021 Multi-zone Residential HVAC Control with Satisfying Occupants' Thermal Comfort Requirements and Saving Energy via Reinforcement Learning
Zhengkai Ding, Qiming Fu 0001, Hongjie Wu, You Lu 0004, Fuyuan Hu
PDCAT4
2021 Research on RNA secondary structure predicting via bidirectional recurrent neural network
abstract
BACKGROUND: RNA secondary structure prediction is an important research content in the field of biological information. Predicting RNA secondary structure with pseudoknots has been proved to be an NP-hard problem. Traditional machine learning methods can not effectively apply protein sequence information with different sequence lengths to the prediction process due to the constraint of the self model when predicting the RNA secondary structure. In addition, there is a large difference between the number of paired bases and the number of unpaired bases in the RNA sequences, which means the problem of positive and negative sample imbalance is easy to make the model fall into a local optimum. To solve the above problems, this paper proposes a variable-length dynamic bidirectional Gated Recurrent Unit(VLDB GRU) model. The model can accept sequences with different lengths through the introduction of flag vector. The model can also make full use of the base information before and after the predicted base and can avoid losing part of the information due to truncation. Introducing a weight vector to predict the RNA training set by dynamically adjusting each base loss function solves the problem of balanced sample imbalance. RESULTS: The algorithm proposed in this paper is compared with the existing algorithms on five representative subsets of the data set RNA STRAND. The experimental results show that the accuracy and Matthews correlation coefficient of the method are improved by 4.7% and 11.4%, respectively. CONCLUSIONS: The flag vector introduced allows the model to effectively use the information before and after the protein sequence; the introduced weight vector solves the problem of unbalanced sample balance. Compared with other algorithms, the LVDB GRU algorithm proposed in this paper has the best detection results.
Weizhong Lu, Hongjie Wu, Yijie Ding, Zhengwei Song, Yu Zhang 0027, Qiming Fu 0001, Haiou Li
BMC Bioinform.3
2021 Empirical Potential Energy Function Toward ab Initio Folding G Protein-Coupled Receptors
abstract
Approximately 40-50 percent of all drugs targets are G protein-coupled receptors (GPCRs). Three-dimensional structure of GPCRs is important to probe their biophysical and biochemical functions and their pharmaceutical applications. Lacking reliable and high quality free function is one of the ugent problems of computational predicting the three-dimensional structure in this community. We proposed a GPCR-specified energy function composed of four novel empirical potential energy terms: a two-dimensional contact energy force field, knowledge-based helix pair connection distance energy term, knowledge-based helix pair angle restraint energy term and a disulfide bond energy term. To validate the energy function, we employed an ab initio GPCR three-dimensional structure predictor to test if the energy function improved the accuracy of prediction. We evaluated 28 solved GPCRs and found that 21(75 percent) targets were correctly folded (TM-score>0.5). Also, the average TM-score using the energy function was 0.54, which was improved 134 percent than the TM-score 0.23 for MODELLER energy function and 170 percent than the TM-score 0.20 for Rosetta membrane energy function. The results confirmed that our empirical potential energy function toward ab initio folding is competitive to state-of-the-art solutions for structural prediction of GPCRs.
Hongjie Wu, Huajing Ling, Qiming Fu 0001, Weizhong Lu, Yijie Ding, Min Jiang 0009, Haiou Li
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 Three-Layer Dynamic Transfer Learning Language Model for E. Coli Promoter Classification
Qinhu Zhang, Siguo Wang, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao
ICIC (2)7
2020 Prediction of Membrane Protein Interaction Based on Deep Residual Learning
Tengsheng Jiang, Hongjie Wu, Yuhui Chen, Haiou Li, Jin Qiu, Weizhong Lu, Qiming Fu 0001
ICIC (2)2
2020 License Plate Detection and Recognition Technology for Complex Real Scenarios
Zhipeng Li 0002, Hamdan Taleb, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao
ICIC (1)6
2020 A New Method Combining DNA Shape Features to Improve the Prediction Accuracy of Transcription Factor Binding Sites
Siguo Wang, Qinhu Zhang, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao
ICIC (2)7
2020 Position Attention-Guided Learning for Infrared-Visible Person Re-identification
Yong Wu 0006, Si-Zhe Wan, Di Wu 0030, Chao Wang 0071, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao
ICIC (1)7
2020 Plant Leaf Recognition Network Based on Feature Learning and Metric Learning
Di Wu 0030, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao, Zhong-Qiu Zhao
ICIC (1)5
2020 Random Occlusion Recovery with Noise Channel for Person Re-identification
Di Wu 0030, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao, Yuchuan Du, Hanli Wang
ICIC (1)5
2020 Predicting in-Vitro Transcription Factor Binding Sites with Deep Embedding Convolution Network
Yindong Zhang, Qinhu Zhang, Chang-an Yuan 0001, Xiao Qin 0005, Hongjie Wu, Xingming Zhao
ICIC (2)5
2020 A Building Energy Consumption Prediction Method Based on Integration of a Deep Neural Network and Transfer Reinforcement Learning
abstract
With respect to the problem of the low accuracy of traditional building energy prediction methods, this paper proposes a novel prediction method for building energy consumption, which is based on the seamless integration of the deep neural network and transfer reinforcement learning (DNN-TRL). The method introduces a stack denoising autoencoder to extract the deep features of the building energy consumption, and shares the hidden layer structure to transfer the common information between different building energy consumption problems. The output of the DNN model is used as the input of the Sarsa algorithm to improve the prediction performance of the target building energy consumption. To verify the performance of the DNN-TRL algorithm, based on the data recorded by American Power Balti Gas and Electric Power Company, and compared with Sarsa, ADE-BPNN, and BP-Adaboost algorithms, the experimental results show that the DNN-TRL algorithm can effectively improve the prediction accuracy of the building energy consumption.
Qiming Fu 0001, QingSong Liu, Hongjie Wu, Baochuan Fu
Int. J. Pattern Recognit. Artif. Intell.4
2019 Knowledge Based Helix Angle and Residue Distance Restraint Free Energy Terms of GPCRs
Huajing Ling, Hongjie Wu, Jiayan Han, Jiwen Ding, Weizhong Lu, Qiming Fu 0001
ICIC (2)2
2019 Research on RNA Secondary Structure Prediction Based on Decision Tree
Weizhong Lu, Hongjie Wu, Hongmei Huang, Yijie Ding
ICIC (2)3
2019 A Prediction Method of DNA-Binding Proteins Based on Evolutionary Information
Weizhong Lu, Zhengwei Song, Yijie Ding, Hongjie Wu, Hongmei Huang
ICIC (2)4
2019 Predicting RNA secondary structure via adaptive deep recurrent neural networks with energy-based filter
abstract
BACKGROUND: RNA secondary structure prediction is an important issue in structural bioinformatics, and RNA pseudoknotted secondary structure prediction represents an NP-hard problem. Recently, many different machine-learning methods, Markov models, and neural networks have been employed for this problem, with encouraging results regarding their predictive accuracy; however, their performances are usually limited by the requirements of the learning model and over-fitting, which requires use of a fixed number of training features. Because most natural biological sequences have variable lengths, the sequences have to be truncated before the features are employed by the learning model, which not only leads to the loss of information but also destroys biological-sequence integrity. RESULTS: To address this problem, we propose an adaptive sequence length based on deep-learning model and integrate an energy-based filter to remove the over-fitting base pairs. CONCLUSIONS: Comparative experiments conducted on an authoritative dataset RNA STRAND (RNA secondary STRucture and statistical Analysis Database) revealed a 12% higher accuracy relative to three currently used methods.
Weizhong Lu, Hongjie Wu, Hongmei Huang, Qiming Fu 0001, Haiou Li
BMC Bioinform.3
2019 Ranking near-native candidate protein structures via random forest classification
abstract
BACKGROUND: In ab initio protein-structure predictions, a large set of structural decoys are often generated, with the requirement to select best five or three candidates from the decoys. The clustered central structures with the most number of neighbors are frequently regarded as the near-native protein structures with the lowest free energy; however, limitations in clustering methods and three-dimensional structural-distance assessments make identifying exact order of the best five or three near-native candidate structures difficult. RESULTS: To address this issue, we propose a method that re-ranks the candidate structures via random forest classification using intra- and inter-cluster features from the results of the clustering. Comparative analysis indicated that our method was better able to identify the order of the candidate structures as comparing with current methods SPICKR, Calibur, and Durandal. The results confirmed that the identification of the first model were closer to the native structure in 12 of 43 cases versus four for SPICKER, and the same as the native structure in up to 27 of 43 cases versus 14 for Calibur and up to eight of 43 cases versus two for Durandal. CONCLUSIONS: In this study, we presented an improved method based on random forest classification to transform the problem of re-ranking the candidate structures by an binary classification. Our results indicate that this method is a powerful method for the problem and the effect of this method is better than other methods.
Hongjie Wu, Hongmei Huang, Weizhong Lu, Qiming Fu 0001, Yijie Ding, Haiou Li
BMC Bioinform.1
2019 Research on predicting 2D-HP protein folding using reinforcement learning with full state space
abstract
BACKGROUND: Protein structure prediction has always been an important issue in bioinformatics. Prediction of the two-dimensional structure of proteins based on the hydrophobic polarity model is a typical non-deterministic polynomial hard problem. Currently reported hydrophobic polarity model optimization methods, greedy method, brute-force method, and genetic algorithm usually cannot converge robustly to the lowest energy conformations. Reinforcement learning with the advantages of continuous Markov optimal decision-making and maximizing global cumulative return is especially suitable for solving global optimization problems of biological sequences. RESULTS: In this study, we proposed a novel hydrophobic polarity model optimization method derived from reinforcement learning which structured the full state space, and designed an energy-based reward function and a rigid overlap detection rule. To validate the performance, sixteen sequences were selected from the classical data set. The results indicated that reinforcement learning with full states successfully converged to the lowest energy conformations against all sequences, while the reinforcement learning with partial states folded 50% sequences to the lowest energy conformations. Reinforcement learning with full states hits the lowest energy on an average 5 times, which is 40 and 100% higher than the three and zero hit by the greedy algorithm and reinforcement learning with partial states respectively in the last 100 episodes. CONCLUSIONS: Our results indicate that reinforcement learning with full states is a powerful method for predicting two-dimensional hydrophobic-polarity protein structure. It has obvious competitive advantages compared with greedy algorithm and reinforcement learning with partial states.
Hongjie Wu, Qiming Fu 0001, Weizhong Lu, Haiou Li
BMC Bioinform.1
2019 Variational Bayesian Exploration-Based Active Sarsa Algorithm
abstract
We proposed an improved variational Bayesian exploration-based active Sarsa (VBE-ASAR) algorithm, which tries to balance the exploration and exploitation dilemma, and speeds up the convergence rate. First, in the learning process, variational Bayesian method is adopted to measure the information gain, which is used as an exploration factor to construct an internal reward function for heuristic exploration. In addition, before the learning process, in order to improve the exploration performance, transfer learning is used to initialize the value function, where Bisimulation metric is introduced to measure the distance between two states from the source MDP and the target MDP, respectively. Finally, we apply the proposed algorithm to the cliff walking problem, and compare with the Sarsa algorithm, the Q-Learning algorithm, the VFT-Sarsa algorithm and the Bayesian Sarsa (BS) algorithm. Experimental results show that the VBE-ASAR algorithm has a faster learning rate.
Qiming Fu 0001, Zhengxia Yang, You Lu 0004, Hongjie Wu, Fuyuan Hu
Int. J. Pattern Recognit. Artif. Intell.4
2018 Prediction of Indoor PM2.5 Index Using Genetic Neural Network Model
Hongjie Wu, Weisheng Liu, Qiming Fu 0001, Baochuan Fu, Dadong Dai
ICIC (1)1
2018 Optimizing GPCR Two-Dimensional Topology from Contact Map
Hongjie Wu, Dadong Dai, Huaxiang Shen, Weizhong Lu, Qiming Fu 0001
ICIC (3)1
2018 RNA Secondary Structure Prediction Based on Long Short-Term Memory Model
Hongjie Wu, Weizhong Lu, Hongmei Huang, Qiming Fu 0001
ICIC (1)1
2018 Optimizing HP Model Using Reinforcement Learning
Hongjie Wu, Qiming Fu 0001
ICIC (2)2
2018 Single Trajectory Learning: Exploration Versus Exploitation
abstract
In reinforcement learning (RL), the exploration/exploitation (E/E) dilemma is a very crucial issue, which can be described as searching between the exploration of the environment to find more profitable actions, and the exploitation of the best empirical actions for the current state. We focus on the single trajectory RL problem where an agent is interacting with a partially unknown MDP over single trajectories, and try to deal with the E/E in this setting. Given the reward function, we try to find a good E/E strategy to address the MDPs under some MDP distribution. This is achieved by selecting the best strategy in mean over a potential MDP distribution from a large set of candidate strategies, which is done by exploiting single trajectories drawn from plenty of MDPs. In this paper, we mainly make the following contributions: (1) We discuss the strategy-selector algorithm based on formula set and polynomial function. (2) We provide the theoretical and experimental regret analysis of the learned strategy under an given MDP distribution. (3) We compare these methods with the “state-of-the-art” Bayesian RL method experimentally.
Qiming Fu 0001, Quan Liu 0004, Hongjie Wu
Int. J. Pattern Recognit. Artif. Intell.5
2018 Unified Deep Learning Architecture for Modeling Biology Sequence
abstract
Prediction of the spatial structure or function of biological macromolecules based on their sequences remains an important challenge in bioinformatics. When modeling biological sequences using traditional sequencing models, long-range interaction, complicated and variable output of labeled structures, and variable length of biological sequences usually lead to different solutions on a case-by-case basis. This study proposed a unified deep learning architecture based on long short-term memory or a gated recurrent unit to capture long-range interactions. The architecture designs the optional reshape operator to adapt to the diversity of the output labels and implements a training algorithm to support the training of sequence models capable of processing variable-length sequences. The merging and pooling operators enhances the ability of capturing short-range interactions between basic units of biological sequences. The proposed deep-learning architecture and its training algorithm might be capable of solving currently variable biological sequence-modeling problems under a unified framework. We validated the model on one of the most difficult biological sequence-modeling problems, protein residue interaction prediction. The results indicate that the accuracy of obtaining the residue interactions of the model exceeded popular approaches by 10 percent on multiple widely-used benchmarks.
Hongjie Wu, Chengyuan Cao, Xiaoyan Xia
IEEE ACM Trans. Comput. Biol. Bioinform.1
2017 β-Barrel Transmembrane Protein Predicting Using Support Vector Machine
Hongjie Wu, Kaihui Bian
ICIC (3)2
2017 Deep Conditional Random Field Approach to Transmembrane Topology Prediction and Application to GPCR Three-Dimensional Structure Modeling
abstract
Transmembrane proteins play important roles in cellular energy production, signal transmission, and metabolism. Many shallow machine learning methods have been applied to transmembrane topology prediction, but the performance was limited by the large size of membrane proteins and the complex biological evolution information behind the sequence. In this paper, we proposed a novel deep approach based on conditional random fields named as dCRF-TM for predicting the topology of transmembrane proteins. Conditional random fields take into account more complicated interrelation between residue labels in full-length sequence than HMM and SVM-based methods. Three widely-used datasets were employed in the benchmark. DCRF-TM had the accuracy 95 percent over helix location prediction and the accuracy 78 percent over helix number prediction. DCRF-TM demonstrated a more robust performance on large size proteins (>350 residues) against 11 state-of-the-art predictors. Further dCRF-TM was applied to ab initio modeling three-dimensional structures of seven-transmembrane receptors, also known as G protein-coupled receptors. The predictions on 24 solved G protein-coupled receptors and unsolved vasopressin V2 receptor illustrated that dCRF-TM helped abGPCR-I-TASSER to improve TM-score 34.3 percent rather than using the random transmembrane definition. Two out of five predicted models caught the experimental verified disulfide bonds in vasopressin V2 receptor.
Hongjie Wu, Kun Wang 0005, Liyao Lu, Yu Xue 0003, Qiang Lyu, Min Jiang 0009
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 Modeling of fuzzy comprehensive evaluation based on cloud model
abstract
A fuzzy comprehensive evaluation model based on cloud model which can be applied to evaluate air quality was established in this paper. Firstly, the basic content of the cloud model was introduced. The evaluation set, the weight set and the membership degree matrix based on the cloud model were discussed, and the evaluation model was established, which overcame the shortage of traditional fuzzy mathematics that describes these factors with precise numbers. Secondly, the result of the comprehensive evaluation was expressed by the cloud model, in which fuzziness and randomness were considered. Thus, it can avoid the large deviation resulting from maximum membership degree principle used in traditional fuzzy comprehensive evaluation, leading to more precious evaluation results, which correspond to human understanding. Finally, a case study of the comprehensive evaluation model was conducted, in which the air quality can be evaluated effectively.
Ning Feng, Baochuan Fu, Hongjie Wu, Xiuhua Wang 0004
ICARCV3
2016 A Parallel Multiple K-Means Clustering and Application on Detect Near Native Model
Hongjie Wu, Longfei Song, Min Jiang 0009
ICIC (2)1
2015 Predicting Helix Boundaries of α-Helix Transmembrane Protein with Feedback Conditional Random Fields
Kun Wang 0005, Hongjie Wu, Weizhong Lu, Baochuan Fu
ICIC (1)2
2014 Improved packing of protein side chains with parallel ant colonies
abstract
INTRODUCTION: The accurate packing of protein side chains is important for many computational biology problems, such as ab initio protein structure prediction, homology modelling, and protein design and ligand docking applications. Many of existing solutions are modelled as a computational optimisation problem. As well as the design of search algorithms, most solutions suffer from an inaccurate energy function for judging whether a prediction is good or bad. Even if the search has found the lowest energy, there is no certainty of obtaining the protein structures with correct side chains. METHODS: We present a side-chain modelling method, pacoPacker, which uses a parallel ant colony optimisation strategy based on sharing a single pheromone matrix. This parallel approach combines different sources of energy functions and generates protein side-chain conformations with the lowest energies jointly determined by the various energy functions. We further optimised the selected rotamers to construct subrotamer by rotamer minimisation, which reasonably improved the discreteness of the rotamer library. RESULTS: We focused on improving the accuracy of side-chain conformation prediction. For a testing set of 442 proteins, 87.19% of X1 and 77.11% of X12 angles were predicted correctly within 40° of the X-ray positions. We compared the accuracy of pacoPacker with state-of-the-art methods, such as CIS-RR and SCWRL4. We analysed the results from different perspectives, in terms of protein chain and individual residues. In this comprehensive benchmark testing, 51.5% of proteins within a length of 400 amino acids predicted by pacoPacker were superior to the results of CIS-RR and SCWRL4 simultaneously. Finally, we also showed the advantage of using the subrotamers strategy. All results confirmed that our parallel approach is competitive to state-of-the-art solutions for packing side chains. CONCLUSIONS: This parallel approach combines various sources of searching intelligence and energy functions to pack protein side chains. It provides a frame-work for combining different inaccuracy/usefulness objective functions by designing parallel heuristic search algorithms.
Lijun Quan, Haiou Li, Xiaoyan Xia, Hongjie Wu
BMC Bioinform.5
2013 A parallel ant colonies approach to de novo prediction of protein backbone in CASP8/9
Hongjie Wu, Jinzhen Wu, Xiaohu Luo, Peide Qian
Sci. China Inf. Sci.2