Jieyue He

dblp:25/4245 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0002-3265-3351ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 AGHINT: Attribute-guided representation learning on heterogeneous information networks with transformer
Jinhui Yuan, Shan Lu 0014, Peibo Duan, Jieyue He
Knowl. Based Syst.4
2024 ACDNet: Attention-guided Collaborative Decision Network for effective medication recommendation
Jiacong Mi, Yi Zu, Jieyue He
J. Biomed. Informatics4
2023 RoKEPG: RoBERTa and Knowledge Enhancement for Prescription Generation of Traditional Chinese Medicine
abstract
Traditional Chinese medicine (TCM) prescription is the most critical form of TCM treatment, and uncovering the complex nonlinear relationship between symptoms and TCM is of great significance for clinical practice and assisting physicians in diagnosis and treatment. Although there have been some studies on TCM prescription generation, these studies consider a single factor and directly model the symptom-prescription generation problem mainly based on symptom descriptions, lacking guidance from TCM knowledge. To this end, we propose a RoBERTa and Knowledge Enhancement model for Prescription Generation of Traditional Chinese Medicine (RoKEPG). RoKEPG is firstly pre-trained by our constructed TCM corpus, followed by fine-tuning the pre-trained model, and the model is guided to generate TCM prescriptions by introducing four classes of knowledge of TCM through the attention mask matrix. Experimental results on the publicly available TCM prescription dataset show that RoKEPG improves the F1metric by about 2% over the baseline model with the best results.
Hua Pu, Jiacong Mi, Shan Lu 0014, Jieyue He
BIBM4
2023 Finformer: A Static-dynamic Spatiotemporal Framework for Stock Trend Prediction
abstract
The core of quantitative investment lies in predicting future trends in stock prices. The future trend of a stock is closely related to the industry it belongs to and its relationship with other stocks. Although some research has focused on stock trend prediction in recent years, most studies have only considered the stock’s own time series feature, neglecting the spatial features between stocks. Some research has incorporated spatial information, but typically only considered predefined static relationships. At the same time, capturing dynamic spatial information in the market has been a long-standing challenge. Thus, we propose a spatio-temporal model, Finformer, in order to go beyond traditional time series models. We designed a sparse static-dynamic transformer to capture dynamic market spatial information as it changes over time and combined predefined relationships to extract highly correlated spatial features in the stock market. To effectively integrate spatial and temporal features, we introduced an adaptive spatio-temporal fusion module that dynamically fuses spatio-temporal features based on market conditions at different periods. Experiments on two real-world stock market datasets show that our proposed model outperforms the state-of-the-art baselines in the signal-based and portfolio-based metrics, which are widely concerned in the financial field. Ablation study and hyper-parameter study further reveal the effectiveness of each module in the model and the impact of hyper-parameters. The code will be made publicly available.1
Yi Zu, Jiacong Mi, Lingning Song, Shan Lu 0014, Jieyue He
IEEE Big Data5
2023 HARPA: hierarchical attention with relation paths for knowledge graph embedding adversarial learning
Naixin Zhang, Jinmeng Wang, Jieyue He
Data Min. Knowl. Discov.3
2022 AFGSL: Automatic Feature Generation based on Graph Structure Learning
Xin Xi, Jieyue He
Knowl. Based Syst.3
2021 KGAPG: Knowledge-Aware Neural Group Representation Learning for Attentive Prescription Generation of Traditional Chinese Medicine
abstract
Prescriptions play an essential role in the process of Traditional Chinese Medicine (TCM) diagnosis and treatment. Prescription generation is to generate a set of herbs to treat the symptoms of a patient by analyzing the relationship between symptoms and herbs. Although there have been a couple of studies to generate prescriptions, they have ignored the implicit relationship between the different symptoms of the patients. In addition, abundant semantic information and interpretability of the knowledge graph can help to portray the complicated relationships between the various modules of TCM. Therefore, this paper proposes a Knowledge-Aware neural Group representation learning model for Attentive Prescription Generation of Traditional Chinese Medicine (KGAPG), which regards the prescription generation task as a group recommendation problem. More specifically, multiple symptoms of a patient are considered as a symptom group and the complicated semantic information between symptoms and herbs can be captured by the knowledge graph. The syndrome information of multiple symptoms which is summarized by a group aggregation method based on the attention mechanism is applied to simulate the actual process of TCM diagnosis and treatment. The experiment results demonstrate that KGAPG is effective on a TCM prescription benchmark dataset, and its evaluation indicators of Precision, Recall and NDCG exceed other state-of-the-art methods.
Shuchen Li, Jieyue He
BIBM3
2021 HGNA-HTI: Heterogeneous graph neural network with attention mechanism for prediction of herb-target interactions
abstract
Herb-target interactions (HTIs) prediction plays a key role in exploring the mechanism of Traditional Chinese Medicine (TCM). There have been many studies on the prediction of drug-target interactions. However, these methods cannot be directly applied to TCM, which has the characteristics of multiple-component, multiple-target and multiple-pathway. Although network-based methods began to be used to study the prediction of HTIs, the aggregate information from the high-order neighborhood on the heterogeneous herb-target network is underutilized. Therefore, this paper proposes a Heterogeneous Graph Neural Network with Attention Mechanism for Prediction of Herb-Target Interactions (HGNA-HTI). Specifically, based on the heterogeneous herb-target graph, HGNA-HTI uses attention mechanism to give high attention values to important nodes and edges based on different types of meta relations, and applies message passing process to incorporate information from different types of links. Moreover, HGNA-HTI aggregates the above information to extract the semantic information and high-level structure of the heterogeneous herb-target graph to obtain the final feature representation to enhance the ability for prediction of HTIs. The experiment results on two datasets, show that the HGNA-HTI model is better than state-of-the-art approaches.
Jieyue He
BIBM3
2021 KGRN: Knowledge Graph Relational Path Network for Target Prediction of TCM Prescriptions
Zhuo Gong, Naixin Zhang, Jieyue He
ICIC (3)3
2021 Attention based adversarially regularized learning for network embedding
Jieyue He, Jinmeng Wang, Zhizhou Yu
Data Min. Knowl. Discov.1
2020 Hybrid attentional memory network for computational drug repositioning
abstract
BACKGROUND: Drug repositioning has been an important and efficient method for discovering new uses of known drugs. Researchers have been limited to one certain type of collaborative filtering (CF) models for drug repositioning, like the neighborhood based approaches which are good at mining the local information contained in few strong drug-disease associations, or the latent factor based models which are effectively capture the global information shared by a majority of drug-disease associations. Few researchers have combined these two types of CF models to derive a hybrid model which can offer the advantages of both. Besides, the cold start problem has always been a major challenge in the field of computational drug repositioning, which restricts the inference ability of relevant models. RESULTS: Inspired by the memory network, we propose the hybrid attentional memory network (HAMN) model, a deep architecture combining two classes of CF models in a nonlinear manner. First, the memory unit and the attention mechanism are combined to generate a neighborhood contribution representation to capture the local structure of few strong drug-disease associations. Then a variant version of the autoencoder is used to extract the latent factor of drugs and diseases to capture the overall information shared by a majority of drug-disease associations. During this process, ancillary information of drugs and diseases can help alleviate the cold start problem. Finally, in the prediction stage, the neighborhood contribution representation is coupled with the drug latent factor and disease latent factor to produce predicted values. Comprehensive experimental results on two data sets demonstrate that our proposed HAMN model outperforms other comparison models based on the AUC, AUPR and HR indicators. CONCLUSIONS: Through the performance on two drug repositioning data sets, we believe that the HAMN model proposes a new solution to improve the prediction accuracy of drug-disease associations and give pharmaceutical personnel a new perspective to develop new drugs.
Jieyue He, Xinxing Yang, Zhuo Gong, Ibrahim Zamit
BMC Bioinform.1
2019 Additional Neural Matrix Factorization model for computational drug repositioning
abstract
BACKGROUND: Computational drug repositioning, which aims to find new applications for existing drugs, is gaining more attention from the pharmaceutical companies due to its low attrition rate, reduced cost, and shorter timelines for novel drug discovery. Nowadays, a growing number of researchers are utilizing the concept of recommendation systems to answer the question of drug repositioning. Nevertheless, there still lie some challenges to be addressed: 1) Learning ability deficiencies; the adopted model cannot learn a higher level of drug-disease associations from the data. 2) Data sparseness limits the generalization ability of the model. 3)Model is easy to overfit if the effect of negative samples is not taken into consideration. RESULTS: In this study, we propose a novel method for computational drug repositioning, Additional Neural Matrix Factorization (ANMF). The ANMF model makes use of drug-drug similarities and disease-disease similarities to enhance the representation information of drugs and diseases in order to overcome the matter of data sparsity. By means of a variant version of the autoencoder, we were able to uncover the hidden features of both drugs and diseases. The extracted hidden features will then participate in a collaborative filtering process by incorporating the Generalized Matrix Factorization (GMF) method, which will ultimately give birth to a model with a stronger learning ability. Finally, negative sampling techniques are employed to strengthen the training set in order to minimize the likelihood of model overfitting. The experimental results on the Gottlieb and Cdataset datasets show that the performance of the ANMF model outperforms state-of-the-art methods. CONCLUSIONS: Through performance on two real-world datasets, we believe that the proposed model will certainly play a role in answering to the major challenge in drug repositioning, which lies in predicting and choosing new therapeutic indications to prospectively test for a drug of interest.
Xinxing Yang, Ibrahim Zamit, Jieyue He
BMC Bioinform.4
2013 Discovering frequent probability pattern in uncertain biological networks by circuit simulation method
abstract
In the field of bioinformatics, many types of data can be represented as the topological graph, such as protein-protein interaction network. Milo proposed the concept of biological motif 0 on Science, which is referred as a substructure that appears in different parts of a network, and appears significantly more frequently than in a random network. Research shows that the motif recognition is important for many biological studies. As the life process itself is a dynamic process, the motif of the same function may be made up of the subgraphs which may slightly differ in topology, so Berg etc. [2] proposed probability motif mining algorithms in the biological network. And science graph data are obtained with the inevitable experimental error or noise data, and some biological network data carries probability information. Since biological evolution itself is a mutant selection process, the input of biological networks should also be a probabilistic network. Therefore, it is more intuitively and practically significantly to mine probability motif in the probability biological network.
Chunyan Wang 0006, Kunpu Qiu, Jieyue He
BIBM4
2012 Efficient and accurate greedy search methods for mining functional modules in protein interaction networks
abstract
BACKGROUND: Most computational algorithms mainly focus on detecting highly connected subgraphs in PPI networks as protein complexes but ignore their inherent organization. Furthermore, many of these algorithms are computationally expensive. However, recent analysis indicates that experimentally detected protein complexes generally contain Core/attachment structures. METHODS: In this paper, a Greedy Search Method based on Core-Attachment structure (GSM-CA) is proposed. The GSM-CA method detects densely connected regions in large protein-protein interaction networks based on the edge weight and two criteria for determining core nodes and attachment nodes. The GSM-CA method improves the prediction accuracy compared to other similar module detection approaches, however it is computationally expensive. Many module detection approaches are based on the traditional hierarchical methods, which is also computationally inefficient because the hierarchical tree structure produced by these approaches cannot provide adequate information to identify whether a network belongs to a module structure or not. In order to speed up the computational process, the Greedy Search Method based on Fast Clustering (GSM-FC) is proposed in this work. The edge weight based GSM-FC method uses a greedy procedure to traverse all edges just once to separate the network into the suitable set of modules. RESULTS: The proposed methods are applied to the protein interaction network of S. cerevisiae. Experimental results indicate that many significant functional modules are detected, most of which match the known complexes. Results also demonstrate that the GSM-FC algorithm is faster and more accurate as compared to other competing algorithms. CONCLUSIONS: Based on the new edge weight definition, the proposed algorithm takes advantages of the greedy search procedure to separate the network into the suitable set of modules. Experimental analysis shows that the identified modules are statistically significant. The algorithm can reduce the computational time significantly while keeping high prediction accuracy.
Jieyue He
BMC Bioinform.1
2012 Clinical charge profiles prediction for patients diagnosed with chronic diseases using Multi-level Support Vector Machine
Rick Chow, Jieyue He
Expert Syst. Appl.3
2011 A Novel Core-Attachment Based Greedy Search Method for Mining Functional Modules in Protein Interaction Networks
Jieyue He
ISBRA2
2009 Tri-Cluster-Tri-Scheme-Training: Exploiting Unlabeled Data for Transmembrane Segments Prediction
abstract
Recent work using supervised learning for protein structure prediction has achieved state-of-the-art classification performance. However, such methods are based only on labeled data, while in practice the labeled data is so few and expensive to obtain and unlabeled data is far more plentiful. An effective way to enhance the performance of the learned hypothesis by using the labeled and unlabeled data together is known as semi-supervised learning. Although there are lots of semi-supervised learning methods, those approaches could not always achieve the acceptable results for bioinformatics application, especially when there is only very few labeled instances. Therefore, in this paper, we present a novel, more effective method tri-cluster-tri-scheme-training (TCTS) which firstly uses tri-cluster to label some high confidence unlabeled instances and then refines the classifiers by utilizing both of the label data and unlabeled data in the Tri-Scheme-training by different schemes. The encouraging experimental results indicate that TCTS algorithm opens a new way to solve the complex classification problem when very few labeled datasets are available.
Jieyue He, Robert W. Harrison, Phang C. Tai, Yi Pan 0001
BIBE1
2008 Protein Sequence Motif Super-Rule-Tree (SRT) Structure Constructed by Hybrid Hierarchical K-Means Clustering Algorithm
abstract
Protein sequence motifs information is crucial to the analysis of biologically significant regions. The conserved regions have the potential to determine the role of the proteins. Many algorithms or techniques to discover motifs require a predefined fixed window size in advance. Due to the fixed size, these approaches often deliver a number of similar motifs simply shifted by some bases or including mismatches. To confront the mismatched motifs problem, we use the super-rule concept to construct a Super-Rule-Tree (SRT) by a modified HHK clustering which requires no parameter setup to identify the similarities and dissimilarities between the motifs. By analyzing the motifs results generated by our approach, they are not only significant in sequence area but secondary structure similarity. We believe new proposed HHK clustering algorithm and SRT can play an important role in similar researches which requires predefined fixed window size.
Bernard Chen 0001, Jieyue He, Stephen Pellicer, Yi Pan 0001
BIBM2
2008 Hierarchical Clustering Support Vector Machines for Classifying Type-2 Diabetes Patients
Rick Chow, Richard Stolz, Jieyue He, Marsha Dowell
ISBRA4
2007 Multiclass Fuzzy Clustering Support Vector Machines for Protein Local Structure Prediction
abstract
Local protein structure prediction is a central task in bioinformatics research. Local protein structure prediction can be transformed into the multiclass problem for huge datasets. In previous study, multiclass clustering support vector machines (CSVMs) was proposed for local protein structure prediction. The greedy algorithm is utilized to select the next closest class if CSVM modeled for the assigned class predicts the sequence segment as negative. However, the greedy algorithm may not be optimal. If all CSVM predict the sequence segment as negative, this sequence segment cannot be classified. In order to further improve performance of the multiclass problem, we propose fuzzy clustering support vector machines (FCSVMs) in this study. The FCSVMs model calculates the class membership value of the given sequence segment for each class and assigns the representative structure of the finally selected class to the sequence segment. Values of the fuzzy membership function are based on testing accuracy of decision function outputs from FCSVMs. Under this mechanism, values of different fuzzy membership functions can be compared. FCSVMs are built specifically for each class partitioned intelligently by the clustering algorithm. This feature makes learning tasks for each FCSVM more specific and simpler. Furthermore, FCSVM modeled for each class can be easily parallelized to handle the complex multiclass problems for huge datasets. Using fuzzy membership functions, all sequence segments can be classified. Compared with the conventional clustering algorithm and CSVMs, testing accuracy for local structure prediction has been improved noticeably when the FCSVMs model is applied.
Jieyue He, Yi Pan 0001
BIBE2
2007 Mutual Information based Minimum Spanning Trees Model for Selecting Discriminative Genes
abstract
Recent studies have shown that gene selection is a crucial technology in microarray data analysis as a result of its large number of genes and relatively small number of samples. Filter methods are fast convergent algorithms with low time complexity. However, filter methods neglect correlation among genes. Other methods for gene selection also have disadvantages. For example, the measurement used to calculate the correlation in other methods can not effectively reflect function similarity among genes, the time complexity will be high based on the whole gene set. Therefore, we propose a novel selection model called mutual information based minimum spanning trees (MIMST) which considers both gene interaction and complementary genes. In this new model, we first use filter methods to remove non-relevant genes, and then compute the interdependence of top-ranked genes. Finally, we construct MST to remove the redundant genes. The experiment results show that MIMST can find the smallest siMIMSTgnificant genes subset with higher classification accuracy compared with other methods.
Jieyue He
BIBE2
2007 Clustering support vector machines for protein local structure prediction
Jieyue He, Robert W. Harrison, Phang C. Tai, Yi Pan 0001
Expert Syst. Appl.2
2006 Transmembrane segments prediction and understanding using support vector machine and decision tree
Jieyue He, Hae-Jin Hu, Robert W. Harrison, Phang C. Tai, Yi Pan 0001
Expert Syst. Appl.1