Hanjing Jiang

dblp:328/3609 · also Han-Jing Jiang · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0003-4208-2027ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Representation Learning Framework for CircRNA-Disease Association Prediction Based on Symmetric Convolutional Networks
Meineng Wang, Hanjing Jiang, Zhi-Xiao Wang, Lei Wang 0121
ICIC (30)2
2025 Surrogate-Assisted Multiobjective Gene Selection for Cell Classification From Large-Scale Single-Cell RNA Sequencing Data
abstract
Accurate cell classification is crucial but expensive for large-scale single-cell RNA sequencing (scRNA-seq) analysis. Gene selection (GS) emerges as a pivotal technique in identifying gene subsets of scRNA-seq for classification accuracy improvement and gene scale reduction. Nevertheless, the rising scale of scRNA-seq data presents challenges to existing GS methods regarding performance and computational time. Thus, we propose a surrogate-assisted evolutionary algorithm for multiobjective GS to address these deficiencies. An innovative two-phase initialization method is proposed to select sparse solutions to provide preliminary insights into gene contributions. Then, a binary competitive swarm optimizer is proposed for effective global search, where a local search method is embedded to eliminate irrelevant genes for efficiency consideration. Additionally, a surrogate model is adopted to forecast classification accuracy efficiently and substitutes part of the computationally expensive classification process. Experiments are conducted on eight large-scale scRNA-seq datasets with more than 20 000 genes. The effectiveness of the proposed GS method for scRNA-seq cell classification compared with eight state-of-the-art methods is validated. Gene expression analysis results of selected genes further validated the significance of the genes selected by the proposed method in the classification of scRNA-seq data.
Jianqing Lin, Cheng He 0001, Hanjing Jiang, Yabing Huang, Yaochu Jin
IEEE Trans. Evol. Comput.3
2024 A Single-Cell Clustering Algorithm Based on Structure Perturbation Non-Negative Matrix Factorization
abstract
The advent of single-cell RNA sequencing (scRNA-seq) has facilitated the acquisition of high-resolution data regarding cell heterogeneity across various tissues. A fundamental and critical step in the analysis of scRNA-seq data is cell type identification. An appropriate graph construction strategy can reflect the topological structure between cells and allows for detailed exploration of the relationships between cells and genes. Given the sparsity of scRNA-seq data and the pronounced noise generated by shallow sequencing, choosing an appropriate graph construction strategy has become a significant challenge. In this paper, we propose a single-cell clustering method, named Sc-PNNMF, based on structural perturbation of nonnegative matrix factorization. Sc-PNNMF employs a structural perturbation algorithm to optimize the gene expression matrix. The optimized matrix guides the matrix factorization process, effectively overcoming the impact of data sparsity and significant noise in gene expression data on the graph construction strategy. We compared the clustering abilities of Sc-PNNMF with five other methods on ten real datasets.
Hanjing Jiang, Meineng Wang, Yangyuan Li, Yabing Huang
BIBM1
2024 Graph-Regularized Non-Negative Matrix Factorization for Single-Cell Clustering in scRNA-Seq Data
abstract
The advent of single-cell RNA sequencing (scRNA-seq) has brought forth fresh perspectives on intricate biological processes, revealing the nuances and divergences present among distinct cells. Accurate single-cell analysis is a crucial prerequisite for in-depth investigation into the underlying mechanisms of heterogeneity. Due to various technical noises, like the impact of dropout values, scRNA-seq data remains challenging to interpret. In this work, we propose an unsupervised learning framework for scRNA-seq data analysis (aka Sc-GNNMF). Based on the non-negativity and sparsity of scRNA-seq data, we propose employing graph-regularized non-negative matrix factorization (GNNMF) algorithm for the analysis of scRNA-seq data, which involves estimating cell-cell sparse similarity and gene-gene sparse similarity through Laplacian kernels and p-nearest neighbor graphs ( p-NNG). By assuming intrinsic geometric local invariance, we use a weighted p-nearest known neighbors ( p-NKN) to optimize the scRNA-seq data. The optimized scRNA-seq data then participates in the matrix decomposition process, promoting the closeness of cells with similar types in cell-gene data space and determining a more suitable embedding space for clustering. Sc-GNNMF demonstrates superior performance compared to other methods and maintains satisfactory compatibility and robustness, as evidenced by experiments on 11 real scRNA-seq datasets. Furthermore, Sc-GNNMF yields excellent results in clustering tasks, extracting useful gene markers, and pseudo-temporal analysis.
Hanjing Jiang, Meineng Wang, Yabing Huang
IEEE J. Biomed. Health Informatics1
2023 Selection of Cancer Biomarkers from Microarray Gene Expression Data Utilizing the Bi-Objective Optimization Method
abstract
Cancer-associated biomarker genes play an indispensable role in the intricate tapestry of cancer development and manifestation. The expression of biomarkers in different types of tumor cells has beneficial implications for shedding light on the development of various cancers, guiding clinical diagnosis, and treatment. Microarray technology enables the expression levels of thousands of genes in samples to be sequenced simultaneously. However, sparse and high-dimensional microarray data present a formidable challenge in identifying biomarker genes. This study presents EREF-NSGA2, a novel method for cancer biomarker selection from microarray data, employing a hybrid gene selection approach. Firstly, the combination of the wrapper and embedded gene selection methods is proposed to filter the microarray data, which efficiently decreases the search space of the algorithm. After that, the improved NSGA-II algorithm is used to search the genes subset obtained from the previous step to reach the optimal subset of cancer biomarker genes. The proposed EREF-NSGA2 is compared with other reported methods on six cancer benchmark gene expression datasets. A detailed biological analysis is performed to analyze the relationship between the selected genes and the cancer data sets they belong to. To summarize, EREF-NSGA2 proves its effectiveness in selecting a feature subset comprising the fewest genes while maintaining the highest classification accuracy.
Hanjing Jiang, Jianqing Lin, Huachao Zhu, Yabing Huang
BIBM1
2023 ScLSTM: single-cell type detection by siamese recurrent network and hierarchical clustering
abstract
MOTIVATION: Categorizing cells into distinct types can shed light on biological tissue functions and interactions, and uncover specific mechanisms under pathological conditions. Since gene expression throughout a population of cells is averaged out by conventional sequencing techniques, it is challenging to distinguish between different cell types. The accumulation of single-cell RNA sequencing (scRNA-seq) data provides the foundation for a more precise classification of cell types. It is crucial building a high-accuracy clustering approach to categorize cell types since the imbalance of cell types and differences in the distribution of scRNA-seq data affect single-cell clustering and visualization outcomes. RESULT: To achieve single-cell type detection, we propose a meta-learning-based single-cell clustering model called ScLSTM. Specifically, ScLSTM transforms the single-cell type detection problem into a hierarchical classification problem based on feature extraction by the siamese long-short term memory (LSTM) network. The similarity matrix derived from the improved sigmoid kernel is mapped to the siamese LSTM feature space to analyze the differences between cells. ScLSTM demonstrated superior classification performance on 8 scRNA-seq data sets of different platforms, species, and tissues. Further quantitative analysis and visualization of the human breast cancer data set validated the superiority and capability of ScLSTM in recognizing cell types.
Hanjing Jiang, Yabing Huang, Qianpeng Li, Boyuan Feng
BMC Bioinform.1
2022 Spectral clustering of single cells using Siamese nerual network combined with improved affinity matrix
abstract
Limitations of bulk sequencing techniques on cell heterogeneity and diversity analysis have been pushed with the development of single-cell RNA-sequencing (scRNA-seq). To detect clusters of cells is a key step in the analysis of scRNA-seq. However, the high-dimensionality of scRNA-seq data and the imbalances in the number of different subcellular types are ubiquitous in real scRNA-seq data sets, which poses a huge challenge to the single-cell-type detection.We propose a meta-learning-based model, SiaClust, which is the combination of Siamese Convolutional Neural Network (CNN) and improved spectral clustering, to achieve scRNA-seq cell type detection. To be specific, with the help of the constrained Sigmoid kernel, the raw high-dimensionality data is mapped to a low-dimensional space, and the Siamese CNN learns the differences between the cell types in the low-dimensional feature space. The similarity matrix learned by Siamese CNN is used in combination with improved spectral clustering and t-distribution Stochastic Neighbor Embedding (t-SNE) for visualization. SiaClust highlights the differences between cell types by comparing the similarity of the samples, whereas blurring the differences within the cell types is better in processing high-dimensional and imbalanced data. SiaClust significantly improves clustering accuracy by using data generated by nine different species and tissues through different scNA-seq protocols for extensive evaluation, as well as analogies to state-of-the-art single-cell clustering models. More importantly, SiaClust accurately locates the exact site of dropout gene, and is more flexible with data size and cell type.
Hanjing Jiang, Yabing Huang, Qianpeng Li
Briefings Bioinform.1
2022 An effective drug-disease associations prediction model based on graphic representation learning over multi-biomolecular network
abstract
BACKGROUND: Drug-disease associations (DDAs) can provide important information for exploring the potential efficacy of drugs. However, up to now, there are still few DDAs verified by experiments. Previous evidence indicates that the combination of information would be conducive to the discovery of new DDAs. How to integrate different biological data sources and identify the most effective drugs for a certain disease based on drug-disease coupled mechanisms is still a challenging problem. RESULTS: In this paper, we proposed a novel computation model for DDA predictions based on graph representation learning over multi-biomolecular network (GRLMN). More specifically, we firstly constructed a large-scale molecular association network (MAN) by integrating the associations among drugs, diseases, proteins, miRNAs, and lncRNAs. Then, a graph embedding model was used to learn vector representations for all drugs and diseases in MAN. Finally, the combined features were fed to a random forest (RF) model to predict new DDAs. The proposed model was evaluated on the SCMFDD-S data set using five-fold cross-validation. Experiment results showed that GRLMN model was very accurate with the area under the ROC curve (AUC) of 87.9%, which outperformed all previous works in terms of both accuracy and AUC in benchmark dataset. To further verify the high performance of GRLMN, we carried out two case studies for two common diseases. As a result, in the ranking of drugs that were predicted to be related to certain diseases (such as kidney disease and fever), 15 of the top 20 drugs have been experimentally confirmed. CONCLUSIONS: The experimental results show that our model has good performance in the prediction of DDA. GRLMN is an effective prioritization tool for screening the reliable DDAs for follow-up studies concerning their participation in drug reposition.
Hanjing Jiang, Yabing Huang
BMC Bioinform.1
2021 Prognostic Prediction for Non-small-Cell Lung Cancer Based on Deep Neural Network and Multimodal Data
Zhongsi Zhang, Hanjing Jiang
ICIC (3)3
2020 A Highly Efficient Biomolecular Network Representation Model for Predicting Drug-Disease Associations
Hanjing Jiang, Zhu-Hong You, Lun Hu, Zhen-Hao Guo, Leon Wong
ICIC (3)1
2019 Combining LSTM Network Model and Wavelet Transform for Predicting Self-interacting Proteins
Zhu-Hong You, Liping Li 0003, Zhen-Hao Guo, Pengwei Hu 0001, Hanjing Jiang
ICIC (1)6
2019 Predicting of Drug-Disease Associations via Sparse Auto-Encoder-Based Rotation Forest
Hanjing Jiang, Zhu-Hong You, Kai Zheng 0020
ICIC (3)1
2019 MISSIM: Improved miRNA-Disease Association Prediction Model Based on Chaos Game Representation and Broad Learning System
Kai Zheng 0020, Zhu-Hong You, Lei Wang 0121, Hanjing Jiang
ICIC (3)6