Yabing Huang

dblp:20/8319 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Surrogate-Assisted Multiobjective Gene Selection for Cell Classification From Large-Scale Single-Cell RNA Sequencing Data
abstract
Accurate cell classification is crucial but expensive for large-scale single-cell RNA sequencing (scRNA-seq) analysis. Gene selection (GS) emerges as a pivotal technique in identifying gene subsets of scRNA-seq for classification accuracy improvement and gene scale reduction. Nevertheless, the rising scale of scRNA-seq data presents challenges to existing GS methods regarding performance and computational time. Thus, we propose a surrogate-assisted evolutionary algorithm for multiobjective GS to address these deficiencies. An innovative two-phase initialization method is proposed to select sparse solutions to provide preliminary insights into gene contributions. Then, a binary competitive swarm optimizer is proposed for effective global search, where a local search method is embedded to eliminate irrelevant genes for efficiency consideration. Additionally, a surrogate model is adopted to forecast classification accuracy efficiently and substitutes part of the computationally expensive classification process. Experiments are conducted on eight large-scale scRNA-seq datasets with more than 20 000 genes. The effectiveness of the proposed GS method for scRNA-seq cell classification compared with eight state-of-the-art methods is validated. Gene expression analysis results of selected genes further validated the significance of the genes selected by the proposed method in the classification of scRNA-seq data.
Jianqing Lin, Cheng He 0001, Hanjing Jiang, Yabing Huang, Yaochu Jin
IEEE Trans. Evol. Comput.4
2024 A Single-Cell Clustering Algorithm Based on Structure Perturbation Non-Negative Matrix Factorization
abstract
The advent of single-cell RNA sequencing (scRNA-seq) has facilitated the acquisition of high-resolution data regarding cell heterogeneity across various tissues. A fundamental and critical step in the analysis of scRNA-seq data is cell type identification. An appropriate graph construction strategy can reflect the topological structure between cells and allows for detailed exploration of the relationships between cells and genes. Given the sparsity of scRNA-seq data and the pronounced noise generated by shallow sequencing, choosing an appropriate graph construction strategy has become a significant challenge. In this paper, we propose a single-cell clustering method, named Sc-PNNMF, based on structural perturbation of nonnegative matrix factorization. Sc-PNNMF employs a structural perturbation algorithm to optimize the gene expression matrix. The optimized matrix guides the matrix factorization process, effectively overcoming the impact of data sparsity and significant noise in gene expression data on the graph construction strategy. We compared the clustering abilities of Sc-PNNMF with five other methods on ten real datasets.
Hanjing Jiang, Meineng Wang, Yangyuan Li, Yabing Huang
BIBM6
2024 Invariance Principles for Nonlinear Discrete-Time Switched Systems and Its Application to Output Synchronization of Dynamical Networks
abstract
In this article, we develop two invariance principles for nonlinear discrete-time switched systems based on multiple Lyapunov functions and multiple weak Lyapunov functions, respectively, which allow the first differences of multiple weak Lyapunov functions to be positive on certain sets. It is shown that the solution of the system is attracted to the largest weakly invariant set in a certain specific region. Then, based on the invariance principle developed and geometrical dissipativity, we obtain the generalized output synchronization for discrete-time dynamical networks with nonidentical nodes by an appropriate switching among several communication topologies. Finally, two examples are provided to demonstrate the effectiveness of the main results.
Jun Fu 0001, Chensong Li, Yabing Huang, Yuzhe Li 0003, Tianyou Chai
IEEE Trans. Cybern.3
2024 Active Interdiction Defence Scheme Against False Data-Injection Attacks: A Stackelberg Game Perspective
abstract
The security issue is of the fundamental importance for cyber-physical systems (CPSs) due to the vulnerability to cyber attacks. In this article, an active interdiction defence scheme is proposed as an extra safeguard against the strategic false data-injection (FDI) attacker. To tackle the coupling between the decision-making processes of the FDI attacker and the active interdiction defence scheme, a framework of dynamic Stackelberg game is formulated to characterize the hierarchical interaction due to their asymmetric information structures. Subsequently, the Stackelberg strategies of both players are constructed. Furthermore, a sufficient condition is derived to guarantee asymptotic stability of the closed-loop system with the FDI attack and the active interdiction defence scheme. Finally, two simulation examples are provided to illustrate the correctness and effectiveness of the proposed active interdiction defence scheme.
Yabing Huang, Jun Zhao 0002
IEEE Trans. Cybern.1
2024 Graph-Regularized Non-Negative Matrix Factorization for Single-Cell Clustering in scRNA-Seq Data
abstract
The advent of single-cell RNA sequencing (scRNA-seq) has brought forth fresh perspectives on intricate biological processes, revealing the nuances and divergences present among distinct cells. Accurate single-cell analysis is a crucial prerequisite for in-depth investigation into the underlying mechanisms of heterogeneity. Due to various technical noises, like the impact of dropout values, scRNA-seq data remains challenging to interpret. In this work, we propose an unsupervised learning framework for scRNA-seq data analysis (aka Sc-GNNMF). Based on the non-negativity and sparsity of scRNA-seq data, we propose employing graph-regularized non-negative matrix factorization (GNNMF) algorithm for the analysis of scRNA-seq data, which involves estimating cell-cell sparse similarity and gene-gene sparse similarity through Laplacian kernels and p-nearest neighbor graphs ( p-NNG). By assuming intrinsic geometric local invariance, we use a weighted p-nearest known neighbors ( p-NKN) to optimize the scRNA-seq data. The optimized scRNA-seq data then participates in the matrix decomposition process, promoting the closeness of cells with similar types in cell-gene data space and determining a more suitable embedding space for clustering. Sc-GNNMF demonstrates superior performance compared to other methods and maintains satisfactory compatibility and robustness, as evidenced by experiments on 11 real scRNA-seq datasets. Furthermore, Sc-GNNMF yields excellent results in clustering tasks, extracting useful gene markers, and pseudo-temporal analysis.
Hanjing Jiang, Meineng Wang, Yabing Huang
IEEE J. Biomed. Health Informatics4
2023 Selection of Cancer Biomarkers from Microarray Gene Expression Data Utilizing the Bi-Objective Optimization Method
abstract
Cancer-associated biomarker genes play an indispensable role in the intricate tapestry of cancer development and manifestation. The expression of biomarkers in different types of tumor cells has beneficial implications for shedding light on the development of various cancers, guiding clinical diagnosis, and treatment. Microarray technology enables the expression levels of thousands of genes in samples to be sequenced simultaneously. However, sparse and high-dimensional microarray data present a formidable challenge in identifying biomarker genes. This study presents EREF-NSGA2, a novel method for cancer biomarker selection from microarray data, employing a hybrid gene selection approach. Firstly, the combination of the wrapper and embedded gene selection methods is proposed to filter the microarray data, which efficiently decreases the search space of the algorithm. After that, the improved NSGA-II algorithm is used to search the genes subset obtained from the previous step to reach the optimal subset of cancer biomarker genes. The proposed EREF-NSGA2 is compared with other reported methods on six cancer benchmark gene expression datasets. A detailed biological analysis is performed to analyze the relationship between the selected genes and the cancer data sets they belong to. To summarize, EREF-NSGA2 proves its effectiveness in selecting a feature subset comprising the fewest genes while maintaining the highest classification accuracy.
Hanjing Jiang, Jianqing Lin, Huachao Zhu, Yabing Huang
BIBM4
2023 ScLSTM: single-cell type detection by siamese recurrent network and hierarchical clustering
abstract
MOTIVATION: Categorizing cells into distinct types can shed light on biological tissue functions and interactions, and uncover specific mechanisms under pathological conditions. Since gene expression throughout a population of cells is averaged out by conventional sequencing techniques, it is challenging to distinguish between different cell types. The accumulation of single-cell RNA sequencing (scRNA-seq) data provides the foundation for a more precise classification of cell types. It is crucial building a high-accuracy clustering approach to categorize cell types since the imbalance of cell types and differences in the distribution of scRNA-seq data affect single-cell clustering and visualization outcomes. RESULT: To achieve single-cell type detection, we propose a meta-learning-based single-cell clustering model called ScLSTM. Specifically, ScLSTM transforms the single-cell type detection problem into a hierarchical classification problem based on feature extraction by the siamese long-short term memory (LSTM) network. The similarity matrix derived from the improved sigmoid kernel is mapped to the siamese LSTM feature space to analyze the differences between cells. ScLSTM demonstrated superior classification performance on 8 scRNA-seq data sets of different platforms, species, and tissues. Further quantitative analysis and visualization of the human breast cancer data set validated the superiority and capability of ScLSTM in recognizing cell types.
Hanjing Jiang, Yabing Huang, Qianpeng Li, Boyuan Feng
BMC Bioinform.2
2023 A Block Cipher Algorithm Identification Scheme Based on Hybrid Random Forest and Logistic Regression Model
Yabing Huang, Chunfu Jia, Daoming Yu
Neural Process. Lett.2
2022 Spectral clustering of single cells using Siamese nerual network combined with improved affinity matrix
abstract
Limitations of bulk sequencing techniques on cell heterogeneity and diversity analysis have been pushed with the development of single-cell RNA-sequencing (scRNA-seq). To detect clusters of cells is a key step in the analysis of scRNA-seq. However, the high-dimensionality of scRNA-seq data and the imbalances in the number of different subcellular types are ubiquitous in real scRNA-seq data sets, which poses a huge challenge to the single-cell-type detection.We propose a meta-learning-based model, SiaClust, which is the combination of Siamese Convolutional Neural Network (CNN) and improved spectral clustering, to achieve scRNA-seq cell type detection. To be specific, with the help of the constrained Sigmoid kernel, the raw high-dimensionality data is mapped to a low-dimensional space, and the Siamese CNN learns the differences between the cell types in the low-dimensional feature space. The similarity matrix learned by Siamese CNN is used in combination with improved spectral clustering and t-distribution Stochastic Neighbor Embedding (t-SNE) for visualization. SiaClust highlights the differences between cell types by comparing the similarity of the samples, whereas blurring the differences within the cell types is better in processing high-dimensional and imbalanced data. SiaClust significantly improves clustering accuracy by using data generated by nine different species and tissues through different scNA-seq protocols for extensive evaluation, as well as analogies to state-of-the-art single-cell clustering models. More importantly, SiaClust accurately locates the exact site of dropout gene, and is more flexible with data size and cell type.
Hanjing Jiang, Yabing Huang, Qianpeng Li
Briefings Bioinform.2
2022 An effective drug-disease associations prediction model based on graphic representation learning over multi-biomolecular network
abstract
BACKGROUND: Drug-disease associations (DDAs) can provide important information for exploring the potential efficacy of drugs. However, up to now, there are still few DDAs verified by experiments. Previous evidence indicates that the combination of information would be conducive to the discovery of new DDAs. How to integrate different biological data sources and identify the most effective drugs for a certain disease based on drug-disease coupled mechanisms is still a challenging problem. RESULTS: In this paper, we proposed a novel computation model for DDA predictions based on graph representation learning over multi-biomolecular network (GRLMN). More specifically, we firstly constructed a large-scale molecular association network (MAN) by integrating the associations among drugs, diseases, proteins, miRNAs, and lncRNAs. Then, a graph embedding model was used to learn vector representations for all drugs and diseases in MAN. Finally, the combined features were fed to a random forest (RF) model to predict new DDAs. The proposed model was evaluated on the SCMFDD-S data set using five-fold cross-validation. Experiment results showed that GRLMN model was very accurate with the area under the ROC curve (AUC) of 87.9%, which outperformed all previous works in terms of both accuracy and AUC in benchmark dataset. To further verify the high performance of GRLMN, we carried out two case studies for two common diseases. As a result, in the ranking of drugs that were predicted to be related to certain diseases (such as kidney disease and fever), 15 of the top 20 drugs have been experimentally confirmed. CONCLUSIONS: The experimental results show that our model has good performance in the prediction of DDA. GRLMN is an effective prioritization tool for screening the reliable DDAs for follow-up studies concerning their participation in drug reposition.
Hanjing Jiang, Yabing Huang
BMC Bioinform.2
2021 Cyber-Physical Systems With Multiple Denial-of-Service Attackers: A Game-Theoretic Framework
abstract
Cyber-physical systems (CPSs) under denial-of-service (DoS) attacks are investigated in the present work. Besides multiple controllers in the physical-layer, there exist multiple DoS attackers in the cyber-layer congesting communication channels simultaneously, which deteriorates the control performance of the physical process. To tackle this coupling decision making process, a game-theoretic framework is established to characterize the interaction among the multiple controllers and the multiple DoS attackers. Specifically, a static cooperative game is formulated for the collaboration among the multiple DoS attackers in the cyber-layer, while a noncooperative differential game for the competition among the multiple controllers in the physical world, and the probabilistic data packet dropout determined by the coordination of DoS attackers plays as the bond connecting these two games. Subsequently, a selection scheme based on Pareto front is proposed to achieve an agreement of all DoS attackers without any preference information, and robust feedback Nash equilibrium strategies with respect to environmental disturbances and packet dropouts are constructed for all controllers. Finally, a quadruple-tank system is employed to demonstrate the effectiveness and applicability of the proposed results.
Yabing Huang, Jun Zhao 0002
IEEE Trans. Circuits Syst. I Regul. Pap.1