VLDB 2026 Research / reviewers in the wild / expert
Yong Liang 0001
dblp:l/YongLiang1
· DBLP profile ↗
48ranked-venue papers
9as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 2 first-author · 16 since 2021Artificial intelligence and machine learning · 22 · 7 first-author · 9 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DADA-EV: domain-adaptive diffusion autoencoder for estimating tissue- and cell-type-specific origin in extracellular vesicle transcriptomesabstractTracing the tissue and cell-type origins of extracellular vesicles (EVs) in blood is critical for liquid biopsy and precision medicine, yet existing deconvolution methods remain limited by the need for labor-intensive reference signatures and poor adaptability to distribution shifts between tissue/cell-type datasets and EV transcriptomes. We introduce DADA-EV (Domain-Adaptive Diffusion Autoencoder for EVs), a hybrid deep learning framework that combines an autoencoder backbone with a generative simulation module and adversarial domain adaptation. DADA-EV features three key innovations: (1) a reference-free design that eliminates reliance on predefined signatures; (2) cross-domain generalization by aligning feature distributions between source (tissue/cell-type) and target (EV) domain; and (3) reduced dependence on source data during target-domain training. Extensive evaluations on pseudo-EV data show that DADA-EV consistently outperforms existing approaches, yielding accurate fraction estimates across diverse tissues and gene sets. Validation using in vitro cell-line mixtures further confirms its reliability in resolving complex compositions, demonstrating high sensitivity in detecting low-abundance targets. Applied to real EV transcriptomes, it reveals tissue- and cell-type heterogeneity across patient groups. In summary, DADA-EV provides a robust, reference-free, and generalizable solution for EV origin tracing, with strong potential to advance diagnosis, prognosis, and treatment monitoring via liquid biopsy. Shuilin Liao, Haoxiang Yang, Shuting Xiao, Shanghui Lu, Yong Liang 0001 |
Briefings Bioinform. | 8 |
| 2026 | scMVAF: a multi-view adaptive fusion clustering approach for single-cell RNA-sequencing dataabstractSingle-cell RNA-sequencing (scRNA-seq) can excavate cellular heterogeneity and distinguish different types of cells. Clustering cells into subpopulations is essential in analyzing scRNA-seq data as it can help subsequent downstream analysis. However, scRNA-seq data are high-dimensional, sparse, and contain erroneous zero counts, which poses a great challenge for clustering. Although various methods have emerged in recent years, they cannot fully grasp the information of cells by characterizing scRNA-seq data from a single perspective, resulting in poor learned embedding representation and poor clustering performance. In this paper, we propose a multi-view clustering framework scMVAF for scRNA-seq data, which can learn more discriminative embedding representations by integrating feature information from multiple cell views. First, to comprehensively capture the data information, we generate multiple diverse views by down-sampling features, and then scMVAF learns a strong embedding representation for each cell view using an autoencoder based on a denoising zero-inflated negative binomial model. Next, to explore the correlation between cells in different views, a multi-view fusion module is introduced to fuse the embeddings from different views into a unified feature space. Concurrently, the fused embeddings are clustered to generate pseudo labels to improve the embedding process, and finally updating the embedding features and pseudo-labels in turn to obtain better clustering performance. Experiments are implemented on 16 real datasets and verify that scMVAF is superior to the other eight advanced technologies. Our code script can be obtained at https://github.com/LQXLE/scMVAF/. Jinfeng Wang 0003, Qixiong Long, Deyu Tang, Jin Deng, Yong Liang 0001 |
Briefings Bioinform. | 5 |
| 2026 | CALNet: an align-then-integrate architecture for navigating the entanglement-bottleneck dilemma in DTI prediction
Dingkui Kang, Yong Liang 0001, Hai-Hui Huang |
Expert Syst. Appl. | 3 |
| 2025 | ADM: adaptive graph diffusion for meta-dimension reductionabstractDimension reduction is essential for analyzing high-dimensional data, with various techniques developed to address diverse data characteristics. However, individual methods often struggle to capture all intricate patterns and complex structures simultaneously. To overcome this limitation, we introduce ADM (Adaptive graph Diffusion for Meta-dimension reduction), a novel meta-dimension reduction method grounded in graph diffusion theory. ADM integrates results from multiple dimension reduction techniques, leveraging their individual strengths while mitigating their specific weaknesses.ADM utilizes dynamic Markov processes to transform Euclidean space results into an information space, revealing intrinsic nonlinear manifold structures that are hard to capture by conventional methods. A critical advancement in ADM is its adaptive diffusion mechanism, which dynamically selects optimal diffusion time scales for each sample, enabling effective representation of multi-scale structures. This approach generates robust, high-quality low-dimensional representations that capture both local and global data structures while reducing noise and technique-specific distortions. We demonstrate ADM's efficacy on simulated and real-world datasets, including various omics data types. Results show that ADM provides clearer separation between biological groups and reveals more meaningful patterns compared to existing methods, advancing the analysis and visualization of complex biological data. Junning Feng 0001, Yong Liang 0001, Tianwei Yu |
Briefings Bioinform. | 2 |
| 2025 | A Simplified Input Strategy for Predicting Multi-Type Associations in miRNA-LncRNA-Disease Network via Stacked Deep Matrix FactorizationabstractUnderstanding the associations among microRNAs (miRNAs), long non-coding RNAs (lncRNAs), and various diseases as biomarkers holds significant biological importance. Developing efficient, straightforward prediction models is essential to reduce the high cost of experimental research. However, most existing methods typically predict miRNA-disease associations (MDAs), lncRNA-disease associations (LDAs), and lncRNA miRNA interactions (LMIs) separately, often relying on both similarities and associations as inputs. These approaches complicate their application across diverse biological and medical domains. Moreover, few models are capable of simultaneously predicting all three types of associations in a unified framework. In this work, we propose a novel and simplified model, called Simplified input strategy for Multiple Associations Prediction (SimpleMAP). Unlike previous approaches, SimpleMAP eliminates the need for similarity networks or external biological data and instead uses only known associations as input, reducing feature contamination and ensuring better generalization. SimpleMAP is designed to predict MDAs, LDAs, and LMIs concurrently, by constructing a three-layer heterogeneous biomolecular network that captures the associations among miRNAs, lncRNAs, and diseases. Our method employs a single, end-to-end architecture based on stacked deep matrix factorization (SDMF) to process sparse input data and learn latent features effectively. SimpleMAP is designed to concurrently predict MDAs, LDAs, and LMIs by constructing a three-layer heterogeneous biomolecular network that captures multi-relational associations among miRNAs, lncRNAs, and diseases. To enhance predictive performance, we incorporate multiple feature integration strategies to fuse representations extracted by SDMF. This streamlined design makes SimpleMAP one of the first models to predict multiple bio-entity associations jointly using only minimal input data, offering a highly scalable and biologically meaningful solution. SimpleMAP demonstrates superior performance against strong baselines. Further validation on two additional datasets involving miRNA-circRNA-disease associations confirms the models robustness and adaptability. Finally, biologically validated case studies underscore the realworld applicability of SimpleMAP for biomarker discovery in complex biological systems. Overall, SimpleMAP introduces a new paradigm in bio-entity association predictionłachieving multi-type, high-performance prediction with minimal input complexityłmaking it a valuable tool for computational biology and biomedical research. Ning Ai, Zhonghua Lu, Yong Liang 0001, Qi Hong Lai, Loi Lei Lai, Hongmin Cai, Dong Ouyang |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | Swin-UMamba†: Adapting Mamba-Based Vision Foundation Models for Medical Image SegmentationabstractVision foundation models have shown great potential in improving generalizability and data efficiency, especially for medical image segmentation since medical image datasets are relatively small due to high annotation costs and privacy concerns. However, current research on foundation models predominantly relies on transformers. The high quadratic complexity and large parameter counts make these models computationally expensive, limiting their potential for clinical applications. In this work, we introduce Swin-UMamba†, a novel Mamba-based model for medical image segmentation that seamlessly leverages the power of the vision foundation model, which is also computationally efficient with the linear complexity of Mamba. Moreover, we investigated and verified the impact of the vision foundation model on medical image segmentation, in which a self-supervised model adaptation scheme was designed to bridge the gap between natural and medical data. Notably, Swin-UMamba† outperforms 7 state-of-the-art methods, including CNN-based, transformer-based, and Mamba-based approaches across AbdomenMRI, Encoscopy, and Microscopy datasets. The code and models are publicly available at: https://github.com/JiarunLiu/Swin-UMamba. Jiarun Liu, Hao Yang 0026, Lequan Yu, Yong Liang 0001, Yizhou Yu, Shaoting Zhang 0001, Hairong Zheng, Shanshan Wang 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Swin-UMamba: Mamba-Based UNet with ImageNet-Based Pretraining
Jiarun Liu, Hao Yang 0026, Yan Xi, Lequan Yu, Cheng Li 0008, Yong Liang 0001, Guangming Shi, Yizhou Yu, Shaoting Zhang 0001, Hairong Zheng, Shanshan Wang 0002 |
MICCAI (9) | 7 |
| 2024 | Contention With Collision Detection in Wireless Full-Duplex NetworksabstractConventional wireless networks are half-duplex and most of them use contention-based protocols. These protocols usually adopt a principle of contention with collision avoidance and infer a collision occurrence very late from the absence of an acknowledgment after data transmission, causing low network performance. Wireless full-duplex (FD) enables simultaneous transmission (TX) and reception (RX) on the same channel. Exploiting this functionality, this article proposes the first design that enables contention with collision detection (CCD) to improve the network performance. We call the proposed design FD-CCD. With FD-CCD, in contention, a node exploits the TX antenna to transmit a signal for channel contention, while exploiting the RX antenna to sense if other nodes are transmitting too. By checking the status of the TX and RX antennas, the node can detect the contention collision before data transmission and, hence, obtain an opportunity to avoid the data collision effectively. FD-CCD also supports priority-based contentions, is of very low contention overhead, and is compatible with conventional 802.11 networks. This article then develops a theoretical model to analyze the system performance and optimize protocol parameter settings. Extensive simulations verify the effectiveness of our design and the accuracy of our model. This study is very helpful in designing efficient FD protocols. Qinglin Zhao, Fangxin Xu, Lian Zhao, Li Feng 0001, Yong Liang 0001 |
IEEE Internet Things J. | 6 |
| 2024 | PMFN-SSL: Self-supervised learning-based progressive multimodal fusion network for cancer diagnosis and prognosis
Hudan Pan, Yong Liang 0001, Ming-Wen Shao, Shengli Xie 0001, Shanghui Lu, Shuilin Liao |
Knowl. Based Syst. | 3 |
| 2024 | Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
Weijian Huang, Cheng Li 0008, Hao Yang 0026, Jiarun Liu, Yong Liang 0001, Hairong Zheng, Shanshan Wang 0010 |
Medical Image Anal. | 5 |
| 2024 | HGCLAMIR: Hypergraph contrastive learning with attention mechanism and integrated multi-view representation for predicting miRNA-disease associationsabstractExisting studies have shown that the abnormal expression of microRNAs (miRNAs) usually leads to the occurrence and development of human diseases. Identifying disease-related miRNAs contributes to studying the pathogenesis of diseases at the molecular level. As traditional biological experiments are time-consuming and expensive, computational methods have been used as an effective complement to infer the potential associations between miRNAs and diseases. However, most of the existing computational methods still face three main challenges: (i) learning of high-order relations; (ii) insufficient representation learning ability; (iii) importance learning and integration of multi-view embedding representation. To this end, we developed a HyperGraph Contrastive Learning with view-aware Attention Mechanism and Integrated multi-view Representation (HGCLAMIR) model to discover potential miRNA-disease associations. First, hypergraph convolutional network (HGCN) was utilized to capture high-order complex relations from hypergraphs related to miRNAs and diseases. Then, we combined HGCN with contrastive learning to improve and enhance the embedded representation learning ability of HGCN. Moreover, we introduced view-aware attention mechanism to adaptively weight the embedded representations of different views, thereby obtaining the importance of multi-view latent representations. Next, we innovatively proposed integrated representation learning to integrate the embedded representation information of multiple views for obtaining more reasonable embedding information. Finally, the integrated representation information was fed into a neural network-based matrix completion method to perform miRNA-disease association prediction. Experimental results on the cross-validation set and independent test set indicated that HGCLAMIR can achieve better prediction performance than other baseline models. Furthermore, the results of case studies and enrichment analysis further demonstrated the accuracy of HGCLAMIR and unconfirmed potential associations had biological significance. Dong Ouyang, Yong Liang 0001, Ning Ai, Junning Feng 0001, Shanghui Lu, Shuilin Liao, Xiao-Ying Liu 0003, Shengli Xie 0001 |
PLoS Comput. Biol. | 2 |
| 2024 | Multi-View Multiattention Graph Learning With Stack Deep Matrix Factorization for circRNA-Drug Sensitivity Association IdentificationabstractIdentifying circular RNA (circRNA)-drug sensitivity association (CDsA) is crucial for advancing drug development. As conducting traditional wet experiments for determining CDsA is costly and inefficient, calculation methods have already proven to be a valid approach to cope with this problem. However, there exists limited research addressing the prediction of the CDsA prediction problem, and certain discrepancies persist, particularly concerning false-negative associations. As a consequence, we present a multi-view framework, called MAGSDMF, for identifying latent CDsA. Firstly, MAGSDMF applies ultiple ttention mechanisms and raph learning methods to dynamically extract features and strengthen the features of inside and across multi-similarity networks of circRNA and drug. Secondly, the tack eep atrix Factorization (SDMF) is devised to directly extract features from CDsAs. We consider multi-similarity networks with the original CDsAs as multi-view information. Thirdly, MAGSDMF utilizes a multi-attention channel mechanism to integrate these features for the purpose of reconstructing CDsA. Finally, MAGSDMF performs another DMF based on the reconstruction to identify the latent CDsAs. Simultaneously, contrastive learning (CL) is implemented to enhance the generalization capability of MAGSDMF and oversee the learning process of the underlying links prediction task. In comparative experiments, MAGSDMF achieves superior performance on two datasets with AUC values of 0.9743 and 0.9739 based on 5-fold cross-validation. Moreover, in case studies, the achievements further validate the identification reliability of MAGSDMF. Ning Ai, Yong Liang 0001, Shanghui Lu, Dong Ouyang, Qi Hong Lai, Loi Lei Lai |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | MUMA: A Multi-Omics Meta-Learning Algorithm for Data Interpretation and ClassificationabstractMulti-omics data integration is a promising field combining various types of omics data, such as genomics, transcriptomics, and proteomics, to comprehensively understand the molecular mechanisms underlying life and disease. However, the inherent noise, heterogeneity, and high dimensionality of multi-omics data present challenges for existing methods to extract meaningful biological information without overfitting. This paper introduces a novel Multi-Omics Meta-learning Algorithm (MUMA) that employs self-adaptive sample weighting and interaction-based regularization for enhanced diagnostic performance and interpretability in multi-omics data analysis. Specifically, MUMA captures crucial biological processes across different omics layers by learning a flexible sample reweighting function adaptable to various noise scenarios. Additionally, MUMA incorporates an interaction-based regularization term, encouraging the model to learn from the relationships among different omics modalities. We evaluate MUMA using simulations and eighteen real datasets, demonstrating its superior performance compared to state-of-the-art methods in classifying biological samples (e.g., cancer subtypes) and selecting relevant biomarkers from noisy multi-omics data. As a powerful tool for multi-omics data integration, MUMA can assist researchers in achieving a deeper understanding of the biological systems involved. The source code for MUMA is available at https://github.com/bio-ai-source/MUMA. Hai-Hui Huang, Yong Liang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | RCDNet: An Interpretable Rain Convolutional Dictionary Network for Single Image DerainingabstractAs common weather, rain streaks adversely degrade the image quality and tend to negatively affect the performance of outdoor computer vision systems. Hence, removing rains from an image has become an important issue in the field. To handle such an ill-posed single image deraining task, in this article, we specifically build a novel deep architecture, called rain convolutional dictionary network (RCDNet), which embeds the intrinsic priors of rain streaks and has clear interpretability. In specific, we first establish a rain convolutional dictionary (RCD) model for representing rain streaks and utilize the proximal gradient descent technique to design an iterative algorithm only containing simple operators for solving the model. By unfolding it, we then build the RCDNet in which every network module has clear physical meanings and corresponds to each operation involved in the algorithm. This good interpretability greatly facilitates an easy visualization and analysis of what happens inside the network and why it works well in the inference process. Moreover, taking into account the domain gap issue in real scenarios, we further design a novel dynamic RCDNet, where the rain kernels can be dynamically inferred corresponding to input rainy images and then help shrink the space for rain layer estimation with few rain maps, so as to ensure a fine generalization performance in the inconsistent scenarios of rain types between training and testing data. By end-to-end training such an interpretable network, all involved rain kernels and proximal operators can be automatically extracted, faithfully characterizing the features of both rain and clean background layers and, thus, naturally leading to better deraining performance. Comprehensive experiments implemented on a series of representative synthetic and real datasets substantiate the superiority of our method, especially on its well generality to diverse testing scenarios and good interpretability for all its modules, compared with state-of-the-art single image derainers both visually and quantitatively. Code is available at https://github.com/hongwang01/DRCDNet. Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Yuexiang Li, Yong Liang 0001, Yefeng Zheng 0001, Deyu Meng |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Predicting potential microbe-disease associations based on auto-encoder and graph convolution networkabstractThe increasing body of research has consistently demonstrated the intricate correlation between the human microbiome and human well-being. Microbes can impact the efficacy and toxicity of drugs through various pathways, as well as influence the occurrence and metastasis of tumors. In clinical practice, it is crucial to elucidate the association between microbes and diseases. Although traditional biological experiments accurately identify this association, they are time-consuming, expensive, and susceptible to experimental conditions. Consequently, conducting extensive biological experiments to screen potential microbe-disease associations becomes challenging. The computational methods can solve the above problems well, but the previous computational methods still have the problems of low utilization of node features and the prediction accuracy needs to be improved. To address this issue, we propose the DAEGCNDF model predicting potential associations between microbes and diseases. Our model calculates four similar features for each microbe and disease. These features are fused to obtain a comprehensive feature matrix representing microbes and diseases. Our model first uses the graph convolutional network module to extract low-rank features with graph information of microbes and diseases, and then uses a deep sparse Auto-Encoder to extract high-rank features of microbe-disease pairs, after which the low-rank and high-rank features are spliced to improve the utilization of node features. Finally, Deep Forest was used for microbe-disease potential relationship prediction. The experimental results show that combining low-rank and high-rank features helps to improve the model performance and Deep Forest has better classification performance than the baseline model. Shanghui Lu, Yong Liang 0001, Rui Miao 0002, Shuilin Liao, Yongfu Zou, Chengjun Yang, Dong Ouyang |
BMC Bioinform. | 2 |
| 2023 | UAMPnet: Unrolled approximate message passing network for nonconvex regularization
Hui Zhang 0032, Shoujiang Li, Yong Liang 0001, Hai Zhang 0001, Mengmeng Du |
Expert Syst. Appl. | 3 |
| 2023 | A distributed sparse logistic regression with L1/2 regularization for microarray biomarker discovery in cancer classification
Ning Ai, Ziyi Yang 0007, Hao-Laing Yuan, Dong Ouyang, Rui Miao 0002, Yu-Han Ji, Yong Liang 0001 |
Soft Comput. | 7 |
| 2023 | Improved Computational Drug-Repositioning by Self-Paced Non-Negative Matrix Tri-FactorizationabstractDrug repositioning (DR) is a strategy to find new targets for existing drugs, which plays an important role in reducing the costs, time, and risk of traditional drug development. Recently, the matrix factorization approach has been widely used in the field of DR prediction. Nevertheless, there are still two challenges: 1) Learning ability deficiencies, the model cannot accurately predict more potential associations. 2) Easy to fall into a bad local optimal solution, the model tends to get a suboptimal result. In this study, we propose a self-paced non-negative matrix tri-factorization (SPLNMTF) model, which integrates three types of different biological data from patients, genes, and drugs into a heterogeneous network through non-negative matrix tri-factorization, thereby learning more information to improve the learning ability of the model. In the meantime, the SPLNMTF model sequentially includes samples into training from easy (high-quality) to complex (low-quality) in the soft weighting way, which effectively alleviates falling into a bad local optimal solution to improve the prediction performance of the model. The experimental results on two real datasets of ovarian cancer and acute myeloid leukemia (AML) show that SPLNMTF outperforms the other eight state-of-the-art models and gets better prediction performance in drug repositioning. The data and source code are available at: https://github.com/qi0906/SPLNMTF. Qi Dang, Yong Liang 0001, Dong Ouyang, Rui Miao 0002, Caijin Ling, Xiao-Ying Liu 0003, Shengli Xie 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2022 | SSGL1/2: An Improved SVM with Smooth GroupL1/2 for Predicting ADabstractAlzheimer’s disease (AD) is currently one of the mainstream senile diseases recognized in the world. It is the key problem how to automatically identify the early AD based on structed Magnetic Resonance Imaging (sMRI). In order to achieve accurate recognition of AD and obtain highly relevant brain lesions, an improved SVM with group L1/2 sparse regularization and smoothing function (SGL1/2) is proposed. It can achieve sparseness within the group, and approximate the non-smooth absolute value function to a smooth function. The improved model adopts a calibrated hinge to replace the hinge loss function in traditional SVM which is abbreviated as SSGL1/2. In the experiment, the proposed model is applied to different sMRI datasets for training and testing. Compared to other regularization of the non-group level and the group level, the classification accuracy of the proposed method reaches up to 96.03%. At the same time, the algorithm can point out the important brain areas in the MRI group, which has important reference value for the doctor’s predictive work. Jinfeng Wang 0003, ShuaiHui Hang, Yong Liang 0001, Wenzhong Wang |
BIBM | 3 |
| 2022 | Predicting multiple types of miRNA-disease associations using adaptive weighted nonnegative tensor factorization with self-paced learning and hypergraph regularizationabstractMore and more evidence indicates that the dysregulations of microRNAs (miRNAs) lead to diseases through various kinds of underlying mechanisms. Identifying the multiple types of disease-related miRNAs plays an important role in studying the molecular mechanism of miRNAs in diseases. Moreover, compared with traditional biological experiments, computational models are time-saving and cost-minimized. However, most tensor-based computational models still face three main challenges: (i) easy to fall into bad local minima; (ii) preservation of high-order relations; (iii) false-negative samples. To this end, we propose a novel tensor completion framework integrating self-paced learning, hypergraph regularization and adaptive weight tensor into nonnegative tensor factorization, called SPLDHyperAWNTF, for the discovery of potential multiple types of miRNA-disease associations. We first combine self-paced learning with nonnegative tensor factorization to effectively alleviate the model from falling into bad local minima. Then, hypergraphs for miRNAs and diseases are constructed, and hypergraph regularization is used to preserve the high-order complex relations of these hypergraphs. Finally, we innovatively introduce adaptive weight tensor, which can effectively alleviate the impact of false-negative samples on the prediction performance. The average results of 5-fold and 10-fold cross-validation on four datasets show that SPLDHyperAWNTF can achieve better prediction performance than baseline models in terms of Top-1 precision, Top-1 recall and Top-1 F1. Furthermore, we implement case studies to further evaluate the accuracy of SPLDHyperAWNTF. As a result, 98 (MDAv2.0) and 98 (MDAv2.0-2) of top-100 are confirmed by HMDDv3.2 dataset. Moreover, the results of enrichment analysis illustrate that unconfirmed potential associations have biological significance. Dong Ouyang, Yong Liang 0001, Xiao-Ying Liu 0003, Shengli Xie 0001, Rui Miao 0002, Ning Ai, Qi Dang |
Briefings Bioinform. | 2 |
| 2022 | A novel meta-analysis based on data augmentation and elastic data shared lasso regularization for gene expressionabstractBACKGROUND: Gene expression analysis can provide useful information for analyzing complex biological mechanisms. However, many reported findings are unrepeatable due to small sample sizes relative to a large number of genes and the low signal-to-noise ratios of most gene expression datasets. RESULTS: Meta-analysis of multi-data sets is an efficient method for tackling the above problem. To improve the performance of meta-analysis, we propose a novel meta-analysis framework. It consists of two parts: (1) a novel data augmentation strategy. Various cross-platform normalization methods exist, which can preserve original biological information of gene expression datasets from different angles and add different "perturbations" to the dataset. Using such perturbation, we provide a feasible means for gene expression data augmentation; (2) elastic data shared lasso (DSL-[Formula: see text]). The DSL-[Formula: see text] method spans the continuum between individual models for each dataset and one model for all datasets. It also overcomes the shortcomings of the data shared lasso method when dealing with highly correlated features. Comprehensive simulation experiment results show that the proposed method has high prediction and gene selection performance. We then apply the proposed method to non-small cell lung cancer (NSCLC) blood gene expression data in order to identify key tumor-related genes. The outcomes of our experiment indicate that the method could be used for identifying a set of robust disease-related gene signatures that may be used for NSCLC early diagnosis or prognosis or even targeting. CONCLUSION: We propose a novel and effective meta-analysis method for biological research, extrapolating and integrating information from multiple gene expression datasets. Hai-Hui Huang, Hao Rao, Rui Miao 0002, Yong Liang 0001 |
BMC Bioinform. | 4 |
| 2022 | SLNL: A novel method for gene selection and phenotype classificationabstractOne of the central tasks of genome research is to predict phenotypes and discover some important gene biomarkers. However, there are three main problems in analyzing genomics data to predict phenotypes and gene marker selection. Such as large p and small n, low reproducibility of the selected biomarkers, and high noise. To provide a unified solution to alleviate the problems as mentioned above, we propose a self-paced learning L 1 / 2 ${{\rm{L}}}_{1/2}$ absolute network-based logistic regression model, called SLNL. Through the L 1 / 2 ${L}_{1/2}$ regularization, the model can get a more sparse result, which provides better interpretability. The absolute network-based penalty enables the model to integrate the feature network knowledge and helps select higher reproducibility genes. Moreover, this proposed penalty overcomes the drawback of a traditional network penalty without considering the sign of the coefficient. By the self-paced learning strategy, the model can now consider the noise level in gene expression data, lower the impact of high noise samples in data to model training, and provide better prediction accuracy. We compare the proposed method with six alternative approaches in various experimental scenarios, including a comprehensive simulation, four benchmark gene expression datasets, one lung cancer data set, and three lung cancer validation sets. Results show that SLNL can identify fewer meaningful biomarkers and obtain the best or equivalent prediction performance. Moreover, biological analysis shows that the genes selected by the SLNL might be helpful to tumor diagnosis and treatment. Hai-Hui Huang, Yong Liang 0001, Xindong Peng |
Int. J. Intell. Syst. | 3 |
| 2022 | SMSPL: Robust Multimodal Approach to Integrative Analysis of Multiomics DataabstractWith the recent advancement of technologies, it is progressively easier to collect diverse types of genome-wide data. It is commonly expected that by analyzing these data in an integrated way, one can improve the understanding of a complex biological system. Current methods, however, are prone to overfitting heavy noise such that their applications are limited. High noise is one of the major challenges for multiomics data integration. This may be the main cause of overfitting and poor performance in generalization. A sample reweighting strategy is typically used to cope with this problem. In this article, we propose a robust multimodal data integration method, called SMSPL, which can simultaneously predict subtypes of cancers and identify potentially significant multiomics signatures. Especially, the proposed method leverages the linkages between different types of data to interactively recommend high-confidence samples, adopts a new soft weighting scheme to assign weights to the training samples of each type, and then iterates between weights recalculating and classifiers updating. Simulation and five real experiments substantiate the capability of the proposed method for classification and identification of significant multiomics signatures with heavy noise. We expect SMSPL to take a small step in the multiomics data integration and help researchers comprehensively understand the biological process. Ziyi Yang 0007, Yong Liang 0001, Hui Zhang 0032 |
IEEE Trans. Cybern. | 3 |
| 2021 | SPLSN: An efficient tool for survival analysis and biomarker selectionabstractIn genome research, it is a fundamental issue to identify few but important survival-related biomarkers. The Cox model is a widely used survival analysis technique, which is used to study the relationship between characteristics and survival response. However, limitations of the existing Cox methods for genomic data are as follows: (1) a typical gene expression data set consists of tens of thousands of genes, and the result of current methods may not be sparse enough; (2) a wealth of structural information about many biological processes, such as regulatory networks and pathways, has often been ignored; (3) genomic data is usually considered as high noise, which is usually ignored in current methods. To alleviate the above problems, in this paper, we study a novel sparse Cox regression model, called SPLSN, which combines self-paced learning (SPL) and a log-sum absolute network-based penalty (Logsum-Net), especially for biomarker selection in survival analysis. SPL is embedded in curriculum design, and the model is trained by gradually increasing samples from low noise to high noise during the training process. The Logsum-Net encourages smoothness among the coefficients of adjacent genes on a specific biological network. We compare the proposed method with five alternative approaches in various experimental scenarios, including a comprehensive simulation, seven benchmark gene expression data sets, and one large validation data set. Results show that the SPLSN can identify fewer meaningful biomarkers and obtain the best or equivalent prediction performance. Moreover, the biological analysis shows that the genes selected by the SPLSN might be helpful to tumor treatment. Hai-Hui Huang, Xindong Peng, Yong Liang 0001 |
Int. J. Intell. Syst. | 3 |
| 2021 | ERFR-CTC: Exploiting Residual Frequency Resources in Physical-Level Cross-Technology CommunicationabstractIn Internet of Things (IoT), physical-level cross-technology communication (CTC) enables IoT gateways to communicate with heterogeneous nodes economically. However, because of bandwidth asymmetry between heterogeneous technologies, many residual frequency resources are often not fully utilized. Without modification on hardware, in this article, we consider the coexistence of ultralow power (ULP) and WiFi nodes, and propose ERFR-CTC that enables an IoT gateway to fully exploit residual frequency resources without adding additional cost. With ERFR-CTC, the gateway can simultaneously communicate with ULP and WiFi nodes only via a single WiFi network interface card (NIC), which is not only economic but also very efficient. In particular, ERFR-CTC enables ULP nodes to correctly demodulate ULP signals without being interfered by WiFi signals. We then develop theoretical models to quantify available residual frequency resources and analyze the system throughput. Finally, extensive simulations verify that our model is very accurate and show that ERFR-CTC can increase the system throughput by up to 51.4%. Shumin Yao, Li Feng 0001, Qinglin Zhao, Qiyu Yang, Yong Liang 0001 |
IEEE Internet Things J. | 5 |
| 2021 | Structural residual learning for single image rain removal
Hong Wang 0021, Qi Xie 0002, Qian Zhao 0002, Yong Liang 0001, Deyu Meng |
Knowl. Based Syst. | 5 |
| 2021 | A Novel Cox Proportional Hazards Model for High-Dimensional Genomic Data in Cancer PrognosisabstractThe Cox proportional hazards model is a popular method to study the connection between feature and survival time. Because of the high-dimensionality of genomic data, existing Cox models trained on any specific dataset often generalize poorly to other independent datasets. In this paper, we suggest a novel strategy for the Cox model. This strategy is included a new learning technique, self-paced learning (SPL), and a new gene selection method, SCAD-Net penalty. The SPL method is adopted to aid to build a more accurate prediction with its built-in mechanism of learning from easy samples first and adaptively learning from hard samples. The SCAD-Net penalty has fixed the problem of the SCAD method without an inherent mechanism to fuse the prior graphical information. We combined the SPL with the SCAD-Net penalty to the Cox model (SSNC). The simulation shows that the SSNC outperforms the benchmark in terms of prediction and gene selection. The analysis of a large-scale experiment across several cancer datasets shows that the SSNC method not only results in higher prediction accuracies but also identifies markers that satisfactory stability across another validation dataset. The demo code for the proposed method is provided in supplemental file. Hai-Hui Huang, Yong Liang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2020 | A greedy screening test strategy to accelerate solving LASSO problems with small regularization parameters
Hai-Wei Shen, Liang-Yong Xia, Sheng-Bing Wu, Yong Liang 0001, Xiang-Tao Liu |
Soft Comput. | 6 |
| 2020 | Novel Regularization Method for Biomarker Selection and Cancer Classificationabstractpenalty to evaluate the grouping effect and penalty with a fixed value of q to evaluate the variable sparsity, respectively. These methods typically produce a good performance with high efficiency, but they often require the data to satisfy a certain probability distribution. In this paper, we propose a novel complex harmonic regularization (CHR) penalty function, which can approximate the combination of [Formula: see text] and regularizations with adjustable p and q to select the groups of the relevant variables. The CHR penalty function can be effectively solved by a direct path seeking algorithm. We demonstrate that the proposed CHR penalty function performs better than the state-of-the-art regularization methods in selecting groups of relevant variables and classification. Xiao-Ying Liu 0003, Hai Zhang 0001, Hui Zhang 0032, Ziyi Yang 0007, Yong Liang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2019 | An integrative analysis system of gene expression using self-paced learning and SCAD-Net
Hai-Hui Huang, Yong Liang 0001 |
Expert Syst. Appl. | 2 |
| 2014 | A novel L1/2 regularization shooting method for Cox's proportional hazards model
Xin-Ze Luan, Yong Liang 0001, Kwong-Sak Leung, Tak-Ming Chan, Zongben Xu, Hai Zhang 0001 |
Soft Comput. | 2 |
| 2013 | Genetic algorithm for dimer-led and error-restricted spaced motif discoveryabstractDNA motif discovery is an important problem for deciphering protein-DNA bindings in gene regulation. To discover generic spaced motifs which have multiple conserved patterns separated by wild-cards called spacers, the genetic algorithm (GA) based GASMEN has been proposed and shown to outperform related methods. However, the over-generic modeling of any number of spacers increases the optimization difficulty in practice. In protein-DNA binding case studies, complicated spaced motifs are rare while dimers with single spacers are more common spaced motifs. Moreover, errors (mismatches) in a conserved pattern are not arbitrarily distributed as certain highly conserved nucleotides are essential to maintain bindings. Motivated by better optimization in real applications, we have developed a new method, which is GA for Dimer-led and Error-restricted Spaced Motifs (GADESM). Common spaced motifs are paid special attention to using dimer-led initialization in the population initialization. The results on real datasets show that the dimer-led initialization in GADESM achieves better fitness than GASMEN with statistical significance. With additional error-restricted motif occurrence retrieval, GADESM has shown better performance than GASMEN on both comprehensive simulation data and a real ChIP-seq case study. Tak-Ming Chan, Leung-Yau Lo, Man Leung Wong, Yong Liang 0001, Kwong-Sak Leung |
CIBCB | 4 |
| 2013 | Sparse logistic regression with a L1/2 penalty for gene selection in cancer classificationabstractBACKGROUND: Microarray technology is widely used in cancer diagnosis. Successfully identifying gene biomarkers will significantly help to classify different cancer types and improve the prediction accuracy. The regularization approach is one of the effective methods for gene selection in microarray data, which generally contain a large number of genes and have a small number of samples. In recent years, various approaches have been developed for gene selection of microarray data. Generally, they are divided into three categories: filter, wrapper and embedded methods. Regularization methods are an important embedded technique and perform both continuous shrinkage and automatic gene selection simultaneously. Recently, there is growing interest in applying the regularization techniques in gene selection. The popular regularization technique is Lasso (L1), and many L1 type regularization terms have been proposed in the recent years. Theoretically, the Lq type regularization with the lower value of q would lead to better solutions with more sparsity. Moreover, the L1/2 regularization can be taken as a representative of Lq (0 <q < 1) regularizations and has been demonstrated many attractive properties. RESULTS: In this work, we investigate a sparse logistic regression with the L1/2 penalty for gene selection in cancer classification problems, and propose a coordinate descent algorithm with a new univariate half thresholding operator to solve the L1/2 penalized logistic regression. Experimental results on artificial and microarray data demonstrate the effectiveness of our proposed approach compared with other regularization methods. Especially, for 4 publicly available gene expression datasets, the L1/2 regularization method achieved its success using only about 2 to 14 predictors (genes), compared to about 6 to 38 genes for ordinary L1 and elastic net regularization approaches. CONCLUSIONS: From our evaluations, it is clear that the sparse logistic regression with the L1/2 penalty achieves higher classification accuracy than those of ordinary L1 and elastic net regularization approaches, while fewer but informative genes are selected. This is an important consideration for screening and diagnostic applications, where the goal is often to develop an accurate test using as few features as possible in order to control cost. Therefore, the sparse logistic regression with the L1/2 penalty is effective technique for gene selection in real classification problems. Yong Liang 0001, Xin-Ze Luan, Kwong-Sak Leung, Tak-Ming Chan, Zongben Xu, Hai Zhang 0001 |
BMC Bioinform. | 1 |
| 2012 | The essential ability of sparse reconstruction of different compressive sensing strategies
Hai Zhang 0001, Yong Liang 0001, HaiLiang Gou, Zongben Xu |
Sci. China Inf. Sci. | 2 |
| 2010 | L1/2 regularization
Zongben Xu, Hai Zhang 0001, Yao Wang 0003, Xiangyu Chang, Yong Liang 0001 |
Sci. China Inf. Sci. | 5 |
| 2007 | Neural Network Training Using Genetic Algorithm with a Novel Binary Encoding
Yong Liang 0001, Kwong-Sak Leung, Zongben Xu |
ISNN (2) | 1 |
| 2007 | A Memetic Algorithm for Multiple-Drug Cancer Chemotherapy Schedule OptimizationabstractThis correspondence introduces a multidrug cancer chemotherapy model to simulate the possible response of the tumor cells under drug administration. We formulate the model as an optimal control problem. The algorithm in this correspondence optimizes the multidrug cancer chemotherapy schedule. The objective is to minimize the tumor size under a set of constraints. We combine the adaptive elitist genetic algorithm with a local search algorithm called iterative dynamic programming (IDP) to form a new memetic algorithm (MA-IDP) for solving the problem. MA-IDP has been shown to be very efficient in solving the multidrug scheduling optimization problem. Sui-Man Tse, Yong Liang 0001, Kwong-Sak Leung, Kin-Hong Lee, Tony Shu Kam Mok |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2006 | A Novel Binary Variable Representation for Genetic and Evolutionary AlgorithmsabstractBased on the theoretical guidance and existing recommendations for designing efficient genetic representations, we investigate a novel genetic representation — a splicing/decomposable (S/D) binary encoding in this paper. The S/D binary representation can be spliced and decomposed to describe potential solutions of the problem with different precisions by different number of uniform-salient building blocks (BBs). According to the characteristics of the S/D binary representation, genetic and evolutionary algorithms (GEAs) can be applied from the high scaled to the low scaled BBs sequentially to avoid genetic drift and improve GEAs’ performance. Our theoretical and empirical investigations reveal that the S/D binary representation is more proper than other existing binary encodings for GEAs searching. Yong Liang 0001, Kwong-Sak Leung, Kin-Hong Lee |
IEEE Congress on Evolutionary Computation | 1 |
| 2006 | Optimal Control of a Cancer Chemotherapy Problem with Different Toxic Elimination ProcessesabstractIn this paper, we propose two new anticancer drug scheduling models with different toxicity clearances according to kinetics of enzyme-catalyzed chemical reactions. We also present a sophisticated automating drug scheduling approach based on evolutionary computation and computer modeling. To explore multiple efficient drug scheduling policies, we use a multimodal optimization algorithm — adaptive elitist-population based genetic algorithm (AEGA) to solve the models, and discuss the situation of multiple optimal solutions under different parameter settings. The simulation results obtained by the new models match well with the clinical treatment experience, and can provide much more drug scheduling policies for a doctor to choose depending on the particular conditions of the patients. Yong Liang 0001, Kwong-Sak Leung, Tony Shu Kam Mok |
IEEE Congress on Evolutionary Computation | 1 |
| 2006 | A splicing/decomposable encoding and its novel operators for genetic algorithmsabstractIn this paper, we introduce a new genetic representation --- a splicing/decomposable (S/D) binary encoding, which was proposed based on some theoretical guidance and existing recommendations for designing efficient genetic representations. Our theoretical and empirical investigations reveal that the S/D binary representation is more proper than other existing binary encodings for searching of genetic algorithms (GAs). Moreover, we define a new genotypic distance on the S/D binary space, which is equivalent to the Euclidean distance on the real-valued space during GAs convergence. Based on the new genotypic distance, GAs can reliably and predictably solve problems of bounded complexity and the methods depended on the Euclidean distance for solving different kinds of optimization problems can be directly used on the S/D binary space. Yong Liang 0001, Kwong-Sak Leung, Kin-Hong Lee |
GECCO | 1 |
| 2006 | Automating the drug scheduling with different toxicity clearance in cancer chemotherapy via evolutionary computationabstractThe toxicity of an anticancer drug is cleared from the body by different processes, including saturable metabolic and nonsaturable renal-excretion pathways. According to the principles of toxicokinetics, we propose a new anticancer drug scheduling model with different toxic elimination processes in this paper. We also present a sophisticated automating drug scheduling approach based on evolutionary computation and computer modeling. To explore multiple efficient drug scheduling policies, we use a multimodal optimization algorithm --- adaptive elitist-population based genetic algorithm (AEGA) to solve the new model, and discuss the situation of multiple optimal solutions under different parameter settings. The simulation results obtained by the new model match well with the clinical treatment experience, and can provide much more drug scheduling policies for a doctor to choose depending on the particular conditions of the patients. Yong Liang 0001, Kwong-Sak Leung, Tony Shu Kam Mok |
GECCO | 1 |
| 2006 | A Novel Evolutionary Drug Scheduling Model in Cancer ChemotherapyabstractIn this paper, we introduce a modified optimal control model of drug scheduling in cancer chemotherapy and a new adaptive elitist-population-based genetic algorithm (AEGA) to solve it. Working closely with an oncologist, we first modify the existing model, because its equation for the cumulative drug toxicity is inconsistent with medical knowledge and clinical experience. To explore multiple efficient drug scheduling policies, we propose a novel variable representation--a cycle-wise representation, and modify the elitist genetic search operators in the AEGA. The simulation results obtained by the modified model match well with the clinical treatment experiences, and can provide multiple efficient solutions for oncologists to consider. Moreover, it has been shown that the evolutionary drug scheduling approach is simple, and capable of solving complex cancer chemotherapy problems by adapting multimodal versions of evolutionary algorithms. Yong Liang 0001, Kwong-Sak Leung, Tony Shu Kam Mok |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2005 | Multi-drug cancer chemotherapy scheduling by a new memetic optimization algorithmabstractThis paper proposes a new memetic algorithm (MA) to solve the multi-drug chemotherapy optimization problem. The new MA combines GA with a local search algorithm called iterative dynamic programming (IDP). A multi-drug chemotherapy model is introduced to simulate the possible response of the tumor cells under drugs administration. Optimization of the multiple chemotherapeutic agents' administration schedules is based on this tumor model. We formulate the optimization problem as an optimal control problem (OCP) with a set of dynamic equations. The objective is to design efficient schedules which minimize the tumor size under a set of constraints. Our new MA has been shown to be very efficient on solving our multi-drug model. Sui-Man Tse, Yong Liang 0001, Kwong-Sak Leung, Kin-Hong Lee, Tony Shu Kam Mok |
Congress on Evolutionary Computation | 2 |
| 2004 | Evolutionary Drug Scheduling Model for Cancer Chemotherapy
Yong Liang 0001, Kwong-Sak Leung, Tony Shu Kam Mok |
GECCO (2) | 1 |
| 2003 | Evolution Strategies with Exclusion-Based Selection Operators and a Fourier Series Auxiliary Function
Kwong-Sak Leung, Yong Liang 0001 |
GECCO | 2 |
| 2003 | Adaptive Elitist-Population Based Genetic Algorithm for Multimodal Function Optimization
Kwong-Sak Leung, Yong Liang 0001 |
GECCO | 2 |
| 2003 | Evolution Strategies with a Fourier Series Auxiliary Function for Difficult Function Optimization
Kwong-Sak Leung, Yong Liang 0001 |
IDEAL | 2 |
| 2002 | Two-way mutation evolution strategiesabstractIn this paper, a new two-way adaptive search mutation strategy is proposed, and a "two-way evolution strategy" (TWES) is established. The experimental results show that TWES yields much faster convergence than classical evolution strategies. This paper also discusses the relationship between the parameter setting and the convergent speed by TWES. Yong Liang 0001, Kwong-Sak Leung |
IEEE Congress on Evolutionary Computation | 1 |