VLDB 2026 Research / reviewers in the wild / expert
Ying Li 0004
dblp:22/1805-4
· DBLP profile ↗
35ranked-venue papers
13as first author
25since 2021 · last 2026
0000-0002-7804-149XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Whom to Align With: Progressive Anomaly Combination Detection for Partially View-Aligned ClusteringabstractPartially View-aligned Clustering (PVC) addresses the challenge of partial view alignment in multi-view learning by leveraging complementary and consistent information. While existing PVC methods show promise, most rely on distance-based strategies that are sensitive to view-specific details and noise, limiting their robustness. In this work, we propose a novel view alignment strategy that reformulates the alignment task as an anomaly detection problem. Rather than learning a view-alignment matrix that enforces strict one-to-one correspondences across views, we adopt a progressive approach to identify well-aligned samples. Specifically, we sample subsets of data by generating random view combinations from unaligned samples and propose an anomaly combination detection module to evaluate the alignment consistency of these combinations. In addition, our progressive training framework alternates between updating model parameters and selecting high-confidence view combinations for subsequent optimization. By reformulating view alignment as an anomaly detection task, our approach provides a more robust and effective solution to partial view alignment. Experiments on benchmark datasets demonstrate that our method outperforms state-of-the-art approaches in the PVC problem. Hang Gao 0014, Zuosong Cai, Cheng Liu 0001, Ying Li 0004, Wei Du 0002, You Zhou 0008 |
AAAI | 6 |
| 2026 | Multi-level cross-view feature embedding for partial view-aligned clustering
Hang Gao 0014, Cheng Liu 0001, Ying Li 0004, You Zhou 0008, Wei Du 0002 |
Knowl. Based Syst. | 3 |
| 2026 | Cross-view discrepancy-driven dynamic weighting for missing view completion in incomplete multi-view clustering
Hang Gao 0014, Zuosong Cai, Cheng Liu 0001, Ying Li 0004, You Zhou 0008, Wei Du 0002 |
Neural Networks | 5 |
| 2026 | Incomplete multi-view clustering with cross-view generation via pre-trained transformer
Hang Gao 0014, Cheng Liu 0001, Hongming Sun, Ying Li 0004, You Zhou 0008, Wei Du 0002 |
Pattern Recognit. | 5 |
| 2025 | Contrastive Auxiliary Learning with Structure Transformation for Heterogeneous GraphsabstractIn recent years, methods based on heterogeneous graph neural networks (HGNNs) have been widely used for embedding heterogeneous graphs (HGs) due to their ability to effectively encode the rich information from HGs into low-dimensional node embeddings. Existing HGNNs focus on neighbor aggregation and semantic fusion while neglecting the HG structure and learning paradigms. However, the original HG data might lack node features, which existing models may not effectively account for. Additionally, exclusively relying on a single supervised learning approach may only partially leverage the invariant information in graph data. To address these challenges, we introduce the Contrastive Auxiliary Learning Model for Heterogeneous Graphs (CALHG). This model combines edge perturbation and graph diffusion to enhance graph data, allowing it to capture the inherent structural information within heterogeneous graphs fully. Additionally, we employ a category-guided multi-view contrastive learning approach, which does not rely on positive and negative samples for model training, enabling us to capture the intrinsic invariances in heterogeneous graph data. Extensive experiments and analyses on five benchmark datasets without node features and three benchmark datasets with node features demonstrate the effectiveness and efficiency of our novel method compared with several state-of-the-art methods. Wei Du 0002, Hongmin Sun, Hang Gao 0014, Ying Li 0004 |
AAAI | 5 |
| 2025 | Multi-Manifolds fusing hyperbolic graph network balanced by pareto optimization for identifying spatial domains of spatial transcriptomicsabstractIdentifying spatial domains for spatial transcriptomics is crucial for achieving comprehensive insights into the pathogenesis of gene expression. Increasingly, computational methods based on graph neural networks are being developed for spatial transcriptomics. However, previous methods have solely focused on the Euclidean manifold. To effectively exploit and explore the informative and deeper topological structures of inherent manifolds, we presented a Multi-Manifolds fusing hyperbolic graph network, balanced by Pareto optimization, for identifying spatial domains in Spatial Transcriptomics (MManiST). First, we developed multi-manifolds encoders for distinct manifolds using the hyperbolic neural network. Features from different manifolds were then combined using an attention mechanism, with multiple reconstruction losses balanced by Pareto optimization. Extensive experiments on commonly used benchmark datasets show that our method consistently outperforms seven state-of-the-art methods. Additionally, we investigated the validity of each component and the impact of fusion methods in ablation experiments. Ying Li 0004, Qifeng Hu, Rui Wang-Sattler, Wei Du 0002 |
Briefings Bioinform. | 1 |
| 2025 | A multi-omics integration framework using multi-label guided learning and multi-scale fusionabstractThe rapid development of high-throughput sequencing technologies has generated vast amounts of omics data, making multi-omics integration a crucial approach for understanding complex diseases. Despite the introduction of various multi-omics integration methods in recent years, existing approaches still have limitations, primarily in their reliance on manual feature selection, restricted applicability, and inability to comprehensively capture both inter-sample and cross-omics interactions. To address these challenges, we propose mmMOI, an end-to-end multi-omics integration framework that incorporates multi-label guided learning and multi-scale attention fusion. mmMOI directly processes raw high-dimensional omics data without requiring manual feature selection, thereby enhancing model interpretability and eliminating biases introduced by feature preselection. First, we introduce a multi-label guided multi-view graph neural network, which enables the model to adaptively learn omics data representations across different datasets, thereby improving generalizability and stability. Second, we design a multi-scale attention fusion network, which integrates global attention and local attention. This dual-attention mechanism allows mmMOI to more accurately integrate multi-omics data, enhance cross-omics feature representations, and improve classification performance. Experimental results demonstrate that mmMOI significantly outperforms state-of-the-art methods in classification tasks, exhibiting high stability and adaptability across diverse biological contexts and sequencing technologies. Additionally, mmMOI successfully identifies key disease-associated biomarkers, further enhancing its biological interpretability and practical relevance. The source code, datasets, and detailed hyperparameter configurations for mmMOI are available at https://github.com/mlcb-jlu/mmMOI. Yinghe Wang, Ying Li 0004, Wei Du 0002 |
Briefings Bioinform. | 4 |
| 2025 | Cross-loop contrast of heterogeneous graphs for interdisciplinary journal recommendation
Ying Li 0004, Hongmin Sun, Wei Du 0002, Qin Ma 0003 |
Expert Syst. Appl. | 1 |
| 2025 | A Novel Approach for Effective Partially View-Aligned Clustering With Triple-ConsistencyabstractMulti-view clustering (MVC), which integrates information from multiple views to enhance performance, has garnered increasing attention in recent years. Partially View-aligned Clustering (PVC), which is a particularly critical aspect of this process, requires a thorough exploration of complementary and consistent information under conditions of partial view alignment. However, most existing PVC methods primarily focus on semantic consistency, employing semantic consistency features for both view alignment and clustering tasks. These methods neglect the effects of noise and complementary information across multiple views and the suitability of these features for clustering. To address these limitations, our approach aims to leverage three distinct types of consistency to extract semantic consistency features and clustering consistency features, which are specifically designed for view alignment and clustering tasks, respectively. By omitting the reconstruction process, we mitigate the adverse effects of mutual information and noise on view alignment. Specifically, we first exploit the structural consistency of similarity graphs across different views to guide feature extraction in view-specific autoencoders. This process produces structural consistency features that are both cluster-discriminative and structurally coherent. Subsequently, two separate multilayer perceptrons (MLPs) are trained via contrastive learning to extract semantic consistency features and clustering consistency features from the structural features. These features are optimized for their respective tasks. Ultimately, a self-paced style view alignment strategy is used to iteratively re-align the data based on semantic and clustering consistency while the model is optimized via the re-aligned data. Extensive experiments on multiple real-world benchmark datasets demonstrate that our method outperforms the state-of-the-art multi-view approaches, highlighting its effectiveness in tackling the challenges of PVC. The code is available at https://github.com/kongyiH/TCLPVC. Hang Gao 0014, Cheng Liu 0001, Zuosong Cai, Hongming Sun, Ying Li 0004, Wei Du 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | What's the Situation With Intelligent Mesh Generation: A Survey and PerspectivesabstractIntelligent Mesh Generation (IMG) represents a novel and promising field of research, utilizing machine learning techniques to generate meshes. Despite its relative infancy, IMG has significantly broadened the adaptability and practicality of mesh generation techniques, delivering numerous breakthroughs and unveiling potential future pathways. However, a noticeable void exists in the contemporary literature concerning comprehensive surveys of IMG methods. This paper endeavors to fill this gap by providing a systematic and thorough survey of the current IMG landscape. With a focus on 113 preliminary IMG methods, we undertake a meticulous analysis from various angles, encompassing core algorithm techniques and their application scope, agent learning objectives, data types, targeted challenges, as well as advantages and limitations. We have curated and categorized the literature, proposing three unique taxonomies based on key techniques, output mesh unit elements, and relevant input data types. This paper also underscores several promising future research directions and challenges in IMG. Na Lei, Zezeng Li, Zebin Xu, Ying Li 0004, Xianfeng Gu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | PecidRL: Petition expectation correction and identification based on deep reinforcement learning
Ying Li 0004, Wensi Fang, Wei Du 0002 |
Inf. Process. Manag. | 1 |
| 2023 | SURE: Screening unlabeled samples for reliable negative samples based on reinforcement learning
Ying Li 0004, Wensi Fang, Qin Ma 0003, Rui Wang-Sattler, Wei Du 0002, Qiong Yu |
Inf. Sci. | 1 |
| 2022 | A novel filter feature selection algorithm based on relief
Xue-Ting Cui, Ying Li 0004 |
Appl. Intell. | 2 |
| 2022 | TIGER: technical variation elimination for metabolomics data using ensemble learning architectureabstractLarge metabolomics datasets inevitably contain unwanted technical variations which can obscure meaningful biological signals and affect how this information is applied to personalized healthcare. Many methods have been developed to handle unwanted variations. However, the underlying assumptions of many existing methods only hold for a few specific scenarios. Some tools remove technical variations with models trained on quality control (QC) samples which may not generalize well on subject samples. Additionally, almost none of the existing methods supports datasets with multiple types of QC samples, which greatly limits their performance and flexibility. To address these issues, a non-parametric method TIGER (Technical variation elImination with ensemble learninG architEctuRe) is developed in this study and released as an R package (https://CRAN.R-project.org/package=TIGERr). TIGER integrates the random forest algorithm into an adaptable ensemble learning architecture. Evaluation results show that TIGER outperforms four popular methods with respect to robustness and reliability on three human cohort datasets constructed with targeted or untargeted metabolomics data. Additionally, a case study aiming to identify age-associated metabolites is performed to illustrate how TIGER can be used for cross-kit adjustment in a longitudinal analysis with experimental data of three time-points generated by different analytical kits. A dynamic website is developed to help evaluate the performance of TIGER and examine the patterns revealed in our longitudinal analysis (https://han-siyu.github.io/TIGER_web/). Overall, TIGER is expected to be a powerful tool for metabolomics data analysis. Jialing Huang, Francesco Foppiano, Cornelia Prehn, Jerzy Adamski, Karsten Suhre, Ying Li 0004, Giuseppe Matullo, Freimut Schliess, Christian Gieger, Annette Peters, Rui Wang-Sattler |
Briefings Bioinform. | 7 |
| 2022 | LION: an integrated R package for effective prediction of ncRNA-protein interactionabstractUnderstanding ncRNA-protein interaction is of critical importance to unveil ncRNAs' functions. Here, we propose an integrated package LION which comprises a new method for predicting ncRNA/lncRNA-protein interaction as well as a comprehensive strategy to meet the requirement of customisable prediction. Experimental results demonstrate that our method outperforms its competitors on multiple benchmark datasets. LION can also improve the performance of some widely used tools and build adaptable models for species- and tissue-specific prediction. We expect that LION will be a powerful and efficient tool for the prediction and analysis of ncRNA/lncRNA-protein interaction. The R Package LION is available on GitHub at https://github.com/HAN-Siyu/LION/. Qi Zhang 0061, Wensi Fang, Ying Li 0004 |
Briefings Bioinform. | 8 |
| 2022 | LPInsider: a webserver for lncRNA-protein interaction extraction from the literatureabstractBACKGROUND: Long non-coding RNA (LncRNA) plays important roles in physiological and pathological processes. Identifying LncRNA-protein interactions (LPIs) is essential to understand the molecular mechanism and infer the functions of lncRNAs. With the overwhelming size of the biomedical literature, extracting LPIs directly from the biomedical literature is essential, promising and challenging. However, there is no webserver of LPIs relationship extraction from literature. RESULTS: LPInsider is developed as the first webserver for extracting LPIs from biomedical literature texts based on multiple text features (semantic word vectors, syntactic structure vectors, distance vectors, and part of speech vectors) and logistic regression. LPInsider allows researchers to extract LPIs by uploading PMID, PMCID, PMID List, or biomedical text. A manually filtered and highly reliable LPI corpus is integrated in LPInsider. The performance of LPInsider is optimal by comprehensive experiment on different combinations of different feature and machine learning models. CONCLUSIONS: LPInsider is an efficient analytical tool for LPIs that helps researchers to enhance their comprehension of lncRNAs from text mining, and also saving their time. In addition, LPInsider is freely accessible from http://www.csbg-jlu.info/LPInsider/ with no login requirement. The source code and LPIs corpus can be downloaded from https://github.com/qiufengdiewu/LPInsider . Ying Li 0004, Lizheng Wei, Cankun Wang, Wei Du 0002 |
BMC Bioinform. | 1 |
| 2022 | MAHE-IM: Multiple Aggregation of Heterogeneous Relation Embedding for Influence Maximization on Heterogeneous Information Networkabstract• A embedding method MAHE-IM for IM on heterogeneous network. • MAHE-IM extracts and aggregates multiple heterogeneous relationships. • Eight network embedding and seven GNN methods extended for IM. • MAHE-IM outperforms eighteen state-of-the-art algorithms. • A web server ( https://mahe-im.com/ ) is developed for maximizing user convenience. Influence maximization (IM), as an essential problem in social network analysis, can identify a minimum group of the influential nodes to maximize the spread of information on the network. The majority of IM studies focus on homogeneous networks, whereas in real world, heterogeneous networks are ubiquitous. The existing IM methods based on homogeneous networks do not consider a variety of complex heterogeneous relationships and the attribute of different types of nodes, which don’t accommodate most real scenario. It is of great significance to study IM methods based on heterogeneous networks, where efficient integrating complex multiple semantic relationships and structures embedded in the heterogeneous network is the key breakthrough point. In this paper, a novel deep learning algorithm for influence maximization on heterogeneous networks based on a multiple aggregation of heterogeneous relation embedding named MAHE-IM is proposed, which can capture the heterogeneous high-order structure and semantic features of heterogeneous information networks. For more comprehensive and systematical evaluation of MAHE-IM, we further extend fourteen state-of-the-art homogeneous and heterogeneous network embedding and graph neural network methods for IM problem and propose fourteen extented IM algorithms. Four popular IM algorithms and our extended fourteen IM algorithms are taken as eighteen baseline algorithms, which can be categorized three types: greedy-based, network embedding-based and GNN-based IM algorithms. Compared with eighteen baseline algorithms, the experimental results illustrate that MAHE-IM is significantly efficient and effective. MAHE-IM is more practical on heterogeneous networks. In addition, in order to maximize the convenience for users, a webserver ( https://mahe-im.com/ ) is developed. The users can obtain the IM results by submitting their heterogeneous networks to the webserver. The corresponding source code and used data in this paper can also be available at https://mahe-im.com/ and https://codeocean.com/capsule/4091031/tree/v1 . Ying Li 0004 |
Expert Syst. Appl. | 1 |
| 2022 | Leveraging multidimensional features for policy opinion sentiment prediction
Wenju Hou, Ying Li 0004 |
Inf. Sci. | 2 |
| 2022 | ADR-MVSNet: A cascade network for 3D point cloud reconstruction with pixel occlusion
Ying Li 0004, Zhijie Zhao |
Pattern Recognit. | 1 |
| 2022 | Global chaotic bat algorithm for feature selection
Ying Li 0004, Xueting Cui |
J. Supercomput. | 1 |
| 2021 | Deep forest ensemble learning for classification of alignments of non-coding RNA sequences based on multi-view structure representationsabstractNon-coding RNAs (ncRNAs) play crucial roles in multiple biological processes. However, only a few ncRNAs' functions have been well studied. Given the significance of ncRNAs classification for understanding ncRNAs' functions, more and more computational methods have been introduced to improve the classification automatically and accurately. In this paper, based on a convolutional neural network and a deep forest algorithm, multi-grained cascade forest (GcForest), we propose a novel deep fusion learning framework, GcForest fusion method (GCFM), to classify alignments of ncRNA sequences for accurate clustering of ncRNAs. GCFM integrates a multi-view structure feature representation including sequence-structure alignment encoding, structure image representation and shape alignment encoding of structural subunits, enabling us to capture the potential specificity between ncRNAs. For the classification of pairwise alignment of two ncRNA sequences, the F-value of GCFM improves 6% than an existing alignment-based method. Furthermore, the clustering of ncRNA families is carried out based on the classification matrix generated from GCFM. Results suggest better performance (with 20% accuracy improved) than existing ncRNA clustering methods (RNAclust, Ensembleclust and CNNclust). Additionally, we apply GCFM to construct a phylogenetic tree of ncRNA and predict the probability of interactions between RNAs. Most ncRNAs are located correctly in the phylogenetic tree, and the prediction accuracy of RNA interaction is 90.63%. A web server (http://bmbl.sdstate.edu/gcfm/) is developed to maximize its availability, and the source code and related data are available at the same URL. Ying Li 0004, Qi Zhang 0061, Zhaoqian Liu, Cankun Wang, Qin Ma 0003, Wei Du 0002 |
Briefings Bioinform. | 1 |
| 2021 | Capsule-LPI: a LncRNA-protein interaction predicting tool based on a capsule networkabstractBACKGROUND: Long noncoding RNAs (lncRNAs) play important roles in multiple biological processes. Identifying LncRNA-protein interactions (LPIs) is key to understanding lncRNA functions. Although some LPIs computational methods have been developed, the LPIs prediction problem remains challenging. How to integrate multimodal features from more perspectives and build deep learning architectures with better recognition performance have always been the focus of research on LPIs. RESULTS: We present a novel multichannel capsule network framework to integrate multimodal features for LPI prediction, Capsule-LPI. Capsule-LPI integrates four groups of multimodal features, including sequence features, motif information, physicochemical properties and secondary structure features. Capsule-LPI is composed of four feature-learning subnetworks and one capsule subnetwork. Through comprehensive experimental comparisons and evaluations, we demonstrate that both multimodal features and the architecture of the multichannel capsule network can significantly improve the performance of LPI prediction. The experimental results show that Capsule-LPI performs better than the existing state-of-the-art tools. The precision of Capsule-LPI is 87.3%, which represents a 1.7% improvement. The F-value of Capsule-LPI is 92.2%, which represents a 1.4% improvement. CONCLUSIONS: This study provides a novel and feasible LPI prediction tool based on the integration of multimodal features and a capsule network. A webserver ( http://csbg-jlu.site/lpc/predict ) is developed to be convenient for users. Ying Li 0004, Shiyao Feng, Qi Zhang 0061, Wei Du 0002 |
BMC Bioinform. | 1 |
| 2021 | RGAM: A novel network architecture for 3D point cloud semantic segmentation in indoor scenes
Xuetao Chen, Ying Li 0004 |
Inf. Sci. | 2 |
| 2021 | DeepHBSP: A Deep Learning Framework for Predicting Human Blood-Secretory Proteins Using Transfer Learning
Wei Du 0002, Hui-Min Bao, Liang Chen 0021, Ying Li 0004, Yanchun Liang 0001 |
J. Comput. Sci. Technol. | 5 |
| 2021 | MGRFE: Multilayer Recursive Feature Elimination Based on an Embedded Genetic Algorithm for Cancer ClassificationabstractMicroarray gene expression data have become a topic of great interest for cancer classification and for further research in the field of bioinformatics. Nonetheless, due to the "large p, small n" paradigm of limited biosamples and high-dimensional data, gene selection is becoming a demanding task, which is aimed at selecting a minimal number of discriminatory genes associated closely with a phenotype. Feature or gene selection is still a challenging problem owing to its nondeterministic polynomial time complexity and thus most of the existing feature selection algorithms utilize heuristic rules. A multilayer recursive feature elimination method based on an embedded integer-coded genetic algorithm, MGRFE, is proposed here, which is aimed at selecting the gene combination with minimal size and maximal information. On the basis of 19 benchmark microarray datasets including multiclass and imbalanced datasets, MGRFE outperforms state-of-the-art feature selection algorithms with better cancer classification accuracy and a smaller selected gene number. MGRFE could be regarded as a promising feature selection method for high-dimensional datasets especially gene expression data. Moreover, the genes selected by MGRFE have close biological relevance to cancer phenotypes. The source code of our proposed algorithm and all the 19 datasets used in this paper are available at https://github.com/Pengeace/MGRFE-GaRFE. Ying Li 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2020 | CapsNet-SSP: multilane capsule network for predicting human saliva-secretory proteinsabstractBACKGROUND: Compared with disease biomarkers in blood and urine, biomarkers in saliva have distinct advantages in clinical tests, as they can be conveniently examined through noninvasive sample collection. Therefore, identifying human saliva-secretory proteins and further detecting protein biomarkers in saliva have significant value in clinical medicine. There are only a few methods for predicting saliva-secretory proteins based on conventional machine learning algorithms, and all are highly dependent on annotated protein features. Unlike conventional machine learning algorithms, deep learning algorithms can automatically learn feature representations from input data and thus hold promise for predicting saliva-secretory proteins. RESULTS: We present a novel end-to-end deep learning model based on multilane capsule network (CapsNet) with differently sized convolution kernels to identify saliva-secretory proteins only from sequence information. The proposed model CapsNet-SSP outperforms existing methods based on conventional machine learning algorithms. Furthermore, the model performs better than other state-of-the-art deep learning architectures mostly used to analyze biological sequences. In addition, we further validate the effectiveness of CapsNet-SSP by comparison with human saliva-secretory proteins from existing studies and known salivary protein biomarkers of cancer. CONCLUSIONS: The main contributions of this study are as follows: (1) an end-to-end model based on CapsNet is proposed to identify saliva-secretory proteins from the sequence information; (2) the proposed model achieves better performance and outperforms existing models; and (3) the saliva-secretory proteins predicted by our model are statistically significant compared with existing cancer biomarkers in saliva. In addition, a web server of CapsNet-SSP is developed for saliva-secretory protein identification, and it can be accessed at the following URL: http://www.csbg-jlu.info/CapsNet-SSP/. We believe that our model and web server will be useful for biomedical researchers who are interested in finding salivary protein biomarkers, especially when they have identified candidate proteins for analyzing diseased tissues near or distal to salivary glands using transcriptome or proteomics. Wei Du 0002, Huansheng Cao, Ying Li 0004 |
BMC Bioinform. | 6 |
| 2020 | A hybrid feature selection algorithm for microarray data
Yue-Feng Zheng, Ying Li 0004, Gang Wang 0013, Qian Xu 0016, Xue-Ting Cui |
J. Supercomput. | 2 |
| 2019 | LncFinder: an integrated platform for long non-coding RNA identification utilizing sequence intrinsic composition, structural information and physicochemical propertyabstractDiscovering new long non-coding RNAs (lncRNAs) has been a fundamental step in lncRNA-related research. Nowadays, many machine learning-based tools have been developed for lncRNA identification. However, many methods predict lncRNAs using sequence-derived features alone, which tend to display unstable performances on different species. Moreover, the majority of tools cannot be re-trained or tailored by users and neither can the features be customized or integrated to meet researchers' requirements. In this study, features extracted from sequence-intrinsic composition, secondary structure and physicochemical property are comprehensively reviewed and evaluated. An integrated platform named LncFinder is also developed to enhance the performance and promote the research of lncRNA identification. LncFinder includes a novel lncRNA predictor using the heterologous features we designed. Experimental results show that our method outperforms several state-of-the-art tools on multiple species with more robust and satisfactory results. Researchers can additionally employ LncFinder to extract various classic features, build classifier with numerous machine learning algorithms and evaluate classifier performance effectively and efficiently. LncFinder can reveal the properties of lncRNA and mRNA from various perspectives and further inspire lncRNA-protein interaction prediction and lncRNA evolution analysis. It is anticipated that LncFinder can significantly facilitate lncRNA-related research, especially for the poorly explored species. LncFinder is released as R package (https://CRAN.R-project.org/package=LncFinder). A web server (http://bmbl.sdstate.edu/lncfinder/) is also developed to maximize its availability. Yanchun Liang 0001, Qin Ma 0003, Yangyi Xu, Wei Du 0002, Cankun Wang, Ying Li 0004 |
Briefings Bioinform. | 8 |
| 2019 | Identification and Functional Inference for Tumor-Associated Long Non-Coding RNAabstractGastric cancer is one of the top leading causes of cancer mortality worldwide especially in China. In recent years, some lncRNAs are discovered to be dysregulated in many cancers. The study on long non-coding RNAs (lncRNAs) relationship with cancers has attracted increasing attention. The molecular mechanism of gastric cancer remains largely unclear factors, especially for lncRNAs. Experiments are feasible to obtain related information, however, experimental identification of cancer-related lncRNAs usually possesses high time complexity and high cost. In this paper, a computational method is proposed to determine the relationship between lncRNA and gastric cancer by reusing the exon-based array of gastric cancer. One specific lncRNAs LINC00365 and its target differentially expressed genes whose products are predicted as blood, urine, or salvia-excretory are identified to be candidates for a combined biomarker for gastric cancer. Further biological function and molecular mechanism of the gastric cancer related lncRNAs and coding gene biomarkers are inferred in terms of multi-source biological knowledge. Ying Li 0004, Yanchun Liang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2018 | A Multi-strategy Region Proposal Network
Ying Li 0004, Gang Wang 0013, Qian Xu 0016 |
Expert Syst. Appl. | 2 |
| 2018 | A novel hybrid algorithm for feature selection
Yue-Feng Zheng, Ying Li 0004, Gang Wang 0013, Qian Xu 0016, Xue-Ting Cui |
Pers. Ubiquitous Comput. | 2 |
| 2017 | RNA-TVcurve: a Web server for RNA secondary structure comparison based on a multi-scale similarity of its triple vector curve representationabstractBACKGROUND: RNAs have been found to carry diverse functionalities in nature. Inferring the similarity between two given RNAs is a fundamental step to understand and interpret their functional relationship. The majority of functional RNAs show conserved secondary structures, rather than sequence conservation. Those algorithms relying on sequence-based features usually have limitations in their prediction performance. Hence, integrating RNA structure features is very critical for RNA analysis. Existing algorithms mainly fall into two categories: alignment-based and alignment-free. The alignment-free algorithms of RNA comparison usually have lower time complexity than alignment-based algorithms. RESULTS: An alignment-free RNA comparison algorithm was proposed, in which novel numerical representations RNA-TVcurve (triple vector curve representation) of RNA sequence and corresponding secondary structure features are provided. Then a multi-scale similarity score of two given RNAs was designed based on wavelet decomposition of their numerical representation. In support of RNA mutation and phylogenetic analysis, a web server (RNA-TVcurve) was designed based on this alignment-free RNA comparison algorithm. It provides three functional modules: 1) visualization of numerical representation of RNA secondary structure; 2) detection of single-point mutation based on secondary structure; and 3) comparison of pairwise and multiple RNA secondary structures. The inputs of the web server require RNA primary sequences, while corresponding secondary structures are optional. For the primary sequences alone, the web server can compute the secondary structures using free energy minimization algorithm in terms of RNAfold tool from Vienna RNA package. CONCLUSION: RNA-TVcurve is the first integrated web server, based on an alignment-free method, to deliver a suite of RNA analysis functions, including visualization, mutation analysis and multiple RNAs structure comparison. The comparison results with two popular RNA comparison tools, RNApdist and RNAdistance, showcased that RNA-TVcurve can efficiently capture subtle relationships among RNAs for mutation detection and non-coding RNA classification. All the relevant results were shown in an intuitive graphical manner, and can be freely downloaded from this server. RNA-TVcurve, along with test examples and detailed documents, are available at: http://ml.jlu.edu.cn/tvcurve/ . Ying Li 0004, Xiaohu Shi, Yanchun Liang 0001, Juan Xie, Qin Ma 0003 |
BMC Bioinform. | 1 |
| 2017 | A novel bacterial foraging optimization algorithm for feature selection
Ying Li 0004, Gang Wang 0013, Yue-Feng Zheng, Qian Xu 0016, Xue-Ting Cui |
Expert Syst. Appl. | 2 |
| 2015 | Multiple parameter control for ant colony optimization applied to feature selection problem
Gang Wang 0013, HaiCheng Eric Chu, Huiling Chen 0001, Weitong Hu, Ying Li 0004, Xujun Peng |
Neural Comput. Appl. | 6 |
| 2012 | Multi-scale RNA comparison based on RNA triple vector curve representationabstractBACKGROUND: In recent years, the important functional roles of RNAs in biological processes have been repeatedly demonstrated. Computing the similarity between two RNAs contributes to better understanding the functional relationship between them. But due to the long-range correlations of RNA, many efficient methods of detecting protein similarity do not work well. In order to comprehensively understand the RNA's function, the better similarity measure among RNAs should be designed to consider their structure features (base pairs). Current methods for RNA comparison could be generally classified into alignment-based and alignment-free. RESULTS: In this paper, we propose a novel wavelet-based method based on RNA triple vector curve representation, named multi-scale RNA comparison. Firstly, we designed a novel numerical representation of RNA secondary structure termed as RNA triple vectors curve (TV-Curve). Secondly, we constructed a new similarity metric based on the wavelet decomposition of the TV-Curve of RNA. Finally we also applied our algorithm to the classification of non-coding RNA and RNA mutation analysis. Furthermore, we compared the results to the two well-known RNA comparison tools: RNAdistance and RNApdist. The results in this paper show the potentials of our method in RNA classification and RNA mutation analysis. CONCLUSION: We provide a better visualization and analysis tool named TV-Curve of RNA, especially for long RNA, which can characterize both sequence and structure features. Additionally, based on TV-Curve representation of RNAs, a multi-scale similarity measure for RNA comparison is proposed, which can capture the local and global difference between the information of sequence and structure of RNAs. Compared with the well-known RNA comparison approaches, the proposed method is validated to be outstanding and effective in terms of non-coding RNA classification and RNA mutation analysis. From the numerical experiments, our proposed method can capture more efficient and subtle relationship of RNAs. Ying Li 0004, Ming Duan, Yanchun Liang 0001 |
BMC Bioinform. | 1 |