EDBT 2026 Demo / reviewers in the wild / expert
Wei Du 0002
dblp:69/870-2
· DBLP profile ↗
32ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-9872-4821ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Whom to Align With: Progressive Anomaly Combination Detection for Partially View-Aligned ClusteringabstractPartially View-aligned Clustering (PVC) addresses the challenge of partial view alignment in multi-view learning by leveraging complementary and consistent information. While existing PVC methods show promise, most rely on distance-based strategies that are sensitive to view-specific details and noise, limiting their robustness. In this work, we propose a novel view alignment strategy that reformulates the alignment task as an anomaly detection problem. Rather than learning a view-alignment matrix that enforces strict one-to-one correspondences across views, we adopt a progressive approach to identify well-aligned samples. Specifically, we sample subsets of data by generating random view combinations from unaligned samples and propose an anomaly combination detection module to evaluate the alignment consistency of these combinations. In addition, our progressive training framework alternates between updating model parameters and selecting high-confidence view combinations for subsequent optimization. By reformulating view alignment as an anomaly detection task, our approach provides a more robust and effective solution to partial view alignment. Experiments on benchmark datasets demonstrate that our method outperforms state-of-the-art approaches in the PVC problem. Hang Gao 0014, Zuosong Cai, Cheng Liu 0001, Ying Li 0004, Wei Du 0002, You Zhou 0008 |
AAAI | 7 |
| 2026 | Efficient Few-Step Solution Generation via Discrete Flow Matching for Combinatorial OptimizationabstractCombinatorial optimization problems (COPs) are fundamental to many real-world applications where efficiently producing high-quality solutions is critical. Recent advances in diffusion-based non-autoregressive models have reformulated solving COPs as a generative process, achieving promising results. However, almost all of these methods still suffer from accumulated errors and high inference costs due to the multi-step stochastic denoising process. To address these issues, we propose EFLOCO, an efficient discrete flow matching method for solving COPs, learning structured and deterministic solution trajectories. EFLOCO replaces noise-driven updates with smooth and guided transitions, thereby improves inference stability and quality. Furthermore, we introduce an adaptive time-step scheduler that makes more efforts in critical transition regions, yielding strong performance under few-step constraints. Experiments on standard Traveling Salesman Problems (TSPs) and Asymmetric TSPs (ATSPs) show that our method consistently outperforms both learning-based and heuristic baselines in terms of solution quality and inference speed. Yuanshu Li, Di Wang 0004, Wei Du 0002, Xuan Wu 0004, Peng Zhao 0018, Yubin Xiao, You Zhou 0008 |
AAAI | 3 |
| 2026 | Multi-level cross-view feature embedding for partial view-aligned clustering
Hang Gao 0014, Cheng Liu 0001, Ying Li 0004, You Zhou 0008, Wei Du 0002 |
Knowl. Based Syst. | 5 |
| 2026 | Cross-view discrepancy-driven dynamic weighting for missing view completion in incomplete multi-view clustering
Hang Gao 0014, Zuosong Cai, Cheng Liu 0001, Ying Li 0004, You Zhou 0008, Wei Du 0002 |
Neural Networks | 7 |
| 2026 | Incomplete multi-view clustering with cross-view generation via pre-trained transformer
Hang Gao 0014, Cheng Liu 0001, Hongming Sun, Ying Li 0004, You Zhou 0008, Wei Du 0002 |
Pattern Recognit. | 7 |
| 2025 | Contrastive Auxiliary Learning with Structure Transformation for Heterogeneous GraphsabstractIn recent years, methods based on heterogeneous graph neural networks (HGNNs) have been widely used for embedding heterogeneous graphs (HGs) due to their ability to effectively encode the rich information from HGs into low-dimensional node embeddings. Existing HGNNs focus on neighbor aggregation and semantic fusion while neglecting the HG structure and learning paradigms. However, the original HG data might lack node features, which existing models may not effectively account for. Additionally, exclusively relying on a single supervised learning approach may only partially leverage the invariant information in graph data. To address these challenges, we introduce the Contrastive Auxiliary Learning Model for Heterogeneous Graphs (CALHG). This model combines edge perturbation and graph diffusion to enhance graph data, allowing it to capture the inherent structural information within heterogeneous graphs fully. Additionally, we employ a category-guided multi-view contrastive learning approach, which does not rely on positive and negative samples for model training, enabling us to capture the intrinsic invariances in heterogeneous graph data. Extensive experiments and analyses on five benchmark datasets without node features and three benchmark datasets with node features demonstrate the effectiveness and efficiency of our novel method compared with several state-of-the-art methods. Wei Du 0002, Hongmin Sun, Hang Gao 0014, Ying Li 0004 |
AAAI | 1 |
| 2025 | Dual Operation Aggregation Graph Neural Networks for Solving Flexible Job-Shop Scheduling Problem with Reinforcement LearningabstractWith the widespread adoption of Internet Protocol (IP) communication technology and web-based platforms, cloud manufacturing has become a significant hallmark of Industry 4.0. Integrating graph algorithms into these web-enabled environments is crucial as they facilitate the representation and analysis of complex relationships in manufacturing processes, enabling efficient decision-making and adaptability in dynamic environments. As a key scheduling problem in cloud manufacturing, the flexible job-shop scheduling problem (FJSP) finds extensive applications in real-world scenarios. However, traditional FJSP-solving methods struggle to meet the efficiency and adaptability demands of cloud manufacturing due to generalization issues and excessive computational time, while reinforcement learning-based methods fail to learn relationships between FJSP nodes, such as interactions between operations of different jobs, leading to limited interpretability and performance. To address these issues, we propose a dual operation aggregation graph neural network (GNN) for solving FJSP. Specifically, we decouple the disjunctive graph into two distinct graphs, reducing graph density and clarifying relationships between machines and operations, thus enabling more effective aggregation and understanding by neural networks. We develop two distinct graph aggregation methods to minimize the influence of non-critical machine and operation nodes on decision-making while enhancing the model's ability to account for long-term benefits. Additionally, to achieve more accurate multi-objective estimation and mitigate reward sparsity, we design a reward function that simultaneously considers machine efficiency, schedule balance, and makespan minimization. Extensive experimental results on well-known datasets demonstrate that our model outperforms state-of-the-art models and exhibits excellent generalization capabilities, effectively addressing the challenges of cloud manufacturing. Peng Zhao 0018, You Zhou 0008, Di Wang 0004, Zhiguang Cao, Yubin Xiao, Xuan Wu 0004, Yuanshu Li, Hongjia Liu, Wei Du 0002, Yuan Jiang 0007, Liupu Wang |
WWW | 9 |
| 2025 | HG-search: multi-stage search for heterogeneous graph neural networks
Hongmin Sun, Ao Kan, Jianhao Liu, Wei Du 0002 |
Appl. Intell. | 4 |
| 2025 | Multi-Manifolds fusing hyperbolic graph network balanced by pareto optimization for identifying spatial domains of spatial transcriptomicsabstractIdentifying spatial domains for spatial transcriptomics is crucial for achieving comprehensive insights into the pathogenesis of gene expression. Increasingly, computational methods based on graph neural networks are being developed for spatial transcriptomics. However, previous methods have solely focused on the Euclidean manifold. To effectively exploit and explore the informative and deeper topological structures of inherent manifolds, we presented a Multi-Manifolds fusing hyperbolic graph network, balanced by Pareto optimization, for identifying spatial domains in Spatial Transcriptomics (MManiST). First, we developed multi-manifolds encoders for distinct manifolds using the hyperbolic neural network. Features from different manifolds were then combined using an attention mechanism, with multiple reconstruction losses balanced by Pareto optimization. Extensive experiments on commonly used benchmark datasets show that our method consistently outperforms seven state-of-the-art methods. Additionally, we investigated the validity of each component and the impact of fusion methods in ablation experiments. Ying Li 0004, Qifeng Hu, Rui Wang-Sattler, Wei Du 0002 |
Briefings Bioinform. | 5 |
| 2025 | A multi-omics integration framework using multi-label guided learning and multi-scale fusionabstractThe rapid development of high-throughput sequencing technologies has generated vast amounts of omics data, making multi-omics integration a crucial approach for understanding complex diseases. Despite the introduction of various multi-omics integration methods in recent years, existing approaches still have limitations, primarily in their reliance on manual feature selection, restricted applicability, and inability to comprehensively capture both inter-sample and cross-omics interactions. To address these challenges, we propose mmMOI, an end-to-end multi-omics integration framework that incorporates multi-label guided learning and multi-scale attention fusion. mmMOI directly processes raw high-dimensional omics data without requiring manual feature selection, thereby enhancing model interpretability and eliminating biases introduced by feature preselection. First, we introduce a multi-label guided multi-view graph neural network, which enables the model to adaptively learn omics data representations across different datasets, thereby improving generalizability and stability. Second, we design a multi-scale attention fusion network, which integrates global attention and local attention. This dual-attention mechanism allows mmMOI to more accurately integrate multi-omics data, enhance cross-omics feature representations, and improve classification performance. Experimental results demonstrate that mmMOI significantly outperforms state-of-the-art methods in classification tasks, exhibiting high stability and adaptability across diverse biological contexts and sequencing technologies. Additionally, mmMOI successfully identifies key disease-associated biomarkers, further enhancing its biological interpretability and practical relevance. The source code, datasets, and detailed hyperparameter configurations for mmMOI are available at https://github.com/mlcb-jlu/mmMOI. Yinghe Wang, Ying Li 0004, Wei Du 0002 |
Briefings Bioinform. | 5 |
| 2025 | Cross-loop contrast of heterogeneous graphs for interdisciplinary journal recommendation
Ying Li 0004, Hongmin Sun, Wei Du 0002, Qin Ma 0003 |
Expert Syst. Appl. | 5 |
| 2025 | ICPPNet: A semantic segmentation network model based on inter-class positional prior for scoliosis reconstruction in ultrasound imagesabstractOBJECTIVE: Considering the radiation hazard of X-ray, safer, more convenient and cost-effective ultrasound methods are gradually becoming new diagnostic approaches for scoliosis. For ultrasound images of spine regions, it is challenging to accurately identify spine regions in images due to relatively small target areas and the presence of a lot of interfering information. Therefore, we developed a novel neural network that incorporates prior knowledge to precisely segment spine regions in ultrasound images. MATERIALS AND METHODS: We constructed a dataset of ultrasound images of spine regions for semantic segmentation. The dataset contains 3136 images of 30 patients with scoliosis. And we propose a network model (ICPPNet), which fully utilizes inter-class positional prior knowledge by combining an inter-class positional probability heatmap, to achieve accurate segmentation of target areas. RESULTS: ICPPNet achieved an average Dice similarity coefficient of 70.83% and an average 95% Hausdorff distance of 11.28 mm on the dataset, demonstrating its excellent performance. The average error between the Cobb angle measured by our method and the Cobb angle measured by X-ray images is 1.41 degrees, and the coefficient of determination is 0.9879 with a strong correlation. DISCUSSION AND CONCLUSION: ICPPNet provides a new solution for the medical image segmentation task with positional prior knowledge between target classes. And ICPPNet strongly supports the subsequent reconstruction of spine models using ultrasound images. You Zhou 0008, Yuanshu Li, Wei Pang 0001, Liupu Wang, Wei Du 0002, Hui Yang 0015 |
J. Biomed. Informatics | 6 |
| 2025 | Improving generalization of neural Vehicle Routing Problem solvers through the lens of model architecture
Yubin Xiao, Di Wang 0004, Xuan Wu 0004, Yuesong Wu, Boyang Li 0001, Wei Du 0002, Liupu Wang, You Zhou 0008 |
Neural Networks | 6 |
| 2025 | A Novel Approach for Effective Partially View-Aligned Clustering With Triple-ConsistencyabstractMulti-view clustering (MVC), which integrates information from multiple views to enhance performance, has garnered increasing attention in recent years. Partially View-aligned Clustering (PVC), which is a particularly critical aspect of this process, requires a thorough exploration of complementary and consistent information under conditions of partial view alignment. However, most existing PVC methods primarily focus on semantic consistency, employing semantic consistency features for both view alignment and clustering tasks. These methods neglect the effects of noise and complementary information across multiple views and the suitability of these features for clustering. To address these limitations, our approach aims to leverage three distinct types of consistency to extract semantic consistency features and clustering consistency features, which are specifically designed for view alignment and clustering tasks, respectively. By omitting the reconstruction process, we mitigate the adverse effects of mutual information and noise on view alignment. Specifically, we first exploit the structural consistency of similarity graphs across different views to guide feature extraction in view-specific autoencoders. This process produces structural consistency features that are both cluster-discriminative and structurally coherent. Subsequently, two separate multilayer perceptrons (MLPs) are trained via contrastive learning to extract semantic consistency features and clustering consistency features from the structural features. These features are optimized for their respective tasks. Ultimately, a self-paced style view alignment strategy is used to iteratively re-align the data based on semantic and clustering consistency while the model is optimized via the re-aligned data. Extensive experiments on multiple real-world benchmark datasets demonstrate that our method outperforms the state-of-the-art multi-view approaches, highlighting its effectiveness in tackling the challenges of PVC. The code is available at https://github.com/kongyiH/TCLPVC. Hang Gao 0014, Cheng Liu 0001, Zuosong Cai, Hongming Sun, Ying Li 0004, Wei Du 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2023 | PecidRL: Petition expectation correction and identification based on deep reinforcement learning
Ying Li 0004, Wensi Fang, Wei Du 0002 |
Inf. Process. Manag. | 5 |
| 2023 | SURE: Screening unlabeled samples for reliable negative samples based on reinforcement learning
Ying Li 0004, Wensi Fang, Qin Ma 0003, Rui Wang-Sattler, Wei Du 0002, Qiong Yu |
Inf. Sci. | 7 |
| 2022 | LPInsider: a webserver for lncRNA-protein interaction extraction from the literatureabstractBACKGROUND: Long non-coding RNA (LncRNA) plays important roles in physiological and pathological processes. Identifying LncRNA-protein interactions (LPIs) is essential to understand the molecular mechanism and infer the functions of lncRNAs. With the overwhelming size of the biomedical literature, extracting LPIs directly from the biomedical literature is essential, promising and challenging. However, there is no webserver of LPIs relationship extraction from literature. RESULTS: LPInsider is developed as the first webserver for extracting LPIs from biomedical literature texts based on multiple text features (semantic word vectors, syntactic structure vectors, distance vectors, and part of speech vectors) and logistic regression. LPInsider allows researchers to extract LPIs by uploading PMID, PMCID, PMID List, or biomedical text. A manually filtered and highly reliable LPI corpus is integrated in LPInsider. The performance of LPInsider is optimal by comprehensive experiment on different combinations of different feature and machine learning models. CONCLUSIONS: LPInsider is an efficient analytical tool for LPIs that helps researchers to enhance their comprehension of lncRNAs from text mining, and also saving their time. In addition, LPInsider is freely accessible from http://www.csbg-jlu.info/LPInsider/ with no login requirement. The source code and LPIs corpus can be downloaded from https://github.com/qiufengdiewu/LPInsider . Ying Li 0004, Lizheng Wei, Cankun Wang, Wei Du 0002 |
BMC Bioinform. | 7 |
| 2021 | Combining GCN and Bi-LSTM for Protein Secondary Structure PredictionabstractProtein secondary structure prediction is still a challenging task in bioinformatics, especially for 8-state (Q8) classification. To address this problem, we have proposed a deep learning based model by integrating graph convolutional network(GCN) and bidirectional long short-term memory (Bi-LSTM) network in this paper. In the model, GCN is utilized to synthesize the information of amino acids and their interactions, while Bi-LSTM has strong ability to capture the long-range dependencies of amino acids. For sequence representation, a new protein embedding derived by ProtTrans is used instead of the traditional amino acid one-hot encoding, together with evolutionary features of PSSM and HHM profiles. Amino acid contact potential derived from SPOTContact-Helical is used to construct amino acid graph. To verify the effectiveness of our proposed model, it is applied to several benchmark datasets, and obtained 78.05%, 76.81% 72.84%, 74.46% and 76.04% Q8 accuracy on CASP10, CASP11, CASP12, CB513 and TS115 datasets, respectively. Compared with 8 state-of-the-art competitions, our model obtained the best performance in most of datasets. Hailong Jin, Wei Du 0002, Jiawei Gu, Xiaohu Shi |
BIBM | 2 |
| 2021 | CyanoPATH: a knowledgebase of genome-scale functional repertoire for toxic cyanobacterial bloomsabstractCyanoPATH is a database that curates and analyzes the common genomic functional repertoire for cyanobacteria harmful algal blooms (CyanoHABs) in eutrophic waters. Based on the literature of empirical studies and genome/protein databases, it summarizes four types of information: common biological functions (pathways) driving CyanoHABs, customized pathway maps, classification of blooming type based on databases and the genomes of cyanobacteria. A total of 19 pathways are reconstructed, which are involved in the utilization of macronutrients (e.g. carbon, nitrogen, phosphorus and sulfur), micronutrients (e.g. zinc, magnesium, iron, etc.) and other resources (e.g. light and vitamins) and in stress resistance (e.g. lead and copper). These pathways, comprised of both transport and biochemical reactions, are reconstructed with proteins from NCBI and reactions from KEGG and visualized with self-created transport/reaction maps. The pathways are hierarchical and consist of subpathways, protein/enzyme complexes and constituent proteins. New cyanobacterial genomes can be annotated and visualized for these pathways and compared with existing species. This set of genomic functional repertoire is useful in analyzing aquatic metagenomes and metatranscriptomes in CyanoHAB research. Most importantly, it establishes a link between genome and ecology. All these reference proteins, pathways and maps and genomes are free to download at http://www.csbg-jlu.info/CyanoPATH. Wei Du 0002, Nicholas Ho, Landon Jenkins, Drew Hockaday, Jiankang Tan, Huansheng Cao |
Briefings Bioinform. | 1 |
| 2021 | Deep forest ensemble learning for classification of alignments of non-coding RNA sequences based on multi-view structure representationsabstractNon-coding RNAs (ncRNAs) play crucial roles in multiple biological processes. However, only a few ncRNAs' functions have been well studied. Given the significance of ncRNAs classification for understanding ncRNAs' functions, more and more computational methods have been introduced to improve the classification automatically and accurately. In this paper, based on a convolutional neural network and a deep forest algorithm, multi-grained cascade forest (GcForest), we propose a novel deep fusion learning framework, GcForest fusion method (GCFM), to classify alignments of ncRNA sequences for accurate clustering of ncRNAs. GCFM integrates a multi-view structure feature representation including sequence-structure alignment encoding, structure image representation and shape alignment encoding of structural subunits, enabling us to capture the potential specificity between ncRNAs. For the classification of pairwise alignment of two ncRNA sequences, the F-value of GCFM improves 6% than an existing alignment-based method. Furthermore, the clustering of ncRNA families is carried out based on the classification matrix generated from GCFM. Results suggest better performance (with 20% accuracy improved) than existing ncRNA clustering methods (RNAclust, Ensembleclust and CNNclust). Additionally, we apply GCFM to construct a phylogenetic tree of ncRNA and predict the probability of interactions between RNAs. Most ncRNAs are located correctly in the phylogenetic tree, and the prediction accuracy of RNA interaction is 90.63%. A web server (http://bmbl.sdstate.edu/gcfm/) is developed to maximize its availability, and the source code and related data are available at the same URL. Ying Li 0004, Qi Zhang 0061, Zhaoqian Liu, Cankun Wang, Qin Ma 0003, Wei Du 0002 |
Briefings Bioinform. | 7 |
| 2021 | Capsule-LPI: a LncRNA-protein interaction predicting tool based on a capsule networkabstractBACKGROUND: Long noncoding RNAs (lncRNAs) play important roles in multiple biological processes. Identifying LncRNA-protein interactions (LPIs) is key to understanding lncRNA functions. Although some LPIs computational methods have been developed, the LPIs prediction problem remains challenging. How to integrate multimodal features from more perspectives and build deep learning architectures with better recognition performance have always been the focus of research on LPIs. RESULTS: We present a novel multichannel capsule network framework to integrate multimodal features for LPI prediction, Capsule-LPI. Capsule-LPI integrates four groups of multimodal features, including sequence features, motif information, physicochemical properties and secondary structure features. Capsule-LPI is composed of four feature-learning subnetworks and one capsule subnetwork. Through comprehensive experimental comparisons and evaluations, we demonstrate that both multimodal features and the architecture of the multichannel capsule network can significantly improve the performance of LPI prediction. The experimental results show that Capsule-LPI performs better than the existing state-of-the-art tools. The precision of Capsule-LPI is 87.3%, which represents a 1.7% improvement. The F-value of Capsule-LPI is 92.2%, which represents a 1.4% improvement. CONCLUSIONS: This study provides a novel and feasible LPI prediction tool based on the integration of multimodal features and a capsule network. A webserver ( http://csbg-jlu.site/lpc/predict ) is developed to be convenient for users. Ying Li 0004, Shiyao Feng, Qi Zhang 0061, Wei Du 0002 |
BMC Bioinform. | 6 |
| 2021 | DeepHBSP: A Deep Learning Framework for Predicting Human Blood-Secretory Proteins Using Transfer Learning
Wei Du 0002, Hui-Min Bao, Liang Chen 0021, Ying Li 0004, Yanchun Liang 0001 |
J. Comput. Sci. Technol. | 1 |
| 2020 | CapsNet-SSP: multilane capsule network for predicting human saliva-secretory proteinsabstractBACKGROUND: Compared with disease biomarkers in blood and urine, biomarkers in saliva have distinct advantages in clinical tests, as they can be conveniently examined through noninvasive sample collection. Therefore, identifying human saliva-secretory proteins and further detecting protein biomarkers in saliva have significant value in clinical medicine. There are only a few methods for predicting saliva-secretory proteins based on conventional machine learning algorithms, and all are highly dependent on annotated protein features. Unlike conventional machine learning algorithms, deep learning algorithms can automatically learn feature representations from input data and thus hold promise for predicting saliva-secretory proteins. RESULTS: We present a novel end-to-end deep learning model based on multilane capsule network (CapsNet) with differently sized convolution kernels to identify saliva-secretory proteins only from sequence information. The proposed model CapsNet-SSP outperforms existing methods based on conventional machine learning algorithms. Furthermore, the model performs better than other state-of-the-art deep learning architectures mostly used to analyze biological sequences. In addition, we further validate the effectiveness of CapsNet-SSP by comparison with human saliva-secretory proteins from existing studies and known salivary protein biomarkers of cancer. CONCLUSIONS: The main contributions of this study are as follows: (1) an end-to-end model based on CapsNet is proposed to identify saliva-secretory proteins from the sequence information; (2) the proposed model achieves better performance and outperforms existing models; and (3) the saliva-secretory proteins predicted by our model are statistically significant compared with existing cancer biomarkers in saliva. In addition, a web server of CapsNet-SSP is developed for saliva-secretory protein identification, and it can be accessed at the following URL: http://www.csbg-jlu.info/CapsNet-SSP/. We believe that our model and web server will be useful for biomedical researchers who are interested in finding salivary protein biomarkers, especially when they have identified candidate proteins for analyzing diseased tissues near or distal to salivary glands using transcriptome or proteomics. Wei Du 0002, Huansheng Cao, Ying Li 0004 |
BMC Bioinform. | 1 |
| 2019 | LncFinder: an integrated platform for long non-coding RNA identification utilizing sequence intrinsic composition, structural information and physicochemical propertyabstractDiscovering new long non-coding RNAs (lncRNAs) has been a fundamental step in lncRNA-related research. Nowadays, many machine learning-based tools have been developed for lncRNA identification. However, many methods predict lncRNAs using sequence-derived features alone, which tend to display unstable performances on different species. Moreover, the majority of tools cannot be re-trained or tailored by users and neither can the features be customized or integrated to meet researchers' requirements. In this study, features extracted from sequence-intrinsic composition, secondary structure and physicochemical property are comprehensively reviewed and evaluated. An integrated platform named LncFinder is also developed to enhance the performance and promote the research of lncRNA identification. LncFinder includes a novel lncRNA predictor using the heterologous features we designed. Experimental results show that our method outperforms several state-of-the-art tools on multiple species with more robust and satisfactory results. Researchers can additionally employ LncFinder to extract various classic features, build classifier with numerous machine learning algorithms and evaluate classifier performance effectively and efficiently. LncFinder can reveal the properties of lncRNA and mRNA from various perspectives and further inspire lncRNA-protein interaction prediction and lncRNA evolution analysis. It is anticipated that LncFinder can significantly facilitate lncRNA-related research, especially for the poorly explored species. LncFinder is released as R package (https://CRAN.R-project.org/package=LncFinder). A web server (http://bmbl.sdstate.edu/lncfinder/) is also developed to maximize its availability. Yanchun Liang 0001, Qin Ma 0003, Yangyi Xu, Wei Du 0002, Cankun Wang, Ying Li 0004 |
Briefings Bioinform. | 6 |
| 2016 | PUEPro: A Computational Pipeline for Prediction of Urine Excretory Proteins
Yan Wang 0028, Wei Du 0002, Yanchun Liang 0001, Xin Chen 0113, Chi Zhang 0021, Wei Pang 0001, Ying Xu 0001 |
ADMA | 2 |
| 2014 | Computational Prediction of Human Saliva-Secreted Proteins
Chunguang Zhou, Zhongbo Cao, Wei Du 0002, Yan Wang 0028 |
ISBRA | 5 |
| 2013 | Effective and stable feature selection method based on filter for gene signature identification in paired microarray dataabstractA huge amount of microarray datasets are produced with big number of genes and small samples. Feature selection methods have become a very sharp tool to select the gene signatures from the whole gene set. In recent years, researchers are concerned much about the datasets containing samples of cancer as well as corresponding control tissues. However, few feature selection methods consider the effect of paired samples. In this article, we propose a new feature selection method for paired microarray datasets based on the original paired t-test approach. We apply on the paired datasets across six common cancer types. Through comparison with some widely used methods on the performance of prediction power, stability of gene lists and functional stability, our method shows excellent performance. The proposed method has good effectiveness, stability and consistency, which enables the method to be applicative to feature selection for paired microarray expression data analysis. Zhongbo Cao, Yan Wang 0028, Wei Du 0002, Yanchun Liang 0001 |
BIBM | 4 |
| 2013 | PMTED: a plant microRNA target expression databaseabstractBACKGROUND: MicroRNAs (miRNAs) are identified in nearly all plants where they play important roles in development and stress responses by target mRNA cleavage or translation repression. MiRNAs exert their functions by sequence complementation with target genes and hence their targets can be predicted using bioinformatics algorithms. In the past two decades, microarray technology has been employed to study genes involved in important biological processes such as biotic response, abiotic response, and specific tissues and developmental stages, many of which are miRNA targets. Despite their value in assisting research work for plant biologists, miRNA target genes are difficult to access without pre-processing and assistance of necessary analytical and visualization tools because they are embedded in a large body of microarray data that are scattered around in public databases. DESCRIPTION: Plant MiRNA Target Expression Database (PMTED) is designed to retrieve and analyze expression profiles of miRNA targets represented in the plethora of existing microarray data that are manually curated. It provides a Basic Information query function for miRNAs and their target sequences, gene ontology, and differential expression profiles. It also provides searching and browsing functions for a global Meta-network among species, bioprocesses, conditions, and miRNAs, meta-terms curated from well annotated microarray experiments. Networks are displayed through a Cytoscape Web-based graphical interface. In addition to conserved miRNAs, PMTED provides a target prediction portal for user-defined novel miRNAs and corresponding target expression profile retrieval. Hypotheses that are suggested by miRNA-target networks should provide starting points for further experimental validation. CONCLUSIONS: PMTED exploits value-added microarray data to study the contextual significance of miRNA target genes and should assist functional investigation for both miRNAs and their targets. PMTED will be updated over time and is freely available for non-commercial use at http://pmted.agrinome.org. Xiuli Sun, Boquan Dong, Lingjie Yin, Rongzhi Zhang, Wei Du 0002, Dongfeng Liu, Nan Shi, Aili Li, Yanchun Liang 0001, Long Mao |
BMC Bioinform. | 5 |
| 2009 | Immune Particle Swarm Optimization for Support Vector Regression on Forest Fire Prediction
Yan Wang 0028, Juexin Wang, Wei Du 0002, Chuncai Wang, Yanchun Liang 0001, Chunguang Zhou, Lan Huang 0002 |
ISNN (2) | 3 |
| 2009 | Methods for labeling error detection in microarrays based on the effect of data perturbation on the regression modelabstractMOTIVATION: Mislabeled samples often appear in gene expression profile because of the similarity of different sub-type of disease and the subjective misdiagnosis. The mislabeled samples deteriorate supervised learning procedures. The LOOE-sensitivity algorithm is an approach for mislabeled sample detection for microarray based on data perturbation. However, the failure of measuring the perturbing effect makes the LOOE-sensitivity algorithm a poor performance. The purpose of this article is to design a novel detection method for mislabeled samples of microarray, which could take advantage of the measuring effect of data perturbations. RESULTS: To measure the effect of data perturbation, we define an index named perturbing influence value (PIV), based on the support vector machine (SVM) regression model. The Column Algorithm (CAPIV), Row Algorithm (RAPIV) and progressive Row Algorithm (PRAPIV) based on the PIV value are proposed to detect the mislabeled samples. Experimental results obtained by using six artificial datasets and five microarray datasets demonstrate that all proposed methods in this article are superior to LOOE-sensitivity. Moreover, compared with the simple SVM and CL-stability, the PRAPIV algorithm shows an increase in precision and high recall. AVAILABILITY: The program and source code (in JAVA) are publicly available at http://ccst.jlu.edu.cn/CSBG/PIVS/index.htm Chunguo Wu, Enrico Blanzieri, You Zhou 0008, Yan Wang 0028, Wei Du 0002, Yanchun Liang 0001 |
Bioinform. | 6 |
| 2007 | Operon Prediction Using Neural Network Based on Multiple Information of Log-Likelihoods
Wei Du 0002, Yan Wang 0028, Fangxun Sun, Chunguang Zhou, Chengquan Hu, Yanchun Liang 0001 |
ISNN (1) | 1 |
| 2007 | A multi-approaches-guided genetic algorithm with application to operon prediction
Yan Wang 0028, Wei Du 0002, Fangxun Sun, Chunguang Zhou, Yanchun Liang 0001 |
Artif. Intell. Medicine | 3 |