VLDB 2026 Research / reviewers in the wild / expert
Shan He 0001
dblp:14/1570-1
· DBLP profile ↗
53ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0003-1694-1465ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 22 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowledge-guided hyper-heuristic evolutionary algorithm for large-scale Boolean network inference
Xiang Liu 0019, Yan Wang 0049, Xiayu Jiang, Shan He 0001 |
Expert Syst. Appl. | 5 |
| 2025 | TG-CDDPM: text-guided antimicrobial peptides generation based on conditional denoising diffusion probabilistic modelabstractAntimicrobial peptides (AMPs) have emerged as a promising substitution to antibiotics thanks to their boarder range of activities, less likelihood of drug resistance, and low toxicity. Traditional biochemical methods for AMP discovery are costly and inefficient. Deep generative models, including the long-short term memory model, variational autoencoder model, and generative adversarial model, have been widely introduced to expedite AMP discovery. However, these models tend to suffer from the lack of diversity in generating AMPs. The denoising diffusion probabilistic model serves as a good candidate for solving this issue. We proposed a three-stage Text-Guided Conditional Denoising Diffusion Probabilistic Model (TG-CDDPM) to generate novel and homologous AMPs. In the first two stages, contrastive learning and inferring models are crafted to create better conditions for guiding AMP generation, respectively. In the last stage, a pre-trained conditional denoising diffusion probabilistic model is leveraged to enrich the peptide knowledge and fine-tuned to learn feature representation in downstream. TG-CDDPM was compared to the state-of-the-art generative models for AMP generation, and it demonstrated competitive or better performance with the assistance of text description as supervised information. The membrane penetration capabilities of the identified candidate AMPs by TG-CDDPM were also validated through molecular weight dynamics experiments. Junhang Cao, Jun Zhang 0078, Qiyuan Yu, Junkai Ji, Jianqiang Li 0001, Shan He 0001, Zexuan Zhu 0001 |
Briefings Bioinform. | 6 |
| 2025 | A Survey on Evolutionary Computation-Based Drug DiscoveryabstractDrug discovery is an expensive and risky process. To combat the challenges in drug discovery, an increasing number of researchers and pharmaceutical companies recognize the benefits of utilizing computational techniques. Evolutionary computation (EC) offers promise as most drug discovery problems are essentially complex optimization problems beyond conventional optimization algorithms. EC methods have been widely applied to solve these complex optimization problems especially in lead com-pound generation and molecular virtual evaluation, substantially speeding up the process of drug discovery and development. This article presents a comprehensive survey of EC-based drug discovery methods. Particularly, a new taxonomy of the methods is provided and the advantages and limitations of the methods are reviewed. In addition, the potential future directions of EC-based drug discovery are discussed and the publicly available resources including databases and computational tools are compiled for the convenience of researchers seeking to pursue this field. Qiyuan Yu, Qiuzhen Lin, Junkai Ji, Wei Zhou 0001, Shan He 0001, Zexuan Zhu 0001, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 5 |
| 2025 | MDTL-ACP: Anticancer Peptides Prediction Based on Multi-Domain Transfer LearningabstractAnticancer peptides (ACPs) have emerged as one of the most promising therapeutic agents for cancer treatment. They are bioactive peptides featuring broad-spectrum activity and low drug-resistance. The discovery of ACPs via traditional biochemical methods is laborious and costly. Accordingly, various computational methods have been developed to facilitate the discovery of ACPs. However, the data resources and knowledge of ACPs are still very scarce, and only a few of them are clinically verified, which limits the competence of computational methods. To address this issue, in this article, we propose an ACP prediction model based on multi-domain transfer learning, namely MDTL-ACP, to discriminate novel ACPs from plentiful inactive peptides. In particular, we collect abundant antimicrobial peptides (AMPs) from four well-studied peptide domains and extract their inherent features as the input of MDTL-ACP. The features learned from multiple source domains of AMPs are then transferred into the target prediction task of ACPs via artificial neural network-based shared-extractor and task-specific classifiers in MDTL-ACP. The knowledge captured in the transferred features enhances the prediction of ACPs in the target domain. Experimental results demonstrate that MDTL-ACP can outperform the traditional and state-of-the-art ACP prediction methods. Junhang Cao, Wei Zhou 0001, Qiyuan Yu, Junkai Ji, Jun Zhang 0078, Shan He 0001, Zexuan Zhu 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | MIFuGP: Boolean network inference from multivariate time series using fuzzy genetic programming
Xiang Liu 0019, Yan Wang 0049, Shan He 0001 |
Inf. Sci. | 5 |
| 2023 | Perturbation-Based Two-Stage Multi-Domain Active LearningabstractIn multi-domain learning (MDL) scenarios, high labeling effort is required due to the complexity of collecting data from various domains. Active Learning (AL) presents an encouraging solution to this issue by annotating a smaller number of highly informative instances, thereby reducing the labeling effort. Previous research has relied on conventional AL strategies for MDL scenarios, which underutilize the domain-shared information of each instance during the selection procedure. To mitigate this issue, we propose a novel perturbation-based two-stage multi-domain active learning (P2S-MDAL) method incorporated into the well-regarded ASP-MTL model. Specifically, P2S-MDAL involves allocating budgets for domains and establishing regions for diversity selection, which are further used to select the most cross-domain influential samples in each region. A perturbation metric has been introduced to evaluate the robustness of the shared feature extractor of the model, facilitating the identification of potentially cross-domain influential samples. Experiments are conducted on three real-world datasets, encompassing both texts and images. The superior performance over conventional AL strategies shows the effectiveness of the proposed strategy. Additionally, an ablation study has been carried out to demonstrate the validity of each component. Finally, we outline several intriguing potential directions for future MDAL research, thus catalyzing the field's advancement. Zeyu Dai 0001, Shan He 0001, Ke Tang 0001 |
CIKM | 3 |
| 2023 | Multi-Domain Learning from Insufficient AnnotationsabstractMulti-domain learning (MDL) refers to simultaneously constructing a model or a set of models on datasets collected from different domains. Conventional approaches emphasize domain-shared information extraction and domain-private information preservation, following the shared-private framework (SP models), which offers significant advantages over single-domain learning. However, the limited availability of annotated data in each domain considerably hinders the effectiveness of conventional supervised MDL approaches in real-world applications. In this paper, we introduce a novel method called multi-domain contrastive learning (MDCL) to alleviate the impact of insufficient annotations by capturing both semantic and structural information from both labeled and unlabeled data. Specifically, MDCL comprises two modules: inter-domain semantic alignment and intra-domain contrast. The former aims to align annotated instances of the same semantic category from distinct domains within a shared hidden space, while the latter focuses on learning a cluster structure of unlabeled instances in a private hidden space for each domain. MDCL is readily compatible with many SP models, requiring no additional model parameters and allowing for end-to-end training. Experimental results across five textual and image multi-domain datasets demonstrate that MDCL brings noticeable improvement over various SP models. Furthermore, MDCL can further be employed in multi-domain active learning (MDAL) to achieve a superior initialization, eventually leading to better overall performance. Shengcai Liu, Jiahao Wu 0004, Shan He 0001, Ke Tang 0001 |
ECAI | 4 |
| 2023 | Improving therapeutic synergy score predictions with adverse effects using multi-task heterogeneous network learningabstractDrug combinations could trigger pharmacological therapeutic effects (TEs) and adverse effects (AEs). Many computational methods have been developed to predict TEs, e.g. the therapeutic synergy scores of anti-cancer drug combinations, or AEs from drug-drug interactions. However, most of the methods treated the AEs and TEs predictions as two separate tasks, ignoring the potential mechanistic commonalities shared between them. Based on previous clinical observations, we hypothesized that by learning the shared mechanistic commonalities between AEs and TEs, we could learn the underlying MoAs (mechanisms of actions) and ultimately improve the accuracy of TE predictions. To test our hypothesis, we formulated the TE prediction problem as a multi-task heterogeneous network learning problem that performed TE and AE learning tasks simultaneously. To solve this problem, we proposed Muthene (multi-task heterogeneous network embedding) and evaluated it on our collected drug-drug interaction dataset with both TEs and AEs indications. Our experimental results showed that, by including the AE prediction as an auxiliary task, Muthene generated more accurate TE predictions than standard single-task learning methods, which supports our hypothesis. Using a drug pair Vincristine-Dasatinib as a case study, we demonstrated that our method not only provides a novel way of TE predictions but also helps us gain a deeper understanding of the MoAs of drug combinations. Yang Yue 0008, Yongxuan Liu, Luoying Hao, Huangshu Lei, Shan He 0001 |
Briefings Bioinform. | 5 |
| 2023 | MpbPPI: a multi-task pre-training-based equivariant approach for the prediction of the effect of amino acid mutations on protein-protein interactionsabstractThe accurate prediction of the effect of amino acid mutations for protein-protein interactions (PPI $\Delta \Delta G$) is a crucial task in protein engineering, as it provides insight into the relevant biological processes underpinning protein binding and provides a basis for further drug discovery. In this study, we propose MpbPPI, a novel multi-task pre-training-based geometric equivariance-preserving framework to predict PPI $\Delta \Delta G$. Pre-training on a strictly screened pre-training dataset is employed to address the scarcity of protein-protein complex structures annotated with PPI $\Delta \Delta G$ values. MpbPPI employs a multi-task pre-training technique, forcing the framework to learn comprehensive backbone and side chain geometric regulations of protein-protein complexes at different scales. After pre-training, MpbPPI can generate high-quality representations capturing the effective geometric characteristics of labeled protein-protein complexes for downstream $\Delta \Delta G$ predictions. MpbPPI serves as a scalable framework supporting different sources of mutant-type (MT) protein-protein complexes for flexible application. Experimental results on four benchmark datasets demonstrate that MpbPPI is a state-of-the-art framework for PPI $\Delta \Delta G$ predictions. The data and source code are available at https://github.com/arantir123/MpbPPI. Yang Yue 0008, Huanxiang Liu, Henry H. Y. Tong, Shan He 0001 |
Briefings Bioinform. | 6 |
| 2023 | MIX-TPI: a flexible prediction framework for TCR-pMHC interactions based on multimodal representationsabstractMOTIVATION: The interactions between T-cell receptors (TCR) and peptide-major histocompatibility complex (pMHC) are essential for the adaptive immune system. However, identifying these interactions can be challenging due to the limited availability of experimental data, sequence data heterogeneity, and high experimental validation costs. RESULTS: To address this issue, we develop a novel computational framework, named MIX-TPI, to predict TCR-pMHC interactions using amino acid sequences and physicochemical properties. Based on convolutional neural networks, MIX-TPI incorporates sequence-based and physicochemical-based extractors to refine the representations of TCR-pMHC interactions. Each modality is projected into modality-invariant and modality-specific representations to capture the uniformity and diversities between different features. A self-attention fusion layer is then adopted to form the classification module. Experimental results demonstrate the effectiveness of MIX-TPI in comparison with other state-of-the-art methods. MIX-TPI also shows good generalization capability on mutual exclusive evaluation datasets and a paired TCR dataset. AVAILABILITY AND IMPLEMENTATION: The source code of MIX-TPI and the test data are available at: https://github.com/Wolverinerine/MIX-TPI. Zhi-an Huang, Wei Zhou 0001, Junkai Ji, Jun Zhang 0078, Shan He 0001, Zexuan Zhu 0001 |
Bioinform. | 6 |
| 2023 | Network Biomarker Detection From Gene Co-Expression Network Using Gaussian Mixture Model ClusteringabstractFinding network biomarkers from gene co-expression networks (GCNs) has attracted a lot of research interest. A network biomarker is a topological module, i.e., a group of densely connected nodes in a GCN, in which the gene expression values correlate with sample labels. Compared with biomarkers based on single genes, network biomarkers are not only more robust in separating samples from different categories, but are also able to better interpret the molecular mechanism of the disease. The previous network biomarker detection methods either employ distance based clustering methods or search for cliques in a GCN to detect topological modules. The first strategy assumes that the topological modules should be spherical in shape, and the second strategy requires all nodes to be fully connected. However, the relations between genes are complex, as a result, genes in the same biological process may not be directly, strongly connected. Therefore, the shapes of those modules could be oval or long strips. Hence, the shapes of gene functional modules and gene disease modules may not meet the aforementioned constraints in the previous methods. Thus, previous methods may break up the genes belonging to the same biological process into different topological modules due to those constraints. To address this issue, we propose a novel network biomarker detection method by using Gaussian mixture model clustering which allows more flexibility in the shapes of the topological modules. We have evaluated the performance of our method on a set of eight TCGA cancer datasets. The results show that our method can detect network modules that possess better discriminate power, and provide biological insights. Zexuan Zhu 0001, Hui Li 0117, Shan He 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | GloDyNE: Global Topology Preserving Dynamic Network Embedding (Extended Abstract)abstractDynamic Network Embedding (DNE) is attracting much attention due to the time-evolving nature of many real-world networks. The main objective of DNE is to efficiently update node embeddings while preserving network topology at each timestep. The idea of most existing DNE methods is to capture the topological changes at or around the most affected nodes (instead of all nodes) and accordingly update node embeddings. Unfortunately, this kind of approximation, although can improve efficiency, cannot effectively preserve the global topology of a dynamic network at each timestep, due to not considering the inactive sub-networks that receive accumulated topological changes propagated via the high-order proximity. To address this issue, we propose a new DNE method for better global topology preservation. Extensive experiments demonstrate the effectiveness and efficiency of the proposed method. Chengbin Hou, Shan He 0001, Ke Tang 0001 |
ICDE | 3 |
| 2022 | MPVNN: Mutated Pathway Visible Neural Network architecture for interpretable prediction of cancer-specific survival riskabstractMOTIVATION: Survival risk prediction using gene expression data is important in making treatment decisions in cancer. Standard neural network (NN) survival analysis models are black boxes with a lack of interpretability. More interpretable visible neural network architectures are designed using biological pathway knowledge. But they do not model how pathway structures can change for particular cancer types. RESULTS: We propose a novel Mutated Pathway Visible Neural Network (MPVNN) architecture, designed using prior signaling pathway knowledge and random replacement of known pathway edges using gene mutation data simulating signal flow disruption. As a case study, we use the PI3K-Akt pathway and demonstrate overall improved cancer-specific survival risk prediction of MPVNN over other similar-sized NN and standard survival analysis methods. We show that trained MPVNN architecture interpretation, which points to smaller sets of genes connected by signal flow within the PI3K-Akt pathway that is important in risk prediction for particular cancer types, is reliable. AVAILABILITY AND IMPLEMENTATION: The data and code are available at https://github.com/gourabghoshroy/MPVNN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gourab Ghosh Roy, Nicholas Geard, Karin Verspoor, Shan He 0001 |
Bioinform. | 4 |
| 2022 | CURC: a CUDA-based reference-free read compressorabstractMOTIVATION: The data deluge of high-throughput sequencing (HTS) has posed great challenges to data storage and transfer. Many specific compression tools have been developed to solve this problem. However, most of the existing compressors are based on central processing unit (CPU) platform, which might be inefficient and expensive to handle large-scale HTS data. With the popularization of graphics processing units (GPUs), GPU-compatible sequencing data compressors become desirable to exploit the computing power of GPUs. RESULTS: We present a GPU-accelerated reference-free read compressor, namely CURC, for FASTQ files. Under a GPU-CPU heterogeneous parallel scheme, CURC implements highly efficient lossless compression of DNA stream based on the pseudogenome approach and CUDA library. CURC achieves 2-6-fold speedup of the compression with competitive compression rate, compared with other state-of-the-art reference-free read compressors. AVAILABILITY AND IMPLEMENTATION: CURC can be downloaded from https://github.com/BioinfoSZU/CURC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shaohui Xie, Xiaotian He, Shan He 0001, Zexuan Zhu 0001 |
Bioinform. | 3 |
| 2022 | Data-Driven Boolean Network Inference Using a Genetic Algorithm With Marker-Based EncodingabstractThe inference of Boolean networks is crucial for analyzing the topology and dynamics of gene regulatory networks. Many data-driven approaches using evolutionary algorithms have been proposed based on time-series data. However, the ability to infer both network topology and dynamics is restricted by their inflexible encoding schemes. To address this problem, we propose a novel Boolean network inference algorithm for inferring both network topology and dynamics simultaneously. The main idea is that, we use a marker-based genetic algorithm to encode both regulatory nodes and logical operators in a chromosome. By using the markers and introducing more logical operators, the proposed algorithm can infer more diverse candidate Boolean functions. The proposed algorithm is applied to five networks, including two artificial Boolean networks and three real-world gene regulatory networks. Compared with other algorithms, the experimental results demonstrate that our proposed algorithm infers more accurate topology and dynamics. Xiang Liu 0019, Ning Shi, Yan Wang 0049, Shan He 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2022 | GloDyNE: Global Topology Preserving Dynamic Network EmbeddingabstractLearning low-dimensional topological representation of a network in dynamic environments is attracting much attention due to the time-evolving nature of many real-world networks. The main and common objective of Dynamic Network Embedding (DNE) is to efficiently update node embeddings while preserving network topology at each time step. The idea of most existing DNE methods is to capture the topological changes at or around the most affected nodes (instead of all nodes) and accordingly update node embeddings. Unfortunately, this kind of approximation, although can improve efficiency, cannot effectively preserve the global topology of a dynamic network at each time step, due to not considering the inactive sub-networks that receive accumulated topological changes propagated via the high-order proximity. To tackle this challenge, we propose a novel node selecting strategy to diversely select the representative nodes over a network, which is coordinated with a new incremental learning paradigm of Skip-Gram based embedding approach. The extensive experiments show GloDyNE, with a small fraction of nodes being selected, can already achieve the superior or comparable performance w.r.t. the state-of-the-art DNE methods in three typical downstream tasks. Particularly, GloDyNE significantly outperforms other methods in the graph reconstruction task, which demonstrates its ability of global topology preservation. Chengbin Hou, Shan He 0001, Ke Tang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2021 | A Network Embedding Based Approach to Drug-Target Interaction Prediction Using Additional Implicit Networks
Chengbin Hou, David W. McDonald, Shan He 0001 |
ICANN (1) | 4 |
| 2021 | PoLoBag: Polynomial Lasso Bagging for signed gene regulatory network inference from expression dataabstractMOTIVATION: Inferring gene regulatory networks (GRNs) from expression data is a significant systems biology problem. A useful inference algorithm should not only unveil the global structure of the regulatory mechanisms but also the details of regulatory interactions such as edge direction (from regulator to target) and sign (activation/inhibition). Many popular GRN inference algorithms cannot infer edge signs, and those that can infer signed GRNs cannot simultaneously infer edge directions or network cycles. RESULTS: To address these limitations of existing algorithms, we propose Polynomial Lasso Bagging (PoLoBag) for signed GRN inference with both edge directions and network cycles. PoLoBag is an ensemble regression algorithm in a bagging framework where Lasso weights estimated on bootstrap samples are averaged. These bootstrap samples incorporate polynomial features to capture higher-order interactions. Results demonstrate that PoLoBag is consistently more accurate for signed inference than state-of-the-art algorithms on simulated and real-world expression datasets. AVAILABILITY AND IMPLEMENTATION: Algorithm and data are freely available at https://github.com/gourabghoshroy/PoLoBag. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gourab Ghosh Roy, Nicholas Geard, Karin Verspoor, Shan He 0001 |
Bioinform. | 4 |
| 2021 | DTI-HeNE: a novel method for drug-target interaction prediction based on heterogeneous network embeddingabstractBACKGROUND: Prediction of the drug-target interaction (DTI) is a critical step in the drug repurposing process, which can effectively reduce the following workload for experimental verification of potential drugs' properties. In recent studies, many machine-learning-based methods have been proposed to discover unknown interactions between drugs and protein targets. A recent trend is to use graph-based machine learning, e.g., graph embedding to extract features from drug-target networks and then predict new drug-target interactions. However, most of the graph embedding methods are not specifically designed for DTI predictions; thus, it is difficult for these methods to fully utilize the heterogeneous information of drugs and targets (e.g., the respective vertex features of drugs and targets and path-based interactive features between drugs and targets). RESULTS: We propose a DTI prediction method DTI-HeNE (DTI based on Heterogeneous Network Embedding), which is specifically designed to cope with the bipartite DTI relations for generating high-quality embeddings of drug-target pairs. This method splits a heterogeneous DTI network into a bipartite DTI network, multiple drug homogeneous networks and target homogeneous networks, and extracts features from these sub-networks separately to better utilize the characteristics of bipartite DTI relations as well as the auxiliary similarity information related to drugs and targets. The features extracted from each sub-network are integrated using pathway information between these sub-networks to acquire new features, i.e., embedding vectors of drug-target pairs. Finally, these features are fed into a random forest (RF) model to predict novel DTIs. CONCLUSIONS: Our experimental results show that, the proposed DTI network embedding method can learn higher-quality features of heterogeneous drug-target interaction networks for novel DTIs discovery. Yang Yue 0008, Shan He 0001 |
BMC Bioinform. | 2 |
| 2021 | GAPORE: Boolean network inference using a genetic algorithm with novel polynomial representation and encoding scheme
Xiang Liu 0019, Yan Wang 0049, Ning Shi, Shan He 0001 |
Knowl. Based Syst. | 5 |
| 2021 | Active Module Identification From Multilayer Weighted Gene Co-Expression Networks: A Continuous Optimization ApproachabstractSearching for active modules, i.e., regions showing striking changes in molecular activity in biological networks is important to reveal regulatory and signaling mechanisms of biological systems. Most existing active modules identification methods are based on protein-protein interaction networks or metabolic networks, which require comprehensive and accurate prior knowledge. On the other hand, weighted gene co-expression networks (WGCNs) are purely constructed from gene expression profiles. However, existing WGCN analysis methods are designed for identifying functional modules but not capable of identifying active modules. There is an urgent need to develop an active module identification algorithm for WGCNs to discover regulatory and signaling mechanism associating with a given cellular response. To address this urgent need, we propose a novel algorithm called active modules on the multi-layer weighted (co-expression gene) network, based on a continuous optimization approach (AMOUNTAIN). The algorithm is capable of identifying active modules not only from single-layer WGCNs but also from multilayer WGCNs such as cross-species and dynamic WGCNs. We first validate AMOUNTAIN on a synthetic benchmark dataset. We then apply AMOUNTAIN to WGCNs constructed from Th17 differentiation gene expression datasets of human and mouse, which include a single layer, a cross-species two-layer and a multilayer dynamic WGCNs. The identified active modules from WGCNs are enriched by known protein-protein interactions, and more importantly, they reveal some interesting and important regulatory and signaling mechanisms of Th17 cell differentiation. Dong Li 0002, Zhisong Pan 0003, Guyu Hu, Graham Anderson, Shan He 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2020 | HEAT: Hyperbolic Embedding of Attributed Networks
David W. McDonald, Shan He 0001 |
IDEAL (1) | 2 |
| 2020 | ATEN: And/Or tree ensemble for inferring accurate Boolean network topology and dynamicsabstractMOTIVATION: Inferring gene regulatory networks from gene expression time series data is important for gaining insights into the complex processes of cell life. A popular approach is to infer Boolean networks. However, it is still a pressing open problem to infer accurate Boolean networks from experimental data that are typically short and noisy. RESULTS: To address the problem, we propose a Boolean network inference algorithm which is able to infer accurate Boolean network topology and dynamics from short and noisy time series data. The main idea is that, for each target gene, we use an And/Or tree ensemble algorithm to select prime implicants of which each is a conjunction of a set of input genes. The selected prime implicants are important features for predicting the states of the target gene. Using these important features we then infer the Boolean function of the target gene. Finally, the Boolean functions of all target genes are combined as a Boolean network. Using the data generated from artificial and real-world gene regulatory networks, we show that our algorithm can infer more accurate Boolean network topology and dynamics from short and noisy time series data than other algorithms. Our algorithm enables us to gain better insights into complex regulatory mechanisms of cell life. AVAILABILITY AND IMPLEMENTATION: Package ATEN is freely available at https://github.com/ningshi/ATEN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ning Shi, Zexuan Zhu 0001, Ke Tang 0001, David Parker 0001, Shan He 0001 |
Bioinform. | 5 |
| 2020 | RoSANE: Robust and scalable attributed network embedding for sparse networks
Chengbin Hou, Shan He 0001, Ke Tang 0001 |
Neurocomputing | 2 |
| 2020 | MUMI: Multitask Module Identification for Biological NetworksabstractIdentifying modules from biological networks is important since modules reveal essential mechanisms and dynamic processes in biological systems. Existing algorithms focus on identifying either active modules or topological modules (communities), which represent dynamic and topological units in the network, respectively. However, high-level biological phenomena, e.g., functions are emergent properties from the interplay between network topology and dynamics. Therefore, to fully explain the mechanisms underlying the high-level biological phenomena, it is important to identify the overlaps between communities and active modules, which indicate the topological units with significant changes of dynamics. However, despite the importance, there are no existing methods to do so. In this article, we propose the multitask module identification (MUMI) algorithm to detect the overlaps between active modules and communities simultaneously. The experimental results show that our method provides new insights into biological mechanisms by combining information from active modules and communities. By formulating the problem as a multitasking learning problem which searches for these two types of modules simultaneously, the algorithm can exploit their latent complementarities to obtain better search performance in terms of accuracy and convergence. Our MATLAB implementation of MUMI is available at https://github.com/WeiqiChen/Mumi-multitask-module-identification. Zexuan Zhu 0001, Shan He 0001 |
IEEE Trans. Evol. Comput. | 3 |
| 2018 | RepLong: de novo repeat identification using long read sequencing dataabstractMotivation: The identification of repetitive elements is important in genome assembly and phylogenetic analyses. The existing de novo repeat identification methods exploiting the use of short reads are impotent in identifying long repeats. Since long reads are more likely to cover repeat regions completely, using long reads is more favorable for recognizing long repeats. Results: In this study, we propose a novel de novo repeat elements identification method namely RepLong based on PacBio long reads. Given that the reads mapped to the repeat regions are highly overlapped with each other, the identification of repeat elements is equivalent to the discovery of consensus overlaps between reads, which can be further cast into a community detection problem in the network of read overlaps. In RepLong, we first construct a network of read overlaps based on pair-wise alignment of the reads, where each vertex indicates a read and an edge indicates a substantial overlap between the corresponding two reads. Secondly, the communities whose intra connectivity is greater than the inter connectivity are extracted based on network modularity optimization. Finally, representative reads in each community are extracted to form the repeat library. Comparison studies on Drosophila melanogaster and human long read sequencing data with genome-based and short-read-based methods demonstrate the efficiency of RepLong in identifying long repeats. RepLong can handle lower coverage data and serve as a complementary solution to the existing methods to promote the repeat identification performance on long-read sequencing data. Availability and implementation: The software of RepLong is freely available at https://github.com/ruiguo-bio/replong. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Rui Guo 0012, Yan-Ran Li 0001, Shan He 0001, Le Ou-Yang, Zexuan Zhu 0001 |
Bioinform. | 3 |
| 2017 | Nadir point estimation for many-objective optimization problems based on emphasized critical regions
Handing Wang, Shan He 0001, Xin Yao 0001 |
Soft Comput. | 2 |
| 2016 | Metabolomics biomarker discovery using multimodal memetic algorithm and multivariate mutual information based feature selectionabstractMetabolomics data has the nature of small sample number, high dimensional, and noisy, which poses great challenges on its analysis. In this paper we propose a novel filter feature selection algorithm, namely MMAFS, for the metabolomics biomarker discovery. The MMAFS utilizes a metaheuristics chain based multimodal memetic algorithm to effectively select both local and global optimal feature subgroups that potentially contain biological meanings. A nearest-neighbor graphic based multivariate mutual information estimation is used to calculate fitness values under the max-dependency criterion. Finally, we introduce a semi-wrapper classification to improve the prediction accuracy. The MMAFS is applied on three real-world metabolomics spectrum data sets. Experimental results on 10 runs of 10-fold external cross validation show that the proposed algorithm outperforms other representative feature selection methods. Particularly, some biomarkers found by MMAFS have been proved by previous researches. Zhen Ji, Zexuan Zhu 0001, Shan He 0001 |
CEC | 4 |
| 2016 | A multi-objective memetic algorithm based on locality-sensitive hashing for one-to-many-to-one dynamic pickup-and-delivery problem
Zexuan Zhu 0001, Shan He 0001, Zhen Ji |
Inf. Sci. | 3 |
| 2016 | Cooperative Co-Evolutionary Module Identification With Application to Cancer Disease Module DiscoveryabstractModule identification or community detection in complex networks has become increasingly important in many scientific fields because it provides insight into the relationship and interaction between network function and topology. In recent years, module identification algorithms based on stochastic optimization algorithms such as evolutionary algorithms have been demonstrated to be superior to other algorithms on small- to medium-scale networks. However, the scalability and resolution limit (RL) problems of these module identification algorithms have not been fully addressed, which impeded their application to real-world networks. This paper proposes a novel module identification algorithm called cooperative co-evolutionary module identification to address these two problems. The proposed algorithm employs a cooperative co-evolutionary framework to handle large-scale networks. We also incorporate a recursive partitioning scheme into the algorithm to effectively address the RL problem. The performance of our algorithm is evaluated on 12 benchmark complex networks. As a medical application, we apply our algorithm to identify disease modules that differentiate low- and high-grade glioma tumors to gain insights into the molecular mechanisms that underpin the progression of glioma. Experimental results show that the proposed algorithm has a very competitive performance compared with other state-of-the-art module identification algorithms. Shan He 0001, Guanbo Jia, Zexuan Zhu 0001, Dan A. Tennant, Ke Tang 0001, Jing Liu 0006, Mirco Musolesi, John K. Heath, Xin Yao 0001 |
IEEE Trans. Evol. Comput. | 1 |
| 2015 | High-throughput DNA sequence data compressionabstractThe exponential growth of high-throughput DNA sequence data has posed great challenges to genomic data storage, retrieval and transmission. Compression is a critical tool to address these challenges, where many methods have been developed to reduce the storage size of the genomes and sequencing data (reads, quality scores and metadata). However, genomic data are being generated faster than they could be meaningfully analyzed, leaving a large scope for developing novel compression algorithms that could directly facilitate data analysis beyond data transfer and storage. In this article, we categorize and provide a comprehensive review of the existing compression methods specialized for genomic data and present experimental results on compression ratio, memory usage, time for compression and decompression. We further present the remaining challenges and potential directions for future research. Zexuan Zhu 0001, Yongpeng Zhang, Zhen Ji, Shan He 0001, Xiao Yang 0019 |
Briefings Bioinform. | 4 |
| 2015 | MUSCLE: automated multi-objective evolutionary optimization of targeted LC-MS/MS analysisabstractAbstract Summary: Developing liquid chromatography tandem mass spectrometry (LC-MS/MS) analyses of (bio)chemicals is both time consuming and challenging, largely because of the large number of LC and MS instrument parameters that need to be optimized. This bottleneck significantly impedes our ability to establish new (bio)analytical methods in fields such as pharmacology, metabolomics and pesticide research. We report the development of a multi-platform, user-friendly software tool MUSCLE (multi-platform unbiased optimization of spectrometry via closed-loop experimentation) for the robust and fully automated multi-objective optimization of targeted LC-MS/MS analysis. MUSCLE shortened the analysis times and increased the analytical sensitivities of targeted metabolite analysis, which was demonstrated on two different manufacturer’s LC-MS/MS instruments. Availability and implementation: Available at http://www.muscleproject.org. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. James Bradbury 0001, Grégory Genta-Jouve, James William Allwood, Warwick B. Dunn, Royston Goodacre, Joshua D. Knowles, Shan He 0001, Mark R. Viant |
Bioinform. | 7 |
| 2015 | Light-weight reference-based compression of FASTQ dataabstractBACKGROUND: The exponential growth of next generation sequencing (NGS) data has posed big challenges to data storage, management and archive. Data compression is one of the effective solutions, where reference-based compression strategies can typically achieve superior compression ratios compared to the ones not relying on any reference. RESULTS: This paper presents a lossless light-weight reference-based compression algorithm namely LW-FQZip to compress FASTQ data. The three components of any given input, i.e., metadata, short reads and quality score strings, are first parsed into three data streams in which the redundancy information are identified and eliminated independently. Particularly, well-designed incremental and run-length-limited encoding schemes are utilized to compress the metadata and quality score streams, respectively. To handle the short reads, LW-FQZip uses a novel light-weight mapping model to fast map them against external reference sequence(s) and produce concise alignment results for storage. The three processed data streams are then packed together with some general purpose compression algorithms like LZMA. LW-FQZip was evaluated on eight real-world NGS data sets and achieved compression ratios in the range of 0.111-0.201. This is comparable or superior to other state-of-the-art lossless NGS data compression algorithms. CONCLUSIONS: LW-FQZip is a program that enables efficient lossless FASTQ data compression. It contributes to the state of art applications for NGS data storage and transmission. LW-FQZip is freely available online at: http://csse.szu.edu.cn/staff/zhuzx/LWFQZip. Yongpeng Zhang, Xiao Yang 0019, Shan He 0001, Zexuan Zhu 0001 |
BMC Bioinform. | 5 |
| 2015 | A Memetic Optimization Strategy Based on Dimension Reduction in Decision SpaceabstractThere can be a complicated mapping relation between decision variables and objective functions in multi-objective optimization problems (MOPs). It is uncommon that decision variables influence objective functions equally. Decision variables act differently in different objective functions. Hence, often, the mapping relation is unbalanced, which causes some redundancy during the search in a decision space. In response to this scenario, we propose a novel memetic (multi-objective) optimization strategy based on dimension reduction in decision space (DRMOS). DRMOS firstly analyzes the mapping relation between decision variables and objective functions. Then, it reduces the dimension of the search space by dividing the decision space into several subspaces according to the obtained relation. Finally, it improves the population by the memetic local search strategies in these decision subspaces separately. Further, DRMOS has good portability to other multi-objective evolutionary algorithms (MOEAs); that is, it is easily compatible with existing MOEAs. In order to evaluate its performance, we embed DRMOS in several state of the art MOEAs to facilitate our experiments. The results show that DRMOS has the advantage in terms of convergence speed, diversity maintenance, and portability when solving MOPs with an unbalanced mapping relation between decision variables and objective functions. Handing Wang, Licheng Jiao, Ronghua Shang, Shan He 0001, Fang Liu 0001 |
Evol. Comput. | 4 |
| 2015 | Robust twin boosting for feature selection from high-dimensional omics data with label noise
Shan He 0001, Huanhuan Chen 0001, Zexuan Zhu 0001, Douglas G. Ward, Helen J. Cooper, Mark R. Viant, John K. Heath, Xin Yao 0001 |
Inf. Sci. | 1 |
| 2015 | Three-dimensional Gabor feature extraction for hyperspectral imagery classification using a memetic framework
Zexuan Zhu 0001, Sen Jia 0001, Shan He 0001, Zhen Ji, LinLin Shen |
Inf. Sci. | 3 |
| 2014 | Protein folding estimation using Paired-Bacteria OptimizerabstractProtein folding estimation attracts a large attention in the area of computational biology, due to its benefits on medical research and the challenge of NP-hard objective functions. In order to simulate the protein folding procedure and estimate the structure of the protein after folding, this paper adopts a Paired-Bacteria Optimizer (PBO), which is a biologically-inspired optimization algorithm. Compared with most Evolutionary Algorithms (EAs), the computational complexity of PBO is much less. Therefore, it is suitable to be applied to solve NP-hard problem. The experimental studies is performed on several benchmark lattice protein combination. The experimental results demonstrated that PBO is able to estimate the folded protein structure with a superior convergence. Mengshi Li, Tianyao Ji, Peter Wu, Shan He 0001, Q. Henry Wu |
IEEE Congress on Evolutionary Computation | 4 |
| 2014 | HAMMER: automated operation of mass frontier to construct in silico mass spectral fragmentation librariesabstractSUMMARY: Experimental MS(n) mass spectral libraries currently do not adequately cover chemical space. This limits the robust annotation of metabolites in metabolomics studies of complex biological samples. In silico fragmentation libraries would improve the identification of compounds from experimental multistage fragmentation data when experimental reference data are unavailable. Here, we present a freely available software package to automatically control Mass Frontier software to construct in silico mass spectral libraries and to perform spectral matching. Based on two case studies, we have demonstrated that high-throughput automation of Mass Frontier allows researchers to generate in silico mass spectral libraries in an automated and high-throughput fashion with little or no human intervention required. AVAILABILITY AND IMPLEMENTATION: Documentation, examples, results and source code are available at http://www.biosciences-labs.bham.ac.uk/viant/hammer/. Ralf J. M. Weber, James William Allwood, Robert Mistrik, Zexuan Zhu 0001, Zhen Ji, Siping Chen, Warwick B. Dunn, Shan He 0001, Mark R. Viant |
Bioinform. | 9 |
| 2014 | Compression of next-generation sequencing quality scores using memetic algorithmabstractBACKGROUND: The exponential growth of next-generation sequencing (NGS) derived DNA data poses great challenges to data storage and transmission. Although many compression algorithms have been proposed for DNA reads in NGS data, few methods are designed specifically to handle the quality scores. RESULTS: In this paper we present a memetic algorithm (MA) based NGS quality score data compressor, namely MMQSC. The algorithm extracts raw quality score sequences from FASTQ formatted files, and designs compression codebook using MA based multimodal optimization. The input data is then compressed in a substitutional manner. Experimental results on five representative NGS data sets show that MMQSC obtains higher compression ratio than the other state-of-the-art methods. Particularly, MMQSC is a lossless reference-free compression algorithm, yet obtains an average compression ratio of 22.82% on the experimental data sets. CONCLUSIONS: The proposed MMQSC compresses NGS quality score data effectively. It can be utilized to improve the overall compression ratio on FASTQ formatted files. Zhen Ji, Zexuan Zhu 0001, Shan He 0001 |
BMC Bioinform. | 4 |
| 2014 | An Evolutionary Multiobjective Approach to Sparse ReconstructionabstractThis paper addresses the problem of finding sparse solutions to linear systems. Although this problem involves two competing cost function terms (measurement error and a sparsity-inducing term), previous approaches combine these into a single cost term and solve the problem using conventional numerical optimization methods. In contrast, the main contribution of this paper is to use a multiobjective approach. The paper begins by investigating the sparse reconstruction problem, and presents data to show that knee regions do exist on the Pareto front (PF) for this problem and that optimal solutions can be found in these knee regions. Another contribution of the paper, a new soft-thresholding evolutionary multiobjective algorithm (StEMO), is then presented, which uses a soft-thresholding technique to incorporate two additional heuristics: one with greater chance to increase speed of convergence toward the PF, and another with higher probability to improve the spread of solutions along the PF, enabling an optimal solution to be found in the knee region. Experiments are presented, which show that StEMO significantly outperforms five other well known techniques that are commonly used for sparse reconstruction. Practical applications are also demonstrated to fundamental problems of recovering signals and images from noisy data. Lin Li 0016, Xin Yao 0001, Rustam Stolkin, Maoguo Gong, Shan He 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2013 | Minimal-redundancy-maximal-relevance feature selection using different relevance measures for omics data classificationabstractOmics refers to a field of study in biology such as genomics, proteomics, and metabolomics. Investigating fundamental biological problems based on omics data would increase our understanding of bio-systems as a whole. However, omics data is characterized with high-dimensionality and unbalance between features and samples, which poses big challenges for classical statistical analysis and machine learning methods. This paper studies a minimal-redundancy-maximal-relevance (MRMR) feature selection for omics data classification using three different relevance evaluation measures including mutual information (MI), correlation coefficient (CC), and maximal information coefficient (MIC). A linear forward search method is used to search the optimal feature subset. The experimental results on five real-world omics datasets indicate that MRMR feature selection with CC is more robust to obtain better (or competitive) classification accuracy than the other two measures. Junshan Yang, Zexuan Zhu 0001, Shan He 0001, Zhen Ji |
CIBCB | 3 |
| 2012 | A memory binary particle swarm optimizationabstractThis paper proposes a memory binary particle swarm optimization algorithm (MBPSO) based on a new updating strategy. Unlike the traditional binary PSO, which updates the binary bits of a particle ignoring their previous status, MBPSO memorizes the bit status and updates them according to a new defined velocity. As such, precious historical information could be retained to guide the search. The velocity vector of MBPSO is designed as a probability for deciding whether the particle bits change or not. The proposed algorithm is tested on four discrete benchmark functions. The experimental results reported over 100 runs show that MBPSO is capable of obtaining encouraging performance in discrete optimization problems. Zhen Ji, Tao Tian, Shan He 0001, Zexuan Zhu 0001 |
IEEE Congress on Evolutionary Computation | 3 |
| 2012 | A crown jewel defense strategy based particle swarm optimizationabstractParticle swarm optimization (PSO) is a metaheuristic algorithm that is easy to implement and performs well on various optimization problems. However, PSO is sensitive to initialization due to its rapid convergence which leads to the lack of population diversity and premature convergence. To solve this problem, a jumping-out strategy named crown jewel defense (CJD) is introduced in this paper. CJD is used to relocate the global best position and reinitializes all particles' personal best position when the swarm is trapped in local optima. Taking the advantage of CJD strategy, the swarm can jump out of the local optimal region without being dragged back and the performance of PSO becomes more robust to the initialization. Experimental results on benchmark functions show that the CJD-based PSO are comparable to or better than the other representative state-of-the-art PSO. Zhen Ji, Shan He 0001, Zexuan Zhu 0001 |
IEEE Congress on Evolutionary Computation | 3 |
| 2012 | Survival analysis of gene expression data using PSO based radial basis function networksabstractGene expression data combined with clinical data has emerged as an important source for survival analysis. However, gene expression data is characterized with thousands of features/genes but only tens or hundreds of observations. The high-dimensionality and unbalance between features and samples pose big challenges for the classical survival analysis methods. This paper proposes a particle swarm optimization based radial basis function networks (PSO-RBFN) for the survival analysis on gene expression data. Particularly, PSO-RBFN applies a principle component analysis for dimensionality reduction and optimizes the RBF network using PSO. The experimental results on three gene expression datasets indicate that PSO-RBFN is able to improve the predict accuracy compared to the other classical survival analysis methods. Wenmin Liu, Zhen Ji, Shan He 0001, Zexuan Zhu 0001 |
IEEE Congress on Evolutionary Computation | 3 |
| 2012 | Memetic clustering based on particle swarm optimizer and K-meansabstractThis paper proposes an efficient memetic clustering algorithm (MCA) for clustering based on particle swarm optimizer (PSO) and K-means. Particularly, PSO is used as a global search to allow fast exploration of the candidate cluster centers. PSO has strong ability to find high quality solutions within tractable time, but it suffers from slow-down convergence as the swarm approaching optima. K-means, achieving fast convergence to optimum solutions, is utilized as local search to fine-tune the solutions of PSO in the framework of memetic algorithm. The performance of MCA is evaluated on four synthetic datasets and three high-dimensional gene expression datasets. Comparison study to K-means, PSO, and PSO-KM (jointed PSO and K-means) indicates that MCA is capable of identifying cluster centers more precisely and robustly than the other counterpart algorithms by taking advantage of both PSO and K-means. Zexuan Zhu 0001, Wenmin Liu, Shan He 0001, Zhen Ji |
IEEE Congress on Evolutionary Computation | 3 |
| 2012 | Community Detection Using Cooperative Co-evolutionary Differential Evolution
Thomas White, Guanbo Jia, Mirco Musolesi, Nil Turan, Ke Tang 0001, Shan He 0001, John K. Heath, Xin Yao 0001 |
PPSN (2) | 7 |
| 2012 | MaConDa: a publicly accessible mass spectrometry contaminants databaseabstractUNLABELLED: Mass spectrometry is widely used in bioanalysis, including the fields of metabolomics and proteomics, to simultaneously measure large numbers of molecules in complex biological samples. Contaminants routinely occur within these samples, for example, originating from the solvents or plasticware. Identification of these contaminants is crucial to enable their removal before data analysis, in particular to maintain the validity of conclusions drawn from uni- and multivariate statistical analyses. Although efforts have been made to report contaminants within mass spectra, this information is fragmented and its accessibility is relatively limited. In response to the needs of the bioanalytical community, here we report the creation of an extensive manually well-annotated database of currently known small molecule contaminants. AVAILABILITY: The Mass spectrometry Contaminants Database (MaConDa) is freely available and accessible through all major browsers or by using the MaConDa web service http://www.maconda.bham.ac.uk. Ralf J. M. Weber, Eva Li, Jonathan Bruty, Shan He 0001, Mark R. Viant |
Bioinform. | 4 |
| 2009 | Profiling of Mass Spectrometry Data for Ovarian Cancer Detection Using Negative Correlation Learning
Shan He 0001, Huanhuan Chen 0001, Xiaoli Li 0002, Xin Yao 0001 |
ICANN (2) | 1 |
| 2009 | Group Search Optimizer: An Optimization Algorithm Inspired by Animal Searching BehaviorabstractNature-inspired optimization algorithms, notably evolutionary algorithms (EAs), have been widely used to solve various scientific and engineering problems because of to their simplicity and flexibility. Here we report a novel optimization algorithm, group search optimizer (GSO), which is inspired by animal behavior, especially animal searching behavior. The framework is mainly based on the producer-scrounger model, which assumes that group members search either for “finding” (producer) or for “joining” (scrounger) opportunities. Based on this framework, concepts from animal searching behavior, e.g., animal scanning mechanisms, are employed metaphorically to design optimum searching strategies for solving continuous optimization problems. When tested against benchmark functions, in low and high dimensions, the GSO algorithm has competitive performance to other EAs in terms of accuracy and convergence speed, especially on high-dimensional multimodal problems. The GSO algorithm is also applied to train artificial neural networks. The promising results on three real-world benchmark problems show the applicability of GSO for problem solving. Shan He 0001, Q. Henry Wu, Jon R. Saunders |
IEEE Trans. Evol. Comput. | 1 |
| 2008 | Application of a group search optimization based Artificial Neural Network to machine condition monitoringabstractArtificial Neural Networks (ANNs) have been applied to machine condition monitoring. This paper first addresses a ANN trained by Group Search Optimizer (GSO), which is a novel population based optimization algorithm inspired by animal social foraging behaviour. The global search performance of GSO has been proven to be competitive to other evolutionary algorithms, such as Genetic Algorithms (GAs) and Particle Swarm Optimizer (PSO). Herein, the parameters of a 3-layer feed-forward ANN, including connection weights and bias are tuned by the GSO algorithm. Secondly the GSO based ANN is applied to model and analysis ultrasound data recorded from grinding machines to distinguish different conditions. The real experimental results show that the proposed method is capable to indicate the malfunction of machine condition from the ultrasound data. Shan He 0001, Xiaoli Li 0002 |
ETFA | 1 |
| 2007 | Profiling of High-Throughput Mass Spectrometry Data for Ovarian Cancer Detection
Shan He 0001, Xiaoli Li 0002 |
IDEAL | 1 |
| 2006 | A Group Search Optimizer for Neural Network Training
Shan He 0001, Q. Henry Wu, Jon R. Saunders |
ICCSA (3) | 1 |
| 2005 | A particle swarm optimiser with passive congregation approach to thermal modelling for power transformersabstractThis paper employs an intelligent learning technique based on a particle swarm optimiser with passive congregation (PSOPC) algorithm to identify the thermal parameters of a simplified thermoelectric analogous thermal model (STEATM) for transformers, based upon only a few onsite measurements instead of experimental methods. The model outputs deliver good agreements with the onsite data based upon a single set of parameters obtained from the PSOPC learning with a fast convergence rate. The simulation results are compared with that obtained using an artificial neural network (ANN) approach. Wenhu Tang, Shan He 0001, Emmanuel Prempain, Q. Henry Wu, J. Fitch |
Congress on Evolutionary Computation | 2 |