EDBT 2026 Demo / reviewers in the wild / expert
Zexuan Zhu 0001
dblp:17/4590
· DBLP profile ↗
131ranked-venue papers
11as first author
68since 2021 · last 2026
0000-0001-8479-6904ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 83 · 5 first-author · 45 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 3 first-author · 16 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PepCCD: A Contrastive Conditioned Diffusion Framework for Target-Specific Peptide GenerationabstractPeptide-based drug design targeting “undruggable” proteins remains one of the most critical challenges in modern drug discovery. Conventional peptide-discovery pipelines rely on low-throughput experimental screening, which is both time-consuming and prohibitively expensive. Moreover, existing computational approaches for designing peptides against target proteins typically depend on the availability of high-quality structural information. Although recent structure-prediction tools such as AlphaFold3 have achieved breakthroughs in protein modeling, their accuracy for functional interfaces remains limited. The acquisition of high-resolution structures is often expensive, time-intensive, and particularly challenging for targets with dynamic conformations, further restricting the efficient development of peptide therapeutics. Additionally, current sequence-based generative methods follow a paradigm that relies on known templates, which limits the exploration of sequence space and results in generated peptides lacking diversity and novelty. To address these limitations, we propose a contrastive conditioned diffusion framework for target-specific peptide generation, referred to as PepCCD. It employs a contrastive learning strategy between proteins and peptides to extract sequence-based conditioning representations of target proteins, which serve as precise conditions to guide a pre-trained diffusion model to generate peptide sequences with the desired target specificity. Extensive experiments on multiple benchmark target proteins demonstrate that the peptides designed by PepCCD exhibit strong binding affinity and outperform state-of-the-art methods in terms of diversity and generation efficiency. Jun Zhang 0078, Yangyang Zhou, Zexuan Zhu 0001 |
AAAI | 4 |
| 2026 | Enhanced Reasoning for Biomedical Document-Level Relation Extraction via a Novel Cascade Language Model FrameworkabstractHaohua Song, Wenhao Gu, Zhijing Li, Yunwenyu, Tiantian Zhu, Xiao Yang, Zexuan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Haohua Song, Wenhao Gu, Zhijing Li 0007, Yunwen Yu, Xiao Yang 0019, Zexuan Zhu 0001 |
ACL (1) | 7 |
| 2026 | Combinatorial Medication Recommendation via Temporal-Aware Multi-objective Optimization
Ruobing Wei, Yichun Guan, Qingxia Shang, Zexuan Zhu 0001, Wei Zhou 0001 |
ICIC (16) | 4 |
| 2026 | Polus: a context-aware enhancement framework for DNA storage via transformer-based soft-decision decodingabstractMOTIVATION: DNA storage offers exceptional information density and archival longevity but is constrained by complex biochemical noise inherent to synthesis, storage, and sequencing. Conventional hard-decision error-correction schemes often rely on excessive redundancy to mitigate these imperfections, which significantly compromises storage efficiency and density. RESULTS: We present Polus, a Transformer-based enhancement framework that improves digital reliability through soft-decision decoding (SDD) without requiring encoder modification. At its core is SeqFormer, a Transformer-based channel model that synergizes sequence context with quality signals to generate calibrated per-base confidence scores, effectively transforming uncertain biochemical noise into informative "soft" erasures. In in silico benchmarks, Polus significantly upgrades mainstream DNA storage codecs. It reduces the sequencing coverage required for DNA Fountain by 38.9%-increasing effective physical density by approximately 80%-and eliminates persistent indel-induced errors in the Yin-Yang codec. Furthermore, it enables a targeted resequencing strategy that achieves full recovery with 99.9% less overhead than uniform deepening. Moreover, a nine-metric evaluation suite was employed to provide multi-dimensional quantitative comparisons of DNA storage codecs across reliability, density, and cost. Collectively, Polus provides a reproducible framework for context-aware decoding and system design guidance for DNA storage. AVAILABILITY AND IMPLEMENTATION: All source code of the Polus, including the SeqFormer implementation, codec algorithms, test data used, and the simulation pipeline is available on GitHub (https://github.com/dinglulu/Polus) and Zenodo (https://zenodo.org/communities/bioinfoszu/). A web hosted instance of Polus is available at https://polus.bioailab.net/polls/home. The SeqFormer model is also released as a standalone repository at https://github.com/dinglulu/SeqFormer and https://zenodo.org/communities/bioinfoszu/. Lulu Ding, Kun Wang 0056, Shaohui Xie, Ling Liu 0003, Zexuan Zhu 0001 |
Bioinform. | 9 |
| 2026 | NanoSimFormer: an end-to-end transformer-based nanopore signal simulator with basecaller guidanceabstractMOTIVATION: High-fidelity simulation of nanopore sequencing signals is critical for rigorous benchmarking and validation of the nanopore signal processing pipeline. However, existing signal simulators often fail to capture the non-linear dynamics of nanopore current signals, relying on static pore models or lacking optimization objectives tied to basecalling, resulting in synthetic signals with low basecalling accuracy and fidelity. RESULTS: We introduce NanoSimFormer, an end-to-end Transformer-based signal simulator that integrates basecaller guidance during training to generate high-fidelity nanopore signals. NanoSimFormer achieves a median basecalling accuracy exceeding 99% and Q-scores above 22.8 for Oxford Nanopore Technologies' latest DNA R10.4.1 and direct RNA sequencing, closely mirroring real experimental baselines. It faithfully recapitulates experimental variant calling performance across the five human samples, achieving F1-scores of 0.9953-0.9973 and 0.7862-0.8612 for single-nucleotide polymorphisms and small indels detections, respectively. Compared with previous simulators, NanoSimFormer also substantially reduces false positives in homopolymer and short tandem repeat regions. NanoSimFormer-derived reads enable high-quality de novo bacterial assembly with consensus error rates below one mismatch per 100 kbp and maintain high correlations with experimental abundance in metagenomic and transcriptomic datasets. AVAILABILITY AND IMPLEMENTATION: NanoSimFormer is freely available on GitHub at: https://github.com/BioinfoSZU/NanoSimFormer. Shaohui Xie, Lulu Ding, Yew-Soon Ong, Zexuan Zhu 0001 |
Bioinform. | 6 |
| 2026 | Molecular-level protein semantic learning via structure-aware coarse-grained language modelingabstractMOTIVATION: Protein language models (PLMs) have emerged as pivotal tools for protein representation, enabling significant advances in structure-function prediction and computational biology. However, current PLMs predominantly rely on fine-grained amino acid sequences as input, treating individual residues as tokens. While this approach facilitates semantic learning at the residue level, it struggles to capture molecular-level semantics, particularly for large proteins, where sequence truncation and inefficient local pattern extraction hinder holistic understanding. The spatial structure of a protein determines its function. Despite the critical role of protein function analysis, coarse-grained protein language frameworks that bridge sequence and structural semantics remain underdeveloped. RESULTS: To fill this gap, we introduce a novel structure-aware coarse-grained protein language that discretizes proteins into local structural patterns derived from their secondary structures. By constructing a vocabulary of these patterns as "words," we represent proteins as compact, structure-aware "sentences" significantly shorter than raw amino acid sequences. We benchmark the proposed coarse-grained language against three state-of-the-art fine-grained protein languages and a classical language modeling method in natural language processing, using two architectures: a lightweight Doc2Vec model and a Transformer-based BERT model, and evaluating performance across diverse downstream tasks, including function prediction, enzyme classification, and interaction identification. The proposed method achieves stable performance across three tasks, especially for long proteins. These results demonstrate that the proposed coarse-grained protein language preserves critical structural and functional semantics and improves molecular-level analysis, offering a promising direction for decoding higher-order biological insights. AVAILABILITY AND IMPLEMENTATION: The data and source code of the proposed method are available at GitHub (https://github.com/bug-0x3f/coarse-grained-protein-language) and Zenodo (DOI: 10.5281/zenodo.17674298). Jun Zhang 0078, Xueer Weng, Zexuan Zhu 0001 |
Bioinform. | 5 |
| 2026 | Multifactorial evolutionary algorithm enhanced by symmetry transformation and ridge regression
Guoxing Luo, Yan Wang 0155, Jihua Fan, Lei Wang 0018, Yutao Qi, Zexuan Zhu 0001, Xiaoliang Ma 0001 |
Expert Syst. Appl. | 7 |
| 2026 | Tumor cell fraction estimation based on tissue region segmentation and nuclear density
Lulu Qin, Xiao Yang 0019, Zhigang Pei, Susan Fotheringham, Xianhong Xu, Zexuan Zhu 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Fast heuristic search algorithms for submodular cost submodular cover under routing constraints
Xuefeng Chen 0001, Liang Feng 0001, Xin Cao 0001, Zexuan Zhu 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Optimal Linear Crossover for Mitigating Negative Transfer in Evolutionary MultitaskingabstractEvolutionary multitasking algorithms use information exchange among individuals in a population to solve multiple optimization problems simultaneously. Negative transfer is a critical factor that affects the performance of evolutionary multitasking algorithms. In this study, we propose an innovative approach to mitigate negative transfer in evolutionary multitasking algorithms. The proposed approach is grounded in rigorous theoretical analysis, which provides valuable theoretical insights into the design of an optimal linear crossover operator for mitigating negative transfer. By identifying interpretable conditions, we establish a solid theoretical foundation to prevent negative transfer in diverse scenarios. Building upon these findings, we theoretically derive a closed-form expression for the optimal crossover operator and propose practical design methods based on approximations. Furthermore, we integrate the proposed optimal crossover operator into a fundamental evolutionary multitasking algorithm framework. The resultant algorithm is comparable or superior to other state-of-the-art methods. Empirical validation through comprehensive experiments confirms the effectiveness of our theoretical findings. Zhaobo Liu, Jianhua Yuan, Zexuan Zhu 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2025 | DCSF-KD: Dynamic Channel-wise Spatial Feature Knowledge Distillation for Object DetectionabstractKnowledge distillation (KD) has recently gained great success in the field of object detection. By transferring the knowledge of the spatial or channel domain from the teacher model to the student model, it allows for a more compact representation with minimal performance loss. Despite this progress, existing KD methods typically treat knowledge from spatial or channel domains independently, ignoring the exploitation of the mutual relationship between these domains. In this work, we first explore the connection between spatial and channel domains and find there exists a strong correlation between them, i.e. the salient channels tend to contain significant object regions in the spatial domain. Motivated by this observation, we propose DCSF-KD, a novel Dynamic Channel-wise Spatial Feature Knowledge Distillation framework for object detection by fully exploiting both spatial and channel knowledge. Specifically, we introduce channel-wise spatial feature distillation and global channel attention distillation, using information from both domains to improve the accuracy of the student network. Experiments demonstrate that our DCSF-KD outperforms existing detection methods on both homogeneous and heterogeneous teacher-student network pairs. For example, when using the MaskRCNN-Swin detector as the teacher, and based on RetinaNet and FCOS with ResNet-50 on MS COCO, our DCSF-KD can achieve 41.9% and 44.1% mAP, respectively. Tao Dai 0001, Hang Guo 0002, Jinbao Wang 0001, Zexuan Zhu 0001 |
AAAI | 5 |
| 2025 | GCD-Sampling: A General Cross-scale Decoupled Sampling for Point CloudabstractSampling strategy (e.g., fixed farthest point sampling) of point cloud has been an essential step for developing practical solutions in 3D computer vision tasks. Previous fixed sampling is simple, but suffer from suboptimal performance for downstream tasks. To adapt to target networks properly, adaptive sampling methods with trainable parameters have been recently developed to enhance the performance. However, existing adaptive sampling methods still suffer from the over-coupling problem of target network, and thus become model-specific, which limits their practical applications. To address this issue, we propose a novel general cross-scale decoupled sampling method (GCD-sampling) for point cloud, which consists of original feature cache, cross-scale feature fusion and convex combination learning for better feature extraction. To reduce the coupling relationship with the target task network, our method only utilizes the point cloud coordinates as the input and output of itself. Besides, we introduce an arbitrary scale structure to enable parameter sharing across multi-scale sampling in point cloud networks. Extensive experiments on different architectures demonstrate the effectiveness of our method over other existing adaptive sampling methods. Tao Dai 0001, Yanzi Wang, Jianyu Xiong, Yaohua Zha, Shutao Xia, Zexuan Zhu 0001 |
AAAI | 6 |
| 2025 | CFPLM: Improve Protein-RNA Interaction Prediction with a Collaborative Framework Powered by Language ModelsabstractProtein-RNA interactions (PRIs) play pivotal roles in biological processes such as gene regulation, making their prediction essential for therapeutic and mechanistic studies. While traditional wet-lab methods are timeconsuming and challenging, computational approaches offer efficient alternatives. Graph-based methods show promise by capturing both direct interactive domains (protein-RNA interaction) and indirect collaborative domains (functional similarity among proteins/RNAs). However, integrating these domains and learning meaningful node representations remain critical challenges. To address this, we propose CFPLM, a collaborative framework fusing large language models, graph convolutional networks, and cross-attention mechanisms to improve PRI prediction. Experiment results demonstrate that CFPLM achieves robust, state-of-the-art performance across three benchmark datasets. It's anticipated to have applicability to similar other interaction prediction tasks. The data and codes are available at: https://github.com/HuanchaoFeng/CFPLM. Jun Zhang 0078, Huanchao Feng, Hang Wei 0005, Zexuan Zhu 0001 |
BIBM | 4 |
| 2025 | Multiobjective Clustering-Guided Multitasking for Diversified Sequential RecommendationabstractSequential recommendation (SR) aims to predict users’ next-item preferences by analyzing their historical interaction sequences. Most existing SR models primarily focus on user preferences for a single objective of item relevance; they often overlook users’ personalized preferences for diversified item attributes. Although recent studies have attempted to address this limitation, efficiently resolving the inherent accuracy-diversity trade-off remains an open challenge. Drawing inspirations from multiobjective optimization and evolutionary multitasking, this paper proposes a multiobjective clustering-guided multitasking framework (MCGM-SR) that reconciles the two competing objectives in sequential recommendation. To enable knowledge transfer across homologous preference patterns, a user-level multiobjective clustering mechanism is introduced to group users with similar behavioral trajectories into objective-oriented clusters. Within each cluster, a multiobjective multitasking algorithm facilitates synergistic optimization across tasks to improve convergence efficiency. Extensive experiments on two real-world public datasets demonstrate that MCGM-SR achieves significant improvements in both accuracy and diversity metrics, compared to state-of-the-art sequential models and their diversified baselines. Further ablation studies reveal that the proposed clustering method significantly enhances optimization performance by effectively facilitating knowledge transfer. Ruxin Wu, Wei Zhou 0001, Junkai Ji, Sijun Peng, Siqin Peng, Zexuan Zhu 0001 |
CEC | 6 |
| 2025 | DIIN: Diffusion Iterative Implicit Networks for Arbitrary-scale Super-resolutionabstractImplicit neural representation (INR) aims to represent continuous domain signals via implicit neural functions and has achieved great success in arbitrary-scale image super-resolution (SR). However, most existing INR-based SR methods focus on learning implicit features from independent coordinate, while neglecting interactions of neighborhood coordinates, thus resulting in limited contextual awareness. In this paper, we rethink the forward process of implicit neural functions as a signal diffusion process, we propose a novel Diffusion Iterative Implicit Network (DIIN) for arbitrary-scale SR to promote global signal flow with neighborhood interactions. The DIIN framework mainly consists of stacked Diffusion Iteration Layers with dictionary cross-attention block to enrich the iterative update process with supplementary information. Besides, we develop the Position-Aware Embedding Block to strengthen spatial dependencies between consecutive input samples.Extensive experiments on public datasets demonstrate that our method achieves state-of-the-art or competitive performance, highlighting its effectiveness and efficiency for arbitrary-scale SR. Our code is available at https://github.com/Song-1205/DIIN. Tao Dai 0001, Hang Guo 0002, Zexuan Zhu 0001 |
IJCAI | 5 |
| 2025 | Improved Lossless Compression based on Polar CodesabstractPolar codes have been proven to be capable of achieving the optimal rate for the lossless compression problem. However, their finite-length performance is not satisfactory due to the insufficient polarization effect. In this work, we combine source polarization with several entropy coding techniques to improve the compression efficiency while keeping the additional complexity negligible. In our framework, the standard encoding of polar codes can be treated as a pre-transform on the source data, and only a small proportion of the transformed data needs further compression thanks to the source polarization. We show that our framework is compatible with the mainstream entropy coding schemes such as Huffman coding, arithmetic coding, and asymmetric number system (ANS). To optimize performance, an iterative algorithm is proposed for the set partitioning of the transformed data. Simulation results show that the improved scheme is superior to the original polar source coding. Ling Liu 0003, Chao Chen 0013, Lulu Ding, Zexuan Zhu 0001, Baoming Bai |
ITW | 5 |
| 2025 | TG-CDDPM: text-guided antimicrobial peptides generation based on conditional denoising diffusion probabilistic modelabstractAntimicrobial peptides (AMPs) have emerged as a promising substitution to antibiotics thanks to their boarder range of activities, less likelihood of drug resistance, and low toxicity. Traditional biochemical methods for AMP discovery are costly and inefficient. Deep generative models, including the long-short term memory model, variational autoencoder model, and generative adversarial model, have been widely introduced to expedite AMP discovery. However, these models tend to suffer from the lack of diversity in generating AMPs. The denoising diffusion probabilistic model serves as a good candidate for solving this issue. We proposed a three-stage Text-Guided Conditional Denoising Diffusion Probabilistic Model (TG-CDDPM) to generate novel and homologous AMPs. In the first two stages, contrastive learning and inferring models are crafted to create better conditions for guiding AMP generation, respectively. In the last stage, a pre-trained conditional denoising diffusion probabilistic model is leveraged to enrich the peptide knowledge and fine-tuned to learn feature representation in downstream. TG-CDDPM was compared to the state-of-the-art generative models for AMP generation, and it demonstrated competitive or better performance with the assistance of text description as supervised information. The membrane penetration capabilities of the identified candidate AMPs by TG-CDDPM were also validated through molecular weight dynamics experiments. Junhang Cao, Jun Zhang 0078, Qiyuan Yu, Junkai Ji, Jianqiang Li 0001, Shan He 0001, Zexuan Zhu 0001 |
Briefings Bioinform. | 7 |
| 2025 | INAB: identify nucleic acid binding domain via cross-modal protein language models and multiscale computationabstractProtein-nucleic acid interactions play a crucial role in biological processes, including gene regulation and editing. Accurately identifying nucleic acid-binding domains in proteins is essential to unravel these interactions, yet traditional experimental methods like X-ray crystallography remain costly and time-intensive. Computational approaches have thus emerged as indispensable tools to complement wet-lab techniques. Here, we introduce a framework for nucleic acid-binding domain prediction by integrating cross-modal protein language models with a multiscale computational architecture. The proposed method leverages a structurally annotated benchmark dataset, which quantifies binding likelihood through hierarchical, proximity-based labels derived from experimental complexes. Evaluations demonstrate that the approach achieves state-of-the-art performance, providing a new insight into the design of multimodal learning systems in protein-nucleic acid interaction analysis and an open resource to accelerate discoveries in functional genomics and drug design. Jun Zhang 0078, Junjie Chen 0004, Zexuan Zhu 0001 |
Briefings Bioinform. | 4 |
| 2025 | Quality Scores Compression of Genomic Sequencing Data: A Comprehensive Review and Performance EvaluationabstractAdvanced sequencing technologies have profoundly revolutionized biology and produced vast amounts of raw sequencing data during the past decades. The enormous amount of sequencing data proposed significant challenges of data storage and transmission. Compressing a big file into a small file is an encouraging method to tackle these challenges. Howver, it has been found that traditional text data compression algorithms are not well-suited for handing the vast sequencing datasets. Therefore, several algorithms are designed specifically for the efficient compression of sequencing data. Recently, considerable research has been devoted to compressing quality scores stored in the FASTQ format file, resulting in substantial advances in compression performance. Despite these advances, there has been no systematic review and evaluation of these algorithms or software. In this review, we aim to conduct a broad review of the existing quality score compression algorithms. We mainly discuss those algorithms from two categories, i.e., lossless and lossy compression. Additionally, we benchmark the compression performance of 12 tools using 14 real datasets. We anticipate that our review will provide practical guidance for others seeking to design an appropriate algorithm for compressing quality scores. Yuansheng Liu, Zexuan Zhu 0001, Xiangxiang Zeng, Quan Zou 0001, Keqin Li 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Evolutionary Dynamic Multiobjective Optimization With Learning Across ProblemsabstractDynamic multi-objective optimization problems (DMOPs) are prevalent in many practical applications and have garnered significant attention from both industry and academia, leading to the proposal of numerous dynamic multi-objective optimization algorithms. Among these approaches, learning-and prediction-based evolutionary approaches have achieved remarkable success due to their fast learning and strong optimization capabilities when dynamic occurs. However, existing methods typically focus on learning and transferring knowledge within a single problem, often relying on historical knowledge to aid the optimization process. This could limit the data available for developing a more general prediction model for dynamic optimization. The potential for leveraging knowledge across different DMOPs to enhance problem-solving efficiency and effectiveness remains largely unexplored. Building on this insight, this paper explores the solution of DMOPs with learning not only within a single problem but also across different problems. In particular, by performing dynamic feature extraction and task-specific solution classification, we propose to construct a centralized learning model that captures the correlations across DMOPs and the corresponding optimized solutions. This approach allows optimization data from multiple problems to be effectively utilized, enabling the learning of dynamic knowledge to enhance evolutionary optimization when a change occurs. To assess the effectiveness of the proposed algorithm, extensive empirical studies have been conducted on the commonly used DMOP benchmarks with fixed and shifty dynamic settings. The numerical results demonstrate the effectiveness of the proposed algorithm in recognizing more general change patterns of DMOPs, showcasing its superiority compared to the existing state-of-the-art learning and prediction-based approaches. Yuling Xie, Quanwu Zhao, Wei Zhou 0001, Zexuan Zhu 0001 |
IEEE Trans. Evol. Comput. | 4 |
| 2025 | A Survey on Evolutionary Computation-Based Drug DiscoveryabstractDrug discovery is an expensive and risky process. To combat the challenges in drug discovery, an increasing number of researchers and pharmaceutical companies recognize the benefits of utilizing computational techniques. Evolutionary computation (EC) offers promise as most drug discovery problems are essentially complex optimization problems beyond conventional optimization algorithms. EC methods have been widely applied to solve these complex optimization problems especially in lead com-pound generation and molecular virtual evaluation, substantially speeding up the process of drug discovery and development. This article presents a comprehensive survey of EC-based drug discovery methods. Particularly, a new taxonomy of the methods is provided and the advantages and limitations of the methods are reviewed. In addition, the potential future directions of EC-based drug discovery are discussed and the publicly available resources including databases and computational tools are compiled for the convenience of researchers seeking to pursue this field. Qiyuan Yu, Qiuzhen Lin, Junkai Ji, Wei Zhou 0001, Shan He 0001, Zexuan Zhu 0001, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 6 |
| 2025 | MDTL-ACP: Anticancer Peptides Prediction Based on Multi-Domain Transfer LearningabstractAnticancer peptides (ACPs) have emerged as one of the most promising therapeutic agents for cancer treatment. They are bioactive peptides featuring broad-spectrum activity and low drug-resistance. The discovery of ACPs via traditional biochemical methods is laborious and costly. Accordingly, various computational methods have been developed to facilitate the discovery of ACPs. However, the data resources and knowledge of ACPs are still very scarce, and only a few of them are clinically verified, which limits the competence of computational methods. To address this issue, in this article, we propose an ACP prediction model based on multi-domain transfer learning, namely MDTL-ACP, to discriminate novel ACPs from plentiful inactive peptides. In particular, we collect abundant antimicrobial peptides (AMPs) from four well-studied peptide domains and extract their inherent features as the input of MDTL-ACP. The features learned from multiple source domains of AMPs are then transferred into the target prediction task of ACPs via artificial neural network-based shared-extractor and task-specific classifiers in MDTL-ACP. The knowledge captured in the transferred features enhances the prediction of ACPs in the target domain. Experimental results demonstrate that MDTL-ACP can outperform the traditional and state-of-the-art ACP prediction methods. Junhang Cao, Wei Zhou 0001, Qiyuan Yu, Junkai Ji, Jun Zhang 0078, Shan He 0001, Zexuan Zhu 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | PAIR: protein-aptamer interaction prediction based on language models and contrastive learning frameworkabstractAptamers are single-stranded DNA or RNA oligonucleotides that selectively bind to specific targets, making them valuable for drug design and diagnostic applications. Identifying the interactions between aptamers and target proteins is crucial for these applications. The systematic evolution of ligands by exponential enrichment process, traditionally used for this purpose, is challenging and time-consuming. The resulting aptamers often suffer from limitations in stability and diversity. Computational approaches have shown promise in aiding the discovery of high-performance aptamers, but existing methods are usually constrained by insufficient training data and limited generalizability. Recently, advancements in pre-training large language models have offered a new avenue to mitigate the dependency on large datasets. In this study, we propose a novel method to predict aptamer-protein interactions using large language models within a contrastive learning framework. Experimental results demonstrate that our method exhibits superior generalization and outperforms existing approaches. This method holds promise as a powerful tool for predicting aptamer-protein interactions. Jun Zhang 0078, Zexuan Zhu 0001 |
BIBM | 4 |
| 2024 | An Elite Archive-Assisted Multi-Objective Evolutionary Algorithm for mRNA DesignabstractMessenger RNA (mRNA) vaccines have emerged as highly effective strategies in the prophylaxis and treatment of diseases. mRNA design, a key to the success of mRNA vaccines, in-volves finding optimal codons and increasing secondary structure stability to lengthen mRNA half-life, ultimately enhancing protein expression. Despite receiving widespread attention, most methods primarily rely on manual design, which is time-consuming and labor-intensive. While optimization approaches can alleviate this issue, existing methods still exhibit critical limitations caused by conflicts between codon usage and mRNA structural stability, compounded by the vast design space of mRNA resulting from the presence of synonymous codons. In this paper, a novel multi-objective evolutionary optimization-based mRNA design method is proposed. We first formulate the mRNA design problem as a multi-objective optimization problem and then develop an Elite Archive-Assisted Multi-Objective Evolutionary algorithm for mRNA Design, namely EAA-MOED, by incorporating a novel elite archive-assisted method into a weighted optimization framework to improve search efficiency. Experimental studies, involving two state-of-the-art mRNA design methods and five well-known MOEAs, show the competitiveness of the proposed EAA-MOED in mRNA design. Wenjing Hong, Cheng Chen 0072, Zexuan Zhu 0001, Ke Tang 0001 |
CEC | 3 |
| 2024 | Expensive Optimization Based on Evolutionary Multi-Tasking and Hybrid Restart StrategyabstractEvolutionary Algorithms (EAs) can not handle expensive optimization problems (EOPs) well due to the limited function evaluations in EOPs. To address this challenge, surrogate-assisted evolutionary algorithms (SAEAs) have been widely used and obtained good performance. With the problem dimension increases, SAEAs encounter some challenges in relatively high complexity on the training time and prediction time. To address this, this article proposes a novel expensive optimization algorithm with evolutionary multi-tasking and hybrid restart strategy (HRS-EMT). In the surrogate model construct, two radial basis function (RBF) models with different kernel functions are trained on all evaluated data to provide diversity and then are solved by a multi-tasking optimizer to a better optimization performance. In the surrogate model management, HRS-EMT combines multiple RBF models into an ensemble RBF (ERBF) model, strategically applied in the initial population pre-selection of surrogate model. Based on the prediction of ERBF surrogate model, HRS-EMT can obtain a better initial population in high dimensions. HRS-EMT is validated on twelve benchmark functions and compared with other state-of-the-art SAEAs. Experimental studies have shown the superior or comparable performance to other popular SAEAs in addressing EOPs. Zhenyuan Li, Xiaoliang Ma 0001, Zexuan Zhu 0001, Yueyue Li |
CEC | 3 |
| 2024 | Similar Locality Based Transfer Evolutionary Optimization for Minimalistic AttacksabstractDeep neural networks are powerful and popular learning models; however, recent studies have shown that deep neural network-based policies are susceptible to deception by adversarial attacks. A minimalistic attack is a specialized form of adversarial attack that aims to accomplish successful attacks at the lowest possible cost. Recently, transfer optimization algorithms have been applied to deceive previously trained policies by acquiring knowledge from previously solved tasks. Experiments indicate that the transfer optimization algorithms perform well compared to traditional optimization algorithms. However, current transfer algorithms for addressing minimalistic attacks not only select a single source task for knowledge transfer but also tend to overly rely on identified appropriate source tasks. To address this issue, this paper introduces a similar locality based transfer evolutionary optimization algorithm. It can adaptively select multiple source tasks and extract valuable knowledge from these source tasks. Moreover, by leveraging the concept of similar locality, the algorithm alleviates its excessive dependence on familiar tasks, thereby providing fresh knowledge for the optimization of the target task. On this basis, the algorithm can mine more valuable knowledge from the large source task space to achieve a successful attack in a shorter period. The algorithm is tested on three Atari games-BeamRider, Qbert, and Seaquest-demonstrating its ability and potential to outperform other transfer optimization algorithms currently available in solving this problem. Wenqiang Ma, Yaqing Hou, Hua Yu 0006, Xiangrong Tong, Zexuan Zhu 0001, Qiang Zhang 0008 |
CEC | 5 |
| 2024 | FreqFormer: Frequency-aware Transformer for Lightweight Image Super-resolution
Tao Dai 0001, Hang Guo 0002, Jinmin Li, Jinbao Wang 0001, Zexuan Zhu 0001 |
IJCAI | 6 |
| 2024 | DDN: Dual-domain Dynamic Normalization for Non-stationary Time Series ForecastingabstractDeep neural networks (DNNs) have recently achieved remarkable advancements in time series forecasting (TSF) due to their powerful ability of sequence dependence modeling. To date, existing DNN-based TSF methods still suffer from unreliable predictions for real-world data due to its non-stationarity characteristics, i.e., data distribution varies quickly over time. To mitigate this issue, several normalization methods (e.g., SAN) have recently been specifically designed by normalization in a fixed period/window in the time domain. However, these methods still struggle to capture distribution variations, due to the complex time patterns of time series in the time domain. Based on the fact that wavelet transform can decompose time series into a linear combination of different frequencies, which exhibits distribution variations with time-varying periods, we propose a novel Dual-domain Dynamic Normalization (DDN) to dynamically capture distribution variations in both time and frequency domains. Specifically, our DDN tries to eliminate the non-stationarity of time series via both frequency and time domain normalization in a sliding window way. Besides, our DDN can serve as a plug-in-play module, and thus can be easily incorporated into other forecasting models. Extensive experiments on public benchmark datasets under different forecasting models demonstrate the superiority of our DDN over other normalization methods. Code will be made available following the review process. Tao Dai 0001, Beiliang Wu, Peiyuan Liu, Naiqi Li, Xue Yuerong, Shutao Xia, Zexuan Zhu 0001 |
NeurIPS | 7 |
| 2024 | Parameter Efficient Adaptation for Image Restoration with Heterogeneous Mixture-of-ExpertsabstractDesigning single-task image restoration models for specific degradation has seen great success in recent years. To achieve generalized image restoration, all-in-one methods have recently been proposed and shown potential for multiple restoration tasks using one single model. Despite the promising results, the existing all-in-one paradigm still suffers from high computational costs as well as limited generalization on unseen degradations. In this work, we introduce an alternative solution to improve the generalization of image restoration models. Drawing inspiration from recent advancements in Parameter Efficient Transfer Learning (PETL), we aim to tune only a small number of parameters to adapt pre-trained restoration models to various tasks. However, current PETL methods fail to generalize across varied restoration tasks due to their homogeneous representation nature. To this end, we propose AdaptIR, a Mixture-of-Experts (MoE) with orthogonal multi-branch design to capture local spatial, global spatial, and channel representation bases, followed by adaptive base combination to obtain heterogeneous representation for different degradations. Extensive experiments demonstrate that our AdaptIR achieves stable performance on single-degradation tasks, and excels in hybrid-degradation tasks, with training only 0.6% parameters for 8 hours. Hang Guo 0002, Tao Dai 0001, Yuanchao Bai, Bin Chen 0011, Xudong Ren, Zexuan Zhu 0001, Shutao Xia |
NeurIPS | 6 |
| 2024 | A compressive seeding algorithm in conjunction with reordering-based compressionabstractMOTIVATION: Seeding is a rate-limiting stage in sequence alignment for next-generation sequencing reads. The existing optimization algorithms typically utilize hardware and machine-learning techniques to accelerate seeding. However, an efficient solution provided by professional next-generation sequencing compressors has been largely overlooked by far. In addition to achieving remarkable compression ratios by reordering reads, these compressors provide valuable insights for downstream alignment that reveal the repetitive computations accounting for more than 50% of seeding procedure in commonly used short read aligner BWA-MEM at typical sequencing coverage. Nevertheless, the exploited redundancy information is not fully realized or utilized. RESULTS: In this study, we present a compressive seeding algorithm, named CompSeed, to fill the gap. CompSeed, in collaboration with the existing reordering-based compression tools, finishes the BWA-MEM seeding process in about half the time by caching all intermediate seeding results in compact trie structures to directly answer repetitive inquiries that frequently cause random memory accesses. Furthermore, CompSeed demonstrates better performance as sequencing coverage increases, as it focuses solely on the small informative portion of sequencing reads after compression. The innovative strategy highlights the promising potential of integrating sequence compression and alignment to tackle the ever-growing volume of sequencing data. AVAILABILITY AND IMPLEMENTATION: CompSeed is available at https://github.com/i-xiaohu/CompSeed. Fahu Ji, Jue Ruan, Zexuan Zhu 0001 |
Bioinform. | 4 |
| 2024 | Dynamic constrained evolutionary optimization based on deep Q-network
Zhengping Liang, Ruitai Yang, Jigang Wang, Ling Liu 0003, Xiaoliang Ma 0001, Zexuan Zhu 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Multi-objective multi-task particle swarm optimization based on objective space division and adaptive transfer
Zhengping Liang, Jiabiao Yan, Jigang Wang, Ling Liu 0003, Zexuan Zhu 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Multifactorial Evolutionary Algorithm Based on Diffusion Gradient DescentabstractThe multifactorial evolutionary algorithm (MFEA) is one of the most widely used evolutionary multitasking (EMT) algorithms. The MFEA implements knowledge transfer among optimization tasks via crossover and mutation operators and it obtains high-quality solutions more efficiently than single-task evolutionary algorithms. Despite the effectiveness of MFEA in solving difficult optimization problems, there is no evidence of population convergence or theoretical explanations of how knowledge transfer increases algorithm performance. To fill this gap, we propose a new MFEA based on diffusion gradient descent (DGD), namely, MFEA-DGD in this article. We prove the convergence of DGD for multiple similar tasks and demonstrate that the local convexity of some tasks can help other tasks escape from local optima via knowledge transfer. Based on this theoretical foundation, we design complementary crossover and mutation operators for the proposed MFEA-DGD. As a result, the evolution population is endowed with a dynamic equation that is similar to DGD, that is, convergence is guaranteed, and the benefit from knowledge transfer is explainable. In addition, a hyper-rectangular search strategy is introduced to allow MFEA-DGD to explore more underdeveloped areas in the unified express space of all tasks and the subspace of each task. The proposed MFEA-DGD is verified experimentally on various multitask optimization problems, and the results demonstrate that MFEA-DGD can converge faster to competitive results compared to state-of-the-art EMT algorithms. We also show the possibility of interpreting the experimental results based on the convexity of different tasks. Zhaobo Liu, Zhengping Liang, Zexuan Zhu 0001 |
IEEE Trans. Cybern. | 5 |
| 2024 | Corrections to "Toward Adaptive Knowledge Transfer in Multifactorial Evolutionary Computation"abstractPresents corrections to the paper, (Corrections to "Toward Adaptive Knowledge Transfer in Multifactorial Evolutionary Computation"). Lei Zhou 0020, Liang Feng 0001, Kay Chen Tan, Jinghui Zhong, Zexuan Zhu 0001, Kai Liu 0001, Chao Chen 0004 |
IEEE Trans. Cybern. | 5 |
| 2024 | Evolutionary Multitasking for Costly Task Offloading in Mobile-Edge Computing NetworksabstractThe offloading of computation-intensive tasks to an edge server near resource-constrained mobile devices can provide improved application performance and user experience. However, with the rapid growth of mobile devices connected to the edge server, it is challenging to directly obtain an optimal task offloading scheme due to increasing computational cost and problem scale. In this study, we model the costly task offloading problem (CTOP) in mobile edge computing networks to achieve efficient joint optimization of energy consumption and processing latency for mobile devices. Inspired by the success of evolutionary multitasking in solving complex optimization problems by leveraging the experience of simple optimization problems, we develop a novel multitasking framework whose effectiveness is demonstrated in solving the CTOP. In this framework, auxiliary tasks are created to optimize the local processing overhead and the edge processing overhead of task offloading. On this basis, we propose an effective multitask evolutionary algorithm that includes segmented knowledge transfer and auxiliary task update. Specifically, source and extended decision variables are considered as different knowledge to be utilized, while the auxiliary tasks are allowed to be updated dynamically. Related knowledge that is learned from cheap and simple auxiliary tasks promotes the evolutionary search for CTOP. Experimental results verify the effectiveness of knowledge transfer. Compared to existing multitasking and single-tasking algorithms, the proposed algorithm shows competitive performance in CTOP instances and achieves better comprehensive performance in terms of energy consumption and processing latency. Chen Yang 0011, Qunjian Chen, Zexuan Zhu 0001, Zhi-an Huang, Shulin Lan, Liehuang Zhu |
IEEE Trans. Evol. Comput. | 3 |
| 2024 | Multitask Learning for Joint Diagnosis of Multiple Mental Disorders in Resting-State fMRIabstractFacing the increasing worldwide prevalence of mental disorders, the symptom-based diagnostic criteria struggle to address the urgent public health concern due to the global shortfall in well-qualified professionals. Thanks to the recent advances in neuroimaging techniques, functional magnetic resonance imaging (fMRI) has surfaced as a new solution to characterize neuropathological biomarkers for detecting functional connectivity (FC) anomalies in mental disorders. However, the existing computer-aided diagnosis models for fMRI analysis suffer from unstable performance on large datasets. To address this issue, we propose an efficient multitask learning (MTL) framework for joint diagnosis of multiple mental disorders using resting-state fMRI data. A novel multiobjective evolutionary clustering algorithm is presented to group regions of interests (ROIs) into different clusters for FC pattern analysis. On the optimal clustering solution, the multicluster multigate mixture-of-expert model is used for the final classification by capturing the highly consistent feature patterns among related diagnostic tasks. Extensive simulation experiments demonstrate that the performance of the proposed framework is superior to that of the other state-of-the-art methods. Moreover, the potential for practical application of the framework is also validated in terms of limited computational resources, real-time analysis, and insufficient training data. The proposed model can identify the remarkable interpretative biomarkers associated with specific mental disorders for clinical interpretation analysis. Zhi-an Huang, Rui Liu 0038, Zexuan Zhu 0001, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Attention-Like Multimodality Fusion With Data Augmentation for Diagnosis of Mental Disorders Using MRIabstractThe globally rising prevalence of mental disorders leads to shortfalls in timely diagnosis and therapy to reduce patients' suffering. Facing such an urgent public health problem, professional efforts based on symptom criteria are seriously overstretched. Recently, the successful applications of computer-aided diagnosis approaches have provided timely opportunities to relieve the tension in healthcare services. Particularly, multimodal representation learning gains increasing attention thanks to the high temporal and spatial resolution information extracted from neuroimaging fusion. In this work, we propose an efficient multimodality fusion framework to identify multiple mental disorders based on the combination of functional and structural magnetic resonance imaging. A multioutput conditional generative adversarial network (GAN) is developed to address the scarcity of multimodal data for augmentation. Based on the augmented training data, the multiheaded gating fusion model is proposed for classification by extracting the complementary features across different modalities. The experiments demonstrate that the proposed model can achieve robust accuracies of 75.1 ± 1.5 %, 72.9 ± 1.1 %, and 87.2 ± 1.5 % for autism spectrum disorder (ASD), attention deficit/hyperactivity disorder, and schizophrenia, respectively. In addition, the interpretability of our model is expected to enable the identification of remarkable neuropathology diagnostic biomarkers, leading to well-informed therapeutic decisions. Rui Liu 0038, Zhi-an Huang, Yao Hu 0001, Zexuan Zhu 0001, Ka-Chun Wong, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Spatial-Temporal Co-Attention Learning for Diagnosis of Mental Disorders From Resting-State fMRI DataabstractNeuroimaging techniques have been widely adopted to detect the neurological brain structures and functions of the nervous system. As an effective noninvasive neuroimaging technique, functional magnetic resonance imaging (fMRI) has been extensively used in computer-aided diagnosis (CAD) of mental disorders, e.g., autism spectrum disorder (ASD) and attention deficit/hyperactivity disorder (ADHD). In this study, we propose a spatial-temporal co-attention learning (STCAL) model for diagnosing ASD and ADHD from fMRI data. In particular, a guided co-attention (GCA) module is developed to model the intermodal interactions of spatial and temporal signal patterns. A novel sliding cluster attention module is designed to address global feature dependency of self-attention mechanism in fMRI time series. Comprehensive experimental results demonstrate that our STCAL model can achieve competitive accuracies of 73.0 ± 4.5%, 72.0 ± 3.8%, and 72.5 ± 4.2% on the ABIDE I, ABIDE II, and ADHD-200 datasets, respectively. Moreover, the potential for feature pruning based on the co-attention scores is validated by the simulation experiment. The clinical interpretation analysis of STCAL can allow medical professionals to concentrate on the discriminative regions of interest and key time frames from fMRI data. Rui Liu 0038, Zhi-an Huang, Yao Hu 0001, Zexuan Zhu 0001, Ka-Chun Wong, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | A Knowledge Transfer-Based Genetic Algorithm for Multi-Target Robotic Arm ControlabstractThe ability to swiftly and precisely reach any user-specified target location is necessary for a robotic arm that can be used in real-world scenarios. To date, many evolutionary optimization algorithms have been used to design controllers for robotic arms. However, when designing a robotic arm to reach multiple targets, most existing methods need to evolve the control strategy from scratch for each target, rather than trying to reuse existing experience. Therefore, computational resources are repeatedly and meaninglessly consumed. To this end, this paper proposes a genetic algorithm based on knowledge transfer (GAKT) dedicated to reusing existing knowledge to optimize a new robotic arm control task. Specifically, the knowledge transfer process can be summarized into the following two steps. First, through sequential transfer, GAKT initializes the population with the help of a knowledge base constructed by a quality diversity algorithm. Second, underperforming individuals are encouraged to acquire knowledge from excellent individuals in the same generation during the optimization process. We tested the effectiveness of GAKT and investigated its average performance by selecting multiple target points in different dimensions. The results show that GAKT can find the most advantageous arrival strategy (that is, make the end of the manipulator the closest to the target) on most of the selected targets. Moreover, we conducted ablation experiments and demonstrated the effectiveness of the knowledge transfer processes. Zhaoping Yu, Wenbin Pei, Yaqing Hou, Zexuan Zhu 0001, Xianneng Li |
CEC | 7 |
| 2023 | Adopting Autodock Koto for Virtual Screening of COVID-19
Zhangfan Yang, Junkai Ji, Zexuan Zhu 0001, Jianqiang Li 0001 |
ICIC (3) | 4 |
| 2023 | MIX-TPI: a flexible prediction framework for TCR-pMHC interactions based on multimodal representationsabstractMOTIVATION: The interactions between T-cell receptors (TCR) and peptide-major histocompatibility complex (pMHC) are essential for the adaptive immune system. However, identifying these interactions can be challenging due to the limited availability of experimental data, sequence data heterogeneity, and high experimental validation costs. RESULTS: To address this issue, we develop a novel computational framework, named MIX-TPI, to predict TCR-pMHC interactions using amino acid sequences and physicochemical properties. Based on convolutional neural networks, MIX-TPI incorporates sequence-based and physicochemical-based extractors to refine the representations of TCR-pMHC interactions. Each modality is projected into modality-invariant and modality-specific representations to capture the uniformity and diversities between different features. A self-attention fusion layer is then adopted to form the classification module. Experimental results demonstrate the effectiveness of MIX-TPI in comparison with other state-of-the-art methods. MIX-TPI also shows good generalization capability on mutual exclusive evaluation datasets and a paired TCR dataset. AVAILABILITY AND IMPLEMENTATION: The source code of MIX-TPI and the test data are available at: https://github.com/Wolverinerine/MIX-TPI. Zhi-an Huang, Wei Zhou 0001, Junkai Ji, Jun Zhang 0078, Shan He 0001, Zexuan Zhu 0001 |
Bioinform. | 7 |
| 2023 | An improved Hover-net for nuclear segmentation and classification in histopathology images
Lulu Qin, Bo-Wei Han, Zexuan Zhu 0001, Guangdong Qiao |
Neural Comput. Appl. | 6 |
| 2023 | Network Biomarker Detection From Gene Co-Expression Network Using Gaussian Mixture Model ClusteringabstractFinding network biomarkers from gene co-expression networks (GCNs) has attracted a lot of research interest. A network biomarker is a topological module, i.e., a group of densely connected nodes in a GCN, in which the gene expression values correlate with sample labels. Compared with biomarkers based on single genes, network biomarkers are not only more robust in separating samples from different categories, but are also able to better interpret the molecular mechanism of the disease. The previous network biomarker detection methods either employ distance based clustering methods or search for cliques in a GCN to detect topological modules. The first strategy assumes that the topological modules should be spherical in shape, and the second strategy requires all nodes to be fully connected. However, the relations between genes are complex, as a result, genes in the same biological process may not be directly, strongly connected. Therefore, the shapes of those modules could be oval or long strips. Hence, the shapes of gene functional modules and gene disease modules may not meet the aforementioned constraints in the previous methods. Thus, previous methods may break up the genes belonging to the same biological process into different topological modules due to those constraints. To address this issue, we propose a novel network biomarker detection method by using Gaussian mixture model clustering which allows more flexibility in the shapes of the topological modules. We have evaluated the performance of our method on a set of eight TCGA cancer datasets. The results show that our method can detect network modules that possess better discriminate power, and provide biological insights. Zexuan Zhu 0001, Hui Li 0117, Shan He 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Multiobjectivization of Single-Objective Optimization in Evolutionary Computation: A SurveyabstractMultiobjectivization has emerged as a new promising paradigm to solve single-objective optimization problems (SOPs) in evolutionary computation, where an SOP is transformed into a multiobjective optimization problem (MOP) and solved by an evolutionary algorithm to find the optimal solutions of the original SOP. The transformation of an SOP into an MOP can be done by adding helper-objective(s) into the original objective, decomposing the original objective into multiple subobjectives, or aggregating subobjectives of the original objective into multiple scalar objectives. Multiobjectivization bridges the gap between SOPs and MOPs by transforming an SOP into the counterpart MOP, through which multiobjective optimization methods manage to attain superior solutions of the original SOP. Particularly, using multiobjectivization to solve SOPs can reduce the number of local optima, create new search paths from local optima to global optima, attain more incomparability solutions, and/or improve solution diversity. Since the term "multiobjectivization" was coined by Knowles et al. in 2001, this subject has accumulated plenty of works in the last two decades, yet there is a lack of systematic and comprehensive survey of these efforts. This article presents a comprehensive multifacet survey of the state-of-the-art multiobjectivization methods. Particularly, a new taxonomy of the methods is provided in this article and the advantages, limitations, challenges, theoretical analyses, benchmarks, applications, as well as future directions of the multiobjectivization methods are discussed. Xiaoliang Ma 0001, Xiaodong Li 0001, Yutao Qi, Lei Wang 0018, Zexuan Zhu 0001 |
IEEE Trans. Cybern. | 6 |
| 2023 | Evolutionary Multitasking for Optimization Based on Generative StrategiesabstractEvolutionary multitasking (EMT) is one of the emerging topics in evolutionary computation. EMT can solve multiple related optimization tasks simultaneously and enhance the optimization of each task via knowledge sharing among tasks. Many EMT algorithms have been proposed and achieved success in various problems, yet EMT for multiobjective optimization remains a big challenge. The existing multiobjective EMT algorithms tend to suffer from slow convergence and difficulty in generating high-quality knowledge. To alleviate these issues, this article proposes a new EMT algorithm, namely, EMT-GS for multiobjective optimization based on two generative strategies. Particularly, generative adversarial networks (GANs) and inertial differential evolution (IDE) are introduced to generate transferable knowledge and offspring, respectively. A GAN is trained periodically for each source-target task pair, based on which helpful knowledge is generated from the source task and transferred to boost the solving of the target task. To accelerate the population convergence, the IDE strategy is put forward to generate offspring in a promising direction according to the individuals from the previous generation and the transferred knowledge. The performance of EMT-GS is validated on three multitasking multiobjective benchmark problems. The experimental results highlight the excellent competitiveness of EMT-GS compared to other state-of-the-art multiobjective EMT algorithms. Zhengping Liang, Yingmiao Zhu, Zhi Li 0020, Zexuan Zhu 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2023 | Fast Vehicle Routing via Knowledge Transfer in a Reproducing Kernel Hilbert SpaceabstractVehicle routing problems (VRPs) are essential in logistics. In the literature, many exact and heuristic optimization algorithms have been proposed to solve the VRPs. These traditional approaches, however, generally start the optimization from scratch and ignore the experiences of solving related VRPs, which may lead to unnecessary computational costs in searching repeated problems and reduce the efficiency of vehicle routing. Recently, transfer optimization (TO) has been presented to speed up vehicle routing by reusing the knowledge learned from similarly solved VRPs. However, existing TO methods build connections across VRPs in a low-dimensional Euclidean space, which has limited modeling ability in the cases of having nonlinear correlations. Keeping this in mind, this article presents a study of TO equipped with the kernel method for fast vehicle routing. In contrast to existing TO methods, in this work, the learning of connections across VRPs for knowledge transfer is conducted in a reproducing kernel Hilbert space (RKHS), which thus has greater modeling capacity in nonlinear customer relationships between VPRs. To evaluate the performance of the proposed method, comprehensive empirical studies have been conducted using well-known VRP benchmarks, against existing state-of-the-art TO methods for vehicle routing. Finally, a well-known real-world VRP application given by a routing company (Jingdong), namely, the package delivery problem (PDP), is investigated to further assess the efficacy of our proposed method. Liang Feng 0001, Min Li 0056, Yu Wang 0108, Zexuan Zhu 0001, Kay Chen Tan |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2023 | Expensive Multiobjective Optimization Based on Information Transfer SurrogateabstractObjective value estimation based on computationally efficient surrogate models is widely used to reduce the computational cost in solving expensive multiobjective optimization problems (MOPs). However, due to the scarcity of training data and the lack of data sharing between training tasks in a surrogate-based system, the estimation effectiveness of the surrogate models might not be satisfactory. In this study, we present a novel surrogate methodology based on information transfer to deal with this problem. Particularly, in the proposed framework, the objectives of an MOP that may have little apparent similarity or correlation are linearly mapped to a number of related tasks. Afterward, the related tasks are used to train a multitask Gaussian process (MTGP). MTGP expands the training data leading to more confident learning of the parameters of the model. The predicted values of the objective functions can be obtained by a reverse mapping from the learned MTGP model. In this way, the computational burden of the expensive objective functions of an MOP can be substantially reduced while maintaining good estimation accuracy. MTGP facilitates mutual information transfer across tasks, avoids learning from scratch for new tasks, and captures the underlying structural information between tasks. The proposed surrogate approach is merged into MOEA/D to address MOPs. Experimental tests under various scenarios indicate that the resultant algorithm outperforms other state-of-the-art surrogate-based multiobjective optimization algorithms. Jianping Luo, YongFei Dong, Zexuan Zhu 0001, Wenming Cao 0001, Xia Li 0006 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | A multifactorial differential evolution with hybrid global and local search strategiesabstractEvolutionary multitasking optimization (EMTO) solves multiple optimization tasks meanwhile in the framework of evolutionary algorithm, aiming at improving the solving performance on each task via knowledge transfer among tasks. As one of the representative EMTO algorithms, multifactorial evolutionary algorithm (MFEA) has attracted great attention and has been used to solve many optimization problems. However, most of MFEAs tend to suffer from premature convergence. To deal with this issue, this article designs a novel MFEA by integrating differential evolution, and a hybrid of global and local search strategies, named MFDE-GLS for short. Particularly, the global search strategy is based on an opposition-based learning and a Gaussian perturbation to improve the search ability and maintain population diversity. A local search strategy is introduced by combining 1-dimension search and n-dimension search to accelerate the convergence. Moreover, a new environmental selection mechanism is developed to keep the elite individuals while maintaining the population diversity based on the affinity propagation clustering method. Comprehensive experiments were conducted on both single-objective and multi-objective multi-task benchmark problems to show the effectiveness of the proposed algorithm. Yongjin Zheng, Yew-Soon Ong, Zexuan Zhu 0001, Xiaoliang Ma 0001 |
CEC | 4 |
| 2022 | Prediction of biomarker-disease associations based on graph attention network and text representationabstractMOTIVATION: The associations between biomarkers and human diseases play a key role in understanding complex pathology and developing targeted therapies. Wet lab experiments for biomarker discovery are costly, laborious and time-consuming. Computational prediction methods can be used to greatly expedite the identification of candidate biomarkers. RESULTS: Here, we present a novel computational model named GTGenie for predicting the biomarker-disease associations based on graph and text features. In GTGenie, a graph attention network is utilized to characterize diverse similarities of biomarkers and diseases from heterogeneous information resources. Meanwhile, a pretrained BERT-based model is applied to learn the text-based representation of biomarker-disease relation from biomedical literature. The captured graph and text features are then integrated in a bimodal fusion network to model the hybrid entity representation. Finally, inductive matrix completion is adopted to infer the missing entries for reconstructing relation matrix, with which the unknown biomarker-disease associations are predicted. Experimental results on HMDD, HMDAD and LncRNADisease data sets showed that GTGenie can obtain competitive prediction performance with other state-of-the-art methods. AVAILABILITY: The source code of GTGenie and the test data are available at: https://github.com/Wolverinerine/GTGenie. Zhi-an Huang, Wenhao Gu, Wenying Pan, Xiao Yang 0019, Zexuan Zhu 0001 |
Briefings Bioinform. | 7 |
| 2022 | CURC: a CUDA-based reference-free read compressorabstractMOTIVATION: The data deluge of high-throughput sequencing (HTS) has posed great challenges to data storage and transfer. Many specific compression tools have been developed to solve this problem. However, most of the existing compressors are based on central processing unit (CPU) platform, which might be inefficient and expensive to handle large-scale HTS data. With the popularization of graphics processing units (GPUs), GPU-compatible sequencing data compressors become desirable to exploit the computing power of GPUs. RESULTS: We present a GPU-accelerated reference-free read compressor, namely CURC, for FASTQ files. Under a GPU-CPU heterogeneous parallel scheme, CURC implements highly efficient lossless compression of DNA stream based on the pseudogenome approach and CUDA library. CURC achieves 2-6-fold speedup of the compression with competitive compression rate, compared with other state-of-the-art reference-free read compressors. AVAILABILITY AND IMPLEMENTATION: CURC can be downloaded from https://github.com/BioinfoSZU/CURC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shaohui Xie, Xiaotian He, Shan He 0001, Zexuan Zhu 0001 |
Bioinform. | 4 |
| 2022 | Multi-objective evolutionary multi-tasking algorithm using cross-dimensional and prediction-based knowledge transfer
Qunjian Chen, Xiaoliang Ma 0001, Zexuan Zhu 0001 |
Inf. Sci. | 5 |
| 2022 | Evolutionary Multitasking for Multiobjective Optimization With Subspace Alignment and Adaptive Differential EvolutionabstractIn contrast to the traditional single-tasking evolutionary algorithms, evolutionary multitasking (EMT) travels in the search space of multiple optimization tasks simultaneously. Through sharing knowledge across the tasks, EMT is able to enhance solving the optimization tasks. However, if knowledge transfer is not properly carried out, the performance of EMT might become unsatisfactory. To address this issue and improve the quality of knowledge transfer among the tasks, a novel multiobjective EMT algorithm based on subspace alignment and self-adaptive differential evolution (DE), namely, MOMFEA-SADE, is proposed in this article. Particularly, a mapping matrix obtained by subspace learning is used to transform the search space of the population and reduce the probability of negative knowledge transfer between tasks. In addition, DE characterized by a self-adaptive trial vector generation strategy is introduced to generate promising solutions based on previous experiences. The experimental results on multiobjective multi/many-tasking optimization test suites show that MOMFEA-SADE is superior or comparable to other state-of-the-art EMT algorithms. MOMFEA-SADE also won the Competition on Evolutionary Multitask Optimization (the multitask multiobjective optimization track) within IEEE 2019 Congress on Evolutionary Computation. Zhengping Liang, Weiqi Liang, Zexuan Zhu 0001 |
IEEE Trans. Cybern. | 5 |
| 2022 | A Dynamic Multiobjective Evolutionary Algorithm Based on Decision Variable ClassificationabstractIn recent years, dynamic multiobjective optimization problems (DMOPs) have drawn increasing interest. Many dynamic multiobjective evolutionary algorithms (DMOEAs) have been put forward to solve DMOPs mainly by incorporating diversity introduction or prediction approaches with conventional multiobjective evolutionary algorithms. Maintaining a good balance of population diversity and convergence is critical to the performance of DMOEAs. To address the above issue, a DMOEA based on decision variable classification (DMOEA-DVC) is proposed in this article. DMOEA-DVC divides the decision variables into two and three different groups in static optimization and changes response stages, respectively. In static optimization, two different crossover operators are used for the two decision variable groups to accelerate the convergence while maintaining good diversity. In change response, DMOEA-DVC reinitializes the three decision variable groups by maintenance, prediction, and diversity introduction strategies, respectively. DMOEA-DVC is compared with the other six state-of-the-art DMOEAs on 33 benchmark DMOPs. The experimental results demonstrate that the overall performance of the DMOEA-DVC is superior or comparable to that of the compared algorithms. Zhengping Liang, Xiaoliang Ma 0001, Zexuan Zhu 0001, Shengxiang Yang |
IEEE Trans. Cybern. | 4 |
| 2022 | Enhanced Multifactorial Evolutionary Algorithm With Meme Helper-TasksabstractEvolutionary multitasking (EMT) is an emerging research direction in the field of evolutionary computation. EMT solves multiple optimization tasks simultaneously using evolutionary algorithms with the aim to improve the solution for each task via intertask knowledge transfer. The effectiveness of intertask knowledge transfer is the key to the success of EMT. The multifactorial evolutionary algorithm (MFEA) represents one of the most widely used implementation paradigms of EMT. However, it tends to suffer from noneffective or even negative knowledge transfer. To address this issue and improve the performance of MFEA, we incorporate a prior-knowledge-based multiobjectivization via decomposition (MVD) into MFEA to construct strongly related meme helper-tasks. In the proposed method, MVD creates a related multiobjective optimization problem for each component task based on the corresponding problem structure or decision variable grouping to enhance positive intertask knowledge transfer. MVD can reduce the number of local optima and increase population diversity. Comparative experiments on the widely used test problems demonstrate that the constructed meme helper-tasks can utilize the prior knowledge of the target problems to improve the performance of MFEA. Xiaoliang Ma 0001, Jian Yin 0004, Anmin Zhu, Xiaodong Li 0001, Lei Wang 0018, Yutao Qi, Zexuan Zhu 0001 |
IEEE Trans. Cybern. | 8 |
| 2022 | Evolutionary Many-Task Optimization Based on Multisource Knowledge TransferabstractMultitask optimization aims to solve two or more optimization tasks simultaneously by leveraging intertask knowledge transfer. However, as the number of tasks increases to the extent of many-task optimization, the knowledge transfer between tasks encounters more uncertainty and challenges, thereby resulting in degradation of optimization performance. To give full play to the many-task optimization framework and minimize the potential negative transfer, this article proposes an evolutionary many-task optimization algorithm based on a multisource knowledge transfer mechanism, namely, EMaTO-MKT. Particularly, in each iteration, EMaTO-MKT determines the probability of using knowledge transfer adaptively according to the evolution experience, and balances the self-evolution within each task and the knowledge transfer among tasks. To perform knowledge transfer, EMaTO-MKT selects multiple highly similar tasks in terms of maximum mean discrepancy as the learning sources for each task. Afterward, a knowledge transfer strategy based on local distribution estimation is applied to enable the learning from multiple sources. Compared with the other state-of-the-art evolutionary many-task algorithms on benchmark test suites, EMaTO-MKT shows competitiveness in solving many-task optimization problems. Zhengping Liang, Xiuju Xu, Ling Liu 0003, Yaofeng Tu, Zexuan Zhu 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2022 | Merged Differential Grouping for Large-Scale Global OptimizationabstractThe divide-and-conquer strategy has been widely used in cooperative co-evolutionary algorithms to deal with large-scale global optimization problems, where a target problem is decomposed into a set of lower-dimensional and tractable subproblems to reduce the problem complexity. However, such a strategy usually demands a large number of function evaluations to obtain an accurate variable grouping. To address this issue, a merged differential grouping (MDG) method is proposed in this article based on the subset–subset interaction and binary search. In the proposed method, each variable is first identified as either a separable variable or a nonseparable variable. Afterward, all separable variables are put into the same subset, and the nonseparable variables are divided into multiple subsets using a binary-tree-based iterative merging method. With the proposed algorithm, the computational complexity of interaction detection is reduced to$O(\max \{n,n_{ns}\times \log _{2} k\})$, where$n$,$n_{ns}(\leq n)$, and$k( < n)$indicate the numbers of decision variables, nonseparable variables, and subsets of nonseparable variables, respectively. The experimental results on benchmark problems show that MDG is very competitive with the other state-of-the-art methods in terms of efficiency and accuracy of problem decomposition. Xiaoliang Ma 0001, Xiaodong Li 0001, Lei Wang 0018, Yutao Qi, Zexuan Zhu 0001 |
IEEE Trans. Evol. Comput. | 6 |
| 2022 | Multiobjective Evolutionary Multitasking With Two-Stage Adaptive Knowledge Transfer Based on Population DistributionabstractMultitasking optimization can achieve better performance than traditional single-tasking optimization by leveraging knowledge transfer between tasks. However, the current multitasking optimization algorithms suffer from some deficiencies. Particularly, on high similar problems, the existing algorithms might fail to take full advantage of knowledge transfer to accelerate the convergence of the search, or easily get trapped in the local optima. Whereas, on low similar problems, they tend to suffer from negative transfer, resulting in performance degradation. To solve these issues, this article proposes an evolutionary multitasking optimization algorithm for multiobjective/many-objective optimization with two-stage adaptive knowledge transfer based on population distribution. The resultant algorithm named EMT-PD can improve the convergence performance of the target optimization tasks based on the knowledge extracted from the probability model that reflects the search trend of the whole population. At the first stage of knowledge transfer, an adaptive weight is used to adjust the search step size of each individual, which can reduce the impact of negative transfer. At the second stage of knowledge transfer, the search range of each individual is further adjusted dynamically, which can improve the population diversity and be beneficial for jumping out of the local optima. Experimental results on multitasking multiobjective optimization test suites show that EMT-PD is superior to other state-of-the-art evolutionary multitasking/single-tasking algorithms. To further investigate the effectiveness of EMT-PD on many-objective optimization problems, a multitasking many-objective optimization test suite is also designed in this article. The experimental results on the new test suite also demonstrate the competitiveness of EMT-PD. Zhengping Liang, Weiqi Liang, Xiaoliang Ma 0001, Ling Liu 0003, Zexuan Zhu 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2021 | An Adaptive Multi-objective Multifactorial Evolutionary Algorithm Based on Mixture Gaussian DistributionabstractIn recent decades, multi-objective multifactorial evolutionary algorithm (MOMFEA) has become a very promising research direction. How to achieve effective knowledge transfer between similar tasks is the key issue to affect the performance of the algorithm. In this paper, an adaptive MOMFEA (AMOMFEA) is proposed by exploiting the mixture Gaussian distribution of the population distributions of related tasks to help solve the target task. Wasserstein distance is used to measure the inter-task relevance in that the weight coefficient in the mixture distribution is proportional to the inter-task relevance. Experimental results on benchmark problems validate the effectiveness and efficiency of the proposed method in comparison with MOMFEA and NSGA-II. Mengfan Xu, Zexuan Zhu 0001, Yutao Qi, Lei Wang 0018, Xiaoliang Ma 0001 |
CEC | 2 |
| 2021 | Mix-order Attention Networks for Image RestorationabstractConvolutional neural networks (CNNs) have obtained great success in image restoration tasks, like single image denoising, demosaicing, and super-resolution. However, most existing CNN-based methods neglect the diversity of image contents and degradations in the corrupted images and treat channel-wise features equally, thus hindering the representation ability of CNNs. To address this issue, we propose deep mix-order attention networks (MAN) to extract features that capture rich feature statistics within networks. Our MAN is mainly built on simple residual blocks and our mix-order channel attention (MOCA) module, which further consists of feature gating and feature pooling blocks to capture different types of semantic information. With our MOCA, our MAN can be flexible to handle various types of image contents and degradations. Besides, our MAN can be generalized to different image restoration tasks, like image denoising, super-resolution, and demosaicing. Extensive experiments demonstrate that our method obtains favorably against state-of-the-art methods in terms of quantitative and qualitative metrics. Tao Dai 0001, Yalei Lv, Bin Chen 0011, Zhi Wang 0001, Zexuan Zhu 0001, Shutao Xia |
ACM Multimedia | 5 |
| 2021 | Memetic Algorithm Based on Community Detection for Energy-Efficient Service Migration Optimization in 5G Mobile Edge ComputingabstractMobile edge computing (MEC) can supplement cloud computing by helping to overcome the limitations of long physical transmission distances and accelerating the responsiveness of edge computing servers. In 5G (fifth generation) cellular networks, adopting MEC can guarantee ultralow latency. To enhance the MEC quality, optimization of the user service profile migration according to the user mobility is essential. However, this optimization establishes an NP-hard problem. Moreover, high-speed 5G base stations with MEC servers often experience high energy consumption. As conventional service migration algorithms such as those based on profile tracking and game theory tend to fall in local optima and neglect energy consumption constraints, we propose a memetic algorithm based on community detection local search (MA-CDLS) to continuously optimize the service migration in 5G MEC scenarios. During busy periods or in crowded areas, MA-CDLS adopts a single-objective optimization of user-perceived latency to achieve high-performance 5G services. During light-load periods or in uncrowded areas, MA-CDLS uses two measures, namely the user-perceived latency and energy consumption, to realize energy-efficient 5G services. MA-CDLS effectively reduces the search space and speeds up the elite selection in the meme operator. Experiments in simulated scenarios show that MA-CDLS achieves a lower user-perceived latency and energy consumption, than the traditional profile tracking and game theory methods, especially during congestion. Ling Liu 0003, Zhengping Liang, Xiaoliang Ma 0001, Zexuan Zhu 0001 |
PIMRC | 5 |
| 2021 | A feedback-based prediction strategy for dynamic multi-objective evolutionary optimization
Zhengping Liang, Ya Zou, Shunxiang Zheng, Shengxiang Yang, Zexuan Zhu 0001 |
Expert Syst. Appl. | 5 |
| 2021 | Solving Generalized Vehicle Routing Problem With Occasional Drivers via Evolutionary MultitaskingabstractWith the emergence of crowdshipping and sharing economy, vehicle routing problem with occasional drivers (VRPOD) has been recently proposed to involve occasional drivers with private vehicles for the delivery of goods. In this article, we present a generalized variant of VRPOD, namely, the vehicle routing problem with heterogeneous capacity, time window, and occasional driver (VRPHTO), by taking the capacity heterogeneity and time window of vehicles into consideration. Furthermore, to meet the requirement in today's cloud computing service, wherein multiple optimization tasks may need to be solved at the same time, we propose a novel evolutionary multitasking algorithm (EMA) to optimize multiple VRPHTOs simultaneously with a single population. Finally, 56 new VRPHTO instances are generated based on the existing common vehicle routing benchmarks. Comprehensive empirical studies are conducted to illustrate the benefits of the new VRPHTOs and to verify the efficacy of the proposed EMA for multitasking against a state-of-art single task evolutionary solver. The obtained results showed that the employment of occasional drivers could significantly reduce the routing cost, and the proposed EMA is not only able to solve multiple VRPHTOs simultaneously but also can achieve enhanced optimization performance via the knowledge transfer between tasks along the evolutionary search process. Liang Feng 0001, Lei Zhou 0020, Abhishek Gupta 0001, Jinghui Zhong, Zexuan Zhu 0001, Kay Chen Tan, A. K. Qin 0001 |
IEEE Trans. Cybern. | 5 |
| 2021 | A Many-Objective Evolutionary Algorithm Based on a Two-Round Selection StrategyabstractBalancing population diversity and convergence is critical for evolutionary algorithms to solve many-objective optimization problems (MaOPs). In this paper, a two-round environmental selection strategy is proposed to pursue good tradeoff between population diversity and convergence for many-objective evolutionary algorithms (MaOEAs). Particularly, in the first round, the solutions with small neighborhood density are picked out to form a candidate pool, where the neighborhood density of a solution is calculated based on a novel adaptive position transformation strategy. In the second round, the best solution in terms of convergence is selected from the candidate pool and inserted into the next generation. The procedure is repeated until a new population is generated. The two-round selection strategy is embedded into an MaOEA framework and the resulting algorithm, namely, 2REA, is compared with eight state-of-the-art MaOEAs on various benchmark MaOPs. The experimental results show that 2REA is very competitive with the compared MaOEAs and the two-round selection strategy works well on balancing population diversity and convergence. Zhengping Liang, Kaifeng Hu, Xiaoliang Ma 0001, Zexuan Zhu 0001 |
IEEE Trans. Cybern. | 4 |
| 2021 | An Indicator-Based Many-Objective Evolutionary Algorithm With Boundary ProtectionabstractMany-objective optimization problems (MaOPs) pose a big challenge to the traditional Pareto-based multiobjective evolutionary algorithms (MOEAs). As the number of objectives increases, the number of mutually nondominated solutions explodes and MOEAs become invalid due to the loss of Pareto-based selection pressure. Indicator-based many-objective evolutionary algorithms (MaOEAs) have been proposed to address this issue by enhancing the environmental selection. Indicator-based MaOEAs are easy to implement and of good versatility, however, they are unlikely to maintain the population diversity and coverage very well. In this article, a new indicator-based MaOEA with boundary protection, namely, MaOEA-IBP, is presented to relieve this weakness. In MaOEA-IBP, a worst elimination mechanism based on the${I}_{{\epsilon }^{+}}$indicator and boundary protection strategy is devised to enhance the balance of population convergence, diversity, and coverage. Specifically, a pair of solutions with the smallest${I}_{{\epsilon }^{+}}$value are first identified from the population. If one solution dominates the other, the dominated solution is eliminated. Otherwise, one solution is eliminated by the boundary protection strategy. MaOEA-IBP is compared with four indicator-based algorithms (i.e.,${I}_{{{ {SDE}}}^{+}}$, SRA, MaOEAIGD, and ARMOEA) and other five state-of-the-art MaOEAs (i.e., KnEA, MaOEA-CSS, 1by1EA, RVEA, and EFR-RR) on various benchmark MaOPs. The experimental results demonstrate that MaOEA-IBP can achieve competitive performance with the compared algorithms. Zhengping Liang, Tingting Luo, Kaifeng Hu, Xiaoliang Ma 0001, Zexuan Zhu 0001 |
IEEE Trans. Cybern. | 5 |
| 2021 | Toward Adaptive Knowledge Transfer in Multifactorial Evolutionary ComputationabstractA multifactorial evolutionary algorithm (MFEA) is a recently proposed algorithm for evolutionary multitasking, which optimizes multiple optimization tasks simultaneously. With the design of knowledge transfer among different tasks, MFEA has demonstrated the capability to outperform its single-task counterpart in terms of both convergence speed and solution quality. In MFEA, the knowledge transfer across tasks is realized via the crossover between solutions that possess different skill factors. This crossover is thus essential to the performance of MFEA. However, we note that the present MFEA and most of its existing variants only employ a single crossover for knowledge transfer, and fix it throughout the evolutionary search process. As different crossover operators have a unique bias in generating offspring, the appropriate configuration of crossover for knowledge transfer in MFEA is necessary toward robust search performance, for solving different problems. Nevertheless, to the best of our knowledge, there is no effort being conducted on the adaptive configuration of crossovers in MFEA for knowledge transfer, and this article thus presents an attempt to fill this gap. In particular, here, we first investigate how different types of crossover affect the knowledge transfer in MFEA on both single-objective (SO) and multiobjective (MO) continuous optimization problems. Furthermore, toward robust and efficient multitask optimization performance, we propose a new MFEA with adaptive knowledge transfer (MFEA-AKT), in which the crossover operator employed for knowledge transfer is self-adapted based on the information collected along the evolutionary search process. To verify the effectiveness of the proposed method, comprehensive empirical studies on both SO and MO multitask benchmarks have been conducted. The experimental results show that the proposed MFEA-AKT is able to identify the appropriate knowledge transfer crossover for different optimization problems and even at different optimization stages along the search, which thus leads to superior or competitive performances when compared to the MFEAs with fixed knowledge transfer crossover operators. Lei Zhou 0020, Liang Feng 0001, Kay Chen Tan, Jinghui Zhong, Zexuan Zhu 0001, Kai Liu 0001, Chao Chen 0004 |
IEEE Trans. Cybern. | 5 |
| 2021 | Multimodal Multiobjective Evolutionary Optimization With Dual Clustering in Decision and Objective SpacesabstractThis article suggests a multimodal multiobjective evolutionary algorithm with dual clustering in decision and objective spaces. One clustering is run in decision space to gather nearby solutions, which will classify solutions into multiple local clusters. Nondominated solutions within each local cluster are first selected to maintain local Pareto sets, and the remaining ones with good convergence in objective space are also selected, which will form a temporary population with more than${N}$solutions (${N}$is the population size). After that, a second clustering is run in objective space for this temporary population to get${N}$final clusters with good diversity in objective space. Finally, a pruning process is repeatedly run on the above clusters until each cluster has only one solution, which removes the most crowded solution in decision space from the most crowded cluster in objective space each time. This way, the clustering in decision space can distinguish all Pareto sets and avoid the loss of local Pareto sets, while that in objective space can maintain diversity in objective space. When solving all the benchmark problems from the competition of multimodal multiobjective optimization in the IEEE Congress on Evolutionary Computation 2019, the experiments validate our advantages to maintain diversity in both objective and decision spaces. Qiuzhen Lin, Wu Lin, Zexuan Zhu 0001, Maoguo Gong, Jianqiang Li 0001, Carlos A. Coello Coello |
IEEE Trans. Evol. Comput. | 3 |
| 2021 | Identifying Autism Spectrum Disorder From Resting-State fMRI Using Deep Belief NetworkabstractWith the increasing prevalence of autism spectrum disorder (ASD), it is important to identify ASD patients for effective treatment and intervention, especially in early childhood. Neuroimaging techniques have been used to characterize the complex biomarkers based on the functional connectivity anomalies in the ASD. However, the diagnosis of ASD still adopts the symptom-based criteria by clinical observation. The existing computational models tend to achieve unreliable diagnostic classification on the large-scale aggregated data sets. In this work, we propose a novel graph-based classification model using the deep belief network (DBN) and the Autism Brain Imaging Data Exchange (ABIDE) database, which is a worldwide multisite functional and structural brain imaging data aggregation. The remarkable connectivity features are selected through a graph extension of K -nearest neighbors and then refined by a restricted path-based depth-first search algorithm. Thanks to the feature reduction, lower computational complexity could contribute to the shortening of the training time. The automatic hyperparameter-tuning technique is introduced to optimize the hyperparameters of the DBN by exploring the potential parameter space. The simulation experiments demonstrate the superior performance of our model, which is 6.4% higher than the best result reported on the ABIDE database. We also propose to use the data augmentation and the oversampling technique to identify further the possible subtypes within the ASD. The interpretability of our model enables the identification of the most remarkable autistic neural correlation patterns from the data-driven outcomes. Zhi-an Huang, Zexuan Zhu 0001, Chuen Heung Yau, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Identification of Autistic Risk Candidate Genes and Toxic Chemicals via Multilabel LearningabstractAs a group of complex neurodevelopmental disorders, autism spectrum disorder (ASD) has been reported to have a high overall prevalence, showing an unprecedented spurt since 2000. Due to the unclear pathomechanism of ASD, it is challenging to diagnose individuals with ASD merely based on clinical observations. Without additional support of biochemical markers, the difficulty of diagnosis could impact therapeutic decisions and, therefore, lead to delayed treatments. Recently, accumulating evidence have shown that both genetic abnormalities and chemical toxicants play important roles in the onset of ASD. In this work, a new multilabel classification (MLC) model is proposed to identify the autistic risk genes and toxic chemicals on a large-scale data set. We first construct the feature matrices and partially labeled networks for autistic risk genes and toxic chemicals from multiple heterogeneous biological databases. Based on both global and local measure metrics, the simulation experiments demonstrate that the proposed model achieves superior classification performance in comparison with the other state-of-the-art MLC methods. Through manual validation with existing studies, 60% and 50% out of the top-20 predicted risk genes are confirmed to have associations with ASD and autistic disorder, respectively. To the best of our knowledge, this is the first computational tool to identify ASD-related risk genes and toxic chemicals, which could lead to better therapeutic decisions of ASD. Zhi-an Huang, Jia Zhang 0019, Zexuan Zhu 0001, Qi Wu 0003, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Tracking Moving Optima of Dynamic Multi-objective Problem via Prediction in Objective SpaceabstractSolving dynamic multi-objective optimization problem (DMOP) requires optimizing multiple conflicting objectives simultaneously. When a dynamic is detected in the changing environment, most of existing prediction-based strategies predict the trajectory of changing Pareto-optimal solutions (POS), based on the historical solutions obtained in the solution space. In this paper, we present a new prediction method to track the moving optima for solving DMOP. In contrast to existing approaches, we propose to build the prediction model in the objective space. As the evaluation for solving a DMOP is based on the Pareto-optimal front (POF), to predict directly in the objective space could provide more useful information than the prediction in the solution space. In particular, to efficiently capture the complex relationships among POFs found along the evolutionary search, here we build a prediction model in Reproducing Kernel Hilbert Space, which holds a closed-form solution. To evaluate the performance of the proposed method, empirical studies have been conducted by comparing against three state-of-the-art prediction-based strategies on fourteen commonly used DMOP benchmarks. The results obtained by using different optimization solvers confirmed the superiority of the proposed method for solving DMOP in terms of both solution quality and time efficiency. Wei Zhou 0001, Liang Feng 0001, Zexuan Zhu 0001, Kai Liu 0001, Chao Chen 0004, Zhou Wu 0001 |
CEC | 3 |
| 2020 | Multi-objective multi-factorial memetic algorithm based on bone route and large neighborhood local search for VRPTWabstractMulti-tasking optimization (MTO) has attracted increasing attention in the domain of evolutionary computation. Different from single-tasking optimization, MTO can solve multiple optimization tasks simultaneously to improve the performance of solving each optimization task by inter-task knowledge transfer. Multifactorial evolutionary algorithm (MFEA) is one of the most widely used MTO algorithm based on assortative mating and vertical cultural transmission. This work extends MFEA by integrating bone route and large neighborhood local search to solve multi-objective vehicle routing problem with time window (VRPTW). The VRPTW is modeled as two related tasks, i.e., one is a multi-objective version of VRPTW (the main task), and the other is a single-objective version of VRPTW (the auxiliary task). The resultant new algorithm namely multi-objective multi-factorial memetic algorithm (MOMFMA) solve the two tasks simultaneously where the information between the tasks is exchanged in the evolutionary process. In addition to the implicit information transfer of MFEA, the bone route is introduced to enable explicit information transfer between tasks. Particularly, bone routes are constructed as semi-finished product solutions and used in large neighborhood local search. The bone route and the large neighborhood local search work together to speed up the convergence of the algorithm. MOMFMA is tested on Solomon's 56 datasets and the experimental results demonstrate that the efficiency of MOMFMA. Zifeng Zhou, Xiaoliang Ma 0001, Zhengping Liang, Zexuan Zhu 0001 |
CEC | 4 |
| 2020 | ATEN: And/Or tree ensemble for inferring accurate Boolean network topology and dynamicsabstractMOTIVATION: Inferring gene regulatory networks from gene expression time series data is important for gaining insights into the complex processes of cell life. A popular approach is to infer Boolean networks. However, it is still a pressing open problem to infer accurate Boolean networks from experimental data that are typically short and noisy. RESULTS: To address the problem, we propose a Boolean network inference algorithm which is able to infer accurate Boolean network topology and dynamics from short and noisy time series data. The main idea is that, for each target gene, we use an And/Or tree ensemble algorithm to select prime implicants of which each is a conjunction of a set of input genes. The selected prime implicants are important features for predicting the states of the target gene. Using these important features we then infer the Boolean function of the target gene. Finally, the Boolean functions of all target genes are combined as a Boolean network. Using the data generated from artificial and real-world gene regulatory networks, we show that our algorithm can infer more accurate Boolean network topology and dynamics from short and noisy time series data than other algorithms. Our algorithm enables us to gain better insights into complex regulatory mechanisms of cell life. AVAILABILITY AND IMPLEMENTATION: Package ATEN is freely available at https://github.com/ningshi/ATEN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ning Shi, Zexuan Zhu 0001, Ke Tang 0001, David Parker 0001, Shan He 0001 |
Bioinform. | 2 |
| 2020 | MUMI: Multitask Module Identification for Biological NetworksabstractIdentifying modules from biological networks is important since modules reveal essential mechanisms and dynamic processes in biological systems. Existing algorithms focus on identifying either active modules or topological modules (communities), which represent dynamic and topological units in the network, respectively. However, high-level biological phenomena, e.g., functions are emergent properties from the interplay between network topology and dynamics. Therefore, to fully explain the mechanisms underlying the high-level biological phenomena, it is important to identify the overlaps between communities and active modules, which indicate the topological units with significant changes of dynamics. However, despite the importance, there are no existing methods to do so. In this article, we propose the multitask module identification (MUMI) algorithm to detect the overlaps between active modules and communities simultaneously. The experimental results show that our method provides new insights into biological mechanisms by combining information from active modules and communities. By formulating the problem as a multitasking learning problem which searches for these two types of modules simultaneously, the algorithm can exploit their latent complementarities to obtain better search performance in terms of accuracy and convergence. Our MATLAB implementation of MUMI is available at https://github.com/WeiqiChen/Mumi-multitask-module-identification. Zexuan Zhu 0001, Shan He 0001 |
IEEE Trans. Evol. Comput. | 2 |
| 2020 | A Survey of Weight Vector Adjustment Methods for Decomposition-Based Multiobjective Evolutionary AlgorithmsabstractMultiobjective evolutionary algorithms based on decomposition (MOEA/D) have attracted tremendous attention and achieved great success in the fields of optimization and decision-making. MOEA/Ds work by decomposing the target multiobjective optimization problem (MOP) into multiple single-objective subproblems based on a set of weight vectors. The subproblems are solved cooperatively in an evolutionary algorithm framework. Since weight vectors define the search directions and, to a certain extent, the distribution of the final solution set, the configuration of weight vectors is pivotal to the success of MOEA/Ds. The most straightforward method is to use predefined and uniformly distributed weight vectors. However, it usually leads to the deteriorated performance of MOEA/Ds on solving MOPs with irregular Pareto fronts. To deal with this issue, many weight vector adjustment methods have been proposed by periodically adjusting the weight vectors in a random, predefined, or adaptive way. This article focuses on weight vector adjustment on a simplex and presents a comprehensive survey of these weight vector adjustment methods covering the weight vector adaptation strategies, theoretical analyses, benchmark test problems, and applications. The current limitations, new challenges, and future directions of weight vector adjustment are also discussed. Xiaoliang Ma 0001, Xiaodong Li 0001, Yutao Qi, Zexuan Zhu 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2019 | Multimodal Multi-objective Optimization Using A Density-based One-by-One Update StrategyabstractFor real-world optimization problems, a uniformly and widely distributed Pareto optimal set (PS) in the decision space can provide more choices for decision makers. However, most of multi-objective evolutionary algorithms (MOEAs) only consider convergence and diversity in the objective space, which rarely pay attention to diversity in the decision space. Especially for multimodal multi-objective optimization problems (MMOPs), there may exist multiple distinct PSs corresponding to the same Pareto front (PF). Thus, we propose a novel multimodal multi-objective evolutionary algorithm using a density-based one-by-one update strategy in this paper, which considers diversity in both the objective and decision spaces. In the proposed algorithm, once an offspring is generated during evolution, the most crowded subregion with the largest niche count in the objective space has to be identified again, helpful to maintain diversity in the objective space. Furthermore, the harmonic average distance approach is used to estimate the global density of solutions in the decision space, trying to maintain the population's diversity in the decision space. Our proposed algorithm is compared with several state-of-the-art algorithms on MMOPs. The experimental results demonstrate that our algorithm is capable of preserving promising solutions with even distribution in both of decision space and objective space and also shows the superiority on solving the adopted MMOPs. Ruizhi Shi, Wu Lin, Qiuzhen Lin, Zexuan Zhu 0001, Jianyong Chen |
CEC | 4 |
| 2019 | Multifactorial Evolutionary Algorithm Enhanced with Cross-task Search DirectionabstractRecently, the multifactorial evolutionary algorithm (MFEA) has achieved remarkable success in multi-task optimization (MTO) and received extensive attention from academia and industry. The key idea of MFEA is to use the inter-task knowledge transfer to produce the mutual promotion effect of all tasks. However, MFEA still has some limitations in accelerating convergence and enhancing global search ability, especially when the optima of different optimization tasks are far away. To relieve this issue, this paper integrates a new cross-task knowledge transfer, which is based on a search direction instead of an individual. The proposed knowledge transfer strategy generates offspring by the sum of an elite individual of one task and a difference vector from another task. As a basic vector, the elite individual is used to speed up the population convergence. Adding the elite individual with a difference vector from another task can enhance the search diversity. The experimental studies have shown the effectiveness and efficiency of the proposed cross-task knowledge transfer strategy, compared with the classical MFEA on a set of benchmark problems with different degrees of similarities. Jian Yin 0004, Anmin Zhu, Zexuan Zhu 0001, Xiaoliang Ma 0001 |
CEC | 3 |
| 2019 | Multifactorial Differential Evolution with Opposition-based Learning for Multi-tasking OptimizationabstractRecently, multi-tasking optimization (MTO) has become a rising research topic in the field of evolutionary computation that has attracted increasing attention of academia. Comparing with single-objective optimization (SOO) and multi-objective optimization (MOO), MTO can solve different optimization tasks simultaneously by utilizing inter-task similarities and complementarities. Based on crossover operator, the classical multifactorial evolutionary algorithm (MFEA) transfers inter-task knowledge. To broaden the search region and accelerate the convergence, this paper integrates differential evolution (DE) and opposition-based learning (OBL) into MFEA and hence proposes MFEA/DE-OBL. The motivation of integrating DE and OBL is that they have different search neighborhoods and strong complementarity with simulated binary crossover (SBX) used in MFEA. Furthermore, integrating DE and OBL can help MFEA jump out of local optima. The effectiveness and efficiency of integrating DE and OBL into MFEA are experimentally studied on a set of benchmark problems with different degrees of similarities. Experimental results demonstrate that the proposed MFEA/DE-OBL dramatically improves the performance compared with the MFEA. Anmin Zhu, Zexuan Zhu 0001, Qiuzhen Lin, Jian Yin 0004, Xiaoliang Ma 0001 |
CEC | 3 |
| 2019 | Multi-objective memetic algorithm based on correlation priority for pickup-and-delivery problemsabstractThis paper presents a multi-objective memetic algorithm based on correlation priority to solve route planning of electric vehicles in pickup-and-delivery problems. Four objectives namely route length, waiting time, charging times, and the number of vehicles are optimized using multi-objective memetic algorithm, which is a combination of multi-objective genetic algorithm, greedy strategy, and a correlation priority based local search. The correlation between two customer nodes is used to fine-tune the route to accelerate the convergence of the algorithm. The algorithm is tested on three sets of data with different scales and the experimental results demonstrate the efficiency of the proposed algorithm. Zifeng Zhou, Xiaoliang Ma 0001, Zexuan Zhu 0001 |
CEC | 3 |
| 2019 | Predicting synthetic lethal interactions in human cancers using graph regularized self-representative matrix factorizationabstractBACKGROUND: Synthetic lethality has attracted a lot of attentions in cancer therapeutics due to its utility in identifying new anticancer drug targets. Identifying synthetic lethal (SL) interactions is the key step towards the exploration of synthetic lethality in cancer treatment. However, biological experiments are faced with many challenges when identifying synthetic lethal interactions. Thus, it is necessary to develop computational methods which could serve as useful complements to biological experiments. RESULTS: In this paper, we propose a novel graph regularized self-representative matrix factorization (GRSMF) algorithm for synthetic lethal interaction prediction. GRSMF first learns the self-representations from the known SL interactions and further integrates the functional similarities among genes derived from Gene Ontology (GO). It can then effectively predict potential SL interactions by leveraging the information provided by known SL interactions and functional annotations of genes. Extensive experiments on the synthetic lethal interaction data downloaded from SynLethDB database demonstrate the superiority of our GRSMF in predicting potential synthetic lethal interactions, compared with other competing methods. Moreover, case studies of novel interactions are conducted in this paper for further evaluating the effectiveness of GRSMF in synthetic lethal interaction prediction. CONCLUSIONS: In this paper, we demonstrate that by adaptively exploiting the self-representation of original SL interaction data, and utilizing functional similarities among genes to enhance the learning of self-representation matrix, our GRSMF could predict potential SL interactions more accurately than other state-of-the-art SL interaction prediction methods. Jiang Huang, Min Wu 0008, Le Ou-Yang, Zexuan Zhu 0001 |
BMC Bioinform. | 5 |
| 2019 | A hybrid of genetic transform and hyper-rectangle search strategies for evolutionary multi-tasking
Zhengping Liang, Liang Feng 0001, Zexuan Zhu 0001 |
Expert Syst. Appl. | 4 |
| 2019 | Two new reference vector adaptation strategies for many-objective evolutionary algorithms
Zhengping Liang, Weijun Hou, Zexuan Zhu 0001 |
Inf. Sci. | 4 |
| 2019 | Hybrid of memory and prediction strategies for dynamic multiobjective optimization
Zhengping Liang, Shunxiang Zheng, Zexuan Zhu 0001, Shengxiang Yang |
Inf. Sci. | 3 |
| 2019 | Differential evolution algorithm with dichotomy-based parameter space compression
Laizhong Cui, Genghui Li, Zexuan Zhu 0001, Zhong Ming 0001, Zhenkun Wen |
Soft Comput. | 3 |
| 2019 | A Survey on Cooperative Co-Evolutionary AlgorithmsabstractThe first cooperative co-evolutionary algorithm (CCEA) was proposed by Potter and De Jong in 1994 and since then many CCEAs have been proposed and successfully applied to solving various complex optimization problems. In applying CCEAs, the complex optimization problem is decomposed into multiple subproblems, and each subproblem is solved with a separate subpopulation, evolved by an individual evolutionary algorithm (EA). Through cooperative co-evolution of multiple EA subpopulations, a complete problem solution is acquired by assembling the representative members from each subpopulation. The underlying divide-and-conquer and collaboration mechanisms enable CCEAs to tackle complex optimization problems efficiently, and hence CCEAs have been attracting wide attention in the EA community. This paper presents a comprehensive survey of these CCEAs, covering problem decomposition, collaborator selection, individual fitness evaluation, subproblem resource allocation, implementations, benchmark test problems, control parameters, theoretical analyses, and applications. The unsolved challenges and potential directions for their solutions are discussed. Xiaoliang Ma 0001, Xiaodong Li 0001, Qingfu Zhang 0001, Ke Tang 0001, Zhengping Liang, Weixin Xie, Zexuan Zhu 0001 |
IEEE Trans. Evol. Comput. | 7 |
| 2018 | A Preliminary Study of Adaptive Indicator Based Evolutionary Algorithm for Dynamic Multiobjective Optimization via AutoencodingabstractDynamic multi-objective optimization problem (D-MOP) is widely existed in many real-world applications. Over the years, DMOP has attracted many research attentions in the literature. The adaptive indicator-based evolutionary algorithm (IBEA2) is a recently proposed multi-objective evolutionary algorithm (MOEA). It has demonstrated strong search capability on commonly used multi-objective benchmarks over state-of-the-art MOEAs. However, as the adaptation of parameter$k$is based on the selected solutions with maximum hypervolume, this mechanism will be inappropriate if the problem changes over time. The reason is that the solutions with high hypervolume at one particular time instance may not be with high hypervolume at another if the problem changed. Keeping this in mind, inspired by the recent autoencoding evolutionary search, which is able to transfer the past search experiences to improve the evolutionary search on unseen problems, in this paper, we propose to extend the IBEA2 by adapting k with transferred high hypervolume solutions obtained before the dynamic change occurs, for solving DMOP. To evaluate the proposed method, empirical comparisons on the commonly used Farina-Deb-Amato (FDA) DMOP benchmarks, against both the IBEA2 and one recently proposed dynamic MOEA, are presented. Wei Zhou 0001, Liang Feng 0001, Siwei Jiang, Shu Zhang 0003, Yaqing Hou, Yew-Soon Ong, Zexuan Zhu 0001, Kai Liu 0001 |
CEC | 7 |
| 2018 | RepLong: de novo repeat identification using long read sequencing dataabstractMotivation: The identification of repetitive elements is important in genome assembly and phylogenetic analyses. The existing de novo repeat identification methods exploiting the use of short reads are impotent in identifying long repeats. Since long reads are more likely to cover repeat regions completely, using long reads is more favorable for recognizing long repeats. Results: In this study, we propose a novel de novo repeat elements identification method namely RepLong based on PacBio long reads. Given that the reads mapped to the repeat regions are highly overlapped with each other, the identification of repeat elements is equivalent to the discovery of consensus overlaps between reads, which can be further cast into a community detection problem in the network of read overlaps. In RepLong, we first construct a network of read overlaps based on pair-wise alignment of the reads, where each vertex indicates a read and an edge indicates a substantial overlap between the corresponding two reads. Secondly, the communities whose intra connectivity is greater than the inter connectivity are extracted based on network modularity optimization. Finally, representative reads in each community are extracted to form the repeat library. Comparison studies on Drosophila melanogaster and human long read sequencing data with genome-based and short-read-based methods demonstrate the efficiency of RepLong in identifying long repeats. RepLong can handle lower coverage data and serve as a complementary solution to the existing methods to promote the repeat identification performance on long-read sequencing data. Availability and implementation: The software of RepLong is freely available at https://github.com/ruiguo-bio/replong. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Rui Guo 0012, Yan-Ran Li 0001, Shan He 0001, Le Ou-Yang, Zexuan Zhu 0001 |
Bioinform. | 6 |
| 2018 | DroidDet: Effective and robust detection of android malware using static analysis along with rotation forest model
Zhu-Hong You, Zexuan Zhu 0001, Wei-Lei Shi, Xing Chen 0001 |
Neurocomputing | 3 |
| 2018 | Adaptive multiple-elites-guided composite differential evolution algorithm with a shift mechanism
Laizhong Cui, Genghui Li, Zexuan Zhu 0001, Qiuzhen Lin, Ka-Chun Wong, Jianyong Chen, Jian Lu 0002 |
Inf. Sci. | 3 |
| 2018 | A novel differential evolution algorithm with a self-adaptation parameter control method by differential evolution
Laizhong Cui, Genghui Li, Zexuan Zhu 0001, Zhenkun Wen, Jian Lu 0002 |
Soft Comput. | 3 |
| 2018 | On Tchebycheff Decomposition Approaches for Multiobjective Evolutionary OptimizationabstractTchebycheff decomposition represents one of the most widely used decomposition approaches that can convert a multiobjective optimization problem into a set of scalar optimization subproblems. Nevertheless, the geometric properties of the subproblem objective functions in Tchebycheff decomposition have not been explicitly studied. This paper proposes a Tchebycheff decomposition with lp-norm constraint on direction vectors in which the subproblem objective functions are endowed with clear geometric property. Especially, the Tchebycheff decomposition with l2-norm constraint on direction vectors is taken as an example to illustrate its advantage. A new unary R2indicator is also introduced to approximate the hyper-volume metric and justify the efficiency of the proposed Tchebycheff decomposition. A resultant Tchebycheff decomposition-based multiobjective evolutionary algorithm (MOEA) with l2-norm constraint and a new population update strategy is proposed to solve multiobjective optimization problems. The experimental results on both benchmark and real-world multiobjective optimization problems show that the proposed algorithm is capable of obtaining high quality solutions compared with other state-of-the-art MOEAs. Xiaoliang Ma 0001, Qingfu Zhang 0001, Guangdong Tian, Junshan Yang, Zexuan Zhu 0001 |
IEEE Trans. Evol. Comput. | 5 |
| 2018 | Concept Drift Adaptation by Exploiting Historical KnowledgeabstractIncremental learning with concept drift has often been tackled by ensemble methods, where models built in the past can be retrained to attain new models for the current data. Two design questions need to be addressed in developing ensemble methods for incremental learning with concept drift, i.e., which historical (i.e., previously trained) models should be preserved and how to utilize them. A novel ensemble learning method, namely, Diversity and Transfer-based Ensemble Learning (DTEL), is proposed in this paper. Given newly arrived data, DTEL uses each preserved historical model as an initial model and further trains it with the new data via transfer learning. Furthermore, DTEL preserves a diverse set of historical models, rather than a set of historical models that are merely accurate in terms of classification accuracy. Empirical studies on 15 synthetic data streams and 5 real-world data streams (all with concept drifts) demonstrate that DTEL can handle concept drift more effectively than 4 other state-of-the-art methods. Yu Sun 0019, Ke Tang 0001, Zexuan Zhu 0001, Xin Yao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Multi-objective memetic algorithm based on request prediction for dynamic pickup-and-delivery problemsabstractThis paper presents a multi-objective memetic algorithm based on request prediction for route planning in dynamic pickup-and-delivery problems. Historical data are used to predict the occurrence of new dynamic requests, based on which predictive routes are planned and tuned subsequently as the real requests occur. Two objectives namely route length and response time are optimized using multi-objective memetic algorithm that is a synergy of multi-objective genetic algorithm and a locality-sensitive hashing based local search. The proposed algorithm is tested on three benchmark problems and the experimental results demonstrate the efficiency of the algorithm. Yanming Yang, Zexuan Zhu 0001 |
CEC | 3 |
| 2017 | LW-FQZip 2: a parallelized reference-based compression of FASTQ filesabstractBACKGROUND: The rapid progress of high-throughput DNA sequencing techniques has dramatically reduced the costs of whole genome sequencing, which leads to revolutionary advances in gene industry. The explosively increasing volume of raw data outpaces the decreasing disk cost and the storage of huge sequencing data has become a bottleneck of downstream analyses. Data compression is considered as a solution to reduce the dependency on storage. Efficient sequencing data compression methods are highly demanded. RESULTS: In this article, we present a lossless reference-based compression method namely LW-FQZip 2 targeted at FASTQ files. LW-FQZip 2 is improved from LW-FQZip 1 by introducing more efficient coding scheme and parallelism. Particularly, LW-FQZip 2 is equipped with a light-weight mapping model, bitwise prediction by partial matching model, arithmetic coding, and multi-threading parallelism. LW-FQZip 2 is evaluated on both short-read and long-read data generated from various sequencing platforms. The experimental results show that LW-FQZip 2 is able to obtain promising compression ratios at reasonable time and memory space costs. CONCLUSIONS: The competence enables LW-FQZip 2 to serve as a candidate tool for archival or space-sensitive applications of high-throughput DNA sequencing data. LW-FQZip 2 is freely available at http://csse.szu.edu.cn/staff/zhuzx/LWFQZip2 and https://github.com/Zhuzxlab/LW-FQZip2 . Zhi-an Huang, Zhenkun Wen, Qingjin Deng, Zexuan Zhu 0001 |
BMC Bioinform. | 6 |
| 2017 | A novel artificial bee colony algorithm with an adaptive population size for numerical function optimization
Laizhong Cui, Genghui Li, Zexuan Zhu 0001, Qiuzhen Lin, Zhenkun Wen, Ka-Chun Wong, Jianyong Chen |
Inf. Sci. | 3 |
| 2017 | PBMDA: A novel and effective path-based computational model for miRNA-disease association predictionabstractIn the recent few years, an increasing number of studies have shown that microRNAs (miRNAs) play critical roles in many fundamental and important biological processes. As one of pathogenetic factors, the molecular mechanisms underlying human complex diseases still have not been completely understood from the perspective of miRNA. Predicting potential miRNA-disease associations makes important contributions to understanding the pathogenesis of diseases, developing new drugs, and formulating individualized diagnosis and treatment for diverse human complex diseases. Instead of only depending on expensive and time-consuming biological experiments, computational prediction models are effective by predicting potential miRNA-disease associations, prioritizing candidate miRNAs for the investigated diseases, and selecting those miRNAs with higher association probabilities for further experimental validation. In this study, Path-Based MiRNA-Disease Association (PBMDA) prediction model was proposed by integrating known human miRNA-disease associations, miRNA functional similarity, disease semantic similarity, and Gaussian interaction profile kernel similarity for miRNAs and diseases. This model constructed a heterogeneous graph consisting of three interlinked sub-graphs and further adopted depth-first search algorithm to infer potential miRNA-disease associations. As a result, PBMDA achieved reliable performance in the frameworks of both local and global LOOCV (AUCs of 0.8341 and 0.9169, respectively) and 5-fold cross validation (average AUC of 0.9172). In the cases studies of three important human diseases, 88% (Esophageal Neoplasms), 88% (Kidney Neoplasms) and 90% (Colon Neoplasms) of top-50 predicted miRNAs have been manually confirmed by previous experimental reports from literatures. Through the comparison performance between PBMDA and other previous models in case studies, the reliable performance also demonstrates that PBMDA could serve as a powerful computational tool to accelerate the identification of disease-miRNA associations. Zhu-Hong You, Zhi-an Huang, Zexuan Zhu 0001, Guiying Yan, Zhengwei Li 0001, Zhenkun Wen, Xing Chen 0001 |
PLoS Comput. Biol. | 3 |
| 2017 | A Similarity-Based Multiobjective Evolutionary Algorithm for Deployment Optimization of Near Space Communication SystemabstractThe deployment of the airships plays a key role in maximizing the performance of the near space communication system. The main problem is how to strike a balance between the conflicting network speed and coverage for complex user distribution. In this paper, we propose a multiobjective deployment optimization model considering path loss, user demand, and inner structure. Under the framework of the multiobjective evolutionary algorithm (MOEA) based on decomposition (MOEA/D), we propose a similarity-based MOEA to optimize this problem. The proposed algorithm is motivated by the population's perception on the decision variable space. The proposed algorithm perceives the decision variable space by deploying airships to latent regions. The perceptions of different solutions are related by the similarity between their deployments and utilized differently by crossover and mutation. The proposed algorithm is tested on five designed problems compared with MOEA/D with the other popular reproduction operators. We also test the proposed scheme integrated with another two popular algorithms. The experimental results show that the similarity-based MOEA/D outperforms the other algorithms significantly in detecting hotspots, tracking multiple hotspots and safely deploying airships for most cases. The proposed scheme also works well with the other algorithms. Maoguo Gong, Zhao Wang 0011, Zexuan Zhu 0001, Licheng Jiao |
IEEE Trans. Evol. Comput. | 3 |
| 2016 | A comparative study on decomposition-based multi-objective evolutionary algorithms for many-objective optimizationabstractMany-objective optimization problems pose challenges to the Pareto-based multi-objective optimization algorithms. Recent studies have suggested that decomposition is a promising method to improve the performance of multi-objective evolutionary algorithms on many-objective optimization problem. Various methods based on decomposition have been developed to solve many-objective problems in recent years. However, the existing experimental comparative studies are usually limited to only a few methods based on decomposition. This paper offers a systematic comparison of seven representative decomposition-based approaches tested on two groups of widely used problems. The experimental results have demonstrated that none of the compared algorithms has a clear advantage over the others, although different algorithms are competitive on different test problems. Therefore, a careful selection of algorithms is necessary in handling a many-objective problem in hand. Xiaoliang Ma 0001, Junshan Yang, Nuosi Wu, Zhen Ji, Zexuan Zhu 0001 |
CEC | 5 |
| 2016 | A more efficient method for domain repeat detection in WD-40 proteinsabstractStructural biology is a branch of molecular biology and biochemistry, aiming to understand the interaction of molecules like proteins by observing their structures. Crystallization is one of the most widely used methods to identify protein structure, yet it is laborious and time-consuming. Researchers are seeking assistance from computers. This paper implements an improved WDSP program for recognizing and predicting secondary structure of WD40 repeat proteins, which is a large protein family in eukaryotes. The original WDSP works well on predicting WD40 protein structures but it also suffers from low computational efficiency. We propose a more computationally efficient WDSP namely FWDSP by imposing clustering and a specific local searching to the original WDSP. Experiment results on three datasets of WD40 proteins demonstrate the effectiveness and efficiency of FWDSP. Nuosi Wu, Zexuan Zhu 0001, Zhen Ji |
CEC | 3 |
| 2016 | Multi-objective memetic algorithm for solving pickup and delivery problem with dynamic customer requests and traffic informationabstractThis paper formulates one-to-many-to-one pickup and delivery problems with dynamic customer requests and traffic information. A multi-objective memetic algorithm namely prioLSH-MOMA is proposed to solve the problems. The new algorithm is characterized with a priority and locality-sensitive hashing based local search. prioLSH-MOMA is designed to find an optimal route of a dynamic pickup and delivery problem in terms of route length and workload. Particularly, a re-planning strategy is introduced to handle the dynamic information. Priority and locality-sensitive hashing based local search is applied to fine-tune the candidate routes during the evolution process. prioLSH-MOMA is evaluated with two dynamic pickup and delivery problems simulated on real-world maps and the results demonstrate the efficiency of the proposed algorithm. Yanming Yang, Xiaoliang Ma 0001, Zexuan Zhu 0001 |
CEC | 5 |
| 2016 | Metabolomics biomarker discovery using multimodal memetic algorithm and multivariate mutual information based feature selectionabstractMetabolomics data has the nature of small sample number, high dimensional, and noisy, which poses great challenges on its analysis. In this paper we propose a novel filter feature selection algorithm, namely MMAFS, for the metabolomics biomarker discovery. The MMAFS utilizes a metaheuristics chain based multimodal memetic algorithm to effectively select both local and global optimal feature subgroups that potentially contain biological meanings. A nearest-neighbor graphic based multivariate mutual information estimation is used to calculate fitness values under the max-dependency criterion. Finally, we introduce a semi-wrapper classification to improve the prediction accuracy. The MMAFS is applied on three real-world metabolomics spectrum data sets. Experimental results on 10 runs of 10-fold external cross validation show that the proposed algorithm outperforms other representative feature selection methods. Particularly, some biomarkers found by MMAFS have been proved by previous researches. Zhen Ji, Zexuan Zhu 0001, Shan He 0001 |
CEC | 3 |
| 2016 | A novel adaptive hybrid crossover operator for multiobjective evolutionary algorithm
Qingling Zhu, Qiuzhen Lin, Zhihua Du, Zhengping Liang, Wenjun Wang 0003, Zexuan Zhu 0001, Jianyong Chen, Peizhi Huang, Zhong Ming 0001 |
Inf. Sci. | 6 |
| 2016 | A multi-objective memetic algorithm based on locality-sensitive hashing for one-to-many-to-one dynamic pickup-and-delivery problem
Zexuan Zhu 0001, Shan He 0001, Zhen Ji |
Inf. Sci. | 1 |
| 2016 | Soft computing in remote sensing image processing
Yanfei Zhong, Zexuan Zhu 0001, Yew-Soon Ong |
Soft Comput. | 2 |
| 2016 | Cooperative Co-Evolutionary Module Identification With Application to Cancer Disease Module DiscoveryabstractModule identification or community detection in complex networks has become increasingly important in many scientific fields because it provides insight into the relationship and interaction between network function and topology. In recent years, module identification algorithms based on stochastic optimization algorithms such as evolutionary algorithms have been demonstrated to be superior to other algorithms on small- to medium-scale networks. However, the scalability and resolution limit (RL) problems of these module identification algorithms have not been fully addressed, which impeded their application to real-world networks. This paper proposes a novel module identification algorithm called cooperative co-evolutionary module identification to address these two problems. The proposed algorithm employs a cooperative co-evolutionary framework to handle large-scale networks. We also incorporate a recursive partitioning scheme into the algorithm to effectively address the RL problem. The performance of our algorithm is evaluated on 12 benchmark complex networks. As a medical application, we apply our algorithm to identify disease modules that differentiate low- and high-grade glioma tumors to gain insights into the molecular mechanisms that underpin the progression of glioma. Experimental results show that the proposed algorithm has a very competitive performance compared with other state-of-the-art module identification algorithms. Shan He 0001, Guanbo Jia, Zexuan Zhu 0001, Dan A. Tennant, Ke Tang 0001, Jing Liu 0006, Mirco Musolesi, John K. Heath, Xin Yao 0001 |
IEEE Trans. Evol. Comput. | 3 |
| 2015 | High-throughput DNA sequence data compressionabstractThe exponential growth of high-throughput DNA sequence data has posed great challenges to genomic data storage, retrieval and transmission. Compression is a critical tool to address these challenges, where many methods have been developed to reduce the storage size of the genomes and sequencing data (reads, quality scores and metadata). However, genomic data are being generated faster than they could be meaningfully analyzed, leaving a large scope for developing novel compression algorithms that could directly facilitate data analysis beyond data transfer and storage. In this article, we categorize and provide a comprehensive review of the existing compression methods specialized for genomic data and present experimental results on compression ratio, memory usage, time for compression and decompression. We further present the remaining challenges and potential directions for future research. Zexuan Zhu 0001, Yongpeng Zhang, Zhen Ji, Shan He 0001, Xiao Yang 0019 |
Briefings Bioinform. | 1 |
| 2015 | CompMap: a reference-based compression program to speed up read mapping to related reference sequencesabstractSUMMARY: Exhaustive mapping of next-generation sequencing data to a set of relevant reference sequences becomes an important task in pathogen discovery and metagenomic classification. However, the runtime and memory usage increase as the number of reference sequences and the repeat content among these sequences increase. In many applications, read mapping time dominates the entire application. We developed CompMap, a reference-based compression program, to speed up this process. CompMap enables the generation of a non-redundant representative sequence for the input sequences. We have demonstrated that reads can be mapped to this representative sequence with a much reduced time and memory usage, and the mapping to the original reference sequences can be recovered with high accuracy. AVAILABILITY AND IMPLEMENTATION: CompMap is implemented in C and freely available at http://csse.szu.edu.cn/staff/zhuzx/CompMap/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zexuan Zhu 0001, Yongpeng Zhang, Xiao Yang 0019 |
Bioinform. | 1 |
| 2015 | Light-weight reference-based compression of FASTQ dataabstractBACKGROUND: The exponential growth of next generation sequencing (NGS) data has posed big challenges to data storage, management and archive. Data compression is one of the effective solutions, where reference-based compression strategies can typically achieve superior compression ratios compared to the ones not relying on any reference. RESULTS: This paper presents a lossless light-weight reference-based compression algorithm namely LW-FQZip to compress FASTQ data. The three components of any given input, i.e., metadata, short reads and quality score strings, are first parsed into three data streams in which the redundancy information are identified and eliminated independently. Particularly, well-designed incremental and run-length-limited encoding schemes are utilized to compress the metadata and quality score streams, respectively. To handle the short reads, LW-FQZip uses a novel light-weight mapping model to fast map them against external reference sequence(s) and produce concise alignment results for storage. The three processed data streams are then packed together with some general purpose compression algorithms like LZMA. LW-FQZip was evaluated on eight real-world NGS data sets and achieved compression ratios in the range of 0.111-0.201. This is comparable or superior to other state-of-the-art lossless NGS data compression algorithms. CONCLUSIONS: LW-FQZip is a program that enables efficient lossless FASTQ data compression. It contributes to the state of art applications for NGS data storage and transmission. LW-FQZip is freely available online at: http://csse.szu.edu.cn/staff/zhuzx/LWFQZip. Yongpeng Zhang, Xiao Yang 0019, Shan He 0001, Zexuan Zhu 0001 |
BMC Bioinform. | 6 |
| 2015 | Robust twin boosting for feature selection from high-dimensional omics data with label noise
Shan He 0001, Huanhuan Chen 0001, Zexuan Zhu 0001, Douglas G. Ward, Helen J. Cooper, Mark R. Viant, John K. Heath, Xin Yao 0001 |
Inf. Sci. | 3 |
| 2015 | Three-dimensional Gabor feature extraction for hyperspectral imagery classification using a memetic framework
Zexuan Zhu 0001, Sen Jia 0001, Shan He 0001, Zhen Ji, LinLin Shen |
Inf. Sci. | 1 |
| 2014 | Feature extraction based on trimmed complex network representation for metabolomic data classificationabstractOver the last few decades, metabolomics has been widely used to reveal the linkages between metabolite signal levels and physiological states. Metabolomic data are naturally high dimensional and noisy, which poses computational challenges for data analysis. In this study, a novel feature extraction method based on trimmed complex network representation is proposed for metabolomic data classification. Particularly, the proposed method begins with feature selection on the original data, and then a complex network of the selected features is constructed to represent each data sample. Afterward, the network edges are trimmed and a few topological network metrics are extracted as new features for the classification of the samples. The experimental results on a real-world metabolomic data of clinical liver transplantation demonstrate the efficiency of the proposed feature extraction method. Zexuan Zhu 0001, Zhen Ji |
IEEE Congress on Evolutionary Computation | 2 |
| 2014 | Locality-sensitive hashing based multiobjective memetic algorithm for dynamic pickup and delivery problemsabstractThis paper proposes a locality-sensitive hashing based multiobjective memetic algorithm namely LSH-MOMA for solving pickup and delivery problems with dynamic requests (DPDPs for short). Particularly, LSH-MOMA is designed to find the solution route of a DPDP by optimizing objectives namely workload and route length in an evolutionary manner. In each generation of LSH-MOMA, locality-sensitive hashing based rectification and local search are imposed to repair and refine the individual candidate routes. LSH-MOMA is evaluated on three simulated DPDPs of different scales and the experimental results demonstrate the efficiency of the method. Fangxiao Wang 0001, Zexuan Zhu 0001 |
IEEE Congress on Evolutionary Computation | 3 |
| 2014 | A growing partitional clustering based on particle swarm optimizationabstractThis paper proposes a growing partitional clustering method based on particle swarm optimization (PSO) namely PSOGC for handling data with non-spherical or non-linearly separable distribution. Particularly, PSOGC uses PSO to optimize the cluster centers. In each iteration of PSO, the particles encoding candidate cluster centers are evolved according to their social and personal knowledge. Given the candidate cluster centers, a growing strategy increasingly absorbs nearby data samples into the corresponding cluster based on k-nearest neighbor graph. The fitness of each particle is evaluated in terms of intra-cluster connectivity and inter-cluster disconnectivity of the resultant clustering. The combination of PSO and growing strategy ensures the stability of global search and the robustness of partition on data of different non-spherical shapes. Experimental results on six synthetic and three UCI real-world data sets demonstrate the efficiency of PSOGC. Nuosi Wu, Zexuan Zhu 0001, Zhen Ji |
IEEE Congress on Evolutionary Computation | 2 |
| 2014 | Using Chou's amphiphilic Pseudo-Amino Acid Composition and Extreme Learning Machine for prediction of Protein-protein interactionsabstractProtein-protein interactions (PPIs) play crucial roles in the execution of various cellular processes. Almost every cellular process relies on transient or permanent physical bindings of proteins. Unfortunately, the experimental methods for identifying PPIs are both time-consuming and expensive. Therefore, it is important to develop computational approaches for predicting PPIs. In this study, a novel approach is presented to predict PPIs using only the information of protein sequences. This method is developed based on learning algorithm-Extreme Learning Machine (ELM) combined with the concept of Chous Pseudo-Amino Acid Composition (PseAAC) composition. PseAAC is a combination of a set of discrete sequence correlation factors and the 20 components of the conventional amino acid composition, so this method can observe a remarkable improvement in prediction quality. ELM classifier is selected as prediction engine, which is a kind of accurate and fast-learning innovative classification method based on the random generation of the input-to-hidden-units weights followed by the resolution of the linear equations to obtain the hidden-to-output weights. When performed on the PPIs data of Saccharomyces cerevisiae, the proposed method achieved 79.66% prediction accuracy with 79.16% sensitivity at the precision of 79.96%. Extensive experiments are performed to compare our method with state-of-the-art techniques Support Vector Machine (SVM). Achieved results show that the proposed approach is very promising for predicting PPIs, and it can be a helpful supplement for PPIs prediction. Qiao-Ying Huang, Zhu-Hong You, Shuai Li 0002, Zexuan Zhu 0001 |
IJCNN | 4 |
| 2014 | HAMMER: automated operation of mass frontier to construct in silico mass spectral fragmentation librariesabstractSUMMARY: Experimental MS(n) mass spectral libraries currently do not adequately cover chemical space. This limits the robust annotation of metabolites in metabolomics studies of complex biological samples. In silico fragmentation libraries would improve the identification of compounds from experimental multistage fragmentation data when experimental reference data are unavailable. Here, we present a freely available software package to automatically control Mass Frontier software to construct in silico mass spectral libraries and to perform spectral matching. Based on two case studies, we have demonstrated that high-throughput automation of Mass Frontier allows researchers to generate in silico mass spectral libraries in an automated and high-throughput fashion with little or no human intervention required. AVAILABILITY AND IMPLEMENTATION: Documentation, examples, results and source code are available at http://www.biosciences-labs.bham.ac.uk/viant/hammer/. Ralf J. M. Weber, James William Allwood, Robert Mistrik, Zexuan Zhu 0001, Zhen Ji, Siping Chen, Warwick B. Dunn, Shan He 0001, Mark R. Viant |
Bioinform. | 5 |
| 2014 | Compression of next-generation sequencing quality scores using memetic algorithmabstractBACKGROUND: The exponential growth of next-generation sequencing (NGS) derived DNA data poses great challenges to data storage and transmission. Although many compression algorithms have been proposed for DNA reads in NGS data, few methods are designed specifically to handle the quality scores. RESULTS: In this paper we present a memetic algorithm (MA) based NGS quality score data compressor, namely MMQSC. The algorithm extracts raw quality score sequences from FASTQ formatted files, and designs compression codebook using MA based multimodal optimization. The input data is then compressed in a substitutional manner. Experimental results on five representative NGS data sets show that MMQSC obtains higher compression ratio than the other state-of-the-art methods. Particularly, MMQSC is a lossless reference-free compression algorithm, yet obtains an average compression ratio of 22.82% on the experimental data sets. CONCLUSIONS: The proposed MMQSC compresses NGS quality score data effectively. It can be utilized to improve the overall compression ratio on FASTQ formatted files. Zhen Ji, Zexuan Zhu 0001, Shan He 0001 |
BMC Bioinform. | 3 |
| 2013 | Minimal-redundancy-maximal-relevance feature selection using different relevance measures for omics data classificationabstractOmics refers to a field of study in biology such as genomics, proteomics, and metabolomics. Investigating fundamental biological problems based on omics data would increase our understanding of bio-systems as a whole. However, omics data is characterized with high-dimensionality and unbalance between features and samples, which poses big challenges for classical statistical analysis and machine learning methods. This paper studies a minimal-redundancy-maximal-relevance (MRMR) feature selection for omics data classification using three different relevance evaluation measures including mutual information (MI), correlation coefficient (CC), and maximal information coefficient (MIC). A linear forward search method is used to search the optimal feature subset. The experimental results on five real-world omics datasets indicate that MRMR feature selection with CC is more robust to obtain better (or competitive) classification accuracy than the other two measures. Junshan Yang, Zexuan Zhu 0001, Shan He 0001, Zhen Ji |
CIBCB | 2 |
| 2013 | A SVM-Based System for Predicting Protein-Protein Interactions Using a Novel Representation of Protein Sequences
Zhu-Hong You, Zhong Ming 0001, Suping Deng, Zexuan Zhu 0001 |
ICIC (1) | 5 |
| 2013 | Global Path Planning of Wheeled Robots Using a Multi-Objective Memetic Algorithm
Fangxiao Wang 0001, Zexuan Zhu 0001 |
IDEAL | 2 |
| 2013 | Discriminative Gabor Feature Selection for Hyperspectral Image ClassificationabstractThree-dimensional Gabor wavelets have recently been successfully applied for hyperspectral image classification due to their ability to extract joint spatial and spectrum information. However, the dimension of the extracted Gabor feature is incredibly huge. In this letter, we propose a symmetrical-uncertainty-based and Markov-blanket-based approach to select informative and nonredundant Gabor features for hyperspectral image classification. The extracted Gabor features with large dimension are first ranked by their information contained for classification and then added one by one after investigating the redundancy with already selected features. The proposed approach was fully tested on the widely used Indian Pine site data. The results show that the selected features are much more efficient and can achieve similar performance with previous approach using only hundreds of features. LinLin Shen, Zexuan Zhu 0001, Sen Jia 0001, Jiasong Zhu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2013 | Self-configuration single particle optimizer for DNA sequence compression
Zhen Ji, Zexuan Zhu 0001, Siping Chen |
Soft Comput. | 3 |
| 2012 | A memory binary particle swarm optimizationabstractThis paper proposes a memory binary particle swarm optimization algorithm (MBPSO) based on a new updating strategy. Unlike the traditional binary PSO, which updates the binary bits of a particle ignoring their previous status, MBPSO memorizes the bit status and updates them according to a new defined velocity. As such, precious historical information could be retained to guide the search. The velocity vector of MBPSO is designed as a probability for deciding whether the particle bits change or not. The proposed algorithm is tested on four discrete benchmark functions. The experimental results reported over 100 runs show that MBPSO is capable of obtaining encouraging performance in discrete optimization problems. Zhen Ji, Tao Tian, Shan He 0001, Zexuan Zhu 0001 |
IEEE Congress on Evolutionary Computation | 4 |
| 2012 | A crown jewel defense strategy based particle swarm optimizationabstractParticle swarm optimization (PSO) is a metaheuristic algorithm that is easy to implement and performs well on various optimization problems. However, PSO is sensitive to initialization due to its rapid convergence which leads to the lack of population diversity and premature convergence. To solve this problem, a jumping-out strategy named crown jewel defense (CJD) is introduced in this paper. CJD is used to relocate the global best position and reinitializes all particles' personal best position when the swarm is trapped in local optima. Taking the advantage of CJD strategy, the swarm can jump out of the local optimal region without being dragged back and the performance of PSO becomes more robust to the initialization. Experimental results on benchmark functions show that the CJD-based PSO are comparable to or better than the other representative state-of-the-art PSO. Zhen Ji, Shan He 0001, Zexuan Zhu 0001 |
IEEE Congress on Evolutionary Computation | 4 |
| 2012 | Survival analysis of gene expression data using PSO based radial basis function networksabstractGene expression data combined with clinical data has emerged as an important source for survival analysis. However, gene expression data is characterized with thousands of features/genes but only tens or hundreds of observations. The high-dimensionality and unbalance between features and samples pose big challenges for the classical survival analysis methods. This paper proposes a particle swarm optimization based radial basis function networks (PSO-RBFN) for the survival analysis on gene expression data. Particularly, PSO-RBFN applies a principle component analysis for dimensionality reduction and optimizes the RBF network using PSO. The experimental results on three gene expression datasets indicate that PSO-RBFN is able to improve the predict accuracy compared to the other classical survival analysis methods. Wenmin Liu, Zhen Ji, Shan He 0001, Zexuan Zhu 0001 |
IEEE Congress on Evolutionary Computation | 4 |
| 2012 | Memetic clustering based on particle swarm optimizer and K-meansabstractThis paper proposes an efficient memetic clustering algorithm (MCA) for clustering based on particle swarm optimizer (PSO) and K-means. Particularly, PSO is used as a global search to allow fast exploration of the candidate cluster centers. PSO has strong ability to find high quality solutions within tractable time, but it suffers from slow-down convergence as the swarm approaching optima. K-means, achieving fast convergence to optimum solutions, is utilized as local search to fine-tune the solutions of PSO in the framework of memetic algorithm. The performance of MCA is evaluated on four synthetic datasets and three high-dimensional gene expression datasets. Comparison study to K-means, PSO, and PSO-KM (jointed PSO and K-means) indicates that MCA is capable of identifying cluster centers more precisely and robustly than the other counterpart algorithms by taking advantage of both PSO and K-means. Zexuan Zhu 0001, Wenmin Liu, Shan He 0001, Zhen Ji |
IEEE Congress on Evolutionary Computation | 1 |
| 2011 | DNA Sequence Compression Using Adaptive Particle Swarm Optimization-Based Memetic AlgorithmabstractWith the rapid development of high-throughput DNA sequencing technologies, the amount of DNA sequence data is accumulating exponentially. The huge influx of data creates new challenges for storage and transmission. This paper proposes a novel adaptive particle swarm optimization-based memetic algorithm (POMA) for DNA sequence compression. POMA is a synergy of comprehensive learning particle swarm optimization (CLPSO) and an adaptive intelligent single particle optimizer (AdpISPO)-based local search. It takes advantage of both CLPSO and AdpISPO to optimize the design of approximate repeat vector (ARV) codebook for DNA sequence compression. ARV is first introduced in this paper to represent the repeated fragments across multiple sequences in direct, mirror, pairing, and inverted patterns. In POMA, candidate ARV codebooks are encoded as particles and the optimal solution, which covers the most approximate repeated fragments with the fewest base variations, is identified through the exploration and exploitation of POMA. In each iteration of POMA, the leader particles in the swarm are selected based on weighted fitness values and each leader particle is fine-tuned with an AdpISPO-based local search, so that the convergence of the search in local region is accelerated. A detailed comparison study between POMA and the counterpart algorithms is performed on 29 (23 basic and 6 composite) benchmark functions and 11 real DNA sequences. POMA is observed to obtain better or competitive performance with a limited number of function evaluations. POMA also attains lower bits-per-base than other state-of-the-art DNA-specific algorithms on DNA sequence data. The experimental results suggest that the cooperation of CLPSO and AdpISPO in the framework of memetic algorithm is capable of searching the ARV codebook space efficiently. Zexuan Zhu 0001, Zhen Ji, Yuhui Shi 0001 |
IEEE Trans. Evol. Comput. | 1 |
| 2010 | Affinity propagation based memetic band selection on hyperspectral imagery datasetsabstractThis paper presents a novel affinity propagation (AP) based memetic band selection method (APMA) for hyperspectral imagery classification. The method incorporates AP based local search and genetic algorithm (GA) based global search to take advantage of both. Particularly, the AP based local search fine-tunes the GA individuals by adding relevant bands and eliminating irrelevant/redundant bands. A comparison study to the filters methods (including ReliefF, AP based method, and FCBF) and the counterpart wrapper GA feature selection on two hyperspectral imagery datasets demonstrates that APMA is capable of attaining competitive or better classification accuracy with fewer selected bands, which suggests APMA searches the band subset space more efficiently and identify better band subsets. Zexuan Zhu 0001, Sen Jia 0001, Zhen Ji |
IEEE Congress on Evolutionary Computation | 1 |
| 2010 | Identification of Full and Partial Class Relevant GenesabstractMulticlass cancer classification on microarray data has provided the feasibility of cancer diagnosis across all of the common malignancies in parallel. Using multiclass cancer feature selection approaches, it is now possible to identify genes relevant to a set of cancer types. However, besides identifying the relevant genes for the set of all cancer types, it is deemed to be more informative to biologists if the relevance of each gene to specific cancer or subset of cancer types could be revealed or pinpointed. In this paper, we introduce two new definitions of multiclass relevancy features, i.e., full class relevant (FCR) and partial class relevant (PCR) features. Particularly, FCR denotes genes that serve as candidate biomarkers for discriminating all cancer types. PCR, on the other hand, are genes that distinguish subsets of cancer types. Subsequently, a Markov blanket embedded memetic algorithm is proposed for the simultaneous identification of both FCR and PCR genes. Results obtained on commonly used synthetic and real-world microarray data sets show that the proposed approach converges to valid FCR and PCR genes that would assist biologists in their research work. The identification of both FCR and PCR genes is found to generate improvement in classification accuracy on many microarray data sets. Further comparison study to existing state-of-the-art feature selection algorithms also reveals the effectiveness and efficiency of the proposed approach. Zexuan Zhu 0001, Yew-Soon Ong, Jacek M. Zurada |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2008 | A fast pruned-extreme learning machine for classification problem
Hai-Jun Rong, Yew-Soon Ong, Ah-Hwee Tan, Zexuan Zhu 0001 |
Neurocomputing | 4 |
| 2007 | Memetic Algorithms for Feature Selection on Microarray Data
Zexuan Zhu 0001, Yew-Soon Ong |
ISNN (1) | 1 |
| 2007 | Markov blanket-embedded genetic algorithm for gene selection
Zexuan Zhu 0001, Yew-Soon Ong, Manoranjan Dash |
Pattern Recognit. | 1 |
| 2007 | Wrapper-Filter Feature Selection Algorithm Using a Memetic FrameworkabstractThis correspondence presents a novel hybrid wrapper and filter feature selection algorithm for a classification problem using a memetic framework. It incorporates a filter ranking method in the traditional genetic algorithm to improve classification performance and accelerate the search in identifying the core feature subsets. Particularly, the method adds or deletes a feature from a candidate feature subset based on the univariate feature ranking information. This empirical study on commonly used data sets from the University of California, Irvine repository and microarray data sets shows that the proposed method outperforms existing methods in terms of classification accuracy, number of selected features, and computational efficiency. Furthermore, we investigate several major issues of memetic algorithm (MA) to identify a good balance between local search and genetic search so as to maximize search quality and efficiency in the hybrid filter and wrapper MA. Zexuan Zhu 0001, Yew-Soon Ong, Manoranjan Dash |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | Whole-Genome Functional Classification of Genes by Latent Semantic Analysis on Microarray Data
See-Kiong Ng, Zexuan Zhu 0001, Yew-Soon Ong |
APBC | 2 |