Xiucai Ye

dblp:99/10593 · DBLP profile ↗
← Back
69ranked-venue papers
9as first author
53since 2021 · last 2026
0000-0002-5547-3919ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 27 · 26 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 10 since 2021Computer networks · 10 · 3 first-author · 7 since 2021Systems, architecture and hardware · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Robust graph structure learning to improve multi-omics cancer subtype classification
abstract
BACKGROUND: Classifying cancer patients into consistent subtypes at the multi-omics level remains a significant challenge in advancing precision medicine. Nevertheless, a key problem in integrating multi-omics data lies in concurrently addressing intra-omics and inter-omics information, along with sample networks. RESULTS: In this study, we introduce the Feature and Graph Structure-Learning Integrated Graph Convolutional Network (FaGGCN), which combines feature learning and graph structure learning for multi-omics cancer subtyping. The model employs convolutional autoencoders to learn information-rich latent features, and patient survival information is further leveraged to select key features that are significantly associated with survival outcomes. The graph autoencoder fuses the key features with inter-omics similarity fusion matrices, enabling the model to learn a comprehensive sample network. Finally, the graph convolutional network integrates the key features while incorporating the sample network to precisely classify patients. Additionally, survival analysis, sensitivity analysis, and differential gene expression analysis highlight the interpretability of the FaGGCN model, as well as its ability to identify biomarkers suitable for clinical research. CONCLUSIONS: Experimental results show that our model achieves competitive performance across eight cancer datasets spanning four omics modalities, with generally improved classification performance and exploratory survival prediction results.
Mengke Guo, Xiucai Ye, Tetsuya Sakurai
BMC Bioinform.2
2026 BlackHole-DTP: a distributed trajectory privacy protection strategy with deep trajectory training and data blackhole
abstract
Abstract Trajectory data offers significant potential for personalized services and behavioral analysis, but also raises substantial privacy concerns. However, centralized privacy protection strategies are susceptible to single-point failures, often fail to accommodate user-specific privacy needs, and can lack precision in local privacy protection. To address these concerns, this work utilizes the decentralized philosophy of Web 3.0. A novel strategy is proposed, the Blackhole Model-based Distributed Trajectory Privacy-Preserving Strategy (BlackHole-DTP). First, the Distributed Deep Learning Trajectory Training Algorithm with Multi-head Attention and Variational Adversarial Autoencoders is introduced to improve the simulation of temporal and semantic information in trajectory data, thereby enhancing data utility and accuracy. Additionally, a Local Differential Privacy Algorithm based on the Data Blackhole is designed that dynamically adjusts privacy protection levels according to user requirements. This algorithm incorporates principles inspired by general relativity to determine perturbation values. Finally, experimental results demonstrate that BlackHole-DTP provides superior privacy protection, achieving significantly lower success rates in reconstruction and re-identification attacks compared to baseline models. Specifically, under high-precision requirements, the attack success rate is reduced by $\sim $30%.
Hong-ming Hou, Jing Zhang 0040, Zhenhan Huang, Meirun Zhang, Xiucai Ye
Comput. J.5
2026 PPTRecS-FL: Privacy-preserving task recommendation strategy based on federated learning in mobile crowdsensing
Jing Zhang 0040, Xiangxuan Zhong, Zhenhan Huang, Meirun Zhang, Xiucai Ye
Comput. Networks5
2026 High-performance computing enhanced task recommendation strategy based on mobile prediction in mobile crowdsensing
Jing Zhang 0040, Xiangxuan Zhong, Zhenhan Huang, Li Xu 0002, Xiucai Ye
Eng. Appl. Artif. Intell.5
2026 TPGNN-FedGPR: triple level privacy-preserving graph neural network for federated geographic POI recommendation
Wenlong Shi, Jing Zhang 0040, Youqin Chen, Xiucai Ye, Hao Liao
Frontiers Comput. Sci.4
2026 CADiS: Causality-Driven Transformer for Anomaly Detection and Root Cause Diagnosis in Industrial Internet of Things
abstract
This paper proposes CADiS, a causality-driven anomaly detection framework, to address the challenges of root cause identification in high-dimensional Industrial Internet of Things (IIoT) multivariate time series. The essential difference between CADiS and existing correlation-driven deep models lies in its core innovation: it fundamentally redefines anomalies as the structural decay of an underlying causal mechanism, rather than merely capturing symptomatic deviations or spurious correlations. Specifically, the framework first learns a directed and lag-aware causal prior from normal data, compiling it into a structured attention mask to constrain information flow. Then, a Causal-Phase Decomposition (CPD) technique treats each time window as a micro-experiment, comparing an ante-phase with a post-phase to explicitly capture the dynamics of causal attenuation. Inference relies on a unified Causal-Change Score (CCS), which quantifies the degradation of causal association strength, directly revealing the breakdown of the system’s causal logic. Furthermore, the decomposed causal change matrix allows for fine-grained and auditable root cause diagnosis. Extensive experiments on real-world industrial datasets demonstrate that CADiS significantly outperforms strong baselines, achieving theVROCof 89.93% andVPRof 76.93% on SWaT; TheAPRof 18.04% andVROCof 78.38% on SMD, thereby validating its robustness and diagnostic precision.
Zuanyang Zeng, Xiaoding Wang 0001, Li Xu 0002, Xiucai Ye, Jia Hu 0001, Farooque Hassan Kumbhar, Kapal Dev
IEEE Internet Things J.4
2026 Multisource Graphs and Dual KAN-Transformers for Next POI Recommendation
Jing Zhang 0040, Zhenhan Huang, Tian Wang 0001, Qihan Huang, Li Xu 0002, Xiucai Ye
IEEE Internet Things J.6
2026 SSDP-DCA: Enhancing Data Collaboration Analysis With Privacy Amplification Techniques in Cloud Environments
abstract
The rapid growth of cloud computing has enabled collaborative data analysis across distributed systems, yet ensuring privacy under stringent regulations remains a critical challenge. Data Collaboration Analysis (DCA) facilitates efficient analysis in such environments but struggles to balance robust privacy with model accuracy under strict constraints . Traditional Differential Privacy (DP) methods often degrade performance due to excessive noise (e.g.,$\epsilon = 2$). We propose SSDP-DCA, a novel framework that enhances DCA by integrating DP with privacy amplification techniques, leveraging the Shuffle Model for anonymization andpackage Subsampling (applied post-shuffle)for privacy amplification. Experimental results demonstrate SSDP-DCA's superiority over baselines, achieving high utility under a fixed global privacy budget, thus offering a robust solution for secure federated learning and healthcare applications.
Tingkai Sun, Xiucai Ye, Akira Imakura, Kazumasa Omote, Tetsuya Sakurai
IEEE Trans. Cloud Comput.2
2025 DiCoGRN: Inference of Colorectal Cancer Subtypespecific Gene Regulatory Networks Using Dual-View Contrastive Transformer
abstract
Colorectal cancer (CRC) is a highly heterogeneous disease with distinct molecular subtypes, exhibiting unique transcriptional programs and clinical behaviors. The development of subtype-specific gene regulatory network (GRN) prediction is critical for uncovering the dysregulated transcriptional circuits driving CRC progression, metastasis and therapy resistance. However, the existing GRN prediction methods have limitations in addressing data sparsity and generalization across CRC subtypes, failing to effectively capture subtype-specific regulatory relationships and struggling to handle uncharacterized regulatory factors. In this study, we propose a novel computational framework, Dual-view Contrastive Gene Regulatory Network (DiCoGRN), which integrate structural and semantic information through dual-view contrastive Transformer to infer CRC subtype-specific GRNs. DiCoGRN identifies different cell subpopulations (CRC subtypes) and applies Graph Contrastive Learning (GraphCL) to learn structural embeddings for each subtype. Then, DiCoGRN retrieves the highly variable genes and driver genes of CRC subtypes from the NCBI database, and uses a large language model (LLM) to extract the semantic embeddings of genes. These two-view information are fused through cross-attention and gating interaction mechanisms, ultimately inferring subtype-aware regulatory connections. Experimental results show that the proposed DiCoGRN outperforms existing methods in recovering regulatory interactions and generalization across 12 subtypes. The performed vitro wet experiments illustrate that GATA3 from our predicted regulatory relationships may be a potential CRC driver gene. Our method not only advances CRC subtype-specific GRN inference but also helps to advance toward personalized CRC treatment. Codes and data are available at https://github.com/Fraid-H/DiCoGRN.
Meng Huang 0003, Huijin Hu, Ming Li 0026, Jian Zhang 0082, Heng Zhang 0001, Xiucai Ye
BIBM7
2025 GSToxi: Gated Cross-Modal Modeling With Graph-Sequence Encoders for Peptide Toxicity Prediction
abstract
Peptide-based therapeutics hold great potential, yet their cytotoxicity remains a key challenge in drug development. Most existing toxicity prediction models rely solely on sequence information, often overlooking the fusion of multimodal submolecular patterns. We propose GSToxi, a multimodal deep learning framework that leverage sequence and molecular graph featurizer to enhance peptide toxicity prediction. A shared gating mechanism is employed to facilitate semantic alignment and cross-modal integration, while a contrastive regularization loss further optimizes latent-space consistency throughout the training process. Furthermore, GSToxi incorporates embeddings from pre-trained protein language models alongside low-level compositional priors, enabling the capture of both global contextual semantics and local structural features. Experimental results show that GSToxi outperforms state-of-the-art baselines across multiple evaluation metrics on an independent test set. Ablation studies underscore the critical contributions of each component, with the molecular graph encoder and pre-trained embeddings proving particularly impactful. This work offers a generalizable and robust framework for peptide toxicity prediction and provides valuable insights for future multimodal modeling of biological molecules.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai
BIBM4
2025 Deep Feature Learning for Multi-Omics Clustering Using Supervised Variational Autoencoders with Clinical Information
abstract
Identifying cancer subtypes through multi-omics clustering holds potential to advance cancer research by uncovering subtype-specific mechanisms. Most existing multi-omics clustering methods extract features in an unsupervised manner and integrate data separately, making it difficult to obtain discriminative feature representations and achieve effective data fusion. In this study, we propose a novel multi-omics clustering method that jointly performs feature extraction and data integration within a unified learning framework based on supervised variational autoencoders (VAEs). The proposed method first encodes each omics dataset using VAEs to extract omics-specific latent features, which are then integrated into a shared latent representation. Simultaneously, we utilize the shared latent space to reconstruct the original omics while incorporating patients’ clinical information through a classification task. By jointly optimizing the reconstruction loss and classification loss, the proposed method enables feature extraction that preserves the unique characteristics of each omics while leveraging clinical information to enhance the shared latent representation, thereby improving the integration of multi-omics data with clinical relevance. The shared latent representation is then used for clustering to identify cancer subtypes. Extensive experimental results demonstrate the robustness and effectiveness of the proposed method in identifying cancer subtypes across multiple cancer datasets.
Xiucai Ye, Tetsuya Sakurai
IJCNN2
2025 Understanding Surgical Triplet Videos Through Transferable Visual Models from Natural Language Supervision
Aoying Wang, Yu-Xi Xie, Xiucai Ye, Patrizia Savi
PRCV (18)5
2025 MOFormer: navigating the antimicrobial peptide design space with Pareto-based multi-objective transformer
abstract
Antimicrobial peptide (AMP) design through deep learning holds the potential to revolutionize antibiotic development. Despite recent progress in AMP generation, designing peptide antibiotics with multiple optimal properties remains a significant challenge. We present MOFormer, an advanced multi-objective AMP design pipeline capable of optimizing multiple AMP properties simultaneously. By leveraging a conditional Transformer, the model refines the AMP sequence-property landscape for efficient multi-objective generation. It also incorporates regularization techniques to maintain a highly structured space, enabling the sampling of precise and desirable candidates. Comparative analyses reveal that MOFormer achieves the optimal hypervolume in the multi-objective space, surpassing advanced methods in simultaneously maximizing antimicrobial activity (minimum inhibitory concentration) and minimizing hemolysis and toxicity, thereby yielding the most promising and desirable set of candidate peptides. When extended to a tri-objective scenario, MOFormer continues to exhibit remarkable optimization performance. Finally, we execute a hierarchical and rapid ranking of generated candidates based on Pareto fronts. We conducted a comprehensive validation of the physicochemical properties and target attributes of the candidates, while AlphaFold structure predictions revealed notably reliable predicted local distance difference test scores ranging from 70% to 87%. Our findings suggest that MOFormer holds potential to accelerate the discovery of efficacious peptide antibiotics by optimizing multi-objective trade-offs.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng
Briefings Bioinform.5
2025 PCDP-CRLPPM: a classified regional location privacy-protection model based on personalized clustering with differential privacy in data management
abstract
Abstract Location data management plays a crucial role in facilitating data collection and supporting location-based services. However, the escalating volume of transportation big data has given rise to increased concerns regarding privacy and security issues in data management, potentially posing threats to the lives and property of users. At present, there are two possible attacks in data management, namely Reverse-clustering Inference Attack and Mobile-spatiotemporal Feature Inference Attack. Additionally, the dynamic allocation of privacy budgets emerges as an NP-hard problem. To protect data privacy and maintain utility in data management, a novel protection model for location privacy information in data management, Classified Regional Location Privacy-Protection Model based on Personalized Clustering with Differential Privacy (PCDP-CRLPPM), is proposed. Firstly, a twice-clustering algorithm combined with gridding is proposed, which divides continuous locations into different clusters based on the different privacy protection needs of different users. Subsequently, these clusters are categorized into different spatiotemporal feature regions. Then, a Sensitive-priority algorithm is proposed to allocate privacy budgets adaptively for each region. Finally, a Regional-fuzzy algorithm is presented to introduce Laplacian noise into the centroids of the regions, thereby safeguarding users’ location privacy. The experimental results demonstrate that, compared to other models, PCDP-CRLPPM exhibits superior resistance against two specific attack models and achieves high levels of data utility while preserving privacy effectively.
Wenlong Shi, Jing Zhang 0040, Xiucai Ye
Comput. J.4
2025 LSTM-TRPS: Trajectory reconstruction protection strategy based on semantic information encoding
Jing Zhang 0040, Haoze Hu, Huaxiong Liao, Xiucai Ye
Comput. Networks5
2025 EAFL-ALP: Energy-Efficient Asynchronous Federated Learning With Adaptive Layered Personalization for Vehicular Networks
abstract
Federated Learning (FL) is the standard paradigm for privacy-preserving model training across distributed Industrial IoT (IIoT) devices; however, deployment remains hindered by non-IID data, high communication costs, and unstable asynchronous convergence. We present Energy-Aware Asynchronous Federated Learning with Adaptive Layered Personalization (EAFL-ALP), which achieves a 99.8% reduction in per-round traffic while improving accuracy and robustness. The framework comprises three coordinated modules: (1) Adaptive Fractal-Wave Personalisation Model (AFWPM), which for each client, grows an entropy-conditioned fractal branch and prunes it with wave-collapse, yielding a self-similar, capacity-adaptive head that captures data heterogeneity; (2) Layerwise Quantization-Based Reversible Differential Privacy Gradient Compression (LQGCM), a variance-driven block stratified that transmits 88-bit meta tuples only, enabling codebook resonance replay, invertible vector quantization and Laplace-private gradients without any numeric payload or sparsity mask; (3) Energy-Minimisation Aggregation Model (EMAM), a closed-form update that mixes staleness weights, proxy-gradient correction and EMA momentum for stable convergence on lossy links. Experiments on five IIoT benchmarks show that EAFL-ALP increases accuracy by up to 32.1%, accelerates convergence 3.3×, lowers privacy leakage by 34.1%, and reduces communication volume by two orders of magnitude with no loss of model fidelity.
Jing Zhang 0040, Hong-ming Hou, Meirun Zhang, Li Xu 0002, Xiucai Ye
IEEE Internet Things J.5
2025 BiFDR: Brain-Inspired Federated Diffusion Transformer with Reinforcement for privacy-preserving molecular generation
Hong-ming Hou, Jing Zhang 0040, Meirun Zhang, Xiucai Ye
J. Biomed. Informatics4
2025 DRL-UPPS: User Trajectory Privacy Protection Strategy Based on Deep Reinforcement Learning in Mobile Crowdsensing
abstract
User trajectories are denser and highly dynamic in mobile crowdsensing (MCS) system, rendering traditional privacy budget allocation schemes insufficient. Additionally, the protection of semantic location privacy is often neglected in these schemes, making them vulnerable to inference attacks. To address these deficiencies, a user trajectory privacy protection strategy based on deep reinforcement learning is proposed in this article. First, a differential privacy-based user trajectory privacy protection algorithm (DP-upps) is designed to protect the privacy by perturbing the extracted trajectory feature points. Then, a deep reinforcement learning-based privacy budget allocation algorithm (DRL-pbas) is introduced. The privacy budget is dynamically adjusted by deep reinforcement learning option to continuously adapt to environmental changes and maximize benefits. After that, a DRL-pbas based user privacy protection strategy (DRL-UPPS) is proposed, integrating semantic location privacy protection. This approach combines the previous two algorithms, allowing the privacy budget to be allocated in a way that effectively balances the protection of physical and semantic location privacy and data quality. Ultimately, a large number of simulation experiments are conducted based on real datasets. The experiments demonstrate that DRL-UPPS can effectively balance privacy protection and data quality, resisting the privacy attacks. Compared with other strategies, DRL-UPPS improves comprehensive privacy protection capability by approximately 10% and data utility by approximately 8%.
Jing Zhang 0040, Li Xu 0002, Xiucai Ye
IEEE Trans. Comput. Soc. Syst.5
2025 CSI-FL: Communication-Sensing Integrated Federated Learning Framework for Heterogeneous IoV
abstract
Communication and sensing integration (CSI) technology underpins efficient data acquisition, real-time sensing, and intelligent decision-making in intelligent transportation systems (ITS). By merging communication and sensing, CSI enables seamless data sharing and collaborative learning within the internet of vehicles (IoV), while tackling the complexities of dynamic, heterogeneous environments. However, IoV systems still confront suboptimal resource allocation, synchronization bottlenecks in federated learning (FL), and the delicate balance between privacy and data utility, limiting scalability and deployment. To address these issues, this article presents a CSI-driven federated learning framework comprising three key modules: the polar-driven resource allocation mechanism (PDRAM), the polar-driven asynchronous federated update mechanism (AFUM), and the deep context-aware dynamic privacy budget allocation model (DC-DPBA). Leveraging polar coding principles, PDRAM optimizes communication channel allocation by prioritizing high-fidelity data on high-speed channels. AFUM adopts a vehicle performance score and a progressive aggregation strategy for asynchronous updates, mitigating synchronization challenges. Meanwhile, DC-DPBA uses deep learning and contextual information to dynamically adjust privacy budgets, striking a balance between data utility and privacy protection. Experimental results show that this framework increases transmission efficiency, model training accuracy, and privacy preservation by 22%, 17%, and 17%, respectively, compared to state-of-the-art approaches, offering a scalable and secure solution for CSI-driven IoV environments.
Jing Zhang 0040, Hong-ming Hou, Meirun Zhang, Li Xu 0002, Xiucai Ye
IEEE Trans. Comput. Soc. Syst.5
2025 PKAN: Leveraging Kolmogorov-Arnold Networks and Multi-Modal Learning for Peptide Prediction With Advanced Language Models
abstract
Peptides can offer highly specific biological activities, serving as essential mediators of intercellular signaling, which are critical for advancing precision medicine and drug development. Their primary structure can be depicted either as an amino acid sequence or as a chemical molecules consisting of atoms and chemical bonds. Large language models (LLMs) hold the potential to thoroughly elucidate the intricate intrinsic properties of peptides. Here we present the Peptide Kolmogorov-Arnold Network (PKAN), a framework leveraging multi-modal representations inspired by advanced language models for peptide activity and functionality prediction. Comparative experiments across tasks show that PKAN outperforms state-of-the-art models while maintaining a streamlined design with superior predictive capabilities. The multi-modal feature importance scoring, anchored in global structures and the significant marginal impacts of derived features on the model, coupled with intricate symbolic regression of specific activation functions, further demonstrates the robustness and precision of the PKAN framework in identifying and elucidating key determinants of peptide functionality. This work provides scientific evidence for investigating the complex mechanisms of peptide materials and supports the progression of peptide language paradigms in biology.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng
IEEE J. Biomed. Health Informatics3
2024 Integrating Biological Language Processing and Memory Attention Model for Protein-DNA Binding Residue Prediction
abstract
Proteins are among the most important substances in the human body, and identifying protein-DNA binding sites is crucial for studying their interactions. Although traditional wet-lab methods can accurately identify these sites, they are time-consuming, labor-intensive, and expensive, making it challenging to keep pace with the rapid increase in protein sequence data. In this study, we propose the Memory Attention-Based Protein-DNA Binding Sites Prediction (MAPDB) model, which leverages multi-head Memory Attention for predicting protein-DNA binding sites. Our model employs a pre-trained embedding module to generate numerical representations of protein sequences, followed by a feature extraction module that uses Memory Attention to capture both intra-sequence and inter-sequence relationships. Extensive experiments on five benchmark datasets show that our model outperforms other state-of-the-art methods, especially in improving MCC scores. These results indicate that MAPDB effectively captures complex relationships within protein sequences, leading to more accurate predictions of protein-DNA binding residues.
Xiucai Ye, Tetsuya Sakurai
BIBM2
2024 Selecting interpretable features for cancer subtyping on multi-omics data
abstract
Analyzing multi-omics data is powerful in cancer research, offering essential insights for identifying distinct cancer subtypes. Due to the significant noise and redundancy in multi-omics data, effective feature extraction before data integration is crucial. However, most existing multi-omics clustering methods directly integrate different omics, which may lead to poor clustering results. In this study, we propose a novel multi-omics clustering method which selects interpretable features from different omics before data integration. The proposed method utilizes clinical information and SHAP values to extract interpretable features from different omics. Feature selection is then performed to select the most important interpretable features by grouping and ranking on the SHAP values. We construct similarity network for each omics based on the selected features, and then perform similarity network fusion to integrate the similarities across different omics. Finally, spectral clustering is applied to obtain the clustering result. We conduct experiments on five cancer datasets across three levels of omics to evaluate the proposed method. Experimental results demonstrate the superior performance of our proposed method in multi-omics clustering analysis for cancer subtyping.
Xiucai Ye, Tetsuya Sakurai
BIBM2
2024 Deciphering the Complex Characterization of Coding LncRNA
abstract
This research addresses the intricate nature of coding long non-coding RNAs (lncRNAs), challenging the traditional view of these molecules as merely non-coding elements. By analyzing sequence, physicochemical, and structural features, we have identified distinct characteristics of mRNA, coding lncRNA, and untranslated lncRNA. The CodLncPred model, developed using the XGBoost model, outperforms existing tools in classifying coding lncRNAs. Furthermore, our study evaluates the computational efficiency of various algorithms, including the cost of time and memory, underscoring the practical implications. Our findings offer a new perspective on coding lncRNAs, providing a robust framework for future exploration in genome biology and disease research.
Xiucai Ye, Tetsuya Sakurai
IJCNN2
2024 CyclePermea: Membrane Permeability Prediction of Cyclic Peptides with a Multi-Loss Fusion Network
abstract
Cyclic peptides, known for their unique ring-like structures, show considerable promise in therapeutic applications. Experimentally determining their permeability is time-consuming and labor-intensive. Hence, an efficient and rapid membrane permeability prediction model would greatly expedite the early-stage screening of cyclic peptide drugs. To meet this end, we proposed a novel deep learning model to predict membrane permeability of cyclic peptides, dubbed as CyclePermea. Remarkably, CyclePermea predicts membrane permeability using only the 1D sequence information of cyclic peptides, unlike previous works based on complex spatial descriptors and various physicochemical properties. It incorporates a peptide encoder based on a pre-trained BERT architecture. We also introduced two auxiliary loss functions designed to enhance the model’s comprehension of cyclic peptides’ distinctive characteristics. The first, termed ’Constraint Contrastive Learning Loss’, aims to mitigate the challenge of feature clustering. The second, ’Cyclization Site Prediction Loss’, is proposed to facilitate the model’s recognition of the unique spatial structure inherent in cyclic peptides. Through extensive experiments, CyclePermea demonstrated superior performance over baseline models in the benchmark dataset, both in in-distribution settings and simulated out-of-distribution settings. We hope that CyclePermea would contribute in accelerating the early screening of cyclic peptide drugs in the future.
Yangyang Chen 0006, Xiucai Ye, Tetsuya Sakurai
IJCNN3
2024 Integrated convolution and self-attention for improving peptide toxicity prediction
abstract
MOTIVATION: Peptides are promising agents for the treatment of a variety of diseases due to their specificity and efficacy. However, the development of peptide-based drugs is often hindered by the potential toxicity of peptides, which poses a significant barrier to their clinical application. Traditional experimental methods for evaluating peptide toxicity are time-consuming and costly, making the development process inefficient. Therefore, there is an urgent need for computational tools specifically designed to predict peptide toxicity accurately and rapidly, facilitating the identification of safe peptide candidates for drug development. RESULTS: We provide here a novel computational approach, CAPTP, which leverages the power of convolutional and self-attention to enhance the prediction of peptide toxicity from amino acid sequences. CAPTP demonstrates outstanding performance, achieving a Matthews correlation coefficient of approximately 0.82 in both cross-validation settings and on independent test datasets. This performance surpasses that of existing state-of-the-art peptide toxicity predictors. Importantly, CAPTP maintains its robustness and generalizability even when dealing with data imbalances. Further analysis by CAPTP reveals that certain sequential patterns, particularly in the head and central regions of peptides, are crucial in determining their toxicity. This insight can significantly inform and guide the design of safer peptide drugs. AVAILABILITY AND IMPLEMENTATION: The source code for CAPTP is freely available at https://github.com/jiaoshihu/CAPTP.
Shihu Jiao, Xiucai Ye, Tetsuya Sakurai, Quan Zou 0001
Bioinform.2
2024 GeoPM-DMEIRL: A deep inverse reinforcement learning security trajectory generation framework with serverless computing
Yi-rui Huang, Jing Zhang 0040, Hong-ming Hou, Xiucai Ye
Future Gener. Comput. Syst.4
2024 An interpretable deep learning model predicts RNA-small molecule binding sites
Wen-Yu Xi, Ruheng Wang, Li Wang 0145, Xiucai Ye, Tetsuya Sakurai
Future Gener. Comput. Syst.4
2024 Lightweight privacy-preserving authenticated key agreements using physically unclonable functions for internet of drones
Tian-Fu Lee, Xiucai Ye, Wei-Jie Huang
J. Inf. Secur. Appl.2
2024 Attribute and closeness based scheduling model for vehicle-to-grid network
Jing Zhang 0040, Jian-Yu Hu, Xiucai Ye
Peer Peer Netw. Appl.4
2024 IEA-DP: Information Entropy-driven Adaptive Differential Privacy Protection Scheme for social networks
Jing Zhang 0040, Kunliang Si, Zuanyang Zeng, Xiucai Ye
J. Supercomput.5
2024 Entropy-driven differential privacy protection scheme based on social graphlet attributes
Jing Zhang 0040, Zuanyang Zeng, Kunliang Si, Xiucai Ye
J. Supercomput.4
2023 Multi-omics clustering based on interpretable and discriminative features for cancer subtyping
abstract
Recent advances in multi-omics databases have enabled biomedical researchers to explore complex cancer systems across hierarchical biological levels. Although there are numerous multi-omics clustering methods, most of them directly integrate heterogeneous features of different omics which may include redundancy or noise and lead to poor clustering results. In this paper, we propose a novel multi-omics clustering method for cancer subtyping which extracts interpretable and discriminative features from different omics before data integration. The proposed method utilizes the clinical information of each omics to supervise the process of extracting interpretable and discriminative features based on SHAP (SHapley Additive exPlanation) values. The shared nearest neighbor-based approach is then applied to calculate the similarity matrix of the extracted features. Finally, we integrate the similarity matrices of different omics and apply spectral clustering on the integrated similarity matrix to obtain the clustering result. Experimental results conducted on four different cancer datasets on three levels of omics demonstrate the superior performance of the proposed method in comparison to the existing multi-omics clustering methods.
Xiucai Ye, Tetsuya Sakurai
BIBM2
2023 Multi-view Network Embedding with Structure and Semantic Contrastive Learning
abstract
Multi-view network embedding aims to learn low-dimensional representation vectors for nodes while preserving multiple relationships between nodes. It can substantially reduce downstream network analysis tasks’ time and space complexity. Although previous works have achieved great performance, they suffer from two limitations: (1) they only preserve the network structure and ignore the semantic level information; (2) they only focus on intra-view signals and ignore the powerful influence of inter-view signals. These limitations highlight the need for more comprehensive approaches to multi-view network embedding that can effectively capture the structure and semantic information, as well as the influence of inter-view signals. A new framework, Multi-view Network Embedding with Structure and Semantic Contrastive Learning (MNE-SSCL), is proposed to address these limitations. It can learn high-quality low-dimensional node embeddings in both intra-veiw and interview, while preserving the structure and semantic information simultaneously. Extensive experiments on three real datasets show that MNE-SSCL outperforms the state-of-the-art methods.
Yifan Shang, Xiucai Ye, Tetsuya Sakurai
ICME2
2023 Common and Unique Features Learning in Multi-view Network Embedding
abstract
Network embedding is a powerful representation learning method for graph data, using the learned low-dimensional compact vectors as node features, which are widely used in various tasks, such as link prediction, node clustering, and classification. Compared with traditional network analysis methods, network embedding reduces computational complexity and improves analysis efficiency. Although previous work has achieved outstanding performance, it faces challenges in multi-view network embedding containing multi-type node relations. Since multi-view networks share a node set but different edges, different networks not only have common information but also have their unique information. To simultaneously capture multi-view networks' common and unique information, we propose a new framework, Common and Unique Features Learning for Multi-view Network Embedding (CU-MNE), to integrate multi-type node relations. In this paper, we propose an inter-view contrastive objective to ensure the consistency of the common features of the same node in a different view and an inter-feature contrastive objective to capture the association between the common and unique features of each network node that can learn high-quality node embeddings. Extensive experiments on three real datasets show that CU-MNE outperforms the state-of-the-art methods.
Yifan Shang, Xiucai Ye, Tetsuya Sakurai
IJCNN2
2023 SiameseCPP: a sequence-based Siamese network to predict cell-penetrating peptides by contrastive learning
abstract
BACKGROUND: Cell-penetrating peptides (CPPs) have received considerable attention as a means of transporting pharmacologically active molecules into living cells without damaging the cell membrane, and thus hold great promise as future therapeutics. Recently, several machine learning-based algorithms have been proposed for predicting CPPs. However, most existing predictive methods do not consider the agreement (disagreement) between similar (dissimilar) CPPs and depend heavily on expert knowledge-based handcrafted features. RESULTS: In this study, we present SiameseCPP, a novel deep learning framework for automated CPPs prediction. SiameseCPP learns discriminative representations of CPPs based on a well-pretrained model and a Siamese neural network consisting of a transformer and gated recurrent units. Contrastive learning is used for the first time to build a CPP predictive model. Comprehensive experiments demonstrate that our proposed SiameseCPP is superior to existing baseline models for predicting CPPs. Moreover, SiameseCPP also achieves good performance on other functional peptide datasets, exhibiting satisfactory generalization ability.
Lesong Wei, Xiucai Ye, Saisai Teng, Zhongshen Li, Junru Jin, Min Jae Kim, Tetsuya Sakurai, Li-Zhen Cui 0001, Balachandran Manavalan, Leyi Wei
Briefings Bioinform.3
2023 Adaptive learning embedding features to improve the predictive performance of SARS-CoV-2 phosphorylation sites
abstract
MOTIVATION: The rapid and extensive transmission of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has led to an unprecedented global health emergency, affecting millions of people and causing an immense socioeconomic impact. The identification of SARS-CoV-2 phosphorylation sites plays an important role in unraveling the complex molecular mechanisms behind infection and the resulting alterations in host cell pathways. However, currently available prediction tools for identifying these sites lack accuracy and efficiency. RESULTS: In this study, we presented a comprehensive biological function analysis of SARS-CoV-2 infection in a clonal human lung epithelial A549 cell, revealing dramatic changes in protein phosphorylation pathways in host cells. Moreover, a novel deep learning predictor called PSPred-ALE is specifically designed to identify phosphorylation sites in human host cells that are infected with SARS-CoV-2. The key idea of PSPred-ALE lies in the use of a self-adaptive learning embedding algorithm, which enables the automatic extraction of context sequential features from protein sequences. In addition, the tool uses multihead attention module that enables the capturing of global information, further improving the accuracy of predictions. Comparative analysis of features demonstrated that the self-adaptive learning embedding features are superior to hand-crafted statistical features in capturing discriminative sequence information. Benchmarking comparison shows that PSPred-ALE outperforms the state-of-the-art prediction tools and achieves robust performance. Therefore, the proposed model can effectively identify phosphorylation sites assistant the biomedical scientists in understanding the mechanism of phosphorylation in SARS-CoV-2 infection. AVAILABILITY AND IMPLEMENTATION: PSPred-ALE is available at https://github.com/jiaoshihu/PSPred-ALE and Zenodo (https://doi.org/10.5281/zenodo.8330277).
Shihu Jiao, Xiucai Ye, Chunyan Ao, Tetsuya Sakurai, Quan Zou 0001, Lei Xu 0047
Bioinform.2
2023 Distortion-free PCA on sample space for highly variable gene detection from single-cell RNA-seq data
Momo Matsuda, Yasunori Futamura, Xiucai Ye, Tetsuya Sakurai
Frontiers Comput. Sci.3
2023 Hasse sensitivity level: A sensitivity-aware trajectory privacy-enhanced framework with Reinforcement Learning
Jing Zhang 0040, Yi-rui Huang, Qihan Huang, Yan-zi Li, Xiucai Ye
Future Gener. Comput. Syst.5
2023 LSEC: Large-scale spectral ensemble clustering
abstract
A fundamental problem in machine learning is ensemble clustering, that is, combining multiple base clusterings to obtain improved clustering result. However, most of the existing methods are unsuitable for large-scale ensemble clustering tasks owing to efficiency bottlenecks. In this paper, we propose a large-scale spectral ensemble clustering (LSEC) method to balance efficiency and effectiveness. In LSEC, a large-scale spectral clustering-based efficient ensemble generation framework is designed to generate various base clusterings with low computational complexity. Thereafter, all the base clusterings are combined using a bipartite graph partition-based consensus function to obtain improved consensus clustering results. The LSEC method achieves a lower computational complexity than most existing ensemble clustering methods. Experiments conducted on ten large-scale datasets demonstrate the efficiency and effectiveness of the LSEC method. The MATLAB code of the proposed method and experimental datasets are available at https://github.com/Li-Hongmin/MyPaperWithCode.
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
Intell. Data Anal.2
2023 Unravelling cancer subtype-specific driver genes in single-cell transcriptomics data with CSDGI
abstract
Cancer is known as a heterogeneous disease. Cancer driver genes (CDGs) need to be inferred for understanding tumor heterogeneity in cancer. However, the existing computational methods have identified many common CDGs. A key challenge exploring cancer progression is to infer cancer subtype-specific driver genes (CSDGs), which provides guidane for the diagnosis, treatment and prognosis of cancer. The significant advancements in single-cell RNA-sequencing (scRNA-seq) technologies have opened up new possibilities for studying human cancers at the individual cell level. In this study, we develop a novel unsupervised method, CSDGI (Cancer Subtype-specific Driver Gene Inference), which applies Encoder-Decoder-Framework consisting of low-rank residual neural networks to inferring driver genes corresponding to potential cancer subtypes at the single-cell level. To infer CSDGs, we apply CSDGI to the tumor single-cell transcriptomics data. To filter the redundant genes before driver gene inference, we perform the differential expression genes (DEGs). The experimental results demonstrate CSDGI is effective to infer driver genes that are cancer subtype-specific. Functional and disease enrichment analysis shows these inferred CSDGs indicate the key biological processes and disease pathways. CSDGI is the first method to explore cancer driver genes at the cancer subtype level. We believe that it can be a useful method to understand the mechanisms of cell transformation driving tumours.
Meng Huang 0003, Jiangtao Ma, Guangqi An, Xiucai Ye
PLoS Comput. Biol.4
2023 Dimension-aware under spatiotemporal constraints: an efficient privacy-preserving framework with peak density clustering
Jing Zhang 0040, Qihan Huang, Jian-Yu Hu, Xiucai Ye
J. Supercomput.4
2022 Ensemble Learning for Cluster Number Detection Based on Shared Nearest Neighbor Graph and Spectral Clustering
abstract
Detecting the number of clusters is important for cluster analysis. Many existing methods detect the cluster number by predefining a list of candidate cluster numbers. However, if the candidate cluster numbers are not well predefined, the cluster number cannot be correctly detected. In this paper, we propose a novel clustering method which automatically generates the candidate cluster numbers and the corresponding cluster partitions based on multiple shared nearest neighbor graphs. A shared low-rank similarity matrix is then recovered from the cluster partitions by ensemble learning. Finally, spectral clustering is applied on the shared low-rank similarity matrix with the candidate cluster numbers to detect the cluster number. Experimental results on both synthetic and real-world datasets demonstrate that the proposed method not only correctly detects the cluster numbers, but also obtains better clustering results in comparison to the existing methods.
Xiucai Ye, Testuya Sakurai
IJCNN2
2022 Multiview network embedding for drug-target Interactions prediction by consistent and complementary information preserving
abstract
Accurate prediction of drug-target interactions (DTIs) can reduce the cost and time of drug repositioning and drug discovery. Many current methods integrate information from multiple data sources of drug and target to improve DTIs prediction accuracy. However, these methods do not consider the complex relationship between different data sources. In this study, we propose a novel computational framework, called MccDTI, to predict the potential DTIs by multiview network embedding, which can integrate the heterogenous information of drug and target. MccDTI learns high-quality low-dimensional representations of drug and target by preserving the consistent and complementary information between multiview networks. Then MccDTI adopts matrix completion scheme for DTIs prediction based on drug and target representations. Experimental results on two datasets show that the prediction accuracy of MccDTI outperforms four state-of-the-art methods for DTIs prediction. Moreover, literature verification for DTIs prediction shows that MccDTI can predict the reliable potential DTIs. These results indicate that MccDTI can provide a powerful tool to predict new DTIs and accelerate drug discovery. The code and data are available at: https://github.com/ShangCS/MccDTI.
Yifan Shang, Xiucai Ye, Yasunori Futamura, Liang Yu 0002, Tetsuya Sakurai
Briefings Bioinform.2
2022 iLoc-miRNA: extracellular/intracellular miRNA prediction using deep BiLSTM with attention mechanism
abstract
The location of microRNAs (miRNAs) in cells determines their function in regulation activity. Studies have shown that miRNAs are stable in the extracellular environment that mediates cell-to-cell communication and are located in the intracellular region that responds to cellular stress and environmental stimuli. Though in situ detection techniques of miRNAs have made great contributions to the study of the localization and distribution of miRNAs, miRNA subcellular localization and their role are still in progress. Recently, some machine learning-based algorithms have been designed for miRNA subcellular location prediction, but their performance is still far from satisfactory. Here, we present a new data partitioning strategy that categorizes functionally similar locations for the precise and instructive prediction of miRNA subcellular location in Homo sapiens. To characterize the localization signals, we adopted one-hot encoding with post padding to represent the whole miRNA sequences, and proposed a deep bidirectional long short-term memory with the multi-head self-attention algorithm to model. The algorithm showed high selectivity in distinguishing extracellular miRNAs from intracellular miRNAs. Moreover, a series of motif analyses were performed to explore the mechanism of miRNA subcellular localization. To improve the convenience of the model, a user-friendly web server named iLoc-miRNA was established (http://iLoc-miRNA.lin-group.cn/).
Zhao-Yue Zhang 0002, Lin Ning 0002, Xiucai Ye, Yasunori Futamura, Tetsuya Sakurai, Hao Lin 0001
Briefings Bioinform.3
2022 NerLTR-DTA: drug-target binding affinity prediction based on neighbor relationship and learning to rank
abstract
MOTIVATION: Drug-target interaction prediction plays an important role in new drug discovery and drug repurposing. Binding affinity indicates the strength of drug-target interactions. Predicting drug-target binding affinity is expected to provide promising candidates for biologists, which can effectively reduce the workload of wet laboratory experiments and speed up the entire process of drug research. Given that, numerous new proteins are sequenced and compounds are synthesized, several improved computational methods have been proposed for such predictions, but there are still some challenges. (i) Many methods only discuss and implement one application scenario, they focus on drug repurposing and ignore the discovery of new drugs and targets. (ii) Many methods do not consider the priority order of proteins (or drugs) related to each target drug (or protein). Therefore, it is necessary to develop a comprehensive method that can be used in multiple scenarios and focuses on candidate order. RESULTS: In this study, we propose a method called NerLTR-DTA that uses the neighbor relationship of similarity and sharing to extract features, and applies a ranking framework with regression attributes to predict affinity values and priority order of query drug (or query target) and its related proteins (or compounds). It is worth noting that using the characteristics of learning to rank to set different queries can smartly realize the multi-scenario application of the method, including the discovery of new drugs and new targets. Experimental results on two commonly used datasets show that NerLTR-DTA outperforms some state-of-the-art competing methods. NerLTR-DTA achieves excellent performance in all application scenarios mentioned in this study, and the rm(test)2 values guarantee such excellent performance is not obtained by chance. Moreover, it can be concluded that NerLTR-DTA can provide accurate ranking lists for the relevant results of most queries through the statistics of the association relationship of each query drug (or query protein). In general, NerLTR-DTA is a powerful tool for predicting drug-target associations and can contribute to new drug discovery and drug repurposing. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented in Python and Java. Source codes and datasets are available at https://github.com/RUXIAOQING964914140/NerLTR-DTA.
Xiaoqing Ru, Xiucai Ye, Tetsuya Sakurai, Quan Zou 0001
Bioinform.2
2022 ToxIBTL: prediction of peptide toxicity based on information bottleneck and transfer learning
abstract
MOTIVATION: Recently, peptides have emerged as a promising class of pharmaceuticals for various diseases treatment poised between traditional small molecule drugs and therapeutic proteins. However, one of the key bottlenecks preventing them from therapeutic peptides is their toxicity toward human cells, and few available algorithms for predicting toxicity are specially designed for short-length peptides. RESULTS: We present ToxIBTL, a novel deep learning framework by utilizing the information bottleneck principle and transfer learning to predict the toxicity of peptides as well as proteins. Specifically, we use evolutionary information and physicochemical properties of peptide sequences and integrate the information bottleneck principle into a feature representation learning scheme, by which relevant information is retained and the redundant information is minimized in the obtained features. Moreover, transfer learning is introduced to transfer the common knowledge contained in proteins to peptides, which aims to improve the feature representation capability. Extensive experimental results demonstrate that ToxIBTL not only achieves a higher prediction performance than state-of-the-art methods on the peptide dataset, but also has a competitive performance on the protein dataset. Furthermore, a user-friendly online web server is established as the implementation of the proposed ToxIBTL. AVAILABILITY AND IMPLEMENTATION: The proposed ToxIBTL and data can be freely accessible at http://server.wei-group.net/ToxIBTL. Our source code is available at https://github.com/WLYLab/ToxIBTL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lesong Wei, Xiucai Ye, Tetsuya Sakurai, Zengchao Mu, Leyi Wei
Bioinform.2
2022 Divide-and-conquer based large-scale spectral clustering
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
Neurocomputing2
2022 Anonymous Dynamic Group Authenticated Key Agreements Using Physical Unclonable Functions for Internet of Medical Things
abstract
Group authenticated key agreements (GAKAs) for an Internet of Medical Things (IoMT) enable medical sensor devices to authenticate each other and agree upon a common session key. These medical sensor devices can then establish a secure channel using this session key and exchange information securely. Owing to hardware limitations, sensors do not have sufficient computing power, energy, and storage resources to implement burdensome protocols. Moreover, sensors are easily damaged, causing the number of members of the group to change. Therefore, a group key agreement (GKA) protocol is required to consider dynamic groups and reduce the computational cost of an IoMT. The physical unclonable function (PUF) uses the uniqueness and randomness of its circuit to perform calculations. It has message fingerprints so that the user is not necessary to store a private key. The PUF is less computationally burdensome than the more common modular exponentiation and can be applied in an IoMT environment with limited resources. This article develops lightweight anonymous dynamic GAKA (DGAKA) that uses the PUF to solve the problems of data storage, transmission security, and computing efficiency that are faced by the IoMT. The proposed protocol has higher security and efficiency in computation and storage than those previously developed.
Tian-Fu Lee, Xiucai Ye, Syuan-Han Lin
IEEE Internet Things J.2
2022 Sequential reinforcement active feature learning for gene signature identification in renal cell carcinoma
Meng Huang 0003, Xiucai Ye, Akira Imakura, Tetsuya Sakurai
J. Biomed. Informatics2
2021 Collaborative Novelty Detection for Distributed Data by a Probabilistic Method
abstract
Novelty detection, which detects anomalies based on a training dataset consisting of only the normal data, is an important task in several applications. In addition, in the real world, there may be situations where data is owned by multiple parties in a distributed manner but cannot be shared with each other due to privacy and confidentiality requirements. Therefore, how to develop distributed novelty detection while preserving privacy is essential. To address this challenge, we propose a probabilistic collaborative method that allows distributed novelty detection for multiple parties without sharing the original data. The proposed method constructs a collaborative kernel based on a collaborative data analysis framework, by which intermediate representations are generated from each party and shared for collaborative novelty detection. Numerical experiments demonstrate that the proposed method obtains better performance compared with the individual novelty detection in the local party.
Akira Imakura, Xiucai Ye, Tetsuya Sakurai
ACML2
2021 Spectral Clustering Joint Deep Embedding Learning by Autoencoder
abstract
Spectral clustering has become one of the most popular clustering methods due to its superior performance compared to the traditional clustering methods. However, the performance of spectral clustering would be limited by complex data, such as a huge number of samples and high dimensionality. To address this problem, some existing methods apply deep learning to learn the lower-dimensional representations, spectral clustering is then applied to the representations. Different from the existing methods that separate the two stages of feature representation learning and spectral clustering, in this paper, we propose Spectral Clustering Joint Deep Embedding (SCJDE), a method that simultaneously learns the feature representations and the spectral embedding of spectral clustering via a deep autoencoder. Moreover, a sparsity constraint is imposed to generate better spectral embedding for spectral clustering. Finally,$k$-means is performed on the spectral embedding to obtain clustering result. The proposed method can learn a good spectral embedding for spectral clustering by deep learning to obtain better clustering results. The experimental results on both synthetic and real-world datasets demonstrate the effectiveness of the proposed method.
Xiucai Ye, Chunhao Wang, Akira Imakura, Tetsuya Sakurai
IJCNN1
2021 Application of learning to rank in bioinformatics tasks
abstract
Over the past decades, learning to rank (LTR) algorithms have been gradually applied to bioinformatics. Such methods have shown significant advantages in multiple research tasks in this field. Therefore, it is necessary to summarize and discuss the application of these algorithms so that these algorithms are convenient and contribute to bioinformatics. In this paper, the characteristics of LTR algorithms and their strengths over other types of algorithms are analyzed based on the application of multiple perspectives in bioinformatics. Finally, the paper further discusses the shortcomings of the LTR algorithms, the methods and means to better use the algorithms and some open problems that currently exist.
Xiaoqing Ru, Xiucai Ye, Tetsuya Sakurai, Quan Zou 0001
Briefings Bioinform.2
2021 ATSE: a peptide toxicity predictor by exploiting structural and evolutionary information based on graph neural network and attention mechanism
abstract
MOTIVATION: Peptides have recently emerged as promising therapeutic agents against various diseases. For both research and safety regulation purposes, it is of high importance to develop computational methods to accurately predict the potential toxicity of peptides within the vast number of candidate peptides. RESULTS: In this study, we proposed ATSE, a peptide toxicity predictor by exploiting structural and evolutionary information based on graph neural networks and attention mechanism. More specifically, it consists of four modules: (i) a sequence processing module for converting peptide sequences to molecular graphs and evolutionary profiles, (ii) a feature extraction module designed to learn discriminative features from graph structural information and evolutionary information, (iii) an attention module employed to optimize the features and (iv) an output module determining a peptide as toxic or non-toxic, using optimized features from the attention module. CONCLUSION: Comparative studies demonstrate that the proposed ATSE significantly outperforms all other competing methods. We found that structural information is complementary to the evolutionary information, effectively improving the predictive performance. Importantly, the data-driven features learned by ATSE can be interpreted and visualized, providing additional information for further analysis. Moreover, we present a user-friendly online computational platform that implements the proposed ATSE, which is now available at http://server.malab.cn/ATSE. We expect that it can be a powerful and useful tool for researchers of interest.
Lesong Wei, Xiucai Ye, Yuyang Xue, Tetsuya Sakurai, Leyi Wei
Briefings Bioinform.2
2020 Ensemble Learning for Spectral Clustering
abstract
Ensemble clustering has attracted much attention in machine learning and data mining for the high performance in the task of clustering. Spectral clustering is one of the most popular clustering methods and has superior performance compared with the traditional clustering methods. Existing ensemble clustering methods usually directly use the clustering results of the base clustering algorithms for ensemble learning, which cannot make good use of the intrinsic data structures explored by the graph Laplacians in spectral clustering, thus cannot obtain the desired clustering result. In this paper, we propose a new ensemble learning method for spectral clustering-based clustering algorithms. Instead of directly using the clustering results obtained from each base spectral clustering algorithm, the proposed method learns a robust presentation of graph Laplacian by ensemble learning from the spectral embedding of each base spectral clustering algorithm. Finally, the proposed method applies k-means on the spectral embedding obtain from the learned graph Laplacian to get clusters. Experimental results on both synthetic and real-world datasets show that the proposed method outperforms other existing ensemble clustering methods.
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
ICDM2
2020 Hubness-based Sampling Method for Nyström Spectral Clustering
abstract
Nyström method is widely used for spectral clustering to obtain low-rank approximations of a large matrix. Sampling is crucial to Nyström method, since selecting the representative sample points that can reflect the data structure is important for obtaining good approximation results. To improve the performance of Nyström based spectral clustering, in this paper, we propose a new sampling method by considering the hubness score of sample points. The data points with the high hubness scores, i.e., appearing frequently in the nearest neighbor lists of other data points, have high probabilities to be selected as the sample points. Taking advantage of the topological property of hubs (i.e., data points with high hubness score), the selected sampling points have close relationships with other data points, thus the proposed method is able to achieve scalable and accurate clustering results. We further design fast computation methods, i.e., local hubness approximated methods, to speed up the sampling process. Experimental results on both synthetic and real-world data sets show that the proposed method not only achieves good performance, but also outperforms other sampling methods for Nyström based spectral clustering.
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
IJCNN2
2020 Collaborative Data Analysis: Non-model Sharing-Type Machine Learning for Distributed Data
Akira Imakura, Xiucai Ye, Tetsuya Sakurai
PKAW2
2020 CPPred-FL: a sequence-based predictor for large-scale identification of cell-penetrating peptides by feature representation learning
abstract
Cell-penetrating peptides (CPPs) have been shown to be a transport vehicle for delivering cargoes into live cells, offering great potential as future therapeutics. It is essential to identify CPPs for better understanding of their functional mechanisms. Machine learning-based methods have recently emerged as a main approach for computational identification of CPPs. However, one of the main challenges and difficulties is to propose an effective feature representation model that sufficiently exploits the inner difference and relevance between CPPs and non-CPPs, in order to improve the predictive performance. In this paper, we have developed CPPred-FL, a powerful bioinformatics tool for fast, accurate and large-scale identification of CPPs. In our predictor, we introduce a new feature representation learning scheme that enables one to learn feature representations from totally 45 well-trained random forest models with multiple feature descriptors from different perspectives, such as compositional information, position-specific information and physicochemical properties, etc. We integrate class and probabilistic information into our feature representations. To improve the feature representation ability, we further remove redundant and irrelevant features by feature space optimization. Benchmarking experiments showed that CPPred-FL, using 19 informative features only, is able to achieve better performance than the state-of-the-art predictors. We anticipate that CPPred-FL will be a powerful tool for large-scale identification of CPPs, facilitating the characterization of their functional mechanisms and accelerating their applications in clinical therapy.
Xiaoli Qiang, Xiucai Ye, Pufeng Du, Ran Su, Leyi Wei
Briefings Bioinform.3
2020 Multiclass spectral feature scaling method for dimensionality reduction
abstract
Irregular features disrupt the desired classification. In this paper, we consider aggressively modifying scales of features in the original space according to the label information to form well-separated clusters in low-dimensional space. The proposed method exploits spectral clustering to derive scaling factors that are used to modify the features. Specifically, we reformulate the Laplacian eigenproblem of the spectral clustering as an eigenproblem of a linear matrix pencil whose eigenvector has the scaling factors. Numerical experiments show that the proposed method outperforms well-established supervised dimensionality reduction methods for toy problems with more samples than features and real-world problems with more features than samples.
Momo Matsuda, Keiichi Morikuni, Akira Imakura, Xiucai Ye, Tetsuya Sakurai
Intell. Data Anal.4
2020 An oversampling framework for imbalanced classification based on Laplacian eigenmaps
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
Neurocomputing1
2019 Complex Moment-Based Supervised Eigenmap for Dimensionality Reduction
abstract
Dimensionality reduction methods that project highdimensional data to a low-dimensional space by matrix trace optimization are widely used for clustering and classification. The matrix trace optimization problem leads to an eigenvalue problem for a low-dimensional subspace construction, preserving certain properties of the original data. However, most of the existing methods use only a few eigenvectors to construct the low-dimensional space, which may lead to a loss of useful information for achieving successful classification. Herein, to overcome the deficiency of the information loss, we propose a novel complex moment-based supervised eigenmap including multiple eigenvectors for dimensionality reduction. Furthermore, the proposed method provides a general formulation for matrix trace optimization methods to incorporate with ridge regression, which models the linear dependency between covariate variables and univariate labels. To reduce the computational complexity, we also propose an efficient and parallel implementation of the proposed method. Numerical experiments indicate that the proposed method is competitive compared with the existing dimensionality reduction methods for the recognition performance. Additionally, the proposed method exhibits high parallel efficiency.
Akira Imakura, Momo Matsuda, Xiucai Ye, Tetsuya Sakurai
AAAI3
2019 Distributed Collaborative Feature Selection Based on Intermediate Representation
abstract
Feature selection is an efficient dimensionality reduction technique for artificial intelligence and machine learning. Many feature selection methods learn the data structure to select the most discriminative features for distinguishing different classes. However, the data is sometimes distributed in multiple parties and sharing the original data is difficult due to the privacy requirement. As a result, the data in one party may be lack of useful information to learn the most discriminative features. In this paper, we propose a novel distributed method which allows collaborative feature selection for multiple parties without revealing their original data. In the proposed method, each party finds the intermediate representations from the original data, and shares the intermediate representations for collaborative feature selection. Based on the shared intermediate representations, the original data from multiple parties are transformed to the same low dimensional space. The feature ranking of the original data is learned by imposing row sparsity on the transformation matrix simultaneously. Experimental results on real-world datasets demonstrate the effectiveness of the proposed method.
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
IJCAI1
2018 Spectral clustering with adaptive similarity measure in Kernel space
abstract
The similarity measure for complex data may not precisely reflect the true data structure, which leads to suboptimal clustering performance for spectral clustering. In this paper, we propose a novel spectral clustering method which measures the similarity of data points based on the adaptive neighb orhood in Kernel space. In Kernel space, by assigning the adaptive and optimal neighbors for each data point based on the local structure, the proposed method learns a sparse matrix as the similarity matrix for spectral clustering. The proposed method is able to explore the underlying similarity relationships between data points, and is robust to the complex data. To validate the efficacy of the proposed method, we perform experiments on both synthetic and real datasets in comparison with some existing spectral clustering methods. The experimental results demonstrate that the proposed method obtains quite promising clustering performance.
Xiucai Ye, Tetsuya Sakurai
Intell. Data Anal.1
2016 Spectral clustering and discriminant analysis for unsupervised feature selection
Xiucai Ye, Kaiyang Ji, Tetsuya Sakurai
ESANN1
2015 Spectral clustering using robust similarity measure based on closeness of shared Nearest Neighbors
abstract
Spectral clustering has become one of the main clustering methods and has a wide range of applications. Similarity measure is crucial to correct cluster separation for spectral clustering. Many existing spectral clustering algorithms typically measure similarity based on the undirected k-Nearest Neighbor (kNN) graph or Gaussian kernel function, which can not reveal the real clusters of not well-separated data sets. In this paper, we propose a novel algorithm called Spectral Clustering based on Shared Nearest Neighbors (SC-SNN) to improve the clustering quality of not well-separated data sets. Instead of using distance for the similarity measure, the proposed SC-SNN algorithm measures the similarity by considering the closeness of shared nearest neighbors in the directed kNN graph, which is able to explore the underlying similarity relationships between data points and is robust to the not well-separated data sets. Moreover, SC-SNN has only one parameter, k, and is less sensitive than the spectral clustering algorithms based on the undirected kNN graph. The proposed SC-SNN algorithm is evaluated by using both synthetic and real-world data sets. The experimental results demonstrate that SC-SNN not only achieves good performance, but also outperforms the traditional spectral clustering algorithms.
Xiucai Ye, Tetsuya Sakurai
IJCNN1
2015 LT codes based distributed coding for efficient distributed storage in Wireless Sensor Networks
abstract
Fountain codes are linear codes with low complexities. LT (Luby Transform) codes, which are a special class of Fountain codes, are widely used in Wireless Sensor Networks (WSNs) to increase the robustness of data storage and efficiency of data retrieval. In this paper, we propose a novel LT codes based Distributed Coding (LTDC) scheme for efficient distributed storage in WSNs. In the proposed LTDC scheme, we use random walks to disseminate sensed data from a source sensor node to a random subset of sensor nodes by multicast. As long as a data packet stops at an ending sensor node of a random walk, the ending sensor node encodes this data packet in a main packet (an encoded data packet) with a certain probability. By adjusting the main packet with the un-encoded data packets, the number of data packets encoded in the main packet follows the distribution of LT codes. The data collector is able to decode the original data by querying any subset of sensor nodes. The theoretical analysis and simulation results have demonstrated that the proposed LTDC scheme has lower data dissemination cost and lower storage overhead, while maintains the same level of fault tolerance as the original LT codes.
Xiucai Ye, Jie Li 0002, Wen-Tsuen Chen, Feilong Tang 0001
Networking1
2015 A novel sleep scheduling scheme in green wireless sensor networks
Jing Zhang 0040, Li Xu 0002, Shuming Zhou, Xiucai Ye
J. Supercomput.4
2014 Distributed Separate Coding for Continuous Data Collection in Wireless Sensor Networks
abstract
In this article, we present a novel distributed separate coding (DSC) scheme for continuous data collection in wireless sensor networks with a mobile base station (mBS). By separately encoding a certain number of data segments in a combined segment and doing decoding-free data replacement in the buffers of each sensor node, the DSC scheme is shown as an efficient method for continuously collecting data segments with a high success ratio. The proposed DSC scheme has a salient feature: with a minimum buffer size 2 in each sensor node, by querying any m −1 sensor nodes, the mBS can reconstruct the m latest data segments with high probability, where m is the number of latest data segments in a time interval t in which n ( t ) ( m ≤ n ( t )) data segments are generated. The necessary storage space in each sensor node can be adjusted by changing the number of sensor nodes queried by the mBS. Furthermore, the transmission cost for data submission to the mBS can be reduced with some additional storage space in each sensor node. The comprehensive performance evaluation has been conducted through computer simulation. It is shown that the proposed DSC scheme outperforms the existing scheme significantly.
Xiucai Ye, Jie Li 0002, Li Xu 0002
ACM Trans. Sens. Networks1
2013 Group Data Collection in wireless sensor networks with a mobile base station
abstract
In this paper, we present a novel Group Data Collection (GDC) scheme for wireless sensor networks with a mobile base station by using coding for data storage in the sensor nodes. By separately encoding a certain number of data segments in a combined segment and doing data replacement in each sensor node, the proposed GDC scheme not only provides an efficient storage method for group data, but also achieves a high success ratio of data collection. The number of necessary buffers in each sensor node can be adjusted by changing the frequency of performing data collection. The performance evaluation has been conducted through comprehensive computer simulations. It further demonstrates the feasibility and superiority of the proposed GDC scheme.
Xiucai Ye, Jie Li 0002, Li Xu 0002
WCNC1
2012 A Novel Data Collection Scheme for WSNs
abstract
In this paper, we present a novel data collection scheme for wireless sensor networks by using separate network coding (SNC). By separately encoding a certain number of data segments in a combined data segment and doing decoding-free data replacement, SNC not only provides efficient storage method for continuous data, but also maintains a high success ratio of data collection. The performance evaluation has been conducted through comprehensive computer simulation. It is shown that SNC outperforms the exiting scheme significantly.
Jie Li 0002, Xiucai Ye, Li Xu 0002
VTC Spring2