Tetsuya Sakurai

dblp:89/3511 · DBLP profile ↗
← Back
65ranked-venue papers
2as first author
40since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 23 · 19 since 2021Systems, architecture and hardware · 6 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 4Theory of computation · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Robust graph structure learning to improve multi-omics cancer subtype classification
abstract
BACKGROUND: Classifying cancer patients into consistent subtypes at the multi-omics level remains a significant challenge in advancing precision medicine. Nevertheless, a key problem in integrating multi-omics data lies in concurrently addressing intra-omics and inter-omics information, along with sample networks. RESULTS: In this study, we introduce the Feature and Graph Structure-Learning Integrated Graph Convolutional Network (FaGGCN), which combines feature learning and graph structure learning for multi-omics cancer subtyping. The model employs convolutional autoencoders to learn information-rich latent features, and patient survival information is further leveraged to select key features that are significantly associated with survival outcomes. The graph autoencoder fuses the key features with inter-omics similarity fusion matrices, enabling the model to learn a comprehensive sample network. Finally, the graph convolutional network integrates the key features while incorporating the sample network to precisely classify patients. Additionally, survival analysis, sensitivity analysis, and differential gene expression analysis highlight the interpretability of the FaGGCN model, as well as its ability to identify biomarkers suitable for clinical research. CONCLUSIONS: Experimental results show that our model achieves competitive performance across eight cancer datasets spanning four omics modalities, with generally improved classification performance and exploratory survival prediction results.
Mengke Guo, Xiucai Ye, Tetsuya Sakurai
BMC Bioinform.3
2026 Estimation of conditional average treatment effects on distributed confidential data
abstract
The estimation of conditional average treatment effects (CATEs) is an important topic in many scientific fields. CATEs can be estimated with high accuracy if data distributed across multiple parties are centralized. However, it is difficult to aggregate such data owing to confidentiality or privacy concerns. To address this issue, we propose data collaboration double machine learning, a method for estimating CATE models using privacy-preserving fusion data constructed from distributed sources, and evaluate its performance through simulations. We make three main contributions. First, our method enables estimation and testing of semi-parametric CATE models without iterative communication on distributed data, providing robustness to model mis-specification compared to parametric approaches. Second, it enables collaborative estimation across different time points and parties by accumulating a knowledge base. Third, our method performs as well as or better than existing methods in simulations using synthetic, semi-synthetic, and real-world datasets.
Yuji Kawamata, Ryoki Motai, Yukihiko Okada, Akira Imakura, Tetsuya Sakurai
Expert Syst. Appl.5
2026 SSDP-DCA: Enhancing Data Collaboration Analysis With Privacy Amplification Techniques in Cloud Environments
abstract
The rapid growth of cloud computing has enabled collaborative data analysis across distributed systems, yet ensuring privacy under stringent regulations remains a critical challenge. Data Collaboration Analysis (DCA) facilitates efficient analysis in such environments but struggles to balance robust privacy with model accuracy under strict constraints . Traditional Differential Privacy (DP) methods often degrade performance due to excessive noise (e.g.,$\epsilon = 2$). We propose SSDP-DCA, a novel framework that enhances DCA by integrating DP with privacy amplification techniques, leveraging the Shuffle Model for anonymization andpackage Subsampling (applied post-shuffle)for privacy amplification. Experimental results demonstrate SSDP-DCA's superiority over baselines, achieving high utility under a fixed global privacy budget, thus offering a robust solution for secure federated learning and healthcare applications.
Tingkai Sun, Xiucai Ye, Akira Imakura, Kazumasa Omote, Tetsuya Sakurai
IEEE Trans. Cloud Comput.5
2025 Wasserstein Gradient Flow over Variational Parameter Space for Variational Inference
abstract
Variational Inference (VI) optimizes varia- tional parameters to closely align a variational distribution with the true posterior, being ap- proached through vanilla gradient descent in black-box VI or natural-gradient descent in natural-gradient VI. In this work, we reframe VI as the optimization of an objective that concerns probability distributions defined over a variational parameter space. Subsequently, we propose Wasserstein gradient descent for solving this optimization, where black-box VI and natural-gradient VI can be interpreted as special cases of the proposed Wasserstein gradient descent. To enhance the efficiency of optimization, we develop practical methods for numerically solving the discrete gradient flows. We validate the effectiveness of the pro- posed methods through experiments on syn- thetic and real-world datasets, supplemented by theoretical analyses.
Dai Hai Nguyen, Tetsuya Sakurai, Hiroshi Mamitsuka
AISTATS2
2025 GSToxi: Gated Cross-Modal Modeling With Graph-Sequence Encoders for Peptide Toxicity Prediction
abstract
Peptide-based therapeutics hold great potential, yet their cytotoxicity remains a key challenge in drug development. Most existing toxicity prediction models rely solely on sequence information, often overlooking the fusion of multimodal submolecular patterns. We propose GSToxi, a multimodal deep learning framework that leverage sequence and molecular graph featurizer to enhance peptide toxicity prediction. A shared gating mechanism is employed to facilitate semantic alignment and cross-modal integration, while a contrastive regularization loss further optimizes latent-space consistency throughout the training process. Furthermore, GSToxi incorporates embeddings from pre-trained protein language models alongside low-level compositional priors, enabling the capture of both global contextual semantics and local structural features. Experimental results show that GSToxi outperforms state-of-the-art baselines across multiple evaluation metrics on an independent test set. Ablation studies underscore the critical contributions of each component, with the molecular graph encoder and pre-trained embeddings proving particularly impactful. This work offers a generalizable and robust framework for peptide toxicity prediction and provides valuable insights for future multimodal modeling of biological molecules.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai
BIBM5
2025 Deep Feature Learning for Multi-Omics Clustering Using Supervised Variational Autoencoders with Clinical Information
abstract
Identifying cancer subtypes through multi-omics clustering holds potential to advance cancer research by uncovering subtype-specific mechanisms. Most existing multi-omics clustering methods extract features in an unsupervised manner and integrate data separately, making it difficult to obtain discriminative feature representations and achieve effective data fusion. In this study, we propose a novel multi-omics clustering method that jointly performs feature extraction and data integration within a unified learning framework based on supervised variational autoencoders (VAEs). The proposed method first encodes each omics dataset using VAEs to extract omics-specific latent features, which are then integrated into a shared latent representation. Simultaneously, we utilize the shared latent space to reconstruct the original omics while incorporating patients’ clinical information through a classification task. By jointly optimizing the reconstruction loss and classification loss, the proposed method enables feature extraction that preserves the unique characteristics of each omics while leveraging clinical information to enhance the shared latent representation, thereby improving the integration of multi-omics data with clinical relevance. The shared latent representation is then used for clustering to identify cancer subtypes. Extensive experimental results demonstrate the robustness and effectiveness of the proposed method in identifying cancer subtypes across multiple cancer datasets.
Xiucai Ye, Tetsuya Sakurai
IJCNN3
2025 MOFormer: navigating the antimicrobial peptide design space with Pareto-based multi-objective transformer
abstract
Antimicrobial peptide (AMP) design through deep learning holds the potential to revolutionize antibiotic development. Despite recent progress in AMP generation, designing peptide antibiotics with multiple optimal properties remains a significant challenge. We present MOFormer, an advanced multi-objective AMP design pipeline capable of optimizing multiple AMP properties simultaneously. By leveraging a conditional Transformer, the model refines the AMP sequence-property landscape for efficient multi-objective generation. It also incorporates regularization techniques to maintain a highly structured space, enabling the sampling of precise and desirable candidates. Comparative analyses reveal that MOFormer achieves the optimal hypervolume in the multi-objective space, surpassing advanced methods in simultaneously maximizing antimicrobial activity (minimum inhibitory concentration) and minimizing hemolysis and toxicity, thereby yielding the most promising and desirable set of candidate peptides. When extended to a tri-objective scenario, MOFormer continues to exhibit remarkable optimization performance. Finally, we execute a hierarchical and rapid ranking of generated candidates based on Pareto fronts. We conducted a comprehensive validation of the physicochemical properties and target attributes of the candidates, while AlphaFold structure predictions revealed notably reliable predicted local distance difference test scores ranging from 70% to 87%. Our findings suggest that MOFormer holds potential to accelerate the discovery of efficacious peptide antibiotics by optimizing multi-objective trade-offs.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng
Briefings Bioinform.6
2025 PKAN: Leveraging Kolmogorov-Arnold Networks and Multi-Modal Learning for Peptide Prediction With Advanced Language Models
abstract
Peptides can offer highly specific biological activities, serving as essential mediators of intercellular signaling, which are critical for advancing precision medicine and drug development. Their primary structure can be depicted either as an amino acid sequence or as a chemical molecules consisting of atoms and chemical bonds. Large language models (LLMs) hold the potential to thoroughly elucidate the intricate intrinsic properties of peptides. Here we present the Peptide Kolmogorov-Arnold Network (PKAN), a framework leveraging multi-modal representations inspired by advanced language models for peptide activity and functionality prediction. Comparative experiments across tasks show that PKAN outperforms state-of-the-art models while maintaining a streamlined design with superior predictive capabilities. The multi-modal feature importance scoring, anchored in global structures and the significant marginal impacts of derived features on the model, coupled with intricate symbolic regression of specific activation functions, further demonstrates the robustness and precision of the PKAN framework in identifying and elucidating key determinants of peptide functionality. This work provides scientific evidence for investigating the complex mechanisms of peptide materials and supports the progression of peptide language paradigms in biology.
Li Wang 0145, Xiangzheng Fu, Xiucai Ye, Tetsuya Sakurai, Xiangxiang Zeng
IEEE J. Biomed. Health Informatics4
2024 Integrating Biological Language Processing and Memory Attention Model for Protein-DNA Binding Residue Prediction
abstract
Proteins are among the most important substances in the human body, and identifying protein-DNA binding sites is crucial for studying their interactions. Although traditional wet-lab methods can accurately identify these sites, they are time-consuming, labor-intensive, and expensive, making it challenging to keep pace with the rapid increase in protein sequence data. In this study, we propose the Memory Attention-Based Protein-DNA Binding Sites Prediction (MAPDB) model, which leverages multi-head Memory Attention for predicting protein-DNA binding sites. Our model employs a pre-trained embedding module to generate numerical representations of protein sequences, followed by a feature extraction module that uses Memory Attention to capture both intra-sequence and inter-sequence relationships. Extensive experiments on five benchmark datasets show that our model outperforms other state-of-the-art methods, especially in improving MCC scores. These results indicate that MAPDB effectively captures complex relationships within protein sequences, leading to more accurate predictions of protein-DNA binding residues.
Xiucai Ye, Tetsuya Sakurai
BIBM3
2024 Selecting interpretable features for cancer subtyping on multi-omics data
abstract
Analyzing multi-omics data is powerful in cancer research, offering essential insights for identifying distinct cancer subtypes. Due to the significant noise and redundancy in multi-omics data, effective feature extraction before data integration is crucial. However, most existing multi-omics clustering methods directly integrate different omics, which may lead to poor clustering results. In this study, we propose a novel multi-omics clustering method which selects interpretable features from different omics before data integration. The proposed method utilizes clinical information and SHAP values to extract interpretable features from different omics. Feature selection is then performed to select the most important interpretable features by grouping and ranking on the SHAP values. We construct similarity network for each omics based on the selected features, and then perform similarity network fusion to integrate the similarities across different omics. Finally, spectral clustering is applied to obtain the clustering result. We conduct experiments on five cancer datasets across three levels of omics to evaluate the proposed method. Experimental results demonstrate the superior performance of our proposed method in multi-omics clustering analysis for cancer subtyping.
Xiucai Ye, Tetsuya Sakurai
BIBM4
2024 Deciphering the Complex Characterization of Coding LncRNA
abstract
This research addresses the intricate nature of coding long non-coding RNAs (lncRNAs), challenging the traditional view of these molecules as merely non-coding elements. By analyzing sequence, physicochemical, and structural features, we have identified distinct characteristics of mRNA, coding lncRNA, and untranslated lncRNA. The CodLncPred model, developed using the XGBoost model, outperforms existing tools in classifying coding lncRNAs. Furthermore, our study evaluates the computational efficiency of various algorithms, including the cost of time and memory, underscoring the practical implications. Our findings offer a new perspective on coding lncRNAs, providing a robust framework for future exploration in genome biology and disease research.
Xiucai Ye, Tetsuya Sakurai
IJCNN3
2024 CyclePermea: Membrane Permeability Prediction of Cyclic Peptides with a Multi-Loss Fusion Network
abstract
Cyclic peptides, known for their unique ring-like structures, show considerable promise in therapeutic applications. Experimentally determining their permeability is time-consuming and labor-intensive. Hence, an efficient and rapid membrane permeability prediction model would greatly expedite the early-stage screening of cyclic peptide drugs. To meet this end, we proposed a novel deep learning model to predict membrane permeability of cyclic peptides, dubbed as CyclePermea. Remarkably, CyclePermea predicts membrane permeability using only the 1D sequence information of cyclic peptides, unlike previous works based on complex spatial descriptors and various physicochemical properties. It incorporates a peptide encoder based on a pre-trained BERT architecture. We also introduced two auxiliary loss functions designed to enhance the model’s comprehension of cyclic peptides’ distinctive characteristics. The first, termed ’Constraint Contrastive Learning Loss’, aims to mitigate the challenge of feature clustering. The second, ’Cyclization Site Prediction Loss’, is proposed to facilitate the model’s recognition of the unique spatial structure inherent in cyclic peptides. Through extensive experiments, CyclePermea demonstrated superior performance over baseline models in the benchmark dataset, both in in-distribution settings and simulated out-of-distribution settings. We hope that CyclePermea would contribute in accelerating the early screening of cyclic peptide drugs in the future.
Yangyang Chen 0006, Xiucai Ye, Tetsuya Sakurai
IJCNN4
2024 Integrated convolution and self-attention for improving peptide toxicity prediction
abstract
MOTIVATION: Peptides are promising agents for the treatment of a variety of diseases due to their specificity and efficacy. However, the development of peptide-based drugs is often hindered by the potential toxicity of peptides, which poses a significant barrier to their clinical application. Traditional experimental methods for evaluating peptide toxicity are time-consuming and costly, making the development process inefficient. Therefore, there is an urgent need for computational tools specifically designed to predict peptide toxicity accurately and rapidly, facilitating the identification of safe peptide candidates for drug development. RESULTS: We provide here a novel computational approach, CAPTP, which leverages the power of convolutional and self-attention to enhance the prediction of peptide toxicity from amino acid sequences. CAPTP demonstrates outstanding performance, achieving a Matthews correlation coefficient of approximately 0.82 in both cross-validation settings and on independent test datasets. This performance surpasses that of existing state-of-the-art peptide toxicity predictors. Importantly, CAPTP maintains its robustness and generalizability even when dealing with data imbalances. Further analysis by CAPTP reveals that certain sequential patterns, particularly in the head and central regions of peptides, are crucial in determining their toxicity. This insight can significantly inform and guide the design of safer peptide drugs. AVAILABILITY AND IMPLEMENTATION: The source code for CAPTP is freely available at https://github.com/jiaoshihu/CAPTP.
Shihu Jiao, Xiucai Ye, Tetsuya Sakurai, Quan Zou 0001
Bioinform.3
2024 Collaborative causal inference on distributed data
abstract
In recent years, the development of technologies for causal inference with privacy preservation of distributed data has gained considerable attention. Many existing methods for distributed data focus on resolving the lack of subjects (samples) and can only reduce random errors in estimating treatment effects. In this study, we propose a data collaboration quasi-experiment (DC-QE) that resolves the lack of both subjects and covariates, reducing random errors and biases in the estimation. Our method involves constructing dimensionality-reduced intermediate representations from private data from local parties, sharing intermediate representations instead of private data for privacy preservation, estimating propensity scores from the shared intermediate representations, and finally, estimating the treatment effects from propensity scores. Through numerical experiments on both artificial and real-world data, we confirm that our method leads to better estimation results than individual analyses. While dimensionality reduction loses some information in the private data and causes performance degradation, we observe that sharing intermediate representations with many parties to resolve the lack of subjects and covariates sufficiently improves performance to overcome the degradation caused by dimensionality reduction. Although external validity is not necessarily guaranteed, our results suggest that DC-QE is a promising method. With the widespread use of our method, intermediate representations can be published as open data to help researchers find causalities and accumulate a knowledge base.
Yuji Kawamata, Ryoki Motai, Yukihiko Okada, Akira Imakura, Tetsuya Sakurai
Expert Syst. Appl.5
2024 An interpretable deep learning model predicts RNA-small molecule binding sites
Wen-Yu Xi, Ruheng Wang, Li Wang 0145, Xiucai Ye, Tetsuya Sakurai
Future Gener. Comput. Syst.6
2024 Moreau-Yoshida variational transport: a general framework for solving regularized distributional optimization problems
Dai Hai Nguyen, Tetsuya Sakurai
Mach. Learn.2
2023 Multi-omics clustering based on interpretable and discriminative features for cancer subtyping
abstract
Recent advances in multi-omics databases have enabled biomedical researchers to explore complex cancer systems across hierarchical biological levels. Although there are numerous multi-omics clustering methods, most of them directly integrate heterogeneous features of different omics which may include redundancy or noise and lead to poor clustering results. In this paper, we propose a novel multi-omics clustering method for cancer subtyping which extracts interpretable and discriminative features from different omics before data integration. The proposed method utilizes the clinical information of each omics to supervise the process of extracting interpretable and discriminative features based on SHAP (SHapley Additive exPlanation) values. The shared nearest neighbor-based approach is then applied to calculate the similarity matrix of the extracted features. Finally, we integrate the similarity matrices of different omics and apply spectral clustering on the integrated similarity matrix to obtain the clustering result. Experimental results conducted on four different cancer datasets on three levels of omics demonstrate the superior performance of the proposed method in comparison to the existing multi-omics clustering methods.
Xiucai Ye, Tetsuya Sakurai
BIBM3
2023 Multi-view Network Embedding with Structure and Semantic Contrastive Learning
abstract
Multi-view network embedding aims to learn low-dimensional representation vectors for nodes while preserving multiple relationships between nodes. It can substantially reduce downstream network analysis tasks’ time and space complexity. Although previous works have achieved great performance, they suffer from two limitations: (1) they only preserve the network structure and ignore the semantic level information; (2) they only focus on intra-view signals and ignore the powerful influence of inter-view signals. These limitations highlight the need for more comprehensive approaches to multi-view network embedding that can effectively capture the structure and semantic information, as well as the influence of inter-view signals. A new framework, Multi-view Network Embedding with Structure and Semantic Contrastive Learning (MNE-SSCL), is proposed to address these limitations. It can learn high-quality low-dimensional node embeddings in both intra-veiw and interview, while preserving the structure and semantic information simultaneously. Extensive experiments on three real datasets show that MNE-SSCL outperforms the state-of-the-art methods.
Yifan Shang, Xiucai Ye, Tetsuya Sakurai
ICME3
2023 Common and Unique Features Learning in Multi-view Network Embedding
abstract
Network embedding is a powerful representation learning method for graph data, using the learned low-dimensional compact vectors as node features, which are widely used in various tasks, such as link prediction, node clustering, and classification. Compared with traditional network analysis methods, network embedding reduces computational complexity and improves analysis efficiency. Although previous work has achieved outstanding performance, it faces challenges in multi-view network embedding containing multi-type node relations. Since multi-view networks share a node set but different edges, different networks not only have common information but also have their unique information. To simultaneously capture multi-view networks' common and unique information, we propose a new framework, Common and Unique Features Learning for Multi-view Network Embedding (CU-MNE), to integrate multi-type node relations. In this paper, we propose an inter-view contrastive objective to ensure the consistency of the common features of the same node in a different view and an inter-feature contrastive objective to capture the association between the common and unique features of each network node that can learn high-quality node embeddings. Extensive experiments on three real datasets show that CU-MNE outperforms the state-of-the-art methods.
Yifan Shang, Xiucai Ye, Tetsuya Sakurai
IJCNN3
2023 SiameseCPP: a sequence-based Siamese network to predict cell-penetrating peptides by contrastive learning
abstract
BACKGROUND: Cell-penetrating peptides (CPPs) have received considerable attention as a means of transporting pharmacologically active molecules into living cells without damaging the cell membrane, and thus hold great promise as future therapeutics. Recently, several machine learning-based algorithms have been proposed for predicting CPPs. However, most existing predictive methods do not consider the agreement (disagreement) between similar (dissimilar) CPPs and depend heavily on expert knowledge-based handcrafted features. RESULTS: In this study, we present SiameseCPP, a novel deep learning framework for automated CPPs prediction. SiameseCPP learns discriminative representations of CPPs based on a well-pretrained model and a Siamese neural network consisting of a transformer and gated recurrent units. Contrastive learning is used for the first time to build a CPP predictive model. Comprehensive experiments demonstrate that our proposed SiameseCPP is superior to existing baseline models for predicting CPPs. Moreover, SiameseCPP also achieves good performance on other functional peptide datasets, exhibiting satisfactory generalization ability.
Lesong Wei, Xiucai Ye, Saisai Teng, Zhongshen Li, Junru Jin, Min Jae Kim, Tetsuya Sakurai, Li-Zhen Cui 0001, Balachandran Manavalan, Leyi Wei
Briefings Bioinform.9
2023 Adaptive learning embedding features to improve the predictive performance of SARS-CoV-2 phosphorylation sites
abstract
MOTIVATION: The rapid and extensive transmission of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) has led to an unprecedented global health emergency, affecting millions of people and causing an immense socioeconomic impact. The identification of SARS-CoV-2 phosphorylation sites plays an important role in unraveling the complex molecular mechanisms behind infection and the resulting alterations in host cell pathways. However, currently available prediction tools for identifying these sites lack accuracy and efficiency. RESULTS: In this study, we presented a comprehensive biological function analysis of SARS-CoV-2 infection in a clonal human lung epithelial A549 cell, revealing dramatic changes in protein phosphorylation pathways in host cells. Moreover, a novel deep learning predictor called PSPred-ALE is specifically designed to identify phosphorylation sites in human host cells that are infected with SARS-CoV-2. The key idea of PSPred-ALE lies in the use of a self-adaptive learning embedding algorithm, which enables the automatic extraction of context sequential features from protein sequences. In addition, the tool uses multihead attention module that enables the capturing of global information, further improving the accuracy of predictions. Comparative analysis of features demonstrated that the self-adaptive learning embedding features are superior to hand-crafted statistical features in capturing discriminative sequence information. Benchmarking comparison shows that PSPred-ALE outperforms the state-of-the-art prediction tools and achieves robust performance. Therefore, the proposed model can effectively identify phosphorylation sites assistant the biomedical scientists in understanding the mechanism of phosphorylation in SARS-CoV-2 infection. AVAILABILITY AND IMPLEMENTATION: PSPred-ALE is available at https://github.com/jiaoshihu/PSPred-ALE and Zenodo (https://doi.org/10.5281/zenodo.8330277).
Shihu Jiao, Xiucai Ye, Chunyan Ao, Tetsuya Sakurai, Quan Zou 0001, Lei Xu 0047
Bioinform.4
2023 Top to random shuffles on colored permutations
Fumihiko Nakano, Taizo Sadahiro, Tetsuya Sakurai
Discret. Appl. Math.3
2023 Another use of SMOTE for interpretable data collaboration analysis
abstract
Recently, data collaboration (DC) analysis has been developed for privacy-preserving integrated analysis across multiple institutions. DC analysis centralizes individually constructed dimensionality-reduced intermediate representations and realizes integrated analysis via collaboration representations without sharing the original data. To construct the collaboration representations, each institution generates and shares a shareable anchor dataset and centralizes its intermediate representation. Although, random anchor dataset functions well for DC analysis in general, using an anchor dataset whose distribution is close to that of the raw dataset is expected to improve the recognition performance, particularly for the interpretable DC analysis. Based on an extension of the synthetic minority over-sampling technique (SMOTE), this study proposes an anchor data construction technique to improve the recognition performance without increasing the risk of data leakage. Numerical results demonstrate the efficiency of the proposed SMOTE-based method over the existing anchor data constructions for artificial and real-world datasets. Specifically, the proposed method achieves 6, 4, and 36 percentage point performance improvements regarding NMI, ACC and essential feature selection, respectively, over existing methods for an income dataset. The proposed method provides another use of SMOTE not for imbalanced data classifications but for a key technology of privacy-preserving integrated analysis.
Akira Imakura, Masateru Kihira, Yukihiko Okada, Tetsuya Sakurai
Expert Syst. Appl.4
2023 Distortion-free PCA on sample space for highly variable gene detection from single-cell RNA-seq data
Momo Matsuda, Yasunori Futamura, Xiucai Ye, Tetsuya Sakurai
Frontiers Comput. Sci.4
2023 LSEC: Large-scale spectral ensemble clustering
abstract
A fundamental problem in machine learning is ensemble clustering, that is, combining multiple base clusterings to obtain improved clustering result. However, most of the existing methods are unsuitable for large-scale ensemble clustering tasks owing to efficiency bottlenecks. In this paper, we propose a large-scale spectral ensemble clustering (LSEC) method to balance efficiency and effectiveness. In LSEC, a large-scale spectral clustering-based efficient ensemble generation framework is designed to generate various base clusterings with low computational complexity. Thereafter, all the base clusterings are combined using a bipartite graph partition-based consensus function to obtain improved consensus clustering results. The LSEC method achieves a lower computational complexity than most existing ensemble clustering methods. Experiments conducted on ten large-scale datasets demonstrate the efficiency and effectiveness of the LSEC method. The MATLAB code of the proposed method and experimental datasets are available at https://github.com/Li-Hongmin/MyPaperWithCode.
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
Intell. Data Anal.4
2023 DC-COX: Data collaboration Cox proportional hazards model for privacy-preserving survival analysis on multiple parties
abstract
The demand for the privacy-preserving survival analysis of medical data integrated from multiple institutions or countries has been increased. However, sharing the original medical data is difficult because of privacy concerns, and even if it could be achieved, we have to pay huge costs for cross-institutional or cross-border communications. To tackle these difficulties of privacy-preserving survival analysis on multiple parties, this study proposes a novel data collaboration Cox proportional hazards (DC-COX) model based on a data collaboration framework for horizontally and vertically partitioned data. By integrating dimensionality-reduced intermediate representations instead of the original data, DC-COX obtains a privacy-preserving survival analysis without iterative cross-institutional communications or huge computational costs. DC-COX enables each local party to obtain an approximation of the maximum likelihood model parameter, the corresponding statistic, such as the p-value, and survival curves for subgroups. Based on a bootstrap technique, we introduce a dimensionality reduction method to improve the efficiency of DC-COX. Numerical experiments demonstrate that DC-COX can compute a model parameter and the corresponding statistics with higher performance than the local party analysis. Particularly, DC-COX demonstrates outstanding performance in essential feature selection based on the p-value compared with the existing methods including the federated learning-based method.
Akira Imakura, Ryoya Tsunoda, Rina Kagawa, Kunihiro Yamagata, Tetsuya Sakurai
J. Biomed. Informatics5
2023 Mirror variational transport: a particle-based algorithm for distributional optimization on constrained domains
Dai Hai Nguyen, Tetsuya Sakurai
Mach. Learn.2
2022 Multiview network embedding for drug-target Interactions prediction by consistent and complementary information preserving
abstract
Accurate prediction of drug-target interactions (DTIs) can reduce the cost and time of drug repositioning and drug discovery. Many current methods integrate information from multiple data sources of drug and target to improve DTIs prediction accuracy. However, these methods do not consider the complex relationship between different data sources. In this study, we propose a novel computational framework, called MccDTI, to predict the potential DTIs by multiview network embedding, which can integrate the heterogenous information of drug and target. MccDTI learns high-quality low-dimensional representations of drug and target by preserving the consistent and complementary information between multiview networks. Then MccDTI adopts matrix completion scheme for DTIs prediction based on drug and target representations. Experimental results on two datasets show that the prediction accuracy of MccDTI outperforms four state-of-the-art methods for DTIs prediction. Moreover, literature verification for DTIs prediction shows that MccDTI can predict the reliable potential DTIs. These results indicate that MccDTI can provide a powerful tool to predict new DTIs and accelerate drug discovery. The code and data are available at: https://github.com/ShangCS/MccDTI.
Yifan Shang, Xiucai Ye, Yasunori Futamura, Liang Yu 0002, Tetsuya Sakurai
Briefings Bioinform.5
2022 iLoc-miRNA: extracellular/intracellular miRNA prediction using deep BiLSTM with attention mechanism
abstract
The location of microRNAs (miRNAs) in cells determines their function in regulation activity. Studies have shown that miRNAs are stable in the extracellular environment that mediates cell-to-cell communication and are located in the intracellular region that responds to cellular stress and environmental stimuli. Though in situ detection techniques of miRNAs have made great contributions to the study of the localization and distribution of miRNAs, miRNA subcellular localization and their role are still in progress. Recently, some machine learning-based algorithms have been designed for miRNA subcellular location prediction, but their performance is still far from satisfactory. Here, we present a new data partitioning strategy that categorizes functionally similar locations for the precise and instructive prediction of miRNA subcellular location in Homo sapiens. To characterize the localization signals, we adopted one-hot encoding with post padding to represent the whole miRNA sequences, and proposed a deep bidirectional long short-term memory with the multi-head self-attention algorithm to model. The algorithm showed high selectivity in distinguishing extracellular miRNAs from intracellular miRNAs. Moreover, a series of motif analyses were performed to explore the mechanism of miRNA subcellular localization. To improve the convenience of the model, a user-friendly web server named iLoc-miRNA was established (http://iLoc-miRNA.lin-group.cn/).
Zhao-Yue Zhang 0002, Lin Ning 0002, Xiucai Ye, Yasunori Futamura, Tetsuya Sakurai, Hao Lin 0001
Briefings Bioinform.6
2022 NerLTR-DTA: drug-target binding affinity prediction based on neighbor relationship and learning to rank
abstract
MOTIVATION: Drug-target interaction prediction plays an important role in new drug discovery and drug repurposing. Binding affinity indicates the strength of drug-target interactions. Predicting drug-target binding affinity is expected to provide promising candidates for biologists, which can effectively reduce the workload of wet laboratory experiments and speed up the entire process of drug research. Given that, numerous new proteins are sequenced and compounds are synthesized, several improved computational methods have been proposed for such predictions, but there are still some challenges. (i) Many methods only discuss and implement one application scenario, they focus on drug repurposing and ignore the discovery of new drugs and targets. (ii) Many methods do not consider the priority order of proteins (or drugs) related to each target drug (or protein). Therefore, it is necessary to develop a comprehensive method that can be used in multiple scenarios and focuses on candidate order. RESULTS: In this study, we propose a method called NerLTR-DTA that uses the neighbor relationship of similarity and sharing to extract features, and applies a ranking framework with regression attributes to predict affinity values and priority order of query drug (or query target) and its related proteins (or compounds). It is worth noting that using the characteristics of learning to rank to set different queries can smartly realize the multi-scenario application of the method, including the discovery of new drugs and new targets. Experimental results on two commonly used datasets show that NerLTR-DTA outperforms some state-of-the-art competing methods. NerLTR-DTA achieves excellent performance in all application scenarios mentioned in this study, and the rm(test)2 values guarantee such excellent performance is not obtained by chance. Moreover, it can be concluded that NerLTR-DTA can provide accurate ranking lists for the relevant results of most queries through the statistics of the association relationship of each query drug (or query protein). In general, NerLTR-DTA is a powerful tool for predicting drug-target associations and can contribute to new drug discovery and drug repurposing. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented in Python and Java. Source codes and datasets are available at https://github.com/RUXIAOQING964914140/NerLTR-DTA.
Xiaoqing Ru, Xiucai Ye, Tetsuya Sakurai, Quan Zou 0001
Bioinform.3
2022 ToxIBTL: prediction of peptide toxicity based on information bottleneck and transfer learning
abstract
MOTIVATION: Recently, peptides have emerged as a promising class of pharmaceuticals for various diseases treatment poised between traditional small molecule drugs and therapeutic proteins. However, one of the key bottlenecks preventing them from therapeutic peptides is their toxicity toward human cells, and few available algorithms for predicting toxicity are specially designed for short-length peptides. RESULTS: We present ToxIBTL, a novel deep learning framework by utilizing the information bottleneck principle and transfer learning to predict the toxicity of peptides as well as proteins. Specifically, we use evolutionary information and physicochemical properties of peptide sequences and integrate the information bottleneck principle into a feature representation learning scheme, by which relevant information is retained and the redundant information is minimized in the obtained features. Moreover, transfer learning is introduced to transfer the common knowledge contained in proteins to peptides, which aims to improve the feature representation capability. Extensive experimental results demonstrate that ToxIBTL not only achieves a higher prediction performance than state-of-the-art methods on the peptide dataset, but also has a competitive performance on the protein dataset. Furthermore, a user-friendly online web server is established as the implementation of the proposed ToxIBTL. AVAILABILITY AND IMPLEMENTATION: The proposed ToxIBTL and data can be freely accessible at http://server.wei-group.net/ToxIBTL. Our source code is available at https://github.com/WLYLab/ToxIBTL. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lesong Wei, Xiucai Ye, Tetsuya Sakurai, Zengchao Mu, Leyi Wei
Bioinform.3
2022 Divide-and-conquer based large-scale spectral clustering
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
Neurocomputing4
2022 Sequential reinforcement active feature learning for gene signature identification in renal cell carcinoma
Meng Huang 0003, Xiucai Ye, Akira Imakura, Tetsuya Sakurai
J. Biomed. Informatics4
2021 Collaborative Novelty Detection for Distributed Data by a Probabilistic Method
abstract
Novelty detection, which detects anomalies based on a training dataset consisting of only the normal data, is an important task in several applications. In addition, in the real world, there may be situations where data is owned by multiple parties in a distributed manner but cannot be shared with each other due to privacy and confidentiality requirements. Therefore, how to develop distributed novelty detection while preserving privacy is essential. To address this challenge, we propose a probabilistic collaborative method that allows distributed novelty detection for multiple parties without sharing the original data. The proposed method constructs a collaborative kernel based on a collaborative data analysis framework, by which intermediate representations are generated from each party and shared for collaborative novelty detection. Numerical experiments demonstrate that the proposed method obtains better performance compared with the individual novelty detection in the local party.
Akira Imakura, Xiucai Ye, Tetsuya Sakurai
ACML3
2021 Efficient Contour Integral-based Eigenvalue Computation Using an Iterative Linear Solver with Shift-Invert Preconditioning
abstract
Contour integral-based (CI) eigenvalue solvers are one of the efficient and robust approaches for sparse eigenvalue problems. They have attracted attention owing to their inherent parallelism. For implementing a CI eigensolver, the inner linear systems arising in the algorithm need to be solved using an efficient method. One widely-used method is to use a sparse direct linear solver provided by a well-established numerical library; it is numerically robust and presents good load balancing of parallel execution of the CI eigensolver. However, owing to high total computational and memory cost, the performance of the direct solver approach is suboptimal. In this study, we propose an alternative method that utilizes a block Krylov iterative linear solver and shift-invert preconditioning that can take advantage of the shift-invariance of the block Krylov subspace. Our approach adaptively sets a preconditioning parameter according to the number of parallel processes to reduce the iteration counts. Several numerical examples confirm that our method outperforms the direct solver approach.
Yasunori Futamura, Tetsuya Sakurai
HPC Asia2
2021 Efficient Implementation of a Dimensionality Reduction Method Using a Complex Moment-Based Subspace
abstract
Dimensionality reduction methods are widely used for processing data efficiently. Recently Imakura et al. proposed a novel dimensionality reduction method using a complex moment-based subspace. Their method can use more eigenvectors than the existing matrix trace optimization-based methods which explains its reported higher precision. However, the computational complexity is also higher than that of the existing methods, in particular for the nonlinear kernel version. To reduce the computational complexity, we propose a practical parallel implementation of the method by introducing the Nyström approximation. We evaluate the parallel performance of our implementation using the Oakforest-PACS supercomputer.
Takahiro Yano, Yasunori Futamura, Akira Imakura, Tetsuya Sakurai
HPC Asia4
2021 Spectral Clustering Joint Deep Embedding Learning by Autoencoder
abstract
Spectral clustering has become one of the most popular clustering methods due to its superior performance compared to the traditional clustering methods. However, the performance of spectral clustering would be limited by complex data, such as a huge number of samples and high dimensionality. To address this problem, some existing methods apply deep learning to learn the lower-dimensional representations, spectral clustering is then applied to the representations. Different from the existing methods that separate the two stages of feature representation learning and spectral clustering, in this paper, we propose Spectral Clustering Joint Deep Embedding (SCJDE), a method that simultaneously learns the feature representations and the spectral embedding of spectral clustering via a deep autoencoder. Moreover, a sparsity constraint is imposed to generate better spectral embedding for spectral clustering. Finally,$k$-means is performed on the spectral embedding to obtain clustering result. The proposed method can learn a good spectral embedding for spectral clustering by deep learning to obtain better clustering results. The experimental results on both synthetic and real-world datasets demonstrate the effectiveness of the proposed method.
Xiucai Ye, Chunhao Wang, Akira Imakura, Tetsuya Sakurai
IJCNN4
2021 Application of learning to rank in bioinformatics tasks
abstract
Over the past decades, learning to rank (LTR) algorithms have been gradually applied to bioinformatics. Such methods have shown significant advantages in multiple research tasks in this field. Therefore, it is necessary to summarize and discuss the application of these algorithms so that these algorithms are convenient and contribute to bioinformatics. In this paper, the characteristics of LTR algorithms and their strengths over other types of algorithms are analyzed based on the application of multiple perspectives in bioinformatics. Finally, the paper further discusses the shortcomings of the LTR algorithms, the methods and means to better use the algorithms and some open problems that currently exist.
Xiaoqing Ru, Xiucai Ye, Tetsuya Sakurai, Quan Zou 0001
Briefings Bioinform.3
2021 ATSE: a peptide toxicity predictor by exploiting structural and evolutionary information based on graph neural network and attention mechanism
abstract
MOTIVATION: Peptides have recently emerged as promising therapeutic agents against various diseases. For both research and safety regulation purposes, it is of high importance to develop computational methods to accurately predict the potential toxicity of peptides within the vast number of candidate peptides. RESULTS: In this study, we proposed ATSE, a peptide toxicity predictor by exploiting structural and evolutionary information based on graph neural networks and attention mechanism. More specifically, it consists of four modules: (i) a sequence processing module for converting peptide sequences to molecular graphs and evolutionary profiles, (ii) a feature extraction module designed to learn discriminative features from graph structural information and evolutionary information, (iii) an attention module employed to optimize the features and (iv) an output module determining a peptide as toxic or non-toxic, using optimized features from the attention module. CONCLUSION: Comparative studies demonstrate that the proposed ATSE significantly outperforms all other competing methods. We found that structural information is complementary to the evolutionary information, effectively improving the predictive performance. Importantly, the data-driven features learned by ATSE can be interpreted and visualized, providing additional information for further analysis. Moreover, we present a user-friendly online computational platform that implements the proposed ATSE, which is now available at http://server.malab.cn/ATSE. We expect that it can be a powerful and useful tool for researchers of interest.
Lesong Wei, Xiucai Ye, Yuyang Xue, Tetsuya Sakurai, Leyi Wei
Briefings Bioinform.4
2021 Interpretable collaborative data analysis on distributed data
abstract
This paper proposes an interpretable non-model sharing collaborative data analysis method as a federated learning system, which is an emerging technology for analyzing distributed data. Analyzing distributed data is essential in many applications, such as medicine, finance, and manufacturing, due to privacy and confidentiality concerns. In addition, interpretability of the obtained model plays an important role in the practical applications of federated learning systems. By centralizing intermediate representations , which are individually constructed by each party, the proposed method obtains an interpretable model, achieving collaborative analysis without revealing the individual data and learning models distributed between local parties. Numerical experiments indicate that the proposed method achieves better recognition performance than individual analysis and comparable performance to centralized analysis for both artificial and real-world problems.
Akira Imakura, Hiroaki Inaba, Yukihiko Okada, Tetsuya Sakurai
Expert Syst. Appl.4
2020 A Training Difficulty Schedule for Effective Search of Meta-Heuristic Design
abstract
In the context of optimization problems, the performance of an algorithm depends on the problem. It is difficult to know a priori what algorithm (and what parameters) will perform best on a new problem. For this reason, we previously proposed a framework that uses grammatical evolution to automatically generate swarm intelligence algorithms given a training problem. However, we observed two issues that affected the results of the framework. The first issue was that sometimes the training problems are too easy, and any candidate algorithm could solve it or too difficult that no candidate algorithm could solve them. The second issue was the presence of parameters in the grammar, which causes a significant increase in the search space. In this work, we addressed those issues by investigating three training schedules in which the problems start easy and get harder over time. We also investigated whether numerical parameters should be part of the grammar. We compared these training schedules to the previous one and compared the performance of the generated algorithms against the traditional algorithms, which are DE, PSO, and CS. We found that gradually increasing the difficulty of the training problem produced algorithms that could solve more testing instances than training only in 10-D. The results suggest that a step-by-step increase in difficulty is a better approach overall. We also found that including parameters in the grammar resulted in algorithms on par with the traditional meta-heuristics. Besides, as expected, our results show that removing parameters from the grammar exhibit the worst overall performance. However, interestingly it could solve most of the testing instances within the given testing budget.
Jair Pereira Junior, Claus Aranha, Tetsuya Sakurai
CEC3
2020 Ensemble Learning for Spectral Clustering
abstract
Ensemble clustering has attracted much attention in machine learning and data mining for the high performance in the task of clustering. Spectral clustering is one of the most popular clustering methods and has superior performance compared with the traditional clustering methods. Existing ensemble clustering methods usually directly use the clustering results of the base clustering algorithms for ensemble learning, which cannot make good use of the intrinsic data structures explored by the graph Laplacians in spectral clustering, thus cannot obtain the desired clustering result. In this paper, we propose a new ensemble learning method for spectral clustering-based clustering algorithms. Instead of directly using the clustering results obtained from each base spectral clustering algorithm, the proposed method learns a robust presentation of graph Laplacian by ensemble learning from the spectral embedding of each base spectral clustering algorithm. Finally, the proposed method applies k-means on the spectral embedding obtain from the learned graph Laplacian to get clusters. Experimental results on both synthetic and real-world datasets show that the proposed method outperforms other existing ensemble clustering methods.
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
ICDM4
2020 Hubness-based Sampling Method for Nyström Spectral Clustering
abstract
Nyström method is widely used for spectral clustering to obtain low-rank approximations of a large matrix. Sampling is crucial to Nyström method, since selecting the representative sample points that can reflect the data structure is important for obtaining good approximation results. To improve the performance of Nyström based spectral clustering, in this paper, we propose a new sampling method by considering the hubness score of sample points. The data points with the high hubness scores, i.e., appearing frequently in the nearest neighbor lists of other data points, have high probabilities to be selected as the sample points. Taking advantage of the topological property of hubs (i.e., data points with high hubness score), the selected sampling points have close relationships with other data points, thus the proposed method is able to achieve scalable and accurate clustering results. We further design fast computation methods, i.e., local hubness approximated methods, to speed up the sampling process. Experimental results on both synthetic and real-world data sets show that the proposed method not only achieves good performance, but also outperforms other sampling methods for Nyström based spectral clustering.
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
IJCNN4
2020 Accelerating the Backpropagation Algorithm by Using NMF-Based Method on Deep Neural Networks
Suhyeon Baek, Akira Imakura, Tetsuya Sakurai, Ichiro Kataoka
PKAW3
2020 Collaborative Data Analysis: Non-model Sharing-Type Machine Learning for Distributed Data
Akira Imakura, Xiucai Ye, Tetsuya Sakurai
PKAW3
2020 Multiclass spectral feature scaling method for dimensionality reduction
abstract
Irregular features disrupt the desired classification. In this paper, we consider aggressively modifying scales of features in the original space according to the label information to form well-separated clusters in low-dimensional space. The proposed method exploits spectral clustering to derive scaling factors that are used to modify the features. Specifically, we reformulate the Laplacian eigenproblem of the spectral clustering as an eigenproblem of a linear matrix pencil whose eigenvector has the scaling factors. Numerical experiments show that the proposed method outperforms well-established supervised dimensionality reduction methods for toy problems with more samples than features and real-world problems with more features than samples.
Momo Matsuda, Keiichi Morikuni, Akira Imakura, Xiucai Ye, Tetsuya Sakurai
Intell. Data Anal.5
2020 An oversampling framework for imbalanced classification based on Laplacian eigenmaps
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
Neurocomputing4
2019 Complex Moment-Based Supervised Eigenmap for Dimensionality Reduction
abstract
Dimensionality reduction methods that project highdimensional data to a low-dimensional space by matrix trace optimization are widely used for clustering and classification. The matrix trace optimization problem leads to an eigenvalue problem for a low-dimensional subspace construction, preserving certain properties of the original data. However, most of the existing methods use only a few eigenvectors to construct the low-dimensional space, which may lead to a loss of useful information for achieving successful classification. Herein, to overcome the deficiency of the information loss, we propose a novel complex moment-based supervised eigenmap including multiple eigenvectors for dimensionality reduction. Furthermore, the proposed method provides a general formulation for matrix trace optimization methods to incorporate with ridge regression, which models the linear dependency between covariate variables and univariate labels. To reduce the computational complexity, we also propose an efficient and parallel implementation of the proposed method. Numerical experiments indicate that the proposed method is competitive compared with the existing dimensionality reduction methods for the recognition performance. Additionally, the proposed method exhibits high parallel efficiency.
Akira Imakura, Momo Matsuda, Xiucai Ye, Tetsuya Sakurai
AAAI4
2019 Distributed Collaborative Feature Selection Based on Intermediate Representation
abstract
Feature selection is an efficient dimensionality reduction technique for artificial intelligence and machine learning. Many feature selection methods learn the data structure to select the most discriminative features for distinguishing different classes. However, the data is sometimes distributed in multiple parties and sharing the original data is difficult due to the privacy requirement. As a result, the data in one party may be lack of useful information to learn the most discriminative features. In this paper, we propose a novel distributed method which allows collaborative feature selection for multiple parties without revealing their original data. In the proposed method, each party finds the intermediate representations from the original data, and shares the intermediate representations for collaborative feature selection. Based on the shared intermediate representations, the original data from multiple parties are transformed to the same low dimensional space. The feature ranking of the original data is learned by imposing row sparsity on the transformation matrix simultaneously. Experimental results on real-world datasets demonstrate the effectiveness of the proposed method.
Xiucai Ye, Akira Imakura, Tetsuya Sakurai
IJCAI4
2018 Spectral Feature Scaling Method for Supervised Dimensionality Reduction
abstract
Spectral dimensionality reduction methods enable linear separations of complex data with high-dimensional features in a reduced space. However, these methods do not always give the desired results due to irregularities or uncertainties of the data. Thus, we consider aggressively modifying the scales of the features to obtain the desired classification. Using prior knowledge on the labels of partial samples to specify the Fiedler vector, we formulate an eigenvalue problem of a linear matrix pencil whose eigenvector has the feature scaling factors. The resulting factors can modify the features of entire samples to form clusters in the reduced space, according to the known labels. In this study, we propose new dimensionality reduction methods supervised using the feature scaling associated with the spectral clustering. Numerical experiments show that the proposed methods outperform well-established supervised methods for toy problems with more samples than features, and are more robust regarding clustering than existing methods. Also, the proposed methods outperform existing methods regarding classification for real-world problems with more features than samples of gene expression profiles of cancer diseases. Furthermore, the feature scaling tends to improve the clustering and classification accuracies of existing unsupervised methods, as the proportion of training data increases.
Momo Matsuda, Keiichi Morikuni, Tetsuya Sakurai
IJCAI3
2018 Experimental Analysis of the Tournament Size on Genetic Algorithms
abstract
We perform an experimental study about the effect of the tournament size parameter from the Tournament Selection operator. Tournament Selection is a classic operator for Genetic Algorithms and Genetic Programming. It is simple to implement and has only one control parameter, the tournament size. Even though it is commonly used, most practitioners still rely on rules of thumb when choosing the tournament size. For example, almost all works in the past 15 years use a value of 2 for the tournament size, with little reasoning behind that choice. To understand the role of the tournament size, we run a real-valued GA on 24 BBOB problems with 10, 20 and 40 dimensions. We also vary the crossover operator and the generational policy of the GA. For each combination of the above factors we observe how the quality of the final solution changes with the tournament size. Our findings do not support the indiscriminate use of tournament size 2, and recommend a more careful set up of this parameter.
Yuri Cossich Lavinas, Claus Aranha, Tetsuya Sakurai, Marcelo Ladeira
SMC3
2018 Spectral clustering with adaptive similarity measure in Kernel space
abstract
The similarity measure for complex data may not precisely reflect the true data structure, which leads to suboptimal clustering performance for spectral clustering. In this paper, we propose a novel spectral clustering method which measures the similarity of data points based on the adaptive neighb orhood in Kernel space. In Kernel space, by assigning the adaptive and optimal neighbors for each data point based on the local structure, the proposed method learns a sparse matrix as the similarity matrix for spectral clustering. The proposed method is able to explore the underlying similarity relationships between data points, and is robust to the complex data. To validate the efficacy of the proposed method, we perform experiments on both synthetic and real datasets in comparison with some existing spectral clustering methods. The experimental results demonstrate that the proposed method obtains quite promising clustering performance.
Xiucai Ye, Tetsuya Sakurai
Intell. Data Anal.2
2018 Parallel Implementation of the Nonlinear Semi-NMF Based Alternating Optimization Method for Deep Neural Networks
Akira Imakura, Yuto Inoue, Tetsuya Sakurai, Yasunori Futamura
Neural Process. Lett.3
2018 Block SS-CAA: A complex moment-based parallel nonlinear eigensolver using the block communication-avoiding Arnoldi procedure
Akira Imakura, Tetsuya Sakurai
Parallel Comput.2
2017 Efficient and scalable calculation of complex band structure using Sakurai-Sugiura method
abstract
Complex band structures (CBSs) are useful to characterize the static and dynamical electronic properties of materials. Despite the intensive developments, the first-principles calculation of CBS for over several hundred atoms are still computationally demanding. We here propose an efficient and scalable computational method to calculate CBSs. The basic idea is to express the Kohn-Sham equation of the real-space grid scheme as a quadratic eigenvalue problem and compute only the solutions which are necessary to construct the CBS by Sakurai-Sugiura method. The serial performance of the proposed method shows a significant advantage in both run-time and memory usage compared to the conventional method. Furthermore, owing to the hierarchical parallelism in Sakurai-Sugiura method and the domain-decomposition technique for real-space grids, we can achieve an excellent scalability in the CBS calculation of a boron and nitrogen doped carbon nanotube consisting of more than 10,000 atoms using 2,048 nodes (139,264 cores) of Oakforest-PACS.
Shigeru Iwase, Yasunori Futamura, Akira Imakura, Tetsuya Sakurai, Tomoya Ono
SC4
2016 Spectral clustering and discriminant analysis for unsupervised feature selection
Xiucai Ye, Kaiyang Ji, Tetsuya Sakurai
ESANN3
2016 Alternating Optimization Method Based on Nonnegative Matrix Factorizations for Deep Neural Networks
Tetsuya Sakurai, Akira Imakura, Yuto Inoue, Yasunori Futamura
ICONIP (4)1
2015 Spectral clustering using robust similarity measure based on closeness of shared Nearest Neighbors
abstract
Spectral clustering has become one of the main clustering methods and has a wide range of applications. Similarity measure is crucial to correct cluster separation for spectral clustering. Many existing spectral clustering algorithms typically measure similarity based on the undirected k-Nearest Neighbor (kNN) graph or Gaussian kernel function, which can not reveal the real clusters of not well-separated data sets. In this paper, we propose a novel algorithm called Spectral Clustering based on Shared Nearest Neighbors (SC-SNN) to improve the clustering quality of not well-separated data sets. Instead of using distance for the similarity measure, the proposed SC-SNN algorithm measures the similarity by considering the closeness of shared nearest neighbors in the directed kNN graph, which is able to explore the underlying similarity relationships between data points and is robust to the not well-separated data sets. Moreover, SC-SNN has only one parameter, k, and is less sensitive than the spectral clustering algorithms based on the undirected kNN graph. The proposed SC-SNN algorithm is evaluated by using both synthetic and real-world data sets. The experimental results demonstrate that SC-SNN not only achieves good performance, but also outperforms the traditional spectral clustering algorithms.
Xiucai Ye, Tetsuya Sakurai
IJCNN2
2014 Correlations between predicted protein disorder and post-translational modifications in plants
abstract
MOTIVATION: Protein structural research in plants lags behind that in animal and bacterial species. This lag concerns both the structural analysis of individual proteins and the proteome-wide characterization of structure-related properties. Until now, no systematic study concerning the relationships between protein disorder and multiple post-translational modifications (PTMs) in plants has been presented. RESULTS: In this work, we calculated the global degree of intrinsic disorder in the complete proteomes of eight typical monocotyledonous and dicotyledonous plant species. We further predicted multiple sites for phosphorylation, glycosylation, acetylation and methylation and examined the correlations of protein disorder with the presence of the predicted PTM sites. It was found that phosphorylation, acetylation and O-glycosylation displayed a clear preference for occurrence in disordered regions of plant proteins. In contrast, methylation tended to avoid disordered sequence, whereas N-glycosylation did not show a universal structural preference in monocotyledonous and dicotyledonous plants. In addition, the analysis performed revealed significant differences between the integral characteristics of monocot and dicot proteomes. They included elevated disorder degree, increased rate of O-glycosylation and R-methylation, decreased rate of N-glycosylation, K-acetylation and K-methylation in monocotyledonous plant species, as compared with dicotyledonous species. Altogether, our study provides the most compelling evidence so far for the connection between protein disorder and multiple PTMs in plants. CONTACT: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Atsushi Kurotani, Alexander A. Tokmakov, Yutaka Kuroda, Yasuo Fukami, Kazuo Shinozaki, Tetsuya Sakurai
Bioinform.6
2013 Performance comparison of parallel eigensolvers based on a contour integral method and a Lanczos method
Ichitaro Yamazaki, Hiroto Tadano, Tetsuya Sakurai, Tsutomu Ikegami
Parallel Comput.3
2010 sORF finder: a program package to identify small open reading frames with high coding potential
abstract
SUMMARY: sORF finder is a program package for identifying small open reading frames (sORFs) with high-coding potential. This application allows the identification of coding sORFs according to the nucleotide composition bias among coding sequences and the potential functional constraint at the amino acid level through evaluation of synonymous and non-synonymous substitution rates. AVAILABILITY: Online tools and source codes are freely available at http://evolver.psc.riken.jp/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kousuke Hanada, Kenji Akiyama, Tetsuya Sakurai, Tetsuro Toyoda, Kazuo Shinozaki, Shin-Han Shiu
Bioinform.3
2010 LegumeTFDB: an integrative database of Glycine max, Lotus japonicus and Medicago truncatula transcription factors
abstract
UNLABELLED: We have established a database named LegumeTFDB to provide access to transcription factor (TF) repertoires of three major legume species: soybean (Glycine max), Lotus japonicus and Medicago truncatula. LegumeTFDB integrates unique information for each TF gene and family, including sequence features, gene promoters, domain alignments, gene ontology (GO) assignment and sequence comparison data derived from comparative analysis with TFs found within legumes, in Arabidopsis, rice and poplar as well as with proteins in NCBI nr and UniProt. We also analyzed the promoter regions for all of the TFs to identify all types of cis-motifs provided by the PLACE database. Additionally, we supply hyperlinks to make available expression data of 2411 soybean TF genes. LegumeTFDB provides an important user-friendly public resource for comparative genomics and understanding of transcriptional regulation in agriculturally important legumes. AVAILABILITY: http://legumetfdb.psc.riken.jp/. SUPPLEMENTARY INFORMATION: Supplementary data available at Bioinformatics online.
Keiichi Mochida, Takuhiro Yoshida, Tetsuya Sakurai, Kazuko Yamaguchi-Shinozaki, Kazuo Shinozaki, Lam-Son Phan Tran
Bioinform.3
2008 A parallel method for large sparse generalized eigenvalue problems using a GridRPC system
Tetsuya Sakurai, Yoshihisa Kodaki, Hiroto Tadano, Daisuke Takahashi, Mitsuhisa Sato, Umpei Nagashima
Future Gener. Comput. Syst.1
2006 Performance Improvement by Data Management Layer in a Grid RPC System
Yoshiaki Aida, Yoshihiro Nakajima, Mitsuhisa Sato, Tetsuya Sakurai, Daisuke Takahashi, Taisuke Boku
GPC4
1996 A Methodology of Parsing Mathematical Notation for Mathematical Computation
abstract
Mathematical notationhas not been parsed effectively up to the present.In order to use mathematical notation on computer directly, and thus to develop many potential applications and to determine if common mathematical notation can be parsed by computer, we inquire into the omission mechanism of parentheses and operators, study the extension mechanism of mathematical notation, and establish a methodology to translate a two-dimensional mathematical expression into a textual functional meaning representation.This methodology is a combination of different formalization methods denoted by a defined box language, various context-sensitive grammars written in a defined metalanguage, and knowledge-based parsers.
Tetsuya Sakurai, Hiroshi Sugiura, Tatsuo Torii
ISSAC2