VLDB 2026 Research / reviewers in the wild / expert
Akira Imakura
dblp:143/6633
· DBLP profile ↗
27ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-4994-2499ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 6 first-author · 10 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Estimation of conditional average treatment effects on distributed confidential dataabstractThe estimation of conditional average treatment effects (CATEs) is an important topic in many scientific fields. CATEs can be estimated with high accuracy if data distributed across multiple parties are centralized. However, it is difficult to aggregate such data owing to confidentiality or privacy concerns. To address this issue, we propose data collaboration double machine learning, a method for estimating CATE models using privacy-preserving fusion data constructed from distributed sources, and evaluate its performance through simulations. We make three main contributions. First, our method enables estimation and testing of semi-parametric CATE models without iterative communication on distributed data, providing robustness to model mis-specification compared to parametric approaches. Second, it enables collaborative estimation across different time points and parties by accumulating a knowledge base. Third, our method performs as well as or better than existing methods in simulations using synthetic, semi-synthetic, and real-world datasets. Yuji Kawamata, Ryoki Motai, Yukihiko Okada, Akira Imakura, Tetsuya Sakurai |
Expert Syst. Appl. | 4 |
| 2026 | New solutions based on the generalized eigenvalue problem for the data collaboration analysisabstractThis paper is concerned with the data collaboration (DC) analysis, a privacy-preserving method for analyzing decentralized datasets held by multiple parties. In this method, privacy-preserving intermediate representations of original datasets are collected from multiple parties and then converted into collaboration representations for collaborative data analysis. However, conventional methods for creating collaboration representations suffer from several challenges; namely, the optimization problem being considered is not well defined, and the process of solving it is very difficult to understand. We thus propose a new solution for creating high-quality collaboration representations for the DC analysis. Specifically, we formulate a revised optimization problem for creating collaboration representations and then transform this optimization problem into a generalized eigenvalue problem. We also propose a reduction of the generalized eigenvalue problem to a singular value decomposition through the QR decomposition. Computational experiments using publicly available datasets demonstrate that our method can outperform the conventional methods for the DC analysis in terms of both prediction accuracy and computational efficiency. • Privacy-preserving data collaboration for analyzing decentralized datasets. • Generalized eigenvalue problem for high-quality collaboration representations. • Reduction of the generalized eigenvalue problem to a singular value decomposition. • Superiority of our method evaluated through computational experiments. Yuta Kawakami, Yuichi Takano, Akira Imakura |
Inf. Sci. | 3 |
| 2026 | SSDP-DCA: Enhancing Data Collaboration Analysis With Privacy Amplification Techniques in Cloud EnvironmentsabstractThe rapid growth of cloud computing has enabled collaborative data analysis across distributed systems, yet ensuring privacy under stringent regulations remains a critical challenge. Data Collaboration Analysis (DCA) facilitates efficient analysis in such environments but struggles to balance robust privacy with model accuracy under strict constraints . Traditional Differential Privacy (DP) methods often degrade performance due to excessive noise (e.g.,$\epsilon = 2$). We propose SSDP-DCA, a novel framework that enhances DCA by integrating DP with privacy amplification techniques, leveraging the Shuffle Model for anonymization andpackage Subsampling (applied post-shuffle)for privacy amplification. Experimental results demonstrate SSDP-DCA's superiority over baselines, achieving high utility under a fixed global privacy budget, thus offering a robust solution for secure federated learning and healthcare applications. Tingkai Sun, Xiucai Ye, Akira Imakura, Kazumasa Omote, Tetsuya Sakurai |
IEEE Trans. Cloud Comput. | 3 |
| 2025 | DC-CP: Data Collaboration Conformal Prediction from Model Training to Conformal Prediction in Single-Round Communication While Preserving Privacy
Tomoru Nakayama, Yuji Kawamata, Akira Imakura, Yukihiko Okada |
IEEE Big Data | 3 |
| 2024 | Collaborative causal inference on distributed dataabstractIn recent years, the development of technologies for causal inference with privacy preservation of distributed data has gained considerable attention. Many existing methods for distributed data focus on resolving the lack of subjects (samples) and can only reduce random errors in estimating treatment effects. In this study, we propose a data collaboration quasi-experiment (DC-QE) that resolves the lack of both subjects and covariates, reducing random errors and biases in the estimation. Our method involves constructing dimensionality-reduced intermediate representations from private data from local parties, sharing intermediate representations instead of private data for privacy preservation, estimating propensity scores from the shared intermediate representations, and finally, estimating the treatment effects from propensity scores. Through numerical experiments on both artificial and real-world data, we confirm that our method leads to better estimation results than individual analyses. While dimensionality reduction loses some information in the private data and causes performance degradation, we observe that sharing intermediate representations with many parties to resolve the lack of subjects and covariates sufficiently improves performance to overcome the degradation caused by dimensionality reduction. Although external validity is not necessarily guaranteed, our results suggest that DC-QE is a promising method. With the widespread use of our method, intermediate representations can be published as open data to help researchers find causalities and accumulate a knowledge base. Yuji Kawamata, Ryoki Motai, Yukihiko Okada, Akira Imakura, Tetsuya Sakurai |
Expert Syst. Appl. | 4 |
| 2023 | Another use of SMOTE for interpretable data collaboration analysisabstractRecently, data collaboration (DC) analysis has been developed for privacy-preserving integrated analysis across multiple institutions. DC analysis centralizes individually constructed dimensionality-reduced intermediate representations and realizes integrated analysis via collaboration representations without sharing the original data. To construct the collaboration representations, each institution generates and shares a shareable anchor dataset and centralizes its intermediate representation. Although, random anchor dataset functions well for DC analysis in general, using an anchor dataset whose distribution is close to that of the raw dataset is expected to improve the recognition performance, particularly for the interpretable DC analysis. Based on an extension of the synthetic minority over-sampling technique (SMOTE), this study proposes an anchor data construction technique to improve the recognition performance without increasing the risk of data leakage. Numerical results demonstrate the efficiency of the proposed SMOTE-based method over the existing anchor data constructions for artificial and real-world datasets. Specifically, the proposed method achieves 6, 4, and 36 percentage point performance improvements regarding NMI, ACC and essential feature selection, respectively, over existing methods for an income dataset. The proposed method provides another use of SMOTE not for imbalanced data classifications but for a key technology of privacy-preserving integrated analysis. Akira Imakura, Masateru Kihira, Yukihiko Okada, Tetsuya Sakurai |
Expert Syst. Appl. | 1 |
| 2023 | LSEC: Large-scale spectral ensemble clusteringabstractA fundamental problem in machine learning is ensemble clustering, that is, combining multiple base clusterings to obtain improved clustering result. However, most of the existing methods are unsuitable for large-scale ensemble clustering tasks owing to efficiency bottlenecks. In this paper, we propose a large-scale spectral ensemble clustering (LSEC) method to balance efficiency and effectiveness. In LSEC, a large-scale spectral clustering-based efficient ensemble generation framework is designed to generate various base clusterings with low computational complexity. Thereafter, all the base clusterings are combined using a bipartite graph partition-based consensus function to obtain improved consensus clustering results. The LSEC method achieves a lower computational complexity than most existing ensemble clustering methods. Experiments conducted on ten large-scale datasets demonstrate the efficiency and effectiveness of the LSEC method. The MATLAB code of the proposed method and experimental datasets are available at https://github.com/Li-Hongmin/MyPaperWithCode. Xiucai Ye, Akira Imakura, Tetsuya Sakurai |
Intell. Data Anal. | 3 |
| 2023 | DC-COX: Data collaboration Cox proportional hazards model for privacy-preserving survival analysis on multiple partiesabstractThe demand for the privacy-preserving survival analysis of medical data integrated from multiple institutions or countries has been increased. However, sharing the original medical data is difficult because of privacy concerns, and even if it could be achieved, we have to pay huge costs for cross-institutional or cross-border communications. To tackle these difficulties of privacy-preserving survival analysis on multiple parties, this study proposes a novel data collaboration Cox proportional hazards (DC-COX) model based on a data collaboration framework for horizontally and vertically partitioned data. By integrating dimensionality-reduced intermediate representations instead of the original data, DC-COX obtains a privacy-preserving survival analysis without iterative cross-institutional communications or huge computational costs. DC-COX enables each local party to obtain an approximation of the maximum likelihood model parameter, the corresponding statistic, such as the p-value, and survival curves for subgroups. Based on a bootstrap technique, we introduce a dimensionality reduction method to improve the efficiency of DC-COX. Numerical experiments demonstrate that DC-COX can compute a model parameter and the corresponding statistics with higher performance than the local party analysis. Particularly, DC-COX demonstrates outstanding performance in essential feature selection based on the p-value compared with the existing methods including the federated learning-based method. Akira Imakura, Ryoya Tsunoda, Rina Kagawa, Kunihiro Yamagata, Tetsuya Sakurai |
J. Biomed. Informatics | 1 |
| 2022 | Fast Algorithm for Low-rank Tensor Completion in Delay-embedded SpaceabstractTensor completion using multiway delay-embedding transform (MDT) (or Hankelization) suffers from the large memory requirement and high computational cost in spite of its high potentiality for the image modeling. Recent studies have shown high completion performance with a relatively small window size, but experiments with large window sizes require huge amount of memory and cannot be easily calculated. In this study, we address this serious computational issue, and propose its fast and efficient algorithm. Key techniques of the proposed method are based on two properties: (1) the signal after MDT can be diagonalized by Fourier transform, (2) an inverse MDT can be represented as a convolutional form. To use the properties, we modify MDT-Tucker [26], a method using Tucker decomposition with MDT, and introducing the fast and efficient algorithm. Our experiments show more than 100 times acceleration while maintaining high accuracy, and to realize the computation with large window size. Ryuki Yamamoto, Hidekata Hontani, Akira Imakura, Tatsuya Yokota |
CVPR | 3 |
| 2022 | Divide-and-conquer based large-scale spectral clustering
Xiucai Ye, Akira Imakura, Tetsuya Sakurai |
Neurocomputing | 3 |
| 2022 | Sequential reinforcement active feature learning for gene signature identification in renal cell carcinoma
Meng Huang 0003, Xiucai Ye, Akira Imakura, Tetsuya Sakurai |
J. Biomed. Informatics | 3 |
| 2021 | Collaborative Novelty Detection for Distributed Data by a Probabilistic MethodabstractNovelty detection, which detects anomalies based on a training dataset consisting of only the normal data, is an important task in several applications. In addition, in the real world, there may be situations where data is owned by multiple parties in a distributed manner but cannot be shared with each other due to privacy and confidentiality requirements. Therefore, how to develop distributed novelty detection while preserving privacy is essential. To address this challenge, we propose a probabilistic collaborative method that allows distributed novelty detection for multiple parties without sharing the original data. The proposed method constructs a collaborative kernel based on a collaborative data analysis framework, by which intermediate representations are generated from each party and shared for collaborative novelty detection. Numerical experiments demonstrate that the proposed method obtains better performance compared with the individual novelty detection in the local party. Akira Imakura, Xiucai Ye, Tetsuya Sakurai |
ACML | 1 |
| 2021 | Efficient Implementation of a Dimensionality Reduction Method Using a Complex Moment-Based SubspaceabstractDimensionality reduction methods are widely used for processing data efficiently. Recently Imakura et al. proposed a novel dimensionality reduction method using a complex moment-based subspace. Their method can use more eigenvectors than the existing matrix trace optimization-based methods which explains its reported higher precision. However, the computational complexity is also higher than that of the existing methods, in particular for the nonlinear kernel version. To reduce the computational complexity, we propose a practical parallel implementation of the method by introducing the Nyström approximation. We evaluate the parallel performance of our implementation using the Oakforest-PACS supercomputer. Takahiro Yano, Yasunori Futamura, Akira Imakura, Tetsuya Sakurai |
HPC Asia | 3 |
| 2021 | Spectral Clustering Joint Deep Embedding Learning by AutoencoderabstractSpectral clustering has become one of the most popular clustering methods due to its superior performance compared to the traditional clustering methods. However, the performance of spectral clustering would be limited by complex data, such as a huge number of samples and high dimensionality. To address this problem, some existing methods apply deep learning to learn the lower-dimensional representations, spectral clustering is then applied to the representations. Different from the existing methods that separate the two stages of feature representation learning and spectral clustering, in this paper, we propose Spectral Clustering Joint Deep Embedding (SCJDE), a method that simultaneously learns the feature representations and the spectral embedding of spectral clustering via a deep autoencoder. Moreover, a sparsity constraint is imposed to generate better spectral embedding for spectral clustering. Finally,$k$-means is performed on the spectral embedding to obtain clustering result. The proposed method can learn a good spectral embedding for spectral clustering by deep learning to obtain better clustering results. The experimental results on both synthetic and real-world datasets demonstrate the effectiveness of the proposed method. Xiucai Ye, Chunhao Wang, Akira Imakura, Tetsuya Sakurai |
IJCNN | 3 |
| 2021 | Interpretable collaborative data analysis on distributed dataabstractThis paper proposes an interpretable non-model sharing collaborative data analysis method as a federated learning system, which is an emerging technology for analyzing distributed data. Analyzing distributed data is essential in many applications, such as medicine, finance, and manufacturing, due to privacy and confidentiality concerns. In addition, interpretability of the obtained model plays an important role in the practical applications of federated learning systems. By centralizing intermediate representations , which are individually constructed by each party, the proposed method obtains an interpretable model, achieving collaborative analysis without revealing the individual data and learning models distributed between local parties. Numerical experiments indicate that the proposed method achieves better recognition performance than individual analysis and comparable performance to centralized analysis for both artificial and real-world problems. Akira Imakura, Hiroaki Inaba, Yukihiko Okada, Tetsuya Sakurai |
Expert Syst. Appl. | 1 |
| 2020 | Ensemble Learning for Spectral ClusteringabstractEnsemble clustering has attracted much attention in machine learning and data mining for the high performance in the task of clustering. Spectral clustering is one of the most popular clustering methods and has superior performance compared with the traditional clustering methods. Existing ensemble clustering methods usually directly use the clustering results of the base clustering algorithms for ensemble learning, which cannot make good use of the intrinsic data structures explored by the graph Laplacians in spectral clustering, thus cannot obtain the desired clustering result. In this paper, we propose a new ensemble learning method for spectral clustering-based clustering algorithms. Instead of directly using the clustering results obtained from each base spectral clustering algorithm, the proposed method learns a robust presentation of graph Laplacian by ensemble learning from the spectral embedding of each base spectral clustering algorithm. Finally, the proposed method applies k-means on the spectral embedding obtain from the learned graph Laplacian to get clusters. Experimental results on both synthetic and real-world datasets show that the proposed method outperforms other existing ensemble clustering methods. Xiucai Ye, Akira Imakura, Tetsuya Sakurai |
ICDM | 3 |
| 2020 | Hubness-based Sampling Method for Nyström Spectral ClusteringabstractNyström method is widely used for spectral clustering to obtain low-rank approximations of a large matrix. Sampling is crucial to Nyström method, since selecting the representative sample points that can reflect the data structure is important for obtaining good approximation results. To improve the performance of Nyström based spectral clustering, in this paper, we propose a new sampling method by considering the hubness score of sample points. The data points with the high hubness scores, i.e., appearing frequently in the nearest neighbor lists of other data points, have high probabilities to be selected as the sample points. Taking advantage of the topological property of hubs (i.e., data points with high hubness score), the selected sampling points have close relationships with other data points, thus the proposed method is able to achieve scalable and accurate clustering results. We further design fast computation methods, i.e., local hubness approximated methods, to speed up the sampling process. Experimental results on both synthetic and real-world data sets show that the proposed method not only achieves good performance, but also outperforms other sampling methods for Nyström based spectral clustering. Xiucai Ye, Akira Imakura, Tetsuya Sakurai |
IJCNN | 3 |
| 2020 | Accelerating the Backpropagation Algorithm by Using NMF-Based Method on Deep Neural Networks
Suhyeon Baek, Akira Imakura, Tetsuya Sakurai, Ichiro Kataoka |
PKAW | 2 |
| 2020 | Collaborative Data Analysis: Non-model Sharing-Type Machine Learning for Distributed Data
Akira Imakura, Xiucai Ye, Tetsuya Sakurai |
PKAW | 1 |
| 2020 | Multiclass spectral feature scaling method for dimensionality reductionabstractIrregular features disrupt the desired classification. In this paper, we consider aggressively modifying scales of features in the original space according to the label information to form well-separated clusters in low-dimensional space. The proposed method exploits spectral clustering to derive scaling factors that are used to modify the features. Specifically, we reformulate the Laplacian eigenproblem of the spectral clustering as an eigenproblem of a linear matrix pencil whose eigenvector has the scaling factors. Numerical experiments show that the proposed method outperforms well-established supervised dimensionality reduction methods for toy problems with more samples than features and real-world problems with more features than samples. Momo Matsuda, Keiichi Morikuni, Akira Imakura, Xiucai Ye, Tetsuya Sakurai |
Intell. Data Anal. | 3 |
| 2020 | An oversampling framework for imbalanced classification based on Laplacian eigenmaps
Xiucai Ye, Akira Imakura, Tetsuya Sakurai |
Neurocomputing | 3 |
| 2019 | Complex Moment-Based Supervised Eigenmap for Dimensionality ReductionabstractDimensionality reduction methods that project highdimensional data to a low-dimensional space by matrix trace optimization are widely used for clustering and classification. The matrix trace optimization problem leads to an eigenvalue problem for a low-dimensional subspace construction, preserving certain properties of the original data. However, most of the existing methods use only a few eigenvectors to construct the low-dimensional space, which may lead to a loss of useful information for achieving successful classification. Herein, to overcome the deficiency of the information loss, we propose a novel complex moment-based supervised eigenmap including multiple eigenvectors for dimensionality reduction. Furthermore, the proposed method provides a general formulation for matrix trace optimization methods to incorporate with ridge regression, which models the linear dependency between covariate variables and univariate labels. To reduce the computational complexity, we also propose an efficient and parallel implementation of the proposed method. Numerical experiments indicate that the proposed method is competitive compared with the existing dimensionality reduction methods for the recognition performance. Additionally, the proposed method exhibits high parallel efficiency. Akira Imakura, Momo Matsuda, Xiucai Ye, Tetsuya Sakurai |
AAAI | 1 |
| 2019 | Distributed Collaborative Feature Selection Based on Intermediate RepresentationabstractFeature selection is an efficient dimensionality reduction technique for artificial intelligence and machine learning. Many feature selection methods learn the data structure to select the most discriminative features for distinguishing different classes. However, the data is sometimes distributed in multiple parties and sharing the original data is difficult due to the privacy requirement. As a result, the data in one party may be lack of useful information to learn the most discriminative features. In this paper, we propose a novel distributed method which allows collaborative feature selection for multiple parties without revealing their original data. In the proposed method, each party finds the intermediate representations from the original data, and shares the intermediate representations for collaborative feature selection. Based on the shared intermediate representations, the original data from multiple parties are transformed to the same low dimensional space. The feature ranking of the original data is learned by imposing row sparsity on the transformation matrix simultaneously. Experimental results on real-world datasets demonstrate the effectiveness of the proposed method. Xiucai Ye, Akira Imakura, Tetsuya Sakurai |
IJCAI | 3 |
| 2018 | Parallel Implementation of the Nonlinear Semi-NMF Based Alternating Optimization Method for Deep Neural Networks
Akira Imakura, Yuto Inoue, Tetsuya Sakurai, Yasunori Futamura |
Neural Process. Lett. | 1 |
| 2018 | Block SS-CAA: A complex moment-based parallel nonlinear eigensolver using the block communication-avoiding Arnoldi procedure
Akira Imakura, Tetsuya Sakurai |
Parallel Comput. | 1 |
| 2017 | Efficient and scalable calculation of complex band structure using Sakurai-Sugiura methodabstractComplex band structures (CBSs) are useful to characterize the static and dynamical electronic properties of materials. Despite the intensive developments, the first-principles calculation of CBS for over several hundred atoms are still computationally demanding. We here propose an efficient and scalable computational method to calculate CBSs. The basic idea is to express the Kohn-Sham equation of the real-space grid scheme as a quadratic eigenvalue problem and compute only the solutions which are necessary to construct the CBS by Sakurai-Sugiura method. The serial performance of the proposed method shows a significant advantage in both run-time and memory usage compared to the conventional method. Furthermore, owing to the hierarchical parallelism in Sakurai-Sugiura method and the domain-decomposition technique for real-space grids, we can achieve an excellent scalability in the CBS calculation of a boron and nitrogen doped carbon nanotube consisting of more than 10,000 atoms using 2,048 nodes (139,264 cores) of Oakforest-PACS. Shigeru Iwase, Yasunori Futamura, Akira Imakura, Tetsuya Sakurai, Tomoya Ono |
SC | 3 |
| 2016 | Alternating Optimization Method Based on Nonnegative Matrix Factorizations for Deep Neural Networks
Tetsuya Sakurai, Akira Imakura, Yuto Inoue, Yasunori Futamura |
ICONIP (4) | 2 |