EDBT 2026 Demo / reviewers in the wild / expert
Shohei Shimizu
dblp:07/4552
· DBLP profile ↗
47ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-1931-0733ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 8 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discovering Linear Non-Gaussian Models for All Categories of Missing Data (Student Abstract)abstractCausal discovery is the task of learning causal models, encoding causal relationships, from a source of information, such as a dataset containing observational data. While many algorithms have been developed to discover causal models under varied sets of assumptions, the case in which the dataset is affected by missing data remains significantly underexplored. Naively applying standard causal discovery algorithms to listwise, test-wise, or regression-wise deleted datasets, or imputing the missing data, can introduce spurious associations between variables and bias function estimation in functional causal models. This issue arises when the data is missing at random or not at random. It ultimately invalidates the theoretical guarantees of these algorithms and prevents finding the true underlying causal model, even in the large-sample limit. An established family of causal models is the Linear Non-Gaussian Acyclic Model (LiNGAM), which assumes linear functional relationships and non-Gaussian independent noise terms. We propose a new causal discovery algorithm for LiNGAM, capable of recovering the underlying causal structure and providing unbiased estimates of the model’s parameters, even when the data is affected by MNAR missingness. Matteo Ceriscioli, Shohei Shimizu, Karthika Mohan |
AAAI | 2 |
| 2026 | I-CAM-UV: Integrating Causal Graphs over Non-Identical Variable Sets Using Causal Additive Models with Unobserved VariablesabstractCausal discovery from observational data is a fundamental tool in various fields of science. While existing approaches are typically designed for a single dataset, we often need to handle multiple datasets with non-identical variable sets in practice. One straightforward approach is to estimate a causal graph from each dataset and construct a single causal graph by overlapping. However, this approach identifies limited causal relationships because unobserved variables in each dataset can be confounders, and some variable pairs may be unobserved in any dataset. To address this issue, we leverage Causal Additive Models with Unobserved Variables (CAM-UV) that provide causal graphs having information related to unobserved variables. We show that the ground truth causal graph has structural consistency with the information of CAM-UV on each dataset. As a result, we propose an approach named I-CAM-UV to integrate CAM-UV results by enumerating all consistent causal graphs. We also provide an efficient combinatorial search algorithm and demonstrate the usefulness of I-CAM-UV against existing methods. Hirofumi Suzuki, Kentaro Kanamori, Takuya Takagi, Thong Pham, Takashi Nicholas Maeda, Shohei Shimizu |
AAAI | 6 |
| 2025 | Causal-discovery-based root-cause analysis and its application in time-series prediction error diagnosisabstractRecent rapid advancements of machine learning have greatly enhanced the accuracy of prediction models, but most models remain "black boxes", making prediction error diagnosis challenging, especially with outliers. This lack of transparency hinders trust and reliability in industrial applications. Heuristic attribution methods, while helpful, often fail to capture true causal relationships, leading to inaccurate error attributions. Various root-cause analysis methods have been developed using Shapley values, yet they typically require predefined causal graphs, limiting their applicability for prediction errors in machine learning models. To address these limitations, we introduce the Causal-Discovery-based Root-Cause Analysis (CD-RCA) method that estimates causal relationships between the prediction error and the explanatory variables, without needing a pre-defined causal graph. By simulating synthetic error data, CD-RCA can identify variable contributions to outliers in prediction errors by Shapley values. Extensive experiments show CD-RCA outperforms current heuristic attribution methods. Hiroshi Yokoyama, Ryusei Shingaki, Kaneharu Nishino, Shohei Shimizu, Thong Pham |
IJCNN | 4 |
| 2025 | Causal models and prediction in cell line perturbation experimentsabstractIn cell line perturbation experiments, a collection of cells is perturbed with external agents and responses such as protein expression measured. Due to cost constraints, only a small fraction of all possible perturbations can be tested in vitro. This has led to the development of computational models that can predict cellular responses to perturbations in silico. A central challenge for these models is to predict the effect of new, previously untested perturbations that were not used in the training data. Here we propose causal structural equations for modeling how perturbations effect cells. From this model, we derive two estimators for predicting responses: a Linear Regression (LR) estimator and a causal structure learning estimator that we term Causal Structure Regression (CSR). The CSR estimator requires more assumptions than LR, but can predict the effects of drugs that were not applied in the training data. Next we present Cellbox, a recently proposed system of ordinary differential equations (ODEs) based model that obtained the best prediction performance on a Melanoma cell line perturbation data set (Yuan et al. in Cell Syst 12:128-140, 2021). We derive analytic results that show a close connection between CSR and Cellbox, providing a new causal interpretation for the Cellbox model. We compare LR and CSR/Cellbox in simulations, highlighting the strengths and weaknesses of the two approaches. Finally we compare the performance of LR and CSR/Cellbox on the benchmark Melanoma data set. We find that the LR model has comparable or slightly better performance than Cellbox. James P. Long, Shohei Shimizu, Thong Pham, Kim-Anh Do |
BMC Bioinform. | 3 |
| 2025 | Information Theoretic Learning-Enhanced Dual-Generative Adversarial Networks With Causal Representation for Robust OOD GeneralizationabstractRecently, machine/deep learning techniques are achieving remarkable success in a variety of intelligent control and management systems, promising to change the future of artificial intelligence (AI) scenarios. However, they still suffer from some intractable difficulty or limitations for model training, such as the out-of-distribution (OOD) issue, in modern smart manufacturing or intelligent transportation systems (ITSs). In this study, we newly design and introduce a deep generative model framework, which seamlessly incorporates the information theoretic learning (ITL) and causal representation learning (CRL) in a dual-generative adversarial network (Dual-GAN) architecture, aiming to enhance the robust OOD generalization in modern machine learning (ML) paradigms. In particular, an ITL- and CRL-enhanced Dual-GAN (ITCRL-DGAN) model is presented, which includes an autoencoder with CRL (AE-CRL) structure to aid the dual-adversarial training with causality-inspired feature representations and a Dual-GAN structure to improve the data augmentation in both feature and data levels. Following a newly designed feature separation strategy, a causal graph is built and improved based on the information theory, which can enhance the causally related factors among the separated core features and further enrich the feature representation with the counterfactual features via interventions based on the refined causal relationships. The ITL is incorporated to improve the extraction of low-dimensional feature representations and learn the optimized causal representations based on the idea of "information flow." A dual-adversarial training mechanism is then developed, which not only enables the generator to expand the boundary of feature distribution in accordance with the optimized feature representation from AE-CRL, but also allows the discriminator to further verify and improve the quality of the augmented data for OOD generalization. Experiment and evaluation results based on an open-source dataset demonstrate the outstanding learning efficiency and classification performance of our proposed model for robust OOD generalization in modern smart applications compared with three baseline methods. Xiaokang Zhou, Xuzhe Zheng, Tian Shu, Wei Liang 0006, Kevin I-Kai Wang, Lianyong Qi, Shohei Shimizu, Qun Jin |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Counterfactual Explanations of Black-box Machine Learning Models using Causal Discovery with Applications to Credit RatingabstractExplainable artificial intelligence (XAI) has helped elucidate the internal mechanisms of machine learning algorithms, bolstering their reliability by demonstrating the basis of their predictions. Several XAI models consider causal relationships to explain models by examining the input-output relationships of prediction models and the dependencies between features. The majority of these models have been based their explanations on counterfactual probabilities, assuming that the causal graph is known. However, this assumption complicates the application of such models to real data, given that the causal relationships between features are unknown in most cases. Thus, this study proposed a novel XAI framework that relaxed the constraint that the causal graph is known. This framework leveraged counterfactual probabilities and additional prior information on causal structure, facilitating the integration of a causal graph estimated through causal discovery methods and a black-box classification model. Furthermore, explanatory scores were estimated based on counterfactual probabilities. Numerical experiments conducted employing artificial data confirmed the possibility of estimating the explanatory score more accurately than in the absence of a causal graph. Finally, as an application to real data, we constructed a classification model of credit ratings assigned by Shiga Bank, Shiga prefecture, Japan. We demonstrated the effectiveness of the proposed method in cases where the causal graph is unknown. Daisuke Takahashi, Shohei Shimizu, Takuma Tanaka |
IJCNN | 2 |
| 2024 | Causal-learn: Causal Discovery in PythonabstractCausal discovery aims at revealing causal relations from observational data, which is a fundamental task in science and engineering. We describe causal-learn, an open-source Python library for causal discovery. This library focuses on bringing a comprehensive collection of causal discovery methods to both practitioners and researchers. It provides easy-to-use APIs for non-specialists, modular building blocks for developers, detailed documentation for learners, and comprehensive methods for all. Different from previous packages in R or Java, causal-learn is fully developed in Python, which could be more in tune with the recent preference shift in programming languages within related communities. The library is available at https://github.com/py-why/causal-learn. Yujia Zheng 0001, Biwei Huang, Wei Chen 0103, Joseph D. Ramsey, Mingming Gong, Ruichu Cai, Shohei Shimizu, Peter Spirtes, Kun Zhang 0001 |
J. Mach. Learn. Res. | 7 |
| 2023 | BiLSTM and VAE Enhanced Multi-Task Neural Network for Trust-Aware E-Commerce Product AnalysisabstractRecently, reputation and trust analysis, especially using e-commerce product reviews which could be accumulated as large scale of text data, has contributed increasing greatly to marketing economics. However, it is not easy to find out the useful comments timely, or make the appropriate evaluation automatically, due to their sheer volume and time-varying updates. Existing studies usually treat them based on usefulness or sentiment analysis as a single task, but it is easy to ignore the latent features among different data domain and the potential causal relationships of various tasks. In this study, we focus on the multi-task deep learning, to cope with the reputation and trust analysis in e-commerce systems using product reviews. A so-called BiLSTM (Bidirectional LSTM) and VAE (Variational Autoencoder) Enhanced Multi-Task Neural Network (BiLV-MTNN) model is constructed to facilitate the evaluation, usefulness, and objectivity analysis simultaneously in a semi-supervised way. Specifically, A VAE based latent feature extraction mechanism and a BiLSTM based hidden vector generation mechanism are developed, which can contribute to the extraction of latent features and discovery of causal relationships among different tasks respectively. A labeling method is also proposed to quantify the objectivity of text composition based on the sentiment analysis, so as to improve the feature representation and further benefit the classification accuracy for multi-task learning. Compared with three baseline methods, experiments and evaluations using the Amazon data demonstrate the effectiveness and applicability of the proposed model in dealing with multi-task learning with alleviated negative transfer for trust-aware recommendation applications in e-commerce systems. Shusuke Wani, Xiaokang Zhou, Shohei Shimizu |
TrustCom | 3 |
| 2023 | Python package for causal discovery based on LiNGAMabstractCausal discovery is a methodology for learning causal graphs from data, and LiNGAM is a well-known model for causal discovery. This paper describes an open-source Python package for causal discovery based on LiNGAM. The package implements various LiNGAM methods under different settings like time series cases, multiple-group cases, mixed data cases, and hidden common cause cases, in addition to evaluation of statistical reliability and model assumptions. The source code is freely available under the MIT license at https://github.com/cdt15/lingam. Takashi Ikeuchi, Mayumi Ide, Takashi Nicholas Maeda, Shohei Shimizu |
J. Mach. Learn. Res. | 5 |
| 2023 | Digital Twin Enhanced Federated Reinforcement Learning With Lightweight Knowledge Distillation in Mobile NetworksabstractThe high-speed mobile networks offer great potentials to many future intelligent applications, such as autonomous vehicles in smart transportation systems. Such networks provide the possibility to interconnect mobile devices to achieve fast knowledge sharing for efficient collaborative learning and operations, especially with the help of distributed machine learning, e.g., Federated Learning (FL), and modern digital technologies, e.g., Digital Twin (DT) systems. Typically, FL requires a fixed group of participants that have Independent and Identically Distributed (IID) data for accurate and stable model training, which is highly unlikely in real-world mobile network scenarios. In this paper, in order to facilitate the lightweight model training and real-time processing in high-speed mobile networks, we design and introduce an end-edge-cloud structured three-layer Federated Reinforcement Learning (FRL) framework, incorporated with an edge-cloud structured DT system. A dual-Reinforcement Learning (dual-RL) scheme is devised to support optimizations of client node selection and global aggregation frequency during FL via a cooperative decision-making strategy, which is assisted by a two-layer DT system deployed in the edge-cloud for real-time monitoring of mobile devices and environment changes. A model pruning and federated bidirectional distillation (Bi-distillation) mechanism is then developed locally for the lightweight model training, while a model splitting scheme with a lightweight data augmentation mechanism is developed globally to separately optimize the aggregation weights based on a splitted neural network structure (i.e., the encoder and classifier) in a more targeted manner, which can work together to effectively reduce the overall communication cost and improve the non-IID problem. Experiment and evaluation results compared with three baseline methods using two different real-world datasets demonstrate the usefulness and outstanding performance of our proposed FRL model in communication-efficient model training and non-IID issue alleviation for high-speed mobile network scenarios. Xiaokang Zhou, Xuzhe Zheng, Xuesong Cui, Jiashuai Shi, Wei Liang 0006, Zheng Yan 0002, Laurence T. Yang, Shohei Shimizu, Kevin I-Kai Wang |
IEEE J. Sel. Areas Commun. | 8 |
| 2023 | Hierarchical Federated Learning With Social Context Clustering-Based Participant Selection for Internet of Medical Things ApplicationsabstractThe proliferation in embedded and communication technologies made the concept of the Internet of Medical Things (IoMT) a reality. Individuals’ physical and physiological status can be constantly monitored, and numerous data can be collected through wearable and mobile devices. However, the silo of individual data brings limitations to existing machine learning approaches to correctly identify a user’s health status. Distributed machine learning paradigms, such as federated learning, offer a potential solution for privacy-preserving knowledge sharing without sending raw personal data. However, federated learning is vulnerable to harmful participants that can degrade the overall model quality by sharing low-quality data. Therefore, it is critical to select suitable participants to ensure the accuracy and efficiency of federated learning. In this article, a unique clustering-based approach is proposed to use social context data for participant selection. Different edge participant groups will be established, and group-specific federated learning will be performed. The models of various edge groups will be further aggregated to strengthen the robustness of the global model. The experimental results demonstrated that through participant selection, clustering-based hierarchical federated learning can achieve better results with less participants in two different IoMT applications for ECG and human motion monitoring. This shows the efficacy of the proposed method in improving federated learning performance and efficiency in various IoMT applications. Xiaokang Zhou, Xiaozhou Ye, Kevin I-Kai Wang, Wei Liang 0006, Nirmal-Kumar C. Nair, Shohei Shimizu, Zheng Yan 0002, Qun Jin |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2023 | Nonlinear Causal Discovery for High-Dimensional Deterministic DataabstractNonlinear causal discovery with high-dimensional data where each variable is multidimensional plays a significant role in many scientific disciplines, such as social network analysis. Previous work majorly focuses on exploiting asymmetry in the causal and anticausal directions between two high-dimensional variables (a cause-effect pair). Although there exist some works that concentrate on the causal order identification between multiple variables, i.e., more than two high-dimensional variables, they do not validate the consistency of methods through theoretical analysis on multiple-variable data. In particular, based on the asymmetry for the cause-effect pair, if model assumptions for any pair of the data are violated, the asymmetry condition will not hold, resulting in the deduction of incorrect order identification. Thus, in this article, we propose a causal functional model, namely high-dimensional deterministic model (HDDM), to identify the causal orderings among multiple high-dimensional variables. We derive two candidates' selection rules to alleviate the inconvenient effects resulted from the violated-assumption pairs. The corresponding theoretical justification is provided as well. With these theoretical results, we develop a method to infer causal orderings for nonlinear multiple-variable data. Simulations on synthetic data and real-world data are conducted to verify the efficacy of our proposed method. Since we focus on deterministic relations in our method, we also verify the robustness of the noises in simulations. Yan Zeng 0002, Zhifeng Hao 0004, Ruichu Cai, Feng Xie 0002, Libo Huang 0001, Shohei Shimizu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2022 | CNN-GRU Based Deep Learning Model for Demand Forecast in Retail IndustryabstractA highly accurate demand forecast contributes to higher profits with more sales opportunities and lower waste losses because of excess inventory in the modern retail industry. As forecasting models for chain stores require high generalization performance and operational efficiency, it is still challenging to build an efficient model in the demand side to timely cope with different purchasing characteristics among multiple regions using traditional learning schemes. Hence, this paper proposes a pooling attention and gated recurrent unit (PA-GRU) based learning model, with one convolution neural network (CNN) layer for feature extraction and one GRU for time series forecasting with two improved attention mechanisms, to enhance the demand forecast with higher accuracy. The first attention mechanism captures the trend and periodic patterns of the target variable, while the second attention mechanism refines the core features for further forecasting. Deep learning frameworks enable forecasting that adapts to time-varying patterns in different complex demands, which can also be efficiently applied for cross-learning in multistock keeping unit and multi-store environments, facilitated by an improved transfer learning scheme. Experiments are conducted based on a real-world POS dataset of a food supermarket with 3,065,572 customers in 46 stores in nine prefectures. The evaluation results demonstrate the outstanding performance of our proposed model in dealing with multiple time series data for demand forecasting in the retail industry compared to those of several existing learning methods. Kazuhi Honjo, Xiaokang Zhou, Shohei Shimizu |
IJCNN | 3 |
| 2022 | Hierarchical Adversarial Attacks Against Graph-Neural-Network-Based IoT Network Intrusion Detection SystemabstractThe advancement of Internet of Things (IoT) technologies leads to a wide penetration and large-scale deployment of IoT systems across an entire city or even country. While IoT systems are capable of providing intelligent services, the large amount of data collected and processed in IoT systems also raises serious security concerns. Many research efforts have been devoted to design intelligent network intrusion detection system (NIDS) to prevent misuse of IoT data across smart applications. However, existing approaches may suffer from the issue of limited and imbalanced attack data when training the detection model, which make the system vulnerable especially for those unknown type attacks. In this study, a novel hierarchical adversarial attack (HAA) generation method is introduced to realize the level-aware black-box adversarial attack strategy, targeting the graph neural network (GNN)-based intrusion detection in IoT systems with a limited budget. By constructing a shadow GNN model, an intelligent mechanism based on a saliency map technique is designed to generate adversarial examples by effectively identifying and modifying the critical feature elements with minimal perturbations. A hierarchical node selection algorithm based on random walk with restart (RWR) is developed to select a set of more vulnerable nodes with high attack priority, considering their structural features, and overall loss changes within the targeted IoT network. The proposed HAA generation method is evaluated using the open-source data set UNSW-SOSR2019 with three baseline methods. Comparison results demonstrate its ability in degrading the classification precision by more than 30% in the two state-of-the-art GNN models, GCN and JK-Net, respectively, for NIDS in IoT environments. Xiaokang Zhou, Wei Liang 0006, Weimin Li 0001, Ke Yan 0001, Shohei Shimizu, Kevin I-Kai Wang |
IEEE Internet Things J. | 5 |
| 2022 | A Survey on Integrity Auditing for Data Storage in the Cloud: From Single Copy to Multiple ReplicasabstractThe rapid advancement of cloud computing has promoted the development of cloud storage services. One of the biggest concerns of cloud users is whether the completeness and recoverability of data can be guaranteed when cloud servers encounter problems. Only when the integrity of data is fully guaranteed can users consume cloud storage with confidence, especially in a complicated cloud environment with multiple clouds. However, the literature still lacks a thorough survey on cloud data integrity auditing for both single copy and multiple replicas. In this article, we survey and compare existing auditing schemes for single copy and multiple replicas based on a set of criteria. Based on our review and analysis, we discuss open issues, potential applications and future directions in the field of the integrity auditing in the cloud, including the implications of such trendy topics as merging blockchain and edge computing into data integrity auditing. Angtai Li, Yu Chen 0008, Zheng Yan 0002, Xiaokang Zhou, Shohei Shimizu |
IEEE Trans. Big Data | 5 |
| 2022 | B4SDC: A Blockchain System for Security Data Collection in MANETsabstractSecurity-related data collection is an essential part for attack detection and security measurement in Mobile Ad Hoc Networks (MANETs). A detection node (i.e., collector) should discover available routes to a collection node for data collection and collect security-related data during route discovery for determining reliable routes. However, few studies provide incentives for security-related data collection in MANETs. In this article, we propose B4SDC, a blockchain system for security-related data collection in MANETs. Through controlling the scale of Route REQuest (RREQ) forwarding in route discovery, the collector can constrain its payment and simultaneously make each forwarder of control information (namely RREQs and Route REPlies, in short RREPs) obtain rewards as much as possible to ensure fairness. At the same time, B4SDC avoids collusion attacks with cooperative receipt reporting, and spoofing attacks by adopting a secure digital signature. Based on a novel Proof-of-Stake consensus mechanism by accumulating stakes through message forwarding, B4SDC not only provides incentives for all participating nodes, but also avoids forking and ensures high efficiency and real decentralization. We analyze B4SDC in terms of incentives and security, and evaluate its performance through simulations. The thorough analysis and experimental results show the efficacy and effectiveness of B4SDC. Gao Liu, Huidong Dong, Zheng Yan 0002, Xiaokang Zhou, Shohei Shimizu |
IEEE Trans. Big Data | 5 |
| 2022 | Intelligent Small Object Detection for Digital Twin in Smart Manufacturing With Industrial Cyber-Physical SystemsabstractRecently, along with several technological advancements in cyber-physical systems, the revolution of Industry 4.0 has brought in an emerging concept named digital twin (DT), which shows its potential to break the barrier between the physical and cyber space in smart manufacturing. However, it is still difficult to analyze and estimate the real-time structural and environmental parameters in terms of their dynamic changes in digital twinning, especially when facing detection tasks of multiple small objects from a large-scale scene with complex contexts in modern manufacturing environments. In this article, we focus on a small object detection model for DT, aiming to realize the dynamic synchronization between a physical manufacturing system and its virtual representation. Three significant elements, including equipment, product, and operator, are considered as the basic environmental parameters to represent and estimate the dynamic characteristics and real-time changes in building a generic DT system of smart manufacturing workshop. A hybrid deep neural network model, based on the integration of MobileNetv2, YOLOv4, and Openpose, is constructed to identify the real-time status from physical manufacturing environment to virtual space. A learning algorithm is then developed to realize the efficient multitype small object detection based on the feature integration and fusion from both shallow and deep layers, in order to facilitate the modeling, monitoring, and optimizing of the whole manufacturing process in the DT system. Experiments and evaluations conducted in three different use cases demonstrate the effectiveness and usefulness of our proposed method, which can achieve a higher detection accuracy for DT in smart manufacturing. Xiaokang Zhou, Xuesong Xu, Wei Liang 0006, Shohei Shimizu, Laurence T. Yang, Qun Jin |
IEEE Trans. Ind. Informatics | 5 |
| 2021 | Causal Discovery with Multi-Domain LiNGAM for Latent FactorsabstractDiscovering causal structures among latent factors from observed data is a particularly challenging problem. Despite some efforts for this problem, existing methods focus on the single-domain data only. In this paper, we propose Multi-Domain Linear Non-Gaussian Acyclic Models for LAtent Factors (MD-LiNA), where the causal structure among latent factors of interest is shared for all domains, and we provide its identification results. The model enriches the causal representation for multi-domain data. We propose an integrated two-phase algorithm to estimate the model. In particular, we first locate the latent factors and estimate the factor loading matrix. Then to uncover the causal structure among shared latent factors of interest, we derive a score function based on the characterization of independence relations between external influences and the dependence relations between multi-domain latent factors and latent factors of interest. We show that the proposed method provides locally consistent estimators. Experimental results on both synthetic and real-world data demonstrate the efficacy and robustness of our approach. Yan Zeng 0002, Shohei Shimizu, Ruichu Cai, Feng Xie 0002, Michio Yamamoto, Zhifeng Hao 0004 |
IJCAI | 2 |
| 2021 | Causal additive models with unobserved variablesabstractCausal discovery from data affected by unobserved variables is an important but difficult problem to solve. The effects that unobserved variables have on the relationships between observed variables are more complex in nonlinear cases than in linear cases. In this study, we focus on causal additive models in the presence of unobserved variables. Causal additive models exhibit structural equations that are additive in the variables and error terms. We take into account the presence of not only unobserved common causes but also unobserved intermediate variables. Our theoretical results show that, when the causal relationships are nonlinear and there are unobserved variables, it is not possible to identify all the causal relationships between observed variables through regression and independence tests. However, our theoretical results also show that it is possible to avoid incorrect inferences. We propose a method to identify all the causal relationships that are theoretically possible to identify without being biased by unobserved variables. The empirical results using artificial data and simulated functional magnetic resonance imaging (fMRI) data show that our method effectively infers causal structures in the presence of unobserved variables. Takashi Nicholas Maeda, Shohei Shimizu |
UAI | 2 |
| 2021 | Siamese Neural Network Based Few-Shot Learning for Anomaly Detection in Industrial Cyber-Physical SystemsabstractWith the increasing population of Industry 4.0, both AI and smart techniques have been applied and become hotly discussed topics in industrial cyber-physical systems (CPS). Intelligent anomaly detection for identifying cyber-physical attacks to guarantee the work efficiency and safety is still a challenging issue, especially when dealing with few labeled data for cyber-physical security protection. In this article, we propose a few-shot learning model with Siamese convolutional neural network (FSL-SCNN), to alleviate the over-fitting issue and enhance the accuracy for intelligent anomaly detection in industrial CPS. A Siamese CNN encoding network is constructed to measure distances of input samples based on their optimized feature representations. A robust cost function design including three specific losses is then proposed to enhance the efficiency of training process. An intelligent anomaly detection algorithm is developed finally. Experiment results based on a fully labeled public dataset and a few labeled dataset demonstrate that our proposed FSL-SCNN can significantly improve false alarm rate (FAR) and F1 scores when detecting intrusion signals for industrial CPS security protection. Xiaokang Zhou, Wei Liang 0006, Shohei Shimizu, Jianhua Ma 0002, Qun Jin |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | RCD: Repetitive causal discovery of linear non-Gaussian acyclic models with latent confoundersabstractCausal discovery from data affected by latent confounders is an important and difficult challenge. Causal functional model-based approaches have not been used to present variables whose relationships are affected by latent confounders, while some constraint-based methods can present them. This paper proposes a causal functional model-based method called repetitive causal discovery (RCD) to discover the causal structure of observed variables affected by latent confounders. RCD repeats inferring the causal directions between a small number of observed variables and determines whether the relationships are affected by latent confounders. RCD finally produces a causal graph where a bi-directed arrow indicates the pair of variables that have the same latent confounders, and a directed arrow indicates the causal direction of a pair of variables that are not affected by the same latent confounder. The results of experimental validation using simulated data and real-world data confirmed that RCD is effective in identifying latent confounders and causal directions between observed variables. Takashi Nicholas Maeda, Shohei Shimizu |
AISTATS | 2 |
| 2020 | Estimation of Post-Nonlinear Causal Models Using Autoencoding StructureabstractDiscovering causal relations in complex systems is an important problem in many research fields. To describe such systems involving nonlinear causal relations, the post-nonlinear (PNL) causal model has been proposed. However, despite its identifiability, estimation methods of PNL model have not been developed as well as linear models. In this paper, we proposed a new estimation method of PNL model using an autoencoding structure. Our method estimates the model by minimizing two losses corresponding to two assumptions of PNL model: independence between the cause and the noise and invertibility of a nonlinear distortion. Experimental results on artificial data show that our method estimates underlying model satisfying both assumptions. In addition, the proposed method finds correct causal directions 1.5 times as many real-world problems as the existing method assuming linear causal relations. Kento Uemura, Shohei Shimizu |
ICASSP | 2 |
| 2019 | Multi-Modality Behavioral Influence Analysis for Personalized Recommendations in Health Social Media EnvironmentabstractRecently, health social media have engaged more and more people to share their personal feelings, opinions, and experience in the context of health informatics, which has drawn increasing attention from both academia and industry. In this paper, we focus on the behavioral influence analysis based on heterogeneous health data generated in social media environments. An integrated deep neural network (DNN)-based learning model is designed to analyze and describe the latent behavioral influence hidden across multiple modalities, in which a convolutional neural network (CNN)-based framework is used to extract the time-series features within a certain social context. The learned features based on cross-modality influence analysis are then trained in a SoftMax classifier, which can result in a restructured representation of high-level features for online physician rating and classification in a data-driven way. Finally, two algorithms within two representative application scenarios are developed to provide patients with personalized recommendations in health social media environments. Experiments using the real world data demonstrate the effectiveness of our proposed model and method. Xiaokang Zhou, Wei Liang 0006, Kevin I-Kai Wang, Shohei Shimizu |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2018 | Cause-Effect Inference by Comparing Regression ErrorsabstractWe address the problem of inferring the causal relation between two variables by comparing the least-squares errors of the predictions in both possible causal directions. Under the assumption of an independence between the function relating cause and effect, the conditional noise distribution, and the distribution of the cause, we show that the errors are smaller in causal direction if both variables are equally scaled and the causal relation is close to deterministic. Based on this, we provide an easily applicable method that only requires a regression in both possible causal directions. The performance of this method is compared with different related causal inference methods in various artificial and real-world data sets. Patrick Blöbaum, Dominik Janzing, Takashi Washio, Shohei Shimizu, Bernhard Schölkopf |
AISTATS | 4 |
| 2017 | A novel principle for causal inference in data with small error variance
Patrick Blöbaum, Shohei Shimizu, Takashi Washio |
ESANN | 2 |
| 2017 | Learning Instrumental Variables with Structural and Non-Gaussianity AssumptionsabstractLearning a causal effect from observational data requires strong assumptions. One possible method is to use instrumental variables, which are typically justified by background knowledge. It is possible, under further assumptions, to discover whether a variable is structurally instrumental to a target causal effect $X \rightarrow Y$. However, the few existing approaches are lacking on how general these assumptions can be, and how to express possible equivalence classes of solutions. We present instrumental variable discovery methods that systematically characterize which set of causal effects can and cannot be discovered under local graphical criteria that define instrumental variables, without reconstructing full causal graphs. We also introduce the first methods to exploit non-Gaussianity assumptions, highlighting identifiability problems and solutions. Due to the difficulty of estimating such models from finite data, we investigate how to strengthen assumptions in order to make the statistical problem more manageable. Ricardo Silva 0001, Shohei Shimizu |
J. Mach. Learn. Res. | 2 |
| 2014 | Bayesian estimation of causal direction in acyclic structural equation models with individual-specific confounder variables and non-Gaussian distributions
Shohei Shimizu, Kenneth Bollen |
J. Mach. Learn. Res. | 1 |
| 2014 | ParceLiNGAM: A Causal Ordering Method Robust Against Latent ConfoundersabstractWe consider learning a causal ordering of variables in a linear nongaussian acyclic model called LiNGAM. Several methods have been shown to consistently estimate a causal ordering assuming that all the model assumptions are correct. But the estimation results could be distorted if some assumptions are violated. In this letter, we propose a new algorithm for learning causal orders that is robust against one typical violation of the model assumptions: latent confounders. The key idea is to detect latent confounders by testing independence between estimated external influences and find subsets (parcels) that include variables unaffected by latent confounders. We demonstrate the effectiveness of our method using artificial data and simulated brain imaging data. Tatsuya Tashiro, Shohei Shimizu, Aapo Hyvärinen, Takashi Washio |
Neural Comput. | 2 |
| 2012 | Estimation of Causal Orders in a Linear Non-Gaussian Acyclic Model: A Method Robust against Latent Confounders
Tatsuya Tashiro, Shohei Shimizu, Aapo Hyvärinen, Takashi Washio |
ICANN (1) | 2 |
| 2012 | Joint estimation of linear non-Gaussian acyclic models
Shohei Shimizu |
Neurocomputing | 1 |
| 2011 | Discovering causal structures in binary exclusive-or skew acyclic models
Takanori Inazumi, Takashi Washio, Shohei Shimizu, Joe Suzuki, Akihiro Yamamoto, Yoshinobu Kawahara |
UAI | 3 |
| 2011 | Analyzing relationships among ARMA processes based on non-Gaussianity of external influences
Yoshinobu Kawahara, Shohei Shimizu, Takashi Washio |
Neurocomputing | 2 |
| 2011 | DirectLiNGAM: A Direct Method for Learning a Linear Non-Gaussian Structural Equation Model
Shohei Shimizu, Takanori Inazumi, Yasuhiro Sogawa, Aapo Hyvärinen, Yoshinobu Kawahara, Takashi Washio, Patrik O. Hoyer, Kenneth Bollen |
J. Mach. Learn. Res. | 1 |
| 2011 | Estimating exogenous variables in data with more variables than observations
Yasuhiro Sogawa, Shohei Shimizu, Teppei Shimamura, Aapo Hyvärinen, Takashi Washio, Seiya Imoto |
Neural Networks | 2 |
| 2010 | Assessing Statistical Reliability of LiNGAM via Multiscale Bootstrap
Yusuke Komatsu, Shohei Shimizu, Hidetoshi Shimodaira |
ICANN (3) | 2 |
| 2010 | Discovery of Exogenous Variables in Data with More Variables Than Observations
Yasuhiro Sogawa, Shohei Shimizu, Aapo Hyvärinen, Takashi Washio, Teppei Shimamura, Seiya Imoto |
ICANN (1) | 2 |
| 2010 | An experimental comparison of linear non-Gaussian causal discovery methods and their variantsabstractMany multivariate Gaussianity-based techniques for identifying causal networks of observed variables have been proposed. These methods have several problems such that they cannot uniquely identify the causal networks without any prior knowledge. To alleviate this problem, a non-Gaussianity-based identification method LiNGAM was proposed. Though the LiNGAM potentially identifies a unique causal network without using any prior knowledge, it needs to properly examine independence assumptions of the causal network and search the correct causal network by using finite observed data points only. On another front, a kernel based independence measure that evaluates the independence more strictly was recently proposed. In addition, some advanced generic search algorithms including beam search have been extensively studied in the past. In this paper, we propose some variants of the LiNGAM method which introduce the kernel based method and the beam search enabling more accurate causal network identification. Furthermore, we experimentally characterize the LiNGAM and its variants in terms of accuracy and robustness of their identification. Yasuhiro Sogawa, Shohei Shimizu, Yoshinobu Kawahara, Takashi Washio |
IJCNN | 2 |
| 2010 | Estimation of a Structural Vector Autoregression Model Using Non-Gaussianity
Aapo Hyvärinen, Kun Zhang 0001, Shohei Shimizu, Patrik O. Hoyer |
J. Mach. Learn. Res. | 3 |
| 2009 | A direct method for estimating a causal ordering in a linear non-Gaussian acyclic model
Shohei Shimizu, Aapo Hyvärinen, Yoshinobu Kawahara |
UAI | 1 |
| 2009 | Estimation of linear non-Gaussian acyclic models for latent factors
Shohei Shimizu, Patrik O. Hoyer, Aapo Hyvärinen |
Neurocomputing | 1 |
| 2008 | Causal modelling combining instantaneous and lagged effects: an identifiable model based on non-GaussianityabstractCausal analysis of continuous-valued variables typically uses either autoregressive models or linear Gaussian Bayesian networks with instantaneous effects. Estimation of Gaussian Bayesian networks poses serious identifiability problems, which is why it was recently proposed to use non-Gaussian models. Here, we show how to combine the non-Gaussian instantaneous model with autoregressive models. We show that such a non-Gaussian model is identifiable without prior knowledge of network structure, and we propose an estimation method shown to be consistent. This approach also points out how neglecting instantaneous effects can lead to completely wrong estimates of the autoregressive coefficients. Aapo Hyvärinen, Shohei Shimizu, Patrik O. Hoyer |
ICML | 2 |
| 2008 | Causal discovery of linear acyclic models with arbitrary distributions
Patrik O. Hoyer, Aapo Hyvärinen, Richard Scheines, Peter Spirtes, Joseph D. Ramsey, Gustavo Lacerda, Shohei Shimizu |
UAI | 7 |
| 2008 | Estimation of causal effects using linear non-Gaussian causal models with hidden variables
Patrik O. Hoyer, Shohei Shimizu, Antti J. Kerminen, Markus Palviainen |
Int. J. Approx. Reason. | 2 |
| 2007 | Discovery of Linear Non-Gaussian Acyclic Models in the Presence of Latent Classes
Shohei Shimizu, Aapo Hyvärinen |
ICONIP (1) | 1 |
| 2006 | A Quasi-stochastic Gradient Algorithm for Variance-Dependent Component Analysis
Aapo Hyvärinen, Shohei Shimizu |
ICANN (2) | 2 |
| 2006 | A Linear Non-Gaussian Acyclic Model for Causal DiscoveryabstractIn recent years, several methods have been proposed for the discovery of causal structure from non-experimental data. Such methods make various assumptions on the data generating process to facilitate its identification from purely observational data. Continuing this line of research, we show how to discover the complete causal structure of continuous-valued data, under the assumptions that (a) the data generating process is linear, (b) there are no unobserved confounders, and (c) disturbance variables have non-Gaussian distributions of non-zero variances. The solution relies on the use of the statistical method known as independent component analysis, and does not require any pre-specified time-ordering of the variables. We provide a complete Matlab package for performing this LiNGAM analysis (short for Linear Non-Gaussian Acyclic Model), and demonstrate the effectiveness of the method using artificially generated data and real-world data. Shohei Shimizu, Patrik O. Hoyer, Aapo Hyvärinen, Antti J. Kerminen |
J. Mach. Learn. Res. | 1 |
| 2005 | Discovery of Non-gaussian Linear Causal Models using ICA
Shohei Shimizu, Aapo Hyvärinen, Yutaka Kano, Patrik O. Hoyer |
UAI | 1 |