VLDB 2026 Research / reviewers in the wild / expert
Ruichu Cai
dblp:09/6889
· DBLP profile ↗
178ranked-venue papers
57as first author
117since 2021 · last 2026
0000-0001-8972-167XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 129 · 45 first-author · 92 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 11 first-author · 28 since 2021Databases, data management, data science and information retrieval · 20 · 7 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 5 first-author · 10 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Horizontal and Vertical Federated Causal Structure Learning via Higher-order CumulantsabstractFederated causal discovery aims to uncover causal relationships while protecting data privacy, with significant real-world applications. Existing methods focus on horizontal federated settings where clients share the same variables but have different samples. However, in practice, clients may have different variables, leading to spurious causal relationships. To address this issue, we comprehensively consider causal structure learning methods under both horizontal and vertical federated settings. Interestingly, we find that, higher-order cumulants rely solely on the joint distribution of the relevant variables and are useful to solve the above problem in the linear non-Gaussian case. This motivates us to provide the identification theories for determining the causal order over observed variables, leveraging the difference in the product of the (cross) cumulants of the specific variables. Based on these theories, we develop a method for learning causal order in the horizontal and vertical federated scenarios. Specifically, we first obtain local (cross) cumulant matrices of observed variables from all participating clients to construct a global cumulant matrix. This global cumulant matrix is then used for recursive source variable identification, ultimately yielding a causal strength matrix of the union of variables from all clients. Our algorithm demonstrates superior performance in experiments on both synthetic and real-world data. Wei Chen 0103, Wanyang Gu, Linjun Peng, Ruichu Cai, Zhifeng Hao 0004, Kun Zhang 0001 |
AAAI | 5 |
| 2026 | CAMA: Enhancing Mathematical Reasoning in Large Language Models with Causal KnowledgeabstractLarge Language Models (LLMs) have demonstrated strong performance across a wide range of tasks, yet they still struggle with complex mathematical reasoning, a challenge fundamentally rooted in deep structural dependencies. To address this challenge, we propose CAusal MAthematician (CAMA), a two stage causal framework that equips LLMs with explicit, reusable mathematical structure. In the learning stage, CAMA first constructs the Mathematical Causal Graph (MCG), a high level representation of solution strategies, by combining LLM priors with causal discovery algorithms applied to a corpus of question solution pairs. The resulting MCG encodes essential knowledge points and their causal dependencies. To better align the graph with downstream reasoning tasks, CAMA further refines the MCG through iterative feedback derived from a selected subset of the question solution pairs. In the reasoning stage, given a new question, CAMA dynamically extracts a task relevant subgraph from the MCG, conditioned on both the question content and the LLM’s intermediate reasoning trace. This subgraph, which encodes the most pertinent knowledge points and their causal dependencies, is then injected back into the LLM to guide its reasoning process. Empirical results on real world datasets show that CAMA significantly improves LLM performance on challenging mathematical problems. Furthermore, our experiments demonstrate that structured guidance consistently outperforms unstructured alternatives, and that incorporating asymmetric causal relationships yields greater improvements than using symmetric associations alone. Lei Zan, Keli Zhang, Ruichu Cai, Lujia Pan |
AAAI | 3 |
| 2026 | CMCTS: A Constrained Monte Carlo Tree Search framework for mathematical reasoning in large language model
Qingwen Lin, Guimin Hu, Zijian Li 0001, Zhifeng Hao 0004, Keli Zhang, Ruichu Cai |
Appl. Intell. | 7 |
| 2026 | Learning by doing: an online causal reinforcement learning framework with causal-aware policy
Ruichu Cai, Siyang Huang, Jie Qiao, Wei Chen 0103, Yan Zeng 0002, Keli Zhang, Fuchun Sun 0001, Zhifeng Hao 0004 |
Sci. China Inf. Sci. | 1 |
| 2026 | Learning high-order user-item relation via hyperedge for recommender system
Yuguang Yan, Ruichu Cai, Michael Kwok-Po Ng |
Neurocomputing | 3 |
| 2026 | Targeted mining of non-overlapping high-utility sequential patterns
Wensheng Gan, Zhidong Lin, Zhenlian Qi, Jian Zhu 0001, Ruichu Cai, Zhifeng Hao 0004 |
Inf. Sci. | 6 |
| 2026 | An identifiable cost-aware causal decision-making framework using counterfactual reasoning
Ruichu Cai, Jie Qiao, Zijian Li 0001, Yuequn Liu, Wei Chen 0103, Keli Zhang, Jiale Zheng |
Neural Networks | 1 |
| 2026 | Time Series Domain Adaptation via Latent Invariant Causal MechanismabstractTime series domain adaptation aims to transfer the complex temporal dependence from the labeled source domain to the unlabeled target domain. Recent advances leverage the stable causal mechanism over observed variables to model the domain-invariant temporal dependence. However, modeling precise causal structures in high-dimensional data, such as videos, remains challenging. Additionally, direct causal edges may not exist among observed variables (e.g., pixels). These limitations hinder the applicability of existing approaches to real-world scenarios. To address these challenges, we find that the high-dimension time series data are generated from the low-dimension latent variables, which motivates us to model the causal mechanisms of the temporal latent process. Based on this intuition, we propose a latent causal mechanism identification framework that guarantees the uniqueness of the reconstructed latent causal structures. Specifically, we first identify latent variables by utilizing sufficient changes in historical information. Moreover, by enforcing the sparsity of the relationships of latent variables, we can achieve identifiable latent causal structures. Built on the theoretical results, we develop the Latent Causality Alignment (LCA) model that leverages variational inference, which incorporates an intra-domain latent sparsity constraint for latent structure reconstruction and an inter-domain latent sparsity constraint for domain-invariant structure reconstruction. Experiment results on eight benchmarks show a general improvement in the domain-adaptive time series classification and forecasting tasks, highlighting the effectiveness of our method in real-world scenarios. Ruichu Cai, Junxian Huang 0002, Zhenhui Yang, Zijian Li 0001, Emadeldeen Eldele, Min Wu 0008, Fuchun Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Test-Time Domain Adaptation With Time-Frequency Consistency and Instance-Aware Batch Renormalization for Online Machinery Fault Diagnosis
Jian Zhu 0001, Bairui Long, Lunke Fei, Yutang Xiao, Boyu Wang 0004, Ruichu Cai |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2026 | Temporal Recommendation Based on Adaptive Deep Matrix FactorizationabstractTemporal recommendation is an important class of tasks in recommender systems, which focuses on modeling and capturing temporal patterns in user behavior to achieve finer-grained and higher-quality recommendations. In real-world scenario, users' temporal behaviors are not only characterized by sequential dependencies among consecutive items, but also by periodic correlations of different items and time-varying similarity of different users. In this paper, we propose an Adaptive Temporal Recommendation (AdaTR) algorithm to capture the inherent features of temporal behaviors and dynamic collaborative signals. Firstly, based on the periodic characteristics of user behaviors, the user-item interactions are counted and aggregated in different time segments across multiple periods, which forms the temporal user-item interaction matrix. Then, in order to capture the time-varying collaborative signals between different users, a deep spectral clustering (DSC) method is implemented on the temporal user-item interaction matrix, where the original representation of user-item interaction is projected into a latent space, and users' temporal behaviors are clustered into different groups. Furthermore, an Adaptive Deep Matrix Factorization (AdaDMF) module is designed to learn the time-varying representations of user preferences on each cluster of temporal user behaviors, which incoporate dynamic collaborative signals among different users. Finally, we combine users' short-term and long-term preferences to generate personalized temporal recommendations. Extensive experiments on four datasets demonstrate that AdaTR performs significantly better than the state-of-the-art baselines. Yali Feng, Zhifeng Hao 0004, Wen Wen 0009, Ruichu Cai |
IEEE Trans. Big Data | 4 |
| 2026 | Higher Order Cumulants-Based Method for Direct and Efficient Causal DiscoveryabstractCausal discovery plays a pivotal role in scientific inquiry and subsequent applications in prediction or decision-making. While many methods have been proposed, many of them rely on independence tests. However, these tests are difficult to implement and computationally intensive. In this article, we aim to propose a direct and computationally efficient method to determine the causal relationship between two observed variables in the linear non-Gaussian case. Building on the insight that cumulants provide information about the shape of a probability distribution, we show that interestingly, the (in)dependence between two observed variables can be directly inferred from the difference in the product of certain joint cumulants of these variables. This concept is named the cause difference criterion. Based on this criterion, we introduce two practical methods, high-order cumulant (HC) and HC-linear non-Gaussian acyclic model (LiNGAM), for causal discovery in the high-dimensional case. Theoretical analyses ensure the identifiability of the proposed criteria and methods. Experimental results indicate that our methods outperform most existing methods. Wei Chen 0103, Linjun Peng, Zhiyi Huang 0008, Ruichu Cai, Zhifeng Hao 0004, Kun Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Disentangling Long-Short Term State Under Unknown Interventions for Online Time Series ForecastingabstractCurrent methods for time series forecasting struggle in the online scenario, since it is difficult to preserve long-term dependency while adapting short-term changes when data are arriving sequentially. Although some recent methods solve this problem by controlling the updates of latent states, they cannot disentangle the long/short-term states, leading to the inability to effectively adapt to nonstationary. To tackle this challenge, we propose a general framework to disentangle long/short-term states for online time series forecasting. Our idea is inspired by the observations where short-term changes can be led by unknown interventions like abrupt policies in the stock market. Based on this insight, we formalize a data generation process with unknown interventions on short-term states. Under mild assumptions, we further leverage the independence of short-term states led by unknown interventions to establish the identification theory to achieve the disentanglement of long/short-term states. Built on this theory, we develop a Long Short-Term Disentanglement model (LSTD) to extract the long/short-term states with long/short term encoders, respectively. Furthermore, the LSTD model incorporates a smooth constraint to preserve the long-term dependencies and an interrupted dependency constraint to enforce the forgetting of short-term dependencies, together boosting the disentanglement of long/short-term states. Experimental results on several benchmark datasets show that our LSTD model outperforms existing methods for online time series forecasting, validating its efficacy in real-world applications. Ruichu Cai, Haiqin Huang, Zhifan Jiang, Zijian Li 0001, Changze Zhou, Yuequn Liu, Zhifeng Hao 0004 |
AAAI | 1 |
| 2025 | Hypergraph Learning for Unsupervised Graph Alignment via Optimal TransportabstractUnsupervised graph alignment aims to find corresponding nodes across different graphs without supervision. Existing methods usually leverage the graph structure to aggregate features of nodes to find relations between nodes. However, the graph structure is inherently limited in pairwise relations between nodes without considering higher-order dependencies among multiple nodes. In this paper, we take advantage of the hypergraph structure to characterize higher-order structural information among nodes for better graph alignment. Specifically, we propose an optimal transport model to learn a hypergraph to capture complex relations among nodes, so that the nodes involved in one hyperedge can be adaptively based on local geometric information. In addition, inspired by the Dirichlet energy function of a hypergraph, we further refine our model to enhance the consistency between structural and feature information in each hyperedge. After that, we jointly leverage graphs and hypergraphs to extract structural and feature information to better model the relations between nodes, which is used to find node correspondences across graphs. We conduct experiments on several benchmark datasets with different settings, and the results demonstrate the effectiveness of our proposed method. Yuguang Yan, Canlin Yang, Ruichu Cai, Michael Kwok-Po Ng |
AAAI | 4 |
| 2025 | Improving Cognitive Decline Prediction in Alzheimer's Disease via a Novel Biomarker with Causal InterpretabilityabstractAlzheimer's disease (AD) progresses before symptoms, motivating accurate and interpretable risk prediction. We propose an anatomy-informed$A \beta$-PET biomarker that divides lobe-level SUVRs by the global SUVR, yielding six region-toglobal ratios. On ADNI PET data augmented with demographics and CDR composites, classifiers (SVM, RF, XGBoost) with ablations show consistent gains. Beyond accuracy, we estimate each ratio's average treatment effect (ATE) via Double Machine Learning (DML) and complement it with SHAP to separate causality from correlation. Results indicate notable improvements (e.g.,$\text{RF} \text{AUC}=0.956; \mathrm{F} 1=0.882$) and a stable, positive ATE for the occipital ratio (Abeta_OCC: 0.2520, 95% CI [0.1827,$0.3265]$). The framework integrates domain structure with causal interpretability to support early identification and clinically meaningful explanations. Zezhen Ding, Jiongjia Li, Xuexin Chen, Ruichu Cai |
BIBM | 6 |
| 2025 | Dr.ECI: Infusing Large Language Models with Causal Knowledge for Decomposed Reasoning in Event Causality IdentificationabstractDespite the demonstrated potential of Large Language Models (LLMs) in diverse NLP tasks, their causal reasoning capability appears inadequate when evaluated within the context of the event causality identification (ECI) task. The ECI tasks pose significant complexity for LLMs and necessitate comprehensive causal priors for accurate identification. To improve the performance of LLMs for causal reasoning, we propose a multi-agent Decomposed reasoning framework for Event Causality Identification, designated as Dr.ECI. In the discovery stage, Dr.ECI incorporates specialized agents such as Causal Explorer and Mediator Detector, which capture implicit causality and indirect causality more effectively. In the reasoning stage, Dr.ECI introduces the agents Direct Reasoner and Indirect Reasoner, which leverage the knowledge of the generalized causal structure specific to the ECI. Extensive evaluations demonstrate the state-of-the-art performance of Dr.ECI comparing with baselines based on LLMs and supervised training. Our implementation will be open-sourced at https://github.com/DMIRLAB-Group/Dr.ECI. Ruichu Cai, Shengyin Yu, Keli Zhang |
COLING | 1 |
| 2025 | CACA: Context-Aware Cross-Attention Network for Extractive Aspect Sentiment Quad PredictionabstractAspect Sentiment Quad Prediction(ASQP) enhances the scope of aspect-based sentiment analysis by introducing the necessity to predict both explicit and implicit aspect and opinion terms. Existing leading generative ASQP approaches do not modeling the contextual relationship of the review sentence to predict implicit terms. However, introducing the contextual information into the pre-trained language models framework is non-trivial due to the inflexibility of the generative encoder-decoder architecture. To well utilize the contextual information, we propose an extractive ASQP framework, CACA, which features with Context-Aware Cross-Attention Network. When implicit terms are present, the Context-Aware Cross-Attention Network enhances the alignment of aspects and opinions, through alternating updates of explicit and implicit representations. Additionally, contrastive learning is introduced in the implicit representation learning process. Experimental results on three benchmarks demonstrate the effectiveness of CACA. Our implementation will be open-sourced at https://github.com/DMIRLAB-Group/CACA. Bingfeng Chen, Yongqi Luo, Ruichu Cai |
COLING | 5 |
| 2025 | Rank Constraints of High-Order Cumulants for Learning Linear Non-Gaussian Latent PolytreeabstractWe study the problem of learning the causal structure of a latent tree model only from the observational data. Prior works often assume either a sufficient number of measured variables or the observed variables can only be the child of latent variables (known as measurement assumption). However, they may yield incorrect or uninformative results when some observed variables also cause the latent variable, or when the number of measured variables is less than two. In this paper, we focus on the linear non-Gaussian latent polytree model, where the observed and latent variables can exhibit arbitrary causal dependence and the number of child variables for each latent variable may be only one. By leveraging the non-Gaussianity within the causal model, we introduce rank constraints of high-order cumulants. These constraints align with trek separation within the causal graph and enable the identification of exogenous variables for the relative set. Such properties have intriguing possibilities for identifying the entire latent polytree structure, including not only the number of latent variables but also causal directions. Consequently, we develop an identification algorithm to learn latent polytree by only using the rank constraints of high-order cumulants, and we verify its effectiveness in simulation experiments. Ruichu Cai, Zhengming Chen 0002, Feng Xie 0002, Zhifeng Hao 0004 |
CSCWD | 1 |
| 2025 | GenLink: Generation-Driven Schema-Linking via Multi-Model Learning for Text-to-SQLabstractSchema linking is widely recognized as a key factor in improving text-to-SQL performance.Supervised fine-tuning approaches enhance SQL generation quality by explicitly finetuning schema linking as an extraction task.However, they suffer from two major limitations: (i) The training corpus of small language models restricts their cross-domain generalization ability.(ii) The extraction-based finetuning process struggles to capture complex linking patterns.To address these issues, we propose GenLink, a generation-driven schemalinking framework based on multi-model learning.Instead of explicitly extracting schema elements, GenLink enhances linking through a generation-based learning process, effectively capturing implicit schema relationships.By integrating multiple small language models, GenLink improves schema-linking recall rate and ensures robust cross-domain adaptability.Experimental results on the BIRD and Spider benchmarks validate the effectiveness of GenLink, achieving execution accuracies of 67.34% (BIRD), 89.7% (Spider development set), and 87.8% (Spider test set), demonstrating its superiority in handling diverse and complex database schemas. Shaobin Shi, Ruichu Cai |
EMNLP | 4 |
| 2025 | Chat2DB: Chatting to the Database with Interactive Agent Assisted Language ModelsabstractCross-domain Text-to-SQL necessitates the capability of semantic parsers to generalize to unseen databases, thus simplifying the process of creating natural language interfaces for databases. The existing Text-to-SQL parser exhibits limitations in its adaptability to new databases, and its execution accuracy is not sufficient for building conversational applications, typically necessitating further fine-tuning for specific databases. In this paper, we introduce Chat2DB, a conversational system designed for database interactions that enhances parser capabilities, rendering them applicable in real-world contexts. Within Chat2DB, we implement an interactive schema-ranking agent that optimizes the performance of LMs-based parsers cost-effectively. We further propose an adaptive retraining stage to allow trained Text-to-SQL parsers to quickly adapt to the target database. Experimental evaluations were conducted to validate the performance of the key components of Chat2DB. In the demonstration, we showcase the interactive visualization interface of Chat2DB to achieve more accurate querying of databases by natural language. Yuyuan Cai, Shaobin Shi, Ruichu Cai |
ICDE | 5 |
| 2025 | Synergy Between Sufficient Changes and Sparse Mixing Procedure for Disentangled Representation LearningabstractDisentangled representation learning aims to uncover the latent variables underlying observed data, yet identifying these variables under mild assumptions remains challenging. Some methods rely on sufficient changes in the distribution of latent variables indicated by auxiliary variables, such as domain indices, but acquiring enough domains is often impractical. Alternative approaches exploit the structural sparsity assumption on mixing processes, but this constraint may not hold in practice. Interestingly, we find that these two seemingly unrelated assumptions can actually complement each other. Specifically, when conditioned on auxiliary variables, the sparse mixing process induces independence between latent and observed variables, which simplifies the mapping from estimated to true latent variables and hence compensates for deficiencies of auxiliary variables. Building on this insight, we propose an identifiability theory with less restrictive constraints regarding the auxiliary variables and the sparse mixing process, enhancing applicability to real-world scenarios. Additionally, we develop a generative model framework incorporating a domain encoding network and a sparse mixing constraint and provide two implementations based on variational autoencoders and generative adversarial networks. Experiment results on synthetic and real-world datasets support our theoretical results. Zijian Li 0001, Shunxing Fan, Yujia Zheng 0001, Ignavier Ng, Shaoan Xie, Guangyi Chen 0002, Xinshuai Dong, Ruichu Cai, Kun Zhang 0001 |
ICLR | 8 |
| 2025 | On the Identification of Temporal Causal Representation with Instantaneous DependenceabstractTemporally causal representation learning aims to identify the latent causal process from time series observations, but most methods require the assumption that the latent causal processes do not have instantaneous relations. Although some recent methods achieve identifiability in the instantaneous causality case, they require either interventions on the latent variables or grouping of the observations, which are in general difficult to obtain in real-world scenarios. To fill this gap, we propose an \textbf{ID}entification framework for instantane\textbf{O}us \textbf{L}atent dynamics (\textbf{IDOL}) by imposing a sparse influence constraint that the latent causal processes have sparse time-delayed and instantaneous relations. Specifically, we establish identifiability results of the latent causal process based on sufficient variability and the sparse influence constraint by employing contextual information of time series data. Based on these theories, we incorporate a temporally variational inference architecture to estimate the latent variables and a gradient-based sparsity regularization to identify the latent causal process. Experimental results on simulation datasets illustrate that our method can identify the latent causal process. Furthermore, evaluations on multiple human motion forecasting benchmarks with instantaneous dependencies indicate the effectiveness of our method in real-world settings. Zijian Li 0001, Yifan Shen 0004, Kaitao Zheng, Ruichu Cai, Xiangchen Song, Mingming Gong, Guangyi Chen 0002, Kun Zhang 0001 |
ICLR | 4 |
| 2025 | Identification of Latent Confounders via Investigating the Tensor Ranks of the Nonlinear ObservationsabstractWe study the problem of learning discrete latent variable causal structures from mixed-type observational data. Traditional methods, such as those based on the tensor rank condition, are designed to identify discrete latent structure models and provide robust identification bounds for discrete causal models. However, when observed variables—specifically, those representing the children of latent variables—are collected at various levels with continuous data types, the tensor rank condition is not applicable, limiting further causal structure learning for latent variables. In this paper, we consider a more general case where observed variables can be either continuous or discrete, and further allow for scenarios where multiple latent parents cause the same set of observed variables. We show that, under the completeness condition, it is possible to discretize the data in a way that satisfies the full-rank assumption required by the tensor rank condition. This enables the identifiability of discrete latent structure models within mixed-type observational data. Moreover, we introduce the two-sufficient measurement condition, a more general structural assumption under which the tensor rank condition holds and the underlying latent causal structure is identifiable by a proposed two-stage identification algorithm. Extensive experiments on both simulated and real-world data validate the effectiveness of our method. Zhengming Chen 0002, Yewei Xia, Feng Xie 0002, Jie Qiao, Zhifeng Hao 0004, Ruichu Cai, Kun Zhang 0001 |
ICML | 6 |
| 2025 | Reducing Confounding Bias without Data Splitting for Causal Inference via Optimal TransportabstractCausal inference seeks to estimate the effect given a treatment such as a medicine or the dosage of a medication. To reduce the confounding bias caused by the non-randomized treatment assignment, most existing methods reduce the shift between subpopulations receiving different treatments. However, these methods split limited training samples into smaller groups, which cuts down the number of samples in each group, while precise distribution estimation and alignment highly rely on a sufficient number of training samples. In this paper, we propose a distribution alignment paradigm without data splitting, which can be naturally applied in the settings of binary and continuous treatments. To this end, we characterize the confounding bias by considering different probability measures of the same set including all the training samples, and exploit the optimal transport theory to analyze the confounding bias and outcome estimation error. Based on this, we propose to learn balanced representations by reducing the bias between the marginal distribution and the conditional distribution of a treatment. As a result, data reduction caused by splitting is avoided, and the outcome prediction model trained on one treatment group can be generalized to the entire population. The experiments on both binary and continuous treatment settings demonstrate the effectiveness of our method. Yuguang Yan, Zongyu Li, Zeqin Yang, Ruichu Cai |
ICML | 6 |
| 2025 | Robust Policy Learning for Multi-UAV Collision Avoidance with Causal Feature Selection
Jiafan Zhuang, Gaofei Han, Zihao Xia, Che Lin, Boxi Wang, Wenji Li, Ruichu Cai, Zhun Fan |
AAMAS | 9 |
| 2025 | Long-Term Individual Causal Effect Estimation via Identifiable Latent Representation LearningabstractEstimating long-term causal effects by combining long-term observational and short-term experimental data is a crucial but challenging problem in many real-world scenarios. In existing methods, several ideal assumptions, e.g. latent unconfoundedness assumption or additive equi-confounding bias assumption, are proposed to address the latent confounder problem raised by the observational data. However, in real-world applications, these assumptions are typically violated which limits their practical effectiveness. In this paper, we tackle the problem of estimating the long-term individual causal effects without the aforementioned assumptions. Specifically, we propose to utilize the natural heterogeneity of data, such as data from multiple sources, to identify latent confounders, thereby significantly avoiding reliance on idealized assumptions. Practically, we devise a latent representation learning-based estimator of long-term causal effects. Theoretically, we establish the identifiability of latent confounders, with which we further achieve long-term effect identification. Extensive experimental studies, conducted on multiple synthetic and semi-synthetic datasets, demonstrate the effectiveness of our proposed method. Ruichu Cai, Junjie Wan, Weilin Chen 0001, Zeqin Yang, Zijian Li 0001, Peng Zhen 0001, Jiecheng Guo |
IJCAI | 1 |
| 2025 | Causal View of Time Series Imputation: Some Identification Results on Missing MechanismabstractTime series imputation is one of the most challenging problems and has broad applications in various fields like health care and the Internet of Things. Existing methods mainly aim to model the temporally latent dependencies and the generation process from the observed time series data. In real-world scenarios, different types of missing mechanisms, like MAR (Missing At Random) and MNAR (Missing Not At Random), can occur in time series data. However, existing methods often overlook the difference among the aforementioned missing mechanisms and use a single model for time series imputation, which can easily lead to misleading results due to mechanism mismatching. In this paper, we propose a framework for the time series imputation problem by exploring Different Missing Mechanisms (DMM in short) and tailoring solutions accordingly. Specifically, we first analyze the data generation processes with temporal latent states and missing cause variables for different mechanisms. Sequentially, we model these generation processes via variational inference and estimate prior distributions of latent variables via a normalizing flow-based neural architecture. Furthermore, we establish identifiability results under the nonlinear independent component analysis framework to show that latent variables are identifiable. Experimental results show that our method surpasses existing time series imputation techniques across various datasets with different missing mechanisms, demonstrating its effectiveness in real-world applications. Ruichu Cai, Kaitao Zheng, Junxian Huang 0002, Zijian Li 0001, Zhengming Chen 0002, Zhifeng Hao 0004 |
IJCAI | 1 |
| 2025 | Causal-aware Large Language Models: Enhancing Decision-Making Through Learning, Adapting and ActingabstractLarge language models (LLMs) have shown great potential in decision-making due to the vast amount of knowledge stored within the models.However, these pre-trained models are prone to lack reasoning abilities and are difficult to adapt to new environments, further hindering their application to complex real-world tasks. To address these challenges, inspired by the human cognitive process, we propose Causal-Aware LLMs, which integrate the structural causal model (SCM) into the decision-making process to model, update, and utilize structured knowledge of the environment in a "learning-adapting-acting" paradigm.Specifically, in the learning stage, we first utilize an LLM to extract the environment-specific causal entities and their causal relations to initialize a structured causal model of the environment. Subsequently, in the adapting stage, we update the structured causal model through external feedback about the environment, via an idea of causal intervention. Finally, in the acting stage, Causal-Aware LLMs exploit structured causal knowledge for more efficient policy-making through the reinforcement learning agent. The above processes are performed iteratively to learn causal knowledge, ultimately enabling the causal-aware LLM to achieve a more accurate understanding of the environment and make more efficient decisions. Experimental results across 22 diverse tasks within the open-world game "Crafter" validate the effectiveness of our proposed method. Haipeng Zhu, Keli Zhang, Junjian Ye, Ruichu Cai |
IJCAI | 8 |
| 2025 | Mitigating Q-Value Overestimation through Latent Causal Modeling in Deep Reinforcement LearningabstractWith the development of deep neural networks, deep reinforcement learning (DRL) has shown great potential in solving complex decision-making problems in high-dimensional environments. Among various DRL methods, Q-value function approaches are well-grounded in theory and exhibit strong practical performance, but often suffer from overestimation of the value function. We find that this overestimation primarily stems from unobserved random factors (latent variables) within the environmental dynamics, while existing solutions fail to address this issue, relying mainly on low-noise and unbiased sampling environments. To address this problem, we introduce latent variables to model unobserved randomness behind the environmental dynamics and employ causal models to capture the underlying data generation mechanisms. Subsequently, we propose a causal Q-value function method based on latent causal models. Specifically, we first design a causal encoding-decoding framework for learning the latent causal models, then utilize the inferred latent variables to construct a causal Q-value function, effectively mitigating the Q-value function overestimation problem. Experimental results show that our method significantly improves the stability of policy learning and the speed of convergence, enhancing the robustness of reinforcement learning in complex dynamic environments. Ruichu Cai, Fuyi Lin, Wei Chen 0103, Jie Qiao, Haipeng Zhu, Zhifeng Hao 0004 |
IJCNN | 1 |
| 2025 | MATOT: A Model-Agnostic Constraint for Time Series Forecasting via Optimal TransportabstractThe conventional mean square error for time series forecasting is a point-wise loss function, which ignores the temporal dependency of forecasting data points and results in unstable predictions. Although other methods involve shape information loss functions, such as dynamic time warping, they assign the same weights to each matching pair and result in suboptimal results when a wrong matching pair is chosen. Besides, although the optimal transport can assign different weights for each matching pair, they easily suffer from false alignment due to the time lag. To solve these challenges, we propose a Model-Agnostic loss function via Temporally Sensitive Optimal Transport (MA-TOT) as a differentiable loss function for time series forecasting, which combines temporally sensitive Wasserstein distance for adaptive matching pair chosen and Gromov-Wasserstein distance for multi-level-similarity-measurement. Extensive experiments of several of the latest time series forecasting models with our loss function on seven real-world benchmark datasets reflect the effectiveness of our method. Ruichu Cai, Zhenhui Yang, Yuguang Yan, Haiqin Huang, Kaitao Zheng, Haozhi Chen, Zhifan Jiang, Zijian Li 0001 |
IJCNN | 1 |
| 2025 | SDFDA: Modeling Spectral Distribution for Time-Series Forecasting Domain AdaptationabstractDomain adaptation for time-series forecasting, which spans several real-world applications, aims to transfer temporal patterns from labeled sources to unlabeled target distributions. Due to the complexity of temporal dependencies, it is difficult for the time domain alignment-based methods to capture the invariant information, hence several methods consider the frequency domain alignment and extract the domain-invariant frequencies. However, these methods usually sacrifice the domain-specific frequencies, which can lead to critical changes in the time domain and further result in suboptimal performance in time-series forecasting tasks. Going beyond frequency alignment, we develop the SDFDA as a solution to preserve the domain-specific frequency with an explicit frequency transformation. Technologically, the proposed method captures frequency information via a transformer-based architecture with minimal change constraints. Moreover, it bridges the relationship of frequency between the source and the target domains with a contrastive Whittle Likelihood restriction. Extensive experiments on benchmark datasets demonstrate that the SDFDA outperforms existing domain adaptation methods for time-series data, showcasing its effectiveness in real-world applications. Daoxin Chen, Ruichu Cai, Haozhi Chen, Zhenhui Yang |
IJCNN | 2 |
| 2025 | Cross-Network Relationship Learning via Optimal Transport for Link PredictionabstractCross-network link prediction aims to predict the existence of edges between nodes in a target network by transferring knowledge from a source network with sufficient annotations. Existing methods mainly focus on associating two networks in a shared embedding space, in which the distribution discrepancy of node embeddings between two networks is reduced. However, besides the discrepancy between embeddings, the structure discrepancy also appears between networks, which has not been well investigated in existing studies. In this paper, we propose to learn cross-network relationship between nodes from different networks, and construct an augmented graph for knowledge transfer based on our learned relationship. To achieve this, we propose an optimal transport model to exploit structure information for learning node association between two networks, and then apply barycentric mapping of the graph to combine structural information of both networks according to the obtained relationship. Based on the constructed graph, we design a graph convolutional network to learn node embeddings for link prediction, so that both feature and structural information are leveraged for knowledge transfer. We conduct experiments on benchmark datasets to show the superiority of our method compared with state-of-the-art methods. Yuguang Yan, Ruichu Cai |
IJCNN | 5 |
| 2025 | PEFT Innovations in Text-to-SQL: Adapter and Prefix Tuning Methods with Structural AwarenessabstractThe goal of the Cross-domain Text-to-SQL task is to accurately translate natural language questions into executable SQL queries, even when applied to previously unseen databases. Recently, full fine-tuning of pretrained T5 models has achieved remarkable results. However, as the model size increases, this paradigm incurs significant computational costs. In this paper, we explore the effectiveness of existing parameter-efficient fine-tuning (PEFT) methods for training the T5 model on the Text-to-SQL task. Recognizing the limitations of PEFT methods in structure learning, we propose two structure-aware variants for Adapter and Prefix tuning methods, integrating a relational graph neural network within their architectures. We conducted extensive experiments to demonstrate the effectiveness of our proposed structure-aware adapter and structure-aware prefix tuning methods, which achieve performance comparable to the state-of-the-art T5-3B model while requiring only about 5.43% of the training parameters needed. Yuyuan Cai, Shaobin Shi, Ruichu Cai |
IJCNN | 5 |
| 2025 | Handling Missing Entities in Zero-Shot Named Entity Recognition: Integrated Recall and Retrieval AugmentationabstractRuichu Cai, Junhao Lu, Zhongjie Chen, Boyan Xu, Zhifeng Hao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ruichu Cai, Junhao Lu, Zhongjie Chen |
NAACL (Long Papers) | 1 |
| 2025 | Track-SQL: Enhancing Generative Language Models with Dual-Extractive Modules for Schema and Context Tracking in Multi-turn Text-to-SQLabstractBingfeng Chen, Shaobin Shi, Yongqi Luo, Boyan Xu, Ruichu Cai, Zhifeng Hao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Bingfeng Chen, Shaobin Shi, Yongqi Luo, Ruichu Cai |
NAACL (Long Papers) | 5 |
| 2025 | Towards Identifiability of Hierarchical Temporal Causal Representation LearningabstractModeling hierarchical latent dynamics behind time series data is critical for capturing temporal dependencies across multiple levels of abstraction in real-world tasks. However, existing temporal causal representation learning methods fail to capture such dynamics, as they fail to recover the joint distribution of hierarchical latent variables from \textit{single-timestep observed variables}.
Interestingly, we find that the joint distribution of hierarchical latent variables can be uniquely determined using three conditionally independent observations. Building on this insight, we propose a Causally Hierarchical Latent Dynamic (CHiLD) identification framework. Our approach first employs temporal contextual observed variables to identify the joint distribution of multi-layer latent variables. Sequentially, we exploit the natural sparsity of the hierarchical structure among latent variables to identify latent variables within each layer. Guided by the theoretical results, we develop a time series generative model grounded in variational inference. This model incorporates a contextual encoder to reconstruct multi-layer latent variables and normalize flow-based hierarchical prior networks to impose the independent noise condition of hierarchical latent dynamics. Empirical evaluations on both synthetic and real-world datasets validate our theoretical claims and demonstrate the effectiveness of CHiLD in modeling hierarchical latent dynamics. Zijian Li 0001, Minghao Fu 0002, Junxian Huang 0002, Yifan Shen 0004, Ruichu Cai, Yuewen Sun, Guangyi Chen 0002, Kun Zhang 0001 |
NeurIPS | 5 |
| 2025 | Online Time Series Forecasting with Theoretical GuaranteesabstractThis paper is concerned with online time series forecasting, where unknown distribution shifts occur over time, i.e., latent variables influence the mapping from historical to future observations. To develop an automated way of online time series forecasting, we propose a Theoretical framework for Online Time-series forecasting (TOT in short) with theoretical guarantees. Specifically, we prove that supplying a forecaster with latent variables tightens the Bayes risk—the benefit endures under estimation uncertainty of latent variables and grows as the latent variables achieve a more precise identifiability. To better introduce latent variables into online forecasting algorithms, we further propose to identify latent variables with minimal adjacent observations. Based on these results, we devise a model-agnostic blueprint by employing a temporal decoder to match the distribution of observed variables and two independent noise estimators to model the causal inference of latent variables and mixing procedures of observed variables, respectively. Experiment results on synthetic data support our theoretical claims. Moreover, plug-in implementations built on several baselines yield general improvement across multiple benchmarks, highlighting the effectiveness in real-world applications. Zijian Li 0001, Changze Zhou, Minghao Fu 0002, Sanjay Manjunath, Guangyi Chen 0002, Yingyao Hu, Ruichu Cai, Kun Zhang 0001 |
NeurIPS | 8 |
| 2025 | Balanced Learning for Incremental Multi-view Clustering
Ming Yin 0002, Ruichu Cai, Xuan Xiong |
PRICAI | 5 |
| 2025 | High-Order Information Embedding Transfer for Clustering with Constrained Laplacian Rank
Wenping Xiong, Guangdong Sun, Ruichu Cai, Minghua Zhao |
PRICAI | 6 |
| 2025 | Learning Disentangled Representation for Multi-Modal Time-Series Sensing SignalsabstractMulti-modal time series data is common in web technologies like the Internet of Things (IoT). Existing methods for multi-modal time series representation learning aim to disentangle the modality-shared and modality-specific latent variables. Although achieving notable performances on downstream tasks, they usually assume an orthogonal latent space. However, the modality-specific and modality-shared latent variables might be dependent on real-world scenarios. Therefore, we propose a general generation process, where the modality-shared and modality-specific latent variables are dependent, and further develop a Multi-modAl TEmporal Disentanglement (MATE) model. Specifically, our MATE model is built on a temporally variational inference architecture with the modality-shared and modality-specific prior networks for the disentanglement of latent variables. Furthermore, we establish identifiability results to show that the extracted representation is disentangled. More specifically, we first achieve the subspace identifiability for modality-shared and modality-specific latent variables by leveraging the pairing of multi-modal data. Then we establish the component-wise identifiability of modality-specific latent variables by employing sufficient changes of historical latent variables. Extensive experimental studies on 12 datasets show a general improvement in different downstream tasks, highlighting the effectiveness of our method in real-world scenarios. Ruichu Cai, Zhifan Jiang, Kaitao Zheng, Zijian Li 0001, Weilin Chen 0001, Xuexin Chen, Yifan Shen 0004, Guangyi Chen 0002, Zhifeng Hao 0004, Kun Zhang 0001 |
WWW | 1 |
| 2025 | EVA-MVC: Equitable View-weight Allocation for Generic Multi-View ClusteringabstractContemporary datasets sourced from the web often adopt a multiview format, collecting data from diverse sources, domains, or modules.Existing methodologies employed to analyze such datasets frequently overlook or inaccurately allocate the view-weights, pivotal metrics reflecting each view's significance.This work introduces EVA-MVC, a simple yet effective algorithm designed for Equitable View-weight Allocation (EVA) seamlessly integrated with arbitrary Multi-view Clustering (MVC) methods.Within the EVA module, we establish theoretical connections between view supplementarity and Multi-view Subspace Learning (MSL), leading to the partition of views into View Communities (VCs) based on these foundational principles.These VCs exhibit internal supplementarity similarities, facilitating Equitable View-weights Allocation through VCspecific MSL.The proposed EVA process precedes and operates independently of traditional or SOTA MVC approaches, requiring no additional processing or specialized design, making it an ideal preprocessing step for MVC applications.Through comprehensive evaluations across diverse multi-view datasets, our findings reveal that our EVA significantly enhances the effectiveness of mainstream MVC frameworks, resulting in a notable performance improvement. Yuan Fang 0001, Xiaofeng Feng, Geping Yang, Ruichu Cai, Yiyang Yang, Zhiguo Gong, Zhifeng Hao 0004 |
WWW | 4 |
| 2025 | Time-aware tensor factorization for temporal recommendation
Yali Feng, Wen Wen 0009, Ruichu Cai |
Appl. Intell. | 4 |
| 2025 | Interpretable high-order knowledge graph neural network for predicting synthetic lethality in human cancersabstractSynthetic lethality (SL) is a promising gene interaction for cancer therapy. Recent SL prediction methods integrate knowledge graphs (KGs) into graph neural networks (GNNs) and employ attention mechanisms to extract local subgraphs as explanations for target gene pairs. However, attention mechanisms often lack fidelity, typically generate a single explanation per gene pair, and fail to ensure trustworthy high-order structures in their explanations. To overcome these limitations, we propose Diverse Graph Information Bottleneck for Synthetic Lethality (DGIB4SL), a KG-based GNN that generates multiple faithful explanations for the same gene pair and effectively encodes high-order structures. Specifically, we introduce a novel DGIB objective, integrating a determinant point process constraint into the standard information bottleneck objective, and employ 13 motif-based adjacency matrices to capture high-order structures in gene representations. Experimental results show that DGIB4SL outperforms state-of-the-art baselines and provides multiple explanations for SL prediction, revealing diverse biological mechanisms underlying SL inference. Xuexin Chen, Ruichu Cai, Zhengting Huang, Zijian Li 0001, Jie Zheng 0002, Min Wu 0008 |
Briefings Bioinform. | 2 |
| 2025 | Temporal latent variable structural causal model for causal discovery under external interferences
Ruichu Cai, Xiaokai Huang, Wei Chen 0103, Zijian Li 0001, Zhifeng Hao 0004 |
Neurocomputing | 1 |
| 2025 | StateHPs: State Hawkes processes for Granger causal discovery from non-stationary event sequencesabstractLearning Granger causality from event sequences has important applications in various scenarios. Many methods have been developed based on Hawkes process with a stationarity assumption. However, these methods often fail in real-world scenarios due to violating the stationarity assumption, as an event sequence can be generated under different states at varying times. Although some work tries to model non-stationarity by searching for best segmentation, they still suffer from the lack of robustness and identification guarantee. An intuitive solution is to model the non-stationary generation process in a unified probabilistic generative framework. This presents two significant challenges: how to model the generation process considering both the stationarity of each subsequence and the non-stationarity among the subsequences, and how to identify the Granger causality. To address these challenges, we devise State Hawkes Processes (StateHPs). For the first challenge, StateHPs formulates the state assignments of each subsequence as a Dirichlet distribution and each state as a Hawkes process. For the second challenge, StateHPs introduces a variational Expectation-Maximization algorithm to identify the Granger causal graph. We also develop the identification theories for StateHPs. On real-world data, StateHPs achieves 35.5%, 33.9%, and 36.7% improvement among F1, Precision, and Recall metrics compared to the SOTA baselines. Yuequn Liu, Guangdong Sun, Ruichu Cai, Zijian Li 0001, Keli Zhang, Lujia Pan, Zhifeng Hao 0004 |
Inf. Sci. | 3 |
| 2025 | On the probability of necessity and sufficiency of explaining Graph Neural Networks: A lower bound optimization approach
Ruichu Cai, Yuxuan Zhu 0001, Xuexin Chen, Yuan Fang 0001, Min Wu 0008, Jie Qiao, Zhifeng Hao 0004 |
Neural Networks | 1 |
| 2025 | Unifying invariant and variant features for graph out-of-distribution via probability of necessity and sufficiency
Xuexin Chen, Ruichu Cai, Kaitao Zheng, Zhifan Jiang, Zhengting Huang, Zhifeng Hao 0004, Zijian Li 0001 |
Neural Networks | 2 |
| 2025 | Identifying Semantic Component for Robust Molecular Property PredictionabstractAlthough graph neural networks have achieved great success in the task of molecular property prediction in recent years, their generalization ability under out-of-distribution (OOD) settings is still under-explored. Most of the existing methods rely on learning discriminative representations for prediction, often assuming that the underlying semantic components are correctly identified. However, this assumption does not always hold, leading to potential misidentifications that affect model robustness. Different from these discriminative-based methods, we propose a generative model to ensure the Semantic-Components Identifiability, named SCI. We demonstrate that the latent variables in this generative model can be explicitly identified into semantic-relevant (SR) and semantic-irrelevant (SI) components, which contributes to better OOD generalization by involving minimal change properties of causal mechanisms. Specifically, we first formulate the data generation process from the atom level to the molecular level, where the latent space is split into SI substructures, SR substructures, and SR atom variables. Sequentially, to reduce misidentification, we restrict the minimal changes of the SR atom variables and add a semantic latent substructure regularization to mitigate the variance of the SR substructure under augmented domain changes. Under mild assumptions, we prove the block-wise identifiability of the SR substructure and the comment-wise identifiability of SR atom variables. Experimental studies achieve state-of-the-art performance and show general improvement on 21 datasets in 3 mainstream benchmarks. Moreover, the visualization results of the proposed SCI method provide insightful case studies and explanations for the prediction results. Zijian Li 0001, Zunhong Xu, Ruichu Cai, Zhenhui Yang, Yuguang Yan, Zhifeng Hao 0004, Guangyi Chen 0002, Kun Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | GNNSynergy: A Multi-View Graph Neural Network for Predicting Anti-Cancer Drug SynergyabstractDrug combinations play very important roles in cancer therapy, as they can enhance curative efficacy and overcome drug resistance. Due to the increasing size of combinatorial space, experimental screening for all the drug combinations becomes infeasible in practice. Therefore, there is a great need to develop accurate computational approaches that can predict potential drug combinations to direct the experimental screening. In this paper, we propose a novel method called GNNSynergy to learn drug embeddings for drug synergy prediction. Given a specific cancer cell line, we propose a multi-view graph neural network framework which considers the current cell line as main view while other cell lines from the same tissue as sub-views. In each view, we first construct different graphs to describe drug synergistic and antagonistic interactions, and adopt graph neural network as encoder to learn drug embeddings. We further combine both the main view and sub-views via an attention mechanism to derive the final drug embeddings for drug synergy prediction. We perform extensive experiments on DrugComb database and the experimental results demonstrate that our proposed GNNSynergy significantly outperforms state-of-the-art methods for novel synergistic drug combination prediction. Zhifeng Hao 0004, Jianming Zhan 0003, Yuan Fang 0001, Min Wu 0008, Ruichu Cai |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | Modeling Multi-Seasonal Multi-Behavior Dependency for Temporal RecommendationabstractMining temporal patterns from user behaviors has long been investigated, but most of the existing work centers on single-type user–item interactions, such as purchase or click, which fails to take advantage of the user’s diversified interests revealed by various types of behavior. However, capturing patterns from different behavior sequences and modeling the complex inter-correlation between them are non-trivial tasks, as the high sparsity of type-related interactions, multi-seasonality of individual behaviors, and time-variant dependency of multi-type activities make it really challenging. To address these challenges, we propose a novel framework that aims to model the M ulti-Seasonal M ulti-Behavior Dep endencies (MMDep) both within and across the multi-type behavior sequences. In the proposed model, an item co-occurrence matrix factorization strategy is introduced to alleviate the sparsity issue in type-related behavior sequences. And a temporal dependency module that incorporates multi-scale EMA mechanism is utilized to capture the multi-seasonal dependencies within individual sequences. Moreover, a cross-behavior dependency module is employed to learn the time-variant dependency among different behaviors. Extensive experiments on three real-world datasets demonstrate that the proposed MMDep performs significantly better than the state-of-the-art baselines. And it may provide some new insights and tools on how to leverage multi-behavior data for better temporal recommendation. Shichao Liang, Wen Wen 0009, Yali Feng, Ruichu Cai, Zhifeng Hao 0002 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2025 | MSC-DOLES: Multi-View Subspace Clustering in Diverse Orthogonal Latent Embedding SpacesabstractIn the domain of Multi-view Subspace Clustering (MSC) in Latent Embedding Space (LES), existing methods aim to capture and leverage critical multi-view information by mapping it into a low-dimensional LES. However, several aspects can be further improved: (i) Fusion Strategy: Existing methods adopt either early fusion or late fusion to integrate multi-view information, limiting the effectiveness of the fusion. (ii) Diversity: Current methods often overlook the inherent diversity in the multi-view data by focusing on a single LES. (iii) Efficiency: LES-based methods exhibit high computational complexity, with cubic time and quadratic space requirements based on the number of samples. To address these issues, we propose a novel framework called MSC-DOLES (Multi-view Subspace Clustering in Diverse Orthogonal Latent Embedding Spaces), a novel framework designed to tackle these challenges. MSC-DOLES incorporates a two-stage fusion approach that generates and learns from multiple LES to maximize cross-view diversity. Orthogonality constraints on individual LES ensure view-internal diversity, resulting in a set of Diverse Orthogonal Latent Embedding Spaces (DOLES). The DOLES are then fused into a consensus anchor graph using learnable anchors. The final clustering is induced by partitioning the obtained graph without pre-processing. We develop an eight-step optimization algorithm for MSC-DOLES, which exhibits nearly linear time and space complexities relative to the number of samples. Extensive experiments demonstrate the superiority of MSC-DOLES over state-of-the-art methods. Yuan Fang 0001, Geping Yang, Ruichu Cai, Yiyang Yang, Zhiguo Gong, Zhifeng Hao 0004 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Toward Targeted Mining of RFM PatternsabstractIn today's era of information overload, leveraging data mining techniques to understand and analyze customer behavior has become essential for businesses. Among these techniques, the recency, frequency, and monetary value analysis model serves as a powerful tool for customer segmentation, enabling companies to identify high-value customers. However, traditional recency, frequency, and monetary (RFM) models do not focus on user-specific targets, often struggling to meet the increasing demands for personalization and efficiency. To address this challenge, this article introduces the concept of target RFM patterns, which must satisfy the three dimensions of recency, frequency, and utility while aligning with user interests. Based on this concept, we formulate the problem of mining target RFM patterns. More importantly, we define a mining order, called TaRFM order, and propose an efficient algorithm called TaRFM. This new algorithm is optimized through three pruning strategies based on the TaRFM order, which not only eliminates a significant number of invalid operations, thereby reducing pattern generation, but also accurately extracts all TaRFM patterns without requiring postprocessing techniques. Finally, extensive experiments conducted on multiple datasets demonstrate the accuracy and efficiency of the TaRFM algorithm. Xiaoye Chen, Wensheng Gan, Jian Zhu 0001, Ruichu Cai, Philip S. Yu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | On the Role of Entropy-Based Loss for Learning Causal Structure With Continuous OptimizationabstractCausal discovery from observational data is an important but challenging task in many scientific fields. A recent line of work formulates the structure learning problem as a continuous constrained optimization task using an algebraic characterization of directed acyclic graphs (DAGs) and the least-square loss function. Though the least-square loss function is well justified under the standard Gaussian noise assumption, it is limited if the assumption does not hold. In this work, we theoretically show that the violation of the Gaussian noise assumption will hinder the causal direction identification, making the causal orientation fully determined by the causal strength as well as the variances of noises in the linear case and by the strong non-Gaussian noises in the nonlinear case. Consequently, we propose a more general entropy-based loss that is theoretically consistent with the likelihood score under any noise distribution. We run extensive empirical evaluations on both synthetic data and real-world data to validate the effectiveness of the proposed method and show that our method achieves the best in structure Hamming distance, false discovery rate (FDR), and true-positive rate (TPR) matrices. Weilin Chen 0001, Jie Qiao, Ruichu Cai |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Testing Conditional Independence Between Latent Variables by Independence ResidualsabstractConditional independence (CI) testing is an important problem, especially in causal discovery. Most testing methods assume that all variables are fully observable and then test the CI among the observed data. Such an assumption is often untenable beyond applications dealing with, e.g., psychological analysis about the mental health status and medical diagnosing (researchers need to consider the existence of latent variables in these scenarios); and typically adopted latent CI test schemes mainly suffer from robust or efficient issues. Accordingly, this article investigates the problem of testing CI between latent variables. To this end, we offer an auxiliary regression-based CI (AReCI) test by taking the measured variable as the surrogate variable of the latent variables to conduct the regression over the latent variables under the linear causal models, in which each latent variable has some certain measured variables. Specifically, given a pair of latent variables$L_X$and$L_Y$, and a corresponding latent variable set$\mathcal{L}_{O}$,$L_X \CI L_Y | \mathcal{L}_{O}$holds if and only if$A_{\{L_X\}}-\omega_1^\intercal A^{\prime}_{\{\mathcal{L}_{O}\}}$and$A_{\{L_Y\}}-\omega_2^\intercal A^{\prime\prime}_{\{\mathcal{L}_{O}\}}$are statistically independent, where$A^{\prime}$and$A^{\prime\prime}$are the two disjoint subset of the measured variable for the corresponding latent variables,$A^{\prime}_{\{\mathcal{L}_{O}\}} \cap A^{\prime\prime}_{\{\mathcal{L}_{O}\}} =\emptyset$, and$\omega_1$is a parameter vector characterized from the cross covariance between$A_{\{L_X\}}$and$A^{\prime}_{\{\mathcal{L}_{O}\}}$, and$\omega_{2}$is a parameter vector characterized from the cross covariance between$A_{\{L_Y\}}$and$A^{\prime\prime}_{\{\mathcal{L}_{O}\}}$. We theoretically show that the AReCI test is capable of addressing both Gaussian and non-Gaussian data. In addition, we find that the well-known partial correlation test can be seen as a special case of the AReCI test. Finally, we devise a causal discovery method by using the AReCI test as the CI test. The experimental results on synthetic and real-world data illustrate the effectiveness of our method. Zhengming Chen 0002, Jie Qiao, Feng Xie 0002, Ruichu Cai, Zhifeng Hao 0004, Keli Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | A Survey on Causal Reinforcement LearningabstractWhile reinforcement learning (RL) achieves tremendous success in sequential decision-making problems of many domains, it still faces key challenges of data inefficiency and the lack of interpretability. Interestingly, many researchers have leveraged insights from the causality literature recently, bringing forth flourishing works to unify the merits of causality and address well the challenges from RL. As such, it is of great necessity and significance to collate these causal RL (CRL) works, offer a review of CRL methods, and investigate the potential functionality from causality toward RL. In particular, we divide the existing CRL approaches into two categories according to whether their causality-based information is given in advance or not. We further analyze each category in terms of the formalization of different models, ranging from the Markov decision process (MDP), partially observed MDP (POMDP), multiarmed bandits (MABs), imitation learning (IL), and dynamic treatment regime (DTR). Each of them represents a distinct type of causal graphical illustration. Moreover, we summarize the evaluation matrices and open sources, while we discuss emerging applications, along with promising prospects for the future development of CRL. Yan Zeng 0002, Ruichu Cai, Fuchun Sun 0001, Libo Huang 0001, Zhifeng Hao 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Lightweight image super-resolution via an efficient local-global transformer network with adaptive attention windows
Simin Zheng, Peiyao Chen, Jianfang Hu, Ruichu Cai |
Vis. Comput. | 6 |
| 2024 | Identification of Causal Structure with Latent Variables Based on Higher Order CumulantsabstractCausal discovery with latent variables is a crucial but challenging task. Despite the emergence of numerous methods aimed at addressing this challenge, they are not fully identified to the structure that two observed variables are influenced by one latent variable and there might be a directed edge in between. Interestingly, we notice that this structure can be identified through the utilization of higher-order cumulants. By leveraging the higher-order cumulants of non-Gaussian data, we provide an analytical solution for estimating the causal coefficients or their ratios. With the estimated (ratios of) causal coefficients, we propose a novel approach to identify the existence of a causal edge between two observed variables subject to latent variable influence. In case when such a causal edge exits, we introduce an asymmetry criterion to determine the causal direction. The experimental results demonstrate the effectiveness of our proposed method. Wei Chen 0103, Zhiyi Huang 0008, Ruichu Cai, Zhifeng Hao 0004, Kun Zhang 0001 |
AAAI | 3 |
| 2024 | Where and How to Attack? A Causality-Inspired Recipe for Generating Counterfactual Adversarial ExamplesabstractDeep neural networks (DNNs) have been demonstrated to be vulnerable to well-crafted adversarial examples, which are generated through either well-conceived L_p-norm restricted or unrestricted attacks. Nevertheless, the majority of those approaches assume that adversaries can modify any features as they wish, and neglect the causal generating process of the data, which is unreasonable and unpractical. For instance, a modification in income would inevitably impact features like the debt-to-income ratio within a banking system. By considering the underappreciated causal generating process, first, we pinpoint the source of the vulnerability of DNNs via the lens of causality, then give theoretical results to answer where to attack. Second, considering the consequences of the attack interventions on the current state of the examples to generate more realistic adversarial examples, we propose CADE, a framework that can generate Counterfactual ADversarial Examples to answer how to attack. The empirical results demonstrate CADE's effectiveness, as evidenced by its competitive performance across diverse attack scenarios, including white-box, transfer-based, and random intervention attacks. Ruichu Cai, Yuxuan Zhu 0001, Jie Qiao, Zefeng Liang, Furui Liu, Zhifeng Hao 0004 |
AAAI | 1 |
| 2024 | TNPAR: Topological Neural Poisson Auto-Regressive Model for Learning Granger Causal Structure from Event SequencesabstractLearning Granger causality from event sequences is a challenging but essential task across various applications. Most existing methods rely on the assumption that event sequences are independent and identically distributed (i.i.d.). However, this i.i.d. assumption is often violated due to the inherent dependencies among the event sequences. Fortunately, in practice, we find these dependencies can be modeled by a topological network, suggesting a potential solution to the non-i.i.d. problem by introducing the prior topological network into Granger causal discovery. This observation prompts us to tackle two ensuing challenges: 1) how to model the event sequences while incorporating both the prior topological network and the latent Granger causal structure, and 2) how to learn the Granger causal structure. To this end, we devise a unified topological neural Poisson auto-regressive model with two processes. In the generation process, we employ a variant of the neural Poisson process to model the event sequences, considering influences from both the topological network and the Granger causal structure. In the inference process, we formulate an amortized inference algorithm to infer the latent Granger causal structure. We encapsulate these two processes within a unified likelihood function, providing an end-to-end framework for this task. Experiments on simulated and real-world data demonstrate the effectiveness of our approach. Yuequn Liu, Ruichu Cai, Wei Chen 0103, Jie Qiao, Yuguang Yan, Zijian Li 0001, Keli Zhang, Zhifeng Hao 0004 |
AAAI | 2 |
| 2024 | Identification of Causal Structure in the Presence of Missing Data with Additive Noise ModelabstractMissing data are an unavoidable complication frequently encountered in many causal discovery tasks. When a missing process depends on the missing values themselves (known as self-masking missingness), the recovery of the joint distribution becomes unattainable, and detecting the presence of such self-masking missingness remains a perplexing challenge. Consequently, due to the inability to reconstruct the original distribution and to discern the underlying missingness mechanism, simply applying existing causal discovery methods would lead to wrong conclusions. In this work, we found that the recent advances additive noise model has the potential for learning causal structure under the existence of the self-masking missingness. With this observation, we aim to investigate the identification problem of learning causal structure from missing data under an additive noise model with different missingness mechanisms, where the `no self-masking missingness' assumption can be eliminated appropriately. Specifically, we first elegantly extend the scope of identifiability of causal skeleton to the case with weak self-masking missingness (i.e., no other variable could be the cause of self-masking indicators except itself). We further provide the sufficient and necessary identification conditions of the causal direction under additive noise model and show that the causal structure can be identified up to an IN-equivalent pattern. We finally propose a practical algorithm based on the above theoretical results on learning the causal skeleton and causal direction. Extensive experiments on synthetic and real data demonstrate the efficiency and effectiveness of the proposed algorithms. Jie Qiao, Zhengming Chen 0002, Jianhua Yu, Ruichu Cai, Zhifeng Hao 0004 |
AAAI | 4 |
| 2024 | Causal Discovery from Poisson Branching Structural Causal Model Using High-Order Cumulant with Path AnalysisabstractCount data naturally arise in many fields, such as finance, neuroscience, and epidemiology, and discovering causal structure among count data is a crucial task in various scientific and industrial scenarios. One of the most common characteristics of count data is the inherent branching structure described by a binomial thinning operator and an independent Poisson distribution that captures both branching and noise. For instance, in a population count scenario, mortality and immigration contribute to the count, where survival follows a Bernoulli distribution, and immigration follows a Poisson distribution. However, causal discovery from such data is challenging due to the non-identifiability issue: a single causal pair is Markov equivalent, i.e., X->Y and Y->X are distributed equivalent. Fortunately, in this work, we found that the causal order from X to its child Y is identifiable if X is a root vertex and has at least two directed paths to Y, or the ancestor of X with the most directed path to X has a directed path to Y without passing X. Specifically, we propose a Poisson Branching Structure Causal Model (PB-SCM) and perform a path analysis on PB-SCM using high-order cumulants. Theoretical results establish the connection between the path and cumulant and demonstrate that the path information can be obtained from the cumulant. With the path information, causal order is identifiable under some graphical conditions. A practical algorithm for learning causal structure under PB-SCM is proposed and the experiments demonstrate and verify the effectiveness of the proposed method. Jie Qiao, Zhengming Chen 0002, Ruichu Cai, Zhifeng Hao 0004 |
AAAI | 4 |
| 2024 | Hypergraph Joint Representation Learning for Hypervertices and Hyperedges via Cross ExpansionabstractHypergraph captures high-order information in structured data and obtains much attention in machine learning and data mining. Existing approaches mainly learn representations for hypervertices by transforming a hypergraph to a standard graph, or learn representations for hypervertices and hyperedges in separate spaces. In this paper, we propose a hypergraph expansion method to transform a hypergraph to a standard graph while preserving high-order information. Different from previous hypergraph expansion approaches like clique expansion and star expansion, we transform both hypervertices and hyperedges in the hypergraph to vertices in the expanded graph, and construct connections between hypervertices or hyperedges, so that richer relationships can be used in graph learning. Based on the expanded graph, we propose a learning model to embed hypervertices and hyperedges in a joint representation space. Compared with the method of learning separate spaces for hypervertices and hyperedges, our method is able to capture common knowledge involved in hypervertices and hyperedges, and also improve the data efficiency and computational efficiency. To better leverage structure information, we minimize the graph reconstruction loss to preserve the structure information in the model. We perform experiments on both hypervertex classification and hyperedge classification tasks to demonstrate the effectiveness of our proposed method. Yuguang Yan, Hanrui Wu, Ruichu Cai |
AAAI | 5 |
| 2024 | An Optimal Transport View for Subspace Clustering and Spectral ClusteringabstractClustering is one of the most fundamental problems in machine learning and data mining, and many algorithms have been proposed in the past decades. Among them, subspace clustering and spectral clustering are the most famous approaches. In this paper, we provide an explanation for subspace clustering and spectral clustering from the perspective of optimal transport. Optimal transport studies how to move samples from one distribution to another distribution with minimal transport cost, and has shown a powerful ability to extract geometric information. By considering a self optimal transport model with only one group of samples, we observe that both subspace clustering and spectral clustering can be explained in the framework of optimal transport, and the optimal transport matrix bridges the spaces of features and spectral embeddings. Inspired by this connection, we propose a spectral optimal transport barycenter model, which learns spectral embeddings by solving a barycenter problem equipped with an optimal transport discrepancy and guidance of data. Based on our proposed model, we take advantage of optimal transport to exploit both feature and metric information involved in data for learning coupled spectral embeddings and affinity matrix in a unified model. We develop an alternating optimization algorithm to solve the resultant problems, and conduct experiments in different settings to evaluate the performance of our proposed methods. Yuguang Yan, Canlin Yang, Jie Zhang 0124, Ruichu Cai, Michael Kwok-Po Ng |
AAAI | 5 |
| 2024 | Exploiting Geometry for Treatment Effect Estimation via Optimal TransportabstractEstimating treatment effects from observational data suffers from the issue of confounding bias, which is induced by the imbalanced confounder distributions between the treated and control groups. As an effective approach, re-weighting learns a group of sample weights to balance the confounder distributions. Existing methods of re-weighting highly rely on a propensity score model or moment alignment. However, for complex real-world data, it is difficult to obtain an accurate propensity score prediction. Although moment alignment is free of learning a propensity score model, accurate estimation for high-order moments is computationally difficult and still remains an open challenge, and first and second-order moments are insufficient to align the distributions and easy to be misled by outliers. In this paper, we exploit geometry to capture the intrinsic structure involved in data for balancing the confounder distributions, so that confounding bias can be reduced even with outliers. To achieve this, we construct a connection between treatment effect estimation and optimal transport, a powerful tool to capture geometric information. After that, we propose an optimal transport model to learn sample weights by extracting geometry from confounders, in which geometric information between groups and within groups is leveraged for better confounder balancing. A projected mirror descent algorithm is employed to solve the derived optimization problem. Experimental studies on both synthetic and real-world datasets demonstrate the effectiveness of our proposed method. Yuguang Yan, Zeqin Yang, Weilin Chen 0001, Ruichu Cai, Michael Kwok-Po Ng |
AAAI | 4 |
| 2024 | S²GSL: Incorporating Segment to Syntactic Enhanced Graph Structure Learning for Aspect-based Sentiment AnalysisabstractPrevious graph-based approaches in Aspectbased Sentiment Analysis(ABSA) have demonstrated impressive performance by utilizing graph neural networks and attention mechanisms to learn structures of static dependency trees and dynamic latent trees.However, incorporating both semantic and syntactic information simultaneously within complex global structures can introduce irrelevant contexts and syntactic dependencies during the process of graph structure learning, potentially resulting in inaccurate predictions.In order to address the issues above, we propose S 2 GSL, incorporating Segment to Syntactic enhanced Graph Structure Learning for ABSA.Specifically, S 2 GSL is featured with a segment-aware semantic graph learning and a syntax-based latent graph learning enabling the removal of irrelevant contexts and dependencies, respectively.We further propose a self-adaptive aggregation network that facilitates the fusion of two graph learning branches, thereby achieving complementarity across diverse structures.Experimental results on four benchmarks demonstrate the effectiveness of our framework. Bingfeng Chen, Qihan Ouyang, Yongqi Luo, Ruichu Cai |
ACL (1) | 5 |
| 2024 | Automating the Selection of Proxy Variables of Unmeasured ConfoundersabstractRecently, interest has grown in the use of proxy variables of unobserved confounding for inferring the causal effect in the presence of unmeasured confounders from observational data. One difficulty inhibiting the practical use is finding valid proxy variables of unobserved confounding to a target causal effect of interest. These proxy variables are typically justified by background knowledge. In this paper, we investigate the estimation of causal effects among multiple treatments and a single outcome, all of which are affected by unmeasured confounders, within a linear causal model, without prior knowledge of the validity of proxy variables. To be more specific, we first extend the existing proxy variable estimator, originally addressing a single unmeasured confounder, to accommodate scenarios where multiple unmeasured confounders exist between the treatments and the outcome. Subsequently, we present two different sets of precise identifiability conditions for selecting valid proxy variables of unmeasured confounders, based on the second-order statistics and higher-order statistics of the data, respectively. Moreover, we propose two data-driven methods for the selection of proxy variables and for the unbiased estimation of causal effects. Theoretical analysis demonstrates the correctness of our proposed algorithms. Experimental results on both synthetic and real-world data show the effectiveness of the proposed approach. Feng Xie 0002, Zhengming Chen 0002, Shanshan Luo, Wang Miao, Ruichu Cai, Zhi Geng |
ICML | 5 |
| 2024 | Feature Attribution with Necessity and Sufficiency via Dual-stage Perturbation Test for Causal ExplanationabstractWe investigate the problem of explainability for machine learning models, focusing on Feature Attribution Methods (FAMs) that evaluate feature importance through perturbation tests. Despite their utility, FAMs struggle to distinguish the contributions of different features, when their prediction changes are similar after perturbation. To enhance FAMs’ discriminative power, we introduce Feature Attribution with Necessity and Sufficiency (FANS), which find a neighborhood of the input such that perturbing samples within this neighborhood have a high Probability of being Necessity and Sufficiency (PNS) cause for the change in predictions, and use this PNS as the importance of the feature. Specifically, FANS compute this PNS via a heuristic strategy for estimating the neighborhood and a perturbation test involving two stages (factual and interventional) for counterfactual reasoning. To generate counterfactual samples, we use a resampling-based approach on the observed samples to approximate the required conditional distribution. We demonstrate that FANS outperforms existing attribution methods on six benchmarks. Please refer to the source code via https://github.com/DMIRLAB-Group/FANS. Xuexin Chen, Ruichu Cai, Zhengting Huang, Yuxuan Zhu 0001, Julien Horwood, Zhifeng Hao 0004, Zijian Li 0001, José Miguel Hernández-Lobato |
ICML | 2 |
| 2024 | Doubly Robust Causal Effect Estimation under Networked Interference via Targeted LearningabstractCausal effect estimation under networked interference is an important but challenging problem. Available parametric methods are limited in their model space, while previous semiparametric methods, e.g., leveraging neural networks to fit only one single nuisance function, may still encounter misspecification problems under networked interference without appropriate assumptions on the data generation process. To mitigate bias stemming from misspecification, we propose a novel doubly robust causal effect estimator under networked interference, by adapting the targeted learning technique to the training of neural networks. Specifically, we generalize the targeted learning technique into the networked interference setting and establish the condition under which an estimator achieves double robustness. Based on the condition, we devise an end-to-end causal effect estimator by transforming the identified theoretical condition into a targeted loss. Moreover, we provide a theoretical analysis of our designed estimator, revealing a faster convergence rate compared to a single nuisance model. Extensive experimental results on two real-world networks with semisynthetic data demonstrate the effectiveness of our proposed estimators. Weilin Chen 0001, Ruichu Cai, Zeqin Yang, Jie Qiao, Yuguang Yan, Zijian Li 0001, Zhifeng Hao 0004 |
ICML | 2 |
| 2024 | Reducing Balancing Error for Causal Inference via Optimal TransportabstractMost studies on causal inference tackle the issue of confounding bias by reducing the distribution shift between the control and treated groups. However, it remains an open question to adopt an appropriate metric for distribution shift in practice. In this paper, we define a generic balancing error on reweighted samples to characterize the confounding bias, and study the connection between the balancing error and the Wasserstein discrepancy derived from the theory of optimal transport. We not only regard the Wasserstein discrepancy as the metric of distribution shift, but also explore the association between the balancing error and the underlying cost function involved in the Wasserstein discrepancy. Motivated by this, we propose to reduce the balancing error under the framework of optimal transport with learnable marginal distributions and the cost function, which is implemented by jointly learning weights and representations associated with factual outcomes. The experiments on both synthetic and real-world datasets demonstrate the effectiveness of our proposed method. Yuguang Yan, Zeqin Yang, Weilin Chen 0001, Ruichu Cai |
ICML | 5 |
| 2024 | Individual Causal Structure Learning from Population Data
Wei Chen 0103, Xiaokai Huang, Zijian Li 0001, Ruichu Cai, Zhiyi Huang 0008, Zhifeng Hao 0004 |
IJCAI | 4 |
| 2024 | Combinatorial Routing for Neural Trees
Ruichu Cai, Yuguang Yan |
IJCAI | 2 |
| 2024 | Learning Discrete Latent Variable Structures with Tensor Rank ConditionsabstractUnobserved discrete data are ubiquitous in many scientific disciplines, and how to learn the causal structure of these latent variables is crucial for uncovering data patterns. Most studies focus on the linear latent variable model or impose strict constraints on latent structures, which fail to address cases in discrete data involving non-linear relationships or complex latent structures. To achieve this, we explore a tensor rank condition on contingency tables for an observed variable set $\mathbf{X}_p$, showing that the rank is determined by the minimum support of a specific conditional set (not necessary in $\mathbf{X}_p$) that d-separates all variables in $\mathbf{X}_p$. By this, one can locate the latent variable through probing the rank on different observed variables set, and further identify the latent causal structure under some structure assumptions. We present the corresponding identification algorithm and conduct simulated experiments to verify the effectiveness of our method. In general, our results elegantly extend the identification boundary for causal discovery with discrete latent variables and expand the application scope of causal discovery with latent variables. Zhengming Chen 0002, Ruichu Cai, Feng Xie 0002, Jie Qiao, Anpeng Wu, Zijian Li 0001, Zhifeng Hao 0004, Kun Zhang 0001 |
NeurIPS | 2 |
| 2024 | On the Identifiability of Poisson Branching Structural Causal Model Using Probability Generating FunctionabstractCausal discovery from observational data, especially for count data, is essential across scientific and industrial contexts, such as biology, economics, and network operation maintenance. For this task, most approaches model count data using Bayesian networks or ordinal relations. However, they overlook the inherent branching structures that are frequently encountered, e.g., a browsing event might trigger an adding cart or purchasing event. This can be modeled by a binomial thinning operator (for branching) and an additive independent Poisson distribution (for noising), known as Poisson Branching Structure Causal Model (PB-SCM). There is a provably sound cumulant-based causal discovery method that allows the identification of the causal structure under a branching structure. However, we show that there still remains a gap in that there exist causal directions that are identifiable while the algorithm fails to identify them. In this work, we address this gap by exploring the identifiability of PB-SCM using the Probability Generating Function (PGF). By developing a compact and exact closed-form solution for the PGF of PB-SCM, we demonstrate that each component in this closed-form solution uniquely encodes a specific local structure, enabling the identification of the local structures by testing their corresponding component appearances in the PGF. Building on this, we propose a practical algorithm for learning causal skeletons and identifying causal directions of PB-SCM using PGF. The effectiveness of our method is demonstrated through experiments on both synthetic and real datasets. Jie Qiao, Zefeng Liang, Zihuai Zeng, Ruichu Cai |
NeurIPS | 5 |
| 2024 | Granger causal representation learning for groups of time series
Ruichu Cai, Yunjin Wu, Xiaokai Huang, Wei Chen 0103, Tom Z. J. Fu, Zhifeng Hao 0004 |
Sci. China Inf. Sci. | 1 |
| 2024 | Generalized Independent Noise Condition for Estimating Causal Structure with Latent VariablesabstractWe investigate the challenging task of learning causal structure in the presence of latent variables, including locating latent variables, determining their quantity, and identifying causal relationships among both latent and observed variables. To address this, we propose a Generalized Independent Noise (GIN) condition for linear non-Gaussian acyclic causal models that incorporate latent variables, which establishes the independence between a linear combination of certain measured variables and some other measured variables. Specifically, for two observed random vectors $\bf{Y}$ and $\bf{Z}$, GIN holds if and only if $\omega^{\intercal}\mathbf{Y}$ and $\mathbf{Z}$ are statistically independent, where $\omega$ is a non-zero parameter vector determined by the cross-covariance between $\mathbf{Y}$ and $\mathbf{Z}$. We then give necessary and sufficient graphical criteria of the GIN condition in linear non-Gaussian acyclic causal models. From a graphical perspective, roughly speaking, GIN implies the existence of a set $\mathcal{S}$ such that $\mathcal{S}$ is causally earlier (w.r.t. the causal ordering) than $\mathbf{Y}$, and that every active (collider-free) path between $\mathbf{Y}$ and $\mathbf{Z}$ must contain a node from $\mathcal{S}$. Interestingly, we find that the independent noise condition (i.e., if there is no confounder, causes are independent of the residual derived from regressing the effect on the causes) can be seen as a special case of GIN. With such a connection between GIN and latent causal structures, we further leverage the proposed GIN condition, together with a well-designed search procedure, to efficiently estimate Linear, Non-Gaussian Latent Hierarchical Models (LiNGLaHs), where latent confounders may also be causally related and may even follow a hierarchical structure. We show that the underlying causal structure of a LiNGLaH is identifiable in light of GIN conditions under mild assumptions. Experimental results on both synthetic and three real-world data sets show the effectiveness of the proposed approach. Feng Xie 0002, Biwei Huang, Zhengming Chen 0002, Ruichu Cai, Clark Glymour, Zhi Geng, Kun Zhang 0001 |
J. Mach. Learn. Res. | 4 |
| 2024 | Causal-learn: Causal Discovery in PythonabstractCausal discovery aims at revealing causal relations from observational data, which is a fundamental task in science and engineering. We describe causal-learn, an open-source Python library for causal discovery. This library focuses on bringing a comprehensive collection of causal discovery methods to both practitioners and researchers. It provides easy-to-use APIs for non-specialists, modular building blocks for developers, detailed documentation for learners, and comprehensive methods for all. Different from previous packages in R or Java, causal-learn is fully developed in Python, which could be more in tune with the recent preference shift in programming languages within related communities. The library is available at https://github.com/py-why/causal-learn. Yujia Zheng 0001, Biwei Huang, Wei Chen 0103, Joseph D. Ramsey, Mingming Gong, Ruichu Cai, Shohei Shimizu, Peter Spirtes, Kun Zhang 0001 |
J. Mach. Learn. Res. | 6 |
| 2024 | Counterfactual contextual bandit for recommendation under delayed feedback
Ruichu Cai, Ruming Lu, Wei Chen 0103, Zhifeng Hao 0004 |
Neural Comput. Appl. | 1 |
| 2024 | Long-term causal effects estimation via latent surrogates representation learning
Ruichu Cai, Weilin Chen 0001, Zeqin Yang, Shu Wan 0002, Jiecheng Guo |
Neural Networks | 1 |
| 2024 | Time-series domain adaptation via sparse associative structure alignment: Learning invariance and variance
Zijian Li 0001, Ruichu Cai, Yuguang Yan, Wei Chen 0103, Keli Zhang, Junjian Ye |
Neural Networks | 2 |
| 2024 | Cross-KG Link Prediction by Learning Substructural SemanticsabstractAbstract Link prediction across different knowledge graphs (i.e. Cross-KG link prediction) plays an important role in discovering new triples and fusing multi-source knowledge. Existing cross-KG link prediction methods mainly rely on entity and relation alignment, and are challenged by the problems of KG incompleteness, semantic implicitness and ambiguosness. To deal with these challenges, we propose a learning framework that incorporates both node-level and substructure-level context for cross-KG link prediction. The proposed method mainly consists of a neural-based tensor-completion module and a graph-convolutional-network module, which respectively captures the node-level and substructure-level semantics to enhance the performance of cross-KG link prediction. Extensive experiments are conducted on three benchmark datasets. The results show that our method significantly outperforms the state-of-the-art baselines and some interesting analysis on real cases are also provided in this paper. Wen Wen 0009, Shiyuan Wu, Ruichu Cai |
Neural Process. Lett. | 3 |
| 2024 | Transferable Time-Series Forecasting Under Causal Conditional ShiftabstractThis paper focuses on the problem of semi-supervised domain adaptation for time-series forecasting, which is underexplored in literature, despite being often encountered in practice. Existing methods on time-series domain adaptation mainly follow the paradigm designed for static data, which cannot handle domain-specific complex conditional dependencies raised by data offset, time lags, and variant data distributions. In order to address these challenges, we analyze variational conditional dependencies in time-series data and find that the causal structures are usually stable among domains, and further raise the causal conditional shift assumption. Enlightened by this assumption, we consider the causal generation process for time-series data and propose an end-to-end model for the semi-supervised domain adaptation problem on time-series forecasting. Our method can not only discover the Granger-Causal structures among cross-domain data but also address the cross-domain time-series forecasting problem with accurate and interpretable predicted results. We further theoretically analyze the superiority of the proposed method, where the generalization error on the target domain is bounded by the empirical risks and by the discrepancy between the causal structures from different domains. Experimental results on both synthetic and real data demonstrate the effectiveness of our method for the semi-supervised domain adaptation method on time-series forecasting. Zijian Li 0001, Ruichu Cai, Tom Z. J. Fu, Zhifeng Hao 0004, Kun Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Block diagonal representation learning with local invariance for face clustering
Shaomin Chen, Ming Yin 0002, Ruichu Cai |
Soft Comput. | 5 |
| 2024 | PatchMixing Masked Autoencoders for 3D Point Cloud Self-Supervised LearningabstractRecently, Point-MAE has extended Masked Autoencoders (MAE) to point clouds for 3D self-supervised learning, which however faces two problems: (1) the shape similarity between the masked point cloud and original point cloud is high, and (2) the pretext task of reconstructing the original point cloud is straightforward which fails to compel the network to learn deep representative features. In this paper, we tackle these problems by proposing a PatchMixing strategy and a teacher-student training framework. First, with PatchMixing, we mix selected point patches of multiple point clouds and attempt to infer the object information from the resulting mixed point cloud. Due to the interference of other objects, the task is challenging but facilitates representation learning. Second, rather than directly restoring the original point cloud, we propose a novel pretext task that involves a two-branch teacher model and a student model. These models process the multiple input point clouds in different ways (no mixing, mixing + unmixing, mixing + masking), but are expected to output similar features, thereby compelling the network to extract essential features from the input. Extensive experiments show that our well-designed PatchMixing strategy and effective teacher-student learning architecture yield impressive results. Specifically, our model achieves a remarkable 92.9% classification accuracy in the Linear SVM task on the ModelNet40 dataset. Through pre-training and fine-tuning on downstream tasks, our method achieves an 89.8% classification accuracy on the most challenging split of ScanObjectNN and an outstanding 94.0% on ModelNet40. Chengxing Lin 0001, Wenju Xu, Jian Zhu 0001, Yongwei Nie, Ruichu Cai, Xuemiao Xu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Graph Domain Adaptation: A Generative ViewabstractRecent years have witnessed tremendous interest in deep learning on graph-structured data. Due to the high cost of collecting labeled graph-structured data, domain adaptation is important to supervised graph learning tasks with limited samples. However, current graph domain adaptation methods are generally adopted from traditional domain adaptation tasks, and the properties of graph-structured data are not well utilized. For example, the observed social networks on different platforms are controlled not only by the different crowds or communities but also by domain-specific policies and background noise. Based on these properties in graph-structured data, we first assume that the graph-structured data generation process is controlled by three independent types of latent variables, i.e., the semantic latent variables, the domain latent variables, and the random latent variables. Based on this assumption, we propose a disentanglement-based unsupervised domain adaptation method for the graph-structured data, which applies variational graph auto-encoders to recover these latent variables and disentangles them via three supervised learning modules. Extensive experimental results on two real-world datasets in the graph classification task reveal that our method not only significantly outperforms the traditional domain adaptation methods and the disentangled-based domain adaptation methods but also outperforms the state-of-the-art graph domain adaptation algorithms. The code is available at https://github.com/rynewu224/GraphDA . Ruichu Cai, Fengzhu Wu, Zijian Li 0001, Pengfei Wei 0001, Lingling Yi, Kun Zhang 0001 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2024 | THPs: Topological Hawkes Processes for Learning Causal Structure on Event SequencesabstractLearning causal structure among event types on multitype event sequences is an important but challenging task. Existing methods, such as the Multivariate Hawkes processes, mostly assumed that each sequence is independent and identically distributed. However, in many real-world applications, it is commonplace to encounter a topological network behind the event sequences such that an event is excited or inhibited not only by its history but also by its topological neighbors. Consequently, the failure in describing the topological dependency among the event sequences leads to the error detection of the causal structure. By considering the Hawkes processes from the view of temporal convolution, we propose a topological Hawkes process (THP) to draw a connection between the graph convolution in the topology domain and the temporal convolution in time domains. We further propose a causal structure learning method on THP in a likelihood framework. The proposed method is featured with the graph convolution-based likelihood function of THP and a sparse optimization scheme with an Expectation-Maximization of the likelihood function. Theoretical analysis and experiments on both synthetic and real-world data demonstrate the effectiveness of the proposed method. Ruichu Cai, Jie Qiao, Keli Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Motif Graph Neural NetworkabstractGraphs can model complicated interactions between entities, which naturally emerge in many important applications. These applications can often be cast into standard graph learning tasks, in which a crucial step is to learn low-dimensional graph representations. Graph neural networks (GNNs) are currently the most popular model in graph embedding approaches. However, standard GNNs in the neighborhood aggregation paradigm suffer from limited discriminative power in distinguishing high-order graph structures as opposed to low-order structures. To capture high-order structures, researchers have resorted to motifs and developed motif-based GNNs. However, the existing motif-based GNNs still often suffer from less discriminative power on high-order structures. To overcome the above limitations, we propose motif GNN (MGNN), a novel framework to better capture high-order structures, hinging on our proposed motif redundancy minimization operator and injective motif combination. First, MGNN produces a set of node representations with respect to each motif. The next phase is our proposed redundancy minimization among motifs which compares the motifs with each other and distills the features unique to each motif. Finally, MGNN performs the updating of node representations by combining multiple representations from different motifs. In particular, to enhance the discriminative power, MGNN uses an injective function to combine the representations with respect to different motifs. We further show that our proposed architecture increases the expressive power of GNNs with a theoretical analysis. We demonstrate that MGNN outperforms state-of-the-art methods on seven public benchmarks on both the node classification and graph classification tasks. Xuexin Chen, Ruichu Cai, Yuan Fang 0001, Min Wu 0008, Zijian Li 0001, Zhifeng Hao 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | TEA: A Sequential Recommendation Framework via Temporally Evolving AggregationsabstractSequential recommendation aims to choose the most suitable items for a user at a specific timestamp given historical behaviors. Existing methods usually model the user behavior sequence based on transition-based methods such as Markov chain. However, these methods also implicitly assume that the users are independent of each other without considering the influence between users. In fact, this influence plays an important role in sequence recommendation since the behavior of a user is easily affected by others. Therefore, it is desirable to aggregate both user behaviors and the influence between users, which are evolved temporally and involved in the heterogeneous graph of users and items. In this article, we incorporate dynamic user-item heterogeneous graphs to propose a novel sequential recommendation framework. As a result, the historical behaviors as well as the influence between users can be taken into consideration. To achieve this, we first formalize sequential recommendation as a problem to estimate conditional probability given temporal dynamic heterogeneous graphs and user behavior sequences. After that, we exploit the conditional random field to aggregate the heterogeneous graphs and user behaviors for probability estimation and employ the pseudo-likelihood approach to derive a tractable objective function. Finally, we provide scalable and flexible implementations of the proposed framework. Experimental results on three real-world datasets not only demonstrate the effectiveness of our proposed method but also provide some insightful discoveries on the sequential recommendation. Zijian Li 0001, Ruichu Cai, Fengzhu Wu, Sili Zhang, Yuexing Hao, Yuguang Yan |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Generalization Bound for Estimating Causal Effects from Observational Network DataabstractEstimating causal effects from observational network data is a significant but challenging problem. Existing works in causal inference for observational network data lack an analysis of the generalization bound, which can theoretically provide support for alleviating the complex confounding bias and practically guide the design of learning objectives in a principled manner. To fill this gap, we derive a generalization bound for causal effect estimation in network scenarios by exploiting 1) the reweighting schema based on joint propensity score and 2) the representation learning schema based on Integral Probability Metric (IPM). We provide two perspectives on the generalization bound in terms of reweighting and representation learning, respectively. Motivated by the analysis of the bound, we propose a weighting regression method based on the joint propensity score augmented with representation learning. Extensive experimental studies on two real-world networks with semi-synthetic data demonstrate the effectiveness of our algorithm. Ruichu Cai, Zeqin Yang, Weilin Chen 0001, Yuguang Yan, Zhifeng Hao 0005 |
CIKM | 1 |
| 2023 | Causal Discovery with Latent Confounders Based on Higher-Order CumulantsabstractCausal discovery with latent confounders is an important but challenging task in many scientific areas. Despite the success of some overcomplete independent component analysis (OICA) based methods in certain domains, they are computationally expensive and can easily get stuck into local optima. We notice that interestingly, by making use of higher-order cumulants, there exists a closed-form solution to OICA in specific cases, e.g., when the mixing procedure follows the One-Latent-Component structure. In light of the power of the closed-form solution to OICA corresponding to the One-Latent-Component structure, we formulate a way to estimate the mixing matrix using the higher-order cumulants, and further propose the testable One-Latent-Component condition to identify the latent variables and determine causal orders. By iteratively removing the share identified latent components, we successfully extend the results on the One-Latent-Component structure to the Multi-Latent-Component structure and finally provide a practical and asymptotically correct algorithm to learn the causal structure with latent variables. Experimental results illustrate the asymptotic correctness and effectiveness of the proposed method. Ruichu Cai, Zhiyi Huang 0008, Wei Chen 0103, Zhifeng Hao 0004, Kun Zhang 0001 |
ICML | 1 |
| 2023 | Latent Causal Dynamics Model for Model-Based Reinforcement Learning
Zhifeng Hao 0004, Haipeng Zhu, Wei Chen 0103, Ruichu Cai |
ICONIP (2) | 4 |
| 2023 | Some General Identification Results for Linear Latent Hierarchical Causal StructureabstractWe study the problem of learning hierarchical causal structure among latent variables from measured variables. While some existing methods are able to recover the latent hierarchical causal structure, they mostly suffer from restricted assumptions, including the tree-structured graph constraint, no ``triangle" structure, and non-Gaussian assumptions. In this paper, we relax these restrictions above and consider a more general and challenging scenario where the beyond tree-structured graph, the ``triangle" structure, and the arbitrary noise distribution are allowed. We investigate the identifiability of the latent hierarchical causal structure and show that by using second-order statistics, the latent hierarchical structure can be identified up to the Markov equivalence classes over latent variables. Moreover, some directions in the Markov equivalence classes of latent variables can be further identified using partially non-Gaussian data. Based on the theoretical results above, we design an effective algorithm for learning the latent hierarchical causal structure. The experimental results on synthetic data verify the effectiveness of the proposed method. Zhengming Chen 0002, Feng Xie 0002, Jie Qiao, Zhifeng Hao 0004, Ruichu Cai |
IJCAI | 5 |
| 2023 | Structural Hawkes Processes for Learning Causal Structure from Discrete-Time Event SequencesabstractLearning causal structure among event types from discrete-time event sequences is a particularly important but challenging task. Existing methods, such as the multivariate Hawkes processes based methods, mostly boil down to learning the so-called Granger causality which assumes that the cause event happens strictly prior to its effect event. Such an assumption is often untenable beyond applications, especially when dealing with discrete-time event sequences in low-resolution; and typical discrete Hawkes processes mainly suffer from identifiability issues raised by the instantaneous effect, i.e., the causal relationship that occurred simultaneously due to the low-resolution data will not be captured by Granger causality. In this work, we propose Structure Hawkes Processes (SHPs) that leverage the instantaneous effect for learning the causal structure among events type in discrete-time event sequence. The proposed method is featured with the Expectation-Maximization of the likelihood function and a sparse optimization scheme. Theoretical results show that the instantaneous effect is a blessing rather than a curse, and the causal structure is identifiable under the existence of the instantaneous effect. Experiments on synthetic and real-world data verify the effectiveness of the proposed method. Jie Qiao, Ruichu Cai, Keli Zhang |
IJCAI | 2 |
| 2023 | Subspace Identification for Multi-Source Domain AdaptationabstractMulti-source domain adaptation (MSDA) methods aim to transfer knowledge from multiple labeled source domains to an unlabeled target domain. Although current methods achieve target joint distribution identifiability by enforcing minimal changes across domains, they often necessitate stringent conditions, such as an adequate number of domains, monotonic transformation of latent variables, and invariant label distributions. These requirements are challenging to satisfy in real-world applications. To mitigate the need for these strict assumptions, we propose a subspace identification theory that guarantees the disentanglement of domain-invariant and domain-specific variables under less restrictive constraints regarding domain numbers and transformation properties and thereby facilitating domain adaptation by minimizing the impact of domain shifts on invariant variables. Based on this theory, we develop a Subspace Identification Guarantee (SIG) model that leverages variational inference. Furthermore, the SIG model incorporates class-aware conditional alignment to accommodate target shifts where label distributions change with the domain. Experimental results demonstrate that our SIG model outperforms existing MSDA techniques on various benchmark datasets, highlighting its effectiveness in real-world applications. Zijian Li 0001, Ruichu Cai, Guangyi Chen 0002, Zhifeng Hao 0004, Kun Zhang 0001 |
NeurIPS | 2 |
| 2023 | Learning dynamic causal mechanisms from non-stationary data
Ruichu Cai, Liting Huang, Wei Chen 0103, Jie Qiao, Zhifeng Hao 0004 |
Appl. Intell. | 1 |
| 2023 | A selection-pattern-aware recommendation model with colored-motif attention network
Junbin Chen, Wen Wen 0009, Ruichu Cai |
Neurocomputing | 5 |
| 2023 | Multi-task ordinal regression with labeled and unlabeled data
Yanshan Xiao, Liangwang Zhang, Bo Liu 0002, Ruichu Cai |
Inf. Sci. | 4 |
| 2023 | Factorizing time-heterogeneous Markov transition for temporal recommendation
Wen Wen 0009, Wencui Wang, Ruichu Cai |
Neural Networks | 4 |
| 2023 | FFFN: Frame-By-Frame Feedback Fusion Network for Video Super-ResolutionabstractVideo super-resolution (VSR) is a fundamental and challenging task in computer vision. Many of the existing VSR works focus on how to effectively align neighboring frames to better incorporate temporal information, while little work is devoted to the important subsequent step of inter-frame information fusion, and the existing methods on frame fusion have shortcomings such as not being able to make full use of spatio-temporal information. In this work, we propose a Frame-by-frame Feedback Fusion Network (FFFN) for VSR tasks. By applying the feedback learning mechanism commonly existing in the human cognitive system to the frame fusion stage, FFFN can refine low-level representation of the fused frames with high-level information in a coarse-to-fine manner. Specifically, after the neighboring frames are aligned, we first rearrange them from near to far according to the distance from the reference frame in the temporal space, and then feed them one-by-one into a proposed recurrent structure called Feedback Fusion Module (FFM), which is then able to iteratively generate high-level representation of the fused frames with several Feature Refinement Groups (FRGs) and feedback connections. Finally, we design a Dual-path Residual Reconstruction Module (DRRM) to reconstruct the final high-resolution image. The proposed FFFN comes with a strong frame fusion and reconstruction ability, and extensive experiments on several benchmark data sets show that it achieves favorable performance against state-of-the-art methods. Jian Zhu 0001, Qingwu Zhang, Lunke Fei, Ruichu Cai, Yuan Xie 0006, Bin Sheng 0001, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Nonlinear Causal Discovery for High-Dimensional Deterministic DataabstractNonlinear causal discovery with high-dimensional data where each variable is multidimensional plays a significant role in many scientific disciplines, such as social network analysis. Previous work majorly focuses on exploiting asymmetry in the causal and anticausal directions between two high-dimensional variables (a cause-effect pair). Although there exist some works that concentrate on the causal order identification between multiple variables, i.e., more than two high-dimensional variables, they do not validate the consistency of methods through theoretical analysis on multiple-variable data. In particular, based on the asymmetry for the cause-effect pair, if model assumptions for any pair of the data are violated, the asymmetry condition will not hold, resulting in the deduction of incorrect order identification. Thus, in this article, we propose a causal functional model, namely high-dimensional deterministic model (HDDM), to identify the causal orderings among multiple high-dimensional variables. We derive two candidates' selection rules to alleviate the inconvenient effects resulted from the violated-assumption pairs. The corresponding theoretical justification is provided as well. With these theoretical results, we develop a method to infer causal orderings for nonlinear multiple-variable data. Simulations on synthetic data and real-world data are conducted to verify the efficacy of our proposed method. Since we focus on deterministic relations in our method, we also verify the robustness of the noises in simulations. Yan Zeng 0002, Zhifeng Hao 0004, Ruichu Cai, Feng Xie 0002, Libo Huang 0001, Shohei Shimizu |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Identification of Linear Latent Variable Model with Arbitrary DistributionabstractAn important problem across multiple disciplines is to infer and understand meaningful latent variables. One strategy commonly used is to model the measured variables in terms of the latent variables under suitable assumptions on the connectivity from the latents to the measured (known as measurement model). Furthermore, it might be even more interesting to discover the causal relations among the latent variables (known as structural model). Recently, some methods have been proposed to estimate the structural model by assuming that the noise terms in the measured and latent variables are non-Gaussian. However, they are not suitable when some of the noise terms become Gaussian. To bridge this gap, we investigate the problem of identification of the structural model with arbitrary noise distributions. We provide necessary and sufficient condition under which the structural model is identifiable: it is identifiable iff for each pair of adjacent latent variables Lx, Ly, (1) at least one of Lx and Ly has non-Gaussian noise, or (2) at least one of them has a non-Gaussian ancestor and is not d-separated from the non-Gaussian component of this ancestor by the common causes of Lx and Ly. This identifiability result relaxes the non-Gaussianity requirements to only a (hopefully small) subset of variables, and accordingly elegantly extends the application scope of the structural model. Based on the above identifiability result, we further propose a practical algorithm to learn the structural model. We verify the correctness of the identifiability result and the effectiveness of the proposed method through empirical studies. Zhengming Chen 0002, Feng Xie 0002, Jie Qiao, Zhifeng Hao 0004, Kun Zhang 0001, Ruichu Cai |
AAAI | 6 |
| 2022 | Causal Alignment Based Fault Root Causes Localization for Wireless NetworkabstractLocalizing fault root causes is challenging but critical for wireless network operation and maintenance. Though supervised methods have shown promising results in training samples, most of the existing approaches assume that the training and the testing samples are independent and identical distributed. Such an i.i.d assumption usually does not hold due to network faults that may occur in different devices across different domains (well known as the distribution shift). Thus, it is necessary to align distributions between the training and test data set. Motivated by the stability of the causal mechanism across the domains, a Causal Alignment based Root Cause Localization (CARCL) framework, including the causal alignment and the multi-stage classifier, is proposed. CARCL first offers to align the distributions locally for each causal module but not globally on the complete variable set. We further develop a multi-stage classifier to determine the root causes with the help of predicted pseudo labels. The experiments demonstrate a superior performance of our method. Yuequn Liu, Jie Qiao, Zhiyi Huang 0008, Xuanzhi Chen, Wei Chen 0103, Ruichu Cai |
ICASSP | 8 |
| 2022 | Double embedding-transfer-based multi-view spectral clustering
Ming Yin 0002, Ruichu Cai, Wen Wen 0009 |
Expert Syst. Appl. | 5 |
| 2022 | Shared state space model for background information extraction and time series prediction
Ruichu Cai, Zhaolong Lin, Wei Chen 0103, Zhifeng Hao 0004 |
Neurocomputing | 1 |
| 2022 | Learning granger causality for non-stationary Hawkes processes
Wei Chen 0103, Jibin Chen, Ruichu Cai, Yuequn Liu, Zhifeng Hao 0004 |
Neurocomputing | 3 |
| 2022 | Motif-based memory networks for complex-factoid question answering
Wen Wen 0009, Ruichu Cai |
Neurocomputing | 5 |
| 2022 | Causal discovery from multi-domain data using the independence of modularities
Jie Qiao, Yiming Bai, Ruichu Cai |
Neural Comput. Appl. | 3 |
| 2022 | Causal Discovery in Linear Non-Gaussian Acyclic Model With Multiple Latent ConfoundersabstractCausal discovery from observational data is a fundamental problem in science. Though the linear non-Gaussian acyclic model (LiNGAM) has shown promising results in various applications, it still faces the following challenges in the data with multiple latent confounders: 1) how to detect the latent confounders and 2) how to uncover the causal relations among observed and latent variables. To address these two challenges, we propose a hybrid causal discovery method for the LiNGAM with multiple latent confounders (MLCLiNGAM). First, we utilize the constraint-based method to learn the causal skeleton. Second, we identify the causal directions, by conducting regression and independence tests on the adjacent pairs in the causal skeleton. Third, we detect the latent confounders with the help of the maximal clique patterns raised by the latent confounders and reconstruct the causal structure with latent variables. Theoretical results show the correctness and efficiency of the algorithms. We conduct extensive experiments on synthetic and real data, which illustrates the efficiency and effectiveness of the proposed algorithms. Wei Chen 0103, Ruichu Cai, Kun Zhang 0001, Zhifeng Hao 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Time Series Domain Adaptation via Sparse Associative Structure AlignmentabstractDomain adaptation on time series data is an important but challenging task. Most of the existing works in this area are based on the learning of the domain-invariant representation of the data with the help of restrictions like MMD. However, such extraction of the domain-invariant representation is a non-trivial task for time series data, due to the complex dependence among the timestamps. In detail, in the fully dependent time series, a small change of the time lags or the offsets may lead to difficulty in the domain invariant extraction. Fortunately, the stability of the causality inspired us to explore the domain invariant structure of the data. To reduce the difficulty in the discovery of causal structure, we relax it to the sparse associative structure and propose a novel sparse associative structure alignment model for domain adaptation. First, we generate the segment set to exclude the obstacle of offsets. Second, the intra-variables and inter-variables sparse attention mechanisms are devised to extract associative structure time-series data with considering time lags. Finally, the associative structure alignment is used to guide the transfer of knowledge from the source domain to the target one. Experimental studies not only verify the good performance of our methods on three real-world datasets but also provide some insightful discoveries on the transferred knowledge. Ruichu Cai, Zijian Li 0001, Wei Chen 0103, Keli Zhang, Junjian Ye, Zhuozhang Li |
AAAI | 1 |
| 2021 | Appearance-Motion Memory Consistency Network for Video Anomaly DetectionabstractAbnormal event detection in the surveillance video is an essential but challenging task, and many methods have been proposed to deal with this problem. The previous methods either only consider the appearance information or directly integrate the results of appearance and motion information without considering their endogenous consistency semantics explicitly. Inspired by the rule humans identify the abnormal frames from multi-modality signals, we propose an Appearance-Motion Memory Consistency Network (AMMC-Net). Our method first makes full use of the prior knowledge of appearance and motion signals to explicitly capture the correspondence between them in the high-level feature space. Then, it combines the multi-view features to obtain a more essential and robust feature representation of regular events, which can significantly increase the gap between an abnormal and a regular event. In the anomaly detection phase, we further introduce a commit error in the latent space joint with the prediction error in pixel space to enhance the detection accuracy. Solid experimental results on various standard datasets validate the effectiveness of our approach. Ruichu Cai, Wen Liu 0003, Shenghua Gao |
AAAI | 1 |
| 2021 | Causal Discovery with Multi-Domain LiNGAM for Latent FactorsabstractDiscovering causal structures among latent factors from observed data is a particularly challenging problem. Despite some efforts for this problem, existing methods focus on the single-domain data only. In this paper, we propose Multi-Domain Linear Non-Gaussian Acyclic Models for LAtent Factors (MD-LiNA), where the causal structure among latent factors of interest is shared for all domains, and we provide its identification results. The model enriches the causal representation for multi-domain data. We propose an integrated two-phase algorithm to estimate the model. In particular, we first locate the latent factors and estimate the factor loading matrix. Then to uncover the causal structure among shared latent factors of interest, we derive a score function based on the characterization of independence relations between external influences and the dependence relations between multi-domain latent factors and latent factors of interest. We show that the proposed method provides locally consistent estimators. Experimental results on both synthetic and real-world data demonstrate the efficacy and robustness of our approach. Yan Zeng 0002, Shohei Shimizu, Ruichu Cai, Feng Xie 0002, Michio Yamamoto, Zhifeng Hao 0004 |
IJCAI | 3 |
| 2021 | SADGA: Structure-Aware Dual Graph Aggregation Network for Text-to-SQLabstractThe Text-to-SQL task, aiming to translate the natural language of the questions into SQL queries, has drawn much attention recently. One of the most challenging problems of Text-to-SQL is how to generalize the trained model to the unseen database schemas, also known as the cross-domain Text-to-SQL task. The key lies in the generalizability of (i) the encoding method to model the question and the database schema and (ii) the question-schema linking method to learn the mapping between words in the question and tables/columns in the database schema. Focusing on the above two key issues, we propose a \emph{Structure-Aware Dual Graph Aggregation Network} (SADGA) for cross-domain Text-to-SQL. In SADGA, we adopt the graph structure to provide a unified encoding model for both the natural language question and database schema. Based on the proposed unified modeling, we further devise a structure-aware aggregation method to learn the mapping between the question-graph and schema-graph. The structure-aware aggregation method is featured with \emph{Global Graph Linking}, \emph{Local Graph Linking} and \emph{Dual-Graph Aggregation Mechanism}. We not only study the performance of our proposal empirically but also achieved 3rd place on the challenging Text-to-SQL benchmark Spider at the time of writing. Ruichu Cai, Jinjie Yuan |
NeurIPS | 1 |
| 2021 | Domain Adaptation with Invariant Representation Learning: What Transformations to Learn?abstractUnsupervised domain adaptation, as a prevalent transfer learning setting, spans many real-world applications. With the increasing representational power and applicability of neural networks, state-of-the-art domain adaptation methods make use of deep architectures to map the input features $X$ to a latent representation $Z$ that has the same marginal distribution across domains. This has been shown to be insufficient for generating optimal representation for classification, and to find conditionally invariant representations, usually strong assumptions are needed. We provide reasoning why when the supports of the source and target data from overlap, any map of $X$ that is fixed across domains may not be suitable for domain adaptation via invariant features. Furthermore, we develop an efficient technique in which the optimal map from $X$ to $Z$ also takes domain-specific information as input, in addition to the features $X$. By using the property of minimal changes of causal mechanisms across domains, our model also takes into account the domain-specific information to ensure that the latent representation $Z$ does not discard valuable information about $Y$. We demonstrate the efficacy of our method via synthetic and real-world data experiments. The code is available at: \texttt{https://github.com/DMIRLAB-Group/DSAN}. Petar Stojanov, Zijian Li 0001, Mingming Gong, Ruichu Cai, Jaime G. Carbonell, Kun Zhang 0001 |
NeurIPS | 4 |
| 2021 | Learning causal structures using hidden compact representation
Jie Qiao, Yiming Bai, Ruichu Cai, Zhifeng Hao 0004 |
Neurocomputing | 3 |
| 2021 | Compensating the vorticity loss during advection with an adaptive vorticity confinement forceabstractAbstract The advection step in grid‐based fluid simulation is prone to numerical dissipation, which results in loss of detail. How to improve the advection accuracy to preserve more fluid details is still challenging. On the other hand, a common way to enhance smoke details is to use vorticity confinement. However, most of the previous methods simply used a fine‐tuned scale factor ε to adjust the strength of the confinement force, which can only amplify existing vortex details and is easy to cause instability when ε is large. In this article, we proposed an adaptive vorticity confinement method, which does not suffer from the above problems, to compensate the vorticity loss during advection with little extra cost. The main idea is to first calculate a scale factor whose value depends on the vorticity loss during advection, and then use it to adaptively control the vorticity confinement force for vorticity compensation with high stability. The experiment results show the effectiveness and efficiency of our method. Jian Zhu 0001, Silong Li, Ruichu Cai, Guoheng Huang, Bin Sheng 0001, Enhua Wu |
Comput. Animat. Virtual Worlds | 3 |
| 2021 | Semi-supervised disentangled framework for transferable named entity recognition
Zhifeng Hao 0004, Di Lv, Zijian Li 0001, Ruichu Cai, Wen Wen 0009 |
Neural Networks | 4 |
| 2021 | Causal Mechanism Transfer Network for Time Series Domain Adaptation in Mechanical SystemsabstractData-driven models are becoming essential parts in modern mechanical systems, commonly used to capture the behavior of various equipment and varying environmental characteristics. Despite the advantages of these data-driven models on excellent adaptivity to high dynamics and aging equipment, they are usually hungry for massive labels, mostly contributed by human engineers at a high cost. Fortunately, domain adaptation enhances the model generalization by utilizing the labeled source data and the unlabeled target data. However, the mainstream domain adaptation methods cannot achieve ideal performance on time series data, since they assume that the conditional distributions are equal. This assumption works well in the static data but is inapplicable for the time series data. Even the first-order Markov dependence assumption requires the dependence between any two consecutive time steps. In this article, we assume that the causal mechanism is invariant and present our Causal Mechanism Transfer Network (CMTN) for time series domain adaptation. By capturing causal mechanisms of time series data, CMTN allows the data-driven models to exploit existing data and labels from similar systems, such that the resulting model on a new system is highly reliable even with limited data. We report our empirical results and lessons learned from two real-world case studies, on chiller plant energy optimization and boiler fault detection, which outperform the existing state-of-the-art method. Zijian Li 0001, Ruichu Cai, Hong Wei Ng, Marianne Winslett, Tom Z. J. Fu |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | Causal Discovery with Confounding Cascade Nonlinear Additive Noise ModelsabstractIdentification of causal direction between a causal-effect pair from observed data has recently attracted much attention. Various methods based on functional causal models have been proposed to solve this problem, by assuming the causal process satisfies some (structural) constraints and showing that the reverse direction violates such constraints. The nonlinear additive noise model has been demonstrated to be effective for this purpose, but the model class does not allow any confounding or intermediate variables between a cause pair–even if each direct causal relation follows this model. However, omitting the latent causal variables is frequently encountered in practice. After the omission, the model does not necessarily follow the model constraints. As a consequence, the nonlinear additive noise model may fail to correctly discover causal direction. In this work, we propose a confounding cascade nonlinear additive noise model to represent such causal influences–each direct causal relation follows the nonlinear additive noise model but we observe only the initial cause and final effect. We further propose a method to estimate the model, including the unmeasured confounding and intermediate variables, from data under the variational auto-encoder framework. Our theoretical results show that with our model, the causal direction is identifiable under suitable technical conditions on the data generation process. Simulation results illustrate the power of the proposed method in identifying indirect causal relations across various settings, and experimental results on real data suggest that the proposed model and method greatly extend the applicability of causal discovery based on functional causal models in nonlinear cases. Jie Qiao, Ruichu Cai, Kun Zhang 0001, Zhifeng Hao 0004 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | Prediction of Synthetic Lethal Interactions in Human Cancers Using Multi-View Graph Auto-EncoderabstractSynthetic lethality (SL) is a very important concept for the development of targeted anticancer drugs. However, experimental methods for SL detection often suffer from various issues like high cost and low consistency across cell lines. Hence, computational methods for predicting novel SLs have recently emerged as complements for wet-lab experiments. In addition, SL data can be represented as a graph where nodes are genes and edges are the SL interactions. It is thus motivated to design advanced graph-based machine learning algorithms for SL prediction. In this paper, we propose a novel SL prediction method using Multi-view Graph Auto-Encoder (SLMGAE). We consider the SL graph as the main view and the graphs from other data sources (e.g., PPI, GO, etc.) as support views. Multiple Graph Auto-Encoders (GAEs) are implemented to reconstruct the graphs for different views. We further design an attention mechanism, which assigns different weights for support views, to combine all the reconstructed graphs for SL prediction. The overall SLMGAE model is then trained by minimizing both the reconstruction error and prediction error. Experimental results on the SynLethDB dataset show that SLMGAE outperforms state-of-the-arts. The case studies on novel predicted SLs also illustrate the effectiveness of our SLMGAE method. Zhifeng Hao 0004, Yuan Fang 0001, Min Wu 0008, Ruichu Cai, Xiaoli Li 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | TAG : Type Auxiliary Guiding for Code Comment GenerationabstractExisting leading code comment generation approaches with the structure-to-sequence framework ignores the type information of the interpretation of the code, e.g., operator, string, etc.However, introducing the type information into the existing framework is non-trivial due to the hierarchical dependence among the type information.In order to address the issues above, we propose a Type Auxiliary Guiding encoder-decoder framework for the code comment generation task which considers the source code as an N-ary tree with type information associated with each node.Specifically, our framework is featured with a Typeassociated Encoder and a Type-restricted Decoder which enables adaptive summarization of the source code.We further propose a hierarchical reinforcement learning method to resolve the training difficulties of our proposed framework.Extensive evaluations demonstrate the state-of-the-art performance of our framework with both the auto-evaluated metrics and case studies. Ruichu Cai, Zijian Li 0001, Yuexing Hao, Yao Chen 0008 |
ACL | 1 |
| 2020 | Meta Multi-Task Learning for Speech Emotion Recognition
Ruichu Cai, Kaibin Guo |
INTERSPEECH | 1 |
| 2020 | Generalized Independent Noise Condition for Estimating Latent Variable Causal GraphsabstractCausal discovery aims to recover causal structures or models underlying the observed data. Despite its success in certain domains, most existing methods focus on causal relations between observed variables, while in many scenarios the observed ones may not be the underlying causal variables (e.g., image pixels), but are generated by latent causal variables or confounders that are causally related. To this end, in this paper, we consider Linear, Non-Gaussian Latent variable Models (LiNGLaMs), in which latent confounders are also causally related, and propose a Generalized Independent Noise (GIN) condition to estimate such latent variable graphs. Specifically, for two observed random vectors $\mathbf{Y}$ and $\mathbf{Z}$, GIN holds if and only if $\omega^{\intercal}\mathbf{Y}$ and $\mathbf{Z}$ are statistically independent, where $\omega$ is a parameter vector characterized from the cross-covariance between $\mathbf{Y}$ and $\mathbf{Z}$. From the graphical view, roughly speaking, GIN implies that causally earlier latent common causes of variables in $\mathbf{Y}$ d-separate $\mathbf{Y}$ from $\mathbf{Z}$. Interestingly, we find that the independent noise condition, i.e., if there is no confounder, causes are independent from the error of regressing the effect on the causes, can be seen as a special case of GIN. Moreover, we show that GIN helps locate latent variables and identify their causal structure, including causal directions. We further develop a recursive learning algorithm to achieve these goals. Experimental results on synthetic and real-world data demonstrate the effectiveness of our method. Feng Xie 0002, Ruichu Cai, Biwei Huang, Clark Glymour, Zhifeng Hao 0004, Kun Zhang 0001 |
NeurIPS | 2 |
| 2020 | Dual-dropout graph convolutional network for predicting synthetic lethality in human cancersabstractMOTIVATION: Synthetic lethality (SL) is a promising form of gene interaction for cancer therapy, as it is able to identify specific genes to target at cancer cells without disrupting normal cells. As high-throughput wet-lab settings are often costly and face various challenges, computational approaches have become a practical complement. In particular, predicting SLs can be formulated as a link prediction task on a graph of interacting genes. Although matrix factorization techniques have been widely adopted in link prediction, they focus on mapping genes to latent representations in isolation, without aggregating information from neighboring genes. Graph convolutional networks (GCN) can capture such neighborhood dependency in a graph. However, it is still challenging to apply GCN for SL prediction as SL interactions are extremely sparse, which is more likely to cause overfitting. RESULTS: In this article, we propose a novel dual-dropout GCN (DDGCN) for learning more robust gene representations for SL prediction. We employ both coarse-grained node dropout and fine-grained edge dropout to address the issue that standard dropout in vanilla GCN is often inadequate in reducing overfitting on sparse graphs. In particular, coarse-grained node dropout can efficiently and systematically enforce dropout at the node (gene) level, while fine-grained edge dropout can further fine-tune the dropout at the interaction (edge) level. We further present a theoretical framework to justify our model architecture. Finally, we conduct extensive experiments on human SL datasets and the results demonstrate the superior performance of our model in comparison with state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: DDGCN is implemented in Python 3.7, open-source and freely available at https://github.com/CXX1113/Dual-DropoutGCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ruichu Cai, Xuexin Chen, Yuan Fang 0001, Min Wu 0008, Yuexing Hao, Jonathan D. Wren |
Bioinform. | 1 |
| 2020 | Detail-preserving smoke simulation using an efficient high-order numerical scheme
Jian Zhu 0001, Hanqiu Sun, Enhua Wu, Ruichu Cai |
Sci. China Inf. Sci. | 5 |
| 2020 | Block diagonal representation learning for robust subspace clustering
Ming Yin 0002, Ruichu Cai |
Inf. Sci. | 4 |
| 2020 | Animating turbulent fluid with a robust and efficient high-order advection methodabstractAbstract The accuracy of advection has a great influence on the visual effect of fluid simulation. Constrained interpolation profile (CIP) method has been an important advection scheme because of its third‐order accuracy and the fact that it only needs to be performed over a compact stencil, but extending it to high‐dimensional advection equations is not easy, because it involves complex calculations and large memory overheads, and is usually unstable. In this article, we propose a stable and efficient three‐dimensional (3D) CIP scheme which can maintain high accuracy but requires low computation and memory cost. We first construct an efficient two‐dimensional (2D) CIP scheme based on dimensional splitting and local Taylor expansions, and then propose an effective way to extend it for 3D applications without decreasing the computational accuracy or affecting the stability. The experimental results show the advantages of our method over the state‐of‐the‐art advection schemes. Jian Zhu 0001, Silong Li, Ruichu Cai, Guoheng Huang, Bin Sheng 0001, Enhua Wu |
Comput. Animat. Virtual Worlds | 3 |
| 2020 | Mining hidden non-redundant causal relationships in online social networks
Wei Chen 0103, Ruichu Cai, Zhifeng Hao 0004, Chang Yuan, Feng Xie 0002 |
Neural Comput. Appl. | 2 |
| 2020 | FOM: Fourth-order moment based causal direction identification on the heteroscedastic data
Ruichu Cai, Jincheng Ye, Jie Qiao, Huiyuan Fu |
Neural Networks | 1 |
| 2020 | Multi-context aware user-item embedding for recommendation
Wen Wen 0009, Ruichu Cai |
Neural Networks | 4 |
| 2020 | A causal discovery algorithm based on the prior selection of leaf nodes
Yan Zeng 0002, Zhifeng Hao 0004, Ruichu Cai, Feng Xie 0002, Liang Ou, Ruihui Huang |
Neural Networks | 3 |
| 2020 | DACH: Domain Adaptation Without Domain InformationabstractDomain adaptation is becoming increasingly important for learning systems in recent years, especially with the growing diversification of data domains in real-world applications, such as the genetic data from various sequencing platforms and video feeds from multiple surveillance cameras. Traditional domain adaptation approaches target to design transformations for each individual domain so that the twisted data from different domains follow an almost identical distribution. In many applications, however, the data from diversified domains are simply dumped to an archive even without clear domain labels. In this article, we discuss the possibility of learning domain adaptations even when the data does not contain domain labels. Our solution is based on our new model, named domain adaption using cross-domain homomorphism (DACH in short), to identify intrinsic homomorphism hidden in mixed data from all domains. DACH is generally compatible with existing deep learning frameworks, enabling the generation of nonlinear features from the original data domains. Our theoretical analysis not only shows the universality of the homomorphism, but also proves the convergence of DACH for significant homomorphism structures over the data domains is preserved. Empirical studies on real-world data sets validate the effectiveness of DACH on merging multiple data domains for joint machine learning tasks and the scalability of our algorithm to domain dimensionality. Ruichu Cai, Zhifeng Hao 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | An Efficient Entropy-Based Causal Discovery Method for Linear Structural Equation Models With IID Noise VariablesabstractThe discovery of causal relationships from the observational data is an important task. To identify the unique causal structure belonging to a Markov equivalence class, a number of algorithms, such as the linear non-Gaussian acyclic model (LiNGAM), have been proposed. However, two challenges remain to be met: 1) these algorithms fail to work on the data which follow linear structural equation model with Gaussian noise and 2) they misjudge the causal direction when the data contain additional measurement errors. In this paper, we propose an entropy-based two-phase iterative algorithm for arbitrary distribution data with additional measurement errors under some mild assumptions. In the first phase of the algorithm, based on the property that entropy can measure the amount of information behind the data with arbitrary distribution, we design a general approach for the identification of exogenous variable on both Gaussian and non-Gaussian data, and we give the corresponding theoretical derivation. In the second phase, to eliminate the effects of measurement errors, we revise the value of the exogenous variable by removing its measurement error and further use the revised value to remove its effect on the remaining variables. Experimental results on real-world causal structures are presented to demonstrate the effectiveness and stability of our method. We also apply the proposed algorithm on the mobile-base-station data with measurement errors, and the results further prove the effectiveness of our algorithm. Feng Xie 0002, Ruichu Cai, Yan Zeng 0002, Jiantao Gao, Zhifeng Hao 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Learning Disentangled Semantic Representation for Domain AdaptationabstractDomain adaptation is an important but challenging task. Most of the existing domain adaptation methods struggle to extract the domain-invariant representation on the feature space with entangling domain information and semantic information. Different from previous efforts on the entangled feature space, we aim to extract the domain invariant semantic information in the latent disentangled semantic representation (DSR) of the data. In DSR, we assume the data generation process is controlled by two independent sets of variables, i.e., the semantic latent variables and the domain latent variables. Under the above assumption, we employ a variational auto-encoder to reconstruct the semantic latent variables and domain latent variables behind the data. We further devise a dual adversarial network to disentangle these two sets of reconstructed latent variables. The disentangled semantic latent variables are finally adapted across the domains. Experimental studies testify that our model yields state-of-the-art performance on several domain adaptation benchmark datasets. Ruichu Cai, Zijian Li 0001, Pengfei Wei 0001, Jie Qiao, Kun Zhang 0001, Zhifeng Hao 0004 |
IJCAI | 1 |
| 2019 | Causal Discovery with Cascade Nonlinear Additive Noise ModelabstractIdentification of causal direction between a causal-effect pair from observed data has recently attracted much attention. Various methods based on functional causal models have been proposed to solve this problem, by assuming the causal process satisfies some (structural) constraints and showing that the reverse direction violates such constraints. The nonlinear additive noise model has been demonstrated to be effective for this purpose, but the model class is not transitive--even if each direct causal relation follows this model, indirect causal influences, which result from omitted intermediate causal variables and are frequently encountered in practice, do not necessarily follow the model constraints; as a consequence, the nonlinear additive noise model may fail to correctly discover causal direction. In this work, we propose a cascade nonlinear additive noise model to represent such causal influences--each direct causal relation follows the nonlinear additive noise model but we observe only the initial cause and final effect. We further propose a method to estimate the model, including the unmeasured intermediate variables, from data, under the variational auto-encoder framework. Our theoretical results show that with our model, causal direction is identifiable under suitable technical conditions on the data generation process. Simulation results illustrate the power of the proposed method in identifying indirect causal relations across various settings, and experimental results on real data suggest that the proposed model and method greatly extend the applicability of causal discovery based on functional causal models in nonlinear cases. Ruichu Cai, Jie Qiao, Kun Zhang 0001, Zhifeng Hao 0004 |
IJCAI | 1 |
| 2019 | Triad Constraints for Learning Causal Structure of Latent VariablesabstractLearning causal structure from observational data has attracted much attention, and it is notoriously challenging to find the underlying structure in the presence of confounders (hidden direct common causes of two variables). In this paper, by properly leveraging the non-Gaussianity of the data, we propose to estimate the structure over latent variables with the so-called Triad constraints: we design a form of "pseudo-residual" from three variables, and show that when causal relations are linear and noise terms are non-Gaussian, the causal direction between the latent variables for the three observed variables is identifiable by checking a certain kind of independence relationship. In other words, the Triad constraints help us to locate latent confounders and determine the causal direction between them. This goes far beyond the Tetrad constraints and reveals more information about the underlying structure from non-Gaussian data. Finally, based on the Triad constraints, we develop a two-step algorithm to learn the causal structure corresponding to measurement models. Experimental results on both synthetic and real data demonstrate the effectiveness and reliability of our method. Ruichu Cai, Feng Xie 0002, Clark Glymour, Zhifeng Hao 0004, Kun Zhang 0001 |
NeurIPS | 1 |
| 2019 | A subgraph-representation-based method for answering complex questions over knowledge bases
Wen Wen 0009, Ruichu Cai |
Neural Networks | 4 |
| 2019 | Auto-scaling for real-time stream analytics on HPC cloud
Yingchao Cheng, Ruichu Cai |
Serv. Oriented Comput. Appl. | 3 |
| 2018 | SELF: Structural Equational Likelihood Framework for Causal DiscoveryabstractCausal discovery without intervention is well recognized as a challenging yet powerful data analysis tool, boosting the development of other scientific areas, such as biology, astronomy, and social science. The major technical difficulty behind the observation-based causal discovery is to effectively and efficiently identify causes and effects from correlated variables given the existence of significant noises. Previous studies mostly employ two very different methodologies under Bayesian network framework, namely global likelihood maximization and locally complexity analysis over marginal distributions. While these approaches are effective in their respective problem domains, in this paper, we show that they can be combined to formulate a new global optimization model with local statistical significance, called structural equational likelihood framework (or SELF in short). We provide thorough analysis on the soundness of the model under mild conditions and present efficient heuristic-based algorithms for scalable model training. Empirical evaluations using XGBoost validate the superiority of our proposal over state-of-the-art solutions, on both synthetic and real world causal structures. Ruichu Cai, Jie Qiao |
AAAI | 1 |
| 2018 | HASS: High Accuracy Spike Sorting with Wavelet Package Decomposition and Mutual Information
Yao Chen 0008, Libo Huang 0001, Jiong He, Kunyao Zhao, Ruichu Cai, Zhifeng Hao 0004 |
BIBM | 5 |
| 2018 | Generating Natural Answers on Knowledge Bases and Text by Sequence-to-Sequence Learning
Zhihao Ye, Ruichu Cai, Zhaohui Liao, Jinfen Li |
ICANN (1) | 2 |
| 2018 | Waterwheel: Realtime Indexing and Temporal Range Query Processing over Massive Data StreamsabstractMassive data streams from sensors in Internet of Things (IoT) and smart devices with Global Positioning System (GPS) are now flooding to database systems for further processing and analysis. The capability of real-time retrieval from both fresh and historical data turns out to be the key enabler to the real world applications in smart manufacturing and smart city utilizing these data streams. In this paper, we present a simple and effective distributed solution to achieve millions of tuple insertions per second and ad-hoc temporal range query processing in milliseconds. To this end, we propose a new data partitioning scheme that takes advantage of the workload characteristics and avoids expensive global data merging. Furthermore, to resolve the throughput bottleneck, we adopt a template-based index method to skip unnecessary index structure adjustments over the relatively stable distribution of incoming tuples. To parallelize data insertion and query processing, we propose an efficient dispatching mechanism and effective load balancing strategies to fully utilize computational resources in a workload-aware manner. On both synthetic and real workloads, our solution consistently outperforms state-of-the-art open-source systems by at least an order of magnitude. Ruichu Cai, Tom Z. J. Fu, Jiong He, Zijie Lu, Marianne Winslett |
ICDE | 2 |
| 2018 | Identification of Causality Among Gene Mutations Through Local Causal Association Rule Discovery
Ruichu Cai, Qiqi Zhen |
ICONIP (7) | 1 |
| 2018 | HPC2-ARS: An Architecture for Real-Time Analytic of Big Data StreamsabstractHPC2-ARS supports a high performance cloud computing (HPC2) based streaming data analytic system, which ensures real-time response on unpredictable and fluctuating Big Data Streams by provisioning and scheduling computing resources autonomously. It focuses on parallel high-volume streaming applications, which have stringent real-time constraints and bring Big Data issues. It is a brand-new three-layered architecture, which solves three essential problems: (a) how many resources are needed for each application to achieve real-time analytic on streaming Big Data, (b) where to best place the allocated resources to minimize resource consumption, and (c) how to minimize response time for parallel applications. In summary, HPC2-ARS provides high performance streaming services. Yingchao Cheng, Ruichu Cai, Wen Wen 0009 |
ICWS | 3 |
| 2018 | An Encoder-Decoder Framework Translating Natural Language to Database QueriesabstractMachine translation is going through a radical revolution, driven by the explosive development of deep learning techniques using Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN). In this paper, we consider a special case in machine translation problems, targeting to convert natural language into Structured Query Language (SQL) for data retrieval over relational database. Although generic CNN and RNN learn the grammar structure of SQL when trained with sufficient samples, the accuracy and training efficiency of the model could be dramatically improved, when the translation model is deeply integrated with the grammar rules of SQL. We present a new encoder-decoder framework, with a suite of new approaches, including new semantic features fed into the encoder, grammar-aware states injected into the memory of decoder, as well as recursive state management for sub-queries. These techniques help the neural network better focus on understanding semantics of operations in natural language and save the efforts on SQL grammar learning. The empirical evaluation on real world database and queries show that our approach outperform state-of-the-art solution by a significant margin. Ruichu Cai, Zijian Li 0001 |
IJCAI | 1 |
| 2018 | Causal Discovery from Discrete Data using Hidden Compact RepresentationabstractCausal discovery from a set of observations is one of the fundamental problems across several disciplines. For continuous variables, recently a number of causal discovery methods have demonstrated their effectiveness in distinguishing the cause from effect by exploring certain properties of the conditional distribution, but causal discovery on categorical data still remains to be a challenging problem, because it is generally not easy to find a compact description of the causal mechanism for the true causal direction. In this paper we make an attempt to find a way to solve this problem by assuming a two-stage causal process: the first stage maps the cause to a hidden variable of a lower cardinality, and the second stage generates the effect from the hidden representation. In this way, the causal mechanism admits a simple yet compact representation. We show that under this model, the causal direction is identifiable under some weak conditions on the true causal mechanism. We also provide an effective solution to recover the above hidden compact representation within the likelihood framework. Empirical studies verify the effectiveness of the proposed approach on both synthetic and real-world data. Ruichu Cai, Jie Qiao, Kun Zhang 0001, Zhifeng Hao 0004 |
NeurIPS | 1 |
| 2018 | Synthetic fluid details for the vorticity loss in advectionabstractAbstract In this paper, a novel method with good numerical stability is proposed from the perspective of energy preserving to alleviate the numerical dissipations in the advection step of Eulerian fluid simulation. The main idea is to measure the vorticity loss during advection, calculate the lost angular kinetic energy with a proposed scheme, and then synthesize a high‐frequency incompressible details field to compensate the lost energy in a way that is consistent with Kolmogorov's theory, which prevents the synthetic details from interfering with the existing fluid flow. The method works independently of the advection scheme and can be easily combined with other advection schemes to enhance the effect. It adds only 5% to 10% of the computational overhead while producing convincing fluid details without changing the overall behavior of the original flow. Jian Zhu 0001, Yu Luo 0004, Xiaohua Ren, Ruichu Cai, Hanqiu Sun, Enhua Wu |
Comput. Animat. Virtual Worlds | 4 |
| 2018 | A component-driven distributed framework for real-time video dehazing
Meihua Wang, Jiaming Mai, Yun Liang 0003, Ruichu Cai, Tom Z. J. Fu |
Multim. Tools Appl. | 4 |
| 2018 | Single image deraining using deep convolutional networks
Meihua Wang, Jiaming Mai, Ruichu Cai, Yun Liang 0003, Hua Wan |
Multim. Tools Appl. | 3 |
| 2018 | Sophisticated Merging Over Random Partitions: A Scalable and Robust Causal Discovery ApproachabstractScalable causal discovery is an essential technology to a wide spectrum of applications, including biomedical studies and social network evolution analysis. To tackle the difficulty of high dimensionality, a number of solutions are proposed in the literature, generally dividing the original variable domain into smaller subdomains by computation intensive partitioning strategies. These approaches usually suffer significant structural errors when the partitioning strategies fail to recognize true causal edges across the output subdomains. Such a structural error accumulates quickly with the growing depth of recursive partitioning, due to the lack of correction mechanism over causally connected variables when they are wrongly divided into two subdomains, finally jeopardizing the robustness of the integrated results. This paper proposes a completely different strategy to solve the problem, powered by a lightweight random partitioning scheme together with a carefully designed merging algorithm over results from the random partitions. Based on the randomness properties of the partitioning scheme, we design a suite of tricks for the merging algorithm, in order to support propagation-based significance enhancement, maximal acyclic subgraph causal ordering, and order-sensitive redundancy elimination. Theoretical studies as well as empirical evaluations verify the genericity, effectiveness, and scalability of our proposal on both simulated and real-world causal structures when the scheme is used in combination with a variety of causal solvers known effective on smaller domains. Ruichu Cai, Zhifeng Hao 0004, Marianne Winslett |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | A Robust Noise Resistant Algorithm for POI Identification from Flickr DataabstractPoint of Interests (POI) identification using social media data (e.g. Flickr, Microblog) is one of the most popular research topics in recent years. However, there exist large amounts of noises (POI irrelevant data) in such crowd-contributed collections. Traditional solutions to this problem is to set a global density threshold and remove the data point as noise if its density is lower than the threshold. However, the density values vary significantly among POIs. As the result, some POIs with relatively lower density could not be identified. To solve the problem, we propose a technique based on the local drastic changes of the data density. First we define the local maxima of the density function as the Urban POIs, and the gradient ascent algorithm is exploited to assign data points into different clusters. To remove noises, we incorporate the Laplacian Zero-Crossing points along the gradient ascent process as the boundaries of the POI. Points located outside the POI region are regarded as noises. Then the technique is extended into the geographical and textual joint space so that it can make use of the heterogeneous features of social media. The experimental results show the significance of the proposed approach in removing noises. Yiyang Yang, Zhiguo Gong, Qing Li 0001, Leong Hou U, Ruichu Cai |
IJCAI | 5 |
| 2017 | An efficient kurtosis-based causal discovery method for linear non-Gaussian acyclic dataabstractUnderstanding the causality behind the observational data is of great importance to a lot of real world applications, e.g., the improvement of Quality of Service. Non-Gaussianity has been exploited in numerous causal discovery methods for observational linear acyclic data. Transforming non-Gaussianity into indirect metrics is a conventional solution employed by existing methods, although this usually results in unreliable estimations or locally optimal solutions. In this work, we employs the excess kurtosis, a direct measure of non-Gaussianity, to establish a causal discovery method for linear non-Gaussian acyclic data. Firstly, we theoretically prove that an exogenous variable has the largest excess kurtosis when disturbance variables follow independent and identically distributions. Secondly, based on this property of exogenous variables, we propose an efficient exogenous variable identification algorithm, and develop a causal discovery method. Extensive experiment results verify the effectiveness and efficiency of the proposed approach. Ruichu Cai, Feng Xie 0002, Wei Chen 0103, Zhifeng Hao 0004 |
IWQoS | 1 |
| 2017 | Identification of adverse drug-drug interactions through causal association rule discovery from spontaneous adverse event reports
Ruichu Cai, Yong Hu 0002, Brittany Melton, Michael E. Matheny, Hua Xu 0001, Lemuel R. Waitman |
Artif. Intell. Medicine | 1 |
| 2017 | Recognizing activities from partially observed streams using posterior regularized conditional random fields
Wen Wen 0009, Ruichu Cai, Xiaowei Yang 0003 |
Neurocomputing | 2 |
| 2017 | DITIR: Distributed Index for High Throughput Trajectory Insertion and Real-time Temporal Range QueryabstractThe prosperity of mobile social network and location-based services, e.g., Uber, is backing the explosive growth of spatial temporal streams on the Internet. It raises new challenges to the underlying data store system, which is supposed to support extremely high-throughput trajectory insertion and low-latency querying with spatial and temporal constraints. State-of-the-art solutions, e.g., HBase, do not render satisfactory performance, due to the high overhead on index update. In this demonstration, we present DITIR, our new system prototype tailored to efficiently processing temporal and spacial queries over historical data as well as latest updates. Our system provides better performance guarantee, by physically partitioning the incoming data tuples on their arrivals and exploiting a template-based insertion schema, to reach the desired ingestion throughput. Load balancing mechanism is also introduced to DITIR, by using which the system is capable of achieving reliable performance against workload dynamics. Our demonstration shows that DITIR supports over 1 million tuple insertions in a second, when running on a 10-node cluster. It also significantly outperforms HBase by 7 times on ingestion throughput and 5 times faster on query latency. Ruichu Cai, Zijie Lu, Tom Z. J. Fu, Marianne Winslett |
Proc. VLDB Endow. | 1 |
| 2017 | Understanding Social Causalities Behind Human Action SequencesabstractSocial causality study on human action sequences is useful and important to improve our understandings to human behaviors on online social networks. The redundant indirect causalities and unobserved confounding factors, such as homophily and simultaneity phenomena, contribute to the huge challenges on accurate causal discovery on such human actions. A causal relationship exists between two persons, if the actions of one person are significantly affected by the actions of the other person, while fairly independent of her/his own prior actions. In this paper, we design a systematic approach based on conditional independence testing to detect such asymmetric relations, even when there are latent confounders underneath the observational action sequences. Technically, a group of asymmetric independence tests are conducted to infer the loose causal directions between action sequence pairs, followed by another group of tests to distinguish different types of relationships, e.g., homophily and simultaneity. Finally, a causal structure learning method is employed to output pairwise causalities with redundant indirect causalities eliminated. Empirical evaluations on simulated data verify the effectiveness and scalability of our proposals. We also present four interesting patterns of causal relations found by our algorithm, on real Sina Weibo feeds, including two new patterns never reported in previous studies. Ruichu Cai, Zhifeng Hao 0004, Marianne Winslett |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Multi-Domain Manifold Learning for Drug-Target Interaction PredictionabstractDrug-target interaction (DTI) provides novel insights about the genomic drug discovery, and is a critical technique to drug discovery. Recently, researchers try to incorporate different information about drugs and targets for prediction. However, the heterogeneous and high-dimensional data poses huge challenge to existing machine learning methods. In the last few years, extensive research efforts have been devoted to the utilization of manifold property on high dimensional data, e.g. dimension reduction methods preserving local structures of the manifolds. Motivated by the successes of these studies, we propose a general framework incorporating both manifold structures and known interaction/non-interaction information to predict the drug-target interactions. To overcome the challenges of domain scaling and information inconsistency, we formulate the problem with Semidefinite Programming (SDP), including new constraints to improve the robustness of the learning procedure. A variety of optimization techniques are also designed to enhance the scalability of the problem solver. Effectiveness of the method is evaluated by experiments on the benchmark dataset. Compared with state-of-the-art methods, the proposed methods generate much more accurate drug-target interaction prediction. Ruichu Cai, Srinivasan Parthasarathy 0001, Anthony K. H. Tung, Wen Zhang 0008 |
SDM | 1 |
| 2016 | Multiple-cause discovery combined with structure learning for high-dimensional discrete data and application to stock prediction
Ruichu Cai, Xiangzhou Zhang, Yong Hu 0002 |
Soft Comput. | 3 |
| 2015 | A Semi-supervised Solution for Cold Start Issue on Recommender Systems
Zhifeng Hao 0004, Yingchao Cheng, Ruichu Cai, Wen Wen 0009 |
APWeb | 3 |
| 2015 | Causal discovery on high dimensional data
Ruichu Cai, Wen Wen 0009, Zhihao Li 0001 |
Appl. Intell. | 3 |
| 2015 | Deterministic identification of specific individuals from GWAS resultsabstractMOTIVATION: Genome-wide association studies (GWASs) are commonly applied on human genomic data to understand the causal gene combinations statistically connected to certain diseases. Patients involved in these GWASs could be re-identified when the studies release statistical information on a large number of single-nucleotide polymorphisms. Subsequent work, however, found that such privacy attacks are theoretically possible but unsuccessful and unconvincing in real settings. RESULTS: We derive the first practical privacy attack that can successfully identify specific individuals from limited published associations from the Wellcome Trust Case Control Consortium (WTCCC) dataset. For GWAS results computed over 25 randomly selected loci, our algorithm always pinpoints at least one patient from the WTCCC dataset. Moreover, the number of re-identified patients grows rapidly with the number of published genotypes. Finally, we discuss prevention methods to disable the attack, thus providing a solution for enhancing patient privacy. AVAILABILITY AND IMPLEMENTATION: Proofs of the theorems and additional experimental results are available in the support online documents. The attack algorithm codes are publicly available at https://sites.google.com/site/zhangzhenjie/GWAS_attack.zip. The genomic dataset used in the experiments is available at http://www.wtccc.org.uk/ on request. Ruichu Cai, Zhifeng Hao 0004, Marianne Winslett, Xiaokui Xiao, Yin Yang 0001, Shuigeng Zhou |
Bioinform. | 1 |
| 2015 | An improved clustering ensemble method based link analysis
Zhifeng Hao 0004, Li-Juan Wang, Ruichu Cai, Wen Wen 0009 |
World Wide Web | 3 |
| 2014 | A Causal Model for Disease Pathway Discovery
Ruichu Cai, Chang Yuan, Wen Wen 0009, Zhihao Li 0001 |
ICONIP (1) | 1 |
| 2014 | A general framework of hierarchical clustering and its applications
Ruichu Cai, Anthony K. H. Tung, Chenyun Dai |
Inf. Sci. | 1 |
| 2014 | Determining molecular predictors of adverse drug reactions with causality analysis based on structure learningabstractOBJECTIVE: Adverse drug reaction (ADR) can have dire consequences. However, our current understanding of the causes of drug-induced toxicity is still limited. Hence it is of paramount importance to determine molecular factors of adverse drug responses so that safer therapies can be designed. METHODS: We propose a causality analysis model based on structure learning (CASTLE) for identifying factors that contribute significantly to ADRs from an integration of chemical and biological properties of drugs. This study aims to address two major limitations of the existing ADR prediction studies. First, ADR prediction is mostly performed by assessing the correlations between the input features and ADRs, and the identified associations may not indicate causal relations. Second, most predictive models lack biological interpretability. RESULTS: CASTLE was evaluated in terms of prediction accuracy on 12 organ-specific ADRs using 830 approved drugs. The prediction was carried out by first extracting causal features with structure learning and then applying them to a support vector machine (SVM) for classification. Through rigorous experimental analyses, we observed significant increases in both macro and micro F1 scores compared with the traditional SVM classifier, from 0.88 to 0.89 and 0.74 to 0.81, respectively. Most importantly, identified links between the biological factors and organ-specific drug toxicities were partially supported by evidence in Online Mendelian Inheritance in Man. CONCLUSIONS: The proposed CASTLE model not only performed better in prediction than the baseline SVM but also produced more interpretable results (ie, biological factors responsible for ADRs), which is critical to discovering molecular activators of ADRs. Ruichu Cai, Yong Hu 0002, Michael E. Matheny, Jingchun Sun, Hua Xu 0001 |
J. Am. Medical Informatics Assoc. | 2 |
| 2013 | Chinese Sentiment Classification Based on the Sentiment Drop Point
Ruichu Cai, Wen Wen 0009 |
ICIC (3) | 3 |
| 2013 | A Hybrid Approach for Large Scale Causality Discovery
Ruichu Cai, Wen Wen 0009 |
ICIC (3) | 3 |
| 2013 | SADA: A General Framework to Support Robust Causation DiscoveryabstractCausality discovery without manipulation is considered a crucial problem to a variety of applications, such as genetic therapy. The state-of-the-art solutions, e.g. LiNGAM, return accurate results when the number of labeled samples is larger than the number of variables. These approaches are thus applicable only when large numbers of samples are available or the problem domain is sufficiently small. Motivated by the observations of the local sparsity properties on causal structures, we propose a general Split-and-Merge strategy, named SADA, to enhance the scalability of a wide class of causality discovery algorithms. SADA is able to accurately identify the causal variables, even when the sample size is significantly smaller than the number of variables. In SADA, the variables are partitioned into subsets, by finding cuts on the sparse probabilistic graphical model over the variables. By running mainstream causation discovery algorithms, e.g. LiNGAM, on the subproblems, complete causality can be reconstructed by combining all the partial results. SADA benefits from the recursive division technique, since each small subproblem generates more accurate result under the same number of samples. We theoretically prove that SADA always reduces the scale of problems without significant sacrifice on result accuracy, depending only on the local sparsity condition over the variables. Experiments on real-world datasets verify the improvements on scalability and accuracy by applying SADA on top of existing causation algorithms. Ruichu Cai |
ICML (2) | 1 |
| 2013 | Regularized Gaussian Mixture Model based discretization for gene expression data association mining
Ruichu Cai, Wen Wen 0009 |
Appl. Intell. | 1 |
| 2013 | Software project risk analysis using Bayesian networks with causality constraints
Yong Hu 0002, Xiangzhou Zhang, Eric W. T. Ngai, Ruichu Cai |
Decis. Support Syst. | 4 |
| 2013 | Product named entity recognition for Chinese query questions based on a skip-chain CRF model
Ruichu Cai, Wen Wen 0009 |
Neural Comput. Appl. | 3 |
| 2013 | Two novel interestingness measures for gene association rule mining
Meihua Wang, Shumin Wu, Ruichu Cai |
Neural Comput. Appl. | 3 |
| 2013 | Causal gene identification using combinatorial V-structure search
Ruichu Cai |
Neural Networks | 1 |
| 2011 | A new hybrid method for gene selection
Ruichu Cai, Xiaowei Yang 0003, Han Huang 0002 |
Pattern Anal. Appl. | 1 |
| 2011 | BASSUM: A Bayesian semi-supervised method for classification feature selection
Ruichu Cai |
Pattern Recognit. | 1 |
| 2011 | What is Unequal among the Equals? Ranking Equivalent Rules from Gene Expression DataabstractIn previous studies, association rules have been proven to be useful in classification problems over high dimensional gene expression data. However, due to the nature of such data sets, it is often the case that millions of rules can be derived such that many of them are covered by exactly the same set of training tuples and thus have exactly the same support and confidence. Ranking and selecting useful rules from such equivalent rule groups remain an interesting and unexplored problem. In this paper, we look at two interestingness measures for ranking the interestingness of rules within equivalent rule group: Max-Subrule-Conf and Min-Subrule-Conf. Based on these interestingness measures, an incremental Apriori-like algorithm is designed to select more interesting rules from the lower bound rules of the group. Moreover, we present an improved classification model to fully exploit the potential of the selected rules. Our empirical studies on our proposed methods over five gene expression data sets show that our proposals improve both the efficiency and effectiveness of the rule extraction and classifier construction over gene expression data sets. Ruichu Cai, Anthony K. H. Tung |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2010 | Kernel based gene expression pattern discovery and its application on cancer classification
Ruichu Cai, Wen Wen 0009, Han Huang 0002 |
Neurocomputing | 1 |
| 2009 | Kernel-based skyline cardinality estimationabstractThe skyline of a d-dimensional dataset consists of all points not dominated by others. The incorporation of the skyline operator into practical database systems necessitates an efficient and effective cardinality estimation module. However, existing theoretical work on this problem is limited to the case where all d dimensions are independent of each other, which rarely holds for real datasets. The state of the art Log Sampling (LS) technique simply applies theoretical results for independent dimensions to non-independent data anyway, sometimes leading to large estimation errors. To solve this problem, we propose a novel Kernel-Based (KB) approach that approximates the skyline cardinality with nonparametric methods. Extensive experiments with various real datasets demonstrate that KB achieves high accuracy, even in cases where LS fails. At the same time, despite its numerical nature, the efficiency of KB is comparable to that of LS. Furthermore, we extend both LS and KB to the k-dominant skyline, which is commonly used instead of the conventional skyline for high-dimensional data. Yin Yang 0001, Ruichu Cai, Dimitris Papadias, Anthony K. H. Tung |
SIGMOD Conference | 3 |
| 2009 | An efficient gene selection algorithm based on mutual information
Ruichu Cai, Xiaowei Yang 0003, Wen Wen 0009 |
Neurocomputing | 1 |
| 2007 | A Novel Gene Ranking Algorithm Based on Random Subspace MethodabstractGene selection is to select the most informative genes from the whole gene set. It's an important preprocessing procedure for the discriminant analysis of microarray data, because many of the genes are irrelevant or redundant to the discriminant problem. In this paper, the gene selection problem is considered as a gene ranking problem and a random subspace method based gene ranking (RSM-GR) algorithm is proposed. In RSM-GR, firstly subsets of the genes are randomly generated; then Support Vector Machines are respectively trained on each subset and thus produce the importance factor of each gene; finally, the importance of each gene obtained from these randomly selected subsets is combined to constitute its final importance. Experiments on two public datasets show that RSM-GR obtains gene sets leading to more accurate classification results than other gene selection methods, and it demands less computational time. RSM-GR can also better deal with datasets with a large number of genes and a big number of genes to be selected. Ruichu Cai, Wen Wen 0009 |
IJCNN | 1 |
| 2006 | A Novel ACO Algorithm with Adaptive Parameter
Han Huang 0002, Xiaowei Yang 0003, Ruichu Cai |
ICIC (3) | 4 |