Hao Zhang 0079

dblp:55/2270-79 · DBLP profile ↗
← Back
42ranked-venue papers
12as first author
36since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 8 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Invariant Feature Learning for Counterfactual Watch-time Prediction in Video Recommendation
abstract
Video recommendation systems heavily rely on user watch time feedback, making accurate watch time prediction a crucial task. However, this task inherently suffers from bias, as recommendation models tend to favor long-duration videos to maximize watch time. This issue, known as duration bias in the watch-time prediction context, can be explained from a causal perspective, where video duration acts as a confounder. Recent works address this bias using backdoor adjustment, isolating the direct effect of content on watch time from observational data. These methods typically discretize video duration into groups, estimate group-wise effects, and then aggregate them via a unified prediction model. However, this aggregation strategy is prone to model misspecification due to feature distribution shift across groups. In this paper, we reinterpret the problem through the lens of invariant learning and propose a novel framework: Duration-Invariant Feature Learning (DIFL). DIFL employs a kernel-based regularization that enforces representation invariance across duration groups, reducing sensitivity to group design and improving generalization. This enables more accurate modeling of the direct causal effect and making counterfactual inference. Extensive experiments on both public and real large-scale production datasets demonstrate the effectiveness of our approach, which achieves SOTA performance.
Chenghou Jin, Yixin Ren, Hongxu Ma 0001, Yewei Xia, Yi Guan, Hao Zhang 0079, Jiandong Ding, Jihong Guan, Shuigeng Zhou
AAAI6
2026 Gene-guided multimodal data fusion for cancer patient survival analysis
Mingjie Xu, Zongbao Yang, Ruxin Wang 0001, Hao Zhang 0079
Neurocomputing5
2026 Causal discovery by continuous optimization with weighted superstructure
Yewei Xia, Hao Zhang 0079, Ruxin Wang 0001, Yuzhong Peng, Jihong Guan, Shuigeng Zhou
Neural Networks3
2026 Causal direction discovery via related conditional residual
Shaofan Chen, Guoyuan He, Wentao Ma 0003, Hao Zhang 0079, Tongqing Zhou, Siwei Wang 0001, Lichuan Gu
Pattern Recognit.4
2026 Cyclic Contrastive Representation Learning for Incomplete Multi-Modal Medical Image Segmentation
abstract
Accurate segmentation of multimodal medical images with missing modalities remains a critical challenge due to incomplete data often encountered in clinical practice. Lack of modality-specific information often leads to significant performance degradation in scenarios with severely missing modalities. To address this problem, we focus on modeling the relationships between modality-specific features. We propose a joint representation learning framework, named as Cyclic Contrastive Latent Representation Segmentation (CLRS), which incorporates cyclic modality-specific representation generation and contrastive feature alignment for robust 3D medical image segmentation under missing modality conditions. CLRS first extracts feature from available modalities using a unified encoder, then generates missing latent representations conditioned on the encoded features via an elaborately designed synthesis strategy. Meanwhile, a channel-wise attention mechanism is introduced to enhance the specific features of the modality. In addition, modality-specific contrastive learning enforces cross-modal discrimination between the generated and encoded representations, which effectively disentangles modality-specific information from shared patterns and enhances the segmentation robustness in missing modality scenarios. Extensive experiments on three 3D multimodal datasets demonstrate the superior performance of CLRS, particularly in scenarios with severe modality absence. For instance, with only a single modality available on ProstateZS dataset, CLRS improves the state-of-the-art (SOTA) by over 4.06% for peripheral zone, 2.20% for central gland.
Shihuan He, Zongbao Yang, Hao Zhang 0079, Ruxin Wang 0001
IEEE J. Biomed. Health Informatics4
2026 Causal Discovery by Multi-Level Wavelet Mapping Correlation Based Statistical Dependence Measurement
abstract
This article proposes a new method for causal discovery based on a novel dependence measurement criterion, namely, Multi-level Wavelet Mapping Correlation (MWMC). MWMC captures nonlinear dependencies between variables by measuring their correlations across multiple levels of wavelet mappings. From a theoretical perspective, we show that the empirical estimate of MWMC converges exponentially fast to its population quantity. Under the null hypothesis of independence, we further design a permutation-based independence testing procedure, termed the Wavelet Independence Test (WIT), built upon MWMC. We prove that WIT not only effectively controls the Type I error rate (false positives), but also guarantees that the Type II error rate (false negatives) is upper bounded by \(\mathcal{O}(n^{-1})\) , where \( n \) denotes the sample size, even with a finite number of permutations. Building on these theoretical guarantees, we derive a causal discovery method by integrating MWMC-based WIT into standard causal discovery pipelines. Extensive experiments on (conditional) independence testing and causal discovery using both synthetic and real-world datasets with varying sample sizes demonstrate that our approach consistently outperforms existing independence testing and causal discovery methods in terms of reduced Type II error rates and statistically validated performance improvements. Impact Statement —Causal discovery is a fundamental task in knowledge discovery, aiming to uncover the underlying data-generating mechanisms in order to support more accurate and interpretable predictions. Statistical independence tests and conditional independence (CI) tests have long served as core tools in this area. To improve the reliability of independence testing, we propose a novel test, WIT, which achieves lower Type II error rates in 19 out of 25 distinct experimental scenarios involving diverse data distributions, compared to 15 out of 25 for the strongest existing baseline. We further apply WIT to CI testing and causal discovery, and extensive empirical results show that it consistently improves the performance of multiple causal discovery algorithms across a range of experimental settings.
Yixin Ren, Hao Zhang 0079, Yewei Xia, Feng Xie 0002, Jihong Guan, Shuigeng Zhou
ACM Trans. Knowl. Discov. Data2
2025 AdaMM: An Adaptive Multimodal Model with Learnable Weights for Protein-Ligand Affinity Prediction
abstract
Protein–ligand binding affinity prediction plays a pivotal role in the field of drug discovery, with multimodalbased methods standing out. Existing approaches based on a multimodal framework typically rely on simple concatenation or pooling, which fail to identify and emphasize key features across heterogeneous sources. To tackle this bottleneck, we propose an adaptive fusion module, which injects sequence–structure embeddings into a functional annotation stream while incorporating functional annotation embeddings into the sequence–structure stream. Subsequently, learnable adaptive weights are employed to combine the outputs of these two pathways in a fully data-driven manner. Extensive experiments on the PDBBind benchmark demonstrate that our method achieves the best performance compared with state-of-the-art methods. Ablation study and hyperparameter analysis confirm that sequence, structure, and functional annotation each provide complementary information, and the joint optimization of these modalities via our adaptive fusion strategy yields the highest overall predictive accuracy. Our code is available at GitHub link https://github.com/Jessez2/AdaMM.
Juncai Zhang, Huazhen Huang, Yixin Ren, Yuzhong Peng, Ruxin Wang 0001, Hao Zhang 0079
BIBM7
2025 Data-Driven Selection of Instrumental Variables for Additive Nonlinear, Constant Effects Models
abstract
We consider the problem of selecting instrumental variables from observational data, a fundamental challenge in causal inference. Existing methods mostly focus on additive linear, constant effects models, limiting their applicability in complex real-world scenarios. In this paper, we tackle a more general and challenging setting: the additive non-linear, constant effects model. We first propose a novel testable condition, termed the Cross Auxiliary-based independent Test (CAT) condition, for selecting the valid IV set. We show that this condition is both necessary and sufficient for identifying valid instrumental variable sets within such a model under milder assumptions. Building on this condition, we develop a practical algorithm for selecting the set of valid instrumental variables. Extensive experiments on both synthetic and two real-world datasets demonstrate the effectiveness and robustness of our proposed approach, highlighting its potential for broader applications in causal analysis.
Xichen Guo, Feng Xie 0002, Yan Zeng 0002, Hao Zhang 0079, Zhi Geng
ICML4
2025 Local Identifying Causal Relations in the Presence of Latent Variables
abstract
We tackle the problem of identifying whether a variable is the cause of a specified target using observational data. State-of-the-art causal learning algorithms that handle latent variables typically rely on identifying the global causal structure, often represented as a partial ancestral graph (PAG), to infer causal relationships. Although effective, these approaches are often redundant and computationally expensive when the focus is limited to a specific causal relationship. In this work, we introduce novel local characterizations that are necessary and sufficient for various types of causal relationships between two variables, enabling us to bypass the need for global structure learning. Leveraging these local insights, we develop efficient and fully localized algorithms that accurately identify causal relationships from observational data. We theoretically demonstrate the soundness and completeness of our approach. Extensive experiments on benchmark networks and real-world datasets further validate the effectiveness and efficiency of our method.
Feng Xie 0002, Hao Zhang 0079, Zhi Geng
ICML4
2025 SERENA: A Unified Stochastic Recursive Variance Reduced Gradient Framework for Riemannian Non-Convex Optimization
abstract
Recently, the expansion of Variance Reduction (VR) to Riemannian stochastic non-convex optimization has attracted increasing interest. Inspired by recursive momentum, we first introduce Stochastic Recursive Variance Reduced Gradient (SRVRG) algorithm and further present Stochastic Recursive Gradient Estimator (SRGE) in Euclidean spaces, which unifies the prevailing variance reduction estimators. We then extend SRGE to Riemannian spaces, resulting in a unified Stochastic rEcursive vaRiance reducEd gradieNt frAmework (SERENA) for Riemannian non-convex optimization. This framework includes the proposed R-SRVRG, R-SVRRM, and R-Hybrid-SGD methods, as well as other existing Riemannian VR methods. Furthermore, we establish a unified theoretical analysis for Riemannian non-convex optimization under retraction and vector transport. The IFO complexity of our proposed R-SRVRG and R-SVRRM to converge to $\varepsilon$-accurate solution is $\mathcal{O}\left(\min \{n^{1/2}{\varepsilon^{-2}}, \varepsilon^{-3}\}\right)$ in the finite-sum setting and ${\mathcal{O}\left( \varepsilon^{-3}\right)}$ for the online case, both of which align with the lower IFO complexity bound. Experimental results indicate that the proposed algorithms surpass other existing Riemannian optimization methods.
Chaojie Ji, Hao Zhang 0079, Ruxin Wang 0001
ICML4
2025 Identifying Causal Mechanism Shifts Under Additive Models with Arbitrary Noise
abstract
In many real-world scenarios, the goal is to identify variables whose causal mechanisms change across related datasets. For example, detecting abnormal root nodes in manufacturing, and identifying key genes that influence cancer by analyzing differences in gene regulatory mechanisms between healthy individuals and cancer patients. This can be done by recovering the causal structure for each dataset independently and then comparing them to identify differences, but the performance is often suboptimal. Typically, existing methods directly identify causal mechanism shifts based on linear additive noise models (ANMs) or by imposing restrictive assumptions on the noise distribution. In this paper, we introduce CMSI, a novel and more general algorithm based on nonlinear ANMs that identifies variables with shifting causal mechanisms under arbitrary noise distributions. Evaluated on various synthetic datasets, CMSI consistently outperforms existing baselines in terms of F1 score. Additionally, we demonstrate CMSI's applicability on gene expression datasets of ovarian cancer patients at different disease stages.
Yewei Xia, Xueliang Cui, Hao Zhang 0079, Yixin Ren, Feng Xie 0002, Jihong Guan, Ruxin Wang 0001, Shuigeng Zhou
IJCAI3
2025 Efficient Constraint-based Window Causal Graph Discovery in Time Series with Multiple Time Lags
abstract
We address the identification of direct causes in time series with multiple time lags, and propose a constraint-based window causal graph discovery method. A key advantage of our method is that the number of required conditional independence (CI) tests scales quadratically with the number of sub-series. The method first uses CI tests to find the minimum trek lag between two arbitrary sub-series, followed by designing an efficient CI testing strategy to identify the direct causes between them. We show that the method is both sound and complete under some graph constraints. We compare the proposed method with typical baselines on various datasets. Experimental results show that our method outperforms all the counterparts in both accuracy and running speed.
Yewei Xia, Yixin Ren, Hong Cheng 0001, Hao Zhang 0079, Jihong Guan, Minchuan Xu, Shuigeng Zhou
IJCAI4
2025 Score-based Generative Modeling for Conditional Independence Testing
abstract
Determining conditional independence (CI) relationships between random variables is a fundamental yet challenging task in machine learning and statistics, especially in high-dimensional settings. Existing generative model-based CI testing methods, such as those utilizing generative adversarial networks (GANs), often struggle with undesirable modeling of conditional distributions and training instability, resulting in subpar performance. To address these issues, we propose a novel CI testing method via score-based generative modeling, which achieves precise Type I error control and strong testing power. Concretely, we first employ a sliced conditional score matching scheme to accurately estimate conditional score and use Langevin dynamics conditional sampling to generate null hypothesis samples, ensuring precise Type I error control. Then, we incorporate a goodness-of-fit stage into the method to verify generated samples and enhance interpretability in practice. We theoretically establish the error bound of conditional distributions modeled by score-based generative models and prove the validity of our CI tests. Extensive experiments on both synthetic and real-world datasets show that our method significantly outperforms existing state-of-the-art methods, providing a promising way to revitalize generative model-based CI testing.
Yixin Ren, Chenghou Jin, Yewei Xia, Longtao Huang, Hui Xue 0001, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou
KDD (2)7
2025 Fast Causal Discovery by Approximate Kernel-based Generalized Score Functions with Linear Computational Complexity
Yixin Ren, Haocheng Zhang, Yewei Xia, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou
KDD (1)4
2025 Local Learning for Covariate Selection in Nonparametric Causal Effect Estimation with Latent Variables
abstract
Estimating causal effects from nonexperimental data is a fundamental problem in many fields of science. A key component of this task is selecting an appropriate set of covariates for confounding adjustment to avoid bias. Most existing methods for covariate selection often assume the absence of latent variables and rely on learning the global causal structure among variables. However, identifying the global structure can be unnecessary and inefficient, especially when our primary interest lies in estimating the effect of a treatment variable on an outcome variable. To address this limitation, we propose a novel local learning approach for covariate selection in nonparametric causal effect estimation, which accounts for the presence of latent variables. Our approach leverages testable independence and dependence relationships among observed variables to identify a valid adjustment set for a target causal relationship, ensuring both soundness and completeness under standard assumptions. We validate the effectiveness of our algorithm through extensive experiments on both synthetic and real-world data.
Xichen Guo, Feng Xie 0002, Yan Zeng 0002, Hao Zhang 0079, Zhi Geng
NeurIPS5
2025 Regression-based conditional independence test with adaptive kernels
Yixin Ren, Juncai Zhang, Yewei Xia, Ruxin Wang 0001, Feng Xie 0002, Jihong Guan, Hao Zhang 0079, Shuigeng Zhou
Artif. Intell.7
2025 Both reliable and unreliable predictions matter: Domain adaptation for bearing fault diagnosis without source data
Wenyi Wu, Hao Zhang 0079, Zhisen Wei, Xiaoyuan Jing, Songsong Wu
Neurocomputing2
2025 Dynamic debiasing of multi-hop fact verification via counterfactual reasoning
Yuzhong Peng, Zongbao Yang, Zhichen Chen, Chang-an Yuan 0001, Xiao Qin 0005, Ruxin Wang 0001, Hao Zhang 0079
Knowl. Based Syst.8
2025 GCCNet: A Novel Network Leveraging Gated Cross-Correlation for Multi-View Classification
abstract
Multi-view learning is a machine learning paradigm that utilizes multiple feature sets or data sources to improve learning performance and generalization. However, existing multi-view learning methods often do not capture and utilize information from different views very well, especially when the relationships between views are complex and of varying quality. In this paper, we propose a novel multi-view learning framework for the multi-view classification task, called Gated Cross-Correlation Network (GCCNet), which addresses these challenges by integrating the three key operational levels in multi-view learning: representation, fusion, and decision. Specifically, GCCNet contains a novel component called the Multi-View Gated Information Distributor (MVGID) to enhance noise filtering and optimize the retention of critical information. In addition, GCCNet uses cross-correlation analysis to reveal dependencies and interactions between different views, as well as integrates an adaptive weighted joint decision strategy to mitigate the interference of low-quality views. Thus, GCCNet can not only comprehensively capture and utilize information from different views, but also facilitate information exchange and synergy between views, ultimately improving the overall performance of the model. Extensive experimental results on ten benchmark datasets show GCCNet's outperforms state-of-the-art methods on eight out of ten datasets, validating its effectiveness and superiority in multi-view learning.
Yuanpeng Zeng, Hao Zhang 0079, Shaojie Qiao, Faliang Huang, Qing Tian 0001, Yuzhong Peng
IEEE Trans. Multim.3
2024 Learning Adaptive Kernels for Statistical Independence Tests
abstract
We propose a novel framework for kernel-based statistical independence tests that enable adaptatively learning parameterized kernels to maximize test power. Our framework can effectively address the pitfall inherent in the existing signal-to-noise ratio criterion by modeling the change of the null distribution during the learning process. Based on the proposed framework, we design a new class of kernels that can adaptatively focus on the significant dimensions of variables to judge independence, which makes the tests more flexible than using simple kernels that are adaptive only in length-scale, and especially suitable for high-dimensional complex data. Theoretically, we demonstrate the consistency of our independence tests, and show that the non-convex objective function used for learning fits the L-smoothing condition, thus benefiting the optimization. Experimental results on both synthetic and real data show the superiority of our method. The source code and datasets are available at \url{https://github.com/renyixin666/HSIC-LK.git}.
Yixin Ren, Yewei Xia, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou
AISTATS3
2024 Efficiently Learning Significant Fourier Feature Pairs for Statistical Independence Testing
abstract
We propose a novel method to efficiently learn significant Fourier feature pairs for maximizing the power of Hilbert-Schmidt Independence Criterion~(HSIC) based independence tests. We first reinterpret HSIC in the frequency domain, which reveals its limited discriminative power due to the inability to adapt to specific frequency-domain features under the current inflexible configuration. To remedy this shortcoming, we introduce a module of learnable Fourier features, thereby developing a new criterion. We then derive a finite sample estimate of the test power by modeling the behavior of the criterion, thus formulating an optimization objective for significant Fourier feature pairs learning. We show that this optimization objective can be computed in linear time (with respect to the sample size $n$), which ensures fast independence tests. We also prove the convergence property of the optimization objective and establish the consistency of the independence tests. Extensive empirical evaluation on both synthetic and real datasets validates our method's superiority in effectiveness and efficiency, particularly in handling high-dimensional data and dealing with large-scale scenarios.
Yixin Ren, Yewei Xia, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou
NeurIPS3
2024 CDSC: Causal decomposition based on spectral clustering
Shaofan Chen, Yuzhong Peng, Guoyuan He, Hao Zhang 0079, Chengdong Wei
Inf. Sci.4
2024 CDRM: Causal disentangled representation learning for missing data
Ruxin Wang 0001, Yuzhong Peng, Hao Zhang 0079
Knowl. Based Syst.5
2024 Towards Effective Causal Partitioning by Edge Cutting of Adjoint Graph
abstract
Causal partitioningis an effective approach for causal discovery based on the divide-and-conquer strategy. Up to now, various heuristic methods based on conditional independence (CI) tests have been proposed for causal partitioning. However, most of these methods fail to achieve satisfactory partitioning without violating$d$-separation, leading to poor inference performance. In this work, we transform causal partitioning into an alternative problem that can be more easily solved. Concretely, we first construct a superstructure$G$of the true causal graph$G_{\mathcal {T}}$by performing a set of low-order CI tests on the observed data$D$. Then, we leverage point-line duality to obtain a graph$G_\mathcal {A}$adjoint to$G$. We show that the solution ofminimizing edge-cut ratioon$G_\mathcal {A}$can lead to a valid causal partitioning withsmaller causal-cut ratioon$G$andwithout violating$d$d-separation. We design an efficient algorithm to solve this problem. Extensive experiments show that the proposed method can achieve significantly better causal partitioning without violating$d$-separation than the existing methods.
Hao Zhang 0079, Yixin Ren, Yewei Xia, Shuigeng Zhou, Jihong Guan
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Hybrid Causal Feature Selection for Cancer Biomarker Identification From RNA-Seq Data
abstract
The discovery of cancer biomarkers helps to advance medical diagnosis and plays an important role in biomedical applications. Most of the existing data-driven methods identify biomarkers by ranking-based strategies, which generally return a subset or superset of the actual biomarkers, while some other causal-wise feature selection methods are based on Markov Blanket (MB) learning, facing the challenges of high-dimensionality & low-sample. In this work, we propose a novel hybrid causal feature selection method (called CAFES) to support large-scale cancer biomarker discovery from real RNA-seq data. Concretely, CAFES first uses minimal-redundancy & maximal-relevance strategy for dimensionality reduction that returns a set of candidate features. CAFES then learns the causal skeleton w.r.t. those features by CI tests and further obtains an appropriate superset of the MB of the target variable. Finally, CAFES learns the causal structure of this superset by the DAG-GNN algorithm and then obtains the MB of the target variable, which can be treated as the cancer biomarkers. We conduct experiments to evaluate the proposed method on two real well-known RNA-seq datasets that covering both binary and multi-class cases. We compare our method CAFES with seven recent methods including Semi-HITON-MB, STMB, BAMB, FBED, LCS-FS, EEMB, and EAMB. The results show that CAFES can identify dozens of cancer biomarkers, and of the discovered biomarkers can be verified by existing works that they are really directly related to the corresponding disease. An advantage of CAFES is that its Recall is significantly higher than those of all the counterparts, indicating that the continuous optimization (DAG-GNN) with the returned causal skeleton after feature selection (that can be treated as a conditional independence-based constraint to the optimization problem) is effective in cancer biomarkers identification under high-dimensional and low-sample RNA-seq data.
Wenwei Xu, Hao Zhang 0079, Yewei Xia, Yixin Ren, Jihong Guan, Shuigeng Zhou
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 Differentially Private Nonlinear Causal Discovery from Numerical Data
abstract
Recently, several methods such as private ANM, EM-PC and Priv-PC have been proposed to perform differentially private causal discovery in various scenarios including bivariate, multivariate Gaussian and categorical cases. However, there is little effort on how to conduct private nonlinear causal discovery from numerical data. This work tries to challenge this problem. To this end, we propose a method to infer nonlinear causal relations from observed numerical data by using regression-based conditional independence test (RCIT) that consists of kernel ridge regression (KRR) and Hilbert-Schmidt independence criterion (HSIC) with permutation approximation. Sensitivity analysis for RCIT is given and a private constraint-based causal discovery framework with differential privacy guarantee is developed. Extensive simulations and real-world experiments for both conditional independence test and causal discovery are conducted, which show that our method is effective in handling nonlinear numerical cases and easy to implement. The source code of our method and data are available at https://github.com/Causality-Inference/PCD.
Hao Zhang 0079, Yewei Xia, Yixin Ren, Jihong Guan, Shuigeng Zhou
AAAI1
2023 Multi-Level Wavelet Mapping Correlation for Statistical Dependence Measurement: Methodology and Performance
abstract
We propose a new criterion for measuring dependence between two real variables, namely, Multi-level Wavelet Mapping Correlation (MWMC). MWMC can capture the nonlinear dependencies between variables by measuring their correlation under different levels of wavelet mappings. We show that the empirical estimate of MWMC converges exponentially to its population quantity. To support independence test better with MWMC, we further design a permutation test based on MWMC and prove that our test can not only control the type I error rate (the rate of false positives) well but also ensure that the type II error rate (the rate of false negatives) is upper bounded by O(1/n) (n is the sample size) with finite permutations. By extensive experiments on (conditional) independence tests and causal discovery, we show that our method outperforms existing independence test methods.
Yixin Ren, Hao Zhang 0079, Yewei Xia, Jihong Guan, Shuigeng Zhou
AAAI2
2023 GEP-DL4Mol: A Novel Molecular Deep-learning Model Optimization Framework for Boosting Molecular Properties Prediction*
abstract
High-performance molecular property prediction is one of the essential problems in many research fields. Although Deep Learning models (DLMs) have the strong potential to predict molecular property, the search space of their neural structure and hyperparameters is too huge to enumerate all possibilities, which implies that one will be difficult to get a optimal configuration solution, or even build a DLM with poor performance. To tackle the aforementioned problem, we first abstractly map various types of supervised DLMs for chemical molecular property prediction under a unified molecular DLM conceptualization. Then, we develop a self-learning molecular DLMs construction and optimization framework, called GEP-DL4Mol, based on the unified supervised molecular DLM conceptualization and Gene Expression Programming. Experiments on eight benchmark datasets including 19 property prediction tasks were conducted to evaluate the performance of the proposed method. Experimental results show that the proposed method outperforms other methods.
Yuzhong Peng, Hao Zhang 0079, Ziqiao Zhang, Yanmei Lin, Shuigeng Zhou, Shaojie Qiao
BIBM2
2023 Causal Discovery by Continuous Optimization with Conditional Independence Constraint: Methodology and Performance
abstract
Discovering causal relationships from observational data is a challenging topic in artificial intelligence. Recent works formulate causal discovery as a continuous optimization problem with a differentiable acyclic constraint. Although these methods have achieved considerable performance improvement, they have two drawbacks: 1) they require a relatively large number of training samples; and 2) their performance will substantially deteriorate when facing heterogeneous noise. To address these problems, we first propose a low-order conditional independence (CI) constraint for the continuous optimization problem, and then design a soft version of the constraint by transforming it to a regularization term in the loss function of the continuous optimization problem. We show the convergence of continuous optimization with our constraint under some mild conditions, and the consistency of causal structure learning with the CI regularization. Extensive experiments on both synthetic and real-world datasets show that with our CI constraint or regularization, existing continuous optimization methods can achieve considerable performance improvement of causal discovery, especially when sample size is small.
Yewei Xia, Hao Zhang 0079, Yixin Ren, Jihong Guan, Shuigeng Zhou
ICDM2
2023 Causal Gene Identification Using Non-Linear Regression-Based Independence Tests
abstract
With the development of biomedical techniques in the past decades, causal gene identification has become one of the most promising applications in human genome-based business, which can help doctors to evaluate the risk of certain genetic diseases and provide further treatment recommendations for potential patients. When no controlled experiments can be applied, machine learning techniques like causal inference-based methods are generally used to identify causal genes. Unfortunately, most of the existing methods detect disease-related genes by ranking-based strategies or feature selection techniques, which generally return a superset of the corresponding real causal genes. There are also some causal inference-based methods that can identify a part of real causal genes from those supersets, but they are just able to return a few causal genes. This is contrary to our knowledge, as many results from controlled experiments have demonstrated that a certain disease, especially cancer, is usually related to dozens or hundreds of genes. In this work, we present an effective approach for identifying causal genes from gene expression data by using a new search strategy based on non-linear regression-based independence tests, which is able to greatly reduce the search space, and simultaneously establish the causal relationships from the candidate genes to the disease variable. Extensive experiments on real-world cancer datasets show that our method is superior to the existing causal inference-based methods in three aspects: 1) our method can identify dozens of causal genes, and 1/3 ∼ 1/2 of the discovered causal genes can be verified by existing works that they are really directly related to the corresponding disease; 2) The discovered causal genes are able to distinguish the status or disease subtype of the target patient; 3) Most of the discovered causal genes are closely relevant to the disease variable.
Hao Zhang 0079, Chuanxu Yan, Yewei Xia, Jihong Guan, Shuigeng Zhou
IEEE ACM Trans. Comput. Biol. Bioinform.1
2023 Conditional Independence Test Based on Residual Similarity
abstract
Recently, many regression-based conditional independence (CI) test methods have been proposed to solve the problem of causal discovery. These methods provide alternatives to test CI of x,y given Z by first removing the information of the controlling set Z from x and y , and then testing the independence between the two residuals R x,Z and R y,Z . When the residuals are linearly uncorrelated, the independence test between them is nontrivial. With the ability to calculate inner product in high-dimensional space, kernel-based methods are usually used to achieve this goal, but they are considerably time-consuming. In this paper, we test the independence between two linear combinations under linear structural equation model. We show that the dependence between the two residuals can be captured by the difference between the similarity of R x,Z and R y,Z and that of R x,Z and R r ( R r is an independent copy of R y,Z ) in high-dimensional space. With this result, we provide a new way to test CI based on the similarity between residuals, which is called SCIT — the abbreviation of Similarity-based CI Testing. Furthermore, we develop two versions of the proposal, called Kernel-SCIT and Neural-SCIT, respectively. Kernel-SCIT calculates the similarity by using kernel functions, while Neural-SCIT approximates the upper bound of the similarity by using deep neural networks. In both algorithms, random permutation tests are performed to control Type I error rate. The proposed tests are evaluated on (conditional) independence test and causal discovery with both synthetic and real datasets. Experimental results show that Kernel-SCIT is simpler yet more efficient and effective than the typical existing kernel-based methods HSIC and KCIT in the cases of small sample size, and Neural-SCIT can significantly boost the performance of CI testing when sufficient samples are available. The source code is available at https://github.com/xyw5vplus1/SCIT .
Hao Zhang 0079, Yewei Xia, Kun Zhang 0001, Shuigeng Zhou, Jihong Guan
ACM Trans. Knowl. Discov. Data1
2022 Residual Similarity Based Conditional Independence Test and Its Application in Causal Discovery
abstract
Recently, many regression based conditional independence (CI) test methods have been proposed to solve the problem of causal discovery. These methods provide alternatives to test CI by first removing the information of the controlling set from the two target variables, and then testing the independence between the corresponding residuals Res1 and Res2. When the residuals are linearly uncorrelated, the independence test between them is nontrivial. With the ability to calculate inner product in high-dimensional space, kernel-based methods are usually used to achieve this goal, but still consume considerable time. In this paper, we investigate the independence between two linear combinations under linear non-Gaussian structural equation model. We show that the dependence between the two residuals can be captured by the difference between the similarity of (Res1, Res2) and that of (Res1, Res3) (Res3 is generated by random permutation) in high-dimensional space. With this result, we design a new method called SCIT for CI test, where permutation test is performed to control Type I error rate. The proposed method is simpler yet more efficient and effective than the existing ones. When applied to causal discovery, the proposed method outperforms the counterparts in terms of both speed and Type II error rate, especially in the case of small sample size, which is validated by our extensive experiments on various datasets.
Hao Zhang 0079, Shuigeng Zhou, Kun Zhang 0001, Jihong Guan
AAAI1
2022 An automatic hyperparameter optimization DNN model for precipitation prediction
Yuzhong Peng, Daoqing Gong, Chuyan Deng, Hongya Li, Hongguo Cai, Hao Zhang 0079
Appl. Intell.6
2022 Learning Causal Structures Based on Divide and Conquer
abstract
This article addresses two important issues of causal inference in the high-dimensional situation. One is how to reduce redundant conditional independence (CI) tests, which heavily impact the efficiency and accuracy of existing constraint-based methods. Another is how to construct the true causal graph from a set of Markov equivalence classes returned by these methods. For the first issue, we design a recursive decomposition approach where the original data (a set of variables) are first decomposed into two small subsets, each of which is then recursively decomposed into two smaller subsets until none of these subsets can be decomposed further. Redundant CI tests can be reduced by inferring causalities from these subsets. The advantage of this decomposition scheme lies in two aspects: 1) it requires only low-order CI tests and 2) it does not violate d -separation. The complete causality can be reconstructed by merging all the partial results of the subsets. For the second issue, we employ regression-based CI tests to check CIs in linear non-Gaussian additive noise cases, which can identify more causal directions by [Formula: see text] (or [Formula: see text]). Consequently, causal direction learning is no longer limited by the number of returned V -structures and consistent propagation. Extensive experiments show that the proposed method can not only substantially reduce redundant CI tests but also effectively distinguish the equivalence classes.
Hao Zhang 0079, Shuigeng Zhou, Chuanxu Yan, Jihong Guan, Xin Wang 0004, Ji Zhang 0001, Jun Huan
IEEE Trans. Cybern.1
2021 Testing Independence Between Linear Combinations for Causal Discovery
abstract
Recently, regression based conditional independence (CI) tests have been employed to solve the problem of causal discovery. These methods provide an alternative way to test for CI by transforming CI to independence between residuals. Generally, it is nontrivial to check for independence when these residuals are linearly uncorrelated. With the ability to represent high-order moments, kernel-based methods are usually used to achieve this goal, but at a cost of considerable time. In this paper, we investigate the independence between two linear combinations under linear non-Gaussian structural equation model (SEM). We show that generally the 1-st to 4-th moments of the two linear combinations contain enough information to infer whether or not they are independent. The proposed method provides a simpler but more effective way to measure CIs, with only calculating the 1-st to 4-th moments of the input variables. When applied to causal discovery, the proposed method outperforms kernel-based methods in terms of both speed and accuracy. which is validated by extensive experiments.
Hao Zhang 0079, Kun Zhang 0001, Shuigeng Zhou, Jihong Guan, Ji Zhang 0001
AAAI1
2021 Combined cause inference: Definition, model and performance
abstract
In recent years, many methods have been developed for discovering causal relationships from observed data. However, as an important kind of causes existing in many causal systems, combined causes (e.g. multi-factor causes consisting of two or more component variables that individually might not be a cause) have not received enough attention. The existing approach includes both individual and combined variables in the causal discovery process using constraint-based methods, can neither distinguish a set of Markov equivalence classes nor identify a combined cause containing one (or more) individual cause(s), therefore can output only some combined causes, instead of all combined causes. In this paper, we first subsume all possible combined causes into three types and give them formal definitions, then extend the additive noise model (ANM) to infer combined causes. We show that if a candidate variable set X w.r.t. a target Y satisfies: (1) allowing ANM for only the forward direction X→Y, and (2) no disturbance variable is contained in X, i.e., removing any component of X will weaken the causal relationship between X and Y, then X forms a combined cause. Based on this finding, we develop an efficient method to discover combined causes. Furthermore, we also conduct extensive experiments to validate the proposed method on both synthetic and real-world data sets.
Hao Zhang 0079, Chuanxu Yan, Shuigeng Zhou, Jihong Guan, Ji Zhang 0001
Inf. Sci.1
2019 Recursively Learning Causal Structures Using Regression-Based Conditional Independence Test
abstract
This paper addresses two important issues in causality inference. One is how to reduce redundant conditional independence (CI) tests, which heavily impact the efficiency and accuracy of existing constraint-based methods. Another is how to construct the true causal graph from a set of Markov equivalence classes returned by these methods.For the first issue, we design a recursive decomposition approach where the original data (a set of variables) is first decomposed into three small subsets, each of which is then recursively decomposed into three smaller subsets until none of subsets can be decomposed further. Consequently, redundant CI tests can be reduced by inferring causality from these subsets. Advantage of this decomposition scheme lies in two aspects: 1) it requires only low-order CI tests, and 2) it does not violate d-separation. Thus, the complete causality can be reconstructed by merging all the partial results of the subsets.For the second issue, we employ regression-based conditional independence test to check CIs in linear non-Gaussian additive noise cases, which can identify more causal directions by x−E(x|Z)⊥z (or y−E(y|Z)⊥z). Therefore, causal direction learning is no longer limited by the number of returned Vstructures and the consistent propagation.Extensive experiments show that the proposed method can not only substantially reduce redundant CI tests but also effectively distinguish the equivalence classes, thus is superior to the state of the art constraint-based methods in causality inference.
Hao Zhang 0079, Shuigeng Zhou, Chuanxu Yan, Jihong Guan, Xin Wang 0004
AAAI1
2019 Measuring Conditional Independence by Independent Residuals for Causal Discovery
abstract
We investigate the relationship between conditional independence (CI) x ⫫ y | Z and the independence of two residuals x −E( x | Z )⫫ y −E( y | Z ), where x and y are two random variables and Z is a set of random variables. We show that if x , y , and Z are generated by following linear structural equation models and all external influences follow joint Gaussian distribution, then x ⫫ y | Z if and only if x −E( x | Z )⫫ y −E( y | Z ). That is, the test of x ⫫ y | Z can be relaxed to a simpler unconditional independence test of x −E( x | Z )⫫ y −E( y | Z ). Furthermore, testing x −E( x | Z )⫫ y −E( y | Z ) can be simplified by testing x −E( x | Z )⫫ y or y −E( y | Z )⫫ x . On the other side, if all these external influences follow non-Gaussian distributions and the model satisfies structural faithfulness condition, then we have x ⫫ y | Z ⇔ x −E( x | Z )⫫ y −E( y | Z ). We apply the results above to the causal discovery problem, where the causal directions are generally determined by a set of V -structures and their consistent propagations, so CI test-based methods can return a set of Markov equivalence classes. We show that in the linear non-Gaussian context, in many cases x −E( x | Z )⫫ z or y −E( y | Z )⫫ z (∀ z ∈ Z and Z is a minimal d -separator) is satisfied when x −E( x | Z )⫫ y −E( y | Z ), which implies z causes x (or y ) if z directly connects to x (or y ). Therefore, we conclude that CIs have useful information for distinguishing Markov equivalence classes. In summary, comparing with the existing discretization-based and kernel-based CI testing methods, the proposed method provides a simpler way to measure CI, which needs only one unconditional independence test and two regression operations. When being applied to causal discovery, it can find more causal relationships, which is extensively validated by experiments.
Hao Zhang 0079, Shuigeng Zhou, Jihong Guan, Jun Huan
ACM Trans. Intell. Syst. Technol.1
2018 Measuring Conditional Independence by Independent Residuals: Theoretical Results and Application in Causal Discovery
abstract
We investigate the relationship between conditional independence (CI) x ⊥ y|Z and the independence of two residuals x – E(x|Z) ⊥ –E(y|Z), where x and y are two random variables, and Z is a set of random variables. We show that if x, y and Z are generated by following linear structural equation model and all external influences follow Gaussian distributions, then x ⊥ y|Z if and only if x – E(x|Z) ⊥ y – E(y|Z). That is, the test of x ⊥ y|Z can be relaxed to a simpler unconditional independence test of x – E(x|Z) ⊥ y – E(y|Z). Furthermore, if all these external influences follow non-Gaussian distributions and the model satisfies structural faithfulness condition, then we have x ⊥ y|Z ⇔ x – E(x|Z) ⊥ y – E(y|Z). We apply the results above to the causal discovery problem, where the causal directions are generally determined by a set of V-structures and their consistent propagations, so CI test-based methods can return a set of Markov equivalence classes. We show that in linear non-Gaussian context, x – E(x|Z) ⊥ y – E(y|Z) ⇒ x – E(x|Z) ⊥ z or y – E(y|Z ⊥ z (∀z ∈ Z) if Z is a minimal d-separator, which implies z causes x (or y) if z directly connects to x (or y). Therefore, we conclude that CIs have useful information for distinguishing Markov equivalence classes. In summary, compared with the existing discretization-based and kernel-based CI testing methods, the proposed method provides a simpler way to measure CI, which needs only one unconditional independence test and two regression operations. When being applied to causal discovery, it can find more causal relationships, which is experimentally validated.
Hao Zhang 0079, Shuigeng Zhou, Jihong Guan
AAAI1
2018 Multicellular Gene Expression Programming-Based Hybrid Model for Precipitation Prediction Coupled with EMD
Hongya Li, Yuzhong Peng, Chuyan Deng, Yonghua Pan, Daoqing Gong, Hao Zhang 0079
ICIC (1)6
2017 Causal Discovery Using Regression-Based Conditional Independence Tests
abstract
Conditional independence (CI) testing is an important tool in causal discovery. Generally, by using CI tests, a set of Markov equivalence classes w.r.t. the observed data can be estimated by checking whether each pair of variables x and y is d-separated, given a set of variables Z. Due to the curse of dimensionality, CI testing is often difficult to return a reliable result for high-dimensional Z. In this paper, we propose a regression-based CI test to relax the test of x ⊥ y|Z to simpler unconditional independence tests of x − f(Z) ⊥ y−g(Z), and x−f(Z) ⊥ Z or y−g(Z) ⊥ Z under the assumption that the data-generating procedure follows additive noise models (ANMs). When the ANM is identifiable, we prove that x − f(Z) ⊥ y − g(Z) ⇒ x ⊥ y|Z. We also show that 1) f and g can be easily estimated by regression, 2) our test is more powerful than the state-of-the-art kernel CI tests, and 3) existing causal learning algorithms can infer much more causal directions by using the proposed method.
Hao Zhang 0079, Shuigeng Zhou, Kun Zhang 0001, Jihong Guan
AAAI1
2016 Exploiting Twitter Moods to Boost Financial Trend Prediction Based on Deep Network Models
Yifu Huang, Kai Huang 0011, Yang Wang 0100, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou
ICIC (3)4