Yewei Xia

dblp:339/8598 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0001-5515-5913ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Invariant Feature Learning for Counterfactual Watch-time Prediction in Video Recommendation
abstract
Video recommendation systems heavily rely on user watch time feedback, making accurate watch time prediction a crucial task. However, this task inherently suffers from bias, as recommendation models tend to favor long-duration videos to maximize watch time. This issue, known as duration bias in the watch-time prediction context, can be explained from a causal perspective, where video duration acts as a confounder. Recent works address this bias using backdoor adjustment, isolating the direct effect of content on watch time from observational data. These methods typically discretize video duration into groups, estimate group-wise effects, and then aggregate them via a unified prediction model. However, this aggregation strategy is prone to model misspecification due to feature distribution shift across groups. In this paper, we reinterpret the problem through the lens of invariant learning and propose a novel framework: Duration-Invariant Feature Learning (DIFL). DIFL employs a kernel-based regularization that enforces representation invariance across duration groups, reducing sensitivity to group design and improving generalization. This enables more accurate modeling of the direct causal effect and making counterfactual inference. Extensive experiments on both public and real large-scale production datasets demonstrate the effectiveness of our approach, which achieves SOTA performance.
Chenghou Jin, Yixin Ren, Hongxu Ma 0001, Yewei Xia, Yi Guan, Hao Zhang 0079, Jiandong Ding, Jihong Guan, Shuigeng Zhou
AAAI4
2026 Causal discovery by continuous optimization with weighted superstructure
Yewei Xia, Hao Zhang 0079, Ruxin Wang 0001, Yuzhong Peng, Jihong Guan, Shuigeng Zhou
Neural Networks2
2026 Causal Discovery by Multi-Level Wavelet Mapping Correlation Based Statistical Dependence Measurement
abstract
This article proposes a new method for causal discovery based on a novel dependence measurement criterion, namely, Multi-level Wavelet Mapping Correlation (MWMC). MWMC captures nonlinear dependencies between variables by measuring their correlations across multiple levels of wavelet mappings. From a theoretical perspective, we show that the empirical estimate of MWMC converges exponentially fast to its population quantity. Under the null hypothesis of independence, we further design a permutation-based independence testing procedure, termed the Wavelet Independence Test (WIT), built upon MWMC. We prove that WIT not only effectively controls the Type I error rate (false positives), but also guarantees that the Type II error rate (false negatives) is upper bounded by \(\mathcal{O}(n^{-1})\) , where \( n \) denotes the sample size, even with a finite number of permutations. Building on these theoretical guarantees, we derive a causal discovery method by integrating MWMC-based WIT into standard causal discovery pipelines. Extensive experiments on (conditional) independence testing and causal discovery using both synthetic and real-world datasets with varying sample sizes demonstrate that our approach consistently outperforms existing independence testing and causal discovery methods in terms of reduced Type II error rates and statistically validated performance improvements. Impact Statement —Causal discovery is a fundamental task in knowledge discovery, aiming to uncover the underlying data-generating mechanisms in order to support more accurate and interpretable predictions. Statistical independence tests and conditional independence (CI) tests have long served as core tools in this area. To improve the reliability of independence testing, we propose a novel test, WIT, which achieves lower Type II error rates in 19 out of 25 distinct experimental scenarios involving diverse data distributions, compared to 15 out of 25 for the strongest existing baseline. We further apply WIT to CI testing and causal discovery, and extensive empirical results show that it consistently improves the performance of multiple causal discovery algorithms across a range of experimental settings.
Yixin Ren, Hao Zhang 0079, Yewei Xia, Feng Xie 0002, Jihong Guan, Shuigeng Zhou
ACM Trans. Knowl. Discov. Data3
2025 A New Model for Prototype-based Continual Learning in Hyperspherical Space
abstract
The continuous emergence of new objects in the visual world poses a serious challenge to deep object recognition methods, which sparks the increasing study on continual or incremental learning. However, learning new tasks faces the tough catastrophic forgetting problem, i.e., dramatic performance degradation on old tasks. A good continual learning model should be robustly adapted to the upcoming tasks while effectively handling catastrophic forgetting. In this paper, we focus on the class-incremental learning (CIL) task, and propose a novel prototype-based continual learning model C-HPN that projects the visual features into a hypersphere geometric space, where continual learning is conducted. C-HPN features two-fold contributions. On the one hand, instead of using the popular cross-entropy loss, we develop an instance-prototype compact loss to obtain well-clustered hyperspherical embeddings and a prototype-prototype separability loss to boost the model’s generalization by introducing large angle distance inductive bias between prototypes in the hyperspherical space. On the other hand, prototype construction and adaptation strategies are designed for effectively adapting new classes, and an instance-prototype relationship preservation distillation mechanism is introduced to overcome catastrophic forgetting. Extensive experiments on several image datasets validate the effectiveness of the proposed method.
Yixin Ren, Yewei Xia, Longtao Huang, Hui Xue 0001, Shuigeng Zhou
ICASSP2
2025 Identification of Latent Confounders via Investigating the Tensor Ranks of the Nonlinear Observations
abstract
We study the problem of learning discrete latent variable causal structures from mixed-type observational data. Traditional methods, such as those based on the tensor rank condition, are designed to identify discrete latent structure models and provide robust identification bounds for discrete causal models. However, when observed variables—specifically, those representing the children of latent variables—are collected at various levels with continuous data types, the tensor rank condition is not applicable, limiting further causal structure learning for latent variables. In this paper, we consider a more general case where observed variables can be either continuous or discrete, and further allow for scenarios where multiple latent parents cause the same set of observed variables. We show that, under the completeness condition, it is possible to discretize the data in a way that satisfies the full-rank assumption required by the tensor rank condition. This enables the identifiability of discrete latent structure models within mixed-type observational data. Moreover, we introduce the two-sufficient measurement condition, a more general structural assumption under which the tensor rank condition holds and the underlying latent causal structure is identifiable by a proposed two-stage identification algorithm. Extensive experiments on both simulated and real-world data validate the effectiveness of our method.
Zhengming Chen 0002, Yewei Xia, Feng Xie 0002, Jie Qiao, Zhifeng Hao 0004, Ruichu Cai, Kun Zhang 0001
ICML2
2025 Extracting Rare Dependence Patterns via Adaptive Sample Reweighting
abstract
Discovering dependence patterns between variables from observational data is a fundamental issue in data analysis. However, existing testing methods often fail to detect subtle yet critical patterns that occur within small regions of the data distribution--patterns we term rare dependence. These rare dependencies obscure the true underlying dependence structure in variables, particularly in causal discovery tasks. To address this issue, we propose a novel testing method that combines kernel-based (conditional) independence testing with adaptive sample importance reweighting. By learning and assigning higher importance weights to data points exhibiting significant dependence, our method amplifies the patterns and can detect them successfully. Theoretically, we analyze the asymptotic distributions of the statistics in this method and show the uniform bound of the learning scheme. Furthermore, we integrate our tests into the PC algorithm, a constraint-based approach for causal discovery, equipping it to uncover causal relationships even in the presence of rare dependence. Empirical evaluation of synthetic and real-world datasets comprehensively demonstrates the efficacy of our method.
Yewei Xia, Zhengming Chen 0002, Liuhua Peng, Mingming Gong, Kun Zhang 0001
ICML2
2025 Identifying Causal Mechanism Shifts Under Additive Models with Arbitrary Noise
abstract
In many real-world scenarios, the goal is to identify variables whose causal mechanisms change across related datasets. For example, detecting abnormal root nodes in manufacturing, and identifying key genes that influence cancer by analyzing differences in gene regulatory mechanisms between healthy individuals and cancer patients. This can be done by recovering the causal structure for each dataset independently and then comparing them to identify differences, but the performance is often suboptimal. Typically, existing methods directly identify causal mechanism shifts based on linear additive noise models (ANMs) or by imposing restrictive assumptions on the noise distribution. In this paper, we introduce CMSI, a novel and more general algorithm based on nonlinear ANMs that identifies variables with shifting causal mechanisms under arbitrary noise distributions. Evaluated on various synthetic datasets, CMSI consistently outperforms existing baselines in terms of F1 score. Additionally, we demonstrate CMSI's applicability on gene expression datasets of ovarian cancer patients at different disease stages.
Yewei Xia, Xueliang Cui, Hao Zhang 0079, Yixin Ren, Feng Xie 0002, Jihong Guan, Ruxin Wang 0001, Shuigeng Zhou
IJCAI1
2025 Efficient Constraint-based Window Causal Graph Discovery in Time Series with Multiple Time Lags
abstract
We address the identification of direct causes in time series with multiple time lags, and propose a constraint-based window causal graph discovery method. A key advantage of our method is that the number of required conditional independence (CI) tests scales quadratically with the number of sub-series. The method first uses CI tests to find the minimum trek lag between two arbitrary sub-series, followed by designing an efficient CI testing strategy to identify the direct causes between them. We show that the method is both sound and complete under some graph constraints. We compare the proposed method with typical baselines on various datasets. Experimental results show that our method outperforms all the counterparts in both accuracy and running speed.
Yewei Xia, Yixin Ren, Hong Cheng 0001, Hao Zhang 0079, Jihong Guan, Minchuan Xu, Shuigeng Zhou
IJCAI1
2025 Score-based Generative Modeling for Conditional Independence Testing
abstract
Determining conditional independence (CI) relationships between random variables is a fundamental yet challenging task in machine learning and statistics, especially in high-dimensional settings. Existing generative model-based CI testing methods, such as those utilizing generative adversarial networks (GANs), often struggle with undesirable modeling of conditional distributions and training instability, resulting in subpar performance. To address these issues, we propose a novel CI testing method via score-based generative modeling, which achieves precise Type I error control and strong testing power. Concretely, we first employ a sliced conditional score matching scheme to accurately estimate conditional score and use Langevin dynamics conditional sampling to generate null hypothesis samples, ensuring precise Type I error control. Then, we incorporate a goodness-of-fit stage into the method to verify generated samples and enhance interpretability in practice. We theoretically establish the error bound of conditional distributions modeled by score-based generative models and prove the validity of our CI tests. Extensive experiments on both synthetic and real-world datasets show that our method significantly outperforms existing state-of-the-art methods, providing a promising way to revitalize generative model-based CI testing.
Yixin Ren, Chenghou Jin, Yewei Xia, Longtao Huang, Hui Xue 0001, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou
KDD (2)3
2025 Fast Causal Discovery by Approximate Kernel-based Generalized Score Functions with Linear Computational Complexity
Yixin Ren, Haocheng Zhang, Yewei Xia, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou
KDD (1)3
2025 Regression-based conditional independence test with adaptive kernels
Yixin Ren, Juncai Zhang, Yewei Xia, Ruxin Wang 0001, Feng Xie 0002, Jihong Guan, Hao Zhang 0079, Shuigeng Zhou
Artif. Intell.3
2024 Learning Adaptive Kernels for Statistical Independence Tests
abstract
We propose a novel framework for kernel-based statistical independence tests that enable adaptatively learning parameterized kernels to maximize test power. Our framework can effectively address the pitfall inherent in the existing signal-to-noise ratio criterion by modeling the change of the null distribution during the learning process. Based on the proposed framework, we design a new class of kernels that can adaptatively focus on the significant dimensions of variables to judge independence, which makes the tests more flexible than using simple kernels that are adaptive only in length-scale, and especially suitable for high-dimensional complex data. Theoretically, we demonstrate the consistency of our independence tests, and show that the non-convex objective function used for learning fits the L-smoothing condition, thus benefiting the optimization. Experimental results on both synthetic and real data show the superiority of our method. The source code and datasets are available at \url{https://github.com/renyixin666/HSIC-LK.git}.
Yixin Ren, Yewei Xia, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou
AISTATS2
2024 Efficiently Learning Significant Fourier Feature Pairs for Statistical Independence Testing
abstract
We propose a novel method to efficiently learn significant Fourier feature pairs for maximizing the power of Hilbert-Schmidt Independence Criterion~(HSIC) based independence tests. We first reinterpret HSIC in the frequency domain, which reveals its limited discriminative power due to the inability to adapt to specific frequency-domain features under the current inflexible configuration. To remedy this shortcoming, we introduce a module of learnable Fourier features, thereby developing a new criterion. We then derive a finite sample estimate of the test power by modeling the behavior of the criterion, thus formulating an optimization objective for significant Fourier feature pairs learning. We show that this optimization objective can be computed in linear time (with respect to the sample size $n$), which ensures fast independence tests. We also prove the convergence property of the optimization objective and establish the consistency of the independence tests. Extensive empirical evaluation on both synthetic and real datasets validates our method's superiority in effectiveness and efficiency, particularly in handling high-dimensional data and dealing with large-scale scenarios.
Yixin Ren, Yewei Xia, Hao Zhang 0079, Jihong Guan, Shuigeng Zhou
NeurIPS2
2024 Towards Effective Causal Partitioning by Edge Cutting of Adjoint Graph
abstract
Causal partitioningis an effective approach for causal discovery based on the divide-and-conquer strategy. Up to now, various heuristic methods based on conditional independence (CI) tests have been proposed for causal partitioning. However, most of these methods fail to achieve satisfactory partitioning without violating$d$-separation, leading to poor inference performance. In this work, we transform causal partitioning into an alternative problem that can be more easily solved. Concretely, we first construct a superstructure$G$of the true causal graph$G_{\mathcal {T}}$by performing a set of low-order CI tests on the observed data$D$. Then, we leverage point-line duality to obtain a graph$G_\mathcal {A}$adjoint to$G$. We show that the solution ofminimizing edge-cut ratioon$G_\mathcal {A}$can lead to a valid causal partitioning withsmaller causal-cut ratioon$G$andwithout violating$d$d-separation. We design an efficient algorithm to solve this problem. Extensive experiments show that the proposed method can achieve significantly better causal partitioning without violating$d$-separation than the existing methods.
Hao Zhang 0079, Yixin Ren, Yewei Xia, Shuigeng Zhou, Jihong Guan
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Hybrid Causal Feature Selection for Cancer Biomarker Identification From RNA-Seq Data
abstract
The discovery of cancer biomarkers helps to advance medical diagnosis and plays an important role in biomedical applications. Most of the existing data-driven methods identify biomarkers by ranking-based strategies, which generally return a subset or superset of the actual biomarkers, while some other causal-wise feature selection methods are based on Markov Blanket (MB) learning, facing the challenges of high-dimensionality & low-sample. In this work, we propose a novel hybrid causal feature selection method (called CAFES) to support large-scale cancer biomarker discovery from real RNA-seq data. Concretely, CAFES first uses minimal-redundancy & maximal-relevance strategy for dimensionality reduction that returns a set of candidate features. CAFES then learns the causal skeleton w.r.t. those features by CI tests and further obtains an appropriate superset of the MB of the target variable. Finally, CAFES learns the causal structure of this superset by the DAG-GNN algorithm and then obtains the MB of the target variable, which can be treated as the cancer biomarkers. We conduct experiments to evaluate the proposed method on two real well-known RNA-seq datasets that covering both binary and multi-class cases. We compare our method CAFES with seven recent methods including Semi-HITON-MB, STMB, BAMB, FBED, LCS-FS, EEMB, and EAMB. The results show that CAFES can identify dozens of cancer biomarkers, and of the discovered biomarkers can be verified by existing works that they are really directly related to the corresponding disease. An advantage of CAFES is that its Recall is significantly higher than those of all the counterparts, indicating that the continuous optimization (DAG-GNN) with the returned causal skeleton after feature selection (that can be treated as a conditional independence-based constraint to the optimization problem) is effective in cancer biomarkers identification under high-dimensional and low-sample RNA-seq data.
Wenwei Xu, Hao Zhang 0079, Yewei Xia, Yixin Ren, Jihong Guan, Shuigeng Zhou
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 Differentially Private Nonlinear Causal Discovery from Numerical Data
abstract
Recently, several methods such as private ANM, EM-PC and Priv-PC have been proposed to perform differentially private causal discovery in various scenarios including bivariate, multivariate Gaussian and categorical cases. However, there is little effort on how to conduct private nonlinear causal discovery from numerical data. This work tries to challenge this problem. To this end, we propose a method to infer nonlinear causal relations from observed numerical data by using regression-based conditional independence test (RCIT) that consists of kernel ridge regression (KRR) and Hilbert-Schmidt independence criterion (HSIC) with permutation approximation. Sensitivity analysis for RCIT is given and a private constraint-based causal discovery framework with differential privacy guarantee is developed. Extensive simulations and real-world experiments for both conditional independence test and causal discovery are conducted, which show that our method is effective in handling nonlinear numerical cases and easy to implement. The source code of our method and data are available at https://github.com/Causality-Inference/PCD.
Hao Zhang 0079, Yewei Xia, Yixin Ren, Jihong Guan, Shuigeng Zhou
AAAI2
2023 Multi-Level Wavelet Mapping Correlation for Statistical Dependence Measurement: Methodology and Performance
abstract
We propose a new criterion for measuring dependence between two real variables, namely, Multi-level Wavelet Mapping Correlation (MWMC). MWMC can capture the nonlinear dependencies between variables by measuring their correlation under different levels of wavelet mappings. We show that the empirical estimate of MWMC converges exponentially to its population quantity. To support independence test better with MWMC, we further design a permutation test based on MWMC and prove that our test can not only control the type I error rate (the rate of false positives) well but also ensure that the type II error rate (the rate of false negatives) is upper bounded by O(1/n) (n is the sample size) with finite permutations. By extensive experiments on (conditional) independence tests and causal discovery, we show that our method outperforms existing independence test methods.
Yixin Ren, Hao Zhang 0079, Yewei Xia, Jihong Guan, Shuigeng Zhou
AAAI3
2023 Causal Discovery by Continuous Optimization with Conditional Independence Constraint: Methodology and Performance
abstract
Discovering causal relationships from observational data is a challenging topic in artificial intelligence. Recent works formulate causal discovery as a continuous optimization problem with a differentiable acyclic constraint. Although these methods have achieved considerable performance improvement, they have two drawbacks: 1) they require a relatively large number of training samples; and 2) their performance will substantially deteriorate when facing heterogeneous noise. To address these problems, we first propose a low-order conditional independence (CI) constraint for the continuous optimization problem, and then design a soft version of the constraint by transforming it to a regularization term in the loss function of the continuous optimization problem. We show the convergence of continuous optimization with our constraint under some mild conditions, and the consistency of causal structure learning with the CI regularization. Extensive experiments on both synthetic and real-world datasets show that with our CI constraint or regularization, existing continuous optimization methods can achieve considerable performance improvement of causal discovery, especially when sample size is small.
Yewei Xia, Hao Zhang 0079, Yixin Ren, Jihong Guan, Shuigeng Zhou
ICDM1
2023 Causal Gene Identification Using Non-Linear Regression-Based Independence Tests
abstract
With the development of biomedical techniques in the past decades, causal gene identification has become one of the most promising applications in human genome-based business, which can help doctors to evaluate the risk of certain genetic diseases and provide further treatment recommendations for potential patients. When no controlled experiments can be applied, machine learning techniques like causal inference-based methods are generally used to identify causal genes. Unfortunately, most of the existing methods detect disease-related genes by ranking-based strategies or feature selection techniques, which generally return a superset of the corresponding real causal genes. There are also some causal inference-based methods that can identify a part of real causal genes from those supersets, but they are just able to return a few causal genes. This is contrary to our knowledge, as many results from controlled experiments have demonstrated that a certain disease, especially cancer, is usually related to dozens or hundreds of genes. In this work, we present an effective approach for identifying causal genes from gene expression data by using a new search strategy based on non-linear regression-based independence tests, which is able to greatly reduce the search space, and simultaneously establish the causal relationships from the candidate genes to the disease variable. Extensive experiments on real-world cancer datasets show that our method is superior to the existing causal inference-based methods in three aspects: 1) our method can identify dozens of causal genes, and 1/3 ∼ 1/2 of the discovered causal genes can be verified by existing works that they are really directly related to the corresponding disease; 2) The discovered causal genes are able to distinguish the status or disease subtype of the target patient; 3) Most of the discovered causal genes are closely relevant to the disease variable.
Hao Zhang 0079, Chuanxu Yan, Yewei Xia, Jihong Guan, Shuigeng Zhou
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 Conditional Independence Test Based on Residual Similarity
abstract
Recently, many regression-based conditional independence (CI) test methods have been proposed to solve the problem of causal discovery. These methods provide alternatives to test CI of x,y given Z by first removing the information of the controlling set Z from x and y , and then testing the independence between the two residuals R x,Z and R y,Z . When the residuals are linearly uncorrelated, the independence test between them is nontrivial. With the ability to calculate inner product in high-dimensional space, kernel-based methods are usually used to achieve this goal, but they are considerably time-consuming. In this paper, we test the independence between two linear combinations under linear structural equation model. We show that the dependence between the two residuals can be captured by the difference between the similarity of R x,Z and R y,Z and that of R x,Z and R r ( R r is an independent copy of R y,Z ) in high-dimensional space. With this result, we provide a new way to test CI based on the similarity between residuals, which is called SCIT — the abbreviation of Similarity-based CI Testing. Furthermore, we develop two versions of the proposal, called Kernel-SCIT and Neural-SCIT, respectively. Kernel-SCIT calculates the similarity by using kernel functions, while Neural-SCIT approximates the upper bound of the similarity by using deep neural networks. In both algorithms, random permutation tests are performed to control Type I error rate. The proposed tests are evaluated on (conditional) independence test and causal discovery with both synthetic and real datasets. Experimental results show that Kernel-SCIT is simpler yet more efficient and effective than the typical existing kernel-based methods HSIC and KCIT in the cases of small sample size, and Neural-SCIT can significantly boost the performance of CI testing when sufficient samples are available. The source code is available at https://github.com/xyw5vplus1/SCIT .
Hao Zhang 0079, Yewei Xia, Kun Zhang 0001, Shuigeng Zhou, Jihong Guan
ACM Trans. Knowl. Discov. Data2
2023 Self-Supervised Learning for Multimedia Recommendation
abstract
Learning representations for multimedia content is critical for multimedia recommendation. Current representation learning methods roughly fall into two groups: (1) using the historical interactions to create ID embeddings of users and items, and (2) treating multi-modal data as the side information of items to enrich their ID embeddings. Each user-item interaction offers the supervisory signal to optimize the representation learning by the traditional supervised learning paradigm. Due to the overlook of the multi-modal patterns ($e.g.$, co-occurrence of visual, acoustic, textual features in micro-videos a user saw before, and her behavioral features) hidden in the data, these methods are insufficient to create powerful representations and obtain satisfactory recommendation accuracy. To capture multi-modal patterns in the data itself, we go beyond the supervised learning paradigm, and incorporate the idea of self-supervised learning (SSL) into multimedia recommendation. Specifically, SSL consists of two components: (1) data augmentation upon multi-modal contents, where we design three operators — feature dropout (FD), feature masking (FM), feature fine and coarse spaces (FAC) — to generate multiple views of individual items; and (2) contrastive learning, which differentiates the views of an item from the others’ to distill additional supervisory signals. Clearly, SSL enables us to explore and exhibit the underlying relations among modalities, thereby resulting in powerful representations. We denote the generic framework by Self-supervised Learning-guided Multimedia Recommendation (SLMRec). Extensive experiments are performed on three real-world datasets, showing that SLMRec achieves significant improvements over several state-of-the-art baselines like LightGCN [1], MMGCN [2]. Further analysis shows how SSL affects recommendation performance.
Zhulin Tao, Xiaohao Liu, Yewei Xia, Xiang Wang 0010, Lifang Yang, Xianglin Huang, Tat-Seng Chua
IEEE Trans. Multim.3