EDBT 2026 Demo / reviewers in the wild / expert
Yan Zeng 0002
dblp:83/4665-2
· DBLP profile ↗
17ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-7721-2560ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning by doing: an online causal reinforcement learning framework with causal-aware policy
Ruichu Cai, Siyang Huang, Jie Qiao, Wei Chen 0103, Yan Zeng 0002, Keli Zhang, Fuchun Sun 0001, Zhifeng Hao 0004 |
Sci. China Inf. Sci. | 5 |
| 2025 | Data-Driven Selection of Instrumental Variables for Additive Nonlinear, Constant Effects ModelsabstractWe consider the problem of selecting instrumental variables from observational data, a fundamental challenge in causal inference. Existing methods mostly focus on additive linear, constant effects models, limiting their applicability in complex real-world scenarios. In this paper, we tackle a more general and challenging setting: the additive non-linear, constant effects model. We first propose a novel testable condition, termed the Cross Auxiliary-based independent Test (CAT) condition, for selecting the valid IV set. We show that this condition is both necessary and sufficient for identifying valid instrumental variable sets within such a model under milder assumptions. Building on this condition, we develop a practical algorithm for selecting the set of valid instrumental variables. Extensive experiments on both synthetic and two real-world datasets demonstrate the effectiveness and robustness of our proposed approach, highlighting its potential for broader applications in causal analysis. Xichen Guo, Feng Xie 0002, Yan Zeng 0002, Hao Zhang 0079, Zhi Geng |
ICML | 3 |
| 2025 | Local Learning for Covariate Selection in Nonparametric Causal Effect Estimation with Latent VariablesabstractEstimating causal effects from nonexperimental data is a fundamental problem in many fields of science. A key component of this task is selecting an appropriate set of covariates for confounding adjustment to avoid bias. Most existing methods for covariate selection often assume the absence of latent variables and rely on learning the global causal structure among variables. However, identifying the global structure can be unnecessary and inefficient, especially when our primary interest lies in estimating the effect of a treatment variable on an outcome variable. To address this limitation, we propose a novel local learning approach for covariate selection in nonparametric causal effect estimation, which accounts for the presence of latent variables. Our approach leverages testable independence and dependence relationships among observed variables to identify a valid adjustment set for a target causal relationship, ensuring both soundness and completeness under standard assumptions. We validate the effectiveness of our algorithm through extensive experiments on both synthetic and real-world data. Xichen Guo, Feng Xie 0002, Yan Zeng 0002, Hao Zhang 0079, Zhi Geng |
NeurIPS | 4 |
| 2025 | Learning Counterfactual Outcomes Under Rank PreservationabstractCounterfactual inference aims to estimate the counterfactual outcome at the individual level given knowledge of an observed treatment and the factual outcome, with broad applications in fields such as epidemiology, econometrics, and management science. Previous methods rely on a known structural causal model (SCM) or assume the homogeneity of the exogenous variable and strict monotonicity between the outcome and exogenous variable. In this paper, we propose a principled approach for identifying and estimating the counterfactual outcome. We first introduce a simple and intuitive rank preservation assumption to identify the counterfactual outcome without relying on a known structural causal model. Building on this, we propose a novel ideal loss for theoretically unbiased learning of the counterfactual outcome and further develop a kernel-based estimator for its empirical estimation. Our theoretical analysis shows that the rank preservation assumption is not stronger than the homogeneity and strict monotonicity assumptions, and shows that the proposed ideal loss is convex, and the proposed estimator is unbiased. Extensive semi-synthetic and real-world experiments are conducted to demonstrate the effectiveness of the proposed method. Peng Wu 0012, Haoxuan Li 0001, Chunyuan Zheng 0001, Yan Zeng 0002, Jiawei Chen 0007, Yang Liu 0018, Ruocheng Guo, Kun Zhang 0001 |
NeurIPS | 4 |
| 2025 | A Survey on Causal Reinforcement LearningabstractWhile reinforcement learning (RL) achieves tremendous success in sequential decision-making problems of many domains, it still faces key challenges of data inefficiency and the lack of interpretability. Interestingly, many researchers have leveraged insights from the causality literature recently, bringing forth flourishing works to unify the merits of causality and address well the challenges from RL. As such, it is of great necessity and significance to collate these causal RL (CRL) works, offer a review of CRL methods, and investigate the potential functionality from causality toward RL. In particular, we divide the existing CRL approaches into two categories according to whether their causality-based information is given in advance or not. We further analyze each category in terms of the formalization of different models, ranging from the Markov decision process (MDP), partially observed MDP (POMDP), multiarmed bandits (MABs), imitation learning (IL), and dynamic treatment regime (DTR). Each of them represents a distinct type of causal graphical illustration. Moreover, we summarize the evaluation matrices and open sources, while we discuss emerging applications, along with promising prospects for the future development of CRL. Yan Zeng 0002, Ruichu Cai, Fuchun Sun 0001, Libo Huang 0001, Zhifeng Hao 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | eTag: Class-Incremental Learning via Embedding Distillation and Task-Oriented GenerationabstractClass incremental learning (CIL) aims to solve the notorious forgetting problem, which refers to the fact that once the network is updated on a new task, its performance on previously-learned tasks degenerates catastrophically. Most successful CIL methods store exemplars (samples of learned tasks) to train a feature extractor incrementally, or store prototypes (features of learned tasks) to estimate the incremental feature distribution. However, the stored exemplars would violate the data privacy concerns, while the fixed prototypes might not reasonably be consistent with the incremental feature distribution, hindering the exploration of real-world CIL applications. In this paper, we propose a data-free CIL method with embedding distillation and Task-oriented generation (eTag), which requires neither exemplar nor prototype. Embedding distillation prevents the feature extractor from forgetting by distilling the outputs from the networks' intermediate blocks. Task-oriented generation enables a lightweight generator to produce dynamic features, fitting the needs of the top incremental classifier. Experimental results confirm that the proposed eTag considerably outperforms state-of-the-art methods on several benchmark datasets. Libo Huang 0001, Yan Zeng 0002, Chuanguang Yang, Zhulin An, Boyu Diao, Yongjun Xu 0001 |
AAAI | 2 |
| 2024 | Policy Learning for Balancing Short-Term and Long-Term RewardsabstractEmpirical researchers and decision-makers spanning various domains frequently seek profound insights into the long-term impacts of interventions. While the significance of long-term outcomes is undeniable, an overemphasis on them may inadvertently overshadow short-term gains. Motivated by this, this paper formalizes a new framework for learning the optimal policy that effectively balances both long-term and short-term rewards, where some long-term outcomes are allowed to be missing. In particular, we first present the identifiability of both rewards under mild assumptions. Next, we deduce the semiparametric efficiency bounds, along with the consistency and asymptotic normality of their estimators. We also reveal that short-term outcomes, if associated, contribute to improving the estimator of the long-term reward. Based on the proposed estimators, we develop a principled policy learning approach and further derive the convergence rates of regret and estimation errors associated with the learned policy. Extensive experiments are conducted to validate the effectiveness of the proposed method, demonstrating its practical applicability. Peng Wu 0012, Ziyu Shen, Feng Xie 0002, Zhongyao Wang, Yan Zeng 0002 |
ICML | 6 |
| 2024 | ACE: Off-Policy Actor-Critic with Causality-Aware Entropy RegularizationabstractThe varying significance of distinct primitive behaviors during the policy learning process has been overlooked by prior model-free RL algorithms. Leveraging this insight, we explore the causal relationship between different action dimensions and rewards to evaluate the significance of various primitive behaviors during training. We introduce a causality-aware entropy term that effectively identifies and prioritizes actions with high potential impacts for efficient exploration. Furthermore, to prevent excessive focus on specific primitive behaviors, we analyze the gradient dormancy phenomenon and introduce a dormancy-guided reset mechanism to further enhance the efficacy of our method. Our proposed algorithm, ACE: Off-policy Actor-critic with Causality-aware Entropy regularization, demonstrates a substantial performance advantage across 29 diverse continuous control tasks spanning 7 domains compared to model-free RL baselines, which underscores the effectiveness, versatility, and efficient sample efficiency of our approach. Benchmark results and videos are available at https://ace-rl.github.io/. Tianying Ji, Yongyuan Liang, Yan Zeng 0002, Yu Luo 0021, Guowei Xu 0001, Ruijie Zheng, Furong Huang, Fuchun Sun 0001, Huazhe Xu |
ICML | 3 |
| 2024 | Local Causal Structure Learning in the Presence of Latent VariablesabstractDiscovering causal relationships from observational data, particularly in the presence of latent variables, poses a challenging problem. While current local structure learning methods have proven effective and efficient when the focus lies solely on the local relationships of a target variable, they operate under the assumption of causal sufficiency. This assumption implies that all the common causes of the measured variables are observed, leaving no room for latent variables. Such a premise can be easily violated in various real-world applications, resulting in inaccurate structures that may adversely impact downstream tasks. In light of this, our paper delves into the primary investigation of locally identifying potential parents and children of a target from observational data that may include latent variables. Specifically, we harness the causal information from m-separation and V-structures to derive theoretical consistency results, effectively bridging the gap between global and local structure learning. Together with the newly developed stop rules, we present a principled method for determining whether a variable is a direct cause or effect of a target. Further, we theoretically demonstrate the correctness of our approach under the standard causal Markov and faithfulness conditions, with infinite samples. Experimental results on both synthetic and real-world data validate the effectiveness and efficiency of our approach. Feng Xie 0002, Peng Wu 0012, Yan Zeng 0002, Zhi Geng |
ICML | 4 |
| 2024 | Identification and Estimation of the Bi-Directional MR with Some Invalid InstrumentsabstractWe consider the challenging problem of estimating causal effects from purely observational data in the bi-directional Mendelian randomization (MR), where some invalid instruments, as well as unmeasured confounding, usually exist.
To address this problem, most existing methods attempt to find proper valid instrumental variables (IVs) for the target causal effect by expert knowledge or by assuming that the causal model is a one-directional MR model.
As such, in this paper, we first theoretically investigate the identification of the bi-directional MR from observational data. In particular, we provide necessary and sufficient conditions under which valid IV sets are correctly identified such that the bi-directional MR model is identifiable, including the causal directions of a pair of phenotypes (i.e., the treatment and outcome).
Moreover, based on the identification theory, we develop a cluster fusion-like method to discover valid IV sets and estimate the causal effects of interest.
We theoretically demonstrate the correctness of the proposed algorithm.
Experimental results show the effectiveness of our method for estimating causal effects in both one-directional and bi-directional MR models. Feng Xie 0002, Yan Zeng 0002, Zhi Geng |
NeurIPS | 4 |
| 2024 | Learning the Optimal Policy for Balancing Short-Term and Long-Term RewardsabstractLearning the optimal policy to balance multiple short-term and long-term rewards has extensive applications across various domains. Yet, there is a noticeable scarcity of research addressing policy learning strategies in this context. In this paper, we aim to learn the optimal policy capable of effectively balancing multiple short-term and long-term rewards, especially in scenarios where the long-term outcomes are often missing due to data collection challenges over extended periods. Towards this goal, the conventional linear weighting method, which aggregates multiple rewards into a single surrogate reward through weighted summation, can only achieve sub-optimal policies when multiple rewards are related. Motivated by this, we propose a novel decomposition-based policy learning (DPPL) method that converts the whole problem into subproblems. The DPPL method is capable of obtaining optimal policies even when multiple rewards are interrelated. Nevertheless, the DPPL method requires a set of preference vectors specified in advance, posing challenges in practical applications where selecting suitable preferences is non-trivial. To mitigate this, we further theoretically transform the optimization problem in DPPL into an $\varepsilon$-constraint problem, where $\varepsilon$ represents the minimum acceptable levels of other rewards while maximizing one reward. This transformation provides intuitive into the selection of preference vectors. Extensive experiments are conducted on the proposed method and the results validate the effectiveness of the method. Qinwei Yang, Yan Zeng 0002, Ruocheng Guo, Yang Liu 0018, Peng Wu 0012 |
NeurIPS | 3 |
| 2023 | Causal discovery of 1-factor measurement models in linear latent variable models with arbitrary noise distributions
Feng Xie 0002, Yan Zeng 0002, Zhengming Chen 0002, Yangbo He, Zhi Geng, Kun Zhang 0001 |
Neurocomputing | 2 |
| 2023 | Nonlinear Causal Discovery for High-Dimensional Deterministic DataabstractNonlinear causal discovery with high-dimensional data where each variable is multidimensional plays a significant role in many scientific disciplines, such as social network analysis. Previous work majorly focuses on exploiting asymmetry in the causal and anticausal directions between two high-dimensional variables (a cause-effect pair). Although there exist some works that concentrate on the causal order identification between multiple variables, i.e., more than two high-dimensional variables, they do not validate the consistency of methods through theoretical analysis on multiple-variable data. In particular, based on the asymmetry for the cause-effect pair, if model assumptions for any pair of the data are violated, the asymmetry condition will not hold, resulting in the deduction of incorrect order identification. Thus, in this article, we propose a causal functional model, namely high-dimensional deterministic model (HDDM), to identify the causal orderings among multiple high-dimensional variables. We derive two candidates' selection rules to alleviate the inconvenient effects resulted from the violated-assumption pairs. The corresponding theoretical justification is provided as well. With these theoretical results, we develop a method to infer causal orderings for nonlinear multiple-variable data. Simulations on synthetic data and real-world data are conducted to verify the efficacy of our proposed method. Since we focus on deterministic relations in our method, we also verify the robustness of the noises in simulations. Yan Zeng 0002, Zhifeng Hao 0004, Ruichu Cai, Feng Xie 0002, Libo Huang 0001, Shohei Shimizu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Causal Discovery with Multi-Domain LiNGAM for Latent FactorsabstractDiscovering causal structures among latent factors from observed data is a particularly challenging problem. Despite some efforts for this problem, existing methods focus on the single-domain data only. In this paper, we propose Multi-Domain Linear Non-Gaussian Acyclic Models for LAtent Factors (MD-LiNA), where the causal structure among latent factors of interest is shared for all domains, and we provide its identification results. The model enriches the causal representation for multi-domain data. We propose an integrated two-phase algorithm to estimate the model. In particular, we first locate the latent factors and estimate the factor loading matrix. Then to uncover the causal structure among shared latent factors of interest, we derive a score function based on the characterization of independence relations between external influences and the dependence relations between multi-domain latent factors and latent factors of interest. We show that the proposed method provides locally consistent estimators. Experimental results on both synthetic and real-world data demonstrate the efficacy and robustness of our approach. Yan Zeng 0002, Shohei Shimizu, Ruichu Cai, Feng Xie 0002, Michio Yamamoto, Zhifeng Hao 0004 |
IJCAI | 1 |
| 2020 | Spike Sorting Based On Low-Rank And Sparse RepresentationabstractAs the first step to study the coding mechanism and synergistic behaviour of neurons, spike sorting plays an important role in the neurosciences research community. Despite many empirical successes in spike sorting models, there are still sufferings from the overlapping and noise corruption problems. To ease these situations, in this paper, we present an efficient and effective method with the help of optimization theory. Firstly, by introducing the low-rank strategy, the global structure underlying the spike data could be discovered. Secondly, by engaging the sparse coding to balance the noise, the proposed model is robust in the overlapping and noise spike sorting scenario. We have conducted experiments on the Wave-clus dataset compared with two state of the art models. The results verify the efficacy of our scheme and confirm the claims above. Libo Huang 0001, Bingo Wing-Kuen Ling, Yan Zeng 0002, Lu Gan 0002 |
ICME | 3 |
| 2020 | A causal discovery algorithm based on the prior selection of leaf nodes
Yan Zeng 0002, Zhifeng Hao 0004, Ruichu Cai, Feng Xie 0002, Liang Ou, Ruihui Huang |
Neural Networks | 1 |
| 2020 | An Efficient Entropy-Based Causal Discovery Method for Linear Structural Equation Models With IID Noise VariablesabstractThe discovery of causal relationships from the observational data is an important task. To identify the unique causal structure belonging to a Markov equivalence class, a number of algorithms, such as the linear non-Gaussian acyclic model (LiNGAM), have been proposed. However, two challenges remain to be met: 1) these algorithms fail to work on the data which follow linear structural equation model with Gaussian noise and 2) they misjudge the causal direction when the data contain additional measurement errors. In this paper, we propose an entropy-based two-phase iterative algorithm for arbitrary distribution data with additional measurement errors under some mild assumptions. In the first phase of the algorithm, based on the property that entropy can measure the amount of information behind the data with arbitrary distribution, we design a general approach for the identification of exogenous variable on both Gaussian and non-Gaussian data, and we give the corresponding theoretical derivation. In the second phase, to eliminate the effects of measurement errors, we revise the value of the exogenous variable by removing its measurement error and further use the revised value to remove its effect on the remaining variables. Experimental results on real-world causal structures are presented to demonstrate the effectiveness and stability of our method. We also apply the proposed algorithm on the mobile-base-station data with measurement errors, and the results further prove the effectiveness of our algorithm. Feng Xie 0002, Ruichu Cai, Yan Zeng 0002, Jiantao Gao, Zhifeng Hao 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |