Li Kheng Chai

dblp:350/0334 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0001-5371-4777ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 65% Efficient and distributed learning · 9% Learning theory · 9%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal effect estimation
2.332025
Progressive Generalization Risk Reduction for Data-Efficient Causal Effect Estimation · KDD (1) 2025
Variational Counterfactual Prediction Under Runtime Domain Corruption · IEEE Trans. Knowl. Data Eng. 2024
To Predict or to Reject: Causal Effect Estimation with Uncertainty on Networked Data · ICDM 2023
Machine learning › Probabilistic and Bayesian machine learning
causal inference
1.722025
Progressive Generalization Risk Reduction for Data-Efficient Causal Effect Estimation · KDD (1) 2025
Enhancing Treatment Effect Estimation via Active Learning: A Counterfactual Covering Perspective · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation
1.522025
Enhancing Treatment Effect Estimation via Active Learning: A Counterfactual Covering Perspective · ICML 2025
To Predict or to Reject: Causal Effect Estimation with Uncertainty on Networked Data · ICDM 2023
Machine learning › Efficient and distributed learning
active learning
0.912025
Enhancing Treatment Effect Estimation via Active Learning: A Counterfactual Covering Perspective · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference
counterfactual prediction
0.812024
Variational Counterfactual Prediction Under Runtime Domain Corruption · IEEE Trans. Knowl. Data Eng. 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
0.712023
To Predict or to Reject: Causal Effect Estimation with Uncertainty on Networked Data · ICDM 2023
Machine learning › Trustworthy machine learning
uncertainty estimation
0.712023
To Predict or to Reject: Causal Effect Estimation with Uncertainty on Networked Data · ICDM 2023
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.212024
Variational Counterfactual Prediction Under Runtime Domain Corruption · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Graph learning
graph neural network
0.212023
To Predict or to Reject: Causal Effect Estimation with Uncertainty on Networked Data · ICDM 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
networked observational data
0.212023
To Predict or to Reject: Causal Effect Estimation with Uncertainty on Networked Data · ICDM 2023

Methods — techniques the papers use, named apart from their topics

greedy radius reduction · 0.9empirical risk minimization · 0.9coverage maximization · 0.9variational inference · 0.8adversarial domain adaptation · 0.8lipschitz constraint · 0.7graph deep kernel learning · 0.7gaussian process · 0.7
YearPublicationVenuePosition
2025 Enhancing Treatment Effect Estimation via Active Learning: A Counterfactual Covering Perspective
abstract
Although numerous complex algorithms for treatment effect estimation have been developed in recent years, their effectiveness remains limited when handling insufficiently labeled training sets due to the high cost of labeling the post-treatment effect, e.g., the expensive tumor imaging or biopsy procedures needed to evaluate treatment effects. Therefore, it becomes essential to actively incorporate more high-quality labeled data, all while adhering to a constrained labeling budget. To enable data-efficient treatment effect estimation, we formalize the problem through rigorous theoretical analysis within the active learning context, where the derived key measures -- factual and counterfactual covering radii determine the risk upper bound. To reduce the bound, we propose a greedy radius reduction algorithm, which excels under an idealized, balanced data distribution. To generalize to more realistic data distributions, we further propose FCCM, which transforms the optimization objective into the Factual and Counterfactual Coverage Maximization to ensure effective radius reduction during data acquisition. Furthermore, benchmarking FCCM against other baselines demonstrates its superiority across both fully synthetic and semi-synthetic datasets. Code: https://github.com/uqhwen2/FCCM.
Hechuan Wen, Tong Chen 0005, Mingming Gong, Li Kheng Chai, Shazia Sadiq, Hongzhi Yin
ICML4
2025 Progressive Generalization Risk Reduction for Data-Efficient Causal Effect Estimation
abstract
Causal effect estimation (CEE) provides a crucial tool for predicting the unobserved counterfactual outcome for an entity. As CEE relaxes the requirement for "perfect'' counterfactual samples (e.g., patients with identical attributes and only differ in treatments received) that are impractical to obtain and can instead operate on observational data, it is usually used in high-stake domains like medical treatment effect prediction. Nevertheless, in those high-stake domains, gathering a decently sized, fully labelled observational dataset remains challenging due to hurdles associated with costs, ethics, expertise and time needed, etc., of which medical treatment surveys are a typical example. Consequently, if the training dataset is small in scale, low generalization risks can hardly be achieved on any CEE algorithms.
Hechuan Wen, Tong Chen 0005, Guanhua Ye, Li Kheng Chai, Shazia Sadiq, Hongzhi Yin
KDD (1)4
2024 Variational Counterfactual Prediction Under Runtime Domain Corruption
abstract
To date, various neural methods have been proposed for causal effect estimation based on observational data, where a default assumption is the same distribution and availability of variables at both training and inference (i.e., runtime) stages. However, distribution shift (i.e., domain shift) could happen during runtime, and bigger challenges arise from the impaired accessibility of variables. This is commonly caused by increasing privacy and ethical concerns, which can make arbitrary variables unavailable in the entire runtime data and imputation impractical. We term the co-occurrence of domain shift and inaccessible variablesruntime domain corruption, which seriously impairs the generalizability of a trained counterfactual predictor. To counter runtime domain corruption, we subsume counterfactual prediction under the notion of domain adaptation. Specifically, we upper-bound the error w.r.t. the target domain (i.e., runtime covariates) by the sum of source domain error and inter-domain distribution distance. In addition, we build an adversarially unified variational causal effect model, named VEGAN, with a novel two-stage adversarial domain adaptation scheme to reduce the latent distribution disparity between treated and control groups first, and between training and runtime variables afterwards. We demonstrate that VEGAN outperforms other state-of-the-art baselines on individual-level treatment effect estimation in the presence of runtime domain corruption on benchmark datasets.
Hechuan Wen, Tong Chen 0005, Li Kheng Chai, Shazia Sadiq, Junbin Gao, Hongzhi Yin
IEEE Trans. Knowl. Data Eng.3
2023 To Predict or to Reject: Causal Effect Estimation with Uncertainty on Networked Data
abstract
Due to the imbalanced nature of networked observational data, the causal effect predictions for some individuals can severely violate the positivity/overlap assumption, rendering unreliable estimations. Nevertheless, this potential risk of individual-level treatment effect estimation on networked data has been largely under-explored. To create a more trustworthy causal effect estimator, we propose the uncertainty-aware graph deep kernel learning (GraphDKL) framework with Lipschitz constraint to model the prediction uncertainty with Gaussian process and identify unreliable estimations. To the best of our knowledge, GraphDKL is the first framework to tackle the violation of positivity assumption when performing causal effect estimation with graphs. With extensive experiments, we demonstrate the superiority of our proposed method in uncertainty-aware causal effect estimation on networked data. The code of GraphDKL is available at https://github.com/uqhwen2/GraphDKL.
Hechuan Wen, Tong Chen 0005, Li Kheng Chai, Shazia Sadiq, Kai Zheng 0001, Hongzhi Yin
ICDM3