Yuewen Sun

dblp:219/9893 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CADRE: Contour-Guided ADMM-Optimized Deep Radon Enhancement for Unsupervised CT Reconstruction from Incomplete Data
abstract
In clinical practice, reducing radiation dose in Computed Tomography (CT) is of paramount importance. However, low-dose protocols, such as sparse-view or limited-angle scanning, result in incomplete projection data, which severely degrades reconstructed image quality and can compromise diagnostic accuracy. While supervised deep learning methods have shown promise, their heavy reliance on large-scale paired training datasets limits their clinical applicability and generalizability. Unsupervised methods like Deep Radon Prior (DRP) circumvent the need for training data but often suffer from convergence instability and fail to preserve fine anatomical details. To address these limitations, this paper introduces CADRE (Contourguided ADMM-optimized Deep Radon Enhancement), a novel unsupervised framework that enhances DRP through a synergistic integration of three innovations: (1) an ADMM-inspired optimization strategy to stabilize convergence and regularize the solution space; (2) a contour-guided mechanism to enforce structural integrity using explicit geometric priors; and (3) an attention-enhanced network architecture with CBAM to improve feature extraction. Extensive experiments on medical datasets demonstrate that CADRE achieves state-of-the-art reconstruction quality from incomplete data. Its robustness is further validated across industrial and cultural heritage applications, offering a powerful and generalizable solution for high-fidelity imaging without paired training data.
Jintao Fu, Tianchen Zeng, Yuewen Sun
BIBM3
2025 Learning Hidden Causal Factors from Psychometrics Data Using Distributional Information
Roberto Legaspi, Xinshuai Dong, Donghuo Zeng, Yuewen Sun, Kazushi Ikeda, Peter Spirtes, Kun Zhang 0001
CogSci4
2025 Causal Representation Learning from Multimodal Biomedical Observations
abstract
Prevalent in biomedical applications (e.g., human phenotype research), multimodal datasets can provide valuable insights into the underlying physiological mechanisms. However, current machine learning (ML) models designed to analyze these datasets often lack interpretability and identifiability guarantees, which are essential for biomedical research. Recent advances in causal representation learning have shown promise in identifying interpretable latent causal variables with formal theoretical guarantees. Unfortunately, most current work on multimodal distributions either relies on restrictive parametric assumptions or yields only coarse identification results, limiting their applicability to biomedical research that favors a detailed understanding of the mechanisms. In this work, we aim to develop flexible identification conditions for multimodal data and principled methods to facilitate the understanding of biomedical datasets. Theoretically, we consider a nonparametric latent distribution (c.f., parametric assumptions in previous work) that allows for causal relationships across potentially different modalities. We establish identifiability guarantees for each latent component, extending the subspace identification results from previous work. Our key theoretical contribution is the structural sparsity of causal connections between modalities, which, as we will discuss, is natural for a large collection of biomedical systems. Empirically, we present a practical framework to instantiate our theoretical insights. We demonstrate the effectiveness of our approach through extensive experiments on both numerical and synthetic datasets. Results on a real-world human phenotype dataset are consistent with established biomedical research, validating our theoretical and methodological framework.
Yuewen Sun, Guangyi Chen 0002, Loka Li, Gongxu Luo, Zijian Li 0001, Yixuan Zhang 0001, Yujia Zheng 0001, Mengyue Yang, Petar Stojanov, Eran Segal, Eric P. Xing, Kun Zhang 0001
ICLR1
2025 Towards Identifiability of Hierarchical Temporal Causal Representation Learning
abstract
Modeling hierarchical latent dynamics behind time series data is critical for capturing temporal dependencies across multiple levels of abstraction in real-world tasks. However, existing temporal causal representation learning methods fail to capture such dynamics, as they fail to recover the joint distribution of hierarchical latent variables from \textit{single-timestep observed variables}. Interestingly, we find that the joint distribution of hierarchical latent variables can be uniquely determined using three conditionally independent observations. Building on this insight, we propose a Causally Hierarchical Latent Dynamic (CHiLD) identification framework. Our approach first employs temporal contextual observed variables to identify the joint distribution of multi-layer latent variables. Sequentially, we exploit the natural sparsity of the hierarchical structure among latent variables to identify latent variables within each layer. Guided by the theoretical results, we develop a time series generative model grounded in variational inference. This model incorporates a contextual encoder to reconstruct multi-layer latent variables and normalize flow-based hierarchical prior networks to impose the independent noise condition of hierarchical latent dynamics. Empirical evaluations on both synthetic and real-world datasets validate our theoretical claims and demonstrate the effectiveness of CHiLD in modeling hierarchical latent dynamics.
Zijian Li 0001, Minghao Fu 0002, Junxian Huang 0002, Yifan Shen 0004, Ruichu Cai, Yuewen Sun, Guangyi Chen 0002, Kun Zhang 0001
NeurIPS6
2025 Generative Framework for Personalized Persuasion: Inferring Causal, Counterfactual, and Latent Knowledge
abstract
We hypothesize that optimal system responses emerge from adaptive strategies grounded in causal and counterfactual knowledge.Counterfactual inference allows us to create hypothetical scenarios to examine the effects of alternative system responses.We enhance this process through causal discovery, which identifies the strategies informed by the underlying causal structure that govern system behaviors.Moreover, we consider the psychological constructs and unobservable noises that might be influencing user-system interactions as latent factors.We show that these factors can be effectively estimated.We employ causal discovery to identify strategy-level causal relationships among user and system utterances, guiding the generation of personalized counterfactual dialogues.We model the user utterance strategies as causal factors, enabling system strategies to be treated as counterfactual actions.Furthermore, we optimize policies for selecting system responses based on counterfactual data.Our results using a real-world dataset on social good demonstrate significant improvements in persuasive system outcomes, with increased cumulative rewards validating the efficacy of causal discovery in guiding personalized counterfactual inference and optimizing dialogue policies for a persuasive dialogue system.
Donghuo Zeng, Roberto Legaspi, Yuewen Sun, Xinshuai Dong, Kazushi Ikeda, Peter Spirtes, Kun Zhang 0001
UMAP3
2024 ACAMDA: Improving Data Efficiency in Reinforcement Learning through Guided Counterfactual Data Augmentation
abstract
Data augmentation plays a crucial role in improving the data efficiency of reinforcement learning (RL). However, the generation of high-quality augmented data remains a significant challenge. To overcome this, we introduce ACAMDA (Adversarial Causal Modeling for Data Augmentation), a novel framework that integrates two causality-based tasks: causal structure recovery and counterfactual estimation. The unique aspect of ACAMDA lies in its ability to recover temporal causal relationships from limited non-expert datasets. The identification of the sequential cause-and-effect allows the creation of realistic yet unobserved scenarios. We utilize this characteristic to generate guided counterfactual datasets, which, in turn, substantially reduces the need for extensive data collection. By simulating various state-action pairs under hypothetical actions, ACAMDA enriches the training dataset for diverse and heterogeneous conditions. Our experimental evaluation shows that ACAMDA outperforms existing methods, particularly when applied to novel and unseen domains.
Yuewen Sun, Erli Wang, Biwei Huang, Chaochao Lu, Changyin Sun 0001, Kun Zhang 0001
AAAI1
2024 CaRiNG: Learning Temporal Causal Representation under Non-Invertible Generation Process
abstract
Identifying the underlying time-delayed latent causal processes in sequential data is vital for grasping temporal dynamics and making downstream reasoning. While some recent methods can robustly identify these latent causal variables, they rely on strict assumptions about the invertible generation process from latent variables to observed data. However, these assumptions are often hard to satisfy in real-world applications containing information loss. For instance, the visual perception process translates a 3D space into 2D images, or the phenomenon of persistence of vision incorporates historical data into current perceptions. To address this challenge, we establish an identifiability theory that allows for the recovery of independent latent components even when they come from a nonlinear and non-invertible mix. Using this theory as a foundation, we propose a principled approach, CaRiNG, to learn the Causal Representation of Non-invertible Generative temporal data with identifiability guarantees. Specifically, we utilize temporal context to recover lost latent information and apply the conditions in our theory to guide the training process. Through experiments conducted on synthetic datasets, we validate that our CaRiNG method reliably identifies the causal process, even when the generation process is non-invertible. Moreover, we demonstrate that our approach considerably improves temporal understanding and reasoning in practical applications.
Guangyi Chen 0002, Yifan Shen 0004, Zhenhao Chen, Xiangchen Song, Yuewen Sun, Weiran Yao, Kun Zhang 0001
ICML5
2024 On the Parameter Identifiability of Partially Observed Linear Causal Models
abstract
Linear causal models are important tools for modeling causal dependencies and yet in practice, only a subset of the variables can be observed. In this paper, we examine the parameter identifiability of these models by investigating whether the edge coefficients can be recovered given the causal structure and partially observed data. Our setting is more general than that of prior research—we allow all variables, including both observed and latent ones, to be flexibly related, and we consider the coefficients of all edges, whereas most existing works focus only on the edges between observed variables. Theoretically, we identify three types of indeterminacy for the parameters in partially observed linear causal models. We then provide graphical conditions that are sufficient for all parameters to be identifiable and show that some of them are provably necessary. Methodologically, we propose a novel likelihood-based parameter estimation method that addresses the variance indeterminacy of latent variables in a specific way and can asymptotically recover the underlying parameters up to trivial indeterminacy. Empirical studies on both synthetic and real-world datasets validate our identifiability theory and the effectiveness of the proposed method in the finite-sample regime.
Xinshuai Dong, Ignavier Ng, Biwei Huang, Yuewen Sun, Songyao Jin, Roberto Legaspi, Peter Spirtes, Kun Zhang 0001
NeurIPS4
2024 Identifying Latent State-Transition Processes for Individualized Reinforcement Learning
abstract
The application of reinforcement learning (RL) involving interactions with individuals has grown significantly in recent years. These interactions, influenced by factors such as personal preferences and physiological differences, causally influence state transitions, ranging from health conditions in healthcare to learning progress in education. As a result, different individuals may exhibit different state-transition processes. Understanding individualized state-transition processes is essential for optimizing individualized policies. In practice, however, identifying these state-transition processes is challenging, as individual-specific factors often remain latent. In this paper, we establish the identifiability of these latent factors and introduce a practical method that effectively learns these processes from observed state-action trajectories. Experiments on various datasets show that the proposed method can effectively identify latent state-transition processes and facilitate the learning of individualized RL policies.
Yuewen Sun, Biwei Huang, Yu Yao 0005, Donghuo Zeng, Xinshuai Dong, Songyao Jin, Roberto Legaspi, Kazushi Ikeda, Peter Spirtes, Kun Zhang 0001
NeurIPS1
2024 Counterfactual Reasoning Using Predicted Latent Personality Dimensions for Optimizing Persuasion Outcome
Donghuo Zeng, Roberto Legaspi, Yuewen Sun, Xinshuai Dong, Kazushi Ikeda, Peter Spirtes, Kun Zhang 0001
PERSUASIVE3
2023 Model-Based Transfer Reinforcement Learning Based on Graphical Model Representations
abstract
Reinforcement learning (RL) plays an essential role in the field of artificial intelligence but suffers from data inefficiency and model-shift issues. One possible solution to deal with such issues is to exploit transfer learning. However, interpretability problems and negative transfer may occur without explainable models. In this article, we define Relation Transfer as explainable and transferable learning based on graphical model representations, inferring the skeleton and relations among variables in a causal view and generalizing to the target domain. The proposed algorithm consists of the following three steps. First, we leverage a suitable casual discovery method to identify the causal graph based on the augmented source domain data. After that, we make inferences on the target model based on the prior causal knowledge. Finally, offline RL training on the target model is utilized as prior knowledge to improve the policy training in the target domain. The proposed method can answer the question of what to transfer and realize zero-shot transfer across related domains in a principled way. To demonstrate the robustness of the proposed framework, we conduct experiments on four classical control problems as well as one simulation to the real-world application. Experimental results on both continuous and discrete cases demonstrate the efficacy of the proposed method.
Yuewen Sun, Kun Zhang 0001, Changyin Sun 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Learning Temporally Causal Latent Processes from General Temporal Data
Weiran Yao, Yuewen Sun, Alex Ho, Changyin Sun 0001, Kun Zhang 0001
ICLR2
2021 A Parallel Framework of Adaptive Dynamic Programming Algorithm With Off-Policy Learning
abstract
In this article, a model-free online adaptive dynamic programming (ADP) approach is developed for solving the optimal control problem of nonaffine nonlinear systems. Combining the off-policy learning mechanism with the parallel paradigm, multithread agents are employed to collect the transitions by interacting with the environment that significantly augments the number of sampled data. On the other hand, each thread agent explores the environment with different initial states under its own behavior policy that enhances the exploration capability and alleviates the correlation between the sampled data. After the policy evaluation process, only one step update is required for policy improvement based on the policy gradient method. The stability of the system under iterative control laws is guaranteed. Moreover, the convergence analysis is given to prove that the iterative Q-function is monotonically nonincreasing and finally converges to the solution of the Hamilton-Jacobi-Bellman (HJB) equation. For implementing the algorithm, the actor-critic (AC) structure is utilized with two neural networks (NNs) to approximate the Q-function and the control policy. Finally, the effectiveness of the proposed algorithm is verified by two numerical examples.
Changyin Sun 0001, Xiaofeng Li 0014, Yuewen Sun
IEEE Trans. Neural Networks Learn. Syst.3
2018 A Broad Neural Network Structure for Class Incremental Learning
Wenzhang Liu, Haiqin Yang, Yuewen Sun, Changyin Sun 0001
ISNN3