Xin Wang 0179

dblp:10/5630-179 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Autonomous Causal Discovery: Evaluating LLMs' Priors and Constraint Strategies for Reliability
abstract
Expert-guided Causal Structure Learning (CSL) incorporates prior knowledge to improve the accuracy of causal discovery, yet the acquisition of such knowledge is often restricted by the availability of human experts. While Large Language Models (LLMs) provide an alternative source of causal priors, LLM-derived knowledge can be inconsistent with the true causal structure due to hallucinations or contextual misinterpretations. This paper introduces a structural constraint measurement framework, which defines constraint strength and constraint quality to describe reliability and effectiveness, enabling a systematic evaluation of LLM-derived constraints. Using this framework, we evaluate five categories of structural constraints: Edge Existence (EEC), Edge Forbidden (EFC), Path Existence (PEC), Path Forbidden (PFC), and Order Constraints (OC). Our theoretical and empirical analyses demonstrate that while EEC offers high constraint strength, it exhibits low quality when derived from LLMs; conversely, PFC and OC provide a balanced trade-off between search-space pruning and reliability. Building on these insights, we propose a two-level CSL optimization framework that partitions the search space by node order and refines the structure using global path constraints. The results show that this framework provides an effective way to incorporate noisy LLM-derived priors into CSL, particularly in settings where expert knowledge is limited.
Lyuzhou Chen, Xiangyu Wang 0016, Taiyu Ban, Derui Lyu, Qinrui Zhu, Xin Wang 0179, Huanhuan Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Generation-Augmented and Embedding Fusion in Document-Level Event Argument Extraction
abstract
Document-level event argument extraction is a crucial task that aims to extract arguments from the entire document, beyond sentence-level analysis. Prior classification-based models still fail to explicitly capture significant relationships and heavily relies on large-scale datasets. In this study, we propose a novel approach called Generation-Augmented and Embedding Fusion. This approach first uses predefined templates and generative language models to produce an embedding capturing role relationship information, then integrates it into the foundational embedding derived from a classification model through a noval embedding fusion mechanism. We conduct the extensive experiments on the RAMS and WikiEvents datasets to demonstrate that our approach is more effective than the baselines, and that it is also data-efficient in low-resource scenarios.
Xingjian Lin, Shengfei Lyu, Xin Wang 0179, Qiuju Chen, Huanhuan Chen 0001
COLING3
2025 Variational Counterfactual Intervention Planning to Achieve Target Outcomes
abstract
A key challenge in personalized healthcare is identifying optimal intervention sequences to guide temporal systems toward target outcomes, a novel problem we formalize as counterfactual target achievement. In addressing this problem, directly adopting counterfactual estimation methods face compounding errors due to the unobservability of counterfactuals. To overcome this, we propose Variational Counterfactual Intervention Planning (VCIP), which reformulates the problem by modeling the conditional likelihood of achieving target outcomes, implemented through variational inference. By leveraging the g-formula to bridge the gap between interventional and observational log-likelihoods, VCIP enables reliable training from observational data. Experiments on both synthetic and real-world datasets show that VCIP significantly outperforms existing methods in target achievement accuracy.
Xin Wang 0179, Shengfei Lyu, Chi Luo, Xiren Zhou, Huanhuan Chen 0001
ICML1
2025 Differentiable Structure Learning with Ancestral Constraints
abstract
Differentiable structure learning of causal directed acyclic graphs (DAGs) is an emerging field in causal discovery, leveraging powerful neural learners. However, the incorporation of ancestral constraints, essential for representing abstract prior causal knowledge, remains an open research challenge. This paper addresses this gap by introducing a generalized framework for integrating ancestral constraints. Specifically, we identify two key issues: the non-equivalence of relaxed characterizations for representing path existence and order violations among paths during optimization. In response, we propose a binary-masked characterization method and an order-guided optimization strategy, tailored to address these challenges. We provide theoretical justification for the correctness of our approach, complemented by experimental evaluations on both synthetic and real-world datasets.
Taiyu Ban, Changxin Rong, Xiangyu Wang 0016, Lyuzhou Chen, Xin Wang 0179, Derui Lyu, Qinrui Zhu, Huanhuan Chen 0001
ICML5
2025 Enhancing Counterfactual Estimation: A Focus on Temporal Treatments
abstract
In the medical field, treatment sequences significantly influence future outcomes through complex temporal interactions. Therefore, highlighting the role of temporal treatments within the model is crucial for accurate counterfactual estimation, which is often overlooked in current methods. To address this, we employ Koopman theory, known for its capability to model complex dynamic systems, and introduce a novel model named the Counterfactual Temporal Dynamics Network via Neural Koopman Operators (CTD-NKO). This model utilizes Koopman operators to encapsulate sequential treatment data, aiming to capture the causal dynamics within the system induced by temporal interactions between treatments. Moreover, CTD-NKO implements a weighting strategy that aligns joint and marginal distributions of the system state and the current treatment to mitigate time-varying confounding bias. This deviates from the balanced representation strategy employed by existing methods, as we demonstrate that such a strategy may suffer from the potential information loss of historical treatments. These designs allow CTD-NKO to exploit treatment information more thoroughly and effectively, resulting in superior performance on both synthetic and real-world datasets.
Xin Wang 0179, Shengfei Lyu, Kangyang Luo, Lishan Yang 0004, Huanhuan Chen 0001, Chunyan Miao
IJCAI1
2025 Pattern-Guided Adaptive Prior for Structure Learning
abstract
Learning the causality between variables, known as DAG structure learning, is critical yet challenging due to issues such as insufficient data and noise. While prior knowledge can improve the learning process and refine the DAG structure, incorporating prior knowledge is not without pitfalls. In particular, we find that the gap between the imprecise prior knowledge and the exact weights modeled by existing methods may result in deviation in edge weights. Such deviation can subsequently cause significant inaccuracies when learning the DAG structure. This paper addresses this challenge by providing a theoretical analysis of the impact of deviation in edge weights during the optimization process of structure learning. We identify two special graph patterns that arise due to the deviation and show that their occurrence increases as the degree of deviation grows. Building on this analysis, we propose the Pattern-Guided Adaptive Prior (PGAP) framework. PGAP detects these patterns as structural signals during optimization and adaptively adjusts the structure learning process to counteract the identified weight deviation, thereby improving the integration of prior knowledge. Experiments verify the effectiveness and robustness of the proposed method.
Lyuzhou Chen, Yanze Gao, Xiangyu Wang 0016, Derui Lyu, Taiyu Ban, Xin Wang 0179, Xiren Zhou, Huanhuan Chen 0001
NeurIPS7
2024 A Dual-module Framework for Counterfactual Estimation over Time
abstract
Efficiently and effectively estimating counterfactuals over time is crucial for optimizing treatment strategies. We present the Adversarial Counterfactual Temporal Inference Network (ACTIN), a novel framework with dual modules to enhance counterfactual estimation. The balancing module employs a distribution-based adversarial method to learn balanced representations, extending beyond the limitations of current classification-based methods to mitigate confounding bias across various treatment types. The integrating module adopts a novel Temporal Integration Predicting (TIP) strategy, which has a wider receptive field of treatments and balanced representations from the beginning to the current time for a more profound level of analysis. TIP goes beyond the established Direct Predicting (DP) strategy, which only relies on current treatments and representations, by empowering the integrating module to effectively capture long-range dependencies and temporal treatment interactions. ACTIN exceeds the confines of specific base models, and when implemented with simple base models, consistently delivers state-of-the-art performance and efficiency across both synthetic and real-world datasets.
Xin Wang 0179, Shengfei Lyu, Lishan Yang 0004, Yibing Zhan, Huanhuan Chen 0001
ICML1
2024 Differentiable Structure Learning with Partial Orders
abstract
Differentiable structure learning is a novel line of causal discovery research that transforms the combinatorial optimization of structural models into a continuous optimization problem. However, the field has lacked feasible methods to integrate partial order constraints, a critical prior information typically used in real-world scenarios, into the differentiable structure learning framework. The main difficulty lies in adapting these constraints, typically suited for the space of total orderings, to the continuous optimization context of structure learning in the graph space. To bridge this gap, this paper formalizes a set of equivalent constraints that map partial orders onto graph spaces and introduces a plug-and-play module for their efficient application. This module preserves the equivalent effect of partial order constraints in the graph space, backed by theoretical validations of correctness and completeness. It significantly enhances the quality of recovered structures while maintaining good efficiency, which learns better structures using 90\% fewer samples than the data-based method on a real-world dataset. This result, together with a comprehensive evaluation on synthetic cases, demonstrates our method's ability to effectively improve differentiable structure learning with partial orders.
Taiyu Ban, Lyuzhou Chen, Xiangyu Wang 0016, Xin Wang 0179, Derui Lyu, Huanhuan Chen 0001
NeurIPS4
2023 Temporal knowledge graph embedding via sparse transfer matrix
Xin Wang 0179, Shengfei Lyu, Xiangyu Wang 0016, Huanhuan Chen 0001
Inf. Sci.1
2023 Knowledge Extraction From National Standards for Natural Resources: A Method for Multi-Domain Texts
abstract
National standards for natural resources (NSNR) plays an important role in promoting efficient use of China's natural resources, which sets standards for many domains such as marine and land resources. Its revision is difficult since standards in different domains may overlap or conflict. To facilitate the revision of NSNR, this paper extracts structural knowledge from the NSNR files to assist its revision. NSNR files are in multi-domain texts, where the traditional knowledge extraction methods could fall short in recalling multi-domain entities. To address this issue, this paper proposes a knowledge extraction method for multi-domain texts, including sub-domain relation discovery (SRD) and domain semantic features fusion (DSFF) module. SRD splits NSNR into sub-domains to facilitate the relation discovery. DSFF integrates relation features in the conditional random field (CRF) model to improve the capability of multi-domain entity recognition. Experimental results demonstrate that the proposed method could effectively extract structural knowledge from NSNR.
Taiyu Ban, Xiangyu Wang 0016, Xin Wang 0179, Jiarun Zhu, Lvzhou Chen, Yizhan Fan
J. Database Manag.3
2022 Nonlinear Causal Discovery in Time Series
abstract
Recent years have witnessed the proliferation of the Functional Causal Model (FCM) for causal learning due to its intuitive representation and accurate learning results. However, existing FCM-based algorithms suffer from the ubiquitous nonlinear relations in time-series data, mainly because these algorithms either assume linear relationships, or nonlinear relationships with additive noise, or do not introduce additional assumptions but can only identify nonlinear causality between two variables. This paper contributes in particular to a practical FCM-based causal learning approach, which can maintain effectiveness for real-world nonstationary data with general nonlinear relationships and unlimited variable scale.Specifically, the non-stationarity of time series data is first exploited with the nonlinear independent component analysis, to discover the underlying components or latent disturbances. Then, the conditional independence between variables and these components is studied to obtain a relation matrix, which guides the algorithm to recover the underlying causal graph. The correctness of the proposal is theoretically proved, and extensive experiments further verify its effectiveness. To the best of our knowledge, the proposal is the first so far that can fully identify causal relationships under general nonlinear conditions.
Xin Wang 0179, Shikang Liu, Huanhuan Chen 0001
CIKM3
2022 Generalization Bounds for Estimating Causal Effects of Continuous Treatments
abstract
We focus on estimating causal effects of continuous treatments (e.g., dosage in medicine), also known as dose-response function. Existing methods in causal inference for continuous treatments using neural networks are effective and to some extent reduce selection bias, which is introduced by non-randomized treatments among individuals and might lead to covariate imbalance and thus unreliable inference. To theoretically support the alleviation of selection bias in the setting of continuous treatments, we exploit the re-weighting schema and the Integral Probability Metric (IPM) distance to derive an upper bound on the counterfactual loss of estimating the average dose-response function (ADRF), and herein the IPM distance builds a bridge from a source (factual) domain to an infinite number of target (counterfactual) domains. We provide a discretized approximation of the IPM distance with a theoretical guarantee in the practical implementation. Based on the theoretical analyses, we also propose a novel algorithm, called Average Dose- response estiMatIon via re-weighTing schema (ADMIT). ADMIT simultaneously learns a re-weighting network, which aims to alleviate the selection bias, and an inference network, which makes factual and counterfactual estimations. In addition, the effectiveness of ADMIT is empirically demonstrated in both synthetic and semi-synthetic experiments by outperforming the existing benchmarks.
Xin Wang 0179, Shengfei Lyu, Huanhuan Chen 0001
NeurIPS1
2022 Domain knowledge-enhanced variable selection for biomedical data analysis
Zhenchao Tao, Bingbing Jiang 0001, Xin Wang 0179, Huanhuan Chen 0001
Inf. Sci.5
2022 Dynamic Link Prediction for Discovery of New Impactful COVID-19 Research Approaches
abstract
In fighting the COVID-19 pandemic, the main challenges include the lack of prior research and the urgency to find effective solutions. It is essential to accurately and rapidly summarize the relevant research work and explore potential solutions for diagnosis, treatment and prevention of COVID-19. It is a daunting task to summarize the numerous existing research works and to assess their effectiveness. This paper explores the discovery of new COVID-19 research approaches based on dynamic link prediction, which analyze the dynamic topological network of keywords to predict possible connections of research concepts. A dynamic link prediction method based on multi-granularity feature fusion is proposed. Firstly, a multi-granularity temporal feature fusion method is adopted to extract the temporal evolution of different order subgraphs. Secondly, a hierarchical feature weighting method is proposed to emphasize actively evolving nodes. Thirdly, a semantic repetition sampling mechanism is designed to avoid the negative effect of semantically equivalent medical entities on the real structure of the graph, and to capture the real topological structure features. Experiments are performed on the COVID-19 Open Research Dataset to assess the performance of the model. The results show that the proposed model performs significantly better than existing state-of-the-art models, thereby confirming the effectiveness of the proposed method for the discovery of new COVID-19 research approaches.
Xiangyu Wang 0016, Taiyu Ban, Jiarun Zhu, Lyuzhou Chen, Xin Wang 0179, Huanhuan Chen 0001, Cyril Leung, Chunyan Miao
IEEE J. Biomed. Health Informatics7