Huiyan Sun

dblp:157/0872 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0002-4664-7147ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 7 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Neuro-Symbolic Causal Boosting: A Framework for Interpretable Attribution of Business Fluctuations
abstract
Attributing business fluctuations to actionable drivers is a critical component of decision-making in high-stakes domains. However, prevailing predictive models, predominantly driven by correlations, often yield inconsistent explanations that degrade under distribution shifts or latent confounding. Empirical causal inference to address this limitation remains challenge due to the identifiability gap in purely data-driven discovery and the complexity of encoding domain knowledge into differentiable learning pipelines. To bridge this gap, we propose Neuro-Symbolic Causal Boosting, a unified framework that integrates semantic domain priors with gradient-based causal estimation. First, we introduce the Complete Cause Identification Algorithm (CCIA). Unlike global search methods, CCIA recursively reconstructs the ancestral causal graph of the target variable by employing Kolmogorov-Arnold Networks (KANs) as high-precision filters for low-order independence testing, coupled with a neuro-symbolic adjudication based on Large Language Model to resolve directionality. Subsequently, the identified structure scaffolds the Causal Additive Boosting Network (CABN). Grounded in the theory of structural identification, CABN enforces a reverse topological learning process. It utilizes weighted KANs to sequentially estimate downstream effects and adjust for confounding, thereby isolating invariant causal mechanisms. Empirical evaluation on a real-world telesales dataset and synthetic benchmarks demonstrates that our framework achieves a 57.5% reduction in Out-of-Distribution prediction error compared to strong correlation-based baselines. Additionally, it identifies and quantifies the impact of actionable drivers, providing a structured approach from observational data to trustworthy and interpretable business strategies.
Yonghe Zhao, Yuezhu Wang, Huiyan Sun
KR4
2025 Identifying cancer prognosis genes through causal learning
abstract
Accurate identification of causal genes for cancer prognosis is critical for estimating disease progression and guiding treatment interventions. In this study, we propose CPCG (Cancer Prognosis's Causal Gene), a two-stage framework identifying gene sets causally associated with patient prognosis across diverse cancer types using transcriptomic data. Initially, an ensemble approach models gene expression's impact on survival with parametric and semiparametric hazard models. Subsequently, an iterative conditional independence test combined with graph pruning is utilized to infer the causal skeleton, thereby pinpointing prognosis-related genes. Experiments on transcriptomic data from 18 cancer types sourced from The Cancer Genome Atlas Project demonstrate CPCG's effectiveness in predicting prognosis under four evaluation metrics. Validations on 24 additional datasets covering 12 cancer types from the Gene Expression Omnibus and the Chinese Glioma Genome Atlas Project further demonstrate CPCG's robustness and generalizability. CPCG identifies a concise but reliable set of genes, obviating the need for gene combination enumeration for survival time estimation. These genes are also proved closely linked to crucial biological processes in cancer. Moreover, CPCG constructs a stable causal skeleton and exhibits insensitivity to the order of data shuffling. Overall, CPCG is a powerful tool for extracting cancer prognostic biomarkers, offering interpretability, generalizability, and robustness. CPCG holds promise for facilitating targeted interventions in clinical treatment strategies.
Siwei Wu, Chaoyi Yin, Yuezhu Wang, Huiyan Sun
Briefings Bioinform.4
2025 Cancer gene identification through integrating causal prompting large language model with omics data-driven causal inference
abstract
Identifying genes causally linked to cancer from a multi-omics perspective is essential for understanding the mechanisms of cancer and improving therapeutic strategies. Traditional statistical and machine-learning methods that rely on generalized correlation approaches to identify cancer genes often produce redundant, biased predictions with limited interpretability, largely due to overlooking confounding factors, selection biases, and the nonlinear activation function in neural networks. In this study, we introduce a novel framework for identifying cancer genes across multiple omics domains, named ICGI (Integrative Causal Gene Identification), which leverages a large language model (LLM) prompted with causality contextual cues and prompts, in conjunction with data-driven causal feature selection. This approach demonstrates the effectiveness and potential of LLMs in uncovering cancer genes and comprehending disease mechanisms, particularly at the genomic level. However, our findings also highlight that current LLMs may not capture comprehensive information across all omics levels. By applying the proposed causal feature selection module to transcriptomic datasets from six cancer types in The Cancer Genome Atlas and comparing its performance with state-of-the-art methods, it demonstrates superior capability in identifying cancer genes that distinguish between cancerous and normal samples. Additionally, we have developed an online service platform that allows users to input a gene of interest and a specific cancer type. The platform provides automated results indicating whether the gene plays a significant role in cancer, along with clear and accessible explanations. Moreover, the platform summarizes the inference outcomes obtained from data-driven causal learning methods.
Haolong Zeng, Chaoyi Yin, Chunyang Chai, Yuezhu Wang, Huiyan Sun
Briefings Bioinform.6
2024 Estimating Individual Causal Treatment Effect by Variable Decomposition
abstract
Estimating individual-level causal effects is crucial for decision-making in various domains, such as personalized healthcare, social marketing, and public policy. Addressing confounding bias is a critical step in accurately estimating the causal effects of treatments on outcomes. However, many current causal inference approaches consider all observed variables as confounders without distinguishing them from colliders or indirect (two-order) colliders. This may lead to M-bias when improperly eliminating confounding bias. In this study, we propose a new framework to accurately estimate individual-level treatment effects by considering a causal structure that includes both confounding variables and indirect colliders. Specifically, we first perform a sample reweighting to approximately eliminate confounding bias. Then, we restore the covariate’ potential latent parents and extract the modules solely related to the outcome. Finally, we take both these modules with the treatment variables to infer counterfactuals for causal inference. To validate the effectiveness of our proposed approach, we conduct extensive experiments on synthetic and commonly used semi-synthetic benchmark datasets. The experimental results demonstrate that our method outperforms current state-of-the-art methods.
Hongyang Jiang 0002, Yonghe Zhao, Yangkun Cao, Huiyan Sun, Yi Chang 0001
IJCNN5
2024 ScMOGAE: A Graph Convolutional Autoencoder-Based Multi-omics Data Integration Framework for Single-Cell Clustering
Benjie Zhou, Hongyang Jiang 0002, Yuezhu Wang, Huiyan Sun
ISBRA (1)5
2024 De-confounding representation learning for counterfactual inference on continuous treatment via generative adversarial network
Yonghe Zhao, Haolong Zeng, Huiyan Sun
Data Min. Knowl. Discov.5
2024 Modeling Interference for Individual Treatment Effect Estimation from Networked Observational Data
abstract
Estimating individual treatment effect (ITE) from observational data has attracted great interest in recent years, which plays a crucial role in decision-making across many high-impact domains such as economics, medicine, and e-commerce. Most existing studies of ITE estimation assume that different units at play are independent and do not influence each other. However, many social science experiments have shown that there often exist different levels of interactions between units in observational data, especially in a networked environment. As a result, the treatment assignment of one unit can affect the outcome of other units connected to it in the network, which is referred to as the interference or spillover effect . In this article, we study an important problem of ITE estimation from networked observational data by modeling the interference between different units and provide a principled framework to support such study. Methodologically, we propose a novel framework, SPNet , that first captures the influence of hidden confounders with the aid of graph convolutional network and then models the interference by introducing an environment summary variable and developing a masked attention mechanism. Experimental evaluations on several semi-synthetic datasets based on real-world networks corroborate the superiority of our proposed framework over state-of-the-art individual treatment effect estimation methods.
Jing Ma 0002, Jundong Li, Ruocheng Guo, Huiyan Sun, Yi Chang 0001
ACM Trans. Knowl. Discov. Data5
2023 Contrastive self-supervised graph convolutional network for detecting the relationship among lncRNAs, miRNAs, and diseases
abstract
Inferring potential relationships among long non-coding RNAs (lncRNAs), microRNAs (miRNAs), and diseases play a crucial role in investigation of disease aetiology and pathogenesis. Due to the high cost of laboratory experiments, there is a practical requirement to develop appropriate computational methods that promise to accelerate the experimental screening process for potential lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs), and lncRNA-miRNA interactions (LMIs). However, most existing methods are applied to predict LDAs, MDAs, and LMIs in specific domains, neglecting the important benefits of integrating multiple sources data and limiting the ability of transferring models to other tasks. Furthermore, with the high sparsity of LDA, MDA, and LMI data, it is difficult for many computational models to exploit enough knowledge to learn the comprehensive patterns of node embedding. In this study, inspired by the recent success of graph contrastive learning, we develop a Contrastive Self-supervised Graph convolutional network to identify potential LDAs, MDAs, and LMIs (called CSGLMD). CSGLMD combines supervised learning and self-supervised learning to fully capture node features. Specifically, CSGLMD primarily leverages the rich association and similarity relationships among lncRNA, miRNA, and disease to construct a lncRNA-miRNA-disease heterogeneous graph (LMDHG) that contains three types of biological entities. It can effectively embed multi-source biological data and assist the model extension to other prediction tasks. In addition, we consider applying a label instantiation mechanism to make the LMDHG better adapt graph neural network structures and control the strength of similarity relationships between the same biological entities. Secondly, CSGLMD implements graph convolutional network (GCN) as encoder to extract node embedding features from the LMDHG, and utilizes a multi-relational modelling decoder to predict LDAs, MDAs, or LMIs. Finally, we designed a contrastive self-supervised learning task that guides the learning of node embeddings without relying on labels, and acts as a regularize in a multi-task learning paradigm to enhance the generalization ability of the model. Extensive results on two datasets (from the old and new versions of the database, respectively) show that CSGLMD significantly outperforms 12 state-of-the-art methods (5 LDA prediction and 7 MDA prediction) in predicting disease-associated lncRNAs and miRNAs. Case studies on old and new datasets can further demonstrate the capability of CSGLMD to discover disease-related new candidate lncRNAs and miRNAs. The source data and code for the proposed model are publicly available on https://github.com/sheng-n/CSGLMD.
Nan Sheng, Lan Huang 0002, Yan Wang 0028, Huiyan Sun, Xuping Xie
BIBM5
2023 Improve Robustness of Graph Neural Networks: Multi-hop Neighbors Meet Homophily-based Truncation Defense
abstract
Graph Neural Networks (GNNs), as promising deep learning approaches, have been applied in various areas. However, it is also known that they are vulnerable to adversarial attacks, which raises many concerns in real application. Regarding Graph Structure Attacks (GSAs), though existing homophily-based truncation defenses have showed strong defense capacity, they suffer the problem of losing much effective neighborhood information in the process of removing adversarial edges, causing a limited performance on both clean and attacked graphs. In this paper, we consider the question: Can we capture more effective neighborhood information by utilizing the higher-order network to help improve the performance of the homophily-based truncation defense? To answer it, we first explore the impacts of different GSAs on the 1-hop and 2-hop networks. We theoretically and empirically find that the 2-hop network also has a strong information retention capacity like the 1-hop network after many GSAs. Motivated by this, we combine the 2-hop network with the homophily-based truncation defense, constructing a stronger defender, MHR-GCN. It integrates effective neighborhood information from both 1-hop and 2-hop networks. Extensive experiments demonstrate that MHR-GCN significantly improves the performance of the truncation defense and outperforms the state-of-the-arts under various GSA settings, especially when the graph is heavily perturbed.
Huiyan Sun, Haobo Shi
IJCNN2
2022 SemiITE: Semi-supervised Individual Treatment Effect Estimation via Disagreement-Based Co-training
Jing Ma 0002, Jundong Li, Huiyan Sun, Yi Chang 0001
ECML/PKDD (4)4
2022 Identification of key somatic oncogenic mutation based on a confounder-free causal inference model
abstract
Abnormal cell proliferation and epithelial-mesenchymal transition (EMT) are the essential events that induce cancer initiation and progression. A fundamental goal in cancer research is to develop an efficient method to detect mutational genes capable of driving cancer. Although several computational methods have been proposed to identify these key mutations, many of them focus on the association between genetic mutations and functional changes in relevant biological processes, but not their real causality. Causal effect inference provides a way to estimate the real induce effect of a certain mutation on vital biological processes of cancer initiation and progression, through addressing the confounder bias due to neutral mutations and unobserved latent variables. In this study, integrating genomic and transcriptomic data, we construct a novel causal inference model based on a deep variational autoencoder to identify key oncogenic somatic mutations. Applied to 10 cancer types, our method quantifies the causal effect of genetic mutations on cell proliferation and EMT by reducing both observed and unobserved confounding biases. The experimental results indicate that genes with higher mutation frequency do not necessarily mean they are more potent in inducing cancer and promoting cancer development. Moreover, our study fills a gap in the use of machine learning for causal inference to identify oncogenic mutations.
Huiyan Sun
PLoS Comput. Biol.3
2021 An AutoEncoder-Based Matrix Factorization Approach to Estimating Cell Proportion from Bulk Tumor RNA-seq Data
abstract
The deconvolution of infiltrating immune cells and stromal cells from complex tumor tissues is significant for studying the impact of these cells on tumor development, as well as assisting cancer therapies. Integrating cell-specific marker genes from several references, we put forward an AutoEncoder-based matrix factorization method to estimate the cell proportion in heterogeneous samples. The proportion predicted by our method achieved over 0.95 Pearson correlation coefficient (PCC) with the ground truth. Moreover, through analyzing association between cell proportion of tumor tissues and clinical information of the tumor patients, we found that the proportion of cancer-associated fibroblast (CAF) gradually increased with the progress of tumor, and the proportion of B cell in tissues was significantly related to the five-year survival rate of tumor patients.
Yingze Xu, Yan Wang 0028, Xuping Xie, Huiyan Sun
BIBM6
2021 The bioinformatics toolbox for circRNA discovery and analysis
abstract
Circular RNAs (circRNAs) are a unique class of RNA molecule identified more than 40 years ago which are produced by a covalent linkage via back-splicing of linear RNA. Recent advances in sequencing technologies and bioinformatics tools have led directly to an ever-expanding field of types and biological functions of circRNAs. In parallel with technological developments, practical applications of circRNAs have arisen including their utilization as biomarkers of human disease. Currently, circRNA-associated bioinformatics tools can support projects including circRNA annotation, circRNA identification and network analysis of competing endogenous RNA (ceRNA). In this review, we collected about 100 circRNA-associated bioinformatics tools and summarized their current attributes and capabilities. We also performed network analysis and text mining on circRNA tool publications in order to reveal trends in their ongoing development.
Liang Chen 0021, Changliang Wang, Huiyan Sun, Juexin Wang, Yanchun Liang 0001, Yan Wang 0028, Garry Wong
Briefings Bioinform.3
2020 Unsupervised Nonlinear Feature Selection from High-Dimensional Signed Networks
abstract
With the rapid development of social media services in recent years, relational data are explosively growing. The signed network, which consists of a mixture of positive and negative links, is an effective way to represent the friendly and hostile relations among nodes, which can represent users or items. Because the features associated with a node of a signed network are usually incomplete, noisy, unlabeled, and high-dimensional, feature selection is an important procedure to eliminate irrelevant features. However, existing network-based feature selection methods are linear methods, which means they can only select features that having the linear dependency on the output values. Moreover, in many social data, most nodes are unlabeled; therefore, selecting features in an unsupervised manner is generally preferred. To this end, in this paper, we propose a nonlinear unsupervised feature selection method for signed networks, called SignedLasso. This method can select a small number of important features with nonlinear associations between inputs and output from a high-dimensional data. More specifically, we formulate unsupervised feature selection as a nonlinear feature selection problem with the Hilbert-Schmidt Independence Criterion Lasso (HSIC Lasso), which can find a small number of features in a nonlinear manner. Then, we propose the use of a deep learning-based node embedding to represent node similarity without label information and incorporate the node embedding into the HSIC Lasso. Through experiments on two real world datasets, we show that the proposed algorithm is superior to existing linear unsupervised feature selection methods.
Tingyu Xia, Huiyan Sun, Makoto Yamada, Yi Chang 0001
AAAI3
2020 Identification of Cancer Development Related Pathways Based on Co-Expression Analyses
abstract
Cancer is a rapidly evolving disease, with complex physiological changes throughout its development. Different patients of the same cancer type may require distinct treatments depending on the level of development. Hence it is essential to identify the genes and biological processes strongly associated with cancer progression to design an effective treatment plan. Our study has found that cancer samples of the same development stage (or grade, subtype) tend to share highly unique co-expression patterns, providing considerably stronger discerning power than differential expressions for cancer staging (and grading, subtyping). Based on this, we have developed a framework for identification and analyses of genes and pathways strongly associated with a cancer's development through identification of co-expression patterns that become increasingly stronger or weaker over cancer samples from early through advanced stages. Functional analyses of such co-expressed genes reveal that (1) cell-cycle, immune response, ribosome, proteasome and oxidative phosphorylation, among others, strongly associate with cancer development, (2) the co-expression patterns among ribosome, proteasome and oxidative phosphorylation genes tend to become increasingly weaker as a cancer advances, for almost all cancer types, and (3) the co-expression patterns among cell cycle and immune response genes tend to become increasingly stronger with cancer progression. We anticipate that co-expression-based analyses like we present here will become a key technique for functional studies of cancer development and evolution.
Hongyang Jiang 0002, Zhihang Wang, Chaoyi Yin, Peishuo Sun, Ying Xu 0001, Huiyan Sun
BIBM6
2019 Multi-Classification of Cancer Samples Based on Co-Expression Analyses
abstract
Cancer staging, grading and subtyping all represent important problems for precision diagnosis, treatment and mechanistic studies of cancer. The majority of the existing computational methods solve this problem via multi-classification of differential gene-expressions of cancer samples of specific classes (Stages, Grades and subtypes) vs. controls. However, the performance of such classification techniques is generally not satisfactory since the discerning power of differential expression patterns in such classifications is limited. We present here a multi-classification technique, based on co-expression patterns specific to individual subclasses in provided training data as co-expression patterns tend to be more conserved than differential expressions within each subclass. A challenge in implementing this strategy lies in how to effectively derive co-expression patterns in individual samples, which is solved through comparing co-expression patterns within a subclass and those in the subclass plus a new sample. Compared with the state-of-the-art gene expression-based classification methods, our method outperforms them in cancer staging, grading and subtyping of cancer samples from TCGA in almost all the measures used. In addition, the co-expressed genes computationally selected for classifications are biologically meaningful, which will prove important for diagnostic biomarker design, treatment plan selection and possibly mechanistic studies of cancer.
Hongyang Jiang 0002, Liang Chen 0021, Ying Xu 0001, Huiyan Sun, Yi Chang 0001
BIBM6
2019 Trends in the development of miRNA bioinformatics tools
abstract
MicroRNAs (miRNAs) are small noncoding RNAs that regulate gene expression via recognition of cognate sequences and interference of transcriptional, translational or epigenetic processes. Bioinformatics tools developed for miRNA study include those for miRNA prediction and discovery, structure, analysis and target prediction. We manually curated 95 review papers and ∼1000 miRNA bioinformatics tools published since 2003. We classified and ranked them based on citation number or PageRank score, and then performed network analysis and text mining (TM) to study the miRNA tools development trends. Five key trends were observed: (1) miRNA identification and target prediction have been hot spots in the past decade; (2) manual curation and TM are the main methods for collecting miRNA knowledge from literature; (3) most early tools are well maintained and widely used; (4) classic machine learning methods retain their utility; however, novel ones have begun to emerge; (5) disease-associated miRNA tools are emerging. Our analysis yields significant insight into the past development and future directions of miRNA tools.
Liang Chen 0021, Liisa Heikkinen, Changliang Wang, Huiyan Sun, Garry Wong
Briefings Bioinform.5
2014 Essential protein identification based on essential protein-protein interaction prediction by integrated edge weights
abstract
Essential proteins are crucial to cellular survival and development. Traditionally, essential proteins are identified by knock-out experiments, which are expensive and often fatal to the target organisms. Regarding this, an important approach to essential protein identification is through computational prediction. In this research, we present a novel computational method, Integrated Edge Weights (IEW), to innovatively predict proteins' essentiality based on essential protein-protein interactions. The experimental results on all three organisms: Saccharomyces cere-visiae (Yeast), Escherichia coli (E. coli), and Caenorhabditis ele-gans (C. elegans) show that IEW achieves better performance than the state-of-the-art methods in terms of precision-recall. Furthermore, we have demonstrated that the highly-ranked protein-protein interactions predicted by our approach tend to be biologically significant in Yeast, E. coli, and C. elegans protein-protein interaction (PPI) networks.
Yuexu Jiang, Yan Wang 0028, Wei Pang 0001, Liang Chen 0021, Huiyan Sun, Yanchun Liang 0001, Enrico Blanzieri
BIBM5