Chang Su 0002

dblp:07/2757-2 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0003-4019-6389ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 Democratizing Clinical Risk Prediction with Cross-Cohort Cross-Modal Knowledge Transfer
abstract
Clinical risk prediction plays a crucial role in early disease detection and personalized intervention. While recent models increasingly incorporate multimodal data, their development typically assumes access to large-scale, multimodal datasets and substantial computational resources. In practice, however, most clinical sites operate under resource constraints, with access limited to EHR data alone and insufficient capacity to train complicated models. This gap highlights the urgent need to democratize clinical risk prediction by enabling effective deployment in data- and resource-limited local clinical settings. In this work, we propose a cross-cohort cross-modal knowledge transfer framework that leverages the multimodal model trained on a nationwide cohort and adapts it to local cohorts with only EHR data. We focus on EHR and genetic data as representative multimodal inputs and address two key challenges. First, to mitigate the influence of noisy or less informative biological signals, we propose a novel mixture-of-aggregations design to enhance the modeling of informative and relevant genetic features. Second, to support rapid model adaptation in low-resource sites, we develop a lightweight graph-guided fine-tuning method that adapts pretrained phenotypical EHR representations to target cohorts using limited patient data. Extensive experiments on real-world clinical data validate the effectiveness of our proposed model.
Manqi Zhou, Zilong Bai, Chang Su 0002, Fei Wang 0001
NeurIPS4
2025 Identifying progression subphenotypes of Alzheimer's disease from large-scale electronic health records with machine learning
Manqi Zhou, Alice S. Tang, Alison M. C. Ke, Chang Su 0002, Yu Huang 0018, William G. Mantyh, Michael Jaffee, Katherine P. Rankin, Steven DeKosky, Jiang Bian 0001, Marina Sirota, Fei Wang 0001
J. Biomed. Informatics6
2024 Unified Insights: Harnessing Multi-modal Data for Phenotype Imputation via View Decoupling
abstract
Phenotype imputation plays a crucial role in improving comprehensive and accurate medical evaluation, which in turn can optimize patient treatment and bolster the reliability of clinical research. Despite the adoption of various techniques, multi-modal biological data, which can provide crucial insights into a patient's overall health, is often overlooked. With multi-modal biological data, patient characterization can be enriched from two distinct views: the biological view and the phenotype view. However, the heterogeneity and imprecise nature of the multimodal data still pose challenges in developing an effective method to model from two views. In this paper, we propose a novel framework to incorporate multi-modal biological data via view decoupling. Specifically, we segregate the modeling of biological data from phenotype data in a graph-based learning framework. From the biological view, the latent factors in biological data are discovered to model patient correlation. From the phenotype view, phenotype co-occurrence can be modeled to reveal patterns across patients. Then patients are encoded from these two distinct views. To mitigate the influence of noise and irrelevant information in biological data, we devise a cross-view contrastive knowledge distillation aimed at distilling insights from the biological view to enhance phenotype imputation. We show that phenotype imputation with the proposed model significantly outperforms the state-of-the-art models on the real-world biomedical database.
Weishen Pan, Zilong Bai, Chang Su 0002, Fei Wang 0001
NeurIPS4
2022 Comprehensively modeling heterogeneous symptom progression for Parkinson's disease subtyping
Chang Su 0002, Jielin Xu, Matthew Brendel, Jacqueline R. M. A. Maasch, Zilong Bai, Yingying Zhu 0003, Claire Henchcliffe, Feixiong Cheng, Fei Wang 0001
AMIA1
2022 Unraveling the "Other" Diagnosis of Suspected Suicide Attempts at Emergency Department using Natural Language Processing (NLP)
Yunyu Xiao, Haotan Zhang, Chang Su 0002, Fei Wang 0001
AMIA4
2022 Improving suicide risk prediction via targeted data fusion: proof of concept using medical claims data
abstract
OBJECTIVE: Reducing suicidal behavior among patients in the healthcare system requires accurate and explainable predictive models of suicide risk across diverse healthcare settings. MATERIALS AND METHODS: We proposed a general targeted fusion learning framework that can be used to build a tailored risk prediction model for any specific healthcare setting, drawing on information fusion from a separate more comprehensive dataset with indirect sample linkage through patient similarities. As a proof of concept, we predicted suicide-related hospitalizations for pediatric patients in a limited statewide Hospital Inpatient Discharge Dataset (HIDD) fused with a more comprehensive medical All-Payer Claims Database (APCD) from Connecticut. RESULTS: We built a suicide risk prediction model for the source data (APCD) and calculated patient risk scores. Patient similarity scores between patients in the source and target (HIDD) datasets using their demographic characteristics and diagnosis codes were assessed. A fused risk score was generated for each patient in the target dataset using our proposed targeted fusion framework. With this model, the averaged sensitivities at 90% and 95% specificity improved by 67% and 171%, and the positive predictive values for the combined fusion model improved 64% and 135% compared to the conventional model. DISCUSSION AND CONCLUSIONS: We proposed a general targeted fusion learning framework that can be used to build a tailored predictive model for any specific healthcare setting. Results from this study suggest we can improve the performance of predictive models in specific target settings without complete integration of the raw records from external data sources.
Wanwan Xu, Chang Su 0002, Steven Rogers, Fei Wang 0001, Kun Chen 0002, Robert H. Aseltine
J. Am. Medical Informatics Assoc.2
2021 Recent advances in biomedical literature mining
abstract
The recent years have witnessed a rapid increase in the number of scientific articles in biomedical domain. These literature are mostly available and readily accessible in electronic format. The domain knowledge hidden in them is critical for biomedical research and applications, which makes biomedical literature mining (BLM) techniques highly demanding. Numerous efforts have been made on this topic from both biomedical informatics (BMI) and computer science (CS) communities. The BMI community focuses more on the concrete application problems and thus prefer more interpretable and descriptive methods, while the CS community chases more on superior performance and generalization ability, thus more sophisticated and universal models are developed. The goal of this paper is to provide a review of the recent advances in BLM from both communities and inspire new research directions.
Sendong Zhao, Chang Su 0002, Zhiyong Lu, Fei Wang 0001
Briefings Bioinform.2
2021 Leveraging attention-based visual clue extraction for image classification
abstract
Abstract Deep learning‐based approaches have made considerable progress in image classification tasks, but most of the approaches lack interpretability, especially in revealing the decisive information causing the categorization of images. This paper seeks to answer the question of what clues encode the discriminative visual information between image categories and can help improve the classification performance. To this end, an attention‐based clue extraction network (ACENet) is introduced to mine the decisive local visual information for image classification. ACENet constructs a clue‐attention mechanism, that is global‐local attention, between the image and visual clue proposals extracted from it and then introduces a contrastive loss defined over the achieved discrete attention distribution to increase the discriminability of clue proposals. The loss encourages considerable attention to be devoted to discriminative clue proposals, that is those similar within the same category and dissimilar across categories. The experimental results for the Negative Web Image (NWI) dataset and the public ImageNet2012 dataset demonstrate that ACENet can extract true clues to improve the image classification performance and outperforms the baselines and the state‐of‐the‐art methods.
Yunbo Cui, Youtian Du, Chang Su 0002
IET Image Process.5
2021 Learning Fundamental Visual Concepts Based on Evolved Multi-Edge Concept Graph
abstract
In general, visual media comprises a set of elements of basic semantics, named fundamental visual concepts, that may not be semantically decomposed, such as objects, scenes and actions. This paper proposes a dynamic learning framework for fundamental visual concept learning from image-textual description paired data based on an evolved multi-edge concept graph (EMCG). First, we construct a multi-edge concept graph to represent the relationships between visual concept instances, in which we introduce two types of edges named visual edges and semantic edges to describe the connection strength in terms of visual appearance and semantic content. Second, we evolve the graph by updating connection strength based on the predicted results of concept learning. Finally, we present a growth algorithm for the multi-edge concept graph to handle cross-dataset concept learning. Driven by the predictions, the multi-edge concept graph can dynamically evolve over time by adjusting the connection strength to adapt better to the observations. In addition, our approach can be considered a weakly-supervised learning algorithm since no labeled concepts are employed for learning. Experimental results demonstrate that evolution can significantly improve the learning of fundamental visual concepts by$\text{14.2}\%$,$\text{7.9}\%$and$\text{12.7}\%$in terms of F1-score for the MSRC, VOC2012 and MSCOCO datasets, respectively, and that the proposed EMCG approach largely outperforms the compared approaches.
Youtian Du, Guangxun Zhang, Zhongmin Cai, Chang Su 0002
IEEE Trans. Multim.5
2020 Interactive Attention Networks for Semantic Text Matching
abstract
Semantic text matching, which matches target texts to source texts, is a general problem in many areas, such as information retrieval, question answering, and recommendation. The challenges to existing research on this topic include 1) out-of-vocabulary and low-frequency keywords and 2) direct utilization of sparse matching matrix of source and target. The out-of-vocabulary and low-frequency keywords could lead to the mismatch of similar keywords in source and target texts. The sparse matching matrix cannot provide enough clues to match the source with the target. To address these challenges, we propose a novel deep neural semantic text matching model. Our model adopts an interactive attention network to achieve information exchange between the source text and the target text, and dynamically explores the matching matrix and learns new representations of source and target texts. Experimental results on three different text matching datasets demonstrate that our model can significantly outperform competitive baselines. Furthermore, our model demonstrates great advantage in alleviating the sparse matching problem and learning out-of-vocabulary words with the local context, which widely exists in a broad spectrum of NLP applications.
Sendong Zhao, Chang Su 0002, Yuantong Li, Fei Wang 0001
ICDM3
2020 Network embedding in biomedical data science
abstract
Owning to the rapid development of computer technologies, an increasing number of relational data have been emerging in modern biomedical research. Many network-based learning methods have been proposed to perform analysis on such data, which provide people a deep understanding of topology and knowledge behind the biomedical networks and benefit a lot of applications for human healthcare. However, most network-based methods suffer from high computational and space cost. There remain challenges on handling high dimensionality and sparsity of the biomedical networks. The latest advances in network embedding technologies provide new effective paradigms to solve the network analysis problem. It converts network into a low-dimensional space while maximally preserves structural properties. In this way, downstream tasks such as link prediction and node classification can be done by traditional machine learning methods. In this survey, we conduct a comprehensive review of the literature on applying network embedding to advance the biomedical domain. We first briefly introduce the widely used network embedding models. After that, we carefully discuss how the network embedding approaches were performed on biomedical networks as well as how they accelerated the downstream tasks in biomedical science. Finally, we discuss challenges the existing network embedding applications in biomedical domains are faced with and suggest several promising future directions for a better improvement in human healthcare.
Chang Su 0002, Jie Tong, Yongjun Zhu 0001, Peng Cui 0001, Fei Wang 0001
Briefings Bioinform.1
2020 Kernel-Based Mixture Mapping for Image and Text Association
abstract
Modeling the relationship between multimodal media, including images, videos, and text, can reduce the gap between the modalities and promote cross-media retrieval, image annotation, etc. In this paper, we propose a new approach called kernel-based mixture mapping (KMM) to model the semantic correlations between web images and text. With this approach, we first construct latent high-dimensional feature spaces based on kernel theory to address the nonlinearity of both the data distributions in the input spaces and the cross-model correlation. Second, we present a probabilistic neighborhood model to describe the spatial locality of semantics by assuming that proximate examples in feature spaces generally have the same semantics and a conditional model to describe cross-modal conditional dependency. Finally, we build a probabilistic mixture model to jointly model the spatial locality of semantics and the conditional dependency between different modalities. By combining nonlinear transformation and probabilistic models, KMM can address the nonlinearity of cross-modal correlation, the complexity of semantic distributions at the global scale, and the continuity of semantic distributions at the local scale. We present a hybrid optimization algorithm to find the solution of KMM based on expectation-maximization and subgradient ascent; this algorithm avoids estimating the parameters of KMM in high-dimensional feature space and is proved to converge to an (local) optimal solution. We demonstrate the performance of KMM using four public datasets. The experimental results show that our approach outperforms the compared methods when modeling the relationships between images and text.
Youtian Du, Yunbo Cui, Chang Su 0002
IEEE Trans. Multim.5
2019 GRAPHENE: A Precise Biomedical Literature Retrieval Engine with Graph Augmented Deep Learning and External Knowledge Empowerment
abstract
Effective biomedical literature retrieval (BLR) plays a central role inprecision medicine informatics. In this paper, we propose GRAPHENE,which is a deep learning based framework for precise BLR. GRAPHENEconsists of three main different modules 1) graph-augmented doc-ument representation learning; 2) query expansion and represen-tation learning and 3) learning to rank biomedical articles. Thegraph-augmented document representation learning module con-structs a document-concept graph containing biomedical conceptnodes and document nodes so that global biomedical related con-cept from external knowledge source can be captured, which isfurther connected to a BiLSTM so both local and global topics canbe explored. Query expansion and representation learning moduleexpands the query with abbreviations and different names, and thenbuilds a CNN-based model to convolve the expanded query andobtain a vector representation for each query. Learning to rank min-imizes a ranking loss between biomedical articles with the queryto learn the retrieval function. Experimental results on applyingour system to TREC Precision Medicine track data are provided todemonstrate its effectiveness.
Sendong Zhao, Chang Su 0002, Andrea Sboner, Fei Wang 0001
CIKM2
2018 Toward capturing heterogeneity for inferring diffusion networks: A mixed diffusion pattern model
Chang Su 0002, Xiaohong Guan, Youtian Du, Minhua Zhang
Knowl. Based Syst.1
2013 Maximizing topic propagation driven by multiple user nodes in micro-blogging
abstract
This work investigates the maximization of topic propagation jointly driven by multiple user nodes in micro-blogging. In this paper, we propose a new method to find a set of user nodes that jointly propagate topics approximately the most widely. First, we obtain multiple nodes with strong influence; Second, we exactly compute the breadth of information spread driven by a single node based on probabilistic models; Finally, we analyze the information propagation jointly driven by multiple nodes and derive an approximately optimal set of driving nodes. We find that the breadth of information propagation jointly driven by multiple nodes is approximately linear with both the breadth of information propagation of single driving nodes and the strength of tie among them, which indicates that selecting the optimal driving nodes needs to consider the link information among them as well as the ability of each node. Experimental results demonstrate the effectiveness of our method.
Chang Su 0002, Youtian Du, Xiaohong Guan, Chenhe Wu
LCN1