Xian Yang 0001

dblp:25/10624-1 · DBLP profile ↗
← Back
38ranked-venue papers
3as first author
31since 2021 · last 2026
0000-0002-1496-8923ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Medical Vision-Language Pretraining with LLM-Guided Temporal Supervision
abstract
Medical vision–language pretraining typically relies on static image–text pairs, overlooking temporal cues vital for understanding clinical progression. This limits model sensitivity to evolving semantics and reduces their effectiveness in real-world clinical reasoning. To address this challenge, we propose TAMM—a temporal alignment framework that leverages weak but semantically rich supervision from large language models (LLMs). Given temporally adjacent clinical reports, LLMs automatically generate (i) coarse-grained trend labels (e.g., improving or worsening), and (ii) fine-grained rationales explaining the supporting clinical evidence. These complementary signals inject temporal semantics without requiring manual annotation, and guide vision–language representation learning to capture trend-sensitive cross-modal alignment and rationale-grounded coherence. Experiments on multiple medical benchmarks demonstrate that TAMM improves retrieval and classification performance while yielding more interpretable, temporally consistent embeddings. Our results highlight the potential of leveraging LLM-derived supervision to equip vision–language models with temporal awareness critical for clinical applications.
Liang Bai 0001, Huimin Yan, Xian Yang 0001
AAAI4
2026 Attribute-guided Dynamic Prompt Learning for Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have achieved remarkable success in analyzing graph-structured data, with their performance dependent on the graph structure. However, models trained on high-quality graph structures often suffer a significant performance drop when evaluated on perturbed graphs. Existing methods tackle this problem by improving the robustness of GNNs, but they often overlook representation deviation caused by structural changes. To address this limitation, we propose an attribute-guided dynamic prompt learning model that generates prompt vectors to approximate the intrinsic information of nodes. With these prompt vectors, the trained GNNs are expected to maintain their performance under perturbed graph structures. Unlike previous prompt-based methods that learn unified prompt vectors for all nodes, we obtain node-level prompts by encoding node attributes that provide unique information. Given the diversity of perturbed graph structures during inference, we introduce a structure-aware adaptation mechanism that adjusts the prompt vectors based on the input graph. Furthermore, we apply gradient-based attacks to generate perturbed graphs, encouraging the model to generalize to unseen structures. Extensive experiments on multiple benchmark datasets demonstrate the effectiveness and robustness of our model.
Zhuomin Liang, Liang Bai 0001, Xian Yang 0001
AAAI3
2026 CauVQ: Causal Vector Quantization for Graph OOD Generalization
abstract
Graph Neural Networks (GNNs) perform well on in-distribution data but often fail under out-of-distribution (OOD) shifts due to reliance on spurious patterns. To address this, we propose CauVQ, a causal vector quantization framework that improves OOD generalization by identifying and leveraging invariant substructures that are causally predictive. To construct stable and symbolic graph representations, CauVQ decomposes each input into local substructures and maps them to a discrete codebook of prototypical motifs. This enables consistent and interpretable encoding across diverse graph domains. To isolate the causal substructures, we maximize their mutual information with graph labels and refine their representations using a learnable interaction matrix and a causal attention mechanism. Furthermore, we introduce a counterfactual regularization strategy to enforce prediction stability under substructure perturbations, encouraging the model to focus on truly causal patterns rather than superficial shortcuts. Extensive experiments across standard and OOD benchmarks demonstrate that CauVQ consistently outperforms state-of-the-art baselines in robustness and interpretability. Our framework offers a promising step toward reliable, explainable, and distribution-aware graph learning.
Liang Bai 0001, Hangyuan Du, Xian Yang 0001
AAAI4
2026 Adaptive Evolutionary Fusion for Multi-View Clustering
abstract
Deep multi-view clustering (MVC) methods achieve impressive performance by effectively capturing complementary information across views, where feature fusion serves as the critical mechanism for maximizing cross-view complementarity. However, most existing methods suffer from rigid dependence on non-adaptive predefined fusion operations, resulting in unverifiable and potentially suboptimal fused feature quality. To resolve these limitations, we propose a novel multi-view clustering framework that learns adaptive hierarchical fusion through an unsupervised evolutionary algorithm. Unlike conventional predefined-fusion strategies, our approach employs tree-structured representations (Fusion Trees) for adaptive feature integration. These Fusion Trees are optimized via our evolutionary mechanism, in which models sharing identical architectures but distinct Fusion Trees are conceptualized as evolutionary individuals. Through implementation of the evolutionarily optimized Fusion Tree, the resultant model generates discriminative representations in accordance with biological evolutionary principles. Comprehensive benchmarking across twelve multi-view datasets validates significant performance gains improvement over state-of-the-art baselines.
Yunxiao Zhao, Liang Bai 0001, Xian Yang 0001
AAAI3
2026 Expanding Domain Generalization Theory: Error Bound Beyond Convex Combination Assumption
Liang Bai 0001, Xian Yang 0001, Jiye Liang
Mach. Learn.3
2026 Asymmetric Co-Training With Decoder-Head Decoupling for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning reduces annotation costs in medical image segmentation by leveraging abundant unlabeled data alongside scarce labels. Most models adopt an encoder-decoder architecture with a task-specific segmentation head. While co-training is effective, existing frameworks suffer from intra-network coupling (decoder-head binding) and inter-network coupling (over-aligned predictions), which reduce prediction diversity and amplify confirmation bias-particularly for small structures, ambiguous boundaries, and anatomically variable regions. We propose AsyCo, an asymmetric co-training framework with two components. (1) Asymmetric Decoder Coupling implements decoder-head decoupling by dynamically remapping encoder-decoder features to non-default heads across branches, breaking intra-network coupling and creating diverse prediction paths without additional parameters. (2) Hierarchical Consistency Regularization converts this diversity into stable supervision by aligning (i) the two branches' final outputs along their default paths (branch-output consistency), (ii) predictions from different segmentation heads evaluated on identical decoder features (inter-head consistency), and (iii) intermediate encoder-decoder representations (representation consistency). Through these mechanisms, AsyCo explicitly mitigates both intra- and inter-network coupling, improving training stability and reducing confirmation bias. Extensive experiments on three clinical benchmarks under limited-label regimes demonstrate that AsyCo consistently outperforms nine state-of-the-art semi-supervised learning methods. These results indicate that AsyCo delivers accurate and reliable segmentation with minimal annotation, thereby enhancing the reliability of medical image analysis in real-world clinical practice.
Muhan Shi, Min Qu, Yinxue Shi, Xian Yang 0001
IEEE J. Biomed. Health Informatics7
2026 Graph-Enhanced Visual Prompting for Pre-Trained Models Adaptation in Medical Imaging Classification
abstract
Adapting Vision Transformers (ViTs) for medical imaging is constrained by the scarcity of data and high-quality annotations, hindering effective training and robust generalization. Visual prompt learning offers a parameter-efficient solution for domain adaptation, but its success depends on accurate and task-relevant semantic guidance-a resource rarely available in real-world clinical practice despite its proven benefits. This motivates the need for mechanisms that can automatically extract reliable semantic cues from existing clinical data. To this end, we propose Graph-Enhanced Visual Prompting (GEVP), the first framework to incorporate cross-modal graph learning into prompt generation for medical imaging. GEVP models image patches and report tokens as graph nodes, captures their spatial and semantic relations via a graph neural network, and produces semantically rich prompts. These prompts are injected into a frozen ViT backbone, guiding attention to diagnostically relevant regions without heavy fine-tuning. A consistent downstream prediction mechanism leverages the pretrained prompt generator to handle both report-available and report-absent settings. Experiments on six public downstream datasets show GEVP surpasses strong prompt- and adapter-based baselines by up to +9.65% F1 on imbalanced tasks and delivers superior unseen disease classification.
Liang Bai 0001, Xian Yang 0001, Jiye Liang
IEEE Trans. Medical Imaging3
2025 Knowledge-Augmented Multimodal Clinical Rationale Generation for Disease Diagnosis with Small Language Models
abstract
Shuai Niu, Jing Ma, Hongzhan Lin, Liang Bai, Zhihua Wang, Richard Yi Da Xu, Yunya Song, Xian Yang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jing Ma 0004, Hongzhan Lin 0001, Liang Bai 0001, Zhihua Wang 0008, Yunya Song, Xian Yang 0001
ACL (1)8
2025 Endo-CLIP: Progressive Self-supervised Pre-training on Raw Colonoscopy Records
Yili He, Peiyao Fu, Ruijie Yang, Zhihua Wang 0008, Quanlin Li, Pinghong Zhou, Xian Yang 0001, Shuo Wang 0011
MICCAI (11)9
2025 Multi-Channel Disentangled Graph Neural Networks With Different Types of Self-Constraints
abstract
Graph Neural Network (GNN) is a popular semi-supervised graph representation learning method, whose performance strongly relies on the quality and quantity of labeled nodes. Given the insufficiency of labeled nodes in many real applications, many multi-channel GNNs have been developed to extract self-supervised information by leveraging consistency and complementarity among augmented graphs from different channels. However, these methods often struggle to balance conflicting self-supervised constraints, enhancing certain types of information at the expense of others. To tackle this problem, we propose a Multi-channel Disentangled Graph Neural Network (MD-GraphNet), which effectively classifies self-supervised constraints by learning disentangled representations. Specifically, our model enforces consistency constraints for shared representations, graph reconstruction constraints for complementary (or private) representations, and aligning constraints for fused representations. Our model overcomes the confusion and loss problems of different types of self-supervised signals. Experimental results on benchmark datasets demonstrate the effectiveness of MD-GraphNet for semi-supervised node classification.
Zhuomin Liang, Liang Bai 0001, Xian Yang 0001, Jiye Liang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Label-Semantic-Based Prompt Tuning for Vision Transformer Adaptation in Medical Image Analysis
abstract
Adapting Vision Transformers (ViTs) to medical image analysis is challenging due to the scarcity of annotated data and the significant domain shift from natural to medical images. Traditional fine-tuning approaches, while effective, require storing separate model parameters for each task, leading to high computational costs. Existing prompt tuning methods reduce this overhead by introducing task-specific prompt tokens, but they often fail to fully leverage label semantics, resulting in suboptimal performance for medical tasks. To address these limitations, we propose a label-semantic-based prompt tuning method (LPT), which transforms the visual prompt learning problem into a text-image alignment task. Unlike traditional prompt methods that only focus on visual prompts, LPT incorporates label semantics through a cross-attention-based module to better align image features with the target labels. This approach not only captures rich semantic information from the labels but also enhances the model’s ability to extract fine-grained image details relevant to specific medical conditions. By leveraging label-text alignment during training, LPT improves both label utilization and model adaptability, enabling more accurate predictions. Extensive experiments on eight diverse medical datasets demonstrate that LPT significantly improves diagnostic accuracy and generalization, outperforming both traditional fine-tuning and current prompt-based methods, especially in data-limited scenarios.
Liang Bai 0001, Xian Yang 0001, Jiye Liang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Contrastive Learning With Enhancing Detailed Information for Pre-Training Vision Transformer
abstract
Contrastive Learning (CL) is an effective self-supervised learning method. It performs instance-level contrastiveness based on the image representations, which enables the model to extract abstract information from images. However, when training data is insufficient, abstract information fails to distinguish samples from different classes. This problem is more severe in the pre-training of Vision Transformer (ViT). In general, detailed information is crucial for enhancing the discrimination of representations. Patch representations, which focus on the details of images, are often overlooked in existing methods that train ViT through CL, resulting in the confusion of similar samples. To address this problem, we propose a Contrastive Learning model with Enhancing Detailed Information (CL-EDI) for pre-training ViT. Our model consists of dual ViT contrastive modules. The first module is similar to MoCo V3, which can learn abstract information about images. The role of the second ViT contrastive module is to enhance detailed information in data representations by aggregating patch representations of images. Extensive experiments demonstrate the necessity of learning detailed information. Across several datasets, our model surpasses existing approaches in image classification, transfer learning and object detection tasks.
Zhuomin Liang, Liang Bai 0001, Jinyu Fan, Xian Yang 0001, Jiye Liang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Local Alignment for Medical Vision-Language Pre-Training
abstract
Establishing local semantic correspondences between medical images and their corresponding reports is crucial for effective medical vision-language pre-training. However, existing methods encounter two major challenges: (1) lesion regions in radiological images are often small, blurry, or lack clear boundaries, complicating accurate localization; and (2) medical reports typically contain redundant or non-diagnostic words, hindering precise semantic alignment. To overcome these issues, we propose MedAligner, a specialized local alignment network for medical vision-language pre-training. MedAligner employs dual encoders to extract both global and local representations and uses global contrastive learning to maintain coarse semantic consistency. To enhance local alignment, we introduce a Word-Region Alignment, which generates a learnable word-pixel similarity matrix that is sparsified to identify salient lesion regions accurately. Additionally, our Diagnostic Term Filtering dynamically samples high-importance diagnostic terms from reports, aligning them with identified lesion areas via a local contrastive loss. Importantly, we adopt a progressive training strategy that gradually refines both the input text and semantic alignment. This is achieved by reconstructing concise diagnostic reports and progressively updating word-pixel similarity, generating increasingly accurate image-text pairs. Extensive experiments demonstrate that MedAligner significantly surpasses existing approaches on tasks such as phrase grounding, image-text retrieval, and zero-shot classification, setting new benchmarks in medical vision-language pre-training.
Huimin Yan, Xian Yang 0001, Liang Bai 0001, Jiye Liang
IEEE Trans. Image Process.2
2025 Robust Polyp Detection and Diagnosis Through Compositional Prompt-Guided Diffusion Models
abstract
Colorectal cancer (CRC) is a significant global health concern, and early detection through screening plays a critical role in reducing mortality. While deep learning models have shown promise in improving polyp detection, classification, and segmentation, their generalization across diverse clinical environments, particularly with out-of-distribution (OOD) data, remains a challenge. Multi-center datasets like PolypGen have been developed to address these issues, but their collection is costly and time-consuming. Traditional data augmentation techniques provide limited variability, failing to capture the complexity of medical images. Diffusion models have emerged as a promising solution for generating synthetic polyp images, but the image generation process in current models mainly relies on segmentation masks as the condition, limiting their ability to capture the full clinical context. To overcome these limitations, we propose a Progressive Spectrum Diffusion Model (PSDM) that integrates diverse clinical annotations-such as segmentation masks, bounding boxes, and colonoscopy reports-by transforming them into compositional prompts. These prompts are organized into coarse and fine components, allowing the model to capture both broad spatial structures and fine details, generating clinically accurate synthetic images. By augmenting training data with PSDM-generated samples, our model significantly improves polyp detection, classification, and segmentation. For instance, on the PolypGen dataset, PSDM increases the F1 score by 2.12% and the mean average precision by 3.09%, demonstrating superior performance in OOD scenarios and enhanced generalization.
Peiyao Fu, Junbo Huang, Quanlin Li, Pinghong Zhou, Zhihua Wang 0008, Fei Wu 0001, Shuo Wang 0011, Xian Yang 0001
IEEE Trans. Medical Imaging11
2025 Graph Contrastive Learning for Fusion of Graph Structure and Attribute Information
abstract
Graph Contrastive Learning (GCL) plays a crucial role in multimedia applications due to its effectiveness in analyzing graph-structured data. Existing GCL methods focus on maximizing the agreement of node representations across different augmentations, which leads to the neglect of unique and complementary information in each augmentation. In this paper, we propose a fusion-based GCL model (FB-GCL) that learns fused representations to effectively capture complementary information from both the graph structure and node attributes. Our model consists of two modules: a graph fusion encoder and a graph contrastive module. The graph fusion encoder adaptively fuses the representations learned from the topology graph and the attribute graph. The graph contrastive module extracts supervision signals from the raw graph by leveraging both the pairwise relationships within the graph structure and the multi-label information from the attributes. Extensive experiments on seven benchmark datasets demonstrate that FB-GCL enhances performance in node classification and link prediction tasks. This improvement is especially valuable for multimedia data analysis, as integrating graph structure and attribute information is crucial for effectively understanding and processing complex datasets.
Zhuomin Liang, Liang Bai 0001, Xian Yang 0001, Jiye Liang
IEEE Trans. Multim.3
2025 Multi-Grained Vision-and-Language Model for Medical Image and Text Alignment
abstract
The increasing interest in learning from paired medical images and textual reports highlights the need for methods that can achieve multi-grained alignment between these two modalities. However, most existing approaches overlook finegrained semantic alignment, which can constrain the quality of the generated representations. To tackle this problem, we propose the Multi-Grained Vision-and-Language Alignment (MGVLA) model, which effectively leverages multi-grained correspondences between medical images and texts at different levels, including disease, instance, and token levels. For disease-level alignment, our approach adopts the concept of contrastive learning and uses medical terminologies detected from textual reports as soft labels to guide the alignment process. At the instance level, we propose a strategy for sampling hard negatives, where images and texts with the same disease type but differing in details such as disease locations and severity are considered as hard negatives. This strategy helps our approach to better distinguish between positive and negative image-text pairs, ultimately enhancing the quality of our learned representations. For token-level alignment, we employ a masking and recovery technique to achieve finegrained semantic alignment between patches and sub-words. This approach effectively aligns the different levels of granularity between the image and language modalities. To assess the efficacy of our MGVLA model, we conduct comprehensive experiments on the image-text retrieval and phrase grounding tasks.
Huimin Yan, Xian Yang 0001, Liang Bai 0001, Jiamin Li 0006, Jiye Liang
IEEE Trans. Multim.2
2025 Improving Image Contrastive Clustering Through Self-Learning Pairwise Constraints
abstract
In this article, a new unsupervised contrastive clustering (CC) model is introduced, namely, image CC with self-learning pairwise constraints (ICC-SPC). This model is designed to integrate pairwise constraints into the CC process, enhancing the latent representation learning and improving clustering results for image data. The incorporation of pairwise constraints helps reduce the impact of false negatives and false positives in contrastive learning, while maintaining robust cluster discrimination. However, obtaining prior pairwise constraints from unlabeled data directly is quite challenging in unsupervised scenarios. To address this issue, ICC-SPC designs a pairwise constraints learning module. This module autonomously learns pairwise constraints among data samples by leveraging consensus information between latent representation and pseudo-labels, which are generated by the clustering algorithm. Consequently, there is no requirement for labeled images, offering a practical resolution to the challenge posed by the lack of sufficient supervised information in unsupervised clustering tasks. ICC-SPC's effectiveness is validated through evaluations on multiple benchmark datasets. This contribution is significant, as we present a novel framework for unsupervised clustering by integrating contrastive learning with self-learning pairwise constraints.
Yecheng Guo, Liang Bai 0001, Xian Yang 0001, Jiye Liang
IEEE Trans. Neural Networks Learn. Syst.3
2024 EndoFinder: Online Image Retrieval for Explainable Colorectal Polyp Diagnosis
Ruijie Yang, Peiyao Fu, Yizhe Zhang 0001, Zhihua Wang 0008, Quanlin Li, Pinghong Zhou, Xian Yang 0001, Shuo Wang 0011
MICCAI (10)8
2024 Enhancing healthcare decision support through explainable AI models for risk prediction
abstract
Electronic health records (EHRs) are a valuable source of information that can aid in understanding a patient’s health condition and making informed healthcare decisions. However, modelling longitudinal EHRs with heterogeneous information is a challenging task. Although recurrent neural networks (RNNs), which are current artificial intelligence (AI) models, have the capability to capture longitudinal information, their explanatory power is limited. Predictive clustering is a recent development in this field, which provides cluster-level explainable evidence for disease risk prediction. Nonetheless, the challenge of determining the optimal number of clusters has put a brake on the widespread application of predictive clustering for disease risk prediction. In this paper, we introduce a novel non-parametric predictive clustering-based risk prediction model that integrates the Dirichlet Process Mixture Model (DPMM) with predictive clustering via neural networks. To enhance the model’s interpretability, we integrate attention mechanisms that enable the capture of local-level evidence in addition to the cluster-level evidence provided by predictive clustering. The outcome of this research is the development of a multi-level explainable artificial intelligence (AI) model. We evaluated the proposed model on two real-world datasets and demonstrated its effectiveness in capturing longitudinal EHR information for disease risk prediction. Additionally, the model was successful in generating explainable evidence to support its predictions.
Qing Yin, Jing Ma 0004, Yunya Song, Liang Bai 0001, Wei Pan 0004, Xian Yang 0001
Decis. Support Syst.8
2024 A Data Dissemination Algorithm Based on Maximization of Causal Path Entropy in Vehicular Ad Hoc Networks
abstract
Efficient data dissemination protocols play a crucial role in transmitting messages among vehicles in vehicular ad hoc networks (VANETs). In scenarios where the destination nodes in VANETs are unknown, it becomes imperative to disseminate the data across as many locations as possible, thereby enhancing the likelihood of reaching potential destinations. While multihop broadcasting facilitates rapid data coverage in networks, it often results in challenges, such as broadcast storms, data redundancy, and transmission delays. This work addresses these issues from the perspective of maximizing the diversity of system evolution, consequently increasing the probability of receiving data by potential destination nodes. To achieve this objective, an efficient data dissemination algorithm causal entropy increase-based dissemination (CEID) is proposed based on causal path entropy theory. By modeling the topology of urban VANETs and calculating the degrees of nodes centrality, the algorithm selects the relay nodes that can increase the diversity of the system during the data dissemination process. This approach aims to maximize the increase of causal path entropy, allowing as many vehicles as possible to receive data quickly while reducing redundancy and collisions during the data dissemination. In addition, a multidirectional synchronous data dissemination strategy is designed for the environment of urban road intersections to achieve fast and efficient data dissemination in the network. Simulation experiments show the superiority of CEID in terms of system entropy increase rate, data transmission delay, and data redundancy when compared to existing algorithms.
Dan Gu, Yajie Ma 0001, Feng Dan, Xian Yang 0001, Fengxing Zhou, Baokang Yan, Shaowu Lu, Bowen Ning
IEEE Internet Things J.4
2024 Enhancing Drug Recommendations Via Heterogeneous Graph Representation Learning in EHR Networks
abstract
Electronic health records (EHRs) contain vast medical information like diagnosis, medication, and procedures, enabling personalized drug recommendations and treatment adjustments. However, current drug recommendation methods only model patients' health conditions from EHR data, neglecting the rich relationships within the data. This paper seeks to utilize a heterogeneous information network (HIN) to represent EHR and develop a graph representation learning method for medication recommendation. However, three critical issues need to be investigated: (1) co-occurrence of diagnosis and drug for the same patient does not imply their relevance; (2) patients' directly associated information may not be sufficient to reflect their health conditions; and (3) the cold start problem exists when patients have no historical EHRs. To tackle these challenges, we develop a bi-channel heterogeneous local structural encoder to decouple and extract the diverse information in HIN. Additionally, a global information capture and fusion module, aggregating meta-paths to form a global representation, is introduced to fill the information gaps in records. A longitudinal model using rich structural information available in EHR data is proposed for drug recommendations to new patients. Experimental results on real-world EHR data demonstrate significant improvements over existing approaches.
Xian Yang 0001, Liang Bai 0001, Jiye Liang
IEEE Trans. Knowl. Data Eng.2
2023 RTANet: Recommendation Target-Aware Network Embedding
abstract
Network embedding is a process of encoding nodes into latent vectors by preserving network structure and content information. It is used in various applications, especially in recommender systems. In a social network setting, when recommending new friends to a user, the similarity between the user's embedding and the target friend will be examined. Traditional methods generate user node embedding without considering the recommendation target. No matter which target is to be recommended, the same embedding vector is generated for that particular user. This approach has its limitations. For example, a user can be both a computer scientist and a musician. When recommending music friends with potentially the same taste to him, we are interested in getting his representation that is useful in recommending music friends rather than computer scientists. His corresponding embedding should consider the user's musical features rather than those associated with computer science with the awareness that the recommendation targets are music friends. In order to address this issue, we propose a new framework which we name it as Recommendation Target-Aware Network embedding method (RTANet). Herein, the embedding of each user is no longer fixed to a constant vector, but it can vary according to their specific recommendation target. Concretely, RTANet assigns different attention weights to each neighbour node, allowing us to obtain the user's context information aggregated from its neighbours before transforming this context into its embedding. Different from other graph attention approaches, the attention weights in our work measure the similarity between each user's neighbour node and the target node, which in return generates the target-aware embedding. To demonstrate the effectiveness of our method, we compared RTANet with several state-of-the-art network embedding methods on four real-world datasets and showed that RTANet outperforms other comparative methods in the recommendation tasks.
Qimeng Cao, Qing Yin, Yunya Song, Zhihua Wang 0008, Yujun Chen, Xian Yang 0001
ICWSM7
2023 High-order graph attention network
Liancheng He, Liang Bai 0001, Xian Yang 0001, Hangyuan Du, Jiye Liang
Inf. Sci.3
2023 A new contrastive learning framework for reducing the effect of hard negatives
Liang Bai 0001, Xian Yang 0001, Jiye Liang
Knowl. Based Syst.3
2023 Exploring the role of edge distribution in graph convolutional networks
Liancheng He, Liang Bai 0001, Xian Yang 0001, Zhuomin Liang, Jiye Liang
Neural Networks3
2022 Improving Deep Embedded Clustering via Learning Cluster-level Representations
abstract
Driven by recent advances in neural networks, various Deep Embedding Clustering (DEC) based short text clustering models are being developed. In these works, latent representation learning and text clustering are performed simultaneously. Although these methods are becoming increasingly popular, they use pure cluster-oriented objectives, which can produce meaningless representations. To alleviate this problem, several improvements have been developed to introduce additional learning objectives in the clustering process, such as models based on contrastive learning. However, existing efforts rely heavily on learning meaningful representations at the instance level. They have limited focus on learning global representations, which are necessary to capture the overall data structure at the cluster level. In this paper, we propose a novel DEC model, which we named the deep embedded clustering model with cluster-level representation learning (DECCRL) to jointly learn cluster and instance level representations. Here, we extend the embedded topic modelling approach to introduce reconstruction constraints to help learn cluster-level representations. Experimental results on real-world short text datasets demonstrate that our model produces meaningful clusters.
Qing Yin, Zhihua Wang 0008, Yunya Song, Liang Bai 0001, Yike Guo, Xian Yang 0001
COLING8
2022 Bayesian data assimilation for estimating instantaneous reproduction numbers during epidemics: Applications to COVID-19
abstract
Estimating the changes of epidemiological parameters, such as instantaneous reproduction number, Rt, is important for understanding the transmission dynamics of infectious diseases. Current estimates of time-varying epidemiological parameters often face problems such as lagging observations, averaging inference, and improper quantification of uncertainties. To address these problems, we propose a Bayesian data assimilation framework for time-varying parameter estimation. Specifically, this framework is applied to estimate the instantaneous reproduction number Rt during emerging epidemics, resulting in the state-of-the-art 'DARt' system. With DARt, time misalignment caused by lagging observations is tackled by incorporating observation delays into the joint inference of infections and Rt; the drawback of averaging is overcome by instantaneously updating upon new observations and developing a model selection mechanism that captures abrupt changes; the uncertainty is quantified and reduced by employing Bayesian smoothing. We validate the performance of DARt and demonstrate its power in describing the transmission dynamics of COVID-19. The proposed approach provides a promising solution for making accurate and timely estimation for transmission dynamics based on reported data.
Xian Yang 0001, Shuo Wang 0011, Yuting Xing, Ling Li 0010, Karl J. Friston, Yike Guo
PLoS Comput. Biol.1
2021 Label-dependent and event-guided interpretable disease risk prediction using EHRs
abstract
Electronic health records (EHRs) contain patients’ heterogeneous data that are collected from medical providers involved in the patient’s care, including medical notes, clinical events, laboratory test results, symptoms, and diagnoses. In the field of modern healthcare, predicting whether patients would experience any risks based on their EHRs has emerged as a promising research area, in which artificial intelligence (AI) plays a key role. To make AI models practically applicable, it is required that the prediction results should be both accurate and interpretable. To achieve this goal, this paper proposed a label-dependent and event-guided risk prediction model (LERP) to predict the presence of multiple disease risks by mainly extracting information from unstructured medical notes. Our model is featured in the following aspects. First, we adopt a label-dependent mechanism that gives greater attention to words from medical notes that are semantically similar to the names of risk labels. Secondly, as the clinical events (e.g., treatments and drugs) can also indicate the health status of patients, our model utilizes the information from events and uses them to generate an event-guided representation of medical notes. Thirdly, both label-dependent and event-guided representations are integrated to make a robust prediction, in which the interpretability is enabled by the attention weights over words from medical notes. To demonstrate the applicability of the proposed method, we apply it to the MIMIC-III dataset, which contains real-world EHRs collected from hospitals. Our method is evaluated in both quantitative and qualitative ways.
Yunya Song, Qing Yin, Yike Guo, Xian Yang 0001
BIBM5
2021 Self-Supervised Detection of Contextual Synonyms in a Multi-Class Setting: Phenotype Annotation Use Case
abstract
Jingqing Zhang, Luis Bolanos Trujillo, Tong Li, Ashwani Tanwar, Guilherme Freire, Xian Yang, Julia Ive, Vibhor Gupta, Yike Guo. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Jingqing Zhang, Luis Bolanos, Ashwani Tanwar, Guilherme Freire, Xian Yang 0001, Julia Ive, Vibhor Gupta, Yike Guo
EMNLP (1)6
2021 Label Dependent Attention Model for Disease Risk Prediction Using Multimodal Electronic Health Records
abstract
Disease risk prediction has attracted increasing attention in the field of modern healthcare, especially with the latest advances in artificial intelligence (AI). Electronic health records (EHRs), which contain heterogeneous patient information, are widely used in disease risk prediction tasks. One challenge of applying AI models for risk prediction lies in generating interpretable evidence to support the prediction results while retaining the prediction ability. In order to address this problem, we propose the method of jointly embedding words and labels whereby attention modules learn the weights of words from medical notes according to their relevance to the names of risk prediction labels. This approach boosts interpretability by employing an attention mechanism and including the names of prediction tasks in the model. However, its application is only limited to the handling of textual inputs such as medical notes. In this paper, we propose a label dependent attention model LDAM to 1) improve the interpretability by exploiting Clinical-BERT (a biomedical language model pre-trained on a large clinical corpus) to encode biomedically meaningful features and labels jointly; 2) extend the idea of joint embedding to the processing of timeseries data, and develop a multi-modal learning framework for integrating heterogeneous information from medical notes and time-series health status indicators. To demonstrate our method, we apply LDAM to the MIMIC-III dataset to predict different disease risks. We evaluate our method both quantitatively and qualitatively. Specifically, the predictive power of LDAM will be shown, and case studies will be carried out to illustrate its interpretability.
Qing Yin, Yunya Song, Yike Guo, Xian Yang 0001
ICDM5
2021 Effective low capacity status prediction for cloud systems
abstract
In cloud systems, an accurate capacity planning is very important for cloud provider to improve service availability. Traditional methods simply predicting "when the available resources is exhausted" are not effective due to customer demand fragmentation and platform allocation constraints. In this paper, we propose a novel prediction approach which proactively predicts the level of resource allocation failures from the perspective of low capacity status. By jointly considering the data from different sources in both time series form and static form, the proposed approach can make accurate LCS predictions in a complex and dynamic cloud environment, and thereby improve the service availability of cloud systems. The proposed approach is evaluated by real-world datasets collected from a large scale public cloud platform, and the results confirm its effectiveness.
Hang Dong 0004, Si Qin, Yong Xu 0010, Bo Qiao 0001, Shandan Zhou, Xian Yang 0001, Chuan Luo 0002, Pu Zhao 0004, Qingwei Lin, Hongyu Zhang 0002, Abulikemu Abuduweili, Sanjay Ramanujan, Karthikeyan Subramanian, Andrew Zhou, Saravanakumar Rajmohan, Dongmei Zhang 0001, Thomas Moscibroda
ESEC/SIGSOFT FSE6
2020 Identifying linked incidents in large-scale online service systems
abstract
In large-scale online service systems, incidents occur frequently due to a variety of causes, from updates of software and hardware to changes in operation environment. These incidents could significantly degrade system’s availability and customers’ satisfaction. Some incidents are linked because they are duplicate or inter-related. The linked incidents can greatly help on-call engineers find mitigation solutions and identify the root causes. In this work, we investigate the incidents and their links in a representative real-world incident management (IcM) system. Based on the identified indicators of linked incidents, we further propose LiDAR (Linked Incident identification with DAta-driven Representation), a deep learning based approach to incident linking. More specifically, we incorporate the textual description of incidents and structural information extracted from historical linked incidents to identify possible links among a large number of incidents. To show the effectiveness of our method, we apply our method to a real-world IcM system and find that our method outperforms other state-of-the-art methods.
Yujun Chen, Xian Yang 0001, Hang Dong 0004, Xiaoting He 0003, Hongyu Zhang 0002, Qingwei Lin, Junjie Chen 0003, Pu Zhao 0004, Yu Kang 0006, Feng Gao 0022, Zhangwei Xu, Dongmei Zhang 0001
ESEC/SIGSOFT FSE2
2019 Unsupervised Annotation of Phenotypic Abnormalities via Semantic Latent Representations on Electronic Health Records
abstract
The extraction of phenotype information which is naturally contained in electronic health records (EHRs) has been found to be useful in various clinical informatics applications such as disease diagnosis. However, due to imprecise descriptions, lack of gold standards and the demand for efficiency, annotating phenotypic abnormalities on millions of EHR narratives is still challenging. In this work, we propose a novel unsupervised deep learning framework to annotate the phenotypic abnormalities from EHRs via semantic latent representations. The proposed framework takes the advantage of Human Phenotype Ontology (HPO), which is a knowledge base of phenotypic abnormalities, to standardize the annotation results. Experiments have been conducted on 52,722 EHRs from MIMIC-III dataset. Quantitative and qualitative analysis have shown the proposed framework achieves state-of-the-art annotation performance and computational efficiency compared with other methods.
Jingqing Zhang, Xiaoyu Zhang 0008, Kai Sun 0005, Xian Yang 0001, Chengliang Dai, Yike Guo
BIBM4
2019 Integrated Multi-omics Analysis Using Variational Autoencoders: Application to Pan-cancer Classification
abstract
Omics data are normally high dimensional with large number of molecular features and relatively small number of available samples with clinical labels. The “curse of dimensionality” makes it challenging to train a machine learning model using high dimensional omics data like DNA methylation and gene expression profiles. Here we propose an end-to-end deep learning model called OmiVAE to extract low dimensional features and classify samples from multi-omics data. OmiVAE combines the basic structure of variational autoencoders with a classifier to achieve task-oriented feature extraction and multi-class classification. The training procedure of OmiVAE is comprised of an unsupervised phase and a supervised phase. During the unsupervised phase, a hierarchical cluster structure of samples can be automatically formed without the need for labels. And in the supervised phase, OmiVAE achieved an average accuracy of 97.49% after 10-fold cross-validation among 33 tumour types and normal samples, which shows better performance than existing methods. The integrated model learned from multi-omics datasets outperformed those using only one type of omics data, which indicates that the complementary information from different omics datatypes provides useful insights for biomedical tasks like cancer classification.
Xiaoyu Zhang 0008, Jingqing Zhang, Kai Sun 0005, Xian Yang 0001, Chengliang Dai, Yike Guo
BIBM4
2019 Outage Prediction and Diagnosis for Cloud Service Systems
abstract
With the rapid growth of cloud service systems and their increasing complexity, service failures become unavoidable. Outages, which are critical service failures, could dramatically degrade system availability and impact user experience. To minimize service downtime and ensure high system availability, we develop an intelligent outage management approach, called AirAlert, which can forecast the occurrence of outages before they actually happen and diagnose the root cause after they indeed occur. AirAlert works as a global watcher for the entire cloud system, which collects all alerting signals, detects dependency among signals and proactively predicts outages that may happen anywhere in the whole cloud system. We analyze the relationships between outages and alerting signals by leveraging Bayesian network and predict outages using a robust gradient boosting tree based classification method. The proposed outage management approach is evaluated using the outage dataset collected from a Microsoft cloud system and the results confirm the effectiveness of the proposed approach.
Yujun Chen, Xian Yang 0001, Qingwei Lin, Hongyu Zhang 0002, Feng Gao 0022, Zhangwei Xu, Yingnong Dang, Dongmei Zhang 0001, Hang Dong 0004, Yong Xu 0010, Yu Kang 0006
WWW2
2014 An iterative parameter estimation method for biological systems and its parallel implementation
abstract
SUMMARY One difficulty in building a mechanistic model of biological systems lies in determining correct parameter values. This paper proposes a novel parameter estimation method to infer unknown parameters, such as kinetic rates, from noisy experimental observations. Derived from the approximate Bayesian computation sequential Monte Carlo algorithm, our method predicts the distribution of each parameter rather than a single value via several intermediate distributions. Motivated by the computational intensity of the method, we improve the approximate Bayesian computation sequential Monte Carlo method in two aspects. First, to increase the efficiency, a windowing method is developed to reduce the parameter‐searching space, and an adaptive sampling weight mechanism is introduced to make the intermediate distributions converge to the target distributions in a much quicker manner. Second, to speed up the estimation process, we implement our method in a parallel computing environment to speed up the sampling process. Copyright © 2013 John Wiley & Sons, Ltd.
Xian Yang 0001, Yike Guo, Li Guo 0002
Concurr. Comput. Pract. Exp.1
2012 Developing a novel integrated model of p38 MAPK and glucocorticoid signalling pathways
abstract
Glucocorticoid (GC) resistance is a key mechanism by which traditional asthma treatments become ineffective for patients, yet the molecular characteristics of the associated regulatory changes are largely unknown. Significant evidence suggests that crosstalk between p38 Mitogen Activated Protein Kinase (MAPK) and GC signalling pathways may contribute to this resistance. Based on a number of studies, a simplified GC signalling pathway model was developed and integrated with a pre-existing model of the p38 MAPK pathway. It is predicted that with experimental data, the validity and use of this model can be confirmed, corrections and updates can be made where necessary, and that through the two pathways' interface points the existence and scale of crosstalk can be examined.
Alex Holehouse, Xian Yang 0001, Ian M. Adcock, Yike Guo
CIBCB2
2012 Modelling and performance analysis of clinical pathways using the stochastic process algebra PEPA
abstract
BACKGROUND: Hospitals nowadays have to serve numerous patients with limited medical staff and equipment while maintaining healthcare quality. Clinical pathway informatics is regarded as an efficient way to solve a series of hospital challenges. To date, conventional research lacks a mathematical model to describe clinical pathways. Existing vague descriptions cannot fully capture the complexities accurately in clinical pathways and hinders the effective management and further optimization of clinical pathways. METHOD: Given this motivation, this paper presents a clinical pathway management platform, the Imperial Clinical Pathway Analyzer (ICPA). By extending the stochastic model performance evaluation process algebra (PEPA), ICPA introduces a clinical-pathway-specific model: clinical pathway PEPA (CPP). ICPA can simulate stochastic behaviours of a clinical pathway by extracting information from public clinical databases and other related documents using CPP. Thus, the performance of this clinical pathway, including its throughput, resource utilisation and passage time can be quantitatively analysed. RESULTS: A typical clinical pathway on stroke extracted from a UK hospital is used to illustrate the effectiveness of ICPA. Three application scenarios are tested using ICPA: 1) redundant resources are identified and removed, thus the number of patients being served is maintained with less cost; 2) the patient passage time is estimated, providing the likelihood that patients can leave hospital within a specific period; 3) the maximum number of input patients are found, helping hospitals to decide whether they can serve more patients with the existing resource allocation. CONCLUSIONS: ICPA is an effective platform for clinical pathway management: 1) ICPA can describe a variety of components (state, activity, resource and constraints) in a clinical pathway, thus facilitating the proper understanding of complexities involved in it; 2) ICPA supports the performance analysis of clinical pathway, thereby assisting hospitals to effectively manage time and resources in clinical pathway.
Xian Yang 0001, Rui Han 0001, Yike Guo, Jeremy T. Bradley, Benita Cox, Robert Dickinson, Richard Kitney
BMC Bioinform.1