EDBT 2026 Demo / reviewers in the wild / expert
Baoyao Yang
dblp:191/1061
· DBLP profile ↗
34ranked-venue papers
10as first author
28since 2021 · last 2026
0000-0001-9092-3164ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 17 since 2021Artificial intelligence and machine learning · 11 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCAN: Self-Calibrated Textual Anchoring for Dual-Granularity Video Screening in Text-Video RetrievalabstractText-Video Retrieval (TVR) is a foundational cross-modal task that aims to align natural language queries with relevant video content, playing a critical role in multimedia search and video understanding. Current methods typically address the cross-modal semantic gap by either augmenting text with generated captions or enhancing video representations at the feature level. However, these strategies often introduce new challenges: generated captions may contain hallucinations that mislead alignment, while dense video representations can retain substantial information redundancy and semantic bias, collectively hindering precise and efficient retrieval. To overcome these limitations, we propose SCAN (Self-Calibrated Textual Anchoring for Dual-Granularity Video Screening). Our framework introduces a dual-granularity video screening mechanism that progressively filters out redundant visual information to obtain semantically condensed video representations. Simultaneously, it employs a self-calibrated textual anchoring module, which leverages the video content to verify and reinforce reliable textual cues, thereby strengthening core text semantics and mitigating caption hallucinations. These two processes are jointly optimized, which aligns text and video semantics in a synergistic manner. Extensive experiments on three benchmarks (MSR-VTT, MSVD, and DiDeMo) demonstrate that SCAN consistently outperforms state-of-the-art methods, validating its effectiveness in reducing hallucinations, eliminating redundancy, and bridging semantic asymmetry for robust text-video retrieval. Dixin Chen, Baoyao Yang, Haifeng Lin, Canrong Du, Wenbin Yao |
ICMR | 2 |
| 2026 | PBA: Persistent backdoor attack on federated learning via distributed generative triggers
Wenyin Yang, Weidong Wu, Baoyao Yang |
Comput. Networks | 6 |
| 2026 | CAM-Interacted Vision GNN for Multi-Label Medical ImagesabstractVision Graph Neural Network (ViG) is designed to recognize different objects through graph-level processing. However, ViG constructs graphs with appearance-level neighbors and neglects the category semantic. The oversight results in the unintentional connection of patches that belong to different objects, thus affecting the distinctiveness of categories in multi-label medical image learning. Since the pixel-level annotations for images are not easily available, category-aware graphs can not be directly built. To solve this problem, we consider localizing category-specific regions using Class Activation Maps (CAMs), an effective way to highlight regions belonging to each category without requiring manual annotations. Specifically, we propose a CAM-interacted Vision GNN (CiV-GNN), in which category-aware graphs are formed to perform intra-category graph processing. CIV-GNN includes a Class-activated Patch Division (CAPD) module, which introduces CAMs as guidance for category-aware graph building. Furthermore, we develop a Multi-graph Interactive Processing (MIP) module to model the relations between category-aware graphs, promoting inter-category interaction learning. Experimental results show that CiV-GNN performs well in surgical tool localization and multi-label medical image classification. Specifically, for m2cai16-localization, CiV-GNN exhibits a 1.43% and 7.02% improvement in mAP50 and mAP50-95, respectively, compared to YOLOv8. Jingchao Wang 0002, Baoyao Yang, Si-Qi Liu 0003, Xiaoqi Zheng, Wenbin Yao, Junxiang Chen |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Uncertainty Reactivation: Dynamic Contrastive Correction for Semi-Supervised Medical Image SegmentationabstractSemi-supervised medical image segmentation has advanced significantly by utilizing pseudo-labeled annotations. However, ensuring pseudo-label accuracy remains challenging, often causing misclassification and confirmation bias. Existing methods mainly use prediction uncertainty to exclude or downweight uncertain regions, but these areas frequently coincide with diagnostically important zones, such as lesion cores or tissue boundaries. Neglecting them can thus degrade segmentation performance. To address this, we propose a Dynamic Contrastive Correction Network (DCCN) that corrects uncertain regions instead of ignoring them. DCCN aligns features from high-uncertainty areas with dynamically assigned classes via contrastive learning, reconstructing the uncertain feature space. Additionally, a Multi-layer Sampling (MLS) module leverages boundary-aware sampling to focus contrastive learning on uncertain tissue boundaries. Experiments on two public datasets show that DCCN surpasses previous SOTA methods and effectively mitigates the challenges of high-uncertainty regions. Kexin Xie, Baoyao Yang, Wanyun Li, Fei Lyu 0004 |
BIBM | 2 |
| 2025 | FairFed++: Closing the Fairness Gap in Federated Learning Through Self-Evolving Clustered OptimizationabstractPerformance fairness in federated learning (FL) aims to ensure that the server treats all clients equitably, thereby encouraging the participation of low-performance clients and enhancing the generalization capabilities of the global model. However, due to client heterogeneity, conflicts may arise among the gradients of different clients during FL, suppressing performance fairness. Although current FL algorithms can achieve convergence, such conflicts lead to performance unfairness when reaching the optimal solution, thereby significantly restricting the generalization ability of the global model learned through FL. To address this issue, strategies such as reweighting and data augmentation have been proposed. However, these approaches often result in performance degradation for certain clients while striving for fairness. Recent studies have highlighted the potential of cluster federated learning (CFL) in achieving performance fairness. Nevertheless, the heavy reliance on a pre-specified number of clusters not only limits its adaptability but also increases the complexity of FL. Inspired by the principle of species evolution, where cells divide under specific internal conditions, we propose a novel FL method, namely FairFed++. Specifically, FairFed++ performs self-evolving clustered optimization, explicitly releasing the reliance on prior knowledge of clustering. By utilizing the accuracy variance within clusters as the splitting criterion, FairFed++ automatically determines the optimal clustering strategy in each round of FL communication until convergence. This approach dynamically adjusts the number of clusters during training without requiring manual intervention, thus improving both adaptability and practicability. Experiments conducted on six datasets demonstrate that FairFed++ achieves superior performance fairness while preserving the generalization ability of the global model. Zhixiang Fang, Baoyao Yang, Weide Zhan, Yanchao Tang, Yiqun Zhang 0006 |
ECAI | 2 |
| 2025 | Unlocking the Potential of mLLMs: Enhancing Video-Text Retrieval Through Caption Supplementation and Conical Embedding OptimizationabstractThe burgeoning field of video-text retrieval has witnessed significant advancements with the advent of deep learning. However, understanding and matching textual descriptions and video data remains a formidable challenge due to the large information gap across textual and video modalities. As observed, the caption of a video is commonly under-described, lacking expressions of minor characters or local details. Some recent advances have attempted to leverage multimodal Large Language Model (mLLM) to bridge the comprehension gap. However, mLLMs’ potential in enhancing video-text retrieval (VTR) is understudied. This paper aims to fill this research vacancy, analyzing the practical significance and model preferences for utilizing mLLMs in VTR enhancement, as well as investigating the effective integration of mLLM-derived information into the retrieval learning. Based on our analytical insights, we innovatively propose treating mLLM as caption supplements rather than substitutes to bridge the expression gap across modalities. To achieve better cross-modal alignment, we systematically generate diverse variations of videos to construct an elastic visual space. By treating mLLM-supplemented captions as out-of-space points, cross-modal representation learning is accomplished through the optimization of a conical-like representation space. Our model achieves state-of-the-art results on various benchmarks, including MSR-VTT, MSVD, and DiDeMo, and analytical experiments suggest appropriate prompt proposals and indicate our method’s robustness to different mLLMs. Baoyao Yang, Junxiang Chen, Wenbin Yao |
ECAI | 1 |
| 2025 | Image-assisted Label Connective Completion for Vessel Segmentation with Insufficient AnnotationsabstractAutomatic and accurate vessel segmentation is crucial for disease diagnosis. Deep learning methods are widely used, but their promising results rely on accurately annotated data. Due to complex vessel morphology and low-contrast image, accurate vessel delineation poses a practical challenge, resulting in insufficient annotations, which is a prominent form of noisy labels. This paper proposes an Image-assisted Label Connective Completion method, which enhances label’s vessel information by images under the supervision of connectivity to address insufficient annotation issue. Specifically, we develop an Image-guided Vessel Enhancement module, which transmits structural information extracted from images based on label navigation to label space, promoting completion of missing annotated parts in original labels. In addition, a branch completion-connectivity loss is designed and introduced as an auxiliary supervision to prevent vessel branch disconnection during label completion. Experimental results on DRIVE, CHASE DB1 and DCA1 datasets demonstrate that our method outperforms existing noisy labels learning methods. Xiaoqi Zheng, Baoyao Yang, Xiuwen Fang, Wenfang Yao, Mang Ye |
ICASSP | 2 |
| 2025 | Harnessing Feature Distribution Consistency for Federated Learning with Noisy LabelsabstractLabel noise in federated learning (FL) is a significant detrimental factor that substantially degrades FL performance. Current methods attempt to mitigate this issue by identifying noisy-labeled samples using rudimentary indicators. However, these indicators fail to distinguish between clean and noisy labels in heavily poisoned scenarios. Although global assessment has been introduced to enhance the recognition of noisy labels, existing methods still struggle to prevent performance degradation in FL systems. Therefore, this paper proposes leveraging the global feature distribution (an integration of local distributions) to improve the detection and correction of noisy labels in FL systems. Specifically, we assess the consistency between the category to which samples belong in the global distribution and their labels to detect noisy labels. Subsequently, a temporal dual-view consistency (TDC) mechanism is designed and introduced to correct the detected noisy labels. TDC evaluates label consistency from the perspectives of sample diversity and label continuity, thereby enhancing the reliability of label correction. Extensive experimental results on both synthetic and real-world noisy datasets demonstrate that the proposed method surpasses current SOTA approaches. Yali Ma, Baoyao Yang, Yanchao Tang, Weide Zhan, Wenyin Yang |
ICIP | 2 |
| 2025 | Action Decomposition-based Actor-Critic for Supply Chain OptimizationabstractIn recent years, deep reinforcement learning (DRL) has demonstrated significant potential to address complex and dynamic supply chain optimization (SCO) problems. However, existing DRL algorithms often encounter challenges when dealing with large-scale and high-dimensional supply chain decisions, making it difficult to effectively coordinate production and transportation. To address these issues, this paper proposes an innovative deep reinforcement learning algorithm—Action Decomposition Actor-Critic (ADAC). This algorithm significantly reduces learning complexity by decomposing complex decision tasks into multiple subtasks. Based on real-world supply chain scenarios, we construct a complex multi-stage supply chain environment. Additionally, we employ action decomposition-based Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) algorithms to learn optimal policies in continuous action spaces, enabling fine-grained control of production and transportation. To verify the effectiveness of the algorithm, we conduct extensive experiments on both simulation and real-world datasets. Experimental results demonstrate that the ADAC algorithm outperforms traditional heuristic algorithms and general DRL algorithms in multiple complex supply chain scenarios. This shows the strong robustness and wide applicability of the ADAC algorithm in SCO problems. Zhengrong Chen, Qinghua Zhu 0001, An Zeng, YuZhu Ji, Baoyao Yang, Dan Pan 0001 |
ICME | 5 |
| 2025 | Unifying Spatio-Temporal Contexts for Advanced Text-Video RetrievalabstractText-to-video retrieval (T2VR) aims to identify the most semantically relevant video based on a text query. Text queries typically involve diverse visual elements and events in video, making it non-trivial to learn a robust video feature representation for different queries. An abundance of spatial information make the model overwhelmed by redundancy and noisy and struggle to focus on linchpin visual elements. Additionally, without effective guidance, models grapple with connecting temporal information across different frames. In this paper, we introduce a Spatial-Temporal Pooling (STP) method to cohesively capture and unify the inherent spatio-temporal context within videos. For spatial information, STP leverages spatial tags such as entities, scenes, and text as attention prompts, steering the model toward salient visual elements while mitigating the impact of redundancies and noise. For temporal information, STP adopts video narratives summarized in captions as temporal prompts to enhance the model’s perception of events. Experimental results show that our approach has achieved improvements on the MSRVTT(1.9%), MSVD(1.2%), and VATEX(0.7%) datasets compared to the SOTA methods. Yanhao Huang, Baoyao Yang, Junxiang Chen, Wenbin Yao, Dixin Chen |
ICME | 2 |
| 2025 | Simple but Effective: Sub-Volume Contrastive Learning for Class-Imbalanced Semi-Supervised 3D Medical Image SegmentationabstractMedical image segmentation is essential for precise anatomical delineation and clinical decision-making. However, fully supervised methods are limited by the substantial cost of acquiring pixel-level annotations, particularly for 3D volumetric data. Semi-supervised learning (SSL) alleviates this challenge by leveraging unlabeled data, yet it remains hindered by severe class imbalance, where dominant structures disproportionately occupy the voxel space, leading to feature degradation and unreliable pseudo-labels. To address this issue, we propose a simple but effective SSL framework, namely Sub-Volume Contrastive Learning (SuVCL), to enhance feature discriminability in imbalanced 3D medical image segmentation. Our approach incorporates localized contrastive learning through sub-volume sampling, which captures small but semantically informative regions to retain fine-grained structural details while mitigating computational overhead. Furthermore, we introduce a balanced memory bank mechanism, which dynamically maintains class-specific feature representations with adaptive updates guided by class-predictive confidence. Extensive experimental evaluations demonstrate that our method substantially enhances segmentation performance for minority classes, demonstrating substantial performance gains over existing SOTAs. Xianrun Xu, Baoyao Yang, Wanyun Li, Jingsong Lin, Yufei Xu |
ACM Multimedia | 2 |
| 2025 | FedCD: A Hybrid Federated Learning Framework for Adaptive Training Under Data Heterogeneity
Weide Zhan, Baoyao Yang, Zhixiang Fang, Dongzhe Li, Yali Ma, Yiqun Zhang 0006 |
PRCV (9) | 2 |
| 2025 | Multi-Agent Reinforcement Learning Algorithm Using Dynamic OW-QMIX in Complex Supply Chain ScenariosabstractHow to effectively optimize the operation for a complex supply chain environment has been high on the agenda. Although the existing deep reinforcement learning methods have achieved success in certain applications, they still face limitations in complex supply chain environments, including difficulties in data sharing and a lack of digital collaboration, especially in multi-agent systems. In order to meet this challenge, we propose a novel multi-agent reinforcement learning algorithm based on dynamic optimistic weights (DO-QMIX), aiming at solving the shortcomings of the traditional weighted QMIX algorithm (WQMIX) in the simplicity of the weighting function. WQMIX employs two weighting schemes to handle multi-agent issues. However, its fixed weighting function restricts algorithm performance and hinders adaptability to dynamic, complex supply chain challenges. Therefore, we propose a dynamic weighting mechanism, which can adjust the weighting function in real time based on the changes in the environment, thus improving the overall efficiency. We construct a complex multi-stage supply chain environment in the real-world supply chain scenario and conduct many experiments using both real-world and simulated datasets. The experimental results demonstrate that DO-QMIX is significantly superior to the traditional multi-agent reinforcement learning algorithm in complex supply chain scenarios, especially in dealing with dynamic changes and complex decisions. Zhiqi Liu, Qinghua Zhu 0001, An Zeng, YuZhu Ji, Baoyao Yang |
SMC | 5 |
| 2025 | Allosteric Feature Collaboration for Model-Heterogeneous Federated LearningabstractAlthough federated learning (FL) has achieved outstanding results in privacy-preserved distributed learning, the setting of model homogeneity among clients restricts its wide application in practice. This article investigates a more general case, namely, model-heterogeneous FL (M-hete FL), where client models are independently designed and can be structurally heterogeneous. M-hete FL faces new challenges in collaborative learning because the parameters of heterogeneous models could not be directly aggregated. In this article, we propose a novel allosteric feature collaboration (AlFeCo) method, which interchanges knowledge across clients and collaboratively updates heterogeneous models on the server. Specifically, an allosteric feature generator is developed to reveal task-relevant information from multiple client models. The revealed information is stored in the client-shared and client-specific codes. We exchange client-specific codes across clients to facilitate knowledge interchange and generate allosteric features that are dimensionally variable for model updates. To promote information communication between different clients, a dual-path (model-model and model-prediction) communication mechanism is designed to supervise the collaborative model updates using the allosteric features. Client models are fully communicated through the knowledge interchange between models and between models and predictions. We further provide theoretical evidence and convergence analysis to support the effectiveness of AlFeCo in M-hete FL. The experimental results show that the proposed AlFeCo method not only performs well on classical FL benchmarks but also is effective in model-heterogeneous federated antispoofing. Our codes are publicly available at https://github.com/ybaoyao/AlFeCo. Baoyao Yang, Pong C. Yuen, Yiqun Zhang 0006, An Zeng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | MMS: Morphology-Mixup Stylized Data Generation for Single Domain Generalization in Medical Image SegmentationabstractSingle-source domain generalization in medical image segmentation is a challenging yet practical task, as domain shift commonly exists across medical datasets. Previous works have attempted to alleviate this problem through adversarial data augmentation or random-style transformation. However, these approaches neither fully leverage medical information nor consider the morphological structure alterations. To address these limitations and enhance the fidelity and diversity of the augmented data, we propose a Morphology-Mixup Stylized data generation (MMS) method, which expands source data from a new morphological perspective, guided by the characteristics of medical imaging. Specifically, we design a Mixed Dual-stream Auto-Encoder (MDs-AE) to simulate the morphology changes between medical image slices and mix the morphology of two slices. In addition, we introduce a feature consistency strategy to improve the effectiveness of morphology mixing. The trained MDs-AE with a random styler is used to generate data that vary in both morphology and style to enhance the generalization ability of the segmentation network. Extensive experimental results demonstrate that MMS is effective and outperforms the state-of-the-art on three cross-domain segmentation tasks. Xiaochen He, Baoyao Yang, Fei Lyu 0004 |
ICASSP | 2 |
| 2024 | Domain Dilation for Single Domain GeneralizationabstractThis work investigates the Single Domain Generalization (SDG), which generalizes a model from a single source domain to multiple unseen target domains. Most existing SDG methods focus on expanding the source domain by either transforming the source samples into different styles or optimizing adversarial noise perturbations applied to the source samples. However, these methods generate fictitious samples using specific image transformation, resulting in insufficient domain expansion. In this paper, we propose a progressive domain expansion method, namely domain dilation (DD) for SDG. This method dilates the source domain from two perspectives: enriching source domain diversity and generating various pseudo domains. To enrich source domain diversity, we generate fictitious samples with diverse styles. To obtain various pseudo domains, this paper generates pseudo domains with a new distribution by maximizing the domain difference from the source domain. Our method outperforms the state-of-the-art methods on prevalent single domain generalization benchmarks through extensive experiments, offering improved results. Yuehui Fan, Baoyao Yang, Fei Lyu 0004 |
ICIP | 2 |
| 2024 | CAM-Guided Translation for Unpaired Weakly-Supervised Medical Image SegmentationabstractMulti-modal learning has shown advantages in improving weakly-supervised medical image segmentation (WS- MIS). However, most current works are based on paired data, which is infeasible to collect in certain scenarios. Although modal translation can be used to generate paired data, it often leads to low-quality translations, such as local deformations or irrational textures, without prior knowledge. This paper proposes a discriminative-aware image translation method, which introduces class activation maps (CAMs) to localize discriminative areas, thus overcoming the lack of pixel-wise annotations in WS-MIS. In addition, we design a CAM-correlation constraint that facilitates multi-modal complementary information exchange to enhance the consistency between CAMs generated from different modalities. Experimental results show that our method outperforms recent weakly-supervised segmentation works when using unpaired multi-modal data. Yuebin Xie, Xiaochen He, Baoyao Yang, Fei Lyu 0004, Si-Qi Liu 0003 |
ICME | 3 |
| 2024 | Multi-category Graph Reasoning for Multi-modal Brain Tumor Segmentation
Dongzhe Li, Baoyao Yang, Weide Zhan, Xiaochen He |
MICCAI (8) | 2 |
| 2024 | Adaptive 3DCNN-Based Interpretable Ensemble Model for Early Diagnosis of Alzheimer's DiseaseabstractAdaptive interpretable ensemble model based on three-dimensional Convolutional Neural Network (3DCNN) and Genetic Algorithm (GA), i.e., 3DCNN+EL+GA, was proposed to differentiate the subjects with Alzheimer's Disease (AD) or Mild Cognitive Impairment (MCI) and further identify the discriminative brain regions significantly contributing to the classifications in a data-driven way. Plus, the discriminative brain sub-regions at a voxel level were further located in these achieved brain regions, with a gradient-based attribution method designed for CNN. Besides disclosing the discriminative brain sub-regions, the testing results on the datasets from the Alzheimer's Disease Neuroimaging Initiative (ADNI) and the Open Access Series of Imaging Studies (OASIS) indicated that 3DCNN+EL+GA outperformed other state-of-the-art deep learning algorithms and that the achieved discriminative brain regions (e.g., the rostral hippocampus, caudal hippocampus, and medial amygdala) were linked to emotion, memory, language, and other essential brain functions impaired early in the AD process. Future research is needed to examine the generalizability of the proposed method and ideas to discern discriminative brain regions for other brain disorders, such as severe depression, schizophrenia, autism, and cerebrovascular diseases, using neuroimaging. Dan Pan 0001, Genqiang Luo, An Zeng, Chao Zou, Haolin Liang, Tong Zhang 0015, Baoyao Yang |
IEEE Trans. Comput. Soc. Syst. | 8 |
| 2024 | DNA-T: Deformable Neighborhood Attention Transformer for Irregular Medical Time SeriesabstractThe real-world Electronic Health Records (EHRs) present irregularities due to changes in the patient's health status, resulting in various time intervals between observations and different physiological variables examined at each observation point. There have been recent applications of Transformer-based models in the field of irregular time series. However, the full attention mechanism in Transformer overly focuses on distant information, ignoring the short-term correlations of the condition. Thereby, the model is not able to capture localized changes or short-term fluctuations in patients' conditions. Therefore, we propose a novel end-to-end Deformable Neighborhood Attention Transformer (DNA-T) for irregular medical time series. The DNA-T captures local features by dynamically adjusting the receptive field of attention and aggregating relevant deformable neighborhoods in irregular time series. Specifically, we design a Deformable Neighborhood Attention (DNA) module that enables the network to attend to relevant neighborhoods by drifting the receiving field of neighborhood attention. The DNA enhances the model's sensitivity to local information and representation of local features, thereby capturing the correlation of localized changes in patients' conditions. We conduct extensive experiments to validate the effectiveness of DNA-T, outperforming existing state-of-the-art methods in predicting the mortality risk of patients. Moreover, we visualize an example to validate the effectiveness of the proposed DNA. Jianxuan Huang, Baoyao Yang, Kejing Yin, Jingwen Xu 0002 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Early Diagnosis of Alzheimer's Disease Based on Multimodal Hypergraph Attention NetworkabstractAlzheimer’s disease (AD) is a typical neurodegenerative disease involving multiple pathogenic factors. Early detection is the key to effective treatment of AD. However, most methods are developed based on data from a single modality, and ignore the relationships among subjects. In machine learning problems, hypergraph can be used to express the relationships between objects. In light of this, a framework for early diagnosis of Alzheimer’s disease based on multimodal hypergraph attention network is proposed in this paper. Specifically, we combine multimodal features to construct cross modal hypergraph, which represents the high-order structural relationships among subjects. Finally, a hypergraph attention network is used to fuse hypergraphs and perform the final classification. Our experimental results on the Alzheimer Disease Neuroimaging Initiative (ADNI) database show that our proposed method has better classification performance than the most advanced methods. Baoyao Yang, Dan Pan 0001, An Zeng, Long Wu |
ICME | 2 |
| 2022 | Cross-Domain Missingness-Aware Time-Series Adaptation With Similarity Distillation in Medical ApplicationsabstractMedical time series of laboratory tests has been collected in electronic health records (EHRs) in many countries. Machine-learning algorithms have been proposed to analyze the condition of patients using these medical records. However, medical time series may be recorded using different laboratory parameters in different datasets. This results in the failure of applying a pretrained model on a test dataset containing a time series of different laboratory parameters. This article proposes to solve this problem with an unsupervised time-series adaptation method that generates time series across laboratory parameters. Specifically, a medical time-series generation network with similarity distillation is developed to reduce the domain gap caused by the difference in laboratory parameters. The relations of different laboratory parameters are analyzed, and the similarity information is distilled to guide the generation of target-domain specific laboratory parameters. To further improve the performance in cross-domain medical applications, a missingness-aware feature extraction network is proposed, where the missingness patterns reflect the health conditions and, thus, serve as auxiliary features for medical analysis. In addition, we also introduce domain-adversarial networks in both feature level and time-series level to enhance the adaptation across domains. Experimental results show that the proposed method achieves good performance on both private and publicly available medical datasets. Ablation studies and distribution visualization are provided to further analyze the properties of the proposed method. Baoyao Yang, Mang Ye, Qingxiong Tan, Pong C. Yuen |
IEEE Trans. Cybern. | 1 |
| 2022 | Revealing Task-Relevant Model Memorization for Source-Protected Unsupervised Domain AdaptationabstractSource-data-free unsupervised domain adaptation (SF-UDA) is an approach to improve model performance in the target domain without accessing the source data. Some SF-UDA methods have been proposed and achieved promising results using the information from source-model parameters. However, current research on information security confirms the ability of a well-trained model to memorize its training data. Therefore, SF-UDA methods that access model parameters remain at risk of privacy disclosure. This paper introduces a new topic of source-protected UDA (SP-UDA) that adapts the source model to the target domain while protecting the source-domain data and model privacy. In SP-UDA, only a black-box source model and a set of unlabeled target data are available for domain adaptation. We consider SP-UDA from a new perspective of model memorization revelation. A Source-Protected Generative Model (SPGM) is developed to reveal task-relevant memorization from the source model. SPGM directly distills the inverse process of the source model without access to source-model parameters to meet the privacy protection objective in SP-UDA. The SPGM is learned under the supervision of a newly designed metric named privacy-protected transfer (PPT). The PPT metric measures the transferability and desensitization of the generated data to encourage the SPGM to extract task-relevant information rather than the unintended memorization. A set of desensitized pseudo data is then generated as substitutes for the real source data in UDA. The performance of the proposed method has been validated in four cross-dataset recognition applications with encouraging results. Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2022 | Model-Induced Generalization Error Bound for Information-Theoretic Representation Learning in Source-Data-Free Unsupervised Domain AdaptationabstractMany unsupervised domain adaptation (UDA) methods have been developed and have achieved promising results in various pattern recognition tasks. However, most existing methods assume that raw source data are available in the target domain when transferring knowledge from the source to the target domain. Due to the emerging regulations on data privacy, the availability of source data cannot be guaranteed when applying UDA methods in a new domain. The lack of source data makes UDA more challenging, and most existing methods are no longer applicable. To handle this issue, this paper analyzes the cross-domain representations in source-data-free unsupervised domain adaptation (SF-UDA). A new theorem is derived to bound the target-domain prediction error using the trained source model instead of the source data. On the basis of the proposed theorem, information bottleneck theory is introduced to minimize the generalization upper bound of the target-domain prediction error, thereby achieving domain adaptation. The minimization is implemented in a variational inference framework using a newly developed latent alignment variational autoencoder (LA-VAE). The experimental results show good performance of the proposed method in several cross-dataset classification tasks without using source data. Ablation studies and feature visualization also validate the effectiveness of our method in SF-UDA. Baoyao Yang, Hao-Wei Yeh, Tatsuya Harada, Pong C. Yuen |
IEEE Trans. Image Process. | 1 |
| 2021 | A Segmentation-Assisted Model for Universal Lesion Detection with Partial Labels
Fei Lyu 0004, Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen |
MICCAI (5) | 2 |
| 2021 | SoFA: Source-data-free Feature Alignment for Unsupervised Domain AdaptationabstractApplying a trained model on a new scenario may suffer from domain shift. Unsupervised domain adaptation (UDA) has been proven to be an effective approach to solve the problem of domain shift by leveraging both data from the scenario that the model was trained on (source) and the new scenario (target). Although the source data are available for training the source model, there is no guarantee that the source data will still be available when applying UDA in the future due to emerging regulations on privacy of data. This results in the in-applicability of most existing UDA methods in the absence of source data. This paper proposes a source-data-free feature alignment (SoFA) method to address this problem by only using the trained source model and unlabeled target data. The source model is used to predict the labels for target data, and we model the generation process from predicted classes to input data to infer the latent features for alignment. Specifically, a mixture of Gaussian distributions is induced from the predicted classes as the reference distribution. The encoded target features are then aligned to the reference distribution via variational inference to extract class semantics without accessing source data. Relationship of the proposed method and the theory of domain adaptation is provided to verify the performance. Experimental results show the proposed method achieves higher or comparable accuracy compared to the existing methods in several cross-dataset classification tasks. Ablation studies are also conducted to confirm the importance of latent feature alignment to adaptation performance. Hao-Wei Yeh, Baoyao Yang, Pong C. Yuen, Tatsuya Harada |
WACV | 2 |
| 2021 | Learning adaptive geometry for unsupervised domain adaptation
Baoyao Yang, Pong C. Yuen |
Pattern Recognit. | 1 |
| 2021 | Explainable Uncertainty-Aware Convolutional Recurrent Neural Network for Irregular Medical Time SeriesabstractInfluenced by the dynamic changes in the severity of illness, patients usually take examinations in hospitals irregularly, producing a large volume of irregular medical time-series data. Performing diagnosis prediction from the irregular medical time series is challenging because the intervals between consecutive records significantly vary along time. Existing methods often handle this problem by generating regular time series from the irregular medical records without considering the uncertainty in the generated data, induced by the varying intervals. Thus, a novel Uncertainty-Aware Convolutional Recurrent Neural Network (UA-CRNN) is proposed in this article, which introduces the uncertainty information in the generated data to boost the risk prediction. To tackle the complex medical time series with subseries of different frequencies, the uncertainty information is further incorporated into the subseries level rather than the whole sequence to seamlessly adjust different time intervals. Specifically, a hierarchical uncertainty-aware decomposition layer (UADL) is designed to adaptively decompose time series into different subseries and assign them proper weights in accordance with their reliabilities. Meanwhile, an Explainable UA-CRNN (eUA-CRNN) is proposed to exploit filters with different passbands to ensure the unity of components in each subseries and the diversity of components in different subseries. Furthermore, eUA-CRNN incorporates with an uncertainty-aware attention module to learn attention weights from the uncertainty information, providing the explainable prediction results. The extensive experimental results on three real-world medical data sets illustrate the superiority of the proposed method compared with the state-of-the-art methods. Qingxiong Tan, Mang Ye, Andy Jinhua Ma, Baoyao Yang, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | DATA-GRU: Dual-Attention Time-Aware Gated Recurrent Unit for Irregular Multivariate Time SeriesabstractDue to the discrepancy of diseases and symptoms, patients usually visit hospitals irregularly and different physiological variables are examined at each visit, producing large amounts of irregular multivariate time series (IMTS) data with missing values and varying intervals. Existing methods process IMTS into regular data so that standard machine learning models can be employed. However, time intervals are usually determined by the status of patients, while missing values are caused by changes in symptoms. Therefore, we propose a novel end-to-end Dual-Attention Time-Aware Gated Recurrent Unit (DATA-GRU) for IMTS to predict the mortality risk of patients. In particular, DATA-GRU is able to: 1) preserve the informative varying intervals by introducing a time-aware structure to directly adjust the influence of the previous status in coordination with the elapsed time, and 2) tackle missing values by proposing a novel dual-attention structure to jointly consider data-quality and medical-knowledge. A novel unreliability-aware attention mechanism is designed to handle the diversity in the reliability of different data, while a new symptom-aware attention mechanism is proposed to extract medical reasons from original clinical records. Extensive experimental results on two real-world datasets demonstrate that DATA-GRU can significantly outperform state-of-the-art methods and provide meaningful clinical interpretation. Qingxiong Tan, Mang Ye, Baoyao Yang, Si-Qi Liu 0003, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen |
AAAI | 3 |
| 2019 | Cross-Domain Visual Representations via Unsupervised Graph AlignmentabstractIn unsupervised domain adaptation, distributions of visual representations are mismatched across domains, which leads to the performance drop of a source model in the target domain. Therefore, distribution alignment methods have been proposed to explore cross-domain visual representations. However, most alignment methods have not considered the difference in distribution structures across domains, and the adaptation would subject to the insufficient aligned cross-domain representations. To avoid the misclassification/misidentification due to the difference in distribution structures, this paper proposes a novel unsupervised graph alignment method that aligns both data representations and distribution structures across the source and target domains. An adversarial network is developed for unsupervised graph alignment, which maps both source and target data to a feature space where data are distributed with unified structure criteria. Experimental results show that the graph-aligned visual representations achieve good performance on both crossdataset recognition and cross-modal re-identification. Baoyao Yang, Pong C. Yuen |
AAAI | 1 |
| 2019 | UA-CRNN: Uncertainty-Aware Convolutional Recurrent Neural Network for Mortality Risk PredictionabstractAccurate prediction of mortality risk is important for evaluating early treatments, detecting high-risk patients and improving healthcare outcomes. Predicting mortality risk from the irregular clinical time series data is challenging due to the varying time intervals in the consecutive records. Existing methods usually solve this issue by generating regular time series data from the original irregular data without considering the uncertainty in the generated data, caused by varying time intervals. In this paper, we propose a novel Uncertainty-Aware Convolutional Recurrent Neural Network (UA-CRNN), which incorporates the uncertainty information in the generated data to improve the mortality risk prediction performance. To handle the complex clinical time series data with sub-series of different frequencies, we propose to incorporate the uncertainty information into the sub-series level rather than the whole time series data. Specifically, we design a novel hierarchical uncertainty-aware decomposition layer (UADL) to adaptively decompose time series into different sub-series and assign them proper weights according to their reliabilities. Experimental results on two real-world clinical datasets demonstrate that the proposed UA-CRNN method significantly outperforms state-of-the-art methods in both short-term and long-term mortality risk predictions. Qingxiong Tan, Andy Jinhua Ma, Mang Ye, Baoyao Yang, Huiqi Deng, Vincent Wai-Sun Wong, Yee-Kit Tse, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Jessica Yuet-Ling Ching, Francis Ka-Leung Chan, Pong C. Yuen |
CIKM | 4 |
| 2019 | Body Parts Synthesis for Cross-Quality Pose EstimationabstractAlthough encouraging results have been obtained in human pose estimation in recent years, the performance may degrade dramatically when the image quality differs between training and testing data sets. This paper addresses problems in cross-image-quality human pose estimation. To achieve this, we follow an unsupervised domain adaptation approach, in which labels in the target domain are unavailable. Unlike existing unsupervised domain adaptation methods that find label information from unlabeled data, the target pose information (label) is instead generated by synthesizing body parts with similar image-quality of the target domain. A translative dictionary is learned to associate the source and target domains, and a cross-quality adaptation model is developed to refine the source pose estimator using the synthesized target body parts. We perform cross-quality experiments on three data sets with different image quality by using two state-of-the-art pose estimators, and compare the proposed method with five unsupervised domain adaptation methods. Our experimental results show that the proposed method outperforms not only the source pose estimators, but also other unsupervised domain adaptation methods. Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Domain-Shared Group-Sparse Dictionary Learning for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation has been proved to be a promising approach to solve the problem of dataset bias. To employ source labels in the target domain, it is required to align the joint distributions of source and target data. To do this, the key research problem is to align conditional distributions across domains without target labels. In this paper, we propose a new criterion of domain-shared group-sparsity that is an equivalent condition for conditional distribution alignment. To solve the problem in joint distribution alignment, a domain-shared group-sparse dictionary learning method is developed towards joint alignment of conditional and marginal distributions. A classifier for target domain is trained using the domain-shared group-sparse coefficients and the target-specific information from the target data. Experimental results on cross-domain face and object recognition show that the proposed method outperforms eight state-of-the-art unsupervised domain adaptation algorithms. Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen |
AAAI | 1 |
| 2018 | Learning domain-shared group-sparse representation for unsupervised domain adaptation
Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen |
Pattern Recognit. | 1 |