Yihang Wu

dblp:325/0689 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
27since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Federated CLIP for Resource-Efficient Heterogeneous Medical Image Classification
abstract
Despite the remarkable performance of deep models in medical imaging, they still require source data for training, which limits their potential in light of privacy concerns. Federated learning (FL), as a decentralized learning framework that trains a shared model with multiple hospitals (a.k.a., FL clients), provides a feasible solution. However, data heterogeneity and resource costs hinder the deployment of FL models, especially when using vision language models (VLM). To address these challenges, we propose a novel contrastive language-image pre-training (CLIP) based FL approach for medical image classification. Specifically, we introduce a masked feature adaptation module (FAM) as a communication module to reduce the communication load while freezing the CLIP encoders to reduce the computational overhead. Furthermore, we propose a masked multi-layer perceptron (MLP) as a private local classifier to adapt to the client tasks. Moreover, we design an adaptive Kullback-Leibler (KL) divergence-based distillation regularization method to enable mutual learning between FAM and MLP. Finally, we incorporate model compression to transmit the FAM parameters while using ensemble predictions for classification. Extensive experiments on four publicly available medical datasets demonstrate that our model provides feasible performance (e.g., 8% higher compared to second best baseline on ISIC2019) with reasonable resource cost (e.g., 120 times faster than FedAVG).
Yihang Wu, Ahmad Chaddad
AAAI1
2026 Breaking the Modality Barrier: Generative Modeling for Accurate Molecule Retrieval from Mass Spectra
abstract
Retrieving molecular structures from tandem mass spectra is a crucial step in rapid compound identification. Existing retrieval methods, such as traditional mass spectral library matching, suffer from limited spectral library coverage, while recent cross-modal representation learning frameworks often encounter modality misalignment, resulting in suboptimal retrieval accuracy and generalization. To address these limitations, we propose GLMR, a Generative Language Model-based Retrieval framework that mitigates the cross-modal misalignment through a two-stage process. In the pre-retrieval stage, a contrastive learning-based model identifies top candidate molecules as contextual priors for the input mass spectrum. In the generative retrieval stage, these candidate molecules are integrated with the input mass spectrum to guide a generative model in producing refined molecular structures, which are then used to re-rank the candidates based on molecular similarity. Experiments on both MassSpecGym and the proposed MassRET-20k dataset demonstrate that GLMR significantly outperforms existing methods, achieving over 40% improvement in top-1 accuracy and exhibiting strong generalizability.
Keyan Ding, Yihang Wu, Xiang Zhuang, Qiang Zhang 0026, Huajun Chen
AAAI3
2026 A systematic analysis of federated learning algorithms for healthcare applications
Ahmad Chaddad, Yihang Wu, Tareef S. Daqqaq, Yousef Katib, Sarah A. Alkhodair, Reem Kateb
Neurocomputing2
2026 GMGaze: MoE-based context-aware gaze estimation with CLIP and multiscale transformer
Yihang Wu, Ahmad Chaddad, Sarah A. Alkhodair, Reem Kateb
Knowl. Based Syst.2
2026 Federated vision transformer with adaptive focal loss for medical image classification
Yihang Wu, Ahmad Chaddad, Tareef S. Daqqaq, Reem Kateb
Knowl. Based Syst.2
2025 A Knowledge Distillation-Based Approach to Enhance Transparency of Classifier Models
abstract
With the rapid development of artificial intelligence (AI), especially in the medical field, the need for its explainability has grown. In medical image analysis, a high degree of transparency and model interpretability can help clinicians better understand and trust the decision-making process of AI models. In this study, we propose a Knowledge Distillation (KD) based approach that aims to enhance the transparency of the AI model in medical image analysis. The initial step is to use traditional CNN to obtain a teacher model and then use KD to simplify the CNN architecture, retain most of the features of the data set, and reduce the number of network layers. It also uses the feature map of the student model to perform hierarchical analysis to identify key features and decision-making processes. This leads to intuitive visual explanations. We selected three public medical data sets (brain tumor, eye disease, and Alzheimer's disease) to test our method. It shows that even when the number of layers is reduced, our model provides a remarkable result in the test set and reduces the time required for the interpretability analysis.
Yihang Wu, Ahmad Chaddad
AAAI3
2025 ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
abstract
Zhigen Li, Jianxiang Peng, Yanmeng Wang, Yong Cao, Tianhao Shen, Minghui Zhang, Linxi Su, Shang Wu, Yihang Wu, YuQian Wang, Ye Wang, Wei Hu, Jianfeng Li, Shaojun Wang, Jing Xiao, Deyi Xiong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhigen Li, Jianxiang Peng, Yanmeng Wang, Tianhao Shen, Linxi Su, Yihang Wu, Jing Xiao 0006, Deyi Xiong
ACL (1)9
2025 MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
abstract
Advancements in generative models have enabled image inpainting models to generate content within specific regions of an image based on provided prompts and masks. However, existing inpainting methods often suffer from problems such as semantic misalignment, structural distortion, and style inconsistency. In this work, we present MTADiffusion, a Mask-Text Alignment diffusion model designed for object inpainting. To enhance the semantic capabilities of the inpainting model, we introduce MTAPipeline, an automatic solution for annotating masks with detailed descriptions. Based on the MTAPipeline, we construct a new MTADataset comprising 5 million images and 25 million mask-text pairs. Furthermore, we propose a multi-task training strategy that integrates both inpainting and edge prediction tasks to improve structural stability. To promote style consistency, we present a novel inpainting style-consistency loss using a pre-trained VGG network and the Gram matrix. Comprehensive evaluations on BrushBench and EditBench demonstrate that MTADiffusion achieves state-of-the-art performance compared to other methods.
Ting Liu 0018, Yihang Wu, Xiaochao Qu, Luoqi Liu
CVPR3
2025 Bargaining-Based Data Markets
abstract
With the prevalence of data-driven business, data markets where data can be commoditized, circulated, and ex-ploited are gaining considerable interest in the data management community. However, the uncertainty in data value poses great challenges for data pricing and thus data trading, which is magnified by the externality arising from the replicable nature of data. In this paper, we present the first bargaining-based data market framework to resolve the externality in data markets. Gearing toward raw data trading, we propose a three-stage bar-gaining model to formulate trading dynamics, which ascertains the data price agreed by both sellers and buyers. With parameters instantiated in preparation stage, an iterative bidding algorithm with provable convergence is designed in negotiation stage to solve the data pricing problem by eliciting equilibrium bids from participants with their profits optimized. Approximation algorithms with guaranteed bounds are presented in settlement stage to solve the NP-hard data allocation problem for profit maximization for the data seller with individual rationality satisfied for data buyers. Experiments on real datasets verify the effectiveness and efficiency of our framework.
Yuran Bi, Jinfei Liu, Kui Ren 0001, Yihang Wu, Yang Cao 0011
ICDE4
2025 Performance Evaluation of the CLIP Model in Classification Tasks
Baosheng Qin, Fenglian Chen, Xiaohan Huang 0013, Yihang Wu, Ahmad Chaddad
SMC6
2025 Analyzing the Calibration of CLIP Models Under Noisy Data Conditions
abstract
Contrastive Language-Image Pretraining (CLIP) has emerged as a powerful paradigm for cross-modal learning, using image-text pairs to achieve remarkable zero-shot classification performance. However, its calibration on noisy data has been less explored, especially in out-of-distribution settings where overconfidence can lead to misclassification. In this work, we evaluate CLIP’s calibration ability under different noise conditions using 10 noise types from the ImageNet-C dataset, including natural noise, digital noise, weather noise, and blur noise. In addition, we propose test-time augmentation (TTA) to improve calibration by increasing prediction diversity and reducing overconfidence. We conduct ~ 600 simulations, and the experimental results show that ViT-B/32 achieves higher accuracy (ACC) and lower expected calibration error (ECE) than ResNet50 in in-distribution settings (e.g., 94.66% ACC vs. 62.37% ACC, 0.59% ECE vs. 1.03% ECE on brightness). However, ResNet50 outperforms ViT-B/32 in out-of-domain (OOD) situations (e.g., 3.14% ECE vs. 20.93% ECE on defocus, fine-tuned on glass). By applying TTA, we reduce ViT-B/32’s ECE to 13.07% on defocus (fine-tuned on glass), demonstrating its effectiveness. Our results highlight the importance of calibration in cross-modal learning and provide a simple yet effective solution for noisy and OOD calibration. Our code is available at: https://github.com/AIPMLab/CLIPSimulationsGao.
Yihang Wu, Ahmad Chaddad
SMC2
2025 Impact of domain adaptation in deep learning for medical image classifications
abstract
Domain adaptation (DA) is a quickly expanding area in machine learning that involves adjusting a model trained in one domain to perform well in another domain. While there have been notable progressions, the fundamental concept of numerous DA methodologies has persisted: aligning the data from various domains into a shared feature space. In this space, knowledge acquired from labeled source data can improve the model training on target data that lacks sufficient labels. In this study, we demonstrate the use of 10 deep learning models to simulate common DA techniques and explore their application in four medical image datasets. We have considered various situations such as multi-modality, noisy data, federated learning (FL), interpretability analysis, and classifier calibration. The experimental results indicate that using DA with ResNet34 in a brain tumor (BT) data set results in an enhancement of 4.7% in model performance. Similarly, the use of DA can reduce the impact of Gaussian noise, as it provides ∼3% accuracy increase using ResNet34 on a BT dataset. Furthermore, simply introducing DA into FL framework shows limited potential (e.g., ∼ 0.3% increase in performance) for skin cancer classification. In addition, the DA method can improve the interpretability of the models using the gradcam++ technique, which offers clinical values. Calibration analysis also demonstrates that using DA provides a lower expected calibration error (ECE) value ∼2% compared to CNN alone on a multi-modality dataset. The codes for our experiments are available at https://github.com/AIPMLab/Domain_Adaptation.
Yihang Wu, Ahmad Chaddad
SMC1
2025 FAA-CLIP: Federated Adversarial Adaptation of CLIP
abstract
Despite the remarkable performance of vision language models (VLMs), such as contrastive language image pretraining (CLIP), the large size of these models is a considerable obstacle to their use in federated learning (FL) systems where the parameters of local client models need to be transferred to a global server for aggregation. Another challenge in FL is the heterogeneity of data from different clients, which affects the generalization performance of the solution. In addition, natural pretrained VLMs exhibit poor generalization ability in the medical datasets, suggests there exists a domain gap. To solve these issues, we introduce a novel method for the federated adversarial adaptation (FAA) of CLIP. Our method, named FAA-CLIP, handles the large communication costs of CLIP using a lightweight feature adaptation module (FAM) for aggregation, effectively adapting this VLM to each client’s data while greatly reducing the number of parameters to transfer. By keeping CLIP frozen and only updating the FAM parameters, our method is also computationally efficient. Unlike existing approaches, our FAA-CLIP method directly addresses the problem of domain shifts across clients via a domain adaptation (DA) module. This module employs a domain classifier to predict if a given sample is from the local client or the global server, allowing the model to learn domain-invariant representations. Extensive experiments on six different datasets containing both natural and medical images demonstrate that FAA-CLIP can generalize well on both natural and medical datasets compared to recent FL approaches. Our codes are available athttps://github.com/AIPMLab/FAA-CLIP.
Yihang Wu, Ahmad Chaddad, Christian Desrosiers, Tareef S. Daqqaq, Reem Kateb
IEEE Internet Things J.1
2025 Domain adaptation techniques for natural and medical image classification
Ahmad Chaddad, Yihang Wu, Reem Kateb, Christian Desrosiers
Inf. Sci.2
2025 Enhancing dual network based semi-supervised medical image segmentation with uncertainty-guided pseudo-labeling
Yunyao Lu, Yihang Wu, Ahmad Chaddad, Tareef S. Daqqaq, Reem Kateb
Knowl. Based Syst.2
2025 Reliable and Private Utility Signaling for Data Markets
abstract
The explosive growth of data has highlighted its critical role in driving economic growth through data marketplaces, which enable extensive data sharing and access to high-quality datasets. To support effective trading, signaling mechanisms provide participants with information about data products before transactions, enabling informed decisions and facilitating trading. However, due to the inherent free-duplication nature of data, commonly practiced signaling methods face a dilemma between privacy and reliability, undermining the effectiveness of signals in guiding decision-making. To address this, this paper explores the benefits and develops a non-TCP-based construction for a desirable signaling mechanism that simultaneously ensures privacy and reliability. We begin by formally defining the desirable utility signaling mechanism and proving its ability to prevent suboptimal decisions for both participants and facilitate informed data trading. To design a protocol to realize its functionality, we propose leveraging maliciously secure multi-party computation (MPC) to ensure the privacy and robustness of signal computation and introduce an MPC-based hash verification scheme to ensure input reliability. In multi-seller scenarios requiring fair data valuation, we further explore the design and optimization of the MPC-based KNN-Shapley method with improved efficiency. Rigorous experiments demonstrate the efficiency and practicality of our approach.
Jiayao Zhang 0006, Yihang Wu, Jinfei Liu, Zheng Yan 0002, Kui Ren 0001, Lei Zhang 0006, Lin Qu
Proc. ACM Manag. Data3
2025 An adaptive network construction for single-cell clustering
Yanmei Hu, Yihang Wu, Yingxi Zhang, Bin Duo, Xiaochuan Tang, Xiangtao Li
J. Supercomput.3
2024 When Data Pricing Meets Non-Cooperative Game Theory
abstract
Driven by the growing field of data intelligence, data market emerges as a promising paradigm for data exchange, enabling the full utilization of data. Data pricing is an essential function in data market that reflects the values or cost of data and is dependent on interactions among multiple market participants including data buyers, data sellers, and data brokers. Game theory presents a promising approach to model the multi-participant interplay in data pricing, yet challenged by the specific nature of data. In this paper, we present a blueprint for applying game theory to data pricing. From a game-theoretic perspective, we highlight the unique characteristics of data (compared to traditional goods) and suggest important desiderata for effective data pricing. We identify four key dimensions (Participant, Object, Action, and Information) to understand the landscape of game theory based data pricing. Within each dimension, data-specific challenges and research gaps are identified. Our work establishes a foundational understanding of data pricing through the lens of game theory and opens up promising research directions in this developing field.
Yuran Bi, Yihang Wu, Jinfei Liu, Kui Ren 0001, Li Xiong 0001
ICDE2
2024 FACMIC: Federated Adaptative CLIP Model for Medical Image Classification
Yihang Wu, Christian Desrosiers, Ahmad Chaddad
MICCAI (12)1
2024 Federated Learning for Healthcare Applications
abstract
Due to the fast advancement of artificial intelligence (AI), centralized-based models have become critical for healthcare tasks like in medical image analysis and human behavior recognition. Although these models exhibit suitable performance, they are frequently constrained by privacy concerns. To attenuate this, a centralized learning strategy cannot be used in cases where there is a risk of data privacy breach, particularly in healthcare centers. Federated learning (FL) is a technique that allows for training a global model without sharing data by training distributed local models and aggregating them. By implementing FL throughout the training process, we can obtain a model with comparable generalization abilities to centralized learning while maintaining data privacy. This survey provides an introduction to the fundamental concepts and categories of FL, highlights the limitations of the centralized healthcare model, and discusses how FL can address these constraints. We also provide a detailed overview of the healthcare applications using FL models, along with commonly used evaluation metrics and public data sets. In this context, we have implemented a case study to demonstrate how FL can be applied in the healthcare field. Furthermore, we outline the key challenges and future trends in FL.
Ahmad Chaddad, Yihang Wu, Christian Desrosiers
IEEE Internet Things J.2
2024 A graph convolutional network model based on regular equivalence for identifying influential nodes in complex networks
Yihang Wu, Yanmei Hu, Siyuan Yin, Xiaochuan Tang, Xiangtao Li
Knowl. Based Syst.1
2023 Enhancing Classification Tasks through Domain Adaptation Strategies
abstract
Domain adaptation (DA) is a technique that uses the knowledge from similar data sets to enhance the generalizability of a model, which proves to be effective in addressing the issue of limited training data. However, due to the complexity of medical data, most advances have occurred in the natural domain and not in the medical field. Furthermore, the majority of advancements are derived from conventional DA datasets, which may introduce result bias. This article presents a comprehensive new analysis of four widely used DA algorithms, tested on more realistic medical data sets to assess the potential of these DA techniques. The examination covers an evaluation of their model performance, discrepancies in data distribution, and the interpretability of the models. For example, the Deep subdomain adaptation network (DSAN) achieves a high level of accuracy on the COVID-19 dataset (89.9%) using Resnet34. Furthermore, the interpretability of the DSAN’s prediction using the skin cancer dataset is implausible. In conclusion, we offer perspectives on the outcomes obtained. Our codes are available at https://github.com/AIPMLab/Domain_Adaptation.
Ahmad Chaddad, Yihang Wu
BIBM2
2023 A Practical Simulation for Domain Adaptation Models
abstract
Domain adaptation (DA) techniques have emerged in machine learning, aiming to alleviate distribution differences between training and test sets using information from related domains. However, due to the intricacies of medical data, most advances have been made in the natural rather than medical field. In this paper, we provide a case study on how DA can be applied to the medical domain using a brain tumor data set. We present a full analysis of four widely used DA algorithms in the natural domain, with a focus on their application to the medical domain. The experiments provide detailed results and insightful observations that highlight the performance and medical applicability of these techniques.
Ahmad Chaddad, Yihang Wu
HealthCom2
2023 Building a Better Metaverse: How Federated Learning is Revolutionizing Virtual Worlds
abstract
The rapid development of artificial intelligence (AI) has been made possible by the swift advancement of technology. Recently, there has been a significant increase in interest in the concept of the Metaverse, which is viewed as the ultimate manifestation of the internet. However, building a vast Metaverse requires a substantial allocation of resources to maintain it. Such circumstances pose several challenges, including the disclosure of sensitive data and the problem of non-independent and homogeneous data distribution. In this study, we will focus on explaining federated learning (FL), a methodology that addresses these issues by training localized models and then consolidating them on a global scale to create the ultimate model without sharing the original data. Additionally, we will also discuss the current challenges and potential developments that should be considered in the upcoming years.
Ahmad Chaddad, Yihang Wu, Reem Kateb
HealthCom2
2023 Domain Adaptation in Machine Learning: A Practical Simulation Study
abstract
Domain adaptation (DA) is a critical technique in machine learning, designed to alleviate distribution differences between training and test sets by leveraging information from similar datasets. This paper presents a practical simulation of both widely-used and recent DA techniques, with a specific focus on unsupervised learning scenarios where labels are only available in the source domain. We thoroughly investigate the impact of these methods on various publicly available datasets, providing detailed explanations of our findings. Furthermore, we conduct an ablation study and elaborate extensively on the simulation results. Our study offers valuable information on the resilience of selected DA algorithms, highlighting the importance of appropriate datasets and neural network architectures with training methodologies to achieve optimal performance in DA tasks.
Ahmad Chaddad, Yihang Wu
ICTAI2
2023 Robust Tracking for Electromagnetic-Actuated Microrobot Via Sliding Mode Control
abstract
Microrobots can work in small, enclosed spaces and complete complex tasks. It has potential applications in the field of biomedical and life sciences. This paper focuses on investigating robust dynamic tracking for electromagnetic-actuated microrobot via sliding mode control. First, we generalize the planar motion to the case of spatial trajectory tracking and then establish the three-dimensional tracking error model of microrobot. Then, we deploy a novel sliding control strategy and ensure the stability of sliding mode dynamics, once the robotic state reaches the sliding surface and sustains there in the sequel. Finally, the numerical simulation is conducted, compared to the traditional proportion-integration-differentiation control strategy, to verify the effectiveness and advantages of our proposed control scheme.
Yihang Wu, Ziyao Qu, Xiaocong Chang, Zhijian Hu, Rongni Yang, Renjie Ma
IECON1
2022 LabVIEW Based Construction and Decoding for 2/3 Polar Codes
abstract
Polar codes have been proved to be a coding method that reach the limit of Shannon channel capacity. In the past ten years, polar codes have attracted wide attention in academia and industry, such that the 5th generation wireless systems (5G) standardization process of the 3rd generation partnership project (3GPP) chose polar codes as a channel coding scheme.However, in the practical communication system, the channel conditions change from time to time, which leads to the uncertainty of the transmitted data. In order to improve the transmission rate and reliability, how to construct the polar code with rate compatible coding has become an urgent problem. In this papar, based on the discussion of the construction methods of 2/3 polar codes, we select the shortening scheme as the rate compatible coding rate and achieve the encoding and decoding based on LabVIEW .2/3 polar codes will be widely used in ground network and non-ground network in control channel.
Yihang Wu, Shuai Han 0002, Guoning Zhi
IWCMC1