VLDB 2026 Research / reviewers in the wild / expert
Chao Wu 0001
dblp:45/3158-1
· DBLP profile ↗
64ranked-venue papers
5as first author
51since 2021 · last 2026
0000-0003-0885-6869ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 2 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 23 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Editing as Unlearning: Are Knowledge Editing Methods Strong Baselines for Large Language Model Unlearning?abstractLarge language Model (LLM) unlearning, i.e., selectively removing information from LLMs, is vital for responsible model deployment. Differently, LLM knowledge editing aims to modify LLM knowledge instead of removing it. Though editing and unlearning seem to be two distinct tasks, we find there is a tight connection between them. In this paper, we conceptualize unlearning as a special case of editing where information is modified to a refusal or "empty set" response, signifying its removal. This paper thus investigates if knowledge editing techniques are strong baselines for LLM unlearning. We evaluate state-of-the-art (SOTA) editing methods (e.g., ROME, MEMIT, GRACE, WISE, and AlphaEdit) against existing unlearning approaches on pretrained and finetuned knowledge. Results show certain editing methods, notably WISE and AlphaEdit, are effective unlearning baselines, especially for pretrained knowledge, and excel in generating human-aligned refusal answers. To better adapt editing methods for unlearning applications, we propose practical recipes including self-improvement and query merging. The former leverages the LLM's own in-context learning ability to craft a more human-aligned unlearning target, and the latter enables ROME and MEMIT to perform well in unlearning longer sample sequences. We advocate for the unlearning community to adopt SOTA editing methods as baselines and explore unlearning from an editing perspective for more holistic LLM memory control. Zexi Li 0001, Xiangzhu Wang, William F. Shen, Meghdad Kurmanji, Xinchi Qiu, Dongqi Cai 0001, Chao Wu 0001, Nicholas D. Lane |
AAAI | 7 |
| 2026 | Improving Model Fusion by Training-Time Neuron Alignment With Fixed Neuron AnchorsabstractModel fusion aims to integrate several deep neural network (DNN) models' knowledge into one by fusing parameters, and it has promising applications, such as improving the generalization of foundation models and parameter averaging in federated learning. However, models under different settings (data, hyperparameter, etc.) have diverse neuron permutations; in other words, from the perspective of loss landscape, they reside in different loss basins, thus hindering model fusion performances. To alleviate this issue, previous studies highlighted the role of permutation invariance and have developed methods to find correct network permutations for neuron alignment after training. Orthogonal to previous attempts, this paper studies training-time neuron alignment, improving model fusion without the need for post-matching. Training-time alignment is cheaper than post-alignment and is applicable in various model fusion scenarios. Starting from fundamental hypotheses and theorems, a simple yet lossless algorithm called TNA-PFN is introduced. TNA-PFN utilizes partially fixed neuron weights as anchors to reduce the potential of training-time permutations, and it is empirically validated in reducing the barriers of linear mode connectivity and multi-model fusion. It is also validated that TNA-PFN can improve the fusion of pretrained models under the setting of model soup (vision transformers) and ColD fusion (pretrained language models). Based on TNA-PFN, two federated learning methods, FedPFN and FedPNU, are proposed, showing the prospects of training-time neuron alignment. FedPFN and FedPNU reach state-of-the-art performances in federated learning under heterogeneous settings and can be compatible with the server-side algorithm. Zexi Li 0001, Zhiqi Li 0004, Tao Shen 0002, Jun Xiao 0001, Yike Guo, Tao Lin 0004, Chao Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | You are Your Own Best Teacher: Achieving Centralized-level Performance in Federated Learning under Heterogeneous and Long-Tailed DataabstractData heterogeneity, stemming from local non-IID data and global long-tailed distributions, is a major challenge in federated learning (FL), leading to significant performance gaps compared to centralized learning. Previous research found that poor representations and biased classifiers are the main problems and proposed neural-collapse-inspired synthetic simplex ETF to help representations be closer to neural collapse optima. However, we find that the neural-collapse-inspired methods are not strong enough to reach neural collapse and still have huge gaps to centralized training. In this paper, we rethink this issue from a self-bootstrap perspective and propose FedYoYo (You Are Your Own Best Teacher), introducing Augmented Self-bootstrap Distillation (ASD) to improve representation learning by distilling knowledge between weakly and strongly augmented local samples, without needing extra datasets or models. We further introduce Distribution-aware Logit Adjustment (DLA) to balance the self-bootstrap process and correct biased feature representations. FedYoYo nearly eliminates the performance gap, achieving centralized-level performance even under mixed heterogeneity. It enhances local representation learning, reducing model drift and improving convergence, with feature prototypes closer to neural collapse optimality. Extensive experiments show FedYoYo achieves state-of-the-art results, even surpassing centralized logit adjustment methods by 5.4\% under global long-tailed settings. Shanshan Yan, Zexi Li 0001, Chao Wu 0001, Yang Lu 0009, Yan Yan 0001, Hanzi Wang |
ICCV | 3 |
| 2025 | REMEDY: Recipe Merging Dynamics in Large Vision-Language ModelsabstractModel merging has emerged as a powerful technique for combining task-specific vision models into a unified and multi-functional model. Previous methods represented by task arithmetic, have demonstrated effectiveness and scalability in this domain. When large vision-language models (LVLMs) arise with model size scaling up, this design becomes challenging to fuse different instruction-tuned LVLMs for generalization enhancement. The large scale and multi-modal nature of LVLMs present unique obstacles, including constructing reusable and modular components to accommodate the multi-component architecture of LVLMs and the requirement for dynamic fusion based on multi-modal input tokens. To address these challenges, we propose the \textbf{RE}cipe \textbf{ME}rging \textbf{DY}namics (REMEDY) method, a scalable and flexible paradigm for model merging in LVLMs. We first define reusable modules termed \textit{recipes} including the projector and shallow LLM layers, enhancing visual-language understanding. Then, we introduce a modality-aware allocator dynamically generates weights in a one-shot manner based on input relevance to existing recipes, enabling efficient cross-modal knowledge integration. REMEDY thus offers an adaptive solution for LVLMs to tackle both seen (i.e., multi-task learning) and unseen (i.e., zero-shot generalization) tasks. Experimental results demonstrate that our method consistently improves performance on both seen and unseen tasks, underscoring the effectiveness of REMEDY in diverse multi-modal scenarios. Didi Zhu, Yibing Song, Tao Shen 0002, Ziyu Zhao 0001, Jinluan Yang, Min Zhang 0068, Chao Wu 0001 |
ICLR | 7 |
| 2025 | Leveraging Pretrained Diffusion Models for Zero-Shot Part Assemblyabstract3D part assembly aims to understand part relationships and predict their 6-DoF poses to construct realistic 3D shapes, addressing the growing demand for autonomous assembly, which is crucial for robots. Existing methods mainly estimate the transformation of each part by training neural networks under supervision, which requires a substantial quantity of manually labeled data. However, the high cost of data collection and the immense variability of real-world shapes and parts make traditional methods impractical for large-scale applications. In this paper, we propose first a zero-shot part assembly method that utilizes pre-trained point cloud diffusion models as discriminators in the assembly process, guiding the manipulation of parts to form realistic shapes. Specifically, we theoretically demonstrate that utilizing a diffusion model for zero-shot part assembly can be transformed into an Iterative Closest Point (ICP) process. Then, we propose a novel pushing-away strategy to address the overlap parts, thereby further enhancing the robustness of the method. To verify our work, we conduct extensive experiments and quantitative comparisons to several strong baseline methods, demonstrating the effectiveness of the proposed approach, which even surpasses the supervised learning method. The code has been released on https://github.com/Ruiyuan-Zhang/Zero-Shot-Assembly. Ruiyuan Zhang, Qi Wang 0111, Yuchi Huo, Chao Wu 0001 |
IJCAI | 5 |
| 2025 | FedGuCci: Making Local Models More Connected in Landscape for Federated LearningabstractFederated learning (FL) involves multiple heterogeneous clients collaboratively training a global model via iterative local updates and model fusion.The generalization of FL's global model has a large gap compared with centralized training, which is its bottleneck for broader applications.In this paper, we study and improve FL's generalization through a fundamental "connectivity" perspective, which means how the local models are connected in the parameter region and fused into a generalized global model.The term "connectivity" is derived from linear mode connectivity (LMC), studying the interpolated loss landscape of two different solutions (e.g., modes) of neural networks.Bridging the gap between LMC and FL, in this paper, we leverage fixed anchor models to empirically and theoretically study the transitivity property of connectivity from two models (LMC) to a group of models (model fusion in FL).Based on the findings, we propose FedGuCci(+), improving group connectivity for better generalization.It is shown that our methods can boost the generalization of FL under client heterogeneity across various tasks (4 CV datasets and 6 NLP datasets) and model architectures (e.g., ViTs and PLMs).The code is available here: FedGuCci Codebase. Zexi Li 0001, Zhiqi Li 0004, Didi Zhu, Tao Shen 0002, Tao Lin 0004, Chao Wu 0001, Nicholas D. Lane |
KDD (2) | 7 |
| 2025 | Towards Generalizable Detector for Generated ImageabstractThe effective detection of generated images is crucial to mitigate potential risks associated with their misuse. Despite significant progress, a fundamental challenge remains: ensuring the generalizability of detectors. To address this, we propose a novel perspective on understanding and improving generated image detection, inspired by the human cognitive process: Humans identify an image as unnatural based on specific patterns because these patterns lie outside the space spanned by those of natural images. This is intrinsically related to out-of-distribution (OOD) detection, which identifies samples whose semantic patterns (i.e., labels) lie outside the semantic pattern space of in-distribution (ID) samples.
By treating patterns of generated images as OOD samples, we demonstrate that models trained merely over natural images bring guaranteed generalization ability under mild assumptions.
This transforms the generalization challenge of generated image detection into the problem of fitting natural image patterns.
Based on this insight, we propose a generalizable detection method through the lens of ID energy. Theoretical results capture the generalization risk of the proposed method. Experimental results across multiple benchmarks demonstrate the effectiveness of our approach. Qianshu Cai, Chao Wu 0001, Yonggang Zhang 0003, Jun Yu 0002, Xinmei Tian 0001 |
NeurIPS | 2 |
| 2025 | Utilizing RBC system for taxation policy evaluation: An adaptive interaction framework based on deep reinforcement learning
Shunyu Liu 0001, Tianrun A. Cai, Chao Wu 0001 |
Expert Syst. Appl. | 4 |
| 2025 | FedMcon: an adaptive aggregation method for federated learning via meta controllerabstractFederated learning (FL) emerged as a novel machine learning setting that enables collaboratively training deep models on decentralized clients with privacy constraints. In the vanilla federated averaging algorithm (FedAvg), the global model is generated by the weighted linear combination of local models, and the weights are proportional to the local data sizes. This methodology, however, encounters challenges when facing heterogeneous and unknown client data distributions, often leading to discrepancies from the intended global objective. The linear combination-based aggregation often fails to address the varied dynamics presented by diverse scenarios, settings, and data distributions inherent in FL, resulting in hindered convergence and compromised generalization. In this paper, we present a new aggregation method, FedMcon, within a framework of meta-learning for FL. We introduce a learnable controller trained on a small proxy dataset and served as an aggregator to learn how to adaptively aggregate heterogeneous local models into a better global model toward the desired objective. The experimental results indicate that the proposed method is effective on extremely non-independent and identically distributed data and it can simultaneously reach 19 times communication speedup in a single FL setting. Tao Shen 0002, Zexi Li 0001, Ziyu Zhao 0001, Didi Zhu, Zheqi Lv, Kun Kuang 0001, Shengyu Zhang 0001, Chao Wu 0001, Fei Wu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 8 |
| 2025 | FediOS: decoupling orthogonal subspaces for personalization in feature-skew federated learning
Lingzhi Gao, Zexi Li 0001, Xinyi Shang, Yang Lu 0009, Chao Wu 0001 |
Mach. Learn. | 5 |
| 2025 | Effective Two-Stage Double Auction for Dynamic Resource Provision Over Edge Networks via Discovering the Power of OverbookingabstractTo facilitate responsive and cost-effective computing service delivery over edge networks, this paper investigates a novel two-stage double auction methodology via discovering an interesting idea of resource overbooking to overcome dynamic and uncertain nature of supply of edge servers (sellers) and demand generated from mobile devices (as buyers). The proposed auction integrates multiple essential goals such as maximizing social welfare as well as accelerating the decision-making process from both short-term and long-term perspectives (e.g., the time required to determine winning seller-buyer pairs), by introducing a stagewise strategy: an overbooking-driven pre-double auction (OPDAuction) for determining long-term cooperations between sellers and buyers before practical resource transactions as Stage I, and a real-time backup double auction (RBDAuction) for quickly coping with residual resource demands during actual transactions. In particular, by embedding a proper overbooking rate, OPDAuction helps with facilitating trading contracts between appropriate sellers and buyers as guidance for future transactions, by allowing the booked resources to exceed theoretical supply. Then, since pre-auctions may cause risks, our RBDAuction adjusts to real-time market changes, further enhancing the overall social welfare. More importantly, we offer an interesting view to show that our proposed two-stage auction can support significant design properties such as truthfulness, individual rationality, and budget balance. Extensive experiments demonstrate that our TwoSAuction achieves up to 76.8% reduction in decision-making time compared to conventional double auctions when considering 150 buyers and 25 sellers, while maintaining superior performance in social welfare and computational scalability over dynamic edge settings. Sicheng Wu, Minghui LiWang, Deqing Wang 0004, Xianbin Wang 0001, Chao Wu 0001, Junyi Tang, Li Li 0008, Xiaoyu Xia 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2024 | Scalable Geometric Fracture Assembly via Co-creation Space among AssemblersabstractGeometric fracture assembly presents a challenging practical task in archaeology and 3D computer vision. Previous methods have focused solely on assembling fragments based on semantic information, which has limited the quantity of objects that can be effectively assembled. Therefore, there is a need to develop a scalable framework for geometric fracture assembly without relying on semantic information. To improve the effectiveness of assembling geometric fractures without semantic information, we propose a co-creation space comprising several assemblers capable of gradually and unambiguously assembling fractures. Additionally, we introduce a novel loss function, i.e., the geometric-based collision loss, to address collision issues during the fracture assembly process and enhance the results. Our framework exhibits better performance on both PartNet and Breaking Bad datasets compared to existing state-of-the-art frameworks. Extensive experiments and quantitative comparisons demonstrate the effectiveness of our proposed framework, which features linear computational complexity, enhanced abstraction, and improved generalization. Our code is publicly available at https://github.com/Ruiyuan-Zhang/CCS. Ruiyuan Zhang, Zexi Li 0001, Hao Dong 0003, Jie Fu 0001, Chao Wu 0001 |
AAAI | 6 |
| 2024 | Distributionally Generative Augmentation for Fair Facial Attribute ClassificationabstractFacial Attribute Classification (FAC) holds substantial promise in widespread applications. However, FAC models trained by traditional methodologies can be unfair by exhibiting accuracy inconsistencies across varied data sub-populations. This unfairness is largely attributed to bias in data, where some spurious attributes (e.g., Male) statistically correlate with the target attribute (e.g., Smiling). Most of existing fairness-aware methods rely on the labels of spurious attributes, which may be unavailable in practice. This work proposes a novel, generation-based two-stage framework to train a fair FAC model on biased data without additional annotation. Initially, we identify the potential spurious attributes based on generative models. Notably, it enhances interpretability by explicitly showing the spurious attributes in image space. Following this, for each image, we first edit the spurious attributes with a random degree sampled from a uniform distribution, while keeping target attribute unchanged. Then we train a fair FAC model by fostering model invariance to these augmentation. Extensive experiments on three common datasets demonstrate the effectiveness of our method in promoting fairness in FAC without compromising accuracy. Codes are in https://github.com/heqianpei/DiGA. Fengda Zhang, Qianpei He, Kun Kuang 0001, Long Chen 0016, Chao Wu 0001, Jun Xiao 0001, Hanwang Zhang |
CVPR | 6 |
| 2024 | More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMsabstractChengyuan Liu, Yangyang Kang, Shihang Wang, Lizhi Qing, Fubang Zhao, Chao Wu, Changlong Sun, Kun Kuang, Fei Wu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yangyang Kang, Lizhi Qing, Fubang Zhao, Chao Wu 0001, Changlong Sun, Kun Kuang 0001, Fei Wu 0001 |
EMNLP | 6 |
| 2024 | Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language ModelsabstractCatastrophic forgetting emerges as a critical challenge when fine-tuning multi-modal large language models (MLLMs), where improving performance on unseen tasks often leads to a significant performance drop on the original tasks. This paper presents a comprehensive analysis of catastrophic forgetting in MLLMs and introduces a post-training adjustment method called Model Tailor. Our method primarily preserves the pre-trained parameters while replacing a small number ($\leq$ 10%) of fine-tuned parameters, maintaining $\sim$ 99% effectiveness on original tasks versus pre-training, and achieving $\sim$ 97% on new tasks compared to standard fine-tuning. Specifically, we derive a sparse mask to identify the model patch, based on a fusion strategy that integrates salience and sensitivity analysis. Subsequently, a compensation mechanism is introduced to decorate the patch, enhancing the model’s performance on both target and original tasks. Additionally, our method is adaptable to multi-task scenarios. Through extensive experiments on InstructBLIP and LLaVA-1.5 in both image captioning and visual question answering tasks, our approach demonstrates significant task adaptability while preserving inherent pre-trained capabilities. Didi Zhu, Zhongyi Sun 0002, Zexi Li 0001, Tao Shen 0002, Shouhong Ding, Chao Wu 0001, Kun Kuang 0001 |
ICML | 7 |
| 2024 | Feature Transformation for Few-Shot LearningabstractThe goal of few-shot learning is to classify new classes of samples with a few labeled training samples. State-of-the-art few-shot learners train a backbone on sufficient datasets and use its extracted features for classification. However, biased data distributions can lead to severe overfitting in few-shot learning. In this paper, we propose a novel feature transformation method that utilizes the statistical characteristics of sufficient data to perform feature transformation on few-shot data to alleviate overfitting caused by biased data distributions. We show an interesting phenomenon that removing the component along the mean feature of the base classes in meta-testing improves the performance for few-shot learning. Our method can be used on off-the-shelf pretrained feature extractors without extra parameters. We show that our method achieves the new state-of-the-art accuracy in the prototype-based method and comparable accuracy with state-of-the-art accuracy in the optimization-based method. Peizheng Wang, Qifei Zhang 0001, Jie Zhang 0081, Gang Li 0050, Chao Wu 0001 |
IJCNN | 5 |
| 2024 | Neural Collapse Anchored Prompt Tuning for Generalizable Vision-Language ModelsabstractLarge-scale vision-language (V-L) models have demonstrated remarkable generalization capabilities for downstream tasks through prompt tuning. However, the mechanisms behind the learned text representations are unknown, limiting further generalization gains, and the limitations are more severe when faced with the prevalent class imbalances seen in web-sourced datasets. Recent advances in the neural collapse (NC) phenomenon of vision-only models suggest that the optimal representation structure is the simplex ETF, which paves the way to study representations in V-L models. In this paper, we make the first attempt to use NC for examining the representations in V-L models via prompt tuning. It is found that NC optimality of text-to-image representations shows a positive correlation with downstream generalizability, which is more severe under class imbalance settings. To improve the representations, we propose Neural-collapse-anchored Prompt Tuning (NPT), a novel method that learns prompts with text and image representations that satisfy the same simplex Equiangular Tight Frame (ETF). NPT incorporates two regularization terms: language-modality collapse and multi-modality isomorphism; and it is compatible with other prompt tuning methods. Extensive experiments show that NPT can consistently help to improve existing prompt tuning techniques across 11 datasets for both balanced and imbalanced settings. Didi Zhu, Zexi Li 0001, Min Zhang 0068, Junkun Yuan, Kun Kuang 0001, Chao Wu 0001 |
KDD | 7 |
| 2024 | Advancing Prompt Learning through an External LayerabstractPrompt learning represents a promising method for adapting pre-trained vision-language models (VLMs) to various downstream tasks by learning a set of text embeddings. One challenge inherent to these methods is the poor generalization performance due to the invalidity of the learned text embeddings for unseen tasks. A straightforward approach to bridge this gap is to freeze the text embeddings in prompts, which results in a lack of capacity to adapt VLMs for downstream tasks. To address this dilemma, we propose a paradigm called EnPrompt with a novel External Layer (EnLa). Specifically, we propose a textual external layer and learnable visual embeddings for adapting VLMs to downstream tasks. The learnable external layer is built upon valid embeddings of pre-trained CLIP. This design considers the balance of learning capabilities between the two branches. To align the textual and visual features, we propose a novel two-pronged approach: i) we introduce the optimal transport as the discrepancy metric to align the vision and text modalities, and ii) we introduce a novel strengthening feature to enhance the interaction between these two modalities. Four representative experiments (i.e., base-to-novel generalization, few-shot learning, cross-dataset generalization, domain shifts generalization) across 15 datasets demonstrate that our method outperforms the existing prompt learning method. Fangming Cui, Xun Yang 0001, Chao Wu 0001, Liang Xiao 0007, Xinmei Tian 0001 |
ACM Multimedia | 3 |
| 2024 | Walking is Matter: A Benchmark for Fine-Grained Gait Segmentation
Zhongguang Zhang, Wenzhu Xu, Qifei Zhang 0001, Chao Wu 0001 |
PRCV (11) | 6 |
| 2024 | Multi-agent Continuous Control with Generative Flow Networks
Yinchuan Li, Shunyu Liu 0001, Xu Zhang 0011, Yunfeng Shao 0001, Chao Wu 0001 |
Neural Networks | 6 |
| 2024 | Towards Effective Clustered Federated Learning: A Peer-to-Peer Framework With Adaptive Neighbor MatchingabstractIn federated learning (FL), clients may have diverse objectives, and merging all clients' knowledge into one global model will cause negative transfer to local performance. Thus, clustered FL is proposed to group similar clients into clusters and maintain several global models. In the literature, centralized clustered FL algorithms require the assumption of the number of clusters and hence are not effective enough to explore the latent relationships among clients. In this paper, without assuming the number of clusters, we propose a peer-to-peer (P2P) FL algorithm namedPANM. InPANM, clients communicate with peers to adaptively form an effective clustered topology. Specifically, we present two novel metrics for measuring client similarity and a two-stage neighbor matching algorithm based Monte Carlo method and Expectation Maximization under the Gaussian Mixture Model assumption. We have conducted theoretical analyses ofPANMon the probability of neighbor estimation and the error gap to the clustered optimum. We have also implemented extensive experiments under both synthetic and real-world clustered heterogeneity. Theoretical analysis and empirical experiments show that the proposed algorithm is superior to the P2P FL counterparts, and it achieves better performance than the centralized cluster FL method.PANMis effective even under extremely low communication budgets. Zexi Li 0001, Jiaxun Lu, Didi Zhu, Yunfeng Shao 0001, Yinchuan Li, Yongheng Wang, Chao Wu 0001 |
IEEE Trans. Big Data | 9 |
| 2023 | Delving into the Adversarial Robustness of Federated LearningabstractIn Federated Learning (FL), models are as fragile as centrally trained models against adversarial examples. However, the adversarial robustness of federated learning remains largely unexplored. This paper casts light on the challenge of adversarial robustness of federated learning. To facilitate a better understanding of the adversarial vulnerability of the existing FL methods, we conduct comprehensive robustness evaluations on various attacks and adversarial training methods. Moreover, we reveal the negative impacts induced by directly adopting adversarial training in FL, which seriously hurts the test accuracy, especially in non-IID settings. In this work, we propose a novel algorithm called Decision Boundary based Federated Adversarial Training (DBFAT), which consists of two components (local re-weighting and global regularization) to improve both accuracy and robustness of FL systems. Extensive experiments on multiple datasets demonstrate that DBFAT consistently outperforms other baselines under both IID and non-IID settings. Jie Zhang 0081, Bo Li 0115, Chen Chen 0043, Lingjuan Lyu, Shuang Wu 0001, Shouhong Ding, Chao Wu 0001 |
AAAI | 7 |
| 2023 | Score-PA: Score-based 3D Part Assembly
Junfeng Cheng, Mingdong Wu, Ruiyuan Zhang, Guanqi Zhan, Chao Wu 0001, Hao Dong 0003 |
BMVC | 5 |
| 2023 | TFSF: Topological and Feature Space Fusion with Spatio-Temporal Modeling for Crop Yield PredictionabstractThe interactions between climate change and geographic conditions have imposed great challenges to agriculture researchers on crop yield predictions. Traditional machine learning algorithms like Lasso and Gradient Boosting Machine often fall short in terms of accuracy. Deep learning has emerged as a promising approach in agriculture modeling: many studies have used convolutional(CNN) and recurrent neural networks(RNN) to effectively capture the nonlinear relationship between yield and various factors such as climate, soil and management. How-ever, these approaches often neglect the spatial relations among different prediction units. In this paper, we introduce Graph Neural Networks(GNNs) to incorporate spatial knowledge in crop yield forecasting. Additionally, a Topology-Feature Space Fusion Graph Neural Network(TFSF-GNN) is proposed to address the limitations of topological graphs based on geographical distance which include only the static spatial information. The network is designed to compute the similarity of meteorological and environmental characteristics in different regions and generates graph structures of the feature space. Multiple graph convolutional networks are then used to extract information from the topological space, feature space, and common space. Extensive experiments on the benchmark dataset demonstrate that our proposed approach outperforms existing network structures on county-level yield prediction tasks. Shifeng Xu, Yijing Zhou, Cuiting Huang, Chao Wu 0001 |
CSCWD | 5 |
| 2023 | No Fear of Classifier Biases: Neural Collapse Inspired Federated Learning with Synthetic and Fixed ClassifierabstractData heterogeneity is an inherent challenge that hinders the performance of federated learning (FL). Recent studies have identified the biased classifiers of local models as the key bottleneck. Previous attempts have used classifier calibration after FL training, but this approach falls short in improving the poor feature representations caused by training-time classifier biases. Resolving the classifier bias dilemma in FL requires a full understanding of the mechanisms behind the classifier. Recent advances in neural collapse have shown that the classifiers and feature prototypes under perfect training scenarios collapse into an optimal structure called simplex equiangular tight frame (ETF). Building on this neural collapse insight, we propose a solution to the FL's classifier bias problem by utilizing a synthetic and fixed ETF classifier during training. The optimal classifier structure enables all clients to learn unified and optimal feature representations even under extremely heterogeneous data. We devise several effective modules to better adapt the ETF structure in FL, achieving both high generalization and personalization. Extensive experiments demonstrate that our method achieves state-of-the-art performances on CIFAR-10, CIFAR-100, and Tiny-ImageNet. The code is available at https://github.com/ZexiLee/ICCV-2023-FedETF. Zexi Li 0001, Xinyi Shang, Tao Lin 0004, Chao Wu 0001 |
ICCV | 5 |
| 2023 | Universal Domain Adaptation via Compressive Attention MatchingabstractUniversal domain adaptation (UniDA) aims to transfer knowledge from the source domain to the target domain without any prior knowledge about the label set. The challenge lies in how to determine whether the target samples belong to common categories. The mainstream methods make judgments based on the sample features, which overemphasizes global information while ignoring the most crucial local objects in the image, resulting in limited accuracy. To address this issue, we propose a Universal Attention Matching (UniAM) framework by exploiting the self-attention mechanism in vision transformer to capture the crucial object information. The proposed framework introduces a novel Compressive Attention Matching (CAM) approach to explore the core information by compressively representing attentions. Furthermore, CAM incorporates a residual-based measurement to determine the sample commonness. By utilizing the measurement, UniAM achieves domain-wise and category-wise Common Feature Alignment (CFA) and Target Class Separation (TCS). Notably, UniAM is the first method utilizing the attention in vision transformer directly to perform classification tasks. Extensive experiments show that UniAM outperforms the current state-of-the-art methods on various benchmark datasets. Didi Zhu, Yinchuan Li, Junkun Yuan, Zexi Li 0001, Kun Kuang 0001, Chao Wu 0001 |
ICCV | 6 |
| 2023 | Target-Discriminability-Induced Multi-Source-Free Domain AdaptationabstractSource-free domain adaptation (SFDA) aims at target adaptation without access to source data, but with only pre-trained source model. Some recent works proposed to automatically combine source models with learnable weights when there are multiple pre-trained source models. In this paper, we propose a simple yet effective framework for multi-source-free domain adaptation (MSFDA). Specifically, based on discriminability towards target samples, we determine transferability of source models before adaptation and generate pseudo-labels during training. To quantify target discriminability, we introduce net confidence which refers to the probability difference between the largest and the second largest probabilities. We empirically show, on several benchmark datasets, our proposed method is competitive to the state-of-the-art methods. Gang Li 0050, Qifei Zhang 0001, Peizheng Wang, Chao Wu 0001 |
ICIP | 5 |
| 2023 | Fairness-aware Contrastive Learning with Partially Annotated Sensitive Attributes
Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Chao Wu 0001, Jun Xiao 0001 |
ICLR | 5 |
| 2023 | Federated Domain Adaptation via Pseudo-label RefinementabstractUnsupervised domain adaptation (UDA) methods usually assume data from multiple domains can be put together for centralized adaptation. Unfortunately, this assumption impairs data privacy, which leads to the failure of traditional methods in practical scenarios. To cope with the above issue, we present a novel decentralized domain adaptation approach which conducts target adaptation in an iterative training process during which only models can be delivered across domains. More specifically, to train a promising target model, we leverage Adversarial Examples (AEs) to filter out error prone predictions of source models towards each target sample based on both robustness and confidence, and then treat the most frequent prediction as the pseudo-label. Besides, to improve central model aggregation, we introduce Knowledge Contribution (KC) to compute reasonable aggregation weights. Extensive experiments conducted on several standard datasets verify the superiority of the proposed method. Gang Li 0050, Qifei Zhang 0001, Peizheng Wang, Jie Zhang 0081, Chao Wu 0001 |
ICME | 5 |
| 2023 | Learning Cautiously in Federated Learning with Noisy and Heterogeneous ClientsabstractFederated learning (FL) is a distributed framework for collaborative training with privacy guarantees. In real-world scenarios, clients may have Non-IID data (local class imbalance) with poor annotation quality (label noise). The co-existence of label noise and class imbalance in FL’s small local datasets renders conventional FL methods and noisy-label learning methods both ineffective. To address the challenges, we propose FEDCNI without using an additional clean proxy dataset. It includes a noise-resilient local solver and a robust global aggregator. For the local solver, we design a more robust prototypical noise detector to distinguish noisy samples. Further to reduce the negative impact brought by the noisy samples, we devise a curriculum pseudo labeling method and a denoise Mixup training strategy. For the global aggregator, we propose a switching re-weighted aggregation method tailored to different learning periods. Extensive experiments demonstrate our method can substantially outperform state-of-the-art solutions in mix-heterogeneous FL environments. Chenrui Wu 0002, Zexi Li 0001, Fangxin Wang 0001, Chao Wu 0001 |
ICME | 4 |
| 2023 | Revisiting Weighted Aggregation in Federated Learning with Neural NetworksabstractIn federated learning (FL), weighted aggregation of local models is conducted to generate a global model, and the aggregation weights are normalized (the sum of weights is 1) and proportional to the local data sizes. In this paper, we revisit the weighted aggregation process and gain new insights into the training dynamics of FL. First, we find that the sum of weights can be smaller than 1, causing global weight shrinking effect (analogous to weight decay) and improving generalization. We explore how the optimal shrinking factor is affected by clients' data heterogeneity and local epochs. Second, we dive into the relative aggregation weights among clients to depict the clients' importance. We develop client coherence to study the learning dynamics and find a critical point that exists. Before entering the critical point, more coherent clients play more essential roles in generalization. Based on the above insights, we propose an effective method for Federated Learning with Learnable Aggregation Weights, named as FedLAW. Extensive experiments verify that our method can improve the generalization of the global model by a large margin on different datasets and models. Zexi Li 0001, Tao Lin 0004, Xinyi Shang, Chao Wu 0001 |
ICML | 4 |
| 2023 | Transformer-Based Multi-Source Domain Adaptation Without Source DataabstractSource-free domain adaptation (SFDA) aims at target adaptation without access to source data, but with only pre-trained source model. A recent line of work proposed to automatically combine the source models with suitable weights so that it performs at least as good as the best source model when there are multiple pre-trained source models. In addition, some works proposed to preserve consistent inter-class relationships across domains to promote more shared transferable knowledge from source domains towards target adaptation. However, these works ignore the generalization ability of the pre-trained source models, which profoundly affects the initial target predictions that are crucial to the target adaptation stage. To this end, we develop a generic and effective framework based on Transformer, called TransMDA, for multi-source-free domain adaptation (MSFDA). Specifically, we inject the Transformer as the attention module into the convolutional network of each source model since it has the ability to encourage the model to turn attention towards the object regions, which can dramatically improve the model's generalization on the target domain. Furthermore, a novel pseudo-label smoothing strategy is proposed to avoid overfitting to the target domain. Experiments on several challenging datasets demonstrate the superiority of our proposed TransMDA method. Gang Li 0050, Chao Wu 0001 |
IJCNN | 2 |
| 2023 | When Masked Image Modeling Meets Source-free Unsupervised Domain Adaptation: Dual-Level Masked Network for Semantic SegmentationabstractSource-Free domain adaptive Semantic Segmentation (SFSS) aims to transfer knowledge from source domain to the target domain with only pre-trained source segmentation model and the unlabeled target dataset. Only a few works have been researched for SFSS, relying on entropy minimization, pseudo-labeling. Nevertheless, due to the domain bias, these methods tend to suffering from the confusion of classes with a similar visual appearance in different domains. To address the above issue, we propose to enhance discriminability towards target samples with masked image modeling to model spatial context relations as additional recognition clues. Specifically, we design a novel Dual-Level Masked Consistency method, which explicitly encourages the model to learn comprehensive context relations, i.e. patch-wise context and channel-wise context, on the target domain. By randomly masking target images and forcing the model to reconstruct predictions of the entire image with left unmasked part, the model has to make full use of spatially contextual information. To take a step further, we propose a novel masking strategy considering both local context and global context information by applying patch-wise masking on image patches and channel-wise masking on latent features. Notably, patch-wise context learning and channel-wise context learning can complement each other. Extensive experiments demonstrate the effectiveness of our proposed method and our method achieves state-of-the-art performance on two synthetic-to-real benchmarks: GTA5→Cityscapes and SYNTHIA→Cityscapes. Gang Li 0050, Xianzheng Ma, Hao Li 0112, Qifei Zhang 0001, Chao Wu 0001 |
ACM Multimedia | 6 |
| 2023 | Generalized Universal Domain Adaptation with Generative Flow NetworksabstractWe introduce a new problem in unsupervised domain adaptation, termed as Generalized Universal Domain Adaptation (GUDA), which aims to achieve precise prediction of all target labels including unknown categories. GUDA bridges the gap between label distribution shift-based and label space mismatch-based variants, essentially categorizing them as a unified problem, guiding to a comprehensive framework for thoroughly solving all the variants. The key challenge of GUDA is developing and identifying novel target categories while estimating the target label distribution. To address this problem, we take advantage of the powerful exploration capability of generative flow networks and propose an active domain adaptation algorithm named GFlowDA, which selects diverse samples with probabilities proportional to a reward function. To enhance the exploration capability and effectively perceive the target label distribution, we tailor the states and rewards, and introduce an efficient solution for parent exploration and state transition. We also propose a training paradigm for GUDA called Generalized Universal Adversarial Network (GUAN), which involves collaborative optimization between GUAN and GFlowNet. Theoretical analysis highlights the importance of exploration, and extensive experiments on benchmark datasets demonstrate the superiority of GFlowDA. Didi Zhu, Yinchuan Li, Yunfeng Shao 0001, Jianye Hao, Fei Wu 0001, Kun Kuang 0001, Jun Xiao 0001, Chao Wu 0001 |
ACM Multimedia | 8 |
| 2023 | Edge-cloud Collaborative Learning with Federated and Centralized FeaturesabstractFederated learning (FL) is a popular way of edge computing that does not compromise user's privacy. Current FL paradigms assume data only resides on the edge, while cloud servers only perform model averaging. However, in real-life situations such as recommender systems, the cloud server usually has abundant features and computation resources. Specifically, the cloud stores historical and interactive features, and the edge stores privacy-sensitive and real-time features. In this paper, our proposed Edge-Cloud Collaborative Knowledge Transfer Framework (ECCT) jointly utilizes the edge-side features and the cloud-side features, enabling bi-directional knowledge transfer between the two by sharing feature embeddings and prediction logits. ECCT consolidates various benefits, including enhancing personalization, enabling model heterogeneity, tolerating training asynchronization, and relieving communication burdens. Extensive experiments on public and industrial datasets demonstrate the effectiveness of ECCT. Zexi Li 0001, Qunwei Li, Yi Zhou 0017, Leon Wenliang Zhong, Chao Wu 0001 |
SIGIR | 6 |
| 2023 | PerimetryNet: A multiscale fine grained deep network for three-dimensional eye gaze estimation using visual field analysisabstractAbstract Three‐dimensional gaze estimation aims to reveal where a person is looking, which plays an important role in identifying users' point‐of‐interest in terms of the direction, attention and interactions. Appearance‐based gaze estimation methods could provide relatively unconstrained gaze tracking from commodity hardware. Inspired by medical perimetry test, we have proposed a multiscale framework with visual field analysis branch to improve estimation accuracy. The model is based on the feature pyramids and predicts vision field to help gaze estimation. In particular, we analysis the effect of the multiscale component and the visual field branch on challenging benchmark datasets: MPIIGaze and EYEDIAP. Based on these studies, our proposed PerimetryNet significantly outperforms state‐of‐the‐art methods. In addition, the multiscale mechanism and visual field branch can be easily applied to existing network architecture for gaze estimation. Related code would be available at public repository https://github.com/gazeEs/PerimetryNet . Shuqing Yu, Shuowen Zhou, Xiaosong Yang, Chao Wu 0001 |
Comput. Animat. Virtual Worlds | 5 |
| 2023 | Federated mutual learning: a collaborative machine learning method for heterogeneous data, models, and objectivesabstractFederated learning (FL) is a novel technique in deep learning that enables clients to collaboratively train a shared model while retaining their decentralized data. However, researchers working on FL face several unique challenges, especially in the context of heterogeneity. Heterogeneity in data distributions, computational capabilities, and scenarios among clients necessitates the development of customized models and objectives in FL. Unfortunately, existing works such as FedAvg may not effectively accommodate the specific needs of each client. To address the challenges arising from heterogeneity in FL, we provide an overview of the heterogeneities in data, model, and objective (DMO). Furthermore, we propose a novel framework called federated mutual learning (FML), which enables each client to train a personalized model that accounts for the data heterogeneity (DH). A “meme model” serves as an intermediary between the personalized and global models to address model heterogeneity (MH). We introduce a knowledge distillation technique called deep mutual learning (DML) to transfer knowledge between these two models on local data. To overcome objective heterogeneity (OH), we design a shared global model that includes only certain parts, and the personalized model is task-specific and enhanced through mutual learning with the meme model. We evaluate the performance of FML in addressing DMO heterogeneities through experiments and compare it with other commonly used FL methods in similar scenarios. The results demonstrate that FML outperforms other methods and effectively addresses the DMO challenges encountered in the FL setting. Tao Shen 0002, Jie Zhang 0081, Xinkang Jia, Fengda Zhang, Zheqi Lv, Kun Kuang 0001, Chao Wu 0001, Fei Wu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2023 | Federated unsupervised representation learningabstractTo leverage the enormous amount of unlabeled data on distributed edge devices, we formulate a new problem in federated learning called federated unsupervised representation learning (FURL) to learn a common representation model without supervision while preserving data privacy. FURL poses two new challenges: (1) data distribution shift (non-independent and identically distributed, non-IID) among clients would make local models focus on different categories, leading to the inconsistency of representation spaces; (2) without unified information among the clients in FURL, the representations across clients would be misaligned. To address these challenges, we propose the federated contrastive averaging with dictionary and alignment (FedCA) algorithm. FedCA is composed of two key modules: a dictionary module to aggregate the representations of samples from each client which can be shared with all clients for consistency of representation space and an alignment module to align the representation of each client on a base model trained on public data. We adopt the contrastive approach for local model training. Through extensive experiments with three evaluation protocols in IID and non-IID settings, we demonstrate that FedCA outperforms all baselines with significant margins. Fengda Zhang, Kun Kuang 0001, Long Chen 0016, Zhaoyang You, Tao Shen 0002, Jun Xiao 0001, Yin Zhang 0006, Chao Wu 0001, Fei Wu 0001, Yueting Zhuang |
Frontiers Inf. Technol. Electron. Eng. | 8 |
| 2023 | Edge-Cloud Polarization and Collaboration: A Comprehensive Survey for AIabstractInfluenced by the great success of deep learning via cloud computing and the rapid development of edge chips, research in artificial intelligence (AI) has shifted to both of the computing paradigms, i.e., cloud computing and edge computing. In recent years, we have witnessed significant progress in developing more advanced AI models on cloud servers that surpass traditional deep learning models owing to model innovations (e.g., Transformers, Pretrained families), explosion of training data and soaring computing capabilities. However, edge computing, especially edge and cloud collaborative computing, are still in its infancy to announce their success due to the resource-constrained IoT scenarios with very limited algorithms deployed. In this survey, we conduct a systematic review for both cloud and edge AI. Specifically, we are the first to set up the collaborative learning mechanism for cloud and edge modeling with a thorough review of the architectures that enable such mechanism. We also discuss potentials and practical experiences of some on-going advanced edge AI topics including pretraining models, graph neural networks and reinforcement learning. Finally, we discuss the promising directions and challenges in this field. Jiangchao Yao, Shengyu Zhang 0001, Feng Wang 0072, Jianwei Zhang 0012, Yunfei Chu, Luo Ji, Kunyang Jia, Tao Shen 0002, Anpeng Wu, Fengda Zhang, Kun Kuang 0001, Chao Wu 0001, Fei Wu 0001, Jingren Zhou 0001, Hongxia Yang |
IEEE Trans. Knowl. Data Eng. | 15 |
| 2022 | SGW-Based Multi-task Learning in Vision Tasks
Ruiyuan Zhang, Yuyao Chen, Dianbing Xi, Yuchi Huo, Chao Wu 0001 |
ACCV (4) | 7 |
| 2022 | Towards Efficient Data Free Blackbox Adversarial AttackabstractClassic black-box adversarial attacks can take advantage of transferable adversarial examples generated by a similar substitute model to successfully fool the target model. However, these substitute models need to be trained by target models' training data, which is hard to acquire due to privacy or transmission reasons. Recognizing the limited availability of real data for adversarial queries, recent works proposed to train substitute models in a data-free black-box scenario. However, their generative adversarial networks (GANs) based framework suffers from the convergence failure and the model collapse, resulting in low efficiency. In this paper, by rethinking the collaborative relationship between the generator and the substitute model, we design a novel black-box attack framework. The proposed method can efficiently imitate the target model through a small number of queries and achieve high attack success rate. The comprehensive experiments over six datasets demonstrate the effectiveness of our method against the state-of-the-art attacks. Especially, we conduct both label-only and probability-only attacks on the Microsoft Azure online model, and achieve a 100% attack success rate with only 0.46% query budget of the SOTA method [49]. Jie Zhang 0081, Bo Li 0115, Jianghe Xu, Shuang Wu 0001, Shouhong Ding, Lei Zhang 0197, Chao Wu 0001 |
CVPR | 7 |
| 2022 | Adversarial Examples for Good: Adversarial Examples Guided Imbalanced LearningabstractAdversarial examples are inputs for machine learning models that have been designed by attackers to cause the model to make mistakes. In this paper, we demonstrate that adversarial examples can also be utilized for good to improve the performance of imbalanced learning. We provide a new perspective on how to deal with imbalanced data: adjust the biased decision boundary by training with Guiding Adversarial Examples (GAEs). Our method can effectively increase the accuracy of minority classes while sacrificing little accuracy on majority classes. We empirically show, on several benchmark datasets, our proposed method is comparable to the state-of-the-art method. To our best knowledge, we are the first to deal with imbalanced learning with adversarial examples. Jie Zhang 0081, Lei Zhang 0197, Gang Li 0050, Chao Wu 0001 |
ICIP | 4 |
| 2022 | Federated Learning with Label Distribution Skew via Logits CalibrationabstractTraditional federated optimization methods perform poorly with heterogeneous data (i.e. , accuracy reduction), especially for highly skewed data. In this paper, we investigate the label distribution skew in FL, where the distribution of labels varies across clients. First, we investigate the label distribution skew from a statistical view. We demonstrate both theoretically and empirically that previous methods based on softmax cross-entropy are not suitable, which can result in local models heavily overfitting to minority classes and missing classes. Additionally, we theoretically introduce a deviation bound to measure the deviation of the gradient after local update. At last, we propose FedLC (\textbf{Fed}erated learning via \textbf{L}ogits \textbf{C}alibration), which calibrates the logits before softmax cross-entropy according to the probability of occurrence of each class. FedLC applies a fine-grained calibrated cross-entropy loss to local update by adding a pairwise label margin. Extensive experiments on federated datasets and real-world datasets demonstrate that FedLC leads to a more accurate global model and much improved performance. Furthermore, integrating other FL methods into our approach can further enhance the performance of the global model. Jie Zhang 0081, Zhiqi Li 0004, Bo Li 0115, Jianghe Xu, Shuang Wu 0001, Shouhong Ding, Chao Wu 0001 |
ICML | 7 |
| 2022 | S2RL: Do We Really Need to Perceive All States in Deep Multi-Agent Reinforcement Learning?abstractCollaborative multi-agent reinforcement learning (MARL) has been widely used in many practical applications, where each agent makes a decision based on its own observation. Most mainstream methods treat each local observation as an entirety when modeling the decentralized local utility functions. However, they ignore the fact that local observation information can be further divided into several entities, and only part of the entities is helpful to model inference. Moreover, the importance of different entities may change over time. To improve the performance of decentralized policies, the attention mechanism is used to capture features of local information. Nevertheless, existing attention models rely on dense fully connected graphs and cannot better perceive important states. To this end, we propose a sparse state based MARL (S2RL) framework, which utilizes a sparse attention mechanism to discard irrelevant information in local observations. The local utility functions are estimated through the self-attention and sparse attention mechanisms separately, then are combined into a standard joint value function and auxiliary joint value function in the central critic. We design the S2RL framework as a plug-and-play module, making it general enough to be applied to various methods. Extensive experiments on StarCraft II show that S2RL can significantly improve the performance of many state-of-the-art methods. Yinchuan Li, Jiahui Li 0003, Kun Kuang 0001, Furui Liu, Yunfeng Shao 0001, Chao Wu 0001 |
KDD | 7 |
| 2022 | DENSE: Data-Free One-Shot Federated LearningabstractOne-shot Federated Learning (FL) has recently emerged as a promising approach, which allows the central server to learn a model in a single communication round. Despite the low communication cost, existing one-shot FL methods are mostly impractical or face inherent limitations, \eg a public dataset is required, clients' models are homogeneous, and additional data/model information need to be uploaded. To overcome these issues, we propose a novel two-stage \textbf{D}ata-fre\textbf{E} o\textbf{N}e-\textbf{S}hot federated l\textbf{E}arning (DENSE) framework, which trains the global model by a data generation stage and a model distillation stage. DENSE is a practical one-shot FL method that can be applied in reality due to the following advantages:(1) DENSE requires no additional information compared with other methods (except the model parameters) to be transferred between clients and the server;(2) DENSE does not require any auxiliary dataset for training;(3) DENSE considers model heterogeneity in FL, \ie different clients can have different model architectures.Experiments on a variety of real-world datasets demonstrate the superiority of our method.For example, DENSE outperforms the best baseline method Fed-ADI by 5.08\% on CIFAR10 dataset. Jie Zhang 0081, Chen Chen 0043, Bo Li 0115, Lingjuan Lyu, Shuang Wu 0001, Shouhong Ding, Chunhua Shen, Chao Wu 0001 |
NeurIPS | 8 |
| 2022 | Research on real-time data transmission and multi-scale video image decomposition of embedded optical sensor array based on machine learning
Mingxin Cai, Chao Wu 0001 |
Multim. Tools Appl. | 3 |
| 2022 | Shuhai: A Tool for Benchmarking High Bandwidth Memory on FPGAsabstractFPGAs are starting to incorporate High Bandwidth Memory (HBM) to both reduce the memory bandwidth bottleneck encountered in some applications and to provide more capacity to store application state. However, the overall performance characteristics of HBMs are still not well understood, especially in the context of FPGAs, making it difficult to optimize designs relying on HBM. In this article, we bridge the gap between nominal specifications and actual performance by characterizing HBM on a state-of-the-art FPGA, i.e., a Xilinx Alveo U280 featuring a two-stack HBM subsystem. To this end, we have developed Shuhai, a benchmarking tool that throws light on all the subtle details of the performance and usage of HBMs on an FPGA. FPGA-based benchmarking should also provide a more accurate picture of HBM than measuring performance on CPUs/GPUs, since CPUs/GPUs are noisier systems due to their complex control logic and cache hierarchy. Since the memory itself is complex, leveraging custom hardware logic to benchmark it directly from an FPGA provides more details as well as more accurate and deterministic measurements. We observe that 1) HBM is able to provide up to 425 GB/s memory bandwidth, and 2) how HBM is used has a significant impact on the achievable throughput, which in turn demonstrates the importance of unveiling the performance characteristics of HBM so as to use HBM in the right manner. To demonstrate the generality of Shuhai, we also show results for other types of memory, e.g., DDR4, and DDR3, and quantitatively compare the performance characteristics of HBM with those of DDR4 and DDR3. Hongjing Huang, Zeke Wang, Jie Zhang 0081, Zhenhao He, Chao Wu 0001, Jun Xiao 0001, Gustavo Alonso |
IEEE Trans. Computers | 5 |
| 2022 | TICS: text-image-based semantic CAPTCHA synthesis via multi-condition adversarial learning
Xinkang Jia, Jun Xiao 0001, Chao Wu 0001 |
Vis. Comput. | 3 |
| 2021 | Web-based Platform for K-12 AI Education in ChinaabstractAs human beings are entering the era in which AI be-comes the new engine to drive social, economic, and sci-entific advancement, education is intensively being re-quired to adapt to this trend, to equip current and next generation with necessary knowledge, skills, and think-ing. Although AI education has achieved relative success in universities and cultivated a large number of talents and companies in the past decade, it hasn't made signifi-cant progress in K-12 education. We identify the key challenges as two gaps, one is about transferring practice from university education to K-12 education, and the other is about the inequal distribution of AI educational resources. To fill these gaps and to efficiently facilitate K-12 AI education, especially in countries like China, we designed and implemented a Web-based platform, which as a focal and sharing point of K-12 educational resources to provide essential AI learning and exercising components to both students and instructors. With this platform, we've successfully conducted a series of initial trials and gained positive feedbacks. We believe a wider-range application of the platform will achieve promising results for K-12 AI education. Chao Wu 0001, Yan Li 0086, Qiongdan Zhang, Fei Wu 0001 |
AAAI | 1 |
| 2021 | Evaluate the Contribution of Multiple Participants in Federated Learning
Zhaoyang You, Xinya Wu, Kexuan Chen, Chao Wu 0001 |
DEXA (2) | 5 |
| 2021 | KD3A: Unsupervised Multi-Source Decentralized Domain Adaptation via Knowledge DistillationabstractConventional unsupervised multi-source domain adaptation (UMDA) methods assume all source domains can be accessed directly. However, this assumption neglects the privacy-preserving policy, where all the data and computations must be kept decentralized. There exist three challenges in this scenario: (1) Minimizing the domain distance requires the pairwise calculation of the data from the source and target domains, while the data on the source domain is not available. (2) The communication cost and privacy security limit the application of existing UMDA methods, such as the domain adversarial training. (3) Since users cannot govern the data quality, the irrelevant or malicious source domains are more likely to appear, which causes negative transfer. To address the above problems, we propose a privacy-preserving UMDA paradigm named Knowledge Distillation based Decentralized Domain Adaptation (KD3A), which performs domain adaptation through the knowledge distillation on models from different source domains. The extensive experiments show that KD3A significantly outperforms state-of-the-art UMDA approaches. Moreover, the KD3A is robust to the negative transfer and brings a 100x reduction of communication cost compared with other decentralized UMDA methods. Haozhe Feng, Zhaoyang You, Tian-Ye Zhang, Minfeng Zhu 0001, Fei Wu 0001, Chao Wu 0001, Wei Chen 0001 |
ICML | 7 |
| 2020 | Multi-platform data collection for public service with Pay-by-Data
Chao Wu 0001, Simon Hu 0001, Chun-Hsiang Lee, Jun Xiao 0001 |
Multim. Tools Appl. | 1 |
| 2019 | Generative Creativity: Adversarial Learning for Bionic Design
Simiao Yu, Hao Dong 0003, Pan Wang 0005, Chao Wu 0001, Yike Guo |
ICANN (3) | 4 |
| 2019 | SIMGAN: Photo-Realistic Semantic Image Manipulation Using Generative Adversarial NetworksabstractSemantic image manipulation (SIM) aims to generate realistic images from an input source image and a target text description, such that the generated images not only match the content of the description, but also maintain text-irrelevant features of the source image. It requires to learn a good mapping between visual features and linguistic features. Previous works on SIM can only generate images of limited resolution that typically lack of fine and clear details. In this work, we aim to generate high-resolution photo-realistic images for SIM. Specifically, we propose SIMGAN, a generative adversarial networks (GAN) based architecture that is capable of generating images of size 256 × 256 for SIM. We demonstrate the effectiveness of SIMGAN and its superiority over existing methods via qualitative and quantitative evaluation on Caltech-200 and Oxford-102 datasets. Simiao Yu, Hao Dong 0003, Felix Liang, Yuanhan Mo, Chao Wu 0001, Yike Guo |
ICIP | 5 |
| 2019 | An artificial intelligence based data-driven approach for design ideation
Liuqing Chen 0002, Pan Wang 0005, Hao Dong 0003, Feng Shi 0007, Yike Guo, Peter R. N. Childs, Jun Xiao 0001, Chao Wu 0001 |
J. Vis. Commun. Image Represent. | 9 |
| 2018 | Deep Sequence Learning with Auxiliary Information for Traffic PredictionabstractPredicting traffic conditions from online route queries is a challenging task as there are many complicated interactions over the roads and crowds involved. In this paper, we intend to improve traffic prediction by appropriate integration of three kinds of implicit but essential factors encoded in auxiliary information. We do this within an encoder-decoder sequence learning framework that integrates the following data: 1) offline geographical and social attributes. For example, the geographical structure of roads or public social events such as national celebrations; 2) road intersection information. In general, traffic congestion occurs at major junctions; 3) online crowd queries. For example, when many online queries issued for the same destination due to a public performance, the traffic around the destination will potentially become heavier at this location after a while. Qualitative and quantitative experiments on a real-world dataset from Baidu have demonstrated the effectiveness of our framework. Binbing Liao, Jingqing Zhang, Chao Wu 0001, Douglas McIlwraith, Tong Chen 0006, Shengwen Yang, Yike Guo, Fei Wu 0001 |
KDD | 3 |
| 2018 | Dest-ResNet: A Deep Spatiotemporal Residual Network for Hotspot Traffic Speed PredictionabstractWith the ever-increasing urbanization process, the traffic jam has become a common problem in the metropolises around the world, making the traffic speed prediction a crucial and fundamental task. This task is difficult due to the dynamic and intrinsic complexity of the traffic environment in urban cities, yet the emergence of crowd map query data sheds new light on it. In general, a burst of crowd map queries for the same destination in a short duration (called "hotspot'') could lead to traffic congestion. For example, queries of the Capital Gym burst on weekend evenings lead to traffic jams around the gym. However, unleashing the power of crowd map queries is challenging due to the innate spatiotemporal characteristics of the crowd queries. To bridge the gap, this paper firstly discovers hotspots underlying crowd map queries. These discovered hotspots address the spatiotemporal variations. Then Dest-ResNet (Deep spatiotemporal Residual Network) is proposed for hotspot traffic speed prediction. Dest-ResNet is a sequence learning framework that jointly deals with two sequences in different modalities, i.e., the traffic speed sequence and the query sequence. The main idea of Dest-ResNet is to learn to explain and amend the errors caused when the unimodal information is applied individually. In this way, Dest-ResNet addresses the temporal causal correlation between queries and the traffic speed. As a result, Dest-ResNet shows a 30% relative boost over the state-of-the-art methods on real-world datasets from Baidu Map. Binbing Liao, Jingqing Zhang, Siliang Tang, Chao Wu 0001, Shengwen Yang, Wenwu Zhu 0001, Yike Guo, Fei Wu 0001 |
ACM Multimedia | 6 |
| 2018 | Dropping Activation Outputs With Localized First-Layer Deep Network for Enhancing User Privacy and Data SecurityabstractDeep learning methods can play a crucial role in anomaly detection, prediction, and supporting decision making for applications like personal health-care, pervasive body sensing, and so on. However, current architecture of deep networks suffers the privacy issue that users need to give out their data to the model (typically hosted in a server or a cluster on Cloud) for training or prediction. This problem is getting more severe for those sensitive health-care or medical data (e.g., fMRI or body sensors measures like EEG signals). In addition to this, there is also a security risk of leaking these data during the data transmission from user to the model (especially when it is through the Internet). Targeting at these issues, in this paper, we proposed a new architecture for deep network in which users do not reveal their original data to the model. In our method, feed-forward propagation and data encryption are combined into one process: we migrate the first layer of deep network to users' local devices and apply the activation functions locally, and then use the “dropping activation output” method to make the output non-invertible. The resulting approach is able to make model prediction without accessing users' sensitive raw data. The experiment conducted in this paper showed that our approach achieves the desirable privacy protection requirement and demonstrated several advantages over the traditional approach with encryption/decryption. Hao Dong 0003, Chao Wu 0001, Yike Guo |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2017 | Semantic Image Synthesis via Adversarial LearningabstractIn this paper, we propose a way of synthesizing realistic images directly with natural language description, which has many useful applications, e.g. intelligent image manipulation. We attempt to accomplish such synthesis: given a source image and a target text description, our model synthesizes images to meet two requirements: 1) being realistic while matching the target text description; 2) maintaining other image features that are irrelevant to the text description. The model should be able to disentangle the semantic information from the two modalities (image and text), and generate new images from the combined semantics. To achieve this, we proposed an end-to-end neural architecture that leverages adversarial learning to automatically learn implicit loss functions, which are optimized to fulfill the aforementioned two requirements. We have evaluated our model by conducting experiments on Caltech-200 bird dataset and Oxford-102 flower dataset, and have demonstrated that our model is capable of synthesizing realistic images that match the given descriptions, while still maintain other features of original images. Hao Dong 0003, Simiao Yu, Chao Wu 0001, Yike Guo |
ICCV | 3 |
| 2013 | Building a generic platform for big sensor data applicationabstractThe drive toward smart cities alongside the rising adoption of personal sensors is leading to a torrent of sensor data. While systems exist for storing and managing sensor data, the real value of such data is the insight which can be generated from it. However there is currently no platform which enables sensor data to be taken from collection, through use in models to produce useful data products. The architecture of such a platform is a current research question in the field of Big Data and Smart Cities. In this paper we explore five key challenges in this field and provide a response through a sensor data platform “Concinnity” which can take sensor data from collection to final product via a data repository and workflow system. This will enable rapid development of applications built on sensor data using data fusion and the integration and composition of models to form novel workflows. We summarize the key features of our approach, exploring how it enables value to be derived from sensor data efficiently. Chun-Hsiang Lee, David Birch, Chao Wu 0001, Dilshan Silva, Orestis Tsinalis, Yang Li 0003, Shulin Yan, Moustafa Ghanem, Yike Guo |
IEEE BigData | 3 |
| 2013 | Enhanced user data privacy with pay-by-data modelabstractPersonal data collection is becoming pervasive these days, these data has the risk of being abused by current application and application marketplace model, because only the price of application is explicitly indicated without clear agreement on usage of data, and the granularity of data access authentication is not enough to protect users privacy. In this short paper, we propose a new model of user data privacy. Data usage of the application is explicitly shown, and controlled by an authentication service, to protect users from the abuse of their data, especially in mobile application. Chao Wu 0001, Yike Guo |
IEEE BigData | 1 |
| 2013 | Sensor Deployment in Bayesian Compressive Sensing Based Environmental Monitoring
Chao Wu 0001, Di Wu 0002, Shulin Yan, Yike Guo |
MobiQuitous | 1 |
| 2012 | Elastic Application Container: A Lightweight Approach for Cloud Resource ProvisioningabstractVirtual machine (VM) based virtual infrastructure has been adopted widely in cloud computing environment for elastic resource provisioning. Performing resource management using VMs, however, is a heavyweight task. In practice, we have identified two scenarios where VM based resource management is less feasible and less resource-efficient. In this paper, we propose a lightweight resource management model that is called Elastic Application Container (EAC). EAC is a virtual resource unit for delivering better resource efficiency and more scalable cloud applications. We describe the EAC system architecture and components, and also present an algorithm for EAC resource provisioning. We also describe an implementation of the EAC-oriented platform to support multi-tenant cloud use. To evaluate our approach and implementation, we conducted experiments and collected performance data by comparing VM-based and EAC-based resource management with regards to their feasibility and resource-efficiency. The experiment results show that our proposed EAC-based resource management approach outperforms the VM-based approach in terms of feasibility and resource-efficiency. Sijin He, Li Guo 0002, Yike Guo, Chao Wu 0001, Moustafa Ghanem, Rui Han 0001 |
AINA | 4 |
| 2012 | A New Paradigm for Web App Development, Deployment, Distribution, and Collaboration
Chao Wu 0001, Yike Guo |
ICSOFT | 1 |