Yang Hua 0001

dblp:34/9988 · DBLP profile ↗
← Back
87ranked-venue papers
2as first author
60since 2021 · last 2026
0000-0001-5536-503XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 61 · 2 first-author · 38 since 2021Artificial intelligence and machine learning · 52 · 2 first-author · 34 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Poisoning with a Pill: Circumventing Detection in Federated Learning
abstract
Federated learning (FL) protects data privacy by enabling distributed model training without direct access to client data. However, its distributed nature makes it vulnerable to model and data poisoning attacks. While numerous defenses filter malicious clients using statistical metrics, they overlook the role of model redundancy, where not all parameters contribute equally to the model and attack performance. Current attacks manipulate all model parameters uniformly, making them more detectable, while defenses focus on the overall statistics of client updates, leaving gaps for more sophisticated attacks. We propose an attack-agnostic augmentation method to enhance the stealthiness and effectiveness of existing poisoning attacks in FL, exposing flaws in current defenses and highlighting the need for fine-grained FL security. Our three-stage methodology, including pill construction, pill poisoning, and pill injection, injects poison into a compact subnet (i.e., pill) of the global model during the iterative FL training. Experimental results show that FL poisoning attacks enhanced by our method can bypass 8 state-of-the-art (SOTA) defenses, gaining an up to 7x error rate increase, as well as on average a more than 2x error rate increase on both IID and non-IID data, in both cross-silo and cross-device FL systems.
Hanxi Guo, Hao Wang 0022, Tao Song 0003, Tianhang Zheng, Yang Hua 0001, Haibing Guan, Xiangyu Zhang 0001
AAAI5
2026 Deep Learning Backdoor Defense via Adaptive Trigger Collisions in Latent Space
abstract
Backdoor attacks in data outsourcing settings pose severe risks to deep neural networks. Specifically, adversaries can manipulate externally sourced training data to implant hidden behaviors in target models (e.g., incorrect predictions on triggered samples). Existing defenses are either pre-processing or post-processing. Since the two approaches are orthogonal and either one can independently strengthen real-world defenses, we focus on the latter in this paper. Yet current post-processing defenses face one or more of the following issues: overemphasis on output logits while overlooking rich information in intermediate layers, injection of uncertain new triggers while requiring alignment with the original triggers, and underuse of poisoned model representations. To overcome the aforementioned limitations, we propose ATClean, an adaptive post-processing defense based on feature collisions in latent space. Specifically, it leverages all layers rather than only output logits to capture backdoor-affected regions using an adaptive loss function, relaxes the need for exact trigger reconstruction by generating adversarial samples that only enforce feature collisions with a theoretical guarantee, and fully exploits poisoned representations with feature-collision-based fine-tuning. Experiments across benchmark datasets, multiple architectures, and seven representative attacks show that ATClean achieves state-of-the-art defense effectiveness with the lowest drop on clean data, including about a 20% improvement in DER, which measures the accuracy-defense trade-off.
Zixun Xiong, Hao Wang 0022, Jian Li 0008, Yang Hua 0001, Miao Pan, Xiaojiang Du
AsiaCCS4
2026 FedCod: An Efficient Coded Communication Protocol for Cross-Silo Federated Learning
Peishen Yan, Jun Li 0004, Hao Wang 0022, Yang Hua 0001, Tao Song 0003, Haibing Guan
IWQoS4
2026 Exploring Diffusion Models' Corruption Stage in Few-Shot Fine-tuning and Mitigating with Bayesian Neural Networks
abstract
Few-shot fine-tuning of Diffusion Models (DMs) is a key advancement, significantly reducing training costs and enabling personalized AI applications. However, we explore the training dynamics of DMs and observe an unanticipated phenomenon: during the training process, image fidelity initially improves, then unexpectedly deteriorates with the emergence of noisy patterns, only to recover later with severe overfitting. We term the stage with generated noisy patterns as corruption stage. To understand this corruption stage, we begin by heuristically modeling the one-shot fine-tuning scenario, and then extend this modeling to more general cases. Through this modeling, we identify the primary cause of this corruption stage: a narrowed learning distribution inherent in the nature of few-shot fine-tuning. To tackle this, we apply Bayesian Neural Networks (BNNs) on DMs with variational inference to implicitly broaden the learned distribution, and present that the learning target of the BNNs can be naturally regarded as an expectation of the diffusion loss and a further regularization with the pretrained DMs. This approach is highly compatible with current few-shot fine-tuning methods in DMs and does not introduce any extra inference costs. Experimental results demonstrate that our method significantly mitigates corruption, and improves the fidelity, quality and diversity of the generated images in both object-driven and subject-driven generation tasks.
Jiaru Zhang, Yang Hua 0001, Bohan Lyu 0001, Hao Wang 0022, Tao Song 0003, Haibing Guan
KDD (1)3
2026 CAMD: Context-Aware Masked Distillation for General Self-Supervised Facial Representation Pre-Training
abstract
Self-supervised pre-training has been shown to effectively learn transferable representations from unlabeled images in many visual tasks. However, existing self-supervised pre-training methods lack sufficient context-awareness and are difficult to obtain fine-grained facial representations, thus resulting in the weak generalization ability of the model to deal with various facial analysis tasks. To address this issue, we propose a Context-Aware Masked Distillation method, termed CAMD, to effectively learn general facial representations for fine-grained facial analysis tasks. The CAMD method first designs an innovative local-to-global masked image modeling framework to learn the contextual spatial structures and semantic relationships between local and global features, enabling effective self-supervised pre-training. In this framework, our pre-training task predicts the dense global feature representations based on the visible local feature representations after masking, so as to achieve semantic alignment across local and global views and significantly enhance spatial sensitivity. Moreover, the CAMD method leverages an attention-driven cross-view hierarchical distillation module to fully distill the features of related regions between different encoder layers of the online and target encoders. This module can learn contextual dependencies and capture discriminative fine-grained facial feature representations. Our method is evaluated on multiple downstream facial analysis tasks, including face alignment, face parsing, facial attribute recognition, facial expression recognition, and head pose estimation, all achieving state-of-the-art results and exhibiting the strong generality and effectiveness. The code is available at: https://github.com/mumumu-wss/CAMD.
Sensen Wang 0001, Si Chen 0002, Dahan Wang, Yang Hua 0001, Yan Yan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Learning Identifiable Structures Helps Avoid Bias in DNN-based Supervised Causal Learning
abstract
Causal discovery is a structured prediction task that aims to predict causal relations among variables based on their data samples. Supervised Causal Learning (SCL) is an emerging paradigm in this field. Existing Deep Neural Network (DNN)-based methods commonly adopt the “Node-Edge approach”, in which the model first computes an embedding vector for each variable-node, then uses these variable-wise representations to concurrently and independently predict for each directed causal-edge. In this paper, we first show that this architecture has some systematic bias that cannot be mitigated regardless of model size and data size. We then propose SiCL, a DNN-based SCL method that predicts a skeleton matrix together with a v-tensor (a third-order tensor representing the v-structures). According to the Markov Equivalence Class (MEC) theory, both the skeleton and the v-structures are \emph{identifiable} causal structures under the canonical MEC setting, so predictions about skeleton and v-structures do not suffer from the identifiability limit in causal discovery, thus SiCL can avoid the systematic bias in Node-Edge architecture, and enable consistent estimators for causal discovery. Moreover, SiCL is also equipped with a specially designed pairwise encoder module with a unidirectional attention layer to model both internal and external relationships of pairs of nodes. Experimental results on both synthetic and real-world benchmarks show that SiCL significantly outperforms other DNN-based SCL approaches.
Jiaru Zhang, Rui Ding 0001, Qiang Fu 0015, Bojun Huang, Zizhen Deng, Yang Hua 0001, Haibing Guan, Shi Han, Dongmei Zhang 0001
AISTATS6
2025 Improving the Training of Data-Efficient GANs via Quality Aware Dynamic Discriminator Rejection Sampling
abstract
Data-Efficient Generative Adversarial Nets (DE-GANs) have become more and more popular in recent years. Existing methods apply data augmentation, noise injection and pre-trained models to maximumly increase the number of training samples thus improving the training of DE-GANs. However, none of these methods considers the sample quality during training, which can also significantly influence the training of DE-GANs. Focusing on sample quality during training, in this paper, we are the first to incorporate discriminator rejection sampling (DRS) into the training process and introduce a novel method, called quality aware dynamic discriminator rejection sampling (QADDRS). Specifically, QADDRS consists of two steps: (1) the sample quality aware step, which aims to obtain the sorted critic scores, i.e., the ordered discriminator outputs, on real/fake samples in the current training stage; (2) the dynamic rejection step that obtains dynamic rejection number N, where N is controlled by the overfitting degree of discriminator (D) during training. When updating the parameters of D, the N high critic score real samples and the N low critic score fake samples in the minibatch are rejected dynamically based on the overfitting degree of D. As a result, QAD-DRS can avoid D becoming overly confident in distinguishing both real and fake samples, thereby alleviating the over-fitting of D issue during training. Extensive experiments on several datasets demonstrate that integrating QADDRS into different DE-GANs can achieve better performance and deliver state-of-the-art results. Codes are available at https://github.com/zzhang05/QADDRS.
Zhaoyu Zhang 0001, Yang Hua 0001, Guanxiong Sun, Hui Wang 0001, Seán F. McLoone
CVPR2
2025 Stealthy Backdoor Attack in Federated Learning via Adaptive Layer-Wise Gradient Alignment
Qingqian Yang, Peishen Yan, Jiaru Zhang, Tao Song 0003, Yang Hua 0001, Hao Wang 0022, Haibing Guan
ICCV6
2025 PCEvolve: Private Contrastive Evolution for Synthetic Dataset Generation via Few-Shot Private Data and Generative APIs
abstract
The rise of generative APIs has fueled interest in privacy-preserving synthetic data generation. While the Private Evolution (PE) algorithm generates Differential Privacy (DP) synthetic images using diffusion model APIs, it struggles with few-shot private data due to the limitations of its DP-protected similarity voting approach. In practice, the few-shot private data challenge is particularly prevalent in specialized domains like healthcare and industry. To address this challenge, we propose a novel API-assisted algorithm, Private Contrastive Evolution (PCEvolve), which iteratively mines inherent inter-class contrastive relationships in few-shot private data beyond individual data points and seamlessly integrates them into an adapted Exponential Mechanism (EM) to optimize DP’s utility in an evolution loop. We conduct extensive experiments on four specialized datasets, demonstrating that PCEvolve outperforms PE and other API-assisted baselines. These results highlight the potential of leveraging API access with private data for quality evaluation, enabling the generation of high-quality DP synthetic images and paving the way for more accessible and effective privacy-preserving generative API applications. Our code is available at https://github.com/TsingZ0/PCEvolve.
Jianqing Zhang, Yang Liu 0165, Yang Hua 0001, Tianyuan Zou, Jian Cao 0001, Qiang Yang 0001
ICML4
2025 Training Diffusion-based Generative Models with Limited Data
abstract
Diffusion-based generative models (diffusion models) often require a large amount of data to train a score-based model that learns the score function of the data distribution through denoising score matching. However, collecting and cleaning such data can be expensive, time-consuming, and even infeasible. In this paper, we present a novel theoretical insight for diffusion models that two factors, i.e., the denoiser function hypothesis space and the number of training samples, can affect the denoising score matching error of all training samples. Based on this theoretical insight, it is evident that minimizing the total denoising score matching error is challenging within the denoiser function hypothesis space in existing methods, when training diffusion models with limited data. To address this, we propose a new diffusion model called Limited Data Diffusion (LD-Diffusion), which consists of two main components: a compressing model and a novel mixed augmentation with fixed probability (MAFP) strategy. Specifically, the compressing model can constrain the complexity of the denoiser function hypothesis space and MAFP can effectively increase the training samples by providing more informative guidance than existing data augmentation methods in the compressed hypothesis space. Extensive experiments on several datasets demonstrate that LD-Diffusion can achieve better performance compared to other diffusion models. Codes are available at https://github.com/zzhang05/LD-Diffusion.
Zhaoyu Zhang 0001, Yang Hua 0001, Guanxiong Sun, Hui Wang 0001, Seán F. McLoone
ICML2
2025 HtFLlib: A Comprehensive Heterogeneous Federated Learning Library and Benchmark
abstract
As AI evolves, collaboration among heterogeneous models helps overcome data scarcity by enabling knowledge transfer across institutions and devices.Traditional Federated Learning (FL) only supports homogeneous models, limiting collaboration among clients with heterogeneous model architectures.To address this, Heterogeneous Federated Learning (HtFL) methods are developed to enable collaboration across diverse heterogeneous models while tackling the data heterogeneity issue at the same time.However, a comprehensive benchmark for standardized evaluation and analysis of the rapidly growing HtFL methods is lacking.Firstly, the highly varied datasets, model heterogeneity scenarios, and different method implementations become hurdles to making easy and fair comparisons among HtFL methods.Secondly, the effectiveness and robustness of HtFL methods are under-explored in various scenarios, such as the medical domain and sensor signal modality.To fill this gap, we introduce the first Heterogeneous Federated Learning Library (HtFLlib), an easy-to-use and extensible framework that integrates multiple datasets and model heterogeneity scenarios, offering a robust benchmark for research and practical applications.Specifically, HtFLlib integrates (1) 12 datasets spanning various domains, modalities, and data heterogeneity scenarios; (2) 40 model architectures, ranging from small to large, across three modalities;(3) a modularized and easy-to-extend HtFL codebase with implementations of 10 representative HtFL methods; and (4) systematic evaluations in terms of accuracy, convergence, computation costs, and communication costs.We emphasize the advantages and potential of state-of-the-art HtFL methods and hope that HtFLlib will catalyze advancing HtFL research and enable its broader applications.The code is released at https://github.com/TsingZ0/HtFLlib.
Jianqing Zhang, Xinghao Wu, Yanbing Zhou, Xiaoting Sun, Qiqi Cai, Yang Liu 0165, Yang Hua 0001, Zhenzhe Zheng 0001, Jian Cao 0001, Qiang Yang 0001
KDD (2)7
2025 PFLlib: A Beginner-Friendly and Comprehensive Personalized Federated Learning Library and Benchmark
abstract
Amid the ongoing advancements in Federated Learning (FL), a machine learning paradigm that allows collaborative learning with data privacy protection, personalized FL (pFL) has gained significant prominence as a research direction within the FL domain. Whereas traditional FL (tFL) focuses on jointly learning a global model, pFL aims to balance each client's global and personalized goals in FL settings. To foster the pFL research community, we started and built PFLlib, a comprehensive pFL library with an integrated benchmark platform. In PFLlib, we implemented 37 state-of-the-art FL algorithms (8 tFL algorithms and 29 pFL algorithms) and provided various evaluation environments with three statistically heterogeneous scenarios and 24 datasets. At present, PFLlib has gained more than 1600 stars and 300 forks on GitHub.
Jianqing Zhang, Yang Liu 0165, Yang Hua 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Jian Cao 0001
J. Mach. Learn. Res.3
2025 DHLA: Dynamic Hybrid Label Assignment for End-to-End Object Detection
abstract
The recent one-to-one label assignment plays a crucial role in removing the last non-differentiable component, i.e., Non-Maximum Suppression (NMS), used in the post-processing step of the one-to-many label assignment, thus building an efficient end-to-end detection system. However, due to the limited number of foreground samples, the one-to-one label assignment often suffers from insufficient representation learning, and its performance is inferior to that of traditional detectors trained using the one-to-many label assignment. To solve these problems, we introduce a novel Dynamic Hybrid Label Assignment (DHLA) method, including a Hybrid Sample Selection (HSS) strategy and a Stage-aware Soft-label Adjustment (SSA) mechanism. In order to enhance the ability of representation learning of the one-to-one label assignment, the HSS strategy subtly integrates the one-to-many and the one-to-one label assignment rules to form a simple and effective hybrid assignment rule, where high-quality samples are selected for training according to an effective task consistency metric. Moreover, the SSA mechanism dynamically adjusts the contributions of different foreground samples at different training stages, thus effectively achieving the transition from one-to-many to one-to-one label assignment. In addition, we leverage a ranking loss function to widen the score gaps between the highest scoring position and surrounding areas for effectively removing duplicate bounding boxes. As a result, our method not only learns robust feature representations during training but also performs efficient end-to-end detection during inference. Extensive experiments demonstrate our method achieves competitive performance compared to state-of-the-art detectors on the challenging COCO and CrowdHuman datasets.
Zhi-Liang Hu, Si Chen 0002, Yang Hua 0001, Dahan Wang, Shunzhi Zhu, Yan Yan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Hierarchical Attention-Enhanced Correlation Refinement for Robust Visual Tracking
abstract
In recent years, visual tracking has witnessed remarkable advancements with the exploration of feature extraction and correlation modeling techniques. However, inadequate robustness of either the backbone network or the correlation operation continues to plague existing trackers, leading to frustrating drift when confronted with similar distractors or cluttered backgrounds. To address this problem, we propose a hierarchical attention-enhanced correlation refinement network (HarNet) for achieving robust visual tracking. Specifically, a gated dual-view attention (GDA) module is first designed to aggregate the intra-layer attention and the inter-layer self-attention based on a fusion gate, so as to enhance hierarchical feature representations of the template. Meanwhile, a target-aware attention (TA) module introduces the template information to the inter-layer self-attention, which can highlight the target information in the search region. Moreover, a graph guided correlation (GGC) module leverages the pixel-to-local and pixel-to-global correlations to fully exploit both local-and global-spatial information between the template and the search region, and then uses the graph convolutional network (GCN) to further learn the node relationships of the correlation map for more finegrained correlations. Thus, with the above three elaborately designed modules, the HarNet is beneficial for the enhancement of feature representation and the precise localization of the target. Extensive experiments on popular visual tracking datasets (including OTB100, VOT2016, VOT2018, VOT2019, UAV123, UAV20L, GOT-10k, and LaSOT) demonstrate the superiority of our proposed method against several state-of-the-art tracking methods.
Si Chen 0002, Rui Xu 0028, Yan Yan 0001, Yang Hua 0001, Dahan Wang, Shunzhi Zhu
IEEE Trans. Intell. Transp. Syst.4
2025 Hierarchical Token-Aware Cross-Modality Reconstruction for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification (VI-ReID) aims to query the same pedestrian's visible (infrared) images in the gallery set from the infrared (visible) images. VI-ReID not only needs to deal with the challenging factors like pose variation and occlusion, but also requires handling the large modality discrepancy. Previous methods mainly focus on learning single-scale modality-shared features and do not effectively explore the multi-scale features of two modalities from both short-range and long-range perspectives. In order to solve these problems, this paper proposes a novel Hierarchical Token-Aware Cross-Modality Reconstruction (HTCR) network to significantly mitigate the modality discrepancy for effective VI-ReID. The HTCR network consists of two main components, i.e., Hierarchical Token-aware Fusion (HTF) and Cross-modality Feature Reconstruction (CFR). The HTF module first bidirectionally exchanges the short-range and long-range multi-scale modality-shared features with a few learnable tokens to achieve discriminative pedestrian features by making full use of the advantages of both Convolutional Neural Network (CNN) and Transformer. Moreover, the CFR module reconstructs global and local pedestrian features of one modality by using the token sequence of the other modality with multi-scale cues to further explore the relationship between the two distinct modalities and alleviate the modality discrepancy. In addition, the Modality-shared feature Reconstruction (MR) loss is leveraged to reduce the noises between the reconstructed and the target features. Experimental results indicate that the proposed HTCR can significantly improve the VI-ReID performance and outperform the state-of-the-art methods on the cross-modality SYSU-MM01, RegDB, and LLCM datasets.
Si Chen 0002, Liuxiang Qiu, Dahan Wang, Wentao Zhu 0002, Yang Hua 0001, Yan Yan 0001
IEEE Trans. Multim.5
2024 Cheaper and Faster: Distributed Deep Reinforcement Learning with Serverless Computing
abstract
Deep reinforcement learning (DRL) has gained immense success in many applications, including gaming AI, robotics, and system scheduling. Distributed algorithms and architectures have been vastly proposed (e.g., actor-learner architecture) to accelerate DRL training with large-scale server-based clusters. However, training on-policy algorithms with the actor-learner architecture unavoidably induces resource wasting due to synchronization between learners and actors, thus resulting in significantly extra billing. As a promising alternative, serverless computing naturally fits on-policy synchronization and alleviates resource wasting in distributed DRL training with pay-as-you-go pricing. Yet, none has leveraged serverless computing to facilitate DRL training. This paper proposes MinionsRL, the first serverless distributed DRL training framework that aims to accelerate DRL training- and cost-efficiency with dynamic actor scaling. We prototype MinionsRL on top of Microsoft Azure Container Instances and evaluate it with popular DRL tasks from OpenAI Gym. Extensive experiments show that MinionsRL reduces total training time by up to 52% and training cost by 86% compared to latest solutions.
Hanfei Yu, Jian Li 0008, Yang Hua 0001, Xu Yuan 0001, Hao Wang 0022
AAAI3
2024 FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning
abstract
Recently, Heterogeneous Federated Learning (HtFL) has attracted attention due to its ability to support heterogeneous models and data. To reduce the high communication cost of transmitting model parameters, a major challenge in HtFL, prototype-based HtFL methods are proposed to solely share class representatives, a.k.a, prototypes, among heterogeneous clients while maintaining the privacy of clients’ models. However, these prototypes are naively aggregated into global prototypes on the server using weighted averaging, resulting in suboptimal global knowledge which negatively impacts the performance of clients. To overcome this challenge, we introduce a novel HtFL approach called FedTGP, which leverages our Adaptive-margin-enhanced Contrastive Learning (ACL) to learn Trainable Global Prototypes (TGP) on the server. By incorporating ACL, our approach enhances prototype separability while preserving semantic meaning. Extensive experiments with twelve heterogeneous models demonstrate that our FedTGP surpasses state-of-the-art methods by up to 9.08% in accuracy while maintaining the communication and privacy advantages of prototype-based HtFL. Our code is available at https://github.com/TsingZ0/FedTGP.
Jianqing Zhang, Yang Liu 0165, Yang Hua 0001, Jian Cao 0001
AAAI3
2024 CGI-DM: Digital Copyright Authentication for Diffusion Models via Contrasting Gradient Inversion
abstract
Diffusion Models (DMs) have evolved into advanced image generation tools, especially for few-shot generation where a pretrained model is fine-tuned on a small set of images to capture a specific style or object. Despite their success, concerns exist about potential copyright violations stemming from the use of unauthorized data in this process. In response, we present Contrasting Gradient Inversion for Diffusion Models (CGI-DM), a novel method featuring vivid visual representations for digital copyright authentication. Our approach involves removing partial information of an image and recovering missing details by exploiting conceptual differences between the pretrained and fine-tuned models. We formulate the differences as KL divergence between latent variables of the two models when given the same input image, which can be maximized through Monte Carlo sampling and Projected Gradient Descent (PGD). The similarity between original and recovered images serves as a strong indicator of potential infringements. Extensive experiments on the WikiArt and Dream-booth datasets demonstrate the high accuracy of CGI-DM in digital copyright authentication, surpassing alternative validation techniques. Code implementation is available at https://github.com/Nicholas0228/Revelio.
Yang Hua 0001, Chumeng Liang, Jiaru Zhang, Hao Wang 0022, Tao Song 0003, Haibing Guan
CVPR2
2024 An Upload-Efficient Scheme for Transferring Knowledge From a Server-Side Pre-trained Generator to Clients in Heterogeneous Federated Learning
abstract
Heterogeneous Federated Learning (HtFL) enables collaborative learning on multiple clients with different model architectures while preserving privacy. Despite recent research progress, knowledge sharing in HtFL is still difficult due to data and model heterogeneity. To tackle this issue, we leverage the knowledge stored in public pretrained generators and propose a new upload-efficient knowledge transfer scheme called Federated Knowledge-Transfer Loop (Fed-KTL). Our FedKTL can produce client-task-related prototypical image-vector pairs via the generator's inference on the server. With these pairs, each client can transfer preexisting knowledge from the generator to its local model through an additional supervised local task. We conduct extensive experiments on four datasets under two types of data heterogeneity with 14 kinds of models including CNNs and ViTs. Results show that our upload-efficient FedKTL surpasses seven state-of-the-art methods by up to 7.31% in accuracy. Moreover, our knowledge transfer scheme is applicable in scenarios with only one edge client. Code: https://github.com/TsingZ0/FedKTL
Jianqing Zhang, Yang Liu 0165, Yang Hua 0001, Jian Cao 0001
CVPR3
2024 SKYMASK: Attack-Agnostic Robust Federated Learning with Fine-Grained Learnable Masks
Peishen Yan, Hao Wang 0022, Tao Song 0003, Yang Hua 0001, Ruhui Ma, Ningxin Hu, Mohammad R. Haghighat, Haibing Guan
ECCV (19)4
2024 Backdoor Federated Learning by Poisoning Backdoor-Critical Layers
abstract
Federated learning (FL) has been widely deployed to enable machine learning training on sensitive data across distributed devices. However, the decentralized learning paradigm and heterogeneity of FL further extend the attack surface for backdoor attacks. Existing FL attack and defense methodologies typically focus on the whole model. None of them recognizes the existence of backdoor-critical (BC) layers-a small subset of layers that dominate the model vulnerabilities. Attacking the BC layers achieves equivalent effects as attacking the whole model but at a far smaller chance of being detected by state-of-the-art (SOTA) defenses. This paper proposes a general in-situ approach that identifies and verifies BC layers from the perspective of attackers. Based on the identified BC layers, we carefully craft a new backdoor attack methodology that adaptively seeks a fundamental balance between attacking effects and stealthiness under various defense strategies. Extensive experiments show that our BC layer-aware backdoor attacks can successfully backdoor FL under seven SOTA defenses with only 10% malicious clients and outperform the latest backdoor attack methods.
Haomin Zhuang, Mingxian Yu, Hao Wang 0022, Yang Hua 0001, Jian Li 0008, Xu Yuan 0001
ICLR4
2024 MVRMLM 2024: Multimodal Video Retrieval and Multimodal Language Modelling
abstract
As the proliferation of video content continues, and many video archives lack suitable metadata, therefore, video retrieval, particularly through example-based search, has become increasingly crucial. Existing metadata often fails to meet the needs of specific types of searches, especially when videos contain elements from different modalities, such as visual and audio. Consequently, developing video retrieval methods that can handle multi-modal content is essential. In designing our novel video retrieval framework named Multi-modal Video Search by Examples (MVSE)1, we focused on accuracy (precision and recall), efficiency (retrieval time in seconds), interactivity, and extensibility, with key components including advanced data processing and a user-friendly interface aimed at enhancing search effectiveness and user experience. With the advent of Large Language Models (LLMs), the interaction between multimodal data, including image and audio has been transformed with a significant leap forward towards a bigger goal of artificial general intelligence. This workshop aims to bring together experts from diverse domains to explore the possibilities of developing novel ways of multimodal data search, understanding and interaction.
Hui Wang 0001, Josef Kittler, Mark J. F. Gales, Rob Cooper, Maurice D. Mulvenna, Wing W. Y. Ng, Yang Hua 0001, Richard Gault, Abbas Haider, Guanfeng Wu
ICMR7
2024 Improving the Training of the GANs with Limited Data via Dual Adaptive Noise Injection
abstract
Recently, many studies have highlighted that training Generative Adversarial Networks (GANs) with limited data suffers from the overfitting of the discriminator (D). Existing studies mitigate the overfitting of D by employing data augmentation, model regularization, or pre-trained models. Despite the success of existing methods in training GANs with limited data, noise injection is another plausible, complementary, yet not well-explored approach to alleviate the overfitting of D issue. In this paper, we propose a simple yet effective method called Dual Adaptive Noise Injection (DANI), to further improve the training of GANs with limited data. Specifically, DANI consists of two adaptive strategies: adaptive injection probability and adaptive noise strength. For the adaptive injection probability, Gaussian noise is injected into both real and fake images for generator (G) and D with a probability p, respectively, where the probability p is controlled by the overfitting degree of D. For the adaptive noise strength, the Gaussian noise is produced by applying the adaptive forward diffusion process to both real and fake images, respectively. As a result, DANI can effectively increase the overlap between the distributions of real and fake data during training, thus alleviating the overfitting of D issue. Extensive experiments on several commonly-used datasets with both StyleGAN2 and FastGAN backbones demonstrate that DANI can further improve the training of GANs with limited data and achieve state-of-the-art results compared with other methods. Codes are available at https://github.com/zzhang05/DANI.
Zhaoyu Zhang 0001, Yang Hua 0001, Guanxiong Sun, Hui Wang 0001, Seán F. McLoone
ACM Multimedia2
2024 Improving the Leaking of Augmentations in Data-Efficient GANs via Adaptive Negative Data Augmentation
abstract
Data augmentation (DA) has shown its effectiveness in training Data-Efficient GANs (DE-GANs). However, applying DA in DE-GANs results in transforming the distributions of generated data and real data to augmented distributions of generated data and real data. This augmentation process could produce some out-of-distribution samples, known as the leaking of augmentations problem, which is highly undesirable in DE-GANs training. Although some methods propose "leaking-free" DAs for DE-GANs, we theoretically and practically argue that the leaking of augmentations problem still exists in these methods. To alleviate the leaking of augmentations in DE-GANs, in this paper, we propose a simple yet effective method called adaptive negative data augmentation (ANDA) for DE-GANs, with a negligible computational cost increase. Specifically, ANDA adaptively augments the augmented distribution of generated data using the augmented distribution of negative real data, where the negative real data is produced by applying negative data augmentation (NDA) on the real data. In this case, potential leaking samples can be presented as "fake" instances to the discriminator adaptively, which avoids the generator (G) learning such samples, thus resulting in better performance. Extensive experiments on several datasets with different DE-GANs demonstrate that ANDA can effectively alleviate the leaking of augmentations problem during training and achieve better performance. Codes are available at https://github.com/zzhang05/ANDA
Zhaoyu Zhang 0001, Yang Hua 0001, Guanxiong Sun, Hui Wang 0001, Seán F. McLoone
WACV2
2024 Improving the Fairness of the Min-Max Game in GANs Training
abstract
Generative adversarial networks (GANs) have achieved great success and become more and more popular in recent years. However, understanding of the min-max game in GANs training is still limited. In this paper, we first utilize information game theory to analyze the min-max game in GANs and introduce a new viewpoint on the GANs training that the min-max game in existing GANs is unfair during training, leading to sub-optimal convergence. To tackle this, we propose a novel GAN called Information Gap GAN (IGGAN), which consists of one generator (G) and two discriminators (D1and D2). Specifically, we apply different data augmentation methods to D1and D2, respectively. The information gap between different data augmentation methods can change the information received by each player in the min-max game and lead to all three players G, D1and D2in IGGAN obtaining incomplete information, which improves the fairness of the min-max game, yielding better convergence. We conduct extensive experiments for large-scale and limited data settings on several common datasets with two backbones, i.e., BigGAN and StyleGAN2. The results demonstrate that IGGAN can achieve a higher Inception Score (IS) and a lower Fréchet Inception Distance (FID) compared with other GANs. Codes are available at https://github.com/zzhang05/IGGAN
Zhaoyu Zhang 0001, Yang Hua 0001, Hui Wang 0001, Seán F. McLoone
WACV2
2024 OFL-W3: A One-shot Federated Learning System on Web 3.0
abstract
Federated Learning (FL) addresses the challenges posed by data silos, which arise from privacy, security regulations, and ownership concerns. Despite these barriers, FL enables these isolated data repositories to participate in collaborative learning without compromising privacy or security. Concurrently, the advancement of blockchain technology and decentralized applications (DApps) within Web 3.0 heralds a new era of transformative possibilities in web development. As such, incorporating FL into Web 3.0 paves the path for overcoming the limitations of data silos through collaborative learning. However, given the transaction speed constraints of core blockchains such as Ethereum (ETH) and the latency in smart contracts, employing one-shot FL, which minimizes client-server interactions in traditional FL to a single exchange, is considered more apt for Web 3.0 environments. This paper presents a practical one-shot FL system for Web 3.0, termed OFL-W3. OFL-W3 capitalizes on blockchain technology by utilizing smart contracts for managing transactions. Meanwhile, OFL-W3 utilizes the Inter-Planetary File System (IPFS) coupled with Flask communication, to facilitate backend server operations to use existing one-shot FL algorithms. With the integration of the incentive mechanism, OFL-W3 showcases an effective implementation of one-shot FL on Web 3.0, offering valuable insights and future directions for AI combined with Web 3.0 studies.
Linshan Jiang, Moming Duan, Bingsheng He, Peishen Yan, Yang Hua 0001, Tao Song 0003
Proc. VLDB Endow.6
2024 MCAS-GP: Deep Learning-Empowered Middle Cerebral Artery Segmentation and Gate Proposition
abstract
With the fast development of AI technologies, deep learning is widely applied for biomedical data analytics and digital healthcare. However, there remain gaps between AI-aided diagnosis and real-world healthcare demands. For example, hemodynamic parameters of the middle cerebral artery (MCA) have significant clinical value for diagnosing adverse perinatal results. Nevertheless, the current measurement procedure is tedious for sonographers. To reduce the workload of sonographers, we propose MCAS-GP, a deep learning-empowered framework that tackles the Middle Cerebral Artery Segmentation and Gate Proposition. MCAS-GP can automatically segment the region of the MCA and detect the corresponding position of the gate in the procedure of fetal MCA Doppler assessment. In MCAS-GP, a novel learnable atrous spatial pyramid pooling (LASPP) module is designed to adaptively learn multi-scale features. We also propose a novel evaluation metric, Affiliation Index, for measuring the effectiveness of the position of the output gate. To evaluate our proposed MCAS-GP, we build a large-scale MCA dataset, collaborating with the International Peace Maternity and Child Health Hospital of China welfare institute (IPMCH). Extensive experiments on the MCA dataset and two other public surgical datasets demonstrate that MCAS-GP can achieve considerable performance improvement in both accuracy and inference time.
Rui Zhang 0087, Shuo Wang 0008, Ruhui Ma, Yang Hua 0001, Tao Song 0003, Yunyun Cao, Haibing Guan
IEEE Trans. Comput. Biol. Bioinform.4
2024 Unpaired Caricature-Visual Face Recognition via Feature Decomposition-Restoration-Decomposition
abstract
Existing caricature-visual face recognition methods train the models based on caricature-visual image pairs from the same identities. Unfortunately, in many real-world applications, facial caricatures and visual facial images are usually unpaired in the training set due to the difficulty of collecting facial caricatures drawn by artists. In this paper, we study caricature-visual face recognition under the practical setting that only unpaired facial caricature and visual facial images are available as training samples, and define this setting as unpaired caricature-visual face recognition. To this end, we develop a novel feature decomposition-restoration-decomposition method (FDRD), which mainly consists of a backbone network, an identity-oriented feature decomposition module, and a modality-oriented feature restoration module, to extract modality-irrelevant identity features. To effectively train FDRD in the case of limited facial caricature training samples, we develop a two-stage learning framework. In the first stage, we perform single-modality restoration, enabling the model to have the basic ability of feature decomposition and restoration for each modality. In the second stage, we perform cross-modality recognition by exchanging new modality features between the two modalities, facilitating the model to focus on the decoupling of identity features and modality features. Experimental results demonstrate that our method performs favorably against several state-of-the-art face recognition methods and cross-modality methods. Our code is available at https://github.com/Capricorn-Karma/FDRD.
Yan Yan 0001, Jing-Hao Xue, Yang Hua 0001, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Siren$^+$+: Robust Federated Learning With Proactive Alarming and Differential Privacy
abstract
Federated learning (FL), an emerging machine learning paradigm that trains a global model across distributed clients without violating data privacy, has recently attracted significant attention. However, FL?s distributed nature and iterative training extensively increase the attacking surface for Byzantine and inference attacks. Existing FL defense methods can hardly protect FL from both Byzantine and inference attacks due to their fundamental conflicts. The noise injected to defend against inference attacks interferes with model weights and training data, obscuring model analysis that Byzantine-robust methods utilize to detect attacks. Besides, the practicability of existing Byzantine-robust methods is limited since they heavily rely on model analysis. In this paper, we present SIREN+, a new robust FL system that defends against a wide spectrum of Byzantine attacks and inference attacks by jointly utilizing a proactive alarming mechanism and local differential privacy (LDP). The proactive alarming mechanism orchestrates clients and the FL server to collaboratively detect attacks using distributed alarms, which is free from the noise interference injected by LDP. Compared with the state-of-the-art defense methods, SIREN+can protect FL from Byzantine and inference attacks from a higher proportion of malicious clients in the system while keeping the global model performing normally. Extensive experiments with diverse settings and attacks on real-world datasets show that SIREN+outperforms existing defense methods when attacked by Byzantine and inference attacks.
Hanxi Guo, Hao Wang 0022, Tao Song 0003, Yang Hua 0001, Ruhui Ma, Xiulang Jin, Zhengui Xue, Haibing Guan
IEEE Trans. Dependable Secur. Comput.4
2024 Robust Searching-Based Gradient Collaborative Management in Intelligent Transportation System
abstract
With the rapid development of big data and the Internet of Things (IoT), traffic data from an Intelligent Transportation System (ITS) is becoming more and more accessible. To understand and simulate the traffic patterns from the traffic data, Multimedia Cognitive Computing (MCC) is an efficient and practical approach. Distributed Machine Learning (DML) has been the trend to provide sufficient computing resources and efficiency for MCC tasks to handle massive data and complex models. DML can speed up computation with those computing resources but introduces communication overhead. Gradient collaborative management or gradient aggregation in DML for MCC tasks is a critical task. An efficient managing algorithm of the communication schedules for gradient aggregation in ITS can improve the performance of MCC tasks. However, existing communication schedules typically rely on specific physical connection matrices, which have low robustness when a malfunction occurs. In this article, we propose Robust Searching-based Gradient Collaborative Management (RSGCM) in Intelligent Transportation System, a practical ring-based gradient managing algorithm for communication schedules across devices to deal with ITS malfunction. RSGCM provides solutions of communication schedules to various kinds of connection matrices with an acceptable amount of training time. Our experimental results have shown that RSGCM can deal with more varieties of connection matrices than existing state-of-the-art communication schedules. RSGCM also increases the robustness of ITS since it can restore the system’s functionality in an acceptable time when device or connection breakdown happens.
Hongjian Shi, Hao Wang 0022, Ruhui Ma, Yang Hua 0001, Tao Song 0003, Honghao Gao, Haibing Guan
ACM Trans. Multim. Comput. Commun. Appl.4
2023 FedALA: Adaptive Local Aggregation for Personalized Federated Learning
abstract
A key challenge in federated learning (FL) is the statistical heterogeneity that impairs the generalization of the global model on each client. To address this, we propose a method Federated learning with Adaptive Local Aggregation (FedALA) by capturing the desired information in the global model for client models in personalized FL. The key component of FedALA is an Adaptive Local Aggregation (ALA) module, which can adaptively aggregate the downloaded global model and local model towards the local objective on each client to initialize the local model before training in each iteration. To evaluate the effectiveness of FedALA, we conduct extensive experiments with five benchmark datasets in computer vision and natural language processing domains. FedALA outperforms eleven state-of-the-art baselines by up to 3.27% in test accuracy. Furthermore, we also apply ALA module to other federated learning methods and achieve up to 24.19% improvement in test accuracy. Code is available at https://github.com/TsingZ0/FedALA.
Jianqing Zhang, Yang Hua 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
AAAI2
2023 Information Bound and Its Applications in Bayesian Neural Networks
abstract
Bayesian neural networks have drawn extensive interest because of their distinctive probabilistic representation framework. However, despite its recent success, little work focuses on the information-theoretic understanding of Bayesian neural networks. In this paper, we propose Information Bound as a metric of the amount of information in Bayesian neural networks. Different from mutual information on deterministic neural networks where modification of network structure or specific input data is usually necessary, Information Bound can be easily estimated on current Bayesian neural networks without any modification of network structures or training processes. By observing the trend of Information Bound during training, we demonstrate the existence of the “critical period” in Bayesian neural networks. Besides, we show that the Information Bound can be used to judge the confidence of the model prediction and to detect out-of-distribution datasets. Based on these observations of model interpretation, we propose Information Bound regularization and Information Bound variance regularization methods. The Information Bound regularization encourages models to learn the minimum necessary information and improves the model generality and robustness. The Information Bound variance regularization encourages models to learn more about complex samples with low Information Bound. Extensive experiments on KMNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100 verify the effectiveness of the proposed regularization methods.
Jiaru Zhang, Yang Hua 0001, Tao Song 0003, Hao Wang 0022, Zhengui Xue, Ruhui Ma, Haibing Guan
ECAI2
2023 Cross-Domain Learning with Normalizing Flow
abstract
Cross-domain learning aims to transfer knowledge learned from one or more datasets to other datasets in different domains, so that less data will be required for learning in new tasks and datasets. One big challenge in cross-domain learning is to effectively synergize the knowledge learning between domains. In this paper, we propose a new solution to address this challenge using normalizing flow, named as DomainFlow, which works as a learned mapping to establish knowledge sharing between source and target domains. The learned flow encourages the posterior distributions in multi-domain learning to be better aligned, leading to better performance in the target domain tasks. We conduct extensive experiments on three representative cross-domain learning tasks: unsupervised domain adaptation, domain generalization and zero-shot sketch-based image retrieval, which demonstrates that with DomainFlow, the overall performance on these diverse tasks can all be improved.
Jian Gao 0018, Yang Hua 0001, Hui Wang 0001
ICASSP3
2023 Flowreg: Latent Space Regularization Using Normalizing Flow For Limited Samples Learning
abstract
Modern deep neural network models have made remarkable success in many areas, supported by large sets of training samples. Yet the hunger for huge data has also become fatal in further expanding the use of deep models. Limited sample learning aims at learning generalized and transferable representations, without requiring large training data. In this paper, we propose FlowReg, a new learnable latent space regularization for limited sample problems. FlowReg modulates the latent space using a Normalizing Flow with a simple prior (such as Gaussian) while maintaining the complexity of the posterior distribution. We conduct thorough experiments on diverse tasks in limited label learning, as well as detailed in-depth analysis to comprehensively demonstrate the effectiveness of FlowReg.
Jian Gao 0018, Yang Hua 0001, Hui Wang 0001
ICASSP3
2023 Online Residual-Based Key Frame Sampling with Self-Coach Mechanism and Adaptive Multi-Level Feature Fusion
abstract
Key frame sampling is a common component in video tasks. Putting more effort into key frames, rather than processing all frames equally, can significantly reduce computational costs and improve processing efficiency. This paper presents ORSampler, an adaptive Online Residual-based key frame Sampler. ORSampler relies on feature residuals to sample key frames and decouples from subsequent video tasks. To facilitate ORSampler, a self-coached mechanism is designed to speed up learning, and an adaptive multi-level feature fusion is proposed to fit the diversity of subsequent video tasks. OR-Sampler has a fast inference speed and can work online. Extensive experiments on two typical video tasks verify the effectiveness and generality of our proposed ORSampler.
Rui Zhang 0087, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
ICASSP2
2023 Spatio-temporal Prompting Network for Robust Video Feature Extraction
abstract
Frame quality deterioration is one of the main challenges in the field of video understanding. To compensate for the information loss caused by deteriorated frames, recent approaches exploit transformer-based integration modules to obtain spatio-temporal information. However, these integration modules are heavy and complex. Furthermore, each integration module is specifically tailored for its target task, making it difficult to generalise to multiple tasks. In this paper, we present a neat and unified framework, called Spatio-Temporal Prompting Network (STPN). It can efficiently extract robust and accurate video features by dynamically adjusting the input features in the backbone network. Specifically, STPN predicts several video prompts containing spatio-temporal information of neighbour frames. Then, these video prompts are prepended to the patch embeddings of the current frame as the updated input for video feature extraction. Moreover, STPN is easy to generalise to various video tasks because it does not contain task-specific modules. Without bells and whistles, STPN achieves state-of-the-art performance on three widely-used datasets for different video understanding tasks, i.e., ImageNetVID for video object detection, YouTubeVIS for video instance segmentation, and GOT-10k for visual object tracking. Codes are available at https://github.com/guanxiongsun/STPN
Guanxiong Sun, Zhaoyu Zhang 0001, Jiankang Deng, Stefanos Zafeiriou, Yang Hua 0001
ICCV6
2023 GPFL: Simultaneously Learning Global and Personalized Feature Information for Personalized Federated Learning
abstract
Federated Learning (FL) is popular for its privacy-preserving and collaborative learning capabilities. Recently, personalized FL (pFL) has received attention for its ability to address statistical heterogeneity and achieve personalization in FL. However, from the perspective of feature extraction, most existing pFL methods only focus on extracting global or personalized feature information during local training, which fails to meet the collaborative learning and personalization goals of pFL. To address this, we propose a new pFL method, named GPFL, to simultaneously learn global and personalized feature information on each client. We conduct extensive experiments on six datasets in three statistically heterogeneous settings and show the superiority of GPFL over ten state-of-the-art methods regarding effectiveness, scalability, fairness, stability, and privacy. Besides, GPFL mitigates overfitting and outperforms the baselines by up to 8.99% in accuracy.
Jianqing Zhang, Yang Hua 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Jian Cao 0001, Haibing Guan
ICCV2
2023 Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial Examples
abstract
Recently, Diffusion Models (DMs) boost a wave in AI for Art yet raise new copyright concerns, where infringers benefit from using unauthorized paintings to train DMs and generate novel paintings in a similar style. To address these emerging copyright violations, in this paper, we are the first to explore and propose to utilize adversarial examples for DMs to protect human-created artworks. Specifically, we first build a theoretical framework to define and evaluate the adversarial examples for DMs. Then, based on this framework, we design a novel algorithm to generate these adversarial examples, named AdvDM, which exploits a Monte-Carlo estimation of adversarial examples for DMs by optimizing upon different latent variables sampled from the reverse process of DMs. Extensive experiments show that the generated adversarial examples can effectively hinder DMs from extracting their features. Therefore, our method can be a powerful tool for human artists to protect their copyright against infringers equipped with DM-based AI-for-Art applications. The code of our method is available on GitHub: https://github.com/mist-project/mist.git.
Chumeng Liang, Yang Hua 0001, Jiaru Zhang, Yiming Xue, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
ICML3
2023 FedCP: Separating Feature Information for Personalized Federated Learning via Conditional Policy
abstract
Recently, personalized federated learning (pFL) has attracted increasing attention in privacy protection, collaborative learning, and tackling statistical heterogeneity among clients, e.g., hospitals, mobile smartphones, etc. Most existing pFL methods focus on exploiting the global information and personalized information in the client-level model parameters while neglecting that data is the source of these two kinds of information. To address this, we propose the Federated Conditional Policy (FedCP) method, which generates a conditional policy for each sample to separate the global information and personalized information in its features and then processes them by a global head and a personalized head, respectively. FedCP is more fine-grained to consider personalization in a sample-specific manner than existing pFL methods. Extensive experiments in computer vision and natural language processing domains show that FedCP outperforms eleven state-of-the-art methods by up to 6.69%. Furthermore, FedCP maintains its superiority when some clients accidentally drop out, which frequently happens in mobile settings. Our code is public at https://github.com/TsingZ0/FedCP.
Jianqing Zhang, Yang Hua 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
KDD2
2023 Self-supervised Multi-object Tracking with Cycle-Consistency
Yuanhang Yin, Yang Hua 0001, Tao Song 0003, Ruhui Ma, Haibing Guan
MMM (2)2
2023 Eliminating Domain Bias for Federated Learning in Representation Space
abstract
Recently, federated learning (FL) is popular for its privacy-preserving and collaborative learning abilities. However, under statistically heterogeneous scenarios, we observe that biased data domains on clients cause a representation bias phenomenon and further degenerate generic representations during local training, i.e., the representation degeneration phenomenon. To address these issues, we propose a general framework Domain Bias Eliminator (DBE) for FL. Our theoretical analysis reveals that DBE can promote bi-directional knowledge transfer between server and client, as it reduces the domain discrepancy between server and client in representation space. Besides, extensive experiments on four datasets show that DBE can greatly improve existing FL methods in both generalization and personalization abilities. The DBE-equipped FL method can outperform ten state-of-the-art personalized FL methods by a large margin. Our code is public at https://github.com/TsingZ0/DBE.
Jianqing Zhang, Yang Hua 0001, Jian Cao 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
NeurIPS2
2023 WH2D2N2: Distributed AI-enabled OK-ASN Service for Web of Things
abstract
Model data-driven ontology and knowledge presentation for evolving semantic Asian social networks (OK-ASN) is a critical strategy for web of things (WoT) services. Meanwhile, Deep Neural Network (DNN)-based OK-ASN service in WoT is growing rapidly. However, most DNN-based services cannot utilize the potential of WoT fully, as heterogeneity exists in WoT. Therefore, this article proposes a novel framework called Web-based Heterogeneous Hierarchical Distributed Deep Neural Network ( WH 2 D 2 N 2 ) to deploy the DNNs for OK-ASN services on WoT, overcoming the heterogeneity. The architecture of the system and the designed Edge-Cloud-Joint execute scheme utilize heterogeneous devices to make DNN inference ubiquitous and output two types of results to meet various requirements. To bring robustness to OK-ASN services, a global scheduling is designed to arrange the workflow dynamically. The results of our experiments prove the efficiency of the execute scheme and the global scheduling in the system.
Ruhui Ma, Yang Hua 0001, Hao Wang 0022, Ningxin Hu, Tao Song 0003, Honghao Gao, Haibing Guan
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2023 Perceptual Data Augmentation for Biomedical Coronary Vessel Segmentation
abstract
Sufficient annotated data is critical to the success of deep learning methods. Annotating for vessel segmentation in X-ray coronary angiograms is extremely difficult because of the small and complex structures to be processed. Although unsupervised domain adaptation methods can be utilized to alleviate the annotation burden by using data in other domains, e.g., eye fundus images, these methods cannot perform well due to the characteristic of medical images. Data augmentation can help improve the similarity of source domain and target domain in unsupervised domain adaptation tasks. Existing data augmentation methods play a limited role in improving domain adaptation performance, especially for special medical image segmentation tasks. In this paper, we propose an effective perceptual data augmentation method to improve the similarity between eye fundus images and coronary angiograms by synthesizing virtual samples. Auto Foreground Augment method is designed to search for geometric transformations that improve the similarity between foreground vessels of eye fundus images and coronary angiograms. The Haar Wavelet-Based Perceptual Similarity Index is utilized to guide the synthesis of virtual samples in foreground and background mixup. Extensive experiments show that our data augmentation method can synthesize high-quality virtual samples and thus improve the domain adaptation performance. To our best knowledge, this is the first work to apply perceptual data augmentation to vessel segmentation in coronary angiograms.
Shuo Wang 0008, Yang Hua 0001, Ruhui Ma, Tao Song 0003, Zhengui Xue, Haibing Guan
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 Discrepancy-Guided Domain-Adaptive Data Augmentation
abstract
Data augmentation has been observed playing a crucial role in achieving better generalization in many machine learning tasks, especially in unsupervised domain adaptation (DA). It is particularly effective on visual object recognition tasks as images are high-dimensional with an enormous range of variations that can be simulated. Existing data augmentation techniques, however, are not explicitly designed to address the differences between different domains. Expert knowledge about the data is required, as well as manual efforts in finding the optimal parameters. In this article, we propose a novel domain-adaptive augmentation method by making use of a state-of-the-art style transfer method and domain discrepancy measurement. Specifically, we measure the discrepancy between source and target domains, and use it as a guide to augment the original source samples using style transferred source-to-target samples. The proposed domain-adaptive augmentation method is data and model agnostic that can be easily incorporated with state-of-the-art DA algorithms. We show empirically that, by using this domain-adaptive augmentation, we are able to gradually reduce the discrepancy between the source and target samples, and further boost the adaptation performance using different DA algorithms on three popular domain adaption datasets.
Jian Gao 0018, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
IEEE Trans. Neural Networks Learn. Syst.2
2022 Improving Bayesian Neural Networks by Adversarial Sampling
abstract
Bayesian neural networks (BNNs) have drawn extensive interest due to the unique probabilistic representation framework. However, Bayesian neural networks have limited publicized deployments because of the relatively poor model performance in real-world applications. In this paper, we argue that the randomness of sampling in Bayesian neural networks causes errors in the updating of model parameters during training and some sampled models with poor performance in testing. To solve this, we propose to train Bayesian neural networks with Adversarial Distribution as a theoretical solution. To avoid the difficulty of calculating Adversarial Distribution analytically, we further present the Adversarial Sampling method as an approximation in practice. We conduct extensive experiments with multiple network structures on different datasets, e.g., CIFAR-10 and CIFAR-100. Experimental results validate the correctness of the theoretical analysis and the effectiveness of the Adversarial Sampling on improving model performance. Additionally, models trained with Adversarial Sampling still keep their ability to model uncertainties and perform better when predictions are retained according to the uncertainties, which further verifies the generality of the Adversarial Sampling approach.
Jiaru Zhang, Yang Hua 0001, Tao Song 0003, Hao Wang 0022, Zhengui Xue, Ruhui Ma, Haibing Guan
AAAI2
2022 Efficient One-Stage Video Object Detection by Exploiting Temporal Consistency
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
ECCV (35)2
2022 TDViT: Temporal Dilated Video Transformer for Dense Video Tasks
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
ECCV (35)2
2022 Hierarchical CADNet: Learning from B-Reps for Machining Feature Recognition
abstract
Deep learning approaches have been shown to be capable of recognizing shape features (e.g. machining features) in Computer-Aided Design (CAD) models in certain circumstances, yet still have issues when the features intersect, and in exploiting the geometric and topological information which comprises the boundary representation (B-Rep) of the typical CAD model. This paper presents a novel hierarchical B-Rep graph shape representation which encodes information about the surface geometry and face topology of the B-Rep. To learn from this new shape representation, a novel hierarchical graph convolutional network called Hierarchical CADNet has been created, which has been shown to outperform other state-of-the-art neural architectures on feature identification, including machining features that intersect, with improvements in accuracy for some more complex CAD models.
Andrew R. Colligan, Trevor T. Robinson, Declan C. Nolan, Yang Hua 0001, Weijuan Cao
Comput. Aided Des.4
2022 Ranked List Loss for Deep Metric Learning
abstract
The objective of deep metric learning (DML) is to learn embeddings that can capture semantic similarity and dissimilarity information among data points. Existing pairwise or tripletwise loss functions used in DML are known to suffer from slow convergence due to a large proportion of trivial pairs or triplets as the model improves. To improve this, ranking-motivated structured losses are proposed recently to incorporate multiple examples and exploit the structured information among them. They converge faster and achieve state-of-the-art performance. In this work, we unveil two limitations of existing ranking-motivated structured losses and propose a novel ranked list loss to solve both of them. First, given a query, only a fraction of data points is incorporated to build the similarity structure. Consequently, some useful examples are ignored and the structure is less informative. To address this, we propose to build a set-based similarity structure by exploiting all instances in the gallery. The learning setting can be interpreted as few-shot retrieval: given a mini-batch, every example is iteratively used as a query, and the rest ones compose the gallery to search, i.e., the support set in few-shot setting. The rest examples are split into a positive set and a negative set. For every mini-batch, the learning objective of ranked list loss is to make the query closer to the positive set than to the negative set by a margin. Second, previous methods aim to pull positive pairs as close as possible in the embedding space. As a result, the intraclass data distribution tends to be extremely compressed. In contrast, we propose to learn a hypersphere for each class in order to preserve useful similarity structure inside it, which functions as regularisation. Extensive experiments demonstrate the superiority of our proposal by comparing with the state-of-the-art methods on the fine-grained image retrieval task. Our source code is available online: https://github.com/XinshaoAmosWang/Ranked-List-Loss-for-DML.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Neil Robertson 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Towards Ubiquitous Intelligent Computing: Heterogeneous Distributed Deep Neural Networks
abstract
For the pursuit of ubiquitous computing, distributed computing systems containing the cloud, edge devices, and Internet-of-Things devices are highly demanded. However, existing distributed frameworks do not tailor for the fast development of Deep Neural Network (DNN), which is the key technique behind many intelligent applications nowadays. Based on prior exploration on distributed deep neural networks (DDNN), we propose Heterogeneous Distributed Deep Neural Network (HDDNN) over the distributed hierarchy, targeting at ubiquitous intelligent computing. While being able to support basic functionalities of DNNs, our framework is optimized for various types of heterogeneity, including heterogeneous computing nodes, heterogeneous neural networks, and heterogeneous system tasks. Besides, our framework features parallel computing, privacy protection and robustness, with other consideration for the combination of heterogeneous distributed system and DNN. Extensive experiments demonstrate that our framework is capable of utilizing hierarchical distributed system better for DNN and tailoring DNN for real-world distributed system properly, which is with low response time, high performance, and better user experience.
Zongpu Zhang, Tao Song 0003, Yang Hua 0001, Xufeng He, Zhengui Xue, Ruhui Ma, Haibing Guan
IEEE Trans. Big Data4
2021 MAMBA: Multi-level Aggregation via Memory Bank for Video Object Detection
abstract
State-of-the-art video object detection methods maintain a memory structure, either a sliding window or a memory queue, to enhance the current frame using attention mechanisms. However, we argue that these memory structures are not efficient or sufficient because of two implied operations: (1) concatenating all features in memory for enhancement, leading to a heavy computational cost; (2) frame-wise memory updating, preventing the memory from capturing more temporal information. In this paper, we propose a multi-level aggregation architecture via memory bank called MAMBA. Specifically, our memory bank employs two novel operations to eliminate disadvantages of existing methods: (1) light-weight key-set construction which can significantly reduce the computational cost; (2) fine-grained feature-wise updating strategy which enables our method to utilize knowledge from the whole video. To better enhance features from complementary levels, i.e., feature maps and proposals, we further propose a generalized enhancement operation (GEO) to aggregate multi-level features in a unified manner. We conduct extensive evaluations on the challenging ImageNetVID dataset. Compared with existing state-of-the-art methods, our method achieves superior performance in terms of both speed and accuracy. More remarkably, MAMBA achieves mAP of 83.7%/84.6% at 12.6/9.1 FPS with ResNet-101.
Guanxiong Sun, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
AAAI2
2021 Temporal Meta-Adaptor for Video Object Detection
Yang Hua 0001, Jian Gao 0018, Neil Robertson 0002
BMVC2
2021 Siren: Byzantine-robust Federated Learning via Proactive Alarming
abstract
With the popularity of machine learning on many applications, data privacy has become a severe issue when machine learning is applied in the real world. Federated learning (FL), an emerging paradigm in machine learning, aims to train a centralized model while distributing training data among a large number of clients in order to avoid data privacy leaking, which has attracted great attention recently. However, the distributed training scheme in FL is susceptible to different kinds of attacks. Existing defense systems mainly utilize model weight analysis to identify malicious clients with many limitations. For example, some defense systems must know the exact number of malicious clients beforehand, which can be easily bypassed by well-designed attack methods and become impractical for real-world scenarios.
Hanxi Guo, Hao Wang 0022, Tao Song 0003, Yang Hua 0001, Zhangcheng Lv, Xiulang Jin, Zhengui Xue, Ruhui Ma, Haibing Guan
SoCC4
2021 ProSelfLC: Progressive Self Label Correction for Training Robust Deep Neural Networks
abstract
To train robust deep neural networks (DNNs), we systematically study several target modification approaches, which include output regularisation, self and non-self label correction (LC). Two key issues are discovered: (1) Self LC is the most appealing as it exploits its own knowledge and requires no extra models. However, how to automatically decide the trust degree of a learner as training goes is not well answered in the literature? (2) Some methods penalise while the others reward low-entropy predictions, prompting us to ask which one is better?To resolve the first issue, taking two well-accepted propositions–deep neural networks learn meaningful patterns before fitting noise [3] and minimum entropy regularisation principle [10]–we propose a novel end-to-end method named ProSelfLC, which is designed according to learning time and entropy. Specifically, given a data point, we progressively increase trust in its predicted label distribution versus its annotated one if a model has been trained for enough time and the prediction is of low entropy (high confidence). For the second issue, according to ProSelfLC, we empirically prove that it is better to redefine a meaningful low-entropy status and optimise the learner toward it. This serves as a defence of entropy minimisation.We demonstrate the effectiveness of ProSelfLC through extensive experiments in both clean and noisy settings. The source code is available at https://github.com/XinshaoAmosWang/ProSelfLC-CVPR2021.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, David A. Clifton, Neil Robertson 0002
CVPR2
2021 Robust Bayesian Neural Networks by Spectral Expectation Bound Regularization
abstract
Bayesian neural networks have been widely used in many applications because of the distinctive probabilistic representation framework. Even though Bayesian neural networks have been found more robust to adversarial attacks compared with vanilla neural networks, their ability to deal with adversarial noises in practice is still limited. In this paper, we propose Spectral Expectation Bound Regularization (SEBR) to enhance the robustness of Bayesian neural networks. Our theoretical analysis reveals that training with SEBR improves the robustness to adversarial noises. We also prove that training with SEBR can reduce the epistemic uncertainty of the model and hence it can make the model more confident with the predictions, which verifies the robustness of the model from another point of view. Experiments on multiple Bayesian neural network structures and different adversarial attacks validate the correctness of the theoretical findings and the effectiveness of the proposed approach.
Jiaru Zhang, Yang Hua 0001, Zhengui Xue, Tao Song 0003, Ruhui Ma, Haibing Guan
CVPR2
2021 Fine-Grained Pose Temporal Memory Module for Video Pose Estimation and Tracking
abstract
The task of video pose estimation and tracking has been largely improved with the development of image pose estimation recently. However, there are still many challenging cases, such as body part occlusion, fast body motion, camera zooming, and complex background. Most existing methods generally use the temporal information to get more precise human bounding boxes or just use it in the tracking stage, but they fail to improve the accuracy of pose estimation tasks. To better solve these problems and utilize the temporal information efficiently and effectively, we present a novel structure, called pose temporal memory module, which is flexible to be transferred into top-down pose estimation frameworks. The temporal information stored in the pose temporal memory is aggregated into the current frame feature in our proposed module. We also transfer compositional de-attention (CoDA) to solve the unique keypoint occlusion problem in this task and propose a novel keypoint feature replacement to recover the extreme error detection under fine-grained keypoint-level guidance. To verify the generality and effectiveness of our proposed method, we integrate our module into two widely used pose estimation frameworks and obtain notable improvement on the PoseTrack dataset with only a few extra computing resources.
Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICASSP2
2021 Self-Supervised Vessel Segmentation via Adversarial Learning
abstract
Vessel segmentation is critically essential for diagnosing a series of diseases, e.g., coronary artery disease and retinal disease. However, annotating vessel segmentation maps of medical images is notoriously challenging due to the tiny and complex vessel structures, leading to insufficient available annotated datasets for existing supervised methods and domain adaptation methods. The subtle structures and con-fusing background of medical images further suppress the efficacy of unsupervised methods. In this paper, we propose a self-supervised vessel segmentation method via adversarial learning. Our method learns vessel representations by training an attention-guided generator and a segmentation generator to simultaneously synthesize fake vessels and segment vessels out of coronary angiograms. To support the research, we also build the first X-ray angiography coronary vessel segmentation dataset, named XCAD. We evaluate our method extensively on multiple vessel segmentation datasets, including the XCAD dataset, the DRIVE dataset, and the STARE dataset. The experimental results show our method suppresses unsupervised methods significantly and achieves competitive performance compared with supervised methods and traditional methods.
Yang Hua 0001, Hanming Deng, Tao Song 0003, Hao Wang 0022, Zhengui Xue, Ruhui Ma, Haibing Guan
ICCV2
2021 Fast and Accurate Scene Parsing via Bi-Direction Alignment Networks
abstract
In this paper, we propose an effective method for fast and accurate scene parsing called Bidirectional Alignment Network (BiAlignNet). Previously, one representative work BiSeNet [1] uses two different paths (Context Path and Spatial Path) to achieve balanced learning of semantics and details, respectively. However, the relationship between the two paths is not well explored. We argue that both paths can benefit each other in a complementary way. Motivated by this, we propose a novel network by aligning two-path information into each other through a learned flow field. To avoid the noise and semantic gaps, we introduce a Gated Flow Alignment Module to align both features in a bidirectional way. Moreover, to make the Spatial Path learn more detailed information, we present an edge-guided hard pixel mining loss to supervise the aligned learning process. Our network achieves 80.1parcent and 78.5parcent mIoU in validation and test set of Cityscapes while running at 30 FPS with full resolution inputs. Code and models will be available at https://github.com/jojacola/BiAlignNet.
Yanran Wu, Xiangtai Li, Yunhai Tong, Yang Hua 0001, Tao Song 0003, Ruhui Ma, Haibing Guan
ICIP5
2021 Themis: A Fair Evaluation Platform for Computer Vision Competitions
abstract
It has become increasingly thorny for computer vision competitions to preserve fairness when participants intentionally fine-tune their models against the test datasets to improve their performance. To mitigate such unfairness, competition organizers restrict the training and evaluation process of participants' models. However, such restrictions introduce massive computation overheads for organizers and potential intellectual property leakage for participants. Thus, we propose Themis, a framework that trains a noise generator jointly with organizers and participants to prevent intentional fine-tuning by protecting test datasets from surreptitious manual labeling. Specifically, with the carefully designed noise generator, Themis adds noise to perturb test sets without twisting the performance ranking of participants' models. We evaluate the validity of Themis with a wide spectrum of real-world models and datasets. Our experimental results show that Themis effectively enforces competition fairness by precluding manual labeling of test sets and preserving the performance ranking of participants' models.
Zinuo Cai, Jianyong Yuan, Yang Hua 0001, Tao Song 0003, Hao Wang 0022, Zhengui Xue, Ningxin Hu, Jonathan Ding, Ruhui Ma, Mohammad R. Haghighat, Haibing Guan
IJCAI3
2021 Semantic-Aware Occlusion-Robust Network for Occluded Person Re-Identification
abstract
In recent years, deep learning-based person re-identification (Re-ID) methods have made significant progress. However, the performance of these methods substantially decreases when dealing with occlusion, which is ubiquitous in realistic scenarios. In this article, we propose a novel semantic-aware occlusion-robust network (SORN) that effectively exploits the intrinsic relationship between the tasks of person Re-ID and semantic segmentation for occluded person Re-ID. Specifically, the SORN is composed of three branches, including a local branch, a global branch, and a semantic branch. In particular, the local branch extracts part-based local features, and the global branch leverages a novel spatial-patch contrastive loss (SPC) to extract occlusion-robust global features. Meanwhile, the semantic branch generates a foreground-background mask for a pedestrian image, which indicates the non-occluded areas of the human body. The three branches are jointly trained in a unified multi-task learning network. Finally, pedestrian matching is performed based on the local features extracted from the non-occluded areas and the global features extracted from the whole pedestrian image. Extensive experimental results on a large-scale occluded person Re-ID dataset (i.e., Occluded-DukeMTMC) and two partial person Re-ID datasets (i.e., Partial-REID and Partial-iLIDS) show the superiority of the proposed method compared with several state-of-the-art methods for occluded and partial person Re-ID. We also demonstrate the effectiveness of the proposed method on two general person Re-ID datasets (i.e., Market-1501 and DukeMTMC-reID).
Yan Yan 0001, Jing-Hao Xue, Yang Hua 0001, Hanzi Wang
IEEE Trans. Circuits Syst. Video Technol.4
2020 Learning Deep Relations to Promote Saliency Detection
abstract
Though saliency detectors has made stunning progress recently. The performances of the state-of-the-art saliency detectors are not acceptable in some confusing areas, e.g., object boundary. We argue that the feature spatial independence should be one of the root cause. This paper explores the ubiquitous relations on the deep features to promote the existing saliency detectors efficiently. We establish the relation by maximizing the mutual information of the deep features of the same category via deep neural networks to break this independence. We introduce a threshold-constrained training pair construction strategy to ensure that we can accurately estimate the relations between different image parts in a self-supervised way. The relation can be utilized to further excavate the salient areas and inhibit confusing backgrounds. The experiments demonstrate that our method can significantly boost the performance of the state-of-the-art saliency detectors on various benchmark datasets. Besides, our model is label-free and extremely efficient. The inference speed is 140 FPS on a single GTX1080 GPU.
Changrui Chen, Xin Sun 0003, Yang Hua 0001, Junyu Dong, Hongwei Xv
AAAI3
2020 Reinforcing Neural Network Stability with Attractor Dynamics
abstract
Recent approaches interpret deep neural works (DNNs) as dynamical systems, drawing the connection between stability in forward propagation and generalization of DNNs. In this paper, we take a step further to be the first to reinforce this stability of DNNs without changing their original structure and verify the impact of the reinforced stability on the network representation from various aspects. More specifically, we reinforce stability by modeling attractor dynamics of a DNN and propose relu-max attractor network (RMAN), a light-weight module readily to be deployed on state-of-the-art ResNet-like networks. RMAN is only needed during training so as to modify a ResNet's attractor dynamics by minimizing an energy function together with the loss of the original learning task. Through intensive experiments, we show that RMAN-modified attractor dynamics bring a more structured representation space to ResNet and its variants, and more importantly improve the generalization ability of ResNet-like networks in supervised tasks due to reinforced stability.
Hanming Deng, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
AAAI2
2020 Reducing Distributional Uncertainty by Mutual Information Maximisation and Transferable Feature Learning
Jian Gao 0018, Yang Hua 0001, Guosheng Hu, Neil Robertson 0002
ECCV (23)2
2020 Improving Detection And Recognition Of Degraded Faces By Discriminative Feature Restoration Using GAN
abstract
Face detection and recognition in the wild is currently one of the most interesting and challenging problems. Many algorithms with high performance have already been proposed and applied in real-world applications. However, the problem of detecting and recognising degraded faces from low-quality images and videos mostly remains unsolved. In this paper, we present an algorithm capable of recovering facial features from low-quality videos and images. The resulting output image boosts the performance of existing face detection and recognition algorithms. It contains an effective method involving metric learning and different loss function components operating on different parts of the generator. This enhances the degraded faces by restoring their lost features rather than its perceptual quality. Our approach has been experimentally proven to enhance face detection and recognition, e.g., the face detection rate is improved by 3.08% for S3FD [1] and the area under the ROC curve for recognition is improved by 2.55% for ArcFace [2] on the SCFace dataset.
Soumya Shubhra Ghosh, Yang Hua 0001, Sankha S. Mukherjee, Neil Robertson 0002
ICIP2
2020 VTT: Long-term Visual Tracking with Transformers
abstract
Long-term visual tracking is a challenging problem. State-of-the-art long-term trackers, e.g., GlobalTrack, utilize region proposal networks (RPNs) to generate target proposals. However, the performance of the trackers is affected by occlusions and large scale or ratio variations. To address these issues, in this paper, we are the first to propose a novel architecture with transformers for long-term visual tracking. Specifically, the proposed Visual Tracking Transformer (VTT) utilizes a transformer encoder-decoder architecture for aggregating global information to deal with occlusion and large scale or ratio variation. Furthermore, it also shows better discriminative power against instance-level distractors without the need for extra labeling and hard-sample mining. We conduct extensive experiments on three large-scale long-term tracking datasets and have achieved state-of-the-art performance.
Tianling Bian, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICPR2
2020 Object-adaptive LSTM network for real-time visual tracking with adversarial data augmentation
Yihan Du, Yan Yan 0001, Si Chen 0002, Yang Hua 0001
Neurocomputing4
2020 Video salient object detection via spatiotemporal attention neural networks
Yi Tang 0008, Wenbin Zou, Yang Hua 0001, Zhi Jin 0002, Xia Li 0006
Neurocomputing3
2019 Deep Metric Learning by Online Soft Mining and Class-Aware Attention
abstract
Deep metric learning aims to learn a deep embedding that can capture the semantic similarity of data points. Given the availability of massive training samples, deep metric learning is known to suffer from slow convergence due to a large fraction of trivial samples. Therefore, most existing methods generally resort to sample mining strategies for selecting nontrivial samples to accelerate convergence and improve performance. In this work, we identify two critical limitations of the sample mining methods, and provide solutions for both of them. First, previous mining methods assign one binary score to each sample, i.e., dropping or keeping it, so they only selects a subset of relevant samples in a mini-batch. Therefore, we propose a novel sample mining method, called Online Soft Mining (OSM), which assigns one continuous score to each sample to make use of all samples in the mini-batch. OSM learns extended manifolds that preserve useful intraclass variances by focusing on more similar positives. Second, the existing methods are easily influenced by outliers as they are generally included in the mined subset. To address this, we introduce Class-Aware Attention (CAA) that assigns little attention to abnormal data samples. Furthermore, by combining OSM and CAA, we propose a novel weighted contrastive loss to learn discriminative embeddings. Extensive experiments on two fine-grained visual categorisation datasets and two video-based person re-identification benchmarks show that our method significantly outperforms the state-of-the-art.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Guosheng Hu, Neil Robertson 0002
AAAI2
2019 SC-RANK: Improving Convolutional Image Captioning with Self-Critical Learning and Ranking Metric-based Reward
Shiyang Yan, Yang Hua 0001, Neil Robertson 0002
BMVC2
2019 Ranked List Loss for Deep Metric Learning
abstract
The objective of deep metric learning (DML) is to learn embeddings that can capture semantic similarity information among data points. Existing pairwise or tripletwise loss functions used in DML are known to suffer from slow convergence due to a large proportion of trivial pairs or triplets as the model improves. To improve this, rankingmotivated structured losses are proposed recently to incorporate multiple examples and exploit the structured information among them. They converge faster and achieve state-of-the-art performance. In this work, we present two limitations of existing ranking-motivated structured losses and propose a novel ranked list loss to solve both of them. First, given a query, only a fraction of data points is incorporated to build the similarity structure. Consequently, some useful examples are ignored and the structure is less informative. To address this, we propose to build a setbased similarity structure by exploiting all instances in the gallery. The samples are split into a positive set and a negative set. Our objective is to make the query closer to the positive set than to the negative set by a margin. Second, previous methods aim to pull positive pairs as close as possible in the embedding space. As a result, the intraclass data distribution might be dropped. In contrast, we propose to learn a hypersphere for each class in order to preserve the similarity structure inside it. Our extensive experiments show that the proposed method achieves state-of-the-art performance on three widely used benchmarks.
Xinshao Wang, Yang Hua 0001, Elyor Kodirov, Guosheng Hu, Romain Garnier, Neil Robertson 0002
CVPR2
2019 Object Guided External Memory Network for Video Object Detection
abstract
Video object detection is more challenging than image object detection because of the deteriorated frame quality. To enhance the feature representation, state-of-the-art methods propagate temporal information into the deteriorated frame by aligning and aggregating entire feature maps from multiple nearby frames. However, restricted by feature map's low storage-efficiency and vulnerable content-address allocation, long-term temporal information is not fully stressed by these methods. In this work, we propose the first object guided external memory network for online video object detection. Storage-efficiency is handled by object guided hard-attention to selectively store valuable features, and long-term information is protected when stored in an addressable external data matrix. A set of read/write operations are designed to accurately propagate/allocate and delete multi-level memory feature under object guidance. We evaluate our method on the ImageNet VID dataset and achieve state-of-the-art performance as well as good speed-accuracy tradeoff. Furthermore, by visualizing the external memory, we show the detailed object-level reasoning process across frames.
Hanming Deng, Yang Hua 0001, Tao Song 0003, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICCV2
2019 Unsupervised Video Summarization with Attentive Conditional Generative Adversarial Networks
abstract
With the rapid growth of video data, video summarization technique plays a key role in reducing people's efforts to explore the content of videos by generating concise but informative summaries. Though supervised video summarization approaches have been well studied and achieved state-of-the-art performance, unsupervised methods are still highly demanded due to the intrinsic difficulty of obtaining high-quality annotations. In this paper, we propose a novel yet simple unsupervised video summarization method with attentive conditional Generative Adversarial Networks (GANs). Firstly, we build our framework upon Generative Adversarial Networks in an unsupervised manner. Specifically, the generator produces high-level weighted frame features and predicts frame-level importance scores, while the discriminator tries to distinguish between weighted frame features and raw frame features. Furthermore, we utilize a conditional feature selector to guide GAN model to focus on more important temporal regions of the whole video frames. Secondly, we are the first to introduce the frame-level multi-head self-attention for video summarization, which learns long-range temporal dependencies along the whole video sequence and overcomes the local constraints of recurrent units, e.g., LSTMs. Extensive evaluations on two datasets, SumMe and TVSum, show that our proposed framework surpasses state-of-the-art unsupervised methods by a large margin, and even outperforms most of the supervised methods. Additionally, we also conduct the ablation study to unveil the influence of each component and parameter settings in our framework.
Xufeng He, Yang Hua 0001, Tao Song 0003, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ACM Multimedia2
2019 GAN-Based Pose-Aware Regulation for Video-Based Person Re-Identification
abstract
Video-based person re-identification deals with the inherent difficulty of matching sequences with different length, unregulated, and incomplete target pose/viewpoint structure. Common approaches operate either by reducing the problem to the still images case, facing a significant information loss, or by exploiting inter-sequence temporal dependencies as in Siamese Recurrent Neural Networks or in gait analysis. However, in all cases, the inter-sequences pose/viewpoint misalignment is considered, and the existing spatial approaches are mostly limited to the still images context. To this end, we propose a novel approach that can exploit more effectively the rich video information, by accounting for the role that the changing pose/viewpoint factor plays in the sequences matching process. In particular, our approach consists of two components. The first one attempts to complement the original pose-incomplete information carried by the sequences with synthetic GAN-generated images, and fuse their features vectors into a more discriminative viewpoint-insensitive embedding, namely Weighted Fusion (WF). Another one performs an explicit pose-based alignment of sequence pairs to promote coherent feature matching, namely Weighted-Pose Regulation (WPR). Extensive experiments on two large video-based benchmark datasets show that our approach outperforms considerably existing methods.
Alessandro Borgia, Yang Hua 0001, Elyor Kodirov, Neil Robertson 0002
WACV2
2019 IEGAN: Multi-Purpose Perceptual Quality Image Enhancement Using Generative Adversarial Network
abstract
Despite the breakthroughs in quality of image enhancement, an end-to-end solution for simultaneous recovery of the finer texture details and sharpness for degraded images with low resolution is still unsolved. Some existing approaches focus on minimizing the pixel-wise reconstruction error which results in a high peak signal-to-noise ratio. The enhanced images fail to provide high-frequency details and are perceptually unsatisfying, i.e., they fail to match the quality expected in a photo-realistic image. In this paper, we present Image Enhancement Generative Adversarial Network (IEGAN), a versatile framework capable of inferring photo-realistic natural images for both artifact removal and super-resolution simultaneously. Moreover, we propose a new loss function consisting of a combination of reconstruction loss, feature loss and an edge loss counterpart. The feature loss helps to push the output image to the natural image manifold and the edge loss preserves the sharpness of the output image. The reconstruction loss provides low-level semantic information to the generator regarding the quality of the generated images compared to the original. Our approach has been experimentally proven to recover photo-realistic textures from heavily compressed low-resolution images on public benchmarks and our proposed high-resolution World100 dataset.
Soumya Shubhra Ghosh, Yang Hua 0001, Sankha S. Mukherjee, Neil Robertson 0002
WACV2
2019 Weakly Supervised Salient Object Detection With Spatiotemporal Cascade Neural Networks
abstract
Recently, deep learning techniques have substantially boosted the performance of salient object detection in still images. However, the salient object detection in videos by using traditional handcrafted features or deep learning features is not fully investigated, probably due to the lack of sufficient manually labeled video data for saliency modeling, especially for the data-driven deep learning. This paper proposes a novel weakly supervised approach to the salient object detection in a video, which can learn a robust saliency prediction model by using very limited manually labeled data and a large amount of weakly labeled data that could be easily generated in a supervised approach. Furthermore, we propose a spatiotemporal cascade neural network architecture for saliency modeling, in which two fully convolutional networks are cascaded to evaluate the visual saliency from both spatial and temporal cues to lead the optimal video saliency prediction. The proposed approach is extensively evaluated on the widely used challenging data sets, and the experiments demonstrate that our proposed approach substantially outperforms the state-of-the-art salient object detection models.
Yi Tang 0008, Wenbin Zou, Zhi Jin 0002, Yuhuan Chen, Yang Hua 0001, Xia Li 0006
IEEE Trans. Circuits Syst. Video Technol.5
2018 Deep Learning based Fetal Middle Cerebral Artery Segmentation in Large-scale Ultrasound Images
Shuo Wang 0008, Yang Hua 0001, Yunyun Cao, Tao Song 0003, Zhengui Xue, Xiaoping Gong, Guanjie Wang, Ruhui Ma, Haibing Guan
BIBM2
2018 Deep Multi-task Learning to Recognise Subtle Facial Expressions of Mental States
Guosheng Hu, Li Liu 0004, Yang Hua 0001, Zhihong Zhang 0001, Fumin Shen, Ling Shao 0001, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang
ECCV (12)5
2018 Object-Adaptive LSTM Network for Visual Tracking
abstract
Convolutional Neural Networks (CNNs) have shown outstanding performance in visual object tracking. However, most of classification-based tracking methods using CNNs are time-consuming due to expensive computation of complex online fine-tuning and massive feature extractions. Besides, these methods suffer from the problem of over-fitting since the training and testing stages of CNN models are based on the videos from the same domain. Recently, matching-based tracking methods (such as Siamese networks) have shown remarkable speed superiority, while they cannot well address target appearance variations and complex scenes for inherent lack of online adaptability and background information. In this paper, we propose a novel object-adaptive LSTM network, which can effectively exploit sequence dependencies and dynamically adapt to the temporal object variations via constructing an intrinsic model for object appearance and motion. In addition, we develop an efficient strategy for proposal selection, where the densely sampled proposals are firstly pre-evaluated using the fast matching-based method and then the well-selected high-quality proposals are fed to the sequence-specific learning LSTM network. This strategy enables our method to adaptively track an arbitrary object and operate faster than conventional CNN-based classification tracking methods. To the best of our knowledge, this is the first work to apply an LSTM network for classification in visual object tracking. Experimental results on OTB and TC-128 benchmarks show that the proposed method achieves state-of-the-art performance, which exhibits great potentials of recurrent structures for visual object tracking.
Yihan Du, Yan Yan 0001, Si Chen 0002, Yang Hua 0001, Hanzi Wang
ICPR4
2018 Tracking-assisted Weakly Supervised Online Visual Object Segmentation in Unconstrained Videos
abstract
This paper tackles the task of online video object segmentation with weak supervision, i.e., labeling the target object and background with pixel-level accuracy in unconstrained videos, given only one bounding box information in the first frame. We present a novel tracking-assisted visual object segmentation framework to achieve this. On the one hand, initialized with a given bounding box in the first frame, the auxiliary object tracking module guides the segmentation module frame by frame by providing motion and region information, which is usually missing in semi-supervised methods. Moreover, compared with the unsupervised approach, our approach with such minimum supervision can focus on the target object without bringing unrelated objects into the final results. On the other hand, the video object segmentation module also improves the robustness of the visual object tracking module by pixel-level localization and objectness information. Thus, segmentation and tracking in our framework can mutually help each other in an online manner. To verify the generality and effectiveness of the proposed framework, we evaluate our weakly supervised method on two cross-domain datasets, i.e., the DAVIS and VOT2016 datasets, with the same configuration and parameter setting. Experimental results show the top performance of our method, which is even better than the leading semi-supervised methods. Furthermore, we conduct the extensive ablation study on our approach to investigate the influence of each component and main parameters.
Zongpu Zhang, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ACM Multimedia2
2018 Cross-View Discriminative Feature Learning for Person Re-Identification
abstract
The viewpoint variability across a network of non-overlapping cameras is a challenging problem affecting person re-identification performance. In this paper, we investigate how to mitigate the cross-view ambiguity by learning highly discriminative deep features under the supervision of a novel loss function. The proposed objective is made up of two terms, the steering meta center term and the enhancing centers dispersion term, that steer the training process to mining effective intra-class and inter-class relationships in the feature domain of the identities. The effect of our loss supervision is to generate a more expanded feature space of compact classes where the overall level of the inter-identities' interference is reduced. Compared with the existing metric learning techniques, this approach has the advantage of achieving a better optimization because it jointly learns the embedding and the metric contextually. Our technique, by dismissing side-sources of performance gain, proves to enhance the CNN invariance to viewpoint without incurring increased training complexity (like in Siamese or triplet networks) and outperforms many related state-of-the-art techniques on Market-1501 and CUHK03.
Alessandro Borgia, Yang Hua 0001, Elyor Kodirov, Neil Robertson 0002
IEEE Trans. Image Process.2
2017 Attribute-Enhanced Face Recognition with Neural Tensor Fusion Networks
abstract
Deep learning has achieved great success in face recognition, however deep-learned features still have limited invariance to strong intra-personal variations such as large pose changes. It is observed that some facial attributes (e.g. eyebrow thickness, gender) are robust to such variations. We present the first work to systematically explore how the fusion of face recognition features (FRF) and facial attribute features (FAF) can enhance face recognition performance in various challenging scenarios. Despite the promise of FAF, we find that in practice existing fusion methods fail to leverage FAF to boost face recognition performance in some challenging scenarios. Thus, we develop a powerful tensor-based framework which formulates feature fusion as a tensor optimisation problem. It is nontrivial to directly optimise this tensor due to the large number of parameters to optimise. To solve this problem, we establish a theoretical equivalence between low-rank tensor optimisation and a two-stream gated neural network. This equivalence allows tractable learning using standard neural network optimisation tools, leading to accurate and stable optimisation. Experimental results show the fused feature works better than individual features, thus proving for the first time that facial attributes aid face recognition. We achieve state-of-the-art performance on three popular databases: MultiPIE (cross pose, lighting and expression), CASIA NIR-VIS2.0 (cross-modality environment) and LFW (uncontrolled environment).
Guosheng Hu, Yang Hua 0001, Zhihong Zhang 0001, Sankha S. Mukherjee, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang
ICCV2
2015 Online Object Tracking with Proposal Selection
abstract
Tracking-by-detection approaches are some of the most successful object trackers in recent years. Their success is largely determined by the detector model they learn initially and then update over time. However, under challenging conditions where an object can undergo transformations, e.g., severe rotation, these methods are found to be lacking. In this paper, we address this problem by formulating it as a proposal selection task and making two contributions. The first one is introducing novel proposals estimated from the geometric transformations undergone by the object, and building a rich candidate set for predicting the object location. The second one is devising a novel selection strategy using multiple cues, i.e., detection score and edgeness score computed from state-of-the-art object edges and motion boundaries. We extensively evaluate our approach on the visual object tracking 2014 challenge and online tracking benchmark datasets, and show the best performance.
Yang Hua 0001, Karteek Alahari, Cordelia Schmid
ICCV1
2015 Contextualizing Object Detection and Classification
abstract
We investigate how to iteratively and mutually boost object classification and detection performance by taking the outputs from one task as the context of the other one. While context models have been quite popular, previous works mainly concentrate on co-occurrence relationship within classes and few of them focus on contextualization from a top-down perspective, i.e. high-level task context. In this paper, our system adopts a new method for adaptive context modeling and iterative boosting. First, the contextualized support vector machine (Context-SVM) is proposed, where the context takes the role of dynamically adjusting the classification score based on the sample ambiguity, and thus the context-adaptive classifier is achieved. Then, an iterative training procedure is presented. In each step, Context-SVM, associated with the output context from one task (object classification or detection), is instantiated to boost the performance for the other task, whose augmented outputs are then further used to improve the former task by Context-SVM. The proposed solution is evaluated on the object classification and detection tasks of PASCAL Visual Object Classes Challenge (VOC) 2007, 2010 and SUN09 data sets, and achieves the state-of-the-art performance.
Qiang Chen 0007, Jian Dong 0011, ZhongYang Huang, Yang Hua 0001, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.5
2014 Occlusion and Motion Reasoning for Long-Term Tracking
Yang Hua 0001, Karteek Alahari, Cordelia Schmid
ECCV (6)1
2013 VideoPuzzle: Descriptive One-Shot Video Composition
abstract
A large amount of short, single-shot videos are created by personal camcorder every day, such as the small video clips in family albums, and thus a solution for presenting and managing these video clips is highly desired. From the perspective of professionalism and artistry, long-take/shot video, also termed one-shot video, is able to present events, persons or scenic spots in an informative manner. This paper presents a novel video composition system “Video Puzzle” which generates aesthetically enhanced long-shot videos from short video clips. Our task here is to automatically composite several related single shots into a virtual long-take video with spatial and temporal consistency. We propose a novel framework to compose descriptive long-take video with content-consistent shots retrieved from a video pool. For each video, frame-by-frame search is performed over the entire pool to find start-end content correspondences through a coarse-to-fine partial matching process. The content correspondence here is general and can refer to the matched regions or objects, such as human body and face. The content consistency of these correspondences enables us to design several shot transition schemes to seamlessly stitch one shot to another in a spatially and temporally consistent manner. The entire long-take video thus comprises several single shots with consistent contents and ίuent transitions. Meanwhile, with the generated matching graph of videos, the proposed system can also provide an efficient video browsing mode. Experiments are conducted on multiple video albums and the results demonstrate the effectiveness and the usefulness of the proposed scheme.
Qiang Chen 0007, Meng Wang 0001, ZhongYang Huang, Yang Hua 0001, Shuicheng Yan
IEEE Trans. Multim.4
2012 Hierarchical matching with side information for image classification
abstract
In this work, we introduce a hierarchical matching framework with so-called side information for image classification based on bag-of-words representation. Each image is expressed as a bag of orderless pairs, each of which includes a local feature vector encoded over a visual dictionary, and its corresponding side information from priors or contexts. The side information is used for hierarchical clustering of the encoded local features. Then a hierarchical matching kernel is derived as the weighted sum of the similarities over the encoded features pooled within clusters at different levels. Finally the new kernel is integrated with popular machine learning algorithms for classification purpose. This framework is quite general and flexible, other practical and powerful algorithms can be easily designed by using this framework as a template and utilizing particular side information for hierarchical clustering of the encoded local features. To tackle the latent spatial mismatch issues in SPM, we design in this work two exemplar algorithms based on two types of side information: object confidence map and visual saliency map, from object detection priors and within-image contexts respectively. The extensive experiments over the Caltech-UCSD Birds 200, Oxford Flowers 17 and 102, PASCAL VOC 2007, and PASCAL VOC 2010 databases show the state-of-the-art performances from these two exemplar algorithms.
Qiang Chen 0007, Yang Hua 0001, ZhongYang Huang, Shuicheng Yan
CVPR3
2011 Contextualizing object detection and classification
abstract
In this paper, we investigate how to iteratively and mutually boost object classification and detection by taking the outputs from one task as the context of the other one. First, instead of intuitive feature and context concatenation or postprocessing with context, the so-called Contextualized Support Vector Machine (Context-SVM) is proposed, where the context takes the responsibility of dynamically adjusting the classification hyperplane, and thus the context-adaptive classifier is achieved. Then, an iterative training procedure is presented. In each step, Context-SVM, associated with the output context from one task (object classification or detection), is instantiated to boost the performance for the other task, whose augmented outputs are then further used to improve the former task by Context-SVM. The proposed solution is evaluated on the object classification and detection tasks of PASCAL Visual Object Challenge (VOC) 2007 and 2010, and achieves the state-of-the-art performance.
Qiang Chen 0007, ZhongYang Huang, Yang Hua 0001, Shuicheng Yan
CVPR4