Tiansheng Huang

dblp:249/2114 · DBLP profile ↗
← Back
35ranked-venue papers
9as first author
34since 2021 · last 2026
0000-0002-4557-1865ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 16 since 2021Computer networks · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
abstract
Fine-tuning-as-a-service, while commercially successful for Large Language Model (LLM) providers, exposes models to harmful finetuning attacks.As a widely explored defense paradigm against such attacks, unlearning attempts to remove malicious knowledge from LLMs, thereby essentially preventing them from being used to perform malicious tasks.However, we highlight a critical flaw: the inherent general adaptability of LLMs allows them to easily bypass selective unlearning by rapidly relearning or repurposing their general capabilities for harmful tasks.To address this fundamental limitation, we propose a paradigm shift: instead of selective removal, we advocate for inducing model collapse, effectively forcing the model to "unlearn everything", specifically in response to updates characteristic of malicious adaptation.This collapse directly neutralizes the very general capabilities that attackers exploit, tackling the core issue unaddressed by selective unlearning.We introduce the Collapse Trap (CTRAP) as a practical mechanism to implement this concept conditionally.Embedded during alignment, CTRAP pre-configures the model's reaction to subsequent fine-tuning dynamics.If updates during fine-tuning constitute a persistent attempt to reverse safety alignment, the pre-configured trap triggers a progressive degradation of the model's core language modeling abilities, ultimately rendering it inert and useless for the attacker.Crucially, this collapse mechanism remains dormant during benign fine-tuning, ensuring the model's utility and general capabilities are preserved.1
Biao Yi, Tiansheng Huang, Baolei Zhang, Tong Li 0011, Lihai Nie, Zheli Liu, Li Shen 0008
ACL (1)2
2026 Prodigal: Backdoor defense for federated learning beyond robust aggregation
Guozhi Liu, Weiwei Lin 0001, Tiansheng Huang, Fang Shi, Xiumin Wang 0005, Li Shen 0008
Knowl. Based Syst.3
2026 Matching Accounts on Blockchain via Pseudo Fine-tuning of Language Models
abstract
Web 3.0, built on blockchain technology, prioritizes user privacy and autonomy, presenting new opportunities for financial systems while also complicating the regulation of illicit activities. In this study, we present a novel infrastructure named Pseudo Fine-tuning (PFT) that provides account matching services to combat financial crimes on account-based blockchains such as money laundering through coin-mixing services. The significance of PFT lies in overcoming the need for real labels to fine-tune language models for account matching, given the limited availability of labeled account pairs for the task. Specifically, our design involves (1) crafting pseudo-labeled pairs from transactions of an account across different periods, and (2) fine-tuning language models to distill knowledge from pseudo pairs, which is transferable to the target task. We provide an in-depth analysis to investigate the inherent knowledge acquired during the PFT process and the conditions conducive to its effectiveness. Comprehensive experiments on real-world datasets collected from coin-mixing services and ENS name services, corroborate that the framework delivers pronounced enhancements over state-of-the-art approaches. Our implementation is released at https://github.com/git-disl/PFT .
Sihao Hu, Tiansheng Huang, Fatih Ilhan, Selim F. Tekin, Greg Eisenhauer, Margaret L. Loper, Ling Liu 0001
ACM Trans. Intell. Syst. Technol.2
2026 PokéLLMon: A Grounding and Reasoning Benchmark for Large Language Models in Pokémon Battles
abstract
Developing grounding techniques for LLMs poses two requirements for interactive environments, i.e., (i) the presence of rich knowledge beyond the scope of existing LLMs and (ii) the complexity of tasks that require strategic reasoning. Existing environments fail to meet both requirements due to their simplicity or reliance on commonsense knowledge already encoded in LLMs for interaction. In this article, we present PokéLLMon, a new benchmark enriched with fictional game knowledge and characterized by the intense, dynamic, and adversarial gameplay of Pokémon battles, setting new challenges for the development of grounding and reasoning techniques in interactive environments. Empirical evaluations demonstrate that existing LLMs lack game knowledge and struggle in Pokémon battles. We investigate grounding techniques that leverage feedback and game knowledge, and provide a thorough analysis of reasoning methods from a new perspective of action consistency. Additionally, we introduce higher-level reasoning challenges when playing against human players. The implementation of our benchmark is released at: https://github.com/git-disl/PokeLLMon .
Sihao Hu, Tiansheng Huang, Gaowen Liu, Ramana Rao Kompella, Ling Liu 0001
ACM Trans. Internet Techn.2
2026 BlockEdge: A Hybrid Blockchain Framework for Secure and Efficient Collaboration in EEC Environments
abstract
In End-Edge-Cloud (EEC) computing environments, the diversity of devices often requires cloud-trained models to be adapted for end/edge devices, complicating decentralized project management. To address this, end/edge devices are increasingly using local model sharing instead of traditional cloud solutions. Popular platforms like GitHub and DockerHub lack the necessary data authenticity and security for high-stakes applications. While blockchain can ensure secure data sharing, permissioned blockchains struggle with the dynamic nature of EEC devices. To solve this, we propose BlockEdge, a hybrid blockchain architecture combining a permissioned blockchain with Practical Byzantine Fault Tolerance (PBFT) for cloud-based data management and a permissionless blockchain with Proof of Work (PoW) for decentralized model sharing at the end/edge. We enhance the PoW process with a dynamic mining algorithm and a lazy-loading Merkle tree structure, improving energy efficiency and computational performance. Experimental results show that BlockEdge reduces energy consumption by over 50% and cuts data update time by 91.73%, effectively addressing the energy and time inefficiencies of mainstream consensus mechanisms.
Wangbo Shen, Weiwei Lin 0001, Tiansheng Huang, Mian Guo, Haijie Wu
ACM Trans. Internet Techn.3
2026 Decentralized Federated Learning With Period Gradient Tracking Over Time-Varying Networks
abstract
To address the communication challenges associated with Federated Learning (FL), Decentralized Federated Learning (DFL) eliminates the central server and trains the model with decentralized method, enabling each client to only communicate with its neighbors. However, per our analysis, model trained with DFL experiences performance degradation because of data-heterogeneity and time-varying topologies. To address these issues, we propose a Dynamic K-step Gradient Tracking (DKGT) method to enhance the performance of DFL over time varying networks. Specifically, DKGT employs K-step local updates and gradient tracking to reduce the communication cost and the variance from heterogeneous data distribution, and we use dynamic gradient tracking parameter to correct gradient over time varying graph. Theoretically, we derive a universal convergence rate for smooth and non-convex problem at the rate of$\mathcal{O}\left(\frac{\left(f(\textbf{x}_0)-f(\textbf{x}^*)\right)}{\sqrt{T}(L\sqrt{KN})^{-1}-\tau(pKL\sqrt{TKN})^{-1}}+\frac{\sigma^2}{KTN\tau(pK-\tau)}\right)$, that τ and p respectively represent the time window length and the connectivity of time-varying networks. Experimentally, we illustrate the robustness and effectiveness of this heterogeneity correction on extensive non-convex neural network training tasks over different topologies and dynamic network settings.
Fang Shi, Yuehong Chen, Qiong Huang 0001, Tiansheng Huang, Guozhi Liu, Li Shen 0008
IEEE Trans. Parallel Distributed Syst.4
2026 GradCloak: Gradient Obfuscation for Privacy-Preserving Distributed Learning as a Service
abstract
Gradient leakage attacks pose a significant privacy threat in distributed learning-as-a-service APIs. Existing literature on gradient leakage defense relies on gradient perturbation for preventing privacy leakage. However, determining where and how much to perturb the gradient offers different capabilities for preventing gradient leakage. This paper presents GradCloak, a principled approach to guiding gradient perturbation with theoretical robustness bounds in federated learning as a service, aiming to find the minimum required noise for simultaneously achieving privacy protection, competitive accuracy, and preventing gradient leakage attacks. The paper is organized into three major components.First, we formulate the gradient leakage threats and their adverse effect. We categorize the attack into two broad types: leakage during local training and leakage before global aggregation.Second, we investigate different gradient perturbation approaches. We analyze and compare these gradient perturbation methods, which are performed at the federated server, with those performed at the participating client(s).Third, we introduce three robustness properties of robust perturbation against gradient leakage threats, formulated bythe anonymization boundfor training data robustness,the perturbation boundfor gradient robustness, andthe distribution robustness boundfor perturbed gradients. We conduct extensive evaluations on eight benchmark datasets to demonstrate that specific settings of gradient perturbation exist that best balance privacy, accuracy, and leakage prevention. Code is available athttps://github.com/git-disl/GradCloak.
Wenqi Wei 0001, Tiansheng Huang, Sihao Hu, Xinxin Fan, Rui Zhang 0066, Jingya Zhou, Ling Liu 0001
IEEE Trans. Serv. Comput.2
2026 MSFFEC-Net: enhanced polyp segmentation via multi-scale feature fusion with edge-aware enhancement and contrastive learning
Qiaohong Liu, Junshi Wang, Luo Zhi, Tiansheng Huang, Wenxia Bai
Vis. Comput.6
2025 Adversarial Attention Perturbations for Large Object Detection Transformers
abstract
Adversarial perturbations are useful tools for exposing vulnerabilities in neural networks. Existing adversarial perturbation methods for object detection are either limited to attacking CNN-based detectors or weak against transformer-based detectors. This paper presents an Attention-Focused Offensive Gradient (AFOG) attack against object detection transformers. By design, AFOG is neural-architecture agnostic and effective for attacking both large transformer-based object detectors and conventional CNN-based detectors with a unified adversarial attention framework. This paper makes three original contributions. First, AFOG utilizes a learnable attention mechanism that focuses perturbations on vulnerable image regions in multi-box detection tasks, increasing performance over non-attention baselines by up to 30.6%. Second, AFOG's attack loss is formulated by integrating two types of feature loss through learnable attention updates with iterative injection of adversarial perturbations. Finally, AFOG is an efficient and stealthy adversarial perturbation method. It probes the weak spots of detection transformers by adding strategically generated and visually imperceptible perturbations which can cause well-trained object detection models to fail. Extensive experiments conducted with twelve large detection transformers on COCO demonstrate the efficacy of AFOG. Our empirical results also show that AFOG outperforms existing attacks on transformer-based and CNN-based object detectors by up to 83% with superior speed and imperceptibility. Code is available at https://github.com/zacharyyahn/AFOG.
Zachary Yahn, Selim F. Tekin, Fatih Ilhan, Sihao Hu, Tiansheng Huang, Yichang Xu, Margaret L. Loper, Ling Liu 0001
ICCV5
2025 Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
abstract
Harmful fine-tuning attack poses serious safety concerns for large language models' fine-tuning-as-a-service. While existing defenses have been proposed to mitigate the issue, their performances are still far away from satisfactory, and the root cause of the problem has not been fully recovered. To this end, we in this paper show that \textit{harmful perturbation} over the model weights could be a probable cause of alignment-broken. In order to attenuate the negative impact of harmful perturbation, we propose an alignment-stage solution, dubbed Booster. Technically, along with the original alignment loss, we append a loss regularizer in the alignment stage's optimization. The regularizer ensures that the model's harmful loss reduction after the simulated harmful perturbation is attenuated, thereby mitigating the subsequent fine-tuning risk. Empirical results show that Booster can effectively reduce the harmful score of the fine-tuned models while maintaining the performance of downstream tasks. Our code is available at https://github.com/git-disl/Booster
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim F. Tekin, Ling Liu 0001
ICLR1
2025 Probe before You Talk: Towards Black-box Defense against Backdoor Unalignment for Large Language Models
abstract
Backdoor unalignment attacks against Large Language Models (LLMs) enable the stealthy compromise of safety alignment using a hidden trigger while evading normal safety auditing. These attacks pose significant threats to the applications of LLMs in the real-world Large Language Model as a Service (LLMaaS) setting, where the deployed model is a fully black-box system that can only interact through text. Furthermore, the sample-dependent nature of the attack target exacerbates the threat. Instead of outputting a fixed label, the backdoored LLM follows the semantics of any malicious command with the hidden trigger, significantly expanding the target space. In this paper, we introduce BEAT, a black-box defense that detects triggered samples during inference to deactivate the backdoor. It is motivated by an intriguing observation (dubbed the **probe concatenate effect**), where concatenated triggered samples significantly reduce the refusal rate of the backdoored LLM towards a malicious probe, while non-triggered samples have little effect. Specifically, BEAT identifies whether an input is triggered by measuring the degree of distortion in the output distribution of the probe before and after concatenation with the input. Our method addresses the challenges of sample-dependent targets from an opposite perspective. It captures the impact of the trigger on the refusal signal (which is sample-independent) instead of sample-specific successful attack behaviors. It overcomes black-box access limitations by using multiple sampling to approximate the output distribution. Extensive experiments are conducted on various backdoor attacks and LLMs (including the closed-source GPT-3.5-turbo), verifying the effectiveness and efficiency of our defense. Besides, we also preliminarily verify that BEAT can effectively defend against popular jailbreak attacks, as they can be regarded as "natural backdoors". Our source code is available at https://github.com/clearloveclearlove/BEAT.
Biao Yi, Tiansheng Huang, Sishuo Chen, Tong Li 0011, Zheli Liu, Zhixuan Chu, Yiming Li 0004
ICLR2
2025 Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
abstract
Safety aligned Large Language Models (LLMs) are vulnerable to harmful fine-tuning attacks – a few harmful data mixed in the fine-tuning dataset can break the LLMs’s safety alignment. While several defenses have been proposed, our evaluation shows that existing defenses fail when some specific training hyper-parameters are chosen – a large learning rate or a large number of training epochs in the fine-tuning stage can easily invalidate the defense. To this end, we propose Antidote, a post-fine-tuning stage solution, which remains agnostic to the training hyper-parameters in the fine-tuning stage. Antidote relies on the philosophy that by removing the harmful parameters, the harmful model can be recovered from the harmful behaviors, regardless of how those harmful parameters are formed in the fine-tuning stage. With this philosophy, we introduce a one-shot pruning stage after harmful fine-tuning to remove the harmful weights that are responsible for the generation of harmful content. Despite its embarrassing simplicity, empirical results show that Antidote can reduce harmful score while maintaining accuracy on downstream tasks.
Tiansheng Huang, Gautam Bhattacharya, Pratik Joshi, Josh Kimball, Ling Liu 0001
ICML1
2025 Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
abstract
Harmful fine-tuning attack introduces significant security risks to the fine-tuning services. Main-stream defenses aim to vaccinate the model such that the later harmful fine-tuning attack is less effective. However, our evaluation results show that such defenses are fragile-- with a few fine-tuning steps, the model still can learn the harmful knowledge. To this end, we do further experiment and find that an embarrassingly simple solution-- adding purely random perturbations to the fine-tuned model, can recover the model from harmful behaviors, though it leads to a degradation in the model’s fine-tuning performance. To address the degradation of fine-tuning performance, we further propose \methodname, which optimizes an adaptive perturbation that will be applied to the model after fine-tuning. \methodname maintains model's safety alignment performance without compromising downstream fine-tuning performance. Comprehensive experiments are conducted on different harmful ratios, fine-tuning tasks and mainstream LLMs, where the average harmful scores are reduced by up-to 21.2%, while maintaining fine-tuning performance. As a by-product, we analyze the adaptive perturbation and show that different layers in various LLMs have distinct safety coefficients. Source code available at https://github.com/w-yibo/Panacea.
Yibo Wang 0039, Tiansheng Huang, Li Shen 0008, Huanjin Yao, Haotian Luo, Naiqiang Tan, Jiaxing Huang 0001, Dacheng Tao
NeurIPS2
2025 Targeted Vaccine: Safety Alignment for Large Language Models Against Harmful Fine-Tuning via Layer-Wise Perturbation
abstract
Harmful fine-tuning attack poses a serious threat to the online fine-tuning service. Vaccine, a recent alignment-stage defense, applies uniform perturbation to all layers of embedding to make the model robust to the simulated embedding drift. However, applying layer-wise uniform perturbation may lead to excess perturbations for some particular non-safety-critical layers, resulting in defense performance degradation and unnecessary memory consumption. To address this limitation, we propose a Targeted Vaccine (T-Vaccine), a memory-efficient safety alignment method that applies perturbation to only selected layers of the model. T-Vaccine follows two core steps: First, it uses the harmful gradient norm as a statistical metric to identify the safety-critical layers. Second, instead of applying uniform perturbation across all layers, T-Vaccine only applies perturbation to the safety-critical layers while keeping other layers frozen during training. Results show that T-Vaccine outperforms Vaccine in terms of both defense effectiveness and resource efficiency. Comparison with other defense baselines, e.g., RepNoise and TAR also demonstrate the superiority of T-Vaccine. Notably, T-Vaccine is the first defense that enables a fine-tuning-based alignment method for 7B pre-trained models trained on consumer GPUs with limited memory (e.g., RTX 4090).
Guozhi Liu, Weiwei Lin 0001, Qi Mu, Tiansheng Huang, Ruichao Mo, Yuren Tao, Li Shen 0008
IEEE Trans. Inf. Forensics Secur.4
2025 Robust Few-Shot Ensemble Learning with Focal Diversity-Based Pruning
abstract
This article presents FusionShot, a focal diversity-optimized few-shot ensemble learning approach for boosting the robustness and generalization performance of pre-trained few-shot models. The article makes three original contributions. First, we explore the unique characteristics of few-shot learning to ensemble multiple few-shot (FS) models by creating three alternative fusion channels. Second, we introduce the concept of focal error diversity to learn the most efficient ensemble teaming strategy, rather than assuming that an ensemble of a larger number of base models will outperform those sub-ensembles of smaller size. We develop a focal diversity ensemble pruning method to effectively prune out the candidate ensembles with low ensemble error diversity and recommend top- \( K \) FS ensembles with the highest focal error diversity. Finally, we capture the complex non-linear patterns of ensemble few-shot predictions by designing the learn-to-combine algorithm, which can learn the diverse weight assignments for robust ensemble fusion over different member models. Extensive experiments on representative few-shot benchmarks show that the top-K ensembles recommended by FusionShot can outperform the representative state-of-the-art (SOTA) few-shot models on novel tasks (different distributions and unknown at training) and can prevail over existing few-shot learners in both cross-domain settings and adversarial settings. For reproducibility purposes, FusionShot trained models, results, and code are made available at https://github.com/sftekin/fusionshot .
Selim F. Tekin, Fatih Ilhan, Tiansheng Huang, Sihao Hu, Margaret L. Loper, Ling Liu 0001
ACM Trans. Intell. Syst. Technol.3
2025 AdaptiveFL: Communication-Adaptive Federated Learning Under Dynamic Bandwidth
abstract
Federated learning (FL) is a distributed machine learning paradigm that enables heterogeneous devices to train a model collaboratively. Recognizing communication as a bottleneck in FL, existing communication-efficient solutions, e.g., HeteroFL and LotteryFL, etc., utilize gradient sparsification to reduce communication costs. However, existing solutions fail to address the dynamic bandwidth issue in which the bandwidth of each client is constantly changing throughout the training process. In this article, we propose AdaptiveFL, a communication-adaptive FL framework, considering the dynamic constraints of bandwidth. The design of AdaptiveFL follows two key steps: 1) in each round, each device selects a best-fit sub-model for communication per currently available bandwidth; and 2) to guarantee the performance of each sub-model sent under dynamic bandwidth constraints, AdaptiveFL employs a local training method that enables each device to train a "tailorable" local model, which can be tailored to any sparsity with competitive accuracy. We compare AdaptiveFL with several communication-efficient SOTA methods and demonstrate that AdaptiveFL outperforms other baselines by a large margin.
Guozhi Liu, Weiwei Lin 0001, Tiansheng Huang, Fang Shi, Wentai Wu, Li Shen 0008
IEEE Trans. Neural Networks Learn. Syst.3
2024 Resource- Efficient Transformer Pruning for Finetuning of Large Models
abstract
With the recent advances in vision transformers and large language models (LLMs),finetuning costly large mod-els on downstream learning tasks poses significant chal-lenges under limited computational resources. This pa-per presents a REsource and ComputAtion-efficient Pruning framework (RECAP) for the finetuning of transformer-based large models. RECAP by design bridges the gap between efficiency and performance through an iterative process cycling between pruning, finetuning, and updating stages to explore different chunks of the given large-scale model. At each iteration, we first prune the model with Taylor-approximation-based importance estimation and then only update a subset of the pruned model weights based on the Fisher-information criterion. In this way, RE-CAP achieves two synergistic and yet conflicting goals: re-ducing the GPU memory footprint while maintaining model performance, unlike most existing pruning methods that re-quire the model to be finetuned beforehand for better preser-vation of model performance. We perform extensive exper-iments with a wide range of large transformer-based archi-tectures on various computer vision and natural language understanding tasks. Compared to recent pruning techniques, we demonstrate that RECAP offers significant im-provements in GPU memory efficiency, capable of reducing the footprint by up to 65%.
Fatih Ilhan, Gong Su, Selim F. Tekin, Tiansheng Huang, Sihao Hu, Ling Liu 0001
CVPR4
2024 Personalized Privacy Protection Mask Against Unauthorized Facial Recognition
Ka-Ho Chow 0001, Sihao Hu, Tiansheng Huang, Ling Liu 0001
ECCV (82)3
2024 Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
abstract
The new paradigm of fine-tuning-as-a-service introduces a new attack surface for Large Language Models (LLMs): a few harmful data uploaded by users can easily trick the fine-tuning to produce an alignment-broken model. We conduct an empirical analysis and uncover a \textit{harmful embedding drift} phenomenon, showing a probable cause of the alignment-broken effect. Inspired by our findings, we propose Vaccine, a perturbation-aware alignment technique to mitigate the security risk of users fine-tuning. The core idea of Vaccine is to produce invariant hidden embeddings by progressively adding crafted perturbation to them in the alignment phase. This enables the embeddings to withstand harmful perturbation from un-sanitized user data in the fine-tuning phase. Our results on open source mainstream LLMs (e.g., Llama2, Opt, Vicuna) demonstrate that Vaccine can boost the robustness of alignment against harmful prompts induced embedding drift while reserving reasoning ability towards benign prompts. Our code is available at https://github.com/git-disl/Vaccine.
Tiansheng Huang, Sihao Hu, Ling Liu 0001
NeurIPS1
2024 Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
abstract
Recent studies show that Large Language Models (LLMs) with safety alignment can be jail-broken by fine-tuning on a dataset mixed with harmful data. For the first time in the literature, we show that the jail-break effect can be mitigated by separating two states in the fine-tuning stage to respectively optimize over the alignment and user datasets. Unfortunately, our subsequent study shows that this simple Bi-State Optimization (BSO) solution experiences convergence instability when steps invested in its alignment state is too small, leading to downgraded alignment performance. By statistical analysis, we show that the \textit{excess drift} towards the switching iterates of the two states could be a probable reason for the instability. To remedy this issue, we propose \textbf{L}azy(\textbf{i}) \textbf{s}afety \textbf{a}lignment (\textbf{Lisa}), which introduces a proximal term to constraint the drift of each state. Theoretically, the benefit of the proximal term is supported by the convergence analysis, wherein we show that a sufficient large proximal factor is necessary to guarantee Lisa's convergence. Empirically, our results on four downstream fine-tuning tasks show that Lisa with a proximal term can significantly increase alignment performance while maintaining the LLM's accuracy on the user tasks. Code is available at https://github.com/git-disl/Lisa.
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim F. Tekin, Ling Liu 0001
NeurIPS1
2024 Adaptive Deep Neural Network Inference Optimization with EENet
abstract
Well-trained deep neural networks (DNNs) treat all test samples equally during prediction. Adaptive DNN inference with early exiting leverages the observation that some test examples can be easier to predict than others. This paper presents EENet, a novel early-exiting scheduling framework for multi-exit DNN models. Instead of having every sample go through all DNN layers during prediction, EENet learns an early exit scheduler, which can intelligently terminate the inference earlier for certain predictions, which the model has high confidence of early exit. As opposed to previous early-exiting solutions with heuristics-based methods, our EENet framework optimizes an early-exiting policy to maximize model accuracy while satisfying the given per-sample average inference budget. Extensive experiments are conducted on four computer vision datasets (CIFAR-10, CIFAR-100, ImageNet, Cityscapes) and two NLP datasets (SST-2, AgNews). The results demonstrate that the adaptive inference by EENet can outperform the representative existing early exit techniques. We also perform a detailed visualization analysis of the comparison results to interpret the benefits of EENet.
Fatih Ilhan, Ka-Ho Chow 0001, Sihao Hu, Tiansheng Huang, Selim F. Tekin, Wenqi Wei 0001, Yanzhao Wu 0001, Myungjin Lee, Ramana Rao Kompella, Hugo Latapie, Gaowen Liu, Ling Liu 0001
WACV4
2024 ZipZap: Efficient Training of Language Models for Large-Scale Fraud Detection on Blockchain
abstract
Language models (LMs) have demonstrated superior performance in detecting fraudulent activities on Blockchains. Nonetheless, the sheer volume of Blockchain data results in excessive memory and computational costs when training LMs from scratch, limiting their capabilities to large-scale applications. In this paper, we present ZipZap, a framework tailored to achieve both parameter and computational efficiency when training LMs on large-scale transaction data. First, with the frequency-aware compression, an LM can be compressed down to a mere 7.5% of its initial size with an imperceptible performance dip. This technique correlates the embedding dimension of an address with its occurrence frequency in the dataset, motivated by the observation that embeddings of low-frequency addresses are insufficiently trained and thus negating the need for a uniformly large dimension for knowledge representation. Second, ZipZap accelerates the speed through the asymmetric training paradigm: It performs transaction dropping and cross-layer parameter-sharing to expedite the pre-training process, while revert to the standard training paradigm for fine-tuning to strike a balance between efficiency and efficacy, motivated by the observation that the optimization goals of pre-training and fine-tuning are inconsistent. Evaluations on real-world, large-scale datasets demonstrate that ZipZap delivers notable parameter and computational efficiency improvements for training LMs. Our implementation is available at: https://github.com/git-disl/ZipZap.
Sihao Hu, Tiansheng Huang, Ka-Ho Chow 0001, Wenqi Wei 0001, Yanzhao Wu 0001, Ling Liu 0001
WWW2
2024 Diversity-driven Privacy Protection Masks Against Unauthorized Face Recognition
abstract
Face recognition (FR) technologies have enabled many life-enriching applications but have also opened doors for potential misuse. Governments, private companies, or even individuals can scrape the web, collect facial images, and build a face database to fuel the FR system to identify human faces without their consent. This paper introduces PMask to combat such a privacy threat against unauthorized FR. It provides a holistic approach to enable privacy-preserving sharing of facial images. PMask preprocesses the facial image and hides its unique facial signature through iterative optimization with dual goals: (i) minimizing the amount of noise to ensure high image quality and (ii) minimizing the perception loss between the privacy-protected face and the original face to ensure the face is recognizable to be the same person by humans. Extensive experiments are conducted on eight representative FR models to evaluate PMask against unauthorized FR. The results validate that PMask provides much stronger protection, introduces less perceptible changes to facial images, and runs faster than state-of-the-art methods to provide privacy protection with a better user experience.
Ka-Ho Chow 0001, Sihao Hu, Tiansheng Huang, Fatih Ilhan, Wenqi Wei 0001, Ling Liu 0001
Proc. Priv. Enhancing Technol.3
2023 FedSpeed: Larger Local Interval, Less Communication Round, and Higher Generalization Accuracy
Li Shen 0008, Tiansheng Huang, Liang Ding 0006, Dacheng Tao
ICLR3
2023 Lockdown: Backdoor Defense for Federated Learning with Isolated Subspace Training
abstract
Federated learning (FL) is vulnerable to backdoor attacks due to its distributed computing nature. Existing defense solution usually requires larger amount of computation in either the training or testing phase, which limits their practicality in the resource-constrain scenarios. A more practical defense, i.e., neural network (NN) pruning based defense has been proposed in centralized backdoor setting. However, our empirical study shows that traditional pruning-based solution suffers \textit{poison-coupling} effect in FL, which significantly degrades the defense performance.This paper presents Lockdown, an isolated subspace training method to mitigate the poison-coupling effect. Lockdown follows three key procedures. First, it modifies the training protocol by isolating the training subspaces for different clients. Second, it utilizes randomness in initializing isolated subspacess, and performs subspace pruning and subspace recovery to segregate the subspaces between malicious and benign clients. Third, it introduces quorum consensus to cure the global model by purging malicious/dummy parameters. Empirical results show that Lockdown achieves \textit{superior} and \textit{consistent} defense performance compared to existing representative approaches against backdoor attacks. Another value-added property of Lockdown is the communication-efficiency and model complexity reduction, which are both critical for resource-constrain FL scenario. Our code is available at \url{https://github.com/git-disl/Lockdown}.
Tiansheng Huang, Sihao Hu, Ka-Ho Chow 0001, Fatih Ilhan, Selim F. Tekin, Ling Liu 0001
NeurIPS1
2023 Joint Task Offloading and Service Placement for Mobile Edge Computing: An Online Two-Timescale Approach
abstract
As a new computing paradigm, mobile edge computing (MEC) pushes the centralized cloud resources close to the edge network, which significantly reduces the pressure of the backbone network and meets the requirements of emerging mobile applications. To achieve high performance of the MEC system, it is essential to design efficient task offloading and service placement schemes, which are responsible for offloading tasks to the edge servers while considering the heterogeneity and diversity of computation services. Our MEC system aims to maximize the long-term average network utility while maintaining the stability of the edge network. Considering that synchronous manner overlooks the scenarios endowed with asymmetric update frequencies for service placement and task offloading, we propose an online algorithm based on the two-timescale Lyapunov optimization in a stochastic network environment without requiring the future information. By making asynchronous decisions on service placement and task offloading with different control parameters$V$, we can achieve a time-average sub-optimal solution that is close to the offline optimum. In addition, we introduce the varying control parameter$V(t)$and$\Omega$-additive approximation to enhance the robustness of the proposed algorithm within an error$\Omega$. Finally, rigorous theoretical analysis and extensive trace-driven experimental results show that the proposed algorithm achieves the$[O(1/V), O(V)]$performance-backlog tradeoff and is more competitive than benchmarks.
Xin Li 0116, Xinglin Zhang 0001, Tiansheng Huang
IEEE Trans. Cloud Comput.3
2022 Contribution-based Federated Learning client selection
abstract
Federated Learning (FL), as a privacy-preserving machine learning paradigm, has been thrusted into the limelight. As a result of the physical bandwidth constraint, only a small number of clients are selected for each round of FL training. However, existing client selection solutions (e.g., the vanilla random selection) typically ignore the heterogeneous data value of the clients. In this paper, we propose the contribution-based selection algorithm (Contribution-Based Exponential-weight algorithm for Exploration and Exploitation, CBE3), which dynamically updates the selection weights according to the impact of clients' data. As a novel component of CBE3, a scaling factor, which helps maintain a good balance between global model accuracy and convergence speed, is proposed to improve the algorithm's adaptability. Theoretically, we proved the regret bound of the proposed CBE3 algorithm, which demonstrates performance gaps between the CBE3 and the optimal choice. Empirically, extensive experiments conducted on Non-Independent Identically Distributed data demonstrate the superior performance of CBE3—with up to 10% accuracy improvement compared with K-Center and Greedy and up to 100% faster convergence compared with the Random algorithm.
Weiwei Lin 0001, Yinhai Xu, Bo Liu 0001, Dongdong Li 0002, Tiansheng Huang, Fang Shi
Int. J. Intell. Syst.5
2022 Adaptive Processor Frequency Adjustment for Mobile-Edge Computing With Intermittent Energy Supply
abstract
With astonishing speed, bandwidth, and scale, mobile-edge computing (MEC) has played an increasingly important role in the next generation of connectivity and service delivery. Yet, along with the massive deployment of MEC servers, the ensuing energy issue is now on an increasingly urgent agenda. In the current context, the large-scale deployment of renewable-energy-supplied MEC servers is perhaps the most promising solution for the incoming energy issue. Nonetheless, as a result of the intermittent nature of their power sources, these special design MEC servers must be more cautious about their energy usage, in a bid to maintain their service sustainability as well as service standard. Targeting optimization on a single-server MEC scenario, we, in this article, propose neural network-based adaptive frequency adjustment (NAFA), an adaptive processor frequency adjustment solution, to enable an effective plan of the server’s energy usage. By learning from the historical data revealing request arrival and energy harvest pattern, the deep reinforcement learning-based solution is capable of making intelligent schedules on the server’s processor frequency, so as to strike a good balance between service sustainability and service quality. The superior performance of NAFA is substantiated by real-data-based experiments, wherein NAFA demonstrates up to 20% increase in the average request acceptance ratio and up to 50% reduction in average request processing time.
Tiansheng Huang, Weiwei Lin 0001, Xiumin Wang 0005, Qingbo Wu 0003, Rui Li 0047, Ching-Hsien Hsu, Albert Y. Zomaya
IEEE Internet Things J.1
2022 Stochastic Client Selection for Federated Learning With Volatile Clients
abstract
Federated learning (FL), arising as a privacy-preserving machine learning paradigm, has received notable attention from the public. In each round of synchronous FL training, only a fraction of available clients are chosen to participate, and the selection decision might have a significant effect on the training efficiency, as well as the final model performance. In this article, we investigate the client selection problem under a volatile context, in which the local training of heterogeneous clients is likely to fail due to various kinds of reasons and in different levels of frequency. Intuitively, too much training failure might potentially reduce the training efficiency, while too much selection on clients with greater stability might introduce bias, thereby resulting in degradation of the training effectiveness. To tackle this tradeoff, we, in this article, formulate the client selection problem under joint consideration of effective participation and fairness. Furthermore, we propose E3CS, a stochastic client selection scheme as a solution. According to our experimental results over a public data set, the proposed selection scheme is able to achieve up to$2\times $faster convergence to a fixed model accuracy while maintaining the same level of final model accuracy, compared with the state-of-the-art selection schemes.
Tiansheng Huang, Weiwei Lin 0001, Li Shen 0008, Keqin Li 0001, Albert Y. Zomaya
IEEE Internet Things J.1
2022 VFedCS: Optimizing Client Selection for Volatile Federated Learning
abstract
Federated learning (FL) has shown great potential as a privacy-preserving solution to training a centralized model based on local data from available clients. However, we argue that, over the course of training, the available clients may exhibit some volatility in terms of the client population, client data, and training status. Considering these volatilities, we propose a new learning scenario termed volatile federated learning (volatile FL) featuring set volatility, statistical volatility, and training volatility. The volatile client set along with the dynamic of clients’ data and the unreliable nature of clients (e.g., unintentional shutdown and network instability) greatly increase the difficulty of client selection. In this article, we formulate and decompose the global problem into two subproblems based on alternating minimization. For an efficient settlement for the proposed selection problem, we quantify the impact of clients’ data and resource heterogeneity for volatile FL and introduce the cumulative effective participation data (CEPD) as an optimization objective. Based on this, we propose upper confidence bound-based greedy selection, dubbed UCB-GS, to address the client selection problem in volatile FL. Theoretically, we prove that the regret of UCB-GS is strictly bounded by a finite constant, justifying its theoretical feasibility. Furthermore, experimental results show that our method significantly reduces the number of training rounds (by up to 62%) while increasing the global model’s accuracy by 7.51%.
Fang Shi, Chunchao Hu, Weiwei Lin 0001, Lisheng Fan, Tiansheng Huang, Wentai Wu
IEEE Internet Things J.5
2022 Energy-Efficient Computation Offloading for UAV-Assisted MEC: A Two-Stage Optimization Scheme
abstract
In addition to the stationary mobile edge computing (MEC) servers, a few MEC surrogates that possess a certain mobility and computation capacity, e.g., flying unmanned aerial vehicles (UAVs) and private vehicles, have risen as powerful counterparts for service provision. In this article, we design a two-stage online scheduling scheme, targeting computation offloading in a UAV-assisted MEC system. On our stage-one formulation, an online scheduling framework is proposed for dynamic adjustment of mobile users' CPU frequency and their transmission power, aiming at producing a socially beneficial solution to users. But the major impediment during our investigation lies in that users might not unconditionally follow the scheduling decision released by servers as a result of their individual rationality. In this regard, we formulate each step of online scheduling on stage one into a non-cooperative game with potential competition over the limited radio resource. As a solution, a centralized online scheduling algorithm, called ONCCO, is proposed, which significantly promotes social benefit on the basis of the users' individual rationality. On our stage-two formulation, we are working towards the optimization of UAV computation resource provision, aiming at minimizing the energy consumption of UAVs during such a process, and correspondingly, another algorithm, called WS-UAV, is given as a solution. Finally, extensive experiments via numerical simulation are conducted for an evaluation purpose, by which we show that our proposed algorithms achieve satisfying performance enhancement in terms of energy conservation and sustainable service provision.
Weiwei Lin 0001, Tiansheng Huang, Xin Li 0116, Fang Shi, Xiumin Wang 0005, Ching-Hsien Hsu
ACM Trans. Internet Techn.2
2021 Asynchronous Online Service Placement and Task Offloading for Mobile Edge Computing
abstract
Mobile edge computing (MEC) pushes the centralized cloud resources close to the edge network, which significantly reduces the pressure of the backbone network and meets the requirements of emerging mobile applications. To achieve high performance of the MEC system, it is essential to design efficient task offloading schemes. Many existing works focus on offloading tasks to the edge servers while ignoring the heterogeneity and diversity of computation services, which is also important in MEC. In this paper, we investigate the joint problem of online task offloading and service placement-downloading and deploying the service-related resources at edge servers-in the dense MEC network. Our MEC system aims to maximize the long-term average network utility while maintaining the stability of the edge network. Due to the uncertainty of task demands, it is impossible to make an online long-term optimal decision. Therefore, we propose an online algorithm based on the two-timescale Lyapunov optimization without requiring the future information. By making asynchronous decisions on service placement and task offloading, we can achieve a time-average sub-optimal solution that is close to the offline optimum. In addition, rigorous theoretical analysis and extensive trace-driven experimental results show that the proposed algorithm is more competitive than benchmarks.
Xin Li 0116, Xinglin Zhang 0001, Tiansheng Huang
SECON3
2021 An Ant Colony Optimization-Based Multiobjective Service Replicas Placement Strategy for Fog Computing
abstract
In recent years, fog computing has emerged as a new paradigm for the future Internet-of-Things (IoT) applications, but at the same time, ensuing new challenges. The geographically vast-distributed architecture in fog computing renders us almost infinite choices in terms of service orchestration. How to properly arrange the service replicas (or service instances) among the nodes remains a critical problem. To be specific, in this article, we investigate a generalized service replicas placement problem that has the potential to be applied to various industrial scenarios. We formulate the problem into a multiobjective model with two scheduling objectives, involving deployment cost and service latency. For problem solving, we propose an ant colony optimization-based solution, called multireplicas Pareto ant colony optimization (MRPACO). We have conducted extensive experiments on MRPACO. The experimental results show that the solutions obtained by our strategy are qualified in terms of both diversity and accuracy, which are the main evaluation metrics of a multiobjective algorithm.
Tiansheng Huang, Weiwei Lin 0001, Chennian Xiong, Jingxuan Huang
IEEE Trans. Cybern.1
2021 An Efficiency-Boosting Client Selection Scheme for Federated Learning With Fairness Guarantee
abstract
The issue of potential privacy leakage during centralized AI's model training has drawn intensive concern from the public. A Parallel and Distributed Computing (or PDC) scheme, termed Federated Learning (FL), has emerged as a new paradigm to cope with the privacy issue by allowing clients to perform model training locally, without the necessity to upload their personal sensitive data. In FL, the number of clients could be sufficiently large, but the bandwidth available for model distribution and re-upload is quite limited, making it sensible to only involve part of the volunteers to participate in the training process. The client selection policy is critical to an FL process in terms of training efficiency, the final model's quality as well as fairness. In this article, we will model the fairness guaranteed client selection as a Lyapunov optimization problem and then a C2MAB-based method is proposed for estimation of the model exchange time between each client and the server, based on which we design a fairness guaranteed algorithm termed RBCS-F for problem-solving. The regret of RBCS-F is strictly bounded by a finite constant, justifying its theoretical feasibility. Barring the theoretical results, more empirical data can be derived from our real training experiments on public datasets.
Tiansheng Huang, Weiwei Lin 0001, Wentai Wu, Ligang He, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Parallel Distributed Syst.1
2020 Virtual Machine Consolidation for NUMA Systems: A Hybrid Heuristic Grey Wolf Approach
abstract
Virtual machines consolidation is known as a powerful means to reduce the number of activated physical machines (PMs), so as to achieve energy-saving for the data centers. Although the consolidation technique is widely studied in non-NUMA systems, we could only trace a few studies targeting NUMA systems. But the virtual machines (VMs) deployment of NUMA systems is quite different from that of non-NUMA systems. More specifically, consolidating VMs in NUMA systems need to decide both target physical machines and NUMA architectures to host the VMs, and more complicated constraints originated from the real usage of NUMA systems that need to be considered. Being motivated by these challenges, we in this paper formally derive the system model according to the real business model of NUMA systems and based on which, we propose a hybrid heuristics swarm intelligence optimization algorithm HHGWA for an efficient solution. To do the evaluation, extensive simulations that integrate real VM and PM information are conducted, the result of which indicates a superior performance of our proposed algorithm.
Kangli Hu, Weiwei Lin 0001, Tiansheng Huang, Keqin Li 0001, Like Ma
ICPADS3