Ka-Ho Chow 0001

dblp:367/9233-1 · also Ka Ho Chow 0001 · DBLP profile ↗
← Back
32ranked-venue papers
11as first author
26since 2021 · last 2026
0000-0001-5917-2577ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 8 since 2021Security and privacy · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GDBR: Label Recovery Attack Against Partial Gradient Encryption in Federated Learning
abstract
The increasing demand for data privacy, alongside the benefits of aggregating data from networked devices, has catalyzed the emergence of federated learning (FL). In FL, clients jointly train a global model by sharing gradients computed over private data. While this paradigm eliminates the need to exchange raw data, inference attacks can still be launched to extract sensitive information from gradients. To this end, partial gradient encryption has emerged as a promising design for balancing privacy and efficiency in practical FL systems, as encrypting only the classification-head gradients is believed to prevent known inference attacks while avoiding the high computational cost of encrypting the entire model. However, this design provides a false sense of privacy. By proposing GDBR, we show that sharing even a single unencrypted layer of gradients can lead to serious privacy leakage. GDBR is the first attack capable of high-fidelity label recovery with partial access to the gradients. It exploits a vulnerability in a commonly used neural building block, constructs a gradient bridge from the unencrypted layer to the final output layer, and approximates the logits information for accurate inference of private labels. These inferred labels not only reveal sensitive information about a client's private dataset but also serve as a prerequisite for many downstream attacks, such as data reconstruction and membership inference. GDBR brings these threats squarely into scope for FL systems employing partial encryption. In addition to theoretical analysis, extensive experiments demonstrate the severity of the problem across a wide variety of datasets and model architectures, including convolutional and transformer-based networks. Overall, our findings challenge the widespread assumption that encrypting only the output layer suffices for privacy protection.
Ka-Ho Chow 0001
EuroS&P2
2025 On the Adversarial Robustness of Graph Neural Networks with Graph Reduction
Kerui Wu, Ka-Ho Chow 0001, Wenqi Wei 0001, Lei Yu 0002
ESORICS (1)2
2025 Harmless Backdoor-based Client-side Watermarking in Federated Learning
abstract
Protecting intellectual property (IP) in federated learning (FL) is increasingly important as clients contribute proprietary data to collaboratively train models. Model watermarking, particularly through backdoor-based methods, has emerged as a popular approach for verifying ownership and contributions in deep neural networks trained via FL. By manipulating their datasets, clients can embed a secret pattern, resulting in non-intuitive predictions that serve as proof of participation, useful for claiming incentives or IP co-ownership. However, this technique faces practical challenges: (i) client watermarks can collide, leading to ambiguous ownership claims, and (ii) malicious clients may exploit watermarks to manipulate model predictions for harmful purposes. To address these issues, we propose Sanitizer, a server-side method that ensures client-embedded backdoors can only be activated in harmless environments but not natural queries. It identifies subnets within client-submitted models, extracts backdoors throughout the FL process, and confines them to harmless, client-specific input subspaces. This approach not only enhances Sanitizer’s efficiency but also resolves conflicts when clients use similar triggers with different target labels. Our empirical results demonstrate that Sanitizer achieves near-perfect success verifying client contributions while mitigating the risks of malicious watermark use. Additionally, it reduces GPU memory consumption by 85% and cuts processing time by at least 5× compared to the baseline. Our code is open-sourced at https://hku-tasr.github.io/Sanitizer/.
Kaijing Luo, Ka-Ho Chow 0001
EuroS&P2
2025 Geminio: Language-Guided Gradient Inversion Attacks in Federated Learning
abstract
Foundation models that bridge vision and language have made significant progress. While they have inspired many life-enriching applications, their potential for abuse in creating new threats remains largely unexplored. In this paper, we reveal that vision-language models (VLMs) can be weaponized to enhance gradient inversion attacks (GIAs) in federated learning (FL), where an FL server attempts to reconstruct private data samples from gradients shared by victim clients. Despite recent advances, existing GIAs struggle to reconstruct high-resolution images when the victim has a large local data batch. One promising direction is to focus reconstruction on valuable samples rather than the entire batch, but current methods lack the flexibility to target specific data of interest. To address this gap, we propose Geminio, the first approach to transform GIAs into semantically meaningful, targeted attacks. It enables a brand new privacy attack experience: attackers can describe, in natural language, the data they consider valuable, and Geminio will prioritize reconstruction to focus on those high-value samples. This is achieved by leveraging a pretrained VLM to guide the optimization of a malicious global model that, when shared with and optimized by a victim, retains only gradients of samples that match the attacker-specified query. Geminio can be launched at any FL round and has no impact on normal training (i.e., the FL server can steal clients' data while still producing a high-utility ML model as in benign scenarios). Extensive experiments demonstrate its effectiveness in pinpointing and reconstructing targeted samples, with high success rates across complex datasets and large batch sizes with resilience against defenses.
Junjie Shan, Jialin Lu, Siu-Ming Yiu, Ka-Ho Chow 0001
ICCV6
2025 OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation
abstract
Retrieval-augmented Generation (RAG) enhances Large Language Models (LLMs) by integrating external knowledge to reduce hallucinations and incorporate up-to-date information without retraining. As an essential part of RAG, external knowledge bases are commonly built by extracting structured data from unstructured PDF documents using Optical Character Recognition (OCR). However, given the imperfect prediction of OCR and the inherent non-uniform representation of structured data, knowledge bases inevitably contain various OCR noises. In this paper, we introduce OHRBench, the first benchmark for understanding the cascading impact of OCR on RAG systems. OHRBench includes 8,561 carefully selected unstructured document images from seven real-world RAG application domains, along with 8,498 Q&A pairs derived from multimodal elements in documents, challenging existing OCR solutions used for RAG. To better understand OCR's impact on RAG systems, we identify two primary types of OCR noise: Semantic Noise and Formatting Noise and apply perturbation to generate a set of structured data with varying degrees of each OCR noise. Using OHRBench, we first conduct a comprehensive evaluation of current OCR solutions and reveal that none is competent for constructing high-quality knowledge bases for RAG systems. We then systematically evaluate the impact of these two noise types and demonstrate the trend relationship between the degree of OCR noise and RAG performance. Our OHRBench, including PDF documents, Q&As, and the ground truth structured data are released at: https://github.com/opendatalab/OHR-Bench
Junyuan Zhang, Qintong Zhang, Bin Wang 0065, Linke Ouyang, Zichen Wen, Ka-Ho Chow 0001, Conghui He, Wentao Zhang 0001
ICCV7
2024 Personalized Privacy Protection Mask Against Unauthorized Facial Recognition
Ka-Ho Chow 0001, Sihao Hu, Tiansheng Huang, Ling Liu 0001
ECCV (82)1
2024 Atlas: Hybrid Cloud Migration Advisor for Interactive Microservices
abstract
Hybrid cloud provides an attractive solution to microservices for better resource elasticity. A subset of application components can be offloaded from the on-premises cluster to the cloud, where they can readily access additional resources. However, the selection of this subset is challenging because of the large number of possible combinations. A poor choice degrades the application performance, disrupts the critical services, and increases the cost to the extent of making the use of hybrid cloud unviable. This paper presents Atlas, a hybrid cloud migration advisor. Atlas uses a data-driven approach to learn how each user-facing API utilizes different components and their network footprints to drive the migration decision. It learns to accelerate the discovery of high-quality migration plans from millions and offers recommendations with customizable trade-offs among three quality indicators: end-to-end latency of user-facing APIs representing application performance, service availability, and cloud hosting costs. Atlas continuously monitors the application even after the migration for proactive recommendations. Our evaluation shows that Atlas can achieve 21% better API performance (latency) and 11% cheaper cost with less service disruption than widely used solutions.
Ka-Ho Chow 0001, Umesh Deshpande, Veera Deenadhayalan, Sangeetha Seshadri, Ling Liu 0001
EuroSys1
2024 Demo: Visualizing the Shadows: Unveiling Data Poisoning Behaviors in Federated Learning
abstract
This demo paper examines the susceptibility of Federated Learning (FL) systems to targeted data poisoning attacks, presenting a novel system for visualizing and mitigating such threats. We simulate targeted data poisoning attacks via label flipping and analyze the impact on model performance, employing a five-component system that includes Simulation and Data Generation, Data Collection and Upload, User-friendly Interface, Analysis and Insight, and Advisory System. Observations from three demo modules: label manipulation, attack timing, and malicious attack availability, and two analysis components: utility and analytical behavior of local model updates highlight the risks to system integrity and offer insight into the resilience of FL systems. The demo is available at https://github.com/CathyXueqingZhang/DataPoisoningVis.
Ka-Ho Chow 0001, Ying Mao 0001, Mohamed Rahouti, Xiang Li 0176, Yuchen Liu 0001, Wenqi Wei 0001
ICDCS3
2024 Adaptive Deep Neural Network Inference Optimization with EENet
abstract
Well-trained deep neural networks (DNNs) treat all test samples equally during prediction. Adaptive DNN inference with early exiting leverages the observation that some test examples can be easier to predict than others. This paper presents EENet, a novel early-exiting scheduling framework for multi-exit DNN models. Instead of having every sample go through all DNN layers during prediction, EENet learns an early exit scheduler, which can intelligently terminate the inference earlier for certain predictions, which the model has high confidence of early exit. As opposed to previous early-exiting solutions with heuristics-based methods, our EENet framework optimizes an early-exiting policy to maximize model accuracy while satisfying the given per-sample average inference budget. Extensive experiments are conducted on four computer vision datasets (CIFAR-10, CIFAR-100, ImageNet, Cityscapes) and two NLP datasets (SST-2, AgNews). The results demonstrate that the adaptive inference by EENet can outperform the representative existing early exit techniques. We also perform a detailed visualization analysis of the comparison results to interpret the benefits of EENet.
Fatih Ilhan, Ka-Ho Chow 0001, Sihao Hu, Tiansheng Huang, Selim F. Tekin, Wenqi Wei 0001, Yanzhao Wu 0001, Myungjin Lee, Ramana Rao Kompella, Hugo Latapie, Gaowen Liu, Ling Liu 0001
WACV2
2024 ZipZap: Efficient Training of Language Models for Large-Scale Fraud Detection on Blockchain
abstract
Language models (LMs) have demonstrated superior performance in detecting fraudulent activities on Blockchains. Nonetheless, the sheer volume of Blockchain data results in excessive memory and computational costs when training LMs from scratch, limiting their capabilities to large-scale applications. In this paper, we present ZipZap, a framework tailored to achieve both parameter and computational efficiency when training LMs on large-scale transaction data. First, with the frequency-aware compression, an LM can be compressed down to a mere 7.5% of its initial size with an imperceptible performance dip. This technique correlates the embedding dimension of an address with its occurrence frequency in the dataset, motivated by the observation that embeddings of low-frequency addresses are insufficiently trained and thus negating the need for a uniformly large dimension for knowledge representation. Second, ZipZap accelerates the speed through the asymmetric training paradigm: It performs transaction dropping and cross-layer parameter-sharing to expedite the pre-training process, while revert to the standard training paradigm for fine-tuning to strike a balance between efficiency and efficacy, motivated by the observation that the optimization goals of pre-training and fine-tuning are inconsistent. Evaluations on real-world, large-scale datasets demonstrate that ZipZap delivers notable parameter and computational efficiency improvements for training LMs. Our implementation is available at: https://github.com/git-disl/ZipZap.
Sihao Hu, Tiansheng Huang, Ka-Ho Chow 0001, Wenqi Wei 0001, Yanzhao Wu 0001, Ling Liu 0001
WWW3
2024 Diversity-driven Privacy Protection Masks Against Unauthorized Face Recognition
abstract
Face recognition (FR) technologies have enabled many life-enriching applications but have also opened doors for potential misuse. Governments, private companies, or even individuals can scrape the web, collect facial images, and build a face database to fuel the FR system to identify human faces without their consent. This paper introduces PMask to combat such a privacy threat against unauthorized FR. It provides a holistic approach to enable privacy-preserving sharing of facial images. PMask preprocesses the facial image and hides its unique facial signature through iterative optimization with dual goals: (i) minimizing the amount of noise to ensure high image quality and (ii) minimizing the perception loss between the privacy-protected face and the original face to ensure the face is recognizable to be the same person by humans. Extensive experiments are conducted on eight representative FR models to evaluate PMask against unauthorized FR. The results validate that PMask provides much stronger protection, introduces less perceptible changes to facial images, and runs faster than state-of-the-art methods to provide privacy protection with a better user experience.
Ka-Ho Chow 0001, Sihao Hu, Tiansheng Huang, Fatih Ilhan, Wenqi Wei 0001, Ling Liu 0001
Proc. Priv. Enhancing Technol.1
2024 Hierarchical Pruning of Deep Ensembles with Focal Diversity
abstract
Deep neural network ensembles combine the wisdom of multiple deep neural networks to improve the generalizability and robustness over individual networks. It has gained increasing popularity to study and apply deep ensemble techniques in the deep learning community. Some mission-critical applications utilize a large number of deep neural networks to form deep ensembles to achieve desired accuracy and resilience, which introduces high time and space costs for ensemble execution. However, it still remains a critical challenge whether a small subset of the entire deep ensemble can achieve the same or better generalizability and how to effectively identify these small deep ensembles for improving the space and time efficiency of ensemble execution. This article presents a novel deep ensemble pruning approach, which can efficiently identify smaller deep ensembles and provide higher ensemble accuracy than the entire deep ensemble of a large number of member networks. Our hierarchical ensemble pruning approach (HQ) leverages three novel ensemble pruning techniques. First, we show that the focal ensemble diversity metrics can accurately capture the complementary capacity of the member networks of an ensemble team, which can guide ensemble pruning. Second, we design a focal ensemble diversity based hierarchical pruning approach, which will iteratively find high quality deep ensembles with low cost and high accuracy. Third, we develop a focal diversity consensus method to integrate multiple focal diversity metrics to refine ensemble pruning results, where smaller deep ensembles can be effectively identified to offer high accuracy, high robustness and high ensemble execution efficiency. Evaluated using popular benchmark datasets, we demonstrate that the proposed hierarchical ensemble pruning approach can effectively identify high quality deep ensembles with better classification generalizability while being more time and space efficient in ensemble decision making. We have released the source codes on GitHub at https://github.com/git-disl/HQ-Ensemble .
Yanzhao Wu 0001, Ka-Ho Chow 0001, Wenqi Wei 0001, Ling Liu 0001
ACM Trans. Intell. Syst. Technol.2
2024 Demystifying Data Poisoning Attacks in Distributed Learning as a Service
abstract
Data Poisoning is a dominating threat in the distributed learning-as-a-service API, where the mediator has limited control over the distributed client contributing to the joint model. Through an in-depth characterization of data poisoning risks in federated learning, this paper presents a comprehensive study towards demystifying data poisoning attacks from three perspectives.First, we formally define the targeted dirty-label data poisoning attack, which aims to cause the trained global model to only misclassify the input from a specific victim class with a designated malicious behavior. Then, we demonstrate theoretical statistical robustness in the eigenvalues of the covariance in the gradient update shared from the client to server when under the data poisoning attack.Second, we study the impact of attack timing and identify the most detrimental attack entry point during the federated training.Last, we examine several existing defenses against data poisoning in addition to the robust statistic detection. Through formal analysis and extensive empirical evidence, we investigate under what conditions the statistical robustness of data poisoning can serve as the forensic evidence for attack mitigation in federated-learning-as-a-service, at what attack timing the attack is most detrimental, and how the attack reacts in the presence of the existing defenses.
Wenqi Wei 0001, Ka-Ho Chow 0001, Yanzhao Wu 0001, Ling Liu 0001
IEEE Trans. Serv. Comput.2
2023 STDLens: Model Hijacking-Resilient Federated Learning for Object Detection
abstract
Federated Learning (FL) has been gaining popularity as a collaborative learning framework to train deep learning-based object detection models over a distributed population of clients. Despite its advantages, FL is vulnerable to model hijacking. The attacker can control how the object detection system should misbehave by implanting Trojaned gradients using only a small number of compromised clients in the collaborative learning process. This paper introduces STDLens, a principled approach to safeguarding FL against such attacks. We first investigate existing mitigation mechanisms and analyze their failures caused by the inherent errors in spatial clustering analysis on gradients. Based on the insights, we introduce a three-tier forensic framework to identify and expel Trojaned gradients and reclaim the performance over the course of FL. We consider three types of adaptive attacks and demonstrate the robustness of STDLens against advanced adversaries. Extensive experiments show that STDLens can protect FL against different model hijacking attacks and outperform existing methods in identifying and removing Trojaned gradients with significantly higher precision and much lower false-positive rates. The source code is available at https://github.com/git-disl/STDLens.
Ka-Ho Chow 0001, Ling Liu 0001, Wenqi Wei 0001, Fatih Ilhan, Yanzhao Wu 0001
CVPR1
2023 Exploring Model Learning Heterogeneity for Boosting Ensemble Robustness
abstract
Deep neural network ensembles hold the potential of improving generalization performance for complex learning tasks. This paper presents formal analysis and empirical evaluation to show that heterogeneous deep ensembles with high ensemble diversity can effectively leverage model learning heterogeneity to boost ensemble robustness. We first show that heterogeneous DNN models trained for solving the same learning problem, e.g., object detection, can significantly strengthen the mean average precision (mAP) through our weighted bounding box ensemble consensus method. Second, we further compose ensembles of heterogeneous models for solving different learning problems, e.g., object detection and semantic segmentation, by introducing the connected component labeling (CCL) based alignment. We show that this two-tier heterogeneity driven ensemble construction method can compose an ensemble team that promotes high ensemble diversity and low negative correlation among member models of the ensemble, strengthening ensemble robustness against both negative examples and adversarial attacks. Third, we provide a formal analysis of the ensemble robustness in terms of negative correlation. Extensive experiments validate the enhanced robustness of heterogeneous ensembles in both benign and adversarial settings. The appendix and source codes are available on GitHub at https://github.com/git-disl/HeteRobust.
Yanzhao Wu 0001, Ka-Ho Chow 0001, Wenqi Wei 0001, Ling Liu 0001
ICDM2
2023 Model Cloaking against Gradient Leakage
abstract
Gradient leakage attacks are dominating privacy threats in federated learning, despite the default privacy that training data resides locally at the clients. Differential privacy has been the de facto standard for privacy protection and is deployed in federated learning to mitigate privacy risks. However, much existing literature points out that differential privacy fails to defend against gradient leakage. The paper presents ModelCloak, a principled approach based on differential privacy noise, aiming for safe-sharing client local model updates. The paper is organized into three major components. First, we introduce the gradient leakage robustness trade-off, in search of the best balance between accuracy and leakage prevention. The trade-off relation is developed based on the behavior of gradient leakage attacks throughout the federated training process. Second, we demonstrate that a proper amount of differential privacy noise can offer the best accuracy performance within the privacy requirement under a fixed differential privacy noise setting. Third, we propose dynamic differential privacy noise and show that the privacy-utility trade-off can be further optimized with dynamic model perturbation, ensuring privacy protection, competitive accuracy, and leakage attack prevention simultaneously.
Wenqi Wei 0001, Ka-Ho Chow 0001, Fatih Ilhan, Yanzhao Wu 0001, Ling Liu 0001
ICDM2
2023 Lockdown: Backdoor Defense for Federated Learning with Isolated Subspace Training
abstract
Federated learning (FL) is vulnerable to backdoor attacks due to its distributed computing nature. Existing defense solution usually requires larger amount of computation in either the training or testing phase, which limits their practicality in the resource-constrain scenarios. A more practical defense, i.e., neural network (NN) pruning based defense has been proposed in centralized backdoor setting. However, our empirical study shows that traditional pruning-based solution suffers \textit{poison-coupling} effect in FL, which significantly degrades the defense performance.This paper presents Lockdown, an isolated subspace training method to mitigate the poison-coupling effect. Lockdown follows three key procedures. First, it modifies the training protocol by isolating the training subspaces for different clients. Second, it utilizes randomness in initializing isolated subspacess, and performs subspace pruning and subspace recovery to segregate the subspaces between malicious and benign clients. Third, it introduces quorum consensus to cure the global model by purging malicious/dummy parameters. Empirical results show that Lockdown achieves \textit{superior} and \textit{consistent} defense performance compared to existing representative approaches against backdoor attacks. Another value-added property of Lockdown is the communication-efficiency and model complexity reduction, which are both critical for resource-constrain FL scenario. Our code is available at \url{https://github.com/git-disl/Lockdown}.
Tiansheng Huang, Sihao Hu, Ka-Ho Chow 0001, Fatih Ilhan, Selim F. Tekin, Ling Liu 0001
NeurIPS3
2023 Implicit Multimodal Crowdsourcing for Joint RF and Geomagnetic Fingerprinting
abstract
In fingerprint-based indoor localization, fusing radio frequency (RF) and geomagnetic signals has been shown to achieve promising results. To efficiently collect fingerprints, implicit crowdsourcing can be used, where signals sampled by pedestrians are automatically labeled with their locations on a map. Previous work on crowdsourced fingerprinting is often based on a single signal, which is susceptible to signal bias and labeling error. We study, for the first time, implicit multimodal crowdsourcing for joint RF and geomagnetic fingerprinting. The scheme, termed UbiFin, exploits the spatial correlation among RF, geomagnetic, and motion signals to mitigate the impact of sensor noise, leading to highly accurate and robust fingerprinting without the need for any explicit manual intervention. Using clustering and dynamic programming, UbiFin correlates spatially different signals and filters effectively mislabeled signals. We conduct extensive experiments on our campus and a large multi-story shopping mall. Efficient and simple to implement, UbiFin outperforms other state-of-the-art crowdsourcing schemes to construct RF and geomagnetic fingerprints in terms of accuracy and robustness (cutting fingerprint error by 40 percent in general).
Jiajie Tan, Ka-Ho Chow 0001, Shueng-Han Gary Chan
IEEE Trans. Mob. Comput.3
2023 Securing Distributed SGD Against Gradient Leakage Threats
abstract
This paper presents a holistic approach to gradient leakage resilient distributed Stochastic Gradient Descent (SGD).First, we analyze two types of strategies for privacy-enhanced federated learning: (i) gradient pruning with random selection or low-rank filtering and (ii) gradient perturbation with additive random noise or differential privacy noise. We analyze the inherent limitations of these approaches and their underlying impact on privacy guarantee, model accuracy, and attack resilience.Next, we present a gradient leakage resilient approach to securing distributed SGD in federated learning, with differential privacy controlled noise as the tool. Unlike conventional methods with the per-client federated noise injection and fixed noise parameter strategy, our approach keeps track of the trend of per-example gradient updates. It makes adaptive noise injection closely aligned throughout the federated model training.Finally, we provide an empirical privacy analysis on the privacy guarantee, model utility, and attack resilience of the proposed approach. Extensive evaluation using five benchmark datasets demonstrates that our gradient leakage resilient approach can outperform the state-of-the-art methods with competitive accuracy performance, strong differential privacy guarantee, and high resilience against gradient leakage attacks.
Wenqi Wei 0001, Ling Liu 0001, Jingya Zhou, Ka-Ho Chow 0001, Yanzhao Wu 0001
IEEE Trans. Parallel Distributed Syst.4
2022 DeepRest: deep resource estimation for interactive microservices
abstract
Interactive microservices expose API endpoints to be invoked by users. For such applications, precisely estimating the resources required to serve specific API traffic is challenging. This is because an API request can interact with different components and consume different resources for each component. The notion of API traffic is vital to application owners since the API endpoints often reflect business logic, e.g., a customer transaction. The existing systems that simply rely on historical resource utilization are not API-aware and thus cannot estimate the resource requirement accurately. This paper presents DeepRest, a deep learning-driven resource estimation system. DeepRest formulates resource estimation as a function of API traffic and learns the causality between user interactions and resource utilization directly in a production environment. Our evaluation shows that DeepRest can estimate resource requirements with over 90% accuracy, even if the API traffic to be estimated has never been observed (e.g., 3× more users than ever or unseen traffic shape). We further apply resource estimation for application sanity checks. DeepRest identifies system anomalies by verifying whether the resource utilization is justifiable by how the application is being used. It can successfully identify two major cyber threats: ransomware and cryptojacking attacks.
Ka-Ho Chow 0001, Umesh Deshpande, Sangeetha Seshadri, Ling Liu 0001
EuroSys1
2022 Boosting Object Detection Ensembles with Error Diversity
abstract
Object detection has played a pivotal role in numerous mission-critical applications. This paper presents a focal error diversity framework, called EDI, for strengthening the robustness of object detection ensembles under benign and adversarial scenarios. We introduce an ensemble pruning method for object detection using a novel focal error diversity measure as the robustness synergy indicator. Given a base model pool, it recommends top sub-ensembles with a smaller ensemble size yet achieving equivalent or even better mAP performance than using all available object detection models as a large ensemble. This is made possible by our negative sampling methods for object detection to capture the degree of negative correlations and the focal error diversity to measure the failure independence of component detection models in an ensemble. Extensive experiments on three object detection benchmark datasets validate that EDI effectively selects space-time efficient object detection ensembles with high mAP performance.
Ka-Ho Chow 0001, Ling Liu 0001
ICDM1
2022 An Adversarial Approach to Protocol Analysis and Selection in Local Differential Privacy
abstract
Local Differential Privacy (LDP) is a popular standard for privacy-preserving data collection. Numerous LDP protocols have been proposed in the literature which differ in how they provide higher utility in different settings. Yet, few have engaged in analyzing the privacy relationships of these protocols under varying settings, and consequently, it is non-trivial to select which LDP protocol is best to use in a newly emerging application. In this paper, we present an adversarial approach to protocol analysis and selection and make three original contributions. First, we introduce a Bayesian adversary to analyze the privacy relationships of LDP protocols under varying settings. We show that different protocols have substantially different responses to the attack effectiveness of the Bayesian adversary, measured in terms of Adversarial Success Rate (ASR). Second, we provide a formal and empirical analysis on a set of privacy and utility-critical factors, including encoding parameters, privacy budget, data domain, adversarial knowledge, and statistical distribution. We show that different settings of these factors have significant effects on the ASRs of LDP protocols, and no protocol provides consistently low ASR across all settings. Third, we design and develop LDPLens, a prototype implementation of our proposed framework. Given a data collection scenario with various factors and constraints, LDPLens enables optimized selection of a desirable LDP protocol for the given scenario. We evaluate the effectiveness of LDPLens using three case studies with real-world datasets. Results show that LDPLens recommends a different protocol in each case study, and the protocol recommended by LDPLens can yield up to 1.5–2 fold reduction in utility loss, ASR or privacy budget compared to a randomly selected protocol.
Mehmet Emre Gursoy, Ling Liu 0001, Ka-Ho Chow 0001, Stacey Truex, Wenqi Wei 0001
IEEE Trans. Inf. Forensics Secur.3
2021 Transparent Network Memory Storage for Efficient Container Execution in Big Data Clouds
abstract
This paper presents a transparent Container Network Memory storage device, coined as CNetMem, aiming to address the open problem of unpredictable performance degradation of containers when the working set of an application no longer fits in container memory. First, CNetMem will enable application tenants running in a container to park their working set memory/file to a faster network memory storage by organizing a group of remote memory nodes as remote memory donors. This allows CNetMem to take advantage of remote idle memory on a cluster before resorting to a slow local I/O subsystem like local disk without any modification of host OS or application. Second, CNetMem provides a hybrid batching technique to remove or alleviate performance bottlenecks in the I/O performance critical path for remote memory read/write with replication or disk backup for fault tolerance. Third, CNetMem introduces a rank-based node selection algorithm to find the optimal node for placing remote memory blocks across cluster. This helps CNetMem to reduce the performance impact due to remote memory eviction. Extensive experiments are conducted on three big data applications and four machine learning workloads. The results show that CNetMem achieves up to 172× throughput improvements compared to vanilla Linux and up to 5.9× completion time improvements over existing approaches in big data and ML workload.
Juhyun Bae, Ling Liu 0001, Ka-Ho Chow 0001, Yanzhao Wu 0001, Gong Su, Arun Iyengar
IEEE BigData3
2021 Boosting Ensemble Accuracy by Revisiting Ensemble Diversity Metrics
abstract
Neural network ensembles are gaining popularity by harnessing the complementary wisdom of multiple base models. Ensemble teams with high diversity promote high failure independence, which is effective for boosting the overall ensemble accuracy. This paper provides an in-depth study on how to design and compute ensemble diversity, which can capture the complementary decision capacity of ensemble member models. We make three original contri-butions. First, we revisit the ensemble diversity metrics in the literature and analyze the inherent problems of poor correlation between ensemble diversity and ensemble ac-curacy, which leads to the low quality ensemble selection using such diversity metrics. Second, instead of computing diversity scores for ensemble teams of different sizes using the same criteria, we introduce focal model based ensemble diversity metrics, coined as FQ-diversity metrics. Our new metrics significantly improve the intrinsic correlation between high ensemble diversity and high ensemble accuracy. Third, we introduce a diversity fusion method, coined as the EQ-diversity metric, by integrating the top three most representative FQ-diversity metrics. Comprehensive experiments on two benchmark datasets (CIFAR-10 and ImageNet) show that our FQ and EQ diversity metrics are effective for selecting high diversity ensemble teams to boost overall ensemble accuracy.
Yanzhao Wu 0001, Ling Liu 0001, Zhongwei Xie, Ka-Ho Chow 0001, Wenqi Wei 0001
CVPR4
2021 Robust Object Detection Fusion Against Deception
abstract
Deep neural network (DNN) based object detection has become an integral part of numerous cyber-physical systems, perceiving physical environments and responding proactively to real-time events. Recent studies reveal that well-trained multi-task learners like DNN-based object detectors perform poorly in the presence of deception. This paper presents FUSE, a deception-resilient detection fusion approach with three novel contributions. First, we develop diversity-enhanced fusion teaming mechanisms, including diversity-enhanced joint training algorithms, for producing high diversity fusion detectors. Second, we introduce a three-tier detection fusion framework and a graph partitioning algorithm to construct fusion-verified detection outputs through three mutually reinforcing components: objectness fusion, bounding box fusion, and classification fusion. Third but not least, we provide a formal analysis of robustness enhancement by FUSE-protected systems. Extensive experiments are conducted on eleven detectors from three families of detection algorithms on two benchmark datasets. We show that FUSE guarantees strong robustness in mitigating the state-of-the-art deception attacks, including adversarial patches - a form of physical attacks using confined visual distortion.
Ka-Ho Chow 0001, Ling Liu 0001
KDD1
2021 SRA: Smart Recovery Advisor for Cyber Attacks
abstract
Continuous Data Protection (CDP) is becoming instrumental in recovering applications from crypto-ransomware attacks. It enables fine-grained recovery through journaling, allowing the applications (its volumes) to recover to any previous state. While zero data loss can be achieved during recovery with CDP, the timestamp of the desired restore point, i.e., the one just prior to the attack, needs to be provided to reconstruct the volume. Such information is often unavailable in practice, and system administrators can only adopt a trial-and-error strategy to narrow down the time range of desired restore points by making multiple time-consuming recovery attempts. The recovery systems offer little guidance in pointing to the restore points containing a valid application state and reducing data loss. To address this problem, we equip the CDP-based recovery with machine intelligence. This demonstration showcases Smart Recovery Advisor (SRA), which offers interpretable, data-driven, and feedback-aware restore point recommendations that reduce the number of recovery attempts while minimizing data loss.
Ka-Ho Chow 0001, Umesh Deshpande, Sangeetha Seshadri, Ling Liu 0001
SIGMOD Conference1
2020 Understanding Object Detection Through an Adversarial Lens
Ka-Ho Chow 0001, Ling Liu 0001, Mehmet Emre Gursoy, Stacey Truex, Wenqi Wei 0001, Yanzhao Wu 0001
ESORICS (2)1
2020 A Framework for Evaluating Client Privacy Leakages in Federated Learning
Wenqi Wei 0001, Ling Liu 0001, Margaret L. Loper, Ka-Ho Chow 0001, Mehmet Emre Gursoy, Stacey Truex, Yanzhao Wu 0001
ESORICS (1)4
2019 Denoising and Verification Cross-Layer Ensemble Against Black-box Adversarial Attacks
abstract
Deep neural networks (DNNs) have demonstrated impressive performance on many challenging machine learning tasks. However, DNNs are vulnerable to adversarial inputs generated by adding maliciously crafted perturbations to the benign inputs. As a growing number of attacks have been reported to generate adversarial inputs of varying sophistication, the defense-attack arms race has been accelerated. In this paper, we present MODEF, a cross-layer model diversity ensemble framework. MODEF intelligently combines unsupervised model denoising ensemble with supervised model verification ensemble by quantifying model diversity, aiming to boost the robustness of the target model against adversarial examples. Evaluated using eleven representative attacks on popular benchmark datasets, we show that MODEF achieves remarkable defense success rates, compared with existing defense methods, and provides a superior capability of repairing adversarial inputs and making correct predictions with high accuracy in the presence of black-box attacks.
Ka-Ho Chow 0001, Wenqi Wei 0001, Yanzhao Wu 0001, Ling Liu 0001
IEEE BigData1
2019 Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural Networks
abstract
Learning Rate (LR) is an important hyper-parameter to tune for effective training of deep neural networks (DNNs). Even for the baseline of a constant learning rate, it is non-trivial to choose a good constant value for training a DNN. Dynamic learning rates involve multi-step tuning of LR values at various stages of the training process and offer high accuracy and fast convergence. However, they are much harder to tune. In this paper, we present a comprehensive study of 13 learning rate functions and their associated LR policies by examining their range parameters, step parameters, and value update parameters. We propose a set of metrics for evaluating and selecting LR policies, including the classification confidence, variance, cost, and robustness, and implement them in LRBench, an LR benchmarking system. LRBench can assist end-users and DNN developers to select good LR policies and avoid bad LR policies for training their DNNs. We tested LRBench on Caffe, an open source deep learning framework, to showcase the tuning optimization of LR policies. Evaluated through extensive experiments, we attempt to demystify the tuning of LR policies by identifying good LR policies with effective LR value ranges and step sizes for LR update schedules.
Yanzhao Wu 0001, Ling Liu 0001, Juhyun Bae, Ka-Ho Chow 0001, Arun Iyengar, Calton Pu, Wenqi Wei 0001, Lei Yu 0002, Qi Zhang 0009
IEEE BigData4
2019 Deep Neural Network Ensembles Against Deception: Ensemble Diversity, Accuracy and Robustness
abstract
Ensemble learning is a methodology that integrates multiple DNN learners for improving prediction performance of individual learners. Diversity is greater when the errors of the ensemble prediction is more uniformly distributed. Greater diversity is highly correlated with the increase in ensemble accuracy. Another attractive property of diversity optimized ensemble learning is its robustness against deception: an adversarial perturbation attack can mislead one DNN model to misclassify but may not fool other ensemble DNN members consistently. In this paper we first give an overview of the concept of ensemble diversity and examine the three types of ensemble diversity in the context of DNN classifiers. We then describe a set of ensemble diversity measures, a suite of algorithms for creating diversity ensembles and for performing ensemble consensus (voted or learned) for generating high accuracy ensemble output by strategically combining outputs of individual members. This paper concludes with a discussion on a set of open issues in quantifying ensemble diversity for robust deep learning.
Ling Liu 0001, Wenqi Wei 0001, Ka-Ho Chow 0001, Margaret L. Loper, Mehmet Emre Gursoy, Stacey Truex, Yanzhao Wu 0001
MASS3
2019 Efficient Locality Classification for Indoor Fingerprint-Based Systems
abstract
Locality classification is an important component to enable location-based services. It entails two sequential queries: 1) whether a target is within the site or not, i.e., inside/outside region decision, and 2) if so, which area in the region the target is located, i.e., area classification. Locality classification is hence more coarse-grained and efficient as compared with pinpointing the exact target location in the region. The classification problem is challenging, because fingerprints may not exist outside the region for training. Furthermore, the target may sample an incomplete RSSI vector due to, say, random signal noise, momentary occlusion, or scanning duration. The algorithm also has to be computationally efficient. We propose INOA, a scalable and practical locality classification overcoming the above challenges. INOA may serve as a plug-in before any fingerprint-based localization, and can be incrementally extended to cover new areas or regions for large-scale deployment. Its preprocessor cherry-picks only those discriminating access points, which greatly enhances computational efficiency and accuracy. By formulating a “one-class” classifier using ensemble learning, INOA accurately decides whether the target is within the region or not. Extensive experimental trials in different sites validate the high efficiency and accuracy of INOA, without the need of full RSSI vectors collected at the target.
Ka-Ho Chow 0001, Suining He, Jiajie Tan, Shueng-Han Gary Chan
IEEE Trans. Mob. Comput.1