VLDB 2026 Research / reviewers in the wild / expert
Yanzhao Wu 0001
dblp:61/9620-1
· DBLP profile ↗
32ranked-venue papers
8as first author
23since 2021 · last 2025
0000-0001-8761-5486ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 14 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Systems, architecture and hardware · 5 · 3 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Security and privacy · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Diversity-Optimized Deep Ensemble Approach for Accurate Plant Leaf Disease Detection
Sai Nath Chowdary Medikonduru, Hongpeng Jin, Yanzhao Wu 0001 |
IEEE Big Data | 3 |
| 2025 | EFT-LR: Benchmarking Learning Rate Policies in Parameter-Efficient Large Language Model Fine-tuningabstractLarge Language Models (LLMs) have achieved extensive impacts across various real-world data mining applications. Given the extremely high cost of training or fine-tuning LLMs, parameter-efficient fine-tuning (e.g., LoRA) has emerged as a popular and practical approach for adapting pre-trained general-purpose LLMs to specific downstream tasks. Among the various hyperparameters involved in parameter-efficient fine-tuning of LLMs, the learning rate (LR) plays a crucial role in determining the overall performance. However, it lacks a systematic benchmark framework to explore and understand how different LR policies influence the effectiveness of parameter-efficient LLM fine-tuning, which makes it challenging to select an optimal LR policy. To address this critical research gap, this paper introduces a systematic benchmark, EFT-LR, for assessing and selecting LR policies for effective parameter-efficient fine-tuning of LLMs. We first present a collection of seven popular LR policies spanning three major categories in the literature. We then perform parameter-efficient fine-tuning of LLMs using these LR policies and assess fine-tuned LLMs on eight downstream tasks. Our empirical analysis using EFT-LR provides an in-depth investigation of the impacts of different LR policies on parameter-efficient LLM fine-tuning, offering practical guidelines for practitioners. We provide the source code at https://github.com/mlsysx/EFT-LR. Md. Tasnim Jawad, Yanzhao Wu 0001 |
CIKM | 2 |
| 2025 | Jailbreaking Large Vision Language Models in Intelligent Transportation SystemsabstractLarge Vision Language Models (LVLMs) demon-strate strong capabilities in multimodal reasoning and many real-world applications, such as visual question answering. However, LVLMs are highly vulnerable to jailbreaking attacks. This paper systematically analyzes the vulnerabilities of LVLMs integrated in Intelligent Transportation Systems (ITS) under carefully crafted jailbreaking attacks. First, we carefully construct a dataset with harmful queries relevant to transportation, following OpenAI’s prohibited categories to which the LVLMs should not respond. Second, we introduce a novel jailbreaking attack that exploits the vulnerabilities of LVLMs through image typography manipulation and multi-turn prompting. Third, we propose a multi-layered response filtering defense technique to prevent the model from generating inappropriate responses. We perform extensive experiments with the proposed attack and defense on the state-of-the-art LVLMs (both open-source and closed-source). To evaluate the attack method and defense technique, we use GPT-4’s judgment to determine the toxicity score of the generated responses, as well as manual verification. Further, we compare our proposed jailbreaking method with existing jailbreaking techniques and highlight severe security risks involved with jailbreaking attacks with image typography manipulation and multi-turn prompting in the LVLMs integrated in ITS. Badhan Chandra Das, Md. Tasnim Jawad, Md. Jueal Mia, M. Hadi Amini, Yanzhao Wu 0001 |
ICMLA | 5 |
| 2025 | CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge CollaborationabstractLarge Language Models (LLMs) exhibit remarkable human-like predictive capabilities. However, it is challenging to deploy LLMs to provide efficient and adaptive inference services at the edge. This paper proposes a novel Cloud-Edge Collaboration framework for LLMs (CE-CoLLM) to tackle these challenges. First, we identify the transmission of LLM contextual data between the cloud and edge as a key performance bottleneck, which introduces substantial communication overhead that dominates overall inference latency and makes naïve cloudedge collaboration for LLMs inefficient. Second, we introduce a suite of novel techniques, including a latency-aware early exit mechanism and efficient cloud context management, into CECoLLM, which collectively reduce communication overhead and preserve LLM inference accuracy. Third, we design two adaptive inference modes to accommodate diverse edge environments: (1) a low-latency standalone edge inference mode that enables reliable edge-side independent LLM inference even under unstable network conditions, and (2) a high-accuracy cloud-edge collaborative inference mode that adaptively leverages cloud resources to enhance prediction accuracy. Extensive experiments on multiple benchmark datasets demonstrate that CE-CoLLM reduces overall inference time by up to 13.81% and offloads over 84.53% of the computational workload from the cloud to the edge, compared to conventional cloud-based LLM deployment, without sacrificing prediction accuracy. The code is provided on GitHub at https://github.com/mlsysx/CE-CoLLM. Hongpeng Jin, Yanzhao Wu 0001 |
ICWS | 2 |
| 2025 | LATTICE: Efficient In-Memory DNN Model VersioningabstractDNN model versions are used for various tasks such as fine-tuning for downstream tasks, explainability, and debugging. Numerous checkpointing solutions exist that can be adapted to persist intermediate versions of a model, as it is being trained, at different storage locations. Additionally, version management tools allow us to log, visualize, compare, and query metadata related to ML, tracking changes made to previously built models. However, the version creation process of existing methods incurs high runtime and storage overheads. In this paper, we introduce LATTICE, a low-latency, direct persistence-based DNN versioning library for Non-Volatile Memory (NVM) expansion devices. LATTICE minimizes stalls during model versioning and reduces end-to-end versioning time by reorganizing the version creation workflow, streamlining memory allocation and deallocation for efficient snapshot creation, and leveraging multi-threaded parallelism. We also develop a user-friendly versioning API that transparently implements direct persistence. Our comprehensive evaluation with diverse DNN models shows that LATTICE can reduce persistence time by as much as 99.99%, decrease end-to-end versioning time by up to 72%, reduce versioning stalls by up to 35%, and increase versioning frequency by 0.2×-3.84× compared to state-of-the-art solutions. LATTICE also reduces space utilization for different workloads. The space savings are from 23.8% to 43.2% for workloads where model layers are progressively frozen and from 84.8% to 98.9% for fine-tuning workloads where only the last layers are tuned. Manoj Pravakar Saha, Ashikee Ghosh, Raju Rangaswami, Yanzhao Wu 0001, Janki Bhimani |
SYSTOR | 4 |
| 2024 | Individual Fairness with Group Awareness Under Uncertainty
Zichong Wang, Jocelyn Dzuong, Xiaoyong Yuan, Zhong Chen 0003, Yanzhao Wu 0001, Wenbin Zhang 0002 |
ECML/PKDD (5) | 5 |
| 2024 | Adaptive Deep Neural Network Inference Optimization with EENetabstractWell-trained deep neural networks (DNNs) treat all test samples equally during prediction. Adaptive DNN inference with early exiting leverages the observation that some test examples can be easier to predict than others. This paper presents EENet, a novel early-exiting scheduling framework for multi-exit DNN models. Instead of having every sample go through all DNN layers during prediction, EENet learns an early exit scheduler, which can intelligently terminate the inference earlier for certain predictions, which the model has high confidence of early exit. As opposed to previous early-exiting solutions with heuristics-based methods, our EENet framework optimizes an early-exiting policy to maximize model accuracy while satisfying the given per-sample average inference budget. Extensive experiments are conducted on four computer vision datasets (CIFAR-10, CIFAR-100, ImageNet, Cityscapes) and two NLP datasets (SST-2, AgNews). The results demonstrate that the adaptive inference by EENet can outperform the representative existing early exit techniques. We also perform a detailed visualization analysis of the comparison results to interpret the benefits of EENet. Fatih Ilhan, Ka-Ho Chow 0001, Sihao Hu, Tiansheng Huang, Selim F. Tekin, Wenqi Wei 0001, Yanzhao Wu 0001, Myungjin Lee, Ramana Rao Kompella, Hugo Latapie, Gaowen Liu, Ling Liu 0001 |
WACV | 7 |
| 2024 | ZipZap: Efficient Training of Language Models for Large-Scale Fraud Detection on BlockchainabstractLanguage models (LMs) have demonstrated superior performance in detecting fraudulent activities on Blockchains. Nonetheless, the sheer volume of Blockchain data results in excessive memory and computational costs when training LMs from scratch, limiting their capabilities to large-scale applications. In this paper, we present ZipZap, a framework tailored to achieve both parameter and computational efficiency when training LMs on large-scale transaction data. First, with the frequency-aware compression, an LM can be compressed down to a mere 7.5% of its initial size with an imperceptible performance dip. This technique correlates the embedding dimension of an address with its occurrence frequency in the dataset, motivated by the observation that embeddings of low-frequency addresses are insufficiently trained and thus negating the need for a uniformly large dimension for knowledge representation. Second, ZipZap accelerates the speed through the asymmetric training paradigm: It performs transaction dropping and cross-layer parameter-sharing to expedite the pre-training process, while revert to the standard training paradigm for fine-tuning to strike a balance between efficiency and efficacy, motivated by the observation that the optimization goals of pre-training and fine-tuning are inconsistent. Evaluations on real-world, large-scale datasets demonstrate that ZipZap delivers notable parameter and computational efficiency improvements for training LMs. Our implementation is available at: https://github.com/git-disl/ZipZap. Sihao Hu, Tiansheng Huang, Ka-Ho Chow 0001, Wenqi Wei 0001, Yanzhao Wu 0001, Ling Liu 0001 |
WWW | 5 |
| 2024 | Hierarchical Pruning of Deep Ensembles with Focal DiversityabstractDeep neural network ensembles combine the wisdom of multiple deep neural networks to improve the generalizability and robustness over individual networks. It has gained increasing popularity to study and apply deep ensemble techniques in the deep learning community. Some mission-critical applications utilize a large number of deep neural networks to form deep ensembles to achieve desired accuracy and resilience, which introduces high time and space costs for ensemble execution. However, it still remains a critical challenge whether a small subset of the entire deep ensemble can achieve the same or better generalizability and how to effectively identify these small deep ensembles for improving the space and time efficiency of ensemble execution. This article presents a novel deep ensemble pruning approach, which can efficiently identify smaller deep ensembles and provide higher ensemble accuracy than the entire deep ensemble of a large number of member networks. Our hierarchical ensemble pruning approach (HQ) leverages three novel ensemble pruning techniques. First, we show that the focal ensemble diversity metrics can accurately capture the complementary capacity of the member networks of an ensemble team, which can guide ensemble pruning. Second, we design a focal ensemble diversity based hierarchical pruning approach, which will iteratively find high quality deep ensembles with low cost and high accuracy. Third, we develop a focal diversity consensus method to integrate multiple focal diversity metrics to refine ensemble pruning results, where smaller deep ensembles can be effectively identified to offer high accuracy, high robustness and high ensemble execution efficiency. Evaluated using popular benchmark datasets, we demonstrate that the proposed hierarchical ensemble pruning approach can effectively identify high quality deep ensembles with better classification generalizability while being more time and space efficient in ensemble decision making. We have released the source codes on GitHub at https://github.com/git-disl/HQ-Ensemble . Yanzhao Wu 0001, Ka-Ho Chow 0001, Wenqi Wei 0001, Ling Liu 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2024 | Demystifying Data Poisoning Attacks in Distributed Learning as a ServiceabstractData Poisoning is a dominating threat in the distributed learning-as-a-service API, where the mediator has limited control over the distributed client contributing to the joint model. Through an in-depth characterization of data poisoning risks in federated learning, this paper presents a comprehensive study towards demystifying data poisoning attacks from three perspectives.First, we formally define the targeted dirty-label data poisoning attack, which aims to cause the trained global model to only misclassify the input from a specific victim class with a designated malicious behavior. Then, we demonstrate theoretical statistical robustness in the eigenvalues of the covariance in the gradient update shared from the client to server when under the data poisoning attack.Second, we study the impact of attack timing and identify the most detrimental attack entry point during the federated training.Last, we examine several existing defenses against data poisoning in addition to the robust statistic detection. Through formal analysis and extensive empirical evidence, we investigate under what conditions the statistical robustness of data poisoning can serve as the forensic evidence for attack mitigation in federated-learning-as-a-service, at what attack timing the attack is most detrimental, and how the attack reacts in the presence of the existing defenses. Wenqi Wei 0001, Ka-Ho Chow 0001, Yanzhao Wu 0001, Ling Liu 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Privacy Risks Analysis and Mitigation in Federated Learning for Medical ImagesabstractFederated learning (FL) is gaining increasing popularity in the medical domain for analyzing medical images, which is considered an effective technique to safeguard sensitive patient data and comply with privacy regulations. However, several recent studies have revealed that the default settings of FL may leak private training data under privacy attacks. Thus, it is still unclear whether and to what extent such privacy risks of FL exist in the medical domain, and if so, “how to mitigate such risks?”. In this paper, first, we propose a holistic framework for Medical data Privacy risk analysis and mitigation in Federated Learning (MedPFL) to analyze privacy risks and develop effective mitigation strategies in FL for protecting private medical data. Second, we demonstrate the substantial privacy risks of using FL to process medical images, where adversaries can easily perform privacy attacks to reconstruct private medical images accurately. Third, we show that the defense approach of adding random noises may not always work effectively to protect medical images against privacy attacks in FL, which poses unique and pressing challenges associated with medical data for privacy protection. Badhan Chandra Das, M. Hadi Amini, Yanzhao Wu 0001 |
BIBM | 3 |
| 2023 | STDLens: Model Hijacking-Resilient Federated Learning for Object DetectionabstractFederated Learning (FL) has been gaining popularity as a collaborative learning framework to train deep learning-based object detection models over a distributed population of clients. Despite its advantages, FL is vulnerable to model hijacking. The attacker can control how the object detection system should misbehave by implanting Trojaned gradients using only a small number of compromised clients in the collaborative learning process. This paper introduces STDLens, a principled approach to safeguarding FL against such attacks. We first investigate existing mitigation mechanisms and analyze their failures caused by the inherent errors in spatial clustering analysis on gradients. Based on the insights, we introduce a three-tier forensic framework to identify and expel Trojaned gradients and reclaim the performance over the course of FL. We consider three types of adaptive attacks and demonstrate the robustness of STDLens against advanced adversaries. Extensive experiments show that STDLens can protect FL against different model hijacking attacks and outperform existing methods in identifying and removing Trojaned gradients with significantly higher precision and much lower false-positive rates. The source code is available at https://github.com/git-disl/STDLens. Ka-Ho Chow 0001, Ling Liu 0001, Wenqi Wei 0001, Fatih Ilhan, Yanzhao Wu 0001 |
CVPR | 5 |
| 2023 | Exploring Model Learning Heterogeneity for Boosting Ensemble RobustnessabstractDeep neural network ensembles hold the potential of improving generalization performance for complex learning tasks. This paper presents formal analysis and empirical evaluation to show that heterogeneous deep ensembles with high ensemble diversity can effectively leverage model learning heterogeneity to boost ensemble robustness. We first show that heterogeneous DNN models trained for solving the same learning problem, e.g., object detection, can significantly strengthen the mean average precision (mAP) through our weighted bounding box ensemble consensus method. Second, we further compose ensembles of heterogeneous models for solving different learning problems, e.g., object detection and semantic segmentation, by introducing the connected component labeling (CCL) based alignment. We show that this two-tier heterogeneity driven ensemble construction method can compose an ensemble team that promotes high ensemble diversity and low negative correlation among member models of the ensemble, strengthening ensemble robustness against both negative examples and adversarial attacks. Third, we provide a formal analysis of the ensemble robustness in terms of negative correlation. Extensive experiments validate the enhanced robustness of heterogeneous ensembles in both benign and adversarial settings. The appendix and source codes are available on GitHub at https://github.com/git-disl/HeteRobust. Yanzhao Wu 0001, Ka-Ho Chow 0001, Wenqi Wei 0001, Ling Liu 0001 |
ICDM | 1 |
| 2023 | Model Cloaking against Gradient LeakageabstractGradient leakage attacks are dominating privacy threats in federated learning, despite the default privacy that training data resides locally at the clients. Differential privacy has been the de facto standard for privacy protection and is deployed in federated learning to mitigate privacy risks. However, much existing literature points out that differential privacy fails to defend against gradient leakage. The paper presents ModelCloak, a principled approach based on differential privacy noise, aiming for safe-sharing client local model updates. The paper is organized into three major components. First, we introduce the gradient leakage robustness trade-off, in search of the best balance between accuracy and leakage prevention. The trade-off relation is developed based on the behavior of gradient leakage attacks throughout the federated training process. Second, we demonstrate that a proper amount of differential privacy noise can offer the best accuracy performance within the privacy requirement under a fixed differential privacy noise setting. Third, we propose dynamic differential privacy noise and show that the privacy-utility trade-off can be further optimized with dynamic model perturbation, ensuring privacy protection, competitive accuracy, and leakage attack prevention simultaneously. Wenqi Wei 0001, Ka-Ho Chow 0001, Fatih Ilhan, Yanzhao Wu 0001, Ling Liu 0001 |
ICDM | 4 |
| 2023 | Selecting and Composing Learning Rate Policies for Deep Neural NetworksabstractThe choice of learning rate (LR) functions and policies has evolved from a simple fixed LR to the decaying LR and the cyclic LR, aiming to improve the accuracy and reduce the training time of Deep Neural Networks (DNNs). This article presents a systematic approach to selecting and composing an LR policy for effective DNN training to meet desired target accuracy and reduce training time within the pre-defined training iterations. It makes three original contributions. First, we develop an LR tuning mechanism for auto-verification of a given LR policy with respect to the desired accuracy goal under the pre-defined training time constraint. Second, we develop an LR policy recommendation system (LRBench) to select and compose good LR policies from the same and/or different LR functions through dynamic tuning, and avoid bad choices, for a given learning task, DNN model, and dataset. Third, we extend LRBench by supporting different DNN optimizers and show the significant mutual impact of different LR policies and different optimizers. Evaluated using popular benchmark datasets and different DNN models (LeNet, CNN3, ResNet), we show that our approach can effectively deliver high DNN test accuracy, outperform the existing recommended default LR policies, and reduce the DNN training time by 1.6-6.7× to meet a targeted model accuracy. Yanzhao Wu 0001, Ling Liu 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2023 | Securing Distributed SGD Against Gradient Leakage ThreatsabstractThis paper presents a holistic approach to gradient leakage resilient distributed Stochastic Gradient Descent (SGD).First, we analyze two types of strategies for privacy-enhanced federated learning: (i) gradient pruning with random selection or low-rank filtering and (ii) gradient perturbation with additive random noise or differential privacy noise. We analyze the inherent limitations of these approaches and their underlying impact on privacy guarantee, model accuracy, and attack resilience.Next, we present a gradient leakage resilient approach to securing distributed SGD in federated learning, with differential privacy controlled noise as the tool. Unlike conventional methods with the per-client federated noise injection and fixed noise parameter strategy, our approach keeps track of the trend of per-example gradient updates. It makes adaptive noise injection closely aligned throughout the federated model training.Finally, we provide an empirical privacy analysis on the privacy guarantee, model utility, and attack resilience of the proposed approach. Extensive evaluation using five benchmark datasets demonstrates that our gradient leakage resilient approach can outperform the state-of-the-art methods with competitive accuracy performance, strong differential privacy guarantee, and high resilience against gradient leakage attacks. Wenqi Wei 0001, Ling Liu 0001, Jingya Zhou, Ka-Ho Chow 0001, Yanzhao Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | Learning Text-image Joint Embedding for Efficient Cross-modal Retrieval with Deep Feature EngineeringabstractThis article introduces a two-phase deep feature engineering framework for efficient learning of semantics enhanced joint embedding, which clearly separates the deep feature engineering in data preprocessing from training the text-image joint embedding model. We use the Recipe1M dataset for the technical description and empirical validation. In preprocessing, we perform deep feature engineering by combining deep feature engineering with semantic context features derived from raw text-image input data. We leverage LSTM to identify key terms, deep NLP models from the BERT family, TextRank, or TF-IDF to produce ranking scores for key terms before generating the vector representation for each key term by using Word2vec. We leverage Wide ResNet50 and Word2vec to extract and encode the image category semantics of food images to help semantic alignment of the learned recipe and image embeddings in the joint latent space. In joint embedding learning, we perform deep feature engineering by optimizing the batch-hard triplet loss function with soft-margin and double negative sampling, taking into account also the category-based alignment loss and discriminator-based alignment loss. Extensive experiments demonstrate that our SEJE approach with deep feature engineering significantly outperforms the state-of-the-art approaches. Zhongwei Xie, Ling Liu 0001, Yanzhao Wu 0001, Luo Zhong, Lin Li 0001 |
ACM Trans. Inf. Syst. | 3 |
| 2022 | A Comparative Measurement Study of Deep Learning as a Service FrameworkabstractBig data powered Deep Learning (DL) and its applications have blossomed in recent years, fueled by three technological trends: a large amount of digitized data openly accessible, a growing number of DL software frameworks in open source and commercial markets, and a selection of affordable parallel computing hardware devices. However, no single DL framework, to date, dominates in terms of performance and accuracy even for baseline classification tasks on standard datasets, making the selection of a DL framework an overwhelming task. This paper takes a holistic approach to conduct empirical comparison and analysis of four representative DL frameworks with three unique contributions.First, given a selection of CPU-GPU configurations, we show that for a specific DL framework, different configurations of its hyper-parameters may have a significant impact on both performance and accuracy of DL applications.Second, to the best of our knowledge, this study is the first to identify the opportunities for improving the training time performance and the accuracy of DL frameworks by configuring parallel computing libraries and tuning individual and multiple hyper-parameters.Third, we also conduct a comparative measurement study on the resource consumption patterns of four DL frameworks and their performance and accuracy implications, including CPU and memory usage, and their correlations to varying settings of hyper-parameters under different configuration combinations of hardware, parallel computing libraries. We argue that this measurement study provides in-depth empirical comparison and analysis of four representative DL frameworks, and offers practical guidance for service providers to deploying and delivering DL as a Service (DLaaS) and for application developers and DLaaS consumers to select the right DL frameworks for the right DL workloads. Yanzhao Wu 0001, Ling Liu 0001, Calton Pu, Wenqi Cao, Semih Sahin, Wenqi Wei 0001, Qi Zhang 0009 |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | Learning TFIDF Enhanced Joint Embedding for Recipe-Image Cross-Modal Retrieval ServiceabstractIt is widely acknowledged that learning joint embeddings of recipes with images is challenging due to the diverse composition and deformation of ingredients in cooking procedures. We present a Multi-modal Semantics enhanced Joint Embedding approach (MSJE) for learning a common feature space between the two modalities (text and image), with the ultimate goal of providing high-performance cross-modal retrieval services. Our MSJE approach has three unique features. First, we extract the TFIDF feature from the title, ingredients and cooking instructions of recipes. By determining the significance of word sequences through combining LSTM learned features with their TFIDF features, we encode a recipe into a TFIDF weighted vector for capturing significant key terms and how such key terms are used in the corresponding cooking instructions. Second, we combine the recipe TFIDF feature with the recipe sequence feature extracted through two-stage LSTM networks, which is effective in capturing the unique relationship between a recipe and its associated image(s). Third, we further incorporate TFIDF enhanced category semantics to improve the mapping of image modality and to regulate the similarity loss function during the iterative learning of cross-modal joint embedding. Experiments on the benchmark dataset Recipe1M show the proposed approach outperforms the state-of-the-art approaches. Zhongwei Xie, Ling Liu 0001, Yanzhao Wu 0001, Lin Li 0001, Luo Zhong |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | Transparent Network Memory Storage for Efficient Container Execution in Big Data CloudsabstractThis paper presents a transparent Container Network Memory storage device, coined as CNetMem, aiming to address the open problem of unpredictable performance degradation of containers when the working set of an application no longer fits in container memory. First, CNetMem will enable application tenants running in a container to park their working set memory/file to a faster network memory storage by organizing a group of remote memory nodes as remote memory donors. This allows CNetMem to take advantage of remote idle memory on a cluster before resorting to a slow local I/O subsystem like local disk without any modification of host OS or application. Second, CNetMem provides a hybrid batching technique to remove or alleviate performance bottlenecks in the I/O performance critical path for remote memory read/write with replication or disk backup for fault tolerance. Third, CNetMem introduces a rank-based node selection algorithm to find the optimal node for placing remote memory blocks across cluster. This helps CNetMem to reduce the performance impact due to remote memory eviction. Extensive experiments are conducted on three big data applications and four machine learning workloads. The results show that CNetMem achieves up to 172× throughput improvements compared to vanilla Linux and up to 5.9× completion time improvements over existing approaches in big data and ML workload. Juhyun Bae, Ling Liu 0001, Ka-Ho Chow 0001, Yanzhao Wu 0001, Gong Su, Arun Iyengar |
IEEE BigData | 4 |
| 2021 | Boosting Ensemble Accuracy by Revisiting Ensemble Diversity MetricsabstractNeural network ensembles are gaining popularity by harnessing the complementary wisdom of multiple base models. Ensemble teams with high diversity promote high failure independence, which is effective for boosting the overall ensemble accuracy. This paper provides an in-depth study on how to design and compute ensemble diversity, which can capture the complementary decision capacity of ensemble member models. We make three original contri-butions. First, we revisit the ensemble diversity metrics in the literature and analyze the inherent problems of poor correlation between ensemble diversity and ensemble ac-curacy, which leads to the low quality ensemble selection using such diversity metrics. Second, instead of computing diversity scores for ensemble teams of different sizes using the same criteria, we introduce focal model based ensemble diversity metrics, coined as FQ-diversity metrics. Our new metrics significantly improve the intrinsic correlation between high ensemble diversity and high ensemble accuracy. Third, we introduce a diversity fusion method, coined as the EQ-diversity metric, by integrating the top three most representative FQ-diversity metrics. Comprehensive experiments on two benchmark datasets (CIFAR-10 and ImageNet) show that our FQ and EQ diversity metrics are effective for selecting high diversity ensemble teams to boost overall ensemble accuracy. Yanzhao Wu 0001, Ling Liu 0001, Zhongwei Xie, Ka-Ho Chow 0001, Wenqi Wei 0001 |
CVPR | 1 |
| 2021 | Gradient-Leakage Resilient Federated LearningabstractFederated learning(FL) is an emerging distributed learning paradigm with default client privacy because clients can keep sensitive data on their devices and only share local training parameter updates with the federated server. However, recent studies reveal that gradient leakages in FL may compromise the privacy of client training data. This paper presents a gradient leakage resilient approach to privacy-preserving federated learning with per training example-based client differential privacy, coined as Fed-CDP. It makes three original contributions. First, we identify three types of client gradient leakage threats in federated learning even with encrypted client-server communications. We articulate when and why the conventional server coordinated differential privacy approach, coined as Fed-SDP, is insufficient to protect the privacy of the training data. Second, we introduce Fed-CDP, the per example-based client differential privacy algorithm, and provide a formal analysis of Fed-CDP with the (∊,δ) differential privacy guarantee, and a formal comparison between Fed-CDP and Fed-SDP in terms of privacy accounting. Third, we formally analyze the privacy-utility tradeoff for providing differential privacy guarantee by Fed-CDP and present a dynamic decay noise-injection policy to further improve the accuracy and resiliency of Fed-CDP. We evaluate and compare Fed-CDP and Fed-CDP(decay) with Fed-SDP in terms of differential privacy guarantee and gradient leakage resilience over five benchmark datasets. The results show that the Fed-CDP approach outperforms conventional Fed-SDP in terms of resilience to client gradient leakages while offering competitive accuracy performance in federated learning. Wenqi Wei 0001, Ling Liu 0001, Yanzhao Wu 0001, Gong Su, Arun Iyengar |
ICDCS | 3 |
| 2021 | Boosting Deep Ensemble Performance with Hierarchical PruningabstractDeep neural network ensembles have become attractive learning techniques with better generalizability over individual models. Some mission critical applications may require a large number of deep neural networks to achieve desirable accuracy and generalizability, making the ensemble execution costly with respect to runtime and space. This paper proposes a novel hierarchical ensemble pruning approach, which can effectively examine a given pool of M base models and identify smaller high quality deep ensembles of size $S(\ll M)$ with higher ensemble accuracy than the entire deep ensemble of all M models. Our hierarchical pruning approach, coined as HQ, combines three novel techniques. First, we show that the focal diversity metrics is innovative and can accurately capture the negative correlation among the member models of an ensemble, and the use of focal diversity metrics can boost ensemble accuracy. Second, we introduce a focal-diversity based hierarchical pruning algorithm to progressively identify low-cost ensembles with high ensemble diversity and accuracy. Third, we design a focal diversity consensus method to find smaller deep ensembles with low negative correlation. We demonstrate such ensembles offer high accuracy and high robustness while being more time and space efficient in ensemble decision making. Evaluated using two benchmark datasets, we show that the proposed focal diversity powered hierarchical pruning can find significantly smaller ensembles of deep neural network models while achieving the same or better classification generalizability. Yanzhao Wu 0001, Ling Liu 0001 |
ICDM | 1 |
| 2020 | Understanding Object Detection Through an Adversarial Lens
Ka-Ho Chow 0001, Ling Liu 0001, Mehmet Emre Gursoy, Stacey Truex, Wenqi Wei 0001, Yanzhao Wu 0001 |
ESORICS (2) | 6 |
| 2020 | A Framework for Evaluating Client Privacy Leakages in Federated Learning
Wenqi Wei 0001, Ling Liu 0001, Margaret L. Loper, Ka-Ho Chow 0001, Mehmet Emre Gursoy, Stacey Truex, Yanzhao Wu 0001 |
ESORICS (1) | 7 |
| 2019 | Denoising and Verification Cross-Layer Ensemble Against Black-box Adversarial AttacksabstractDeep neural networks (DNNs) have demonstrated impressive performance on many challenging machine learning tasks. However, DNNs are vulnerable to adversarial inputs generated by adding maliciously crafted perturbations to the benign inputs. As a growing number of attacks have been reported to generate adversarial inputs of varying sophistication, the defense-attack arms race has been accelerated. In this paper, we present MODEF, a cross-layer model diversity ensemble framework. MODEF intelligently combines unsupervised model denoising ensemble with supervised model verification ensemble by quantifying model diversity, aiming to boost the robustness of the target model against adversarial examples. Evaluated using eleven representative attacks on popular benchmark datasets, we show that MODEF achieves remarkable defense success rates, compared with existing defense methods, and provides a superior capability of repairing adversarial inputs and making correct predictions with high accuracy in the presence of black-box attacks. Ka-Ho Chow 0001, Wenqi Wei 0001, Yanzhao Wu 0001, Ling Liu 0001 |
IEEE BigData | 3 |
| 2019 | Demystifying Learning Rate Policies for High Accuracy Training of Deep Neural NetworksabstractLearning Rate (LR) is an important hyper-parameter to tune for effective training of deep neural networks (DNNs). Even for the baseline of a constant learning rate, it is non-trivial to choose a good constant value for training a DNN. Dynamic learning rates involve multi-step tuning of LR values at various stages of the training process and offer high accuracy and fast convergence. However, they are much harder to tune. In this paper, we present a comprehensive study of 13 learning rate functions and their associated LR policies by examining their range parameters, step parameters, and value update parameters. We propose a set of metrics for evaluating and selecting LR policies, including the classification confidence, variance, cost, and robustness, and implement them in LRBench, an LR benchmarking system. LRBench can assist end-users and DNN developers to select good LR policies and avoid bad LR policies for training their DNNs. We tested LRBench on Caffe, an open source deep learning framework, to showcase the tuning optimization of LR policies. Evaluated through extensive experiments, we attempt to demystify the tuning of LR policies by identifying good LR policies with effective LR value ranges and step sizes for LR update schedules. Yanzhao Wu 0001, Ling Liu 0001, Juhyun Bae, Ka-Ho Chow 0001, Arun Iyengar, Calton Pu, Wenqi Wei 0001, Lei Yu 0002, Qi Zhang 0009 |
IEEE BigData | 1 |
| 2019 | Memory Disaggregation: Research Problems and OpportunitiesabstractMemory usage imbalance has been consistently observed in many virtualized Clouds and production datacenters. Such temporal memory utilization variance is a major root cause for excessive paging and thrashing on virtual servers even though there are sufficient idle memory on the same node or in the Cloud cluster. Memory disaggregation is an emerging research and development endeavor towards addressing these memory usage imbalance problems. This paper first defines and characterizes the concept of memory disaggregation, and discusses the demands and challenges of efficient memory disaggregation in cloud datacenters. It then examines some promising research issues, design choices and directions to overcome some of the challenges posed by memory disaggregation. Specifically, it proposes two major new research challenges and solution directions for enabling elastic, on-demand disaggregated memory orchestration: (1) virtual server memory and node level memory co-design and (2) local memory and remote memory co-design. A brief description of two ongoing research projects is provided for both solution directions. The paper ends with a brief discussion of other advanced and emerging memory and storage technologies and potential opportunities for memory disaggregation. Ling Liu 0001, Wenqi Cao, Semih Sahin, Qi Zhang 0009, Juhyun Bae, Yanzhao Wu 0001 |
ICDCS | 6 |
| 2019 | Deep Neural Network Ensembles Against Deception: Ensemble Diversity, Accuracy and RobustnessabstractEnsemble learning is a methodology that integrates multiple DNN learners for improving prediction performance of individual learners. Diversity is greater when the errors of the ensemble prediction is more uniformly distributed. Greater diversity is highly correlated with the increase in ensemble accuracy. Another attractive property of diversity optimized ensemble learning is its robustness against deception: an adversarial perturbation attack can mislead one DNN model to misclassify but may not fool other ensemble DNN members consistently. In this paper we first give an overview of the concept of ensemble diversity and examine the three types of ensemble diversity in the context of DNN classifiers. We then describe a set of ensemble diversity measures, a suite of algorithms for creating diversity ensembles and for performing ensemble consensus (voted or learned) for generating high accuracy ensemble output by strategically combining outputs of individual members. This paper concludes with a discussion on a set of open issues in quantifying ensemble diversity for robust deep learning. Ling Liu 0001, Wenqi Wei 0001, Ka-Ho Chow 0001, Margaret L. Loper, Mehmet Emre Gursoy, Stacey Truex, Yanzhao Wu 0001 |
MASS | 7 |
| 2018 | Experimental Characterizations and Analysis of Deep Learning FrameworksabstractBig Data has fueled the wide deployment of Deep Learning (DL) in many fields, such as image classification, voice recognition and NLP. The growing number of open source DL software frameworks has put forward high demands on comparative study of their efficiency with respect to both runtime performance and accuracy. This paper presents a brief overview of our empirical evaluation of four representative DL frameworks: TensorFlow, Caffe, Torch and Theano through a comparative analysis and characterization. First, we show that the complex interactions among neural networks (NN), hyper-parameters, their specific runtime implementations and datasets are latent factors for the uncertainty of runtime performance and accuracy. Second, we characterized the CPU/GPU resource usage patterns under different configurations for different frameworks to obtain an in-depth understanding of the impact of different batch sizes. Third, we describe the data loading process of ImageNet for TensorFlow and present an experimental characterization of TensorFlow with respect to its data loading process when the dataset is too large to fit into the main memory of the CPU server. We conjecture that our experimental characterization and analysis can offer empirical guidance for users and application developers to select the right DL frameworks and configurations for their domain-specific learning tasks and datasets. Yanzhao Wu 0001, Wenqi Cao, Semih Sahin, Ling Liu 0001 |
IEEE BigData | 1 |
| 2018 | Benchmarking Deep Learning Frameworks: Design Considerations, Metrics and BeyondabstractWith increasing number of open-source deep learning (DL) software tools made available, benchmarking DL software frameworks and systems is in high demand. This paper presents design considerations, metrics and challenges towards developing an effective benchmark for DL software frameworks and illustrate our observations through a comparative study of three popular DL frameworks: TensorFlow, Caffe, and Torch. First, we show that these deep learning frameworks are optimized with their default configurations settings. However, the default configuration optimized on one specific dataset may not work effectively for other datasets with respect to runtime performance and learning accuracy. Second, the default configuration optimized on a dataset by one DL framework does not work well for another DL framework on the same dataset. Third, experiments show that different DL frameworks exhibit different levels of robustness against adversarial examples. Through this study, we envision that unlike traditional performance-driven benchmarks, benchmarking deep learning software frameworks should take into account of both runtime and accuracy and their latent interaction with hyper-parameters and data-dependent configurations of DL frameworks. Ling Liu 0001, Yanzhao Wu 0001, Wenqi Wei 0001, Wenqi Cao, Semih Sahin, Qi Zhang 0009 |
ICDCS | 2 |
| 2018 | CCAligner: a token based large-gap clone detectorabstractCopying code and then pasting with large number of edits is a common activity in software development, and the pasted code is a kind of complicated Type-3 clone. Due to large number of edits, we consider the clone as a large-gap clone. Large-gap clone can reflect the extension of code, such as change and improvement. The existing state-of-the-art clone detectors suffer from several limitations in detecting large-gap clones. In this paper, we propose a tool, CCAligner, using code window that considers e edit distance for matching to detect large-gap clones. In our approach, a novel e-mismatch index is designed and the asymmetric similarity coefficient is used for similarity measure. We thoroughly evaluate CCAligner both for large-gap clone detection, and for general Type-1, Type-2 and Type-3 clone detection. The results show that CCAligner performs better than other competing tools in large-gap clone detection, and has the best execution time for 10MLOC input with good precision and recall in general Type-1 to Type-3 clone detection. Compared with existing state-of-the-art tools, CCAligner is the best performing large-gap clone detection tool, and remains competitive with the best clone detectors in general Type-1, Type-2 and Type-3 clone detection. Jeffrey Svajlenko, Yanzhao Wu 0001, Chanchal Kumar Roy |
ICSE | 3 |