EDBT 2026 Demo / reviewers in the wild / expert
Li Li 0064
dblp:53/2189-64
· DBLP profile ↗
52ranked-venue papers
4as first author
44since 2021 · last 2026
0000-0002-2044-8289ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-author · 15 since 2021Computer networks · 15 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Security and privacy · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the EdgeabstractFederated fine-tuning enables privacypreserving LLM adaptation but faces a critical bottleneck: the disparity between LLMs' high memory demands and edge devices' limited capacity.To break the memory barrier, we propose Chain Federated Fine-Tuning (CHAINFED), an innovative paradigm that forgoes end-to-end updates in favor of a sequential, layer-by-layer manner.It first trains the initial adapter to convergence, freezes its weights, and then proceeds to the next.This iterative train-and-freeze process forms an optimization chain, gradually enhancing the model's task-specific proficiency.CHAINFED further integrates three core techniques: 1) Dynamic Layer Co-Tuning to bridge semantic gaps between sequentially tuned layers and facilitate information flow; 2) Globally Perceptive Optimization to endow each adapter with foresight beyond its local objective; 3) Function-Oriented Adaptive Tuning to automatically identify the optimal fine-tuning starting point.Extensive experiments on multiple benchmarks demonstrate the superiority of CHAINFED over existing methods, boosting average accuracy by up to 46.46%. Yebo Wu, Jingguang Li, Chunlin Tian, Kahou Tam, Zhijiang Guo, Li Li 0064 |
ACL (1) | 6 |
| 2026 | FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUsabstractTransformer-based models have enabled unprecedented capabilities across language, vision, and multimodal tasks. On-device fine-tuning of transformer models offers a privacy-preserving path to personalized AI, yet remains inefficient on mobile GPUs due to severe memory constraints and frequent layout transformations in attention mechanism during training. Existing mobile training frameworks either use unified layouts for forward and backward passes - leading to fragmented memory access and poor GPU utilization during backpropagation - or rely on explicit layout conversions, which introduce significant transformation overhead.To overcome this, we propose FBLayout, a layout-aware framework that co-designs tensor organization with mobile GPU platforms. FBLayout introduces: (1) a unified R-Tile layout for multidimensional reductions across forward/backward passes; (2) tile-based index transformation to eliminate physical data movement; and (3) activation-guided layout selection to propagate efficient layouts globally. Evaluations on seven transformer models across different mobile phones (including ARM Mali and Qualcomm Adreno GPUs) show that FBLayout achieves 2.2-5.7× speedup over MNN, TFLite, and TVM, while significantly improving cache efficiency and reducing memory footprint, enabling practical on-device large model fine-tuning. Kahou Tam, Wei Niu 0002, Xiaomin Ouyang, Cheng-Zhong Xu 0001, Li Li 0064 |
MobiSys | 6 |
| 2026 | Floe: Federated Specialization for Real-Time LLM-SLM InferenceabstractDeploying large language models (LLMs) in realtime systems is challenging due to their high resource demands and privacy concerns. We propose Floe a hybrid federated learning framework designed for latency-sensitive, resourceconstrained environments. Floeombines a cloud-based blackbox LLM with lightweight small language models (SLMs) on edge devices to enable low-latency, privacy-preserving inference. Personal data and fine-tuning remain on-device, while the cloud LLM contributes general knowledge without exposing proprietary weights. A heterogeneity-aware LoRA adaptation strategy ensures efficient edge deployment across diverse hardware, and a logit-level fusion mechanism enables real-time coordination between edge and cloud models. Experiments demonstrate that Floenhances user privacy and personalization, while significantly improving model performance and reducing inference latency on edge devices under real-time constraints, compared to baseline approaches. Chunlin Tian, Kahou Tam, Yebo Wu, Shuaihang Zhong, Li Li 0064, Nicholas D. Lane, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | EICopilot: Search and Explore Enterprise Information Over Large-Scale Knowledge Graphs with LLM-Driven Agents
Yuhui Yun, Huilong Ye, Jingfeng Deng, Ruojia Li, Li Li 0064, Haoyi Xiong |
IEEE Big Data | 6 |
| 2025 | Breaking the Memory Wall for Heterogeneous Federated Learning via Progressive TrainingabstractFederated Learning (FL) enables multiple devices to collaboratively train a shared model while preserving data privacy. Most existing research assumes that all participating devices have sufficient resources to support the training process. However, the high memory requirements of model training present a significant challenge to deploying FL on resource-constrained devices in practical scenarios. To this end, this paper presents ProFL, a new framework that effectively addresses the memory constraints in FL. Rather than updating the full model during local training, ProFL partitions the model into blocks based on its original architecture and trains each block in a progressive fashion. It first trains the front blocks and safely freezes them after convergence. Training of the next block is then triggered. This process progressively grows the model to be trained until the training of the full model is completed. In this way, the peak memory footprint is effectively reduced for feasible deployment on heterogeneous devices. In order to preserve the feature representation of each block, the training process is divided into two stages: model shrinking and model growing. During the model shrinking stage, we meticulously design corresponding output modules to assist each block in learning the expected feature representation and obtain the initialization model parameters. Subsequently, the obtained output modules and initialization model parameters are utilized in the corresponding model growing stage, which progressively trains the full model. Additionally, a novel metric from the scalar perspective is proposed to assess the learning status of each block, enabling us to securely freeze it after convergence and initiate the training of the next one. Finally, we theoretically prove the convergence of ProFL and conduct extensive experiments on representative models and datasets to evaluate its effectiveness. The results demonstrate that ProFL effectively reduces the peak memory footprint by up to 57.4% and improves model accuracy by up to 82.4%. Yebo Wu, Li Li 0064, Cheng-Zhong Xu 0001 |
KDD (1) | 2 |
| 2025 | HCInfer: Hierarchical Coordination for Real-Time Collaborative Inference of LLM on the EdgeabstractDeploying large language models (LLMs) on edge devices enables real-time responses while preserving user privacy. However, constrained memory and compute resources pose significant challenges for high-quality, single-device inference. To address this, we propose HCInfer, a hierarchical coordination framework for collaborative LLM inference across edge devices. By leveraging idle neighboring devices, HCInfer alleviates performance bottlenecks typical in isolated deployments. HCInfer employs a two-level coordination strategy. At the inter-device level, it leverages idle neighboring devices to collaboratively process attention computations, significantly reducing synchronization overhead. At the intra-device level, it applies finegrained memory and compute optimizations to fully exploit local hardware capabilities. Building on this architecture, HCInfer integrates three key components: (1) Asymmetric Transformer decomposition decouples attention and FFN computation, enabling selective and parallel execution across devices. (2) Layer-wise subdeadline scheduling dynamically profiles execution latency and adapts precision or structure to meet real-time constraints (3) An Overhead Mitigation Module efficiently manages on-device resource usage to support scalability without overwhelming hardware. We evaluate HCInfer on PC, smart home, and mobile platforms using OPT-13B, Qwen2.5-14B, and Llama2-13B models. Experiments show HCInfer achieves 1.67× to 4.3× speedup in TTFT and 1.16× to 17.15× speedup in TPOT compared to existing baselines, maintaining a sub-deadline miss rate (SubDMR) of 15.3% under worst-case conditions while keeping model accuracy degradation within 8% for typical cases and up to 11% in extreme scenarios. These results demonstrate HCInfer's potential to enable efficient and responsive LLM inference in real-world edge environments. Lizi Zhang, Cheng-Zhong Xu 0001, Li Li 0064 |
RTSS | 4 |
| 2025 | CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the Edge
Chunlin Tian, Xinpeng Qin, Kahou Tam, Li Li 0064, Yuanzhe Zhao, Minglei Zhang, Cheng-Zhong Xu 0001 |
USENIX ATC | 4 |
| 2025 | AssyLLM: Efficient Federated Fine-tuning of LLMs via Assembling Pre-trained Blocks
Shichen Zhan, Li Li 0064, Cheng-Zhong Xu 0001 |
USENIX ATC | 2 |
| 2025 | Noise-Robust Federated Learning via Interclient Co-DistillationabstractFederated learning (FL) is a new learning paradigm that enables multiple clients to collaboratively train a high-performance model while preserving user privacy. However, the effectiveness of FL heavily relies on the availability of accurately labeled data, which can be challenging to obtain in real-world scenarios. To address this issue and robustly train shared models using distributed noisy labeled data, we propose FedDQ, a noise-robust FL framework that utilizes co-distillation and quality-aware aggregation techniques. FedDQ incorporates two key features: a noise-adaptive training strategy and an efficient label-correcting mechanism. The noise-adaptive training strategy relies on the estimation of labels' noise levels to dynamically adjust clients' training engagement, which mitigates the impact of wrong labels while efficiently exploring features from clean data. In addition, FedDQ designs a two-head network and employs it for co-distillation. The co-distillation strategy facilitates knowledge transfer among clients to share the representational capabilities. Besides, FedDQ enhances label correction to rectify improper labels through co-filtering and label correction. The experimental results demonstrate the effectiveness of FedDQ in improving model performance and handling noisy data challenges in FL settings. On the CIFAR-100 dataset with noisy labels, FedDQ exhibits a notable improvement of up to 32.4% compared to the baseline method. Liang Gao 0001, Li Li 0064, Yingwen Chen 0001, Shaojing Fu, Dongsheng Wang 0004, Siwei Wang 0001, Cheng-Zhong Xu 0001, Ming Xu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Federated Noisy Client LearningabstractFederated learning (FL) collaboratively trains a shared global model depending on multiple local clients, while keeping the training data decentralized to preserve data privacy. However, standard FL methods ignore the noisy client issue, which may harm the overall performance of the shared model. We first investigate the critical issue caused by noisy clients in FL and quantify the negative impact of the noisy clients in terms of the representations learned by different layers. We have the following two key observations: 1) the noisy clients can severely impact the convergence and performance of the global model in FL and 2) the noisy clients can induce greater bias in the deeper layers than the former layers of the global model. Based on the above observations, we propose federated noisy client learning (Fed-NCL), a framework that conducts robust FL with noisy clients. Specifically, Fed-NCL first identifies the noisy clients through well estimating the data quality and model divergence. Then robust layerwise aggregation is proposed to adaptively aggregate the local models of each client to deal with the data heterogeneity caused by the noisy clients. We further perform label correction on the noisy clients to improve the generalization of the global model. Experimental results on various datasets demonstrate that our algorithm boosts the performances of different state-of-the-art systems with noisy clients. Our code is available at https://github.com/TKH666/Fed-NCL. Kahou Tam, Li Li 0064, Bo Han 0003, Cheng-Zhong Xu 0001, Huazhu Fu |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | m$^{2}$2LLM: A Multi-Dimensional Optimization Framework for LLM Inference on Mobile DevicesabstractLarge Language Models (LLMs) are reshaping mobile AI. Directly deploying LLMs on mobile devices is an emerging paradigm that can widely support different mobile applications while preserving data privacy. However, intensive memory footprint, long inference latency and high energy consumption severely bottlenecks on-device inference of LLM in real-world scenarios. In response to these challenges, this work introduces m$^{2}$LLM, an innovative framework that performs joint optimization from multiple dimensions for on-device LLM inference in order to strike a balance among performance, realtimeliness and energy efficiency. Specifically, m$^{2}$LLM features the following four core components including : 1) Hardware-aware Model Customization, 2) Elastic Chunk-wise Pipeline, 3) Latency-guided Prompt Compression and 4) Layer-wise Resource Scheduling. These four components interact with each other in order to guide the inference process from the following three dimensions. At the model level, m$^{2}$LLM designs an elastic chunk-wise pipeline to expand device memory and customize the model according to the hardware configuration, maximizing performance within the memory budget. At the prompt level, facing the stochastic input, m$^{2}$LLM judiciously compresses the prompts in order to guarantee the first token can be generated in time while maintaining the semantic information. Additionally, at the system level, the layer-wise resource scheduler is employed in order to complete the token generation process with minimized energy consumption while guaranteeing the realtimeness in the highly dynamic mobile environment. m$^{2}$LLM is evaluated on off-the-shelf smartphone with represented models and datasets. Compared to baseline methods, m$^{2}$LLM delivers 2.99–13.5× TTFT acceleration and 2.28–24.3× energy savings, with only a minimal model performance loss of 2% –7% . Xiaobo Zhou 0002, Li Li 0064 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | Ranking-based Client Imitation Selection for Efficient Federated LearningabstractFederated Learning (FL) enables multiple devices to collaboratively train a shared model while ensuring data privacy. The selection of participating devices in each training round critically affects both the model performance and training efficiency, especially given the vast heterogeneity in training capabilities and data distribution across devices. To deal with these challenges, we introduce a novel device selection solution called FedRank, which is based on an end-to-end, ranking-based model that is pre-trained by imitation learning against state-of-the-art analytical approaches. It not only considers data and system heterogeneity at runtime but also adaptively and efficiently chooses the most suitable clients for model training. Specifically, FedRank views client selection in FL as a ranking problem and employs a pairwise training strategy for the smart selection process. Additionally, an imitation learning-based approach is designed to counteract the cold-start issues often seen in state-of-the-art learning-based approaches. Experimental results reveal that FedRank boosts model accuracy by 5.2% to 56.9%, accelerates the training convergence up to $2.01 \times$ and saves the energy consumption up to 40.1%. Chunlin Tian, Xinpeng Qin, Li Li 0064, Cheng-Zhong Xu 0001 |
ICML | 4 |
| 2024 | FedGCS: A Generative Framework for Efficient Client Selection in Federated Learning via Gradient-based Optimization
Zhiyuan Ning 0001, Chunlin Tian, Meng Xiao 0001, Wei Fan 0010, Pengyang Wang, Li Li 0064, Pengfei Wang 0008, Yuanchun Zhou |
IJCAI | 6 |
| 2024 | FedMG: A Federated Multi-Global Optimization Framework for Autonomous Driving ControlabstractControl is a critical module of autonomous driving systems, which ensures safety and enhances the human-machine interface. Due to the diverse control demands dictated by different driving scenarios, autonomous vehicles require a data-intensive, adaptive, and intelligent controller. To speed up the control process and improve the performance in different scenarios, we introduce a novelty federated learning framework FedMG, which efficiently coordinates diverse vehicles to train a collaboratively models while preserving data privacy to tune the control process. Through detailed analysis of driving scenarios, vehicles are clustered to different groups based on driving scenarios to seek a balance between data quality and communication efficiency. It enables the consolidation of several global models, each optimized for peak performance, thereby enhancing the overall system’s effectiveness. Extensive experiments with different numbers of vehicles and a variety of driving scenarios demonstrate the effectiveness of FedMG. The framework significantly reduces cumulative driving errors, achieving reductions ranging from 5.42% to 76.43%, while improving user comfort, with improvements ranging from 2.23% to 34.61% over baselines. Jialiang Ma, Chunlin Tian, Li Li 0064, Cheng-Zhong Xu 0001 |
IWQoS | 3 |
| 2024 | GreenLLM: Towards Efficient Large Language Model via Energy-aware PruningabstractThis paper proposes GreenLLM, a framework that effectively deploys generative Large Language Models (LLMs) on resource-limited edge devices to well meet the memory and timing constraints with minimized energy consumption. Specifically, GreenLLM employs an energy estimation scheme based on physical hardware to guide a pruning-ratio generator incorporating space, weight, and power (SWaP) constraints for optimal pruning ratio. For each layer, we employ a dependency-aware energy-efficient Pruner in a task-agnostic manner, maximally preserving most of the LLM functionality. Finally, we use downstream datasets to fine-tune the pruned model to recover performance. Chunlin Tian, Xinpeng Qin, Li Li 0064 |
IWQoS | 3 |
| 2024 | FedEKT: Ensemble Knowledge Transfer for Model-Heterogeneous Federated LearningabstractFederated Learning (FL) enables multiple clients to collaboratively train a shared server model while preserving data privacy. Most existing FL systems rely on the assumption that the server model and client models have homogeneous architecture. However, intensive resource requirements during the training process prevent low-end devices from contributing to the server model with their own data. On the other hand, the resource constraints on participating clients can significantly limit the size of the server model in the model-homogeneous setting, thereby restricting the application scope of FL. In this work, we propose FedEKT, a novel model-heterogeneous FL system designed to obtain a high-performance large server model while benefiting heterogeneous small client models. Specifically, a new aggregation approach is designed to enable the integration of knowledge from heterogeneous client models to a large server model while mitigating the adverse effects of biases stemming from data heterogeneity. Subsequently, to enhance the performance of client models by benefiting from the high-performance server model, FedEKT distills this large server model into multiple heterogeneous client models, facilitating the transfer of integrated knowledge back to the client models. In addition, we design specialized modules within the model and communication strategy to accomplish aggregation and transfer of knowledge in a data-free manner. The evaluation results demonstrate that FedEKT enhances the accuracy of the server model and client models by up to 53.96% and 12.35%, respectively, compared with the state-of-the-art FL approach on CIFAR-100. Meihan Wu, Li Li 0064, Tao Chang, Peng Qiao, Cui Miao, Jie Zhou 0032, Jingnan Wang, Xiaodong Wang 0002 |
IWQoS | 2 |
| 2024 | PFed-DBA: Distribution Bias Aware Personalized Federated Learning for Data HeterogeneityabstractPersonalized Federated Learning (PFL) aims to learn a custom model for each distributed client while benefiting from collaborative training in order to overcome the detrimental impact of data heterogeneity. Despite the promising benefits, the existing approaches often compromise the generalization performance of personalized models, as they solely focus on enhancing the personalization capability of models or merely aim to strike a balance between personalization and generalization. Indeed, increasing the personalization capability while preserving the strong generalization performance enabled by collaborative training remains a challenge for PFL, as the two objectives seem to compete with each other. To tackle this challenge, we investigate the relationship between model generalization and personalization under different degrees of heterogeneity. We find that besides the client-specific data distribution, the distribution bias between the unique data distribution of each client and that of the whole population is another critical factor that prominently impacts these two performances. Motivated by the above finding, we propose PFed-DBA, a novel PFL framework that effectively perceives this distribution bias to guide the training process. Concretely, we design the PFL models as a skip-connection network between a shared module for learning the shared representations delivering the common distribution of data across all clients and a personalized module for learning the personalized representations of the heterogeneous distribution bias. Then, we devise corresponding loss functions, aggregation strategy, and updating strategy in order to make the two modules intelligently complement each other. Moreover, we conduct extensive experiments to evaluate the effectiveness of PFed-DBA. The results show that PFed-DBA improves model accuracy to 12.34% at best compared with the state-of-the-art. Meihan Wu, Li Li 0064, Tao Chang, Jie Zhou 0032, Eric Rigall, Cui Miao, Xiaodong Wang 0002, Cheng-Zhong Xu 0001 |
IWQoS | 2 |
| 2024 | Heterogeneity-Aware Memory Efficient Federated Learning via Progressive Layer FreezingabstractFederated Learning (FL) emerges as a new learning paradigm that enables multiple devices to collaboratively train a shared model while preserving data privacy. However, intensive memory footprint during the training process severely bottlenecks the deployment of FL on resource-limited mobile devices in real-world cases. Thus, a framework that can effectively reduce the memory footprint while guaranteeing training efficiency and model accuracy is crucial for FL.In this paper, we propose SmartFreeze, a framework that effectively reduces the memory footprint by conducting the training in a progressive manner. Instead of updating the full model in each training round, SmartFreeze divides the shared model into blocks consisting of a specified number of layers. It first trains the front block with a well-designed output module, safely freezes it after convergence, and then triggers the training of the next one. This process iterates until the whole model has been successfully trained. In this way, the backward computation of the frozen blocks and the corresponding memory space for storing the intermediate outputs and gradients are effectively saved. Except for the progressive training framework, SmartFreeze consists of the following two core components: a pace controller and a participant selector. The pace controller is designed to effectively monitor the training progress of each block at runtime and safely freezes them after convergence while the participant selector selects the right devices to participate in the training for each block by jointly considering the memory capacity, the statistical and system heterogeneity. Extensive experiments are conducted to evaluate the effectiveness of SmartFreeze on both simulation and hardware testbeds. The results demonstrate that SmartFreeze effectively reduces average memory usage by up to 82%. Moreover, it simultaneously improves the model accuracy by up to 83.1% and accelerates the training process up to 2.02 ×. Yebo Wu, Li Li 0064, Chunlin Tian, Chang Tao, Wang Cong, Cheng-Zhong Xu 0001 |
IWQoS | 2 |
| 2024 | Heterogeneity-Aware Coordination for Federated Learning via Stitching Pre-trained blocksabstractFederated learning (FL) coordinates multiple devices to collaboratively train a shared model while preserving data privacy. However, large memory footprint and high energy consumption during the training process excludes the low-end devices from contributing to the global model with their own data, which severely deteriorates the model performance in real-world scenarios. In this paper, we propose FedStitch, a hierarchical coordination framework for heterogeneous federated learning with pre-trained blocks. Unlike the traditional approaches that train the global model from scratch, for a new task, FedStitch composes the global model via stitching pre-trained blocks. Specifically, each participating client selects the most suitable block based on their local data from the candidate pool composed of blocks from pre-trained models. The server then aggregates the optimal block for stitching. This process iterates until a new stitched network is generated. Except for the new training paradigm, FedStitch consists of the following three core components: 1) an RL-weighted aggregator, and 2) a search space optimizer deployed on the server side, and 3) a local energy optimizer deployed on each participating client. The RL-weighted aggregator helps to select the right block in the non-IID scenario, while the search space optimizer continuously reduces the size of the candidate block pool during stitching. Meanwhile, the local energy optimizer is designed to minimize the energy consumption of each client while guaranteeing the overall training progress. The results demonstrate that compared to existing approaches, FedStitch improves the model accuracy up to 20.93%. At the same time, it achieves up to 8.12× speedup, reduces the memory footprint up to 79.5%, and achieves 89.41% energy saving at most during the learning procedure. Shichen Zhan, Yebo Wu, Chunlin Tian, Li Li 0064 |
IWQoS | 5 |
| 2024 | When, Where, and What? A Benchmark for Accident Anticipation and Localization with Large Language ModelsabstractAs autonomous driving systems increasingly become part of daily transportation, the ability to accurately anticipate and mitigate potential traffic accidents is paramount. Traditional accident anticipation models primarily utilizing dashcam videos are adept at predicting when an accident may occur but fall short in localizing the incident and identifying involved entities. Addressing this gap, this study introduces a novel framework that integrates Large Language Models (LLMs) to enhance predictive capabilities across multiple dimensions-what, when, and where accidents might occur. We develop an innovative chain-based attention mechanism that dynamically adjusts to prioritize high-risk elements within complex driving scenes. This mechanism is complemented by a three-stage model that processes outputs from smaller models into detailed multimodal inputs for LLMs, thus enabling a more nuanced understanding of traffic dynamics. Empirical validation on the DAD, CCD, and A3D datasets demonstrates superior performance in Average Precision (AP) and Mean Time-To-Accident (mTTA), establishing new benchmarks for accident prediction technology. Our approach not only advances the technological framework for autonomous driving safety but also enhances human-AI interaction, making predictive insights generated by autonomous systems more intuitive and actionable. Haicheng Liao, Yongkang Li 0003, Chengyue Wang 0001, Yanchen Guan, Kahou Tam, Chunlin Tian, Li Li 0064, Cheng-Zhong Xu 0001, Zhenning Li 0001 |
ACM Multimedia | 7 |
| 2024 | CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus AttentionsabstractAccurately and promptly predicting accidents among surrounding traffic agents from camera footage is crucial for the safety of autonomous vehicles (AVs). This task presents substantial challenges stemming from the unpredictable nature of traffic accidents, their long-tail distribution, the intricacies of traffic scene dynamics, and the inherently constrained field of vision of onboard cameras. To address these challenges, this study introduces a novel accident anticipation framework for AVs, termed CRASH. It seamlessly integrates five components: object detector, feature extractor, object-aware module, context-aware module, and multi-layer fusion. Specifically, we develop the object-aware module to prioritize high-risk objects in complex and ambiguous environments by calculating the spatial-temporal relationships between traffic agents. In parallel, the context-aware is also devised to extend global visual information from the temporal to the frequency domain using the Fast Fourier Transform (FFT) and capture fine-grained visual features of potential objects and broader context cues within traffic scenes. To capture a wider range of visual cues, we further propose a multi-layer fusion that dynamically computes the temporal dependencies between different scenes and iteratively updates the correlations between different visual features for accurate and timely accident prediction. Evaluated on real-world datasets-Dashcam Accident Dataset (DAD), Car Crash Dataset (CCD), and AnAn Accident Detection (A3D) datasets-our model surpasses existing top baselines in critical evaluation metrics like Average Precision (AP) and mean Time-To-Accident (mTTA). Importantly, its robustness and adaptability are particularly evident in challenging driving scenarios with missing or limited training data, demonstrating significant potential for application in real-world autonomous driving systems. Haicheng Liao, Huanming Shen, Chengyue Wang 0001, Chunlin Tian, Kahou Tam, Li Li 0064, Cheng-Zhong Xu 0001, Zhenning Li 0001 |
ACM Multimedia | 7 |
| 2024 | HydraLoRA: An Asymmetric LoRA Architecture for Efficient Fine-TuningabstractAdapting Large Language Models (LLMs) to new tasks through fine-tuning has been made more efficient by the introduction of Parameter-Efficient Fine-Tuning (PEFT) techniques, such as LoRA. However, these methods often underperform compared to full fine-tuning, particularly in scenarios involving complex datasets. This issue becomes even more pronounced in complex domains, highlighting the need for improved PEFT approaches that can achieve better performance. Through a series of experiments, we have uncovered two critical insights that shed light on the training and parameter inefficiency of LoRA. Building on these insights, we have developed HydraLoRA, a LoRA framework with an asymmetric structure that eliminates the need for domain expertise. Our experiments demonstrate that HydraLoRA outperforms other PEFT approaches, even those that rely on domain knowledge during the training and inference phases. Our anonymous codes are submitted with the paper and will be publicly available. Code is available: https://github.com/Clin0212/HydraLoRA. Chunlin Tian, Zhijiang Guo, Li Li 0064, Cheng-Zhong Xu 0001 |
NeurIPS | 4 |
| 2024 | FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor ManagementabstractFederated Learning (FL) emerges as a new learning paradigm that enables multiple devices to collaboratively train a shared model while preserving data privacy. However, one fundamental and prevailing challenge that hinders the deployment of FL on mobile devices is the memory limitation. This paper proposes FedHybrid, a novel framework that effectively reduces the memory footprint during the training process while guaranteeing the model accuracy and the overall training progress. Specifically, FedHybrid first selects the participating devices for each training round by jointly evaluating their memory budget, computing capability, and data diversity. After that, it judiciously analyzes the computational graph and generates an execution plan for each selected client in order to meet the corresponding memory budget while minimizing the training delay through employing a hybrid of recomputation and compression techniques according to the characteristic of each tensor. During the local training process, FedHybrid carries out the execution plan with a well-designed activation compression technique to effectively achieve memory reduction with minimum accuracy loss. We conduct extensive experiments to evaluate FedHybrid on both simulation and off-the-shelf mobile devices. The experiment results demonstrate that FedHybrid achieves up to a 39.1% increase in model accuracy and a 15.5X reduction in wall clock time under various memory budgets compared with the baselines. Kahou Tam, Chunlin Tian, Li Li 0064, Haikai Zhao, Cheng-Zhong Xu 0001 |
SenSys | 3 |
| 2024 | Breaking the Memory Wall for Heterogeneous Federated Learning via Model SplittingabstractFederated Learning (FL) enables multiple devices to collaboratively train a shared model while preserving data privacy. Ever-increasing model complexity coupled with limited memory resources on the participating devices severely bottlenecks the deployment of FL in real-world scenarios. Thus, a framework that can effectively break the memory wall while jointly taking into account the hardware and statistical heterogeneity in FL is urgently required. In this article, we proposeSmartSplita framework that effectively reduces the memory footprint on the device side while guaranteeing the training progress and model accuracy for heterogeneous FL through model splitting. Towards this end,SmartSplitemploys a hierarchical structure to adaptively guide the overall training process. In each training round, the central manager, hosted on the server, dynamically selects the participating devices and sets the cutting layer by jointly considering the memory budget, training capacity, and data distribution of each device. The MEC manager, deployed within the edge server, proceeds to split the local model and perform training of the server-side portion. Meanwhile, it fine-tunes the splitting points based on the time-evolving statistical importance. The on-device manager, embedded inside each mobile device, continuously monitors the local training status while employing cost-aware checkpointing to match the runtime dynamic memory budget. Extensive experiments on representative datasets are conducted on both commercial off-the-shelf mobile device testbeds. The experimental results show thatSmartSplitexcels in FL training on highly memory-constrained mobile SoCs, offering up to a 94% peak latency reduction and 100-fold memory savings. It enhances accuracy performance by 1.49%-57.18% and adaptively adjusts to dynamic memory budgets through cost-aware recomputation Chunlin Tian, Li Li 0064, Kahou Tam, Yebo Wu, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | FedCoop: Cooperative Federated Learning for Noisy LabelsabstractFederated Learning coordinates multiple clients to collaboratively train a shared model while preserving data privacy. However, the training data with noisy labels located on the participating clients severely harm the model performance. In this paper, we propose FedCoop, a cooperative Federated Learning framework for noisy labels. FedCoop mainly contains three components and conducts robust training in two phases, data selection and model training. In the data selection phase, in order to mitigate the confirmation bias caused by a single client, the Loss Transformer intelligently estimates the probability of each sample’s label to be clean through cooperating with the helper clients, which have high data trustability and similarity. After that, the Feature Comparator evaluates the label quality for each sample in terms of latent feature space in order to further improve the robustness of noisy label detection. In the model training phase, the Feature Matcher trains the model on both the noisy and clean data in a semi-supervised manner to fully utilize the training data and exploits the feature of global class to increase the consistency of pseudo labeling across the clients. The experimental results show FedCoop outperforms the baselines on various datasets with different noise settings. It effectively improves the model accuracy up to 62% and 27% on average compared with the baselines. Kahou Tam, Li Li 0064, Cheng-Zhong Xu 0001 |
ECAI | 2 |
| 2023 | HCPerf: Driving Performance-Directed Hierarchical Coordination for Autonomous VehiclesabstractThe rapid development of autonomous driving poses new research challenges to the on-vehicle computing system. In particular, the execution time of autonomous driving tasks highly depends on the specific driving environment. For instance, the execution time of configurable sensor fusion increases significantly as the scene becomes complex, which leads to end-to-end deadline misses from sensing to control and may cause accidents. Thus, a framework that can effectively utilize the system resources to guarantee the end-to-end deadlines of autonomous driving tasks as well as effectively prioritize the responsiveness and throughput of the control commands is crucial for autonomous driving. In this paper, we propose HCPerf, a performance-directed hierarchical coordination framework that intelligently coordinates the autonomous driving tasks with high execution time variation and complex dependencies according to the driving performance in real-time. Specifically, HCPerf mainly consists of two coordinators. The internal coordinator intelligently schedules the tasks according to the driving performance of the vehicle in order to help them meet the end-to-end deadlines while well prioritizing the responsiveness and throughput of the control commands. At the same time, the external coordinator dynamically tunes the rates of tasks according to the schedulability in order to efficiently utilize the system resource. We conduct extensive experiments on both simulation and hardware testbeds with the representative autonomous driving application. The results show that HCPerf can effectively improve the driving performance by 7.69%-45.94% in different driving scenarios. Jialiang Ma, Li Li 0064, Zejiang Wang, Jun Wang 0001, Cheng-Zhong Xu 0001 |
ICDCS | 2 |
| 2023 | FedHybrid: Hierarchical Hybrid Training for High-Performance Federated LearningabstractFederated Learning coordinates multiple devices to train a shared model while preserving data privacy. Despite its potential benefit, the increasing number of participating devices poses new challenges to the deployment in real-world cases. The highly limited amount of data located on each device coupled with significantly unbalanced data across different devices severely impede the performance of the shared model and the overall training progress at the same time.In this paper, we propose FedHybrid, a hierarchical hybrid training framework for high-performance Federated Learning on a wide scale. Unlike the existing work that mainly focuses on the statistical challenge, FedHybrid establishes a hierarchical hybrid training framework that effectively utilizes the fragmented and unbalanced data located on the participating devices on a wide scale. Specifically, FedHybrid consists of the following two core components, a global coordinator deployed on the central server and a local coordinator deployed on each participating device. The global coordinator organizes the participating devices into different groups through jointly considering the system heterogeneity and unbalanced training data in order to accelerate the overall training progress while guaranteeing the model performance. Within each group, a novel device-to-device (D2D) sequential training procedure is coordinated by the local coordinator to effectively utilize the fragmented and unbalanced training data in order to intelligently update the local models. At the same time, we provide the theoretical analysis of FedHybrid and conduct extensive experiments to evaluate its effectiveness. The results show that FedHybrid effectively improves model accuracy up to 27% and accelerates the whole training process by 20% on average. Tao Chang, Li Li 0064, Meihan Wu, Wei Yu 0029, Xiaodong Wang 0002 |
SECON | 2 |
| 2023 | PAGroup: Privacy-aware grouping framework for high-performance federated learning
Tao Chang, Li Li 0064, Meihan Wu, Wei Yu 0029, Xiaodong Wang 0002, Cheng-Zhong Xu 0001 |
J. Parallel Distributed Comput. | 2 |
| 2023 | GraphCS: Graph-based client selection for heterogeneity in federated learning
Tao Chang, Li Li 0064, Meihan Wu, Wei Yu 0029, Xiaodong Wang 0002, Cheng-Zhong Xu 0001 |
J. Parallel Distributed Comput. | 2 |
| 2023 | SmartDL: energy-aware decremental learning in a mobile-based federation for geo-spatial system
Wenting Zou, Li Li 0064, Zichen Xu 0001, Dan Wu 0010, Cheng-Zhong Xu 0001, Yuhao Wang 0001, Haoyang Zhu |
Neural Comput. Appl. | 2 |
| 2023 | A Framework for Behavioral Biometric Authentication Using Deep Metric Learning on Mobile DevicesabstractMobile authentication using behavioral biometrics has been an active area of research. Existing research relies on building machine learning classifiers to recognize an individual’s unique patterns. However, these classifiers are not powerful enough to learn the discriminative features. When implemented on the mobile devices, they face new challenges from the behavioral dynamics, data privacy and side-channel leaks. To address these challenges, we present a new framework to incorporate training on battery-powered mobile devices, so private data never leaves the device and training can be flexibly scheduled to adapt the behavioral patterns at runtime. We re-formulate the classification problem into deep metric learning to improve the discriminative power and design an effective countermeasure to thwart side-channel leaks by embedding a noise signature in the sensing signals without sacrificing too much usability. The experiments demonstrate authentication accuracy over 95 percent on three public datasets, a sheer 15 percent gain from multi-class classification with less data and robustness against brute-force and side-channel attacks with 99 and 90 percent success, respectively. We show the feasibility of training with mobile CPUs, where training 100 epochs takes less than 10 mins and can be boosted 3-5 times with feature transfer. Finally, we profile memory, energy and computational overhead. Our results indicate that training consumes lower energy than watching videos and slightly higher energy than playing games. Cong Wang 0006, Yanru Xiao, Xing Gao 0001, Li Li 0064, Jun Wang 0077 |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | AutoRS: Environment-Dependent Real-Time Scheduling for End-to-End Autonomous DrivingabstractThe rapid development of autonomous driving poses new research challenges for on-vehicle computing system. The execution time of autonomous driving tasks heavily depends on the driving environment. As the scene becomes complex, task execution time increases significantly, leading to end-to-end deadline misses and potential accidents. Hence, a framework that can effectively schedule tasks according to the driving environment in order to guarantee end-to-end deadlines is critical for autonomous driving. In this article, we propose AutoRS, an environment-dependent real-time scheduling framework for end-to-end autonomous driving. AutoRS consists of two nested control loops. The inner control loop schedules tasks based on the driving environment to help them meet end-to-end deadlines while prioritizing the responsiveness and throughput of control commands. The outer control loop tunes task rates based on schedulability to efficiently utilize system resources with an RL-based design. We conduct extensive experiments on both simulation and hardware testbeds using representative autonomous driving applications. The results demonstrate that AutoRS effectively improves the driving performance by$7.95\%-56.9\%$in different driving environments. AutoRS can significantly enhance the safety and reliability of autonomous driving systems by providing timely control commands in complex and dynamic driving environments while guaranteeing task deadlines. Jialiang Ma, Li Li 0064, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | FedCDR: Federated Cross-Domain Recommendation for Privacy-Preserving Rating PredictionabstractThe cold-start problem, faced when providing recommendations to newly joined users with no historical interaction record existing in the platform, is one of the most critical problems that negatively impact the performance of a recommendation system. Fortunately, cross-domain recommendation~(CDR) is a promising approach for solving this problem, which can exploit the knowledge of these users from source domains to provide recommendations in the target domain. However, this method requires that the central server has the interaction behaviour data in both domains of all the users, which prevents users from participating due to privacy issues. Meihan Wu, Li Li 0064, Chang Tao, Eric Rigall, Xiaodong Wang 0002, Cheng-Zhong Xu 0001 |
CIKM | 2 |
| 2022 | FedDC: Federated Learning with Non-IID Data via Local Drift Decoupling and CorrectionabstractFederated learning (FL) allows multiple clients to collectively train a high-performance global model without sharing their private data. However, the key challenge in federated learning is that the clients have significant statistical heterogeneity among their local data distributions, which would cause inconsistent optimized local models on the clientside. To address this fundamental dilemma, we propose a novel federated learning algorithm with local drift decoupling and correction (FedDC). Our FedDC only introduces lightweight modifications in the local training phase, in which each client utilizes an auxiliary local drift variable to track the gap between the local model parameter and the global model parameters. The key idea of FedDC is to utilize this learned local drift variable to bridge the gap, i.e., conducting consistency in parameter-level. The experiment results and analysis demonstrate that FedDC yields expediting convergence and better performance on various image classification tasks, robust in partial participation settings, non-iid data, and heterogeneous clients. Liang Gao 0001, Huazhu Fu, Li Li 0064, Yingwen Chen 0001, Ming Xu 0002, Cheng-Zhong Xu 0001 |
CVPR | 3 |
| 2022 | LoADPart: Load-Aware Dynamic Partition of Deep Neural Networks for Edge OffloadingabstractThe emerging edge computing technique provides support for the computation tasks that are delay-sensitive and compute-intensive, such as deep neural network inference, by offloading them from a user-end device to an edge server for fast execution. The increasing offloaded tasks on an edge server are gradually facing the contention of both the network and computation resources. The existing offloading approaches often partition the deep neural network at a place where the amount of data transmission is small to save network resource, but rarely consider the problem caused by computation resource shortage on the edge server. In this paper, we design LoADPart, a deep neural network offloading system. LoADPart can dynamically and jointly analyze both the available network bandwidth and the computation load of the edge server, and make proper decisions of deep neural network partition with a light-weighted algorithm, to minimize the end-to-end inference latency. We implement LoADPart for MindSpore, a deep learning framework supporting edge AI, and compare it with state-of-the-art solutions in the experiments on 6 deep neural networks. The results show that under the variation of server computation load, LoADPart can reduce the end-to-end latency by 14.2% on average and up to 32.3% in some specific cases. Hongzhou Liu, Wenli Zheng, Li Li 0064, Minyi Guo |
ICDCS | 3 |
| 2022 | HARMONY: Heterogeneity-Aware Hierarchical Management for Federated Learning SystemabstractFederated learning (FL) enables multiple devices to collaboratively train a shared model while preserving data privacy. However, despite its emerging applications in many areas, real-world deployment of on-device FL is challenging due to wildly diverse training capability and data distribution across heterogeneous edge devices, which highly impact both model performance and training efficiency. This paper proposes Harmony, a high-performance FL framework with heterogeneity-aware hierarchical management of training devices and training data. Unlike previous work that mainly focuses on heterogeneity in either training capability or data distribution, Harmony adopts a hierarchical structure to jointly handle both heterogeneities in a unified manner. Specifically, the two core components of Harmony are a global coordinator hosted by the central server and a local coordinator deployed on each participating device. Without accessing the raw data, the global coordinator first selects the participants, and then further reorganizes their training samples based on the accurate estimation of the runtime training capability and data distribution of each device. The local coordinator keeps monitoring the local training status and conducts efficient training with guidance from the global coordinator. We conduct extensive experiments to evaluate Harmony using both hardware and simulation testbeds on representative datasets. The experimental results show that Harmony improves the accuracy performance by 1.67% - 27.62%. In addition, Harmony effectively accelerates the training process up to $3.29\times$ and $1.84\times$ on average, and saves energy up to 88.41% and 28.04% on average. Chunlin Tian, Li Li 0064, Jun Wang 0001, Cheng-Zhong Xu 0001 |
MICRO | 2 |
| 2022 | FedGosp: A Novel Framework of Gossip Federated Learning for Data HeterogeneityabstractFederated learning (FL) provides the possibility to solve the problem of data privacy, but it suffers much from the data heterogeneity among different participants. Currently, some promising FL algorithms improve the effectiveness of learning under the non independent-and-identically-distributed (Non-IID) data settings. However, they require a large number of communication rounds between the server and clients for an acceptable accuracy. Inspired by the training paradigm of gossip learning, this paper proposes a new FL framework, named FedGosp. It first classifies the clients into different categories based on the model weights trained by the locally stored data. Then FedGosp utilizes the communication not only between clients and the server, but also between different classes of clients themselves. This training process enables instilling knowledge about various data distributions in the passed models. We evaluate the performance of FedGosp in multiple Non-IID settings on CIFAR10 and MNIST datasets, and compare it with the recently popular algorithms such as SCAFFOLD, FedAvg and FedProx. The experimental results show that FedGosp can improve the model accuracy by 6.53% and save 5.6 × communication costs at best compared to the second-ranked baseline. Yue Hu 0016, Miao Zhang 0037, Li Li 0064, Tao Chang, Quanjun Yin |
SMC | 4 |
| 2022 | Machine Learning in Real-Time Internet of Things (IoT) Systems: A SurveyabstractOver the last decade, machine learning (ML) and deep learning (DL) algorithms have significantly evolved and been employed in diverse applications, such as computer vision, natural language processing, automated speech recognition, etc. Real-time safety-critical embedded and Internet of Things (IoT) systems, such as autonomous driving systems, UAVs, drones, security robots, etc., heavily rely on ML/DL-based technologies, accelerated with the improvement of hardware technologies. The cost of a deadline (required time constraint) missed by ML/DL algorithms would be catastrophic in these safety-critical systems. However, ML/DL algorithm-based applications have more concerns about accuracy than strict time requirements. Accordingly, researchers from the real-time systems (RTSs) community address the strict timing requirements of ML/DL technologies to include in RTSs. This article will rigorously explore the state-of-the-art results emphasizing the strengths and weaknesses in ML/DL-based scheduling techniques, accuracy versus execution time tradeoff policies of ML algorithms, and security and privacy of learning-based algorithms in real-time IoT systems. Jiang Bian 0003, Abdullah Al Arafat, Haoyi Xiong, Jing Li 0025, Li Li 0064, Hongyang Chen 0001, Jun Wang 0001, Dejing Dou, Zhishan Guo |
IEEE Internet Things J. | 5 |
| 2022 | FGFL: A blockchain-based fair incentive governor for Federated Learning
Liang Gao 0001, Li Li 0064, Yingwen Chen 0001, Cheng-Zhong Xu 0001, Ming Xu 0002 |
J. Parallel Distributed Comput. | 2 |
| 2022 | Performance optimization of autonomous driving control under end-to-end deadlines
Yunhao Bai, Li Li 0064, Zejiang Wang, Junmin Wang 0002 |
Real Time Syst. | 2 |
| 2022 | Investigating Security Vulnerabilities in a Hot Data Center with Reduced Cooling RedundancyabstractData centers have been growing rapidly in recent years to meet the surging demand of cloud services. However, the expanding scale and powerful servers generate a great amount of heat, resulting in significant cooling costs. A trend in modern data centers is to raise the temperature and maintain all servers in a relatively hot environment. While this can save on cooling costs given benign workloads running in servers, the hot environment increases the risk of a cooling failure. In this article, we unveil a new vulnerability of existing data centers with aggressive cooling energy saving policies. Such a vulnerability might be exploited to launch thermal attacks that could severely worsen the thermal conditions in a data center. Specifically, we conduct thermal measurements and uncover effective thermal attack vectors at the server, rack, and data center levels. We also present damage assessments of thermal attacks. Our results demonstrate that thermal attacks can (1) largely increase the temperature of victim servers degrading their performance and reliability, (2) negatively impact on thermal conditions of neighboring servers causing local hotspots, (3) raise the cooling cost, and (4) even lead to cooling failures. Finally, we propose and evaluate effective server and data center level defenses to enhance thermal stabilities. Xing Gao 0001, Guannan Liu 0003, Zhang Xu, Haining Wang 0001, Li Li 0064 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2021 | FIFL: A Fair Incentive Mechanism for Federated LearningabstractFederated learning is a novel machine learning framework that enables multiple devices to collaboratively train high-performance models while preserving data privacy. Federated learning is a kind of crowdsourcing computing, where a task publisher shares profit with workers to utilize their data and computing resources. Intuitively, devices have no interest to participate in training without rewards that match their expended resources. In addition, guarding against malicious workers is also essential because they may upload meaningless updates to get undeserving rewards or damage the global model. In order to effectively solve these problems, we propose FIFL, a fair incentive mechanism for federated learning. FIFL rewards workers fairly to attract reliable and efficient ones while punishing and eliminating the malicious ones based on a dynamic real-time worker assessment mechanism. We evaluate the effectiveness of FIFL through theoretical analysis and comprehensive experiments. The evaluation results show that FIFL fairly distributes rewards according to workers’ behaviour and quality. FIFL increases the system revenue by 0.2% to 3.4% in reliable federations compared with baselines. In the unreliable scenario containing attackers which destroy the model’s performance, the system revenue of FIFL outperforms the baselines by more than 46.7%. Liang Gao 0001, Li Li 0064, Yingwen Chen 0001, Wenli Zheng, Cheng-Zhong Xu 0001, Ming Xu 0002 |
ICPP | 2 |
| 2021 | SmartDistance: A Mobile-based Positioning System for Automatically Monitoring Social DistanceabstractCoronavirus disease 2019 (COVID-19) has resulted in an ongoing pandemic. Since COVID-19 spreads mainly via close contact among people, social distancing has become an effective manner to slow down the spread. However, completely forbidding close contact can also lead to unacceptable damage to the society. Thus, a system that can effectively monitor people's social distance and generate corresponding alerts when a high infection probability is detected is in urgent need. In this paper, we propose SmartDistance, a smartphone based software framework that monitors people's interaction in an effective manner, and generates a reminder whenever the infection probability is high. Specifically, SmartDistance dynamically senses both the relative distance and orientation during social interaction with a well-designed relative positioning system. In addition, it recognizes different events (e.g., speaking, coughing) and determines the infection space through a droplet transmission model. With event recognition and relative positioning, SmartDistance effectively detects risky social interaction, generates an alert immediately, and records the relevant data for close contact reporting. We prototype SmartDistance on different Android smartphones, and the evaluation shows it reduces the false positive rate from 33% to 1% and the false negative rate from 5% to 3% in infection risk detection. Li Li 0064, Wenli Zheng, Cheng-Zhong Xu 0001 |
INFOCOM | 1 |
| 2021 | Dynamic traffic bottlenecks identification based on congestion diffusion model by influence maximization in metro-city scalesabstractSummary Traffic bottlenecks dynamically change with the variance of traffic demand. Identifying traffic bottlenecks plays an important role in traffic planning and provides decision making. However, traffic bottlenecks are difficult to identify because of the complexity of traffic road networks and many other factors. In this article, we propose an influence spreading based method to find the dynamic changed traffic bottlenecks, where the influence caused by bottlenecks is maximal. We first build a traffic congestion diffusion (TCD) model to capture traffic flow influence (TFI) spreading over traffic road networks. The bottlenecks identification problem based on TCD is modeled as an influence maximization problem, that is, selecting the most influential nodes such that the deterioration of traffic condition is maximal. With the proof of the submodularity of TFI spreading over traffic networks, a provably near‐optimal algorithm is used to solve the NP‐hard problem. With the exploration of unique properties of TFI spread, an approximate influence maximization method for TCD (TCD‐AIM) is proposed. To the best of our knowledge, this should be the first model for a metro‐city scale from the influence perspective. Experimental results show that TCD‐AIM finds bottlenecks with up to 130% congestion density increase in the future. Baoxin Zhao, Cheng-Zhong Xu 0001, Siyuan Liu 0001, Juanjuan Zhao 0001, Li Li 0064 |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | Multi-layer Coordination for High-Performance Energy-Efficient Federated LearningabstractFederated Learning is designed for multiple mobile devices to collaboratively train an artificial intelligence model while preserving data privacy. Instead of collecting the raw training data from mobile devices to the cloud, Federated Learning coordinates a group of devices to train a shared model in a distributed manner with the training data located on the devices. However, in order to effectively deploy Federated Learning on resource-constrained mobile devices, several critical issues including convergence rate, scalability and energy efficiency should be well addressed. In this paper, we propose MCFL, a multi-layer online coordination framework for high-performance energy efficient federated learning. MCFL consists of two layers: a macro-layer on the central server and a micro-layer on each participating device. In each training round, the macro coordinator performs two tasks, namely, selecting the right devices to participate, and estimating a time limit, such that the overall training time is significantly reduced while still guaranteeing the model accuracy. Unlike existing systems, MCFL removes the restriction that participating devices must be connected to power sources, thus allowing more timely and ubiquitous training. This clearly requires on-device training to be highly energy-efficient. To this end, the micro coordinator determines optimal schedules for hardware resources in order to meet the time limit set by the macro coordinator with the least amount of energy consumption. Tested on real devices as well as simulation testbed, MCFL has shown to be able to effectively balance the convergence rate, model accuracy and energy efficiency. Compared with existing systems, MCFL can achieve a speedup up to 8.66× and reduce energy consumption by up to 76.5% during the training process. Li Li 0064, Jun Wang 0001, Cheng-Zhong Xu 0001 |
IWQoS | 1 |
| 2019 | A Congestion Diffusion Model with Influence Maximization for Traffic Bottlenecks Identification in Metrocity ScalesabstractTraffic bottlenecks identification plays an important role in traffic planning and provides decision-making for prevention of traffic congestion. Although traffic bottlenecks widely exist, they are difficult to predict because of the changing traffic condition and traffic demand. In this paper, we introduce a traffic congestion diffusion (TCD) model with traffic flow influence (TFI) to capture the traffic dynamics and give a panoramic view for the city by cross domain data fusion. We proposed novel definition of bottleneck from the perspective of influence spread under TCD. The bottlenecks identification problem is modeled as an influence maximization problem, i.e., selecting the top K influential nodes in road networks under certain traffic conditions. We establish the submodularity of influence spread and solve the NP-hard optimal seed selection problem by using an efficient heuristic algorithm (TCD-IM) with provable near-optimal performance guarantees. To the best of our knowledge, this should be the first model for a metro-city scale from the influence perspective. The TCD-IM model is able to identify the dynamic traffic bottlenecks. Baoxin Zhao, Cheng-Zhong Xu 0001, Siyuan Liu 0001, Juanjuan Zhao 0001, Li Li 0064 |
IEEE BigData | 5 |
| 2019 | Close the Gap between Deep Learning and Mobile Intelligence by Incorporating Training in the LoopabstractPre-trained deep learning models can be deployed on mobile devices to conduct inference. However, they are usually not updated thereafter. In this paper, we take a step further to incorporate training deep neural networks on battery-powered mobile devices and overcome the difficulties from the lack of labeled data. We design and implement a new framework to enlarge sample space via data paring and learn a deep metric under the privacy, memory and computational constraints. A case study of deep behavioral authentication is conducted. Our experiments demonstrate accuracy over 95% on three public datasets, a sheer 15% gain from traditional multi-class classification with less data and robustness against brute-force attacks with 99% success. We demonstrate the training performance on various smartphone models, where training 100 epochs takes less than 10 mins and can be boosted 3-5 times with feature transfer. We also profile memory, energy and computational overhead. Our results indicate that training consumes lower energy than watching videos so can be scheduled intermittently on mobile devices. Cong Wang 0006, Yanru Xiao, Xing Gao 0001, Li Li 0064, Jun Wang 0077 |
ACM Multimedia | 4 |
| 2019 | SmartPC: Hierarchical Pace Control in Real-Time Federated Learning SystemabstractFederated Learning is a technique for learning AI models through the collaboration of a large number of resourceconstrained mobile devices, while preserving data privacy. Instead of aggregating the training data from devices, Federated Learning uses multiple rounds of parameter aggregation to train a model, wherein the participating devices are coordinated to incrementally update a shared model with their own parameters locally learned. To efficiently deploy Federated Learning system over mobile devices, several critical issues including realtimeliness and energy efficiency should be well addressed. This paper proposes SmartPC, a hierarchical online pace control framework for Federated Learning that balances the training time and model accuracy in an energy-efficient manner. SmartPC consists of two layers of pace control: global and local. Prior to every training round, the global controller first oversees the status (e.g., connectivity, availability, and energy/resource remained) of every participating device, then selects qualified devices and assigns them a well-estimated virtual deadline for task completion. Within such virtual deadline, a statistically significant proportion (e.g., 60%) of the devices are expected to complete one round of their local training and model updates, while the overall progress of multi-round training procedure is kept up adaptively. On each device, a local pace controller then dynamically adjusts device settings such as CPU frequency so that the learning task is able to meet the deadline with the least amount of energy consumption. We performed extensive experiments to evaluate SmartPC on both Android smartphones and simulation platforms using well-known datasets. The experiment results show that SmartPC reduces up to 32:8% energy consumption on mobile devices and achieves a speedup of 2.27 in training time without model accuracy degradation. Li Li 0064, Haoyi Xiong, Zhishan Guo, Jun Wang 0001, Cheng-Zhong Xu 0001 |
RTSS | 1 |
| 2018 | Reduced Cooling Redundancy: A New Security Vulnerability in a Hot Data Center
Xing Gao 0001, Zhang Xu, Haining Wang 0001, Li Li 0064 |
NDSS | 4 |
| 2017 | SceneMan: Bridging mobile apps with system energy manager via scenario notificationabstractPower management on current mobile devices relies on OS modules known as DVFS governors. However, existing governors determine system configuration only based on low-level information such as CPU load without any input about application-level behaviors. In particular, there exists no communication from mobile apps to energy managers. We find that information about app usage scenarios (e.g., gaming, video chatting) can usually help energy manager perform a better job and achieve more energy savings. Although app-level energy optimizations have been proposed, they generally focus on single usage scenarios and do not address optimization across multiple scenarios. In this paper, we propose SceneMan, an energy optimization framework for mobile apps based on usage scenario notification. SceneMan has three components: an API, a scenario notifier, and an energy manager. The key idea is to make energy managers aware of app-level scenarios. At runtime, apps notify the energy manager about their usage scenarios with provided APIs used by developers. The energy manager then takes appropriate actions to minimize energy consumption of the running scenario while meeting performance requirements. Energy optimization across scenarios can thus be easily achieved. The framework requires little extra programming effort and can help apps achieve better energy efficiency in a transparent way. We implement our system on a Nexus 6 smartphone and test it with 13 real-world apps under 2 usage scenarios, namely, gaming and video chatting. We achieve up to 33.2% energy savings with a worst-case performance loss of 5.1%. Li Li 0064, Jun Wang 0077, Handong Ye, Ziang Hu |
ISLPED | 1 |
| 2017 | PowerNetS: Coordinating Data Center Network With Servers and Cooling for Power OptimizationabstractRecently, a lot of research efforts have been made to optimize the large amounts of energy consumed by different devices in data centers, including servers, cooling, and the data center network (DCN). Unfortunately, current research addresses these devices mostly in a separate manner, leading to inferior optimization results. This paper proposes PowerNetS, a power optimization framework that coordinates servers and DCN, as well as cooling, for minimized power consumption of a data center. PowerNetS leverages workload correlation analysis for more energy savings during server and traffic consolidations. More importantly, PowerNetS tries to change the DCN topology during server consolidation, in order to have more intra-server traffic and shorter flows that go through fewer switches. For example, two virtual machines previously located on two different servers can now be migrated to the same server, so that the flow between them no longer needs to use switches, which allows more devices to sleep for energy savings without network performance degradation. PowerNetS has been implemented on a physical testbed with 6 servers and 10 virtual switches that are configured using a production 48-port OpenFlow switch. Our evaluation with Wikipedia, Yahoo!, and IBM traces shows that PowerNetS can save up to 51.6% of energy by coordinating servers and DCN, which is 44.3% and 15.8% more than two state-of-the-art baselines, respectively. By further coordinating with cooling to utilize different cooling efficiencies at different locations within a data center, PowerNetS can achieve 8.8%-14.6% additional energy savings. Kuangyu Zheng, Wenli Zheng, Li Li 0064 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2014 | Joint power optimization of data center network and servers with correlation analysisabstractData center power optimization has recently received a great deal of research attention. For example, server consolidation has been demonstrated as one of the most effective energy saving methodologies. Likewise, traffic consolidation has also been recently proposed to save energy for data center networks (DCNs). However, current research on data center power optimization focuses on servers and DCN separately. As a result, the optimization results are often inferior, because server consolidation without considering the DCN may cause traffic congestion and thus degraded network performance. On the other hand, server consolidation may change the DCN topology, allowing new opportunities for energy savings. In this paper, we propose PowerNetS, a power optimization strategy that leverages workload correlation analysis to jointly minimize the total power consumption of servers and the DCN. The design of PowerNetS is based on the key observations that the workloads of different servers and DCN traffic flows do not peak at exactly the same time. Thus, more energy savings can be achieved if the workload correlations are considered in server and traffic consolidations. In addition, PowerNetS considers the DCN topology during server consolidation, which leads to less inter-server traffic and thus more energy savings and shorter network delays. We implement PowerNetS on a hardware testbed composed of 10 virtual switches configured with a production 48-port OpenFlow switch and 6 servers. Our empirical results with Wikipedia, Yahoo!, and IBM traces demonstrate that PowerNetS can save up to 51.6% of energy for a data center. PowerNetS also outperforms two state-of-the-art baselines by 44.3% and 15.8% on energy savings, respectively. Our simulation results with 72 switches and 122 servers also show the superior energy efficiency of PowerNetS over the baselines. Kuangyu Zheng, Xiaodong Wang 0007, Li Li 0064 |
INFOCOM | 3 |