Zongwei Zhu

dblp:83/11300 · DBLP profile ↗
← Back
49ranked-venue papers
7as first author
34since 2021 · last 2026
0000-0003-3607-2631ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 33 · 1 first-author · 24 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 PPFL: A Parameter Behavior-Driven Plug-in Personalization Engine for Federated Learning
abstract
Personalized Federated Learning (PFL) customizes models for each client to mitigate challenges from non-IID data, wherein a dominant strategy is model decoupling that partitions models into shared and personalized parts based on architectural priors (e.g., backbone vs. head). However, we reveal a critical flaw in this strategy: it induces "intrinsic drift," a performance degradation often more severe than the well-known client drift, which limits final accuracy. We trace this drift to a steep cliff of high loss emerging from the naive stitching of shared and personalized parts. To address this, we shift from architectural partitioning to a parameter behavior-driven paradigm. We introduce PPFL, an approach that employs a novel soft-fusion strategy guided by parameter-wise behavioral perception. PPFL dynamically infers each parameter's functional role—whether it behaves more like a 'personalist' or a 'generalist' in the current context—by synthesizing its multifaceted behavior observed during local training. Extensive experiments on image, text, and multimodal classification benchmarks show that PPFL outperforms eight state-of-the-art baselines by up to 5.3%. Moreover, it can function as a plug-in module, boosting the accuracy of vanilla FedAvg with a 16.82% absolute gain.
Qianyue Cao, Zongwei Zhu, Zirui Lian, Rui Zhang 0040, Boyu Li 0006, Yi Xiong 0003, Xuehai Zhou
AAAI2
2026 FedGAMA: Federated Learning on Heterogeneous and Long-Tailed Data via Group-Wise Asymmetric Masked Aggregation
Chenyue Xu, Zongwei Zhu, Qianyue Cao, Rui Zhang 0040, Xuehai Zhou
KSEM (1)2
2026 CSCL: Bridging the plasticity-stability gap in continuous supervised contrastive learning
Yi Xiong 0003, Liqi Xiang, Qianyue Cao, Zongwei Zhu, Zirui Lian, Xuehai Zhou
Neural Networks4
2026 MultiLens: A Multiobjective Adaptive DVFS Framework for Energy-Efficient DNN Inference
abstract
To tackle power management challenges in deep neural networks (DNNs), dynamic voltage and frequency scaling (DVFS) has gained attention for its ability to enhance energy efficiency without modifying DNN structures. However, current DVFS methods, which rely on historical data such as processor utilization and task load, suffer from issues like frequency ping-pong, response lag, and limited generalizability. These challenges are exacerbated by real-world scenarios that prioritize time, energy, or energy efficiency differently, making it even harder for existing methods to effectively configure DVFS under such multi-objective constraints or trade-offs. This paper presents MultiLens, a multi-objective adaptive DVFS framework. First, we propose a power-sensitive feature extraction method along with multi-objective constraint modeling to characterize DNN inference behavior. Second, critical power blocks are then identified through clustering based on inference behavior similarity, enabling adaptive DVFS instrumentation point settings. Moreover, to enhance the adaptability of multiple platforms and the flexibility of multiple scenarios, MultiLens integrates a complete deployment process. Experimental results demonstrate the effectiveness of the MultiLens in optimizing energy efficiency across different hardware platforms and deployment scenarios.
Jiawei Geng, Zongwei Zhu, Weihong Liu, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 AsyncGrid: An Intralayer and Interlayer Asynchronous Hybrid Parallelism System for Responsive Edge LLM Inference
abstract
Edge deployment of large language models (LLMs) is increasingly attractive due to its advantages in privacy, customization, and availability. However, edge environments face significant challenges in reducing Time-to-First-Token (TTFT). TTFT consists of (1) queuing delay and (2) prefill latency, both of which are exacerbated by edge‑resource constraints: the substantial computational demands of LLM inference grow superlinearly with prompt length, causing high prefill latency; and limited edge resources restrict prefill throughput, preventing the timely handling of incoming requests, thereby exacerbating queuing delays. Model parallelism is a commonly used solution in cloud-based systems, but directly applying it to edge environments proves ineffective. Intra-layer parallelism (e.g., tensor/sequence parallelism) can reduce prefill latency but suffers from frequent global synchronization, which bottlenecks prefill throughput due to edge-limited interconnection bandwidth. Inter-layer parallelism (e.g., pipeline parallelism) improves prefill throughput via fully asynchronous execution but retains high prefill latency due to stage-wise serialized computation. To address this dilemma, this paper leverages the properties of the causal attention mechanism in LLMs and proposes Intra-layer Asynchronous Parallelism (IAP), which performs intra-layer parallel computations to reduce prefill latency while avoiding global synchronization to mitigate prefill throughput bottlenecks. Moreover, considering communication sensitivity in intra-layer parallelism, this paper integrate IAP with inter-layer asynchronous parallelism into a unified plan space. This hybrid parallelism adapts to diverse hardware and request loads, enabling more effective TTFT optimization. To enable the end-to-end implementation of this hybrid parallelism, this paper propose AsyncGrid, an LLM inference system tailored for responsive edge LLM inference. AsyncGrid (1) models runtime overheads through a performance profiler, (2) employs an integer programming (IP) formulation to optimize execution plan, with the objective of minimizing latency while meeting throughput requirements, and (3) implements fine-grained communication optimization during runtime. A comprehensive evaluation on an edge testbed demonstrates AsyncGrid’s significant advantages over existing methods, achieving substantial improvements in both homogeneous and heterogeneous settings.
Yi Xiong 0003, Rui Zhang 0040, Yulong Zu, Weihong Liu, Zongwei Zhu, Jiawei Geng, Boyu Li 0006, Qianyue Cao, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 MedFS: Pursuing Low Update Overhead via Metadata-Enabled Delta Compression for Log-structured File System on Mobile Device
Chao Wu 0006, Cheng Ji 0002, Li-Pin Chang, Zongwei Zhu, Congming Gao, Weichao Guo, Yanzhi Wang 0001
FAST4
2025 Archer: Adaptive Memory Compression with Page-Association-Rule Awareness for High-Speed Response of Mobile Devices
Changlong Li 0006, Zongwei Zhu, Chao Wang 0003, Fangming Liu, Edwin H.-M. Sha, Xuehai Zhou
FAST2
2025 HeterScale: A hierarchical task scheduling framework for intelligent edge collaboration in IIOT
Weihong Liu, Zongwei Zhu, Yulong Zu
J. Syst. Archit.3
2025 Magnifier: A Chiplet Feature-Aware Test Case Generation Method for Deep Learning Accelerators
abstract
The development of deep learning has led to increasing demands for computation and memory, making multi-chiplet accelerators a powerful solution. Multi-chiplet accelerators require more precise consideration of hardware configurations and mapping schemes in terms of computation, memory, and communication patterns compared to monolithic designs, in order to avoid underutilization of performance. However, there is currently a lack of performance testing methods specifically tailored for multi-chiplet accelerators. Existing testing methods primarily focus on correctness testing and do not address potential performance issues from a hardware perspective. To address these issues, this paper proposes Magnifier: a test case generation method for performance testing of multi-chiplet accelerators. Firstly, we analyze typical multi-chiplet accelerator prototype from the perspectives of computation, memory, and communication patterns, and summarize a chiplet feature-aware operator task set. Next, we define the test evaluation metric IPPstd and use a candidate operator set to construct a sampling space for model-level test cases. Finally, we build a GAN to learn the distribution of high-diversity test cases, enabling the rapid generation of high-quality test cases. We validate the proposed method on both simulated and real multi-chiplet accelerators. Experiments show that Magnifier can improve the metric of test cases by up to 3.42 times and significantly reduce generation time, providing valuable insights for optimizing the hardware and software of multi-chiplet accelerators.
Boyu Li 0006, Zongwei Zhu, Weihong Liu, Qianyue Cao, Changlong Li 0006, Cheng Ji 0002, Xi Li 0003, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 HaloFL: Efficient Heterogeneity-Aware Federated Learning Through Optimal Submodel Extraction and Dynamic Sparse Adjustment
abstract
Federated learning (FL) is an advanced framework that enables collaborative training of machine learning models across edge devices. An effective strategy to enhance training efficiency is to allocate the optimal submodel based on each device’s resource capabilities. However, system heterogeneity significantly increases the difficulty of allocating submodel parameter budgets appropriately for each device, leading to the straggler problem. Meanwhile, data heterogeneity complicates the selection of the optimal submodel structure for specific devices, thereby impacting training performance. Furthermore, the dynamic nature of edge environments, such as fluctuations in network communication and computational resources, exacerbates these challenges, making it even more difficult to precisely extract appropriately sized and structured submodels from the global model. To address the challenges in heterogeneous training environments, we propose an efficient FL framework, namely, HaloFL. The framework dynamically adjusts the structure and parameter budget of submodels during training by evaluating three dimensions: 1) model-wise performance; 2) layer-wise performance; and 3) unit-wise performance. First, we design a data-aware model unit importance evaluation method to determine the optimal submodel structure for different data distributions. Next, using this evaluation method, we analyze the importance of model layers and reallocate parameters from noncritical layers to critical layers within a fixed parameter budget, further optimizing the submodel structure. Finally, we introduce a resource-aware dual-UCB multiarmed bandit agent, which dynamically adjusts the total parameter budget of submodels according to changes in the training environment, allowing the framework to better adapt to the performance differences of heterogeneous devices. Experimental results demonstrate that HaloFL exhibits outstanding efficiency in various dynamic and heterogeneous scenarios, achieving up to a 14.80% improvement in accuracy and a$3.06\times $speedup compared to existing FL frameworks.
Zirui Lian, Qianyue Cao, Zongwei Zhu, Cheng Ji 0002, Changlong Li 0006, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2025 Freezing-based Memory and Process Co-design for User Experience on Resource-limited Mobile Devices
abstract
Mobile devices with limited resources are prevalent, as they have a relatively low price. Providing a good user experience with limited resources has been a big challenge. This work finds that foreground applications are often unexpectedly interfered by background applications’ memory activities. Improving user experience on resource-limited mobile devices calls for a strong collaboration between memory and process management. This article proposes Ice , a framework to optimize the user experience on resource-limited mobile devices. With Ice, processes that will cause frequent refaults in the background are identified and frozen accordingly. The frozen application will be thawed when memory condition allows. Based on the proposed Ice, this work shows that the refault can be further reduced by revisiting the LRU lists in the original kernel with app-freezing awareness (called Ice + ). Evaluation of resource-limited mobile devices demonstrates that the user experience is effectively improved with Ice. Specifically, Ice boosts the frame rate by 1.57x on average over the state of the art. The frame rate is further enhanced by 5.14% on average with Ice + .
Changlong Li 0006, Zongwei Zhu, Chun Jason Xue, Yu Liang 0004, Rachata Ausavarungnirun, Liang Shi 0001, Xuehai Zhou
ACM Trans. Comput. Syst.2
2025 A Lightweight I/O Throttling Service to Improve the User Experience of Mobile Devices
abstract
As one of the most frequently occurring operations, I/Os significantly affect the application launching time and frame rate of mobile devices, hence influencing the user experience. However, the response speed of I/O requests is still the bottleneck in practice. This paper shows that high I/O latency is usually due to the congestion inside Flash instead of the system layer. Unfortunately, Flash is treated as a black box device and cannot be modified after delivery. In this paper, we propose a novel service to address this issue without an intra-Flash modification. Specifically, this paper proposes a lightweight I/O throttling framework in mobile systems named FlashDAM. This service throttles the I/O flow to make way for I/Os that may block the foreground application. FlashDAM is the first work that proves that proper I/O throttling positively affects the user experience, contrary to the common belief. Furthermore, this paper proposes FlashDAM$^+$, an enhanced version of FlashDAM. By coordinating I/O throttling and compression, FlashDAM's effect in the system layer is minimized. We have implemented FlashDAM on real mobile devices. Experimental results illustrate that the app launching speed and frame rate are enhanced by 72% and 45% separately compared to the state-of-the-art. When enabling the compression feature of FlashDAM, that is, FlashDAM$^+$, screen jank and application launch latency are further reduced by 9.5% and 11.4%, respectively, under heavy background I/O load.
Changlong Li 0006, Zongwei Zhu, Yuyangjun Lu, Chao Wang 0003, Xuehai Zhou, Edwin H.-M. Sha
IEEE Trans. Serv. Comput.2
2024 PowerLens: An Adaptive DVFS Framework for Optimizing Energy Efficiency in Deep Neural Networks
abstract
To address the power management challenges in deep neural networks (DNNs), dynamic voltage and frequency scaling (DVFS) technology is garnering attention for its ability to enhance energy efficiency without modifying the structure of DNNs. However, current DVFS methods, which depend on historical information such as processor utilization and task computational load, face issues like frequency ping-pong, response lag, and poor generalizability. Therefore, this paper introduces PowerLens, an adaptive DVFS framework. Initially, we develop a power-sensitive feature extraction method for DNNs and identify critical power blocks through clustering based on power behavior similarity, thereby achieving adaptive DVFS instrumentation point settings. Then, the framework adaptively presets the target frequency for each power block through a decision model. Finally, through a refined training and deployment process, we ensure the framework's effective adaptability across different platforms. Experimental results confirm the effectiveness of the framework in energy efficiency optimization.
Jiawei Geng, Zongwei Zhu, Weihong Liu, Xuehai Zhou, Boyu Li 0006
DAC2
2024 EPipe: Pipeline Inference Framework with High-quality Offline Parallelism Planning for Heterogeneous Edge Devices
abstract
Pipeline parallelism is essential for edge computing as it effectively consolidates the limited resources of edge devices, enabling the deployment of large Deep Neural Network (DNN) models and accelerating inference processes without compromising the performance of models. Accurate computation and communication latency estimation on heterogeneous edge devices is essential for searching for a superior parallelism plan. However, existing heterogeneous pipeline inference approaches either incur substantial resource wastage during online parallelism planning, as they utilize profiling strategies that occupy physical devices; or rely on cost models with inadequate representational capabilities, leading to inaccurate predictions, thereby harming the result of pipeline planning. This paper proposes EPipe, a novel pipeline inference framework that supports high-quality offline planning in heterogeneous edge environments. EPipe integrates two core components: the Task-Device Co-analyzer (TDC) and the Multi-pipeline Parallelism Planner (MPP). TDC utilizes an undirected connected graph to depict the compatibility of DNNs across device groups and precisely estimates inference and communication latencies through fine-grained modeling. Based on TDC, MPP utilizes a dynamic programming-based genetic algorithm to explore multi-pipeline solutions, extending beyond traditional single-pipeline methods. A comprehensive experimental evaluation on an edge testbed confirms the effectiveness of EPipe, demonstrating significant speedups in inference tasks for both task streams and single tasks.
Yi Xiong 0003, Weihong Liu, Rui Zhang 0040, Yulong Zu, Zongwei Zhu, Xuehai Zhou
ICCAD5
2024 GOFL: An Accurate and Efficient Federated Learning Framework Based on Gradient Optimization in Heterogeneous IoT Systems
abstract
Federated learning (FL) is designed for training models using data distributed across multiple Internet of Things (IoT) devices or servers, reducing data transfer overhead and ensuring data security. However, the decentralization and diversity of IoT devices introduce statistical and system heterogeneity, which can lead to unstable model training and even system crashes. Although many studies attribute performance issues to client-drift caused by this heterogeneity, there is a lack of insight into how different forms of heterogeneity impact local model gradient variations and model convergence. In this article, we investigate model gradient distribution characteristics in heterogeneous training. We find that the challenge is not solely due to client-drift but is also closely linked to a high degree of model overfitting, which negatively affects local model training and equilibrium convergence. To address this challenge, we introduce an efficient framework called gradient optimization with FL (GOFL). First, GOFL incorporates the federated gradient normalization (FGN) technique to maintain gradient distribution consistency while mitigating client-drift stemming from heterogeneity. We also highlight the benefits of FGN in reducing local model overfitting and improving convergence. Second, GOFL introduces the federated device aggregation (FDA) strategy, a critical addition to FGN. It adaptively guides device selection and aggregation based on device contributions, ensuring a more balanced training approach in the face of system heterogeneity. The experimental results demonstrate that GOFL achieves state-of-the-art training accuracy while reducing the number of training rounds. In particular, it improves the accuracy of the classical FL framework FedAvg by 30.57% and reduces the number of convergence rounds by 5.17 times.
Zirui Lian, Zongwei Zhu, Xuehai Zhou, Weihong Liu
IEEE Internet Things J.3
2024 FedStar: Efficient Federated Learning on Heterogeneous Communication Networks
abstract
The proliferation of multi-media applications and increased computing power of mobile devices have led to the development of personalized artificial intelligent (AI) applications that utilize the massive user-information residing on them. However, the traditional centralized training paradigm is not applicable in this scenario due to potential privacy risks and high communication overhead. Federated learning (FL) provides an option to these applications. Nevertheless, the heterogeneity of computing and communication latency among devices have posed great challenges to building efficient learning frameworks. Existing optimizations on FL either fail to speed up training on heterogeneous devices or suffer from poor communication efficiency. In this paper, we propose FedStar, an efficient FL framework that supports decentralized asynchronous training on heterogeneous communication networks. Considering the heterogeneous computing power in the network, FedStar supports running heterogeneity-aware local steps on each device. What’s more, considering the heterogeneous communication latency and possibly unreachable communication path between some devices, FedStar generates a decentralized communication topology that can achieve maximal training throughput. Finally, it adopts weighted aggregation to guarantee high convergence accuracy of global model. Theoretical analysis results show the convergence behaviour of FedStar under non-convex settings. Experimental results show that FedStar can achieve a speedup of 4.81× than the state-of-the-art FL schemes with high convergence accuracy.
Qianyue Cao, Yongchun Zheng, Zongwei Zhu, Cheng Ji 0002, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 NebulaFL: Self-Organizing Efficient Multilayer Federated Learning Framework With Adaptive Load Tuning in Heterogeneous Edge Systems
abstract
As a promising edge intelligence technology, federated learning (FL) enables Internet of Things (IoT) devices to train the models collaboratively while ensuring the data privacy and security. Recently, hierarchical FL (HFL) has been designed to promote distributed training in the intricate hierarchical structure of IoT. However, the coarse-grained hierarchical schemes usually fail to thoroughly adapt to the hierarchical environment, leading to high training latency. Meanwhile, highly heterogeneous communication and computation delays due to the device diversity (the system heterogeneity) and decentralized data distribution due to the decentralized device distribution (the data heterogeneity) exacerbate the above challenges. This article proposes NebulaFL, a dual heterogeneity-aware multilayer FL framework, to support efficient distributed training in IoT scenarios. NebulaFL proposes an innovative multilayer architecture organization scheme to adapt the complex hierarchical heterogeneous scenarios. Specifically, through a finer-grained division of the HFL hierarchy, hybrid synchronous-asynchronous training is implemented at both the global system and local device-layer levels. More importantly, to adaptively build a heterogeneity-aware hierarchical training architecture, NebulaFL considers the effect of dual heterogeneity in the architectural organization scheme to determine the optimal location of devices in a multilayer environment. To further improve the training efficiency during the training process, NebulaFL employs an augmented multiarmed bandit technique based on the reinforcement learning to adjust the device-layer training load by evaluating the dynamic training utility and convergence uncertainty feedback. Experiments demonstrate that NebulaFL achieves up to a$15.68\times $speed-up ratio and a 23.94% increase in the training accuracy compared to the latest or classic approaches.
Zirui Lian, Qianyue Cao, Weihong Liu, Zongwei Zhu, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Ace-Sniper: Cloud-Edge Collaborative Scheduling Framework With DNN Inference Latency Modeling on Heterogeneous Devices
abstract
The cloud–edge collaborative inference requires efficient scheduling of artificial intelligence (AI) tasks to the appropriate edge intelligence devices. Gls DNN inference latency has become a vital basis for improving scheduling efficiency. However, edge devices exhibit highly heterogeneous due to the differences in hardware architectures, computing power, etc. Meanwhile, the diverse deep neural networks (DNNs) are continuing to iterate over time. The diversity of devices and DNNs introduces high computational costs for measurement methods, while invasive prediction methods face significant development efforts and application limitations. In this article, we propose and develop Ace-Sniper, a scheduling framework with DNN inference latency modeling on heterogeneous devices. First, to address the device heterogeneity, a unified hardware resource modeling (HRM) is designed by considering the platforms as black-box functions that output feature vectors. Second, neural network similarity (NNS) is introduced for feature extraction of diverse and frequently iterated DNNs. Finally, with the results of HRM and NNS as input, the performance characterization network is designed to predict the latencies of the given unseen DNNs on heterogeneous devices, which can be combined into most time-based scheduling algorithms. Experimental results show that the average relative error of DNN inference latency prediction is 11.11%, and the prediction accuracy reaches 93.2%. Compared with the nontime-aware scheduling methods, the average waiting time for tasks is reduced by 82.95%, and the platform throughput is improved by 63% on average.
Weihong Liu, Jiawei Geng, Zongwei Zhu, Cheng Ji 0002, Changlong Li 0006, Zirui Lian, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 Arch2End: Two-Stage Unified System-Level Modeling for Heterogeneous Intelligent Devices
abstract
The surge in intelligent edge computing has propelled the adoption and expansion of the distributed embedded systems (DESs). Numerous scheduling strategies are introduced to improve the DES throughput, such as latency-aware and group-based hierarchical scheduling. Effective device modeling can help in modular and plug-in scheduler design. For uniformity in scheduling interfaces, an unified device performance modeling is adopted, typically involving the system-level modeling that incorporates both the hardware and software stacks, broadly divided into two categories. Fine-grained modeling methods based on the hardware architecture analysis become very difficult when dealing with a large number of heterogeneous devices, mainly because much architecture information is closed-source and costly to analyse. Coarse-grained methods are based on the limited architecture information or benchmark models, resulting in insufficient generalization in the complex inference performance of diverse deep neural networks (DNNs). Therefore, we introduce a two-stage system-level modeling method (Arch2End), combining limited architecture information with scalable benchmark models to achieve an unified performance representation. Stage one leverages public information to analyse architectures in an uniform abstraction and to design the benchmark models for exploring the device performance boundaries, ensuring uniformity. Stage two extracts critical device features from the end-to-end inference metrics of extensive simulation models, ensuring universality and enhancing characterization capacity. Compared to the state-of-the-art methods, Arch2End achieves the lowest DNN latency prediction relative errors in the NAS-Bench-201 (1.7%) and real-world DNNs (8.2%). It also showcases superior performance in intergroup balanced device grouping strategies.
Weihong Liu, Zongwei Zhu, Boyu Li 0006, Yi Xiong 0003, Zirui Lian, Jiawei Geng, Xuehai Zhou
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 Heter-Train: A Distributed Training Framework Based on Semi-Asynchronous Parallel Mechanism for Heterogeneous Intelligent Transportation Systems
abstract
Transportation big data (TBD) are increasingly combined with artificial intelligence to mine novel patterns and information due to the powerful representational capabilities of deep neural networks (DNNs), especially for anti-COVID19 applications. The distributed cloud-edge-vehicle training architecture has been applied to accelerate DNNs training while ensuring low latency and high privacy for TBD processing. However, multiple intelligent devices (e.g., intelligent vehicles, edge computing chips at base stations) and different networks in intelligent transportation systems lead to computing power and communication heterogeneity among distributed nodes. Existing parallel training mechanisms perform poorly on heterogeneous cloud-edge-vehicle clusters. The synchronous parallel mechanism may force fast workers to wait for the slowest worker for synchronization, thus wasting their computing power. The asynchronous mechanism has communication bottlenecks and can exacerbate the straggler problem, causing increased training iterations and even incorrect convergence. In this paper, we introduce a distributed training framework, Heter-Train. First, a communication-efficient semi-asynchronous parallel mechanism (SAP-SGD) is proposed, which can take full advantage of acceleration effect of asynchronous strategy on heterogeneous training and constrain the straggler problem by using global interval synchronization. Second, Considering the difference in node bandwidth, we design a solution for heterogeneous communication. Moreover, a novel weighted aggregation strategy is proposed to aggregate the model parameters with different versions. Finally, experimental results show that our proposed strategy can achieve up to$6.74 \times $speedups on training time, with almost no accuracy decrease.
Jiawei Geng, Haipeng Jia, Zongwei Zhu, Hai Fang, Chengxi Gao, Cheng Ji 0002, Gangyong Jia, Guangjie Han, Xuehai Zhou
IEEE Trans. Intell. Transp. Syst.4
2023 MIATS: A Chinese Spelling Error Correction Algorithm Based on Multimodal Information Alignment of Three-Towers Structure
abstract
Chinese spelling correction (CSC) is a crucial task in natural language processing, aiming to detect and correct spelling errors in Chinese text. The improved performance of Chinese spelling errors correction algorithms can enhance the efficiency and accuracy of upstream and downstream Chinese natural language processing tasks, such as OCR, ASR, and translation.However, current methods based on neural networks are mostly limited to either using only contextual information to correct misspelled words or failing to fully utilize glyph and pinyin information. Therefore, we propose a multimodal approach to address the above issues. Specifically, a three-tower multimodal structure is used to extract glyph, pinyin, and semantic information, and a decoder composing of an error probability prediction network and a transformer network is employed to achieve cross-modal information interaction. Besides, an additional training task is used to achieve cross-modal information alignment. Experiments demonstrate that proposed network outperform most existing motheds.
Guochao Zhao, Zongwei Zhu, Guixing Wu
ECAI4
2023 ICE: Collaborating Memory and Process Management for User Experience on Resource-limited Mobile Devices
abstract
Mobile devices with limited resources are prevalent as they have a relatively low price. Providing a good user experience with limited resources has been a big challenge. This paper found that foreground applications are often unexpectedly interfered by background applications' memory activities. Improving user experience on resource-limited mobile devices calls for a strong collaboration between memory and process management. This paper proposes a framework, Ice, to optimize the user experience on resource-limited mobile devices. With Ice, processes that will cause frequent refaults in the background are identified and frozen accordingly. The frozen application will be thawed when memory condition allows. Evaluation of resource-limited mobile devices demonstrates that the user experience is effectively improved with Ice. Specifically, Ice boosts the frame rate by 1.57x on average over the state-of-the-art.
Changlong Li 0006, Yu Liang 0004, Rachata Ausavarungnirun, Zongwei Zhu, Liang Shi 0001, Chun Jason Xue
EuroSys4
2023 Transparent File Deduplication with Reduced Update Cost on Encryption Enabled Mobile Devices
abstract
Data deduplication has been long studied to achieve data reduction. However, deploying deduplication on encryption enabled mobile systems might consume much memory footprint and computation time for hash calculations. Moreover, frequent file updates on deduplicated files could badly degrade the deduplication efficacy due to the increased file-system metadata penalty. Considering the characteristics of mobile devices, an efficient data deduplication method is proposed in this paper. First, it separates the hash calculation into foreground and background stages. The background stage calculates the hash values of potentially duplicate files while the foreground stage quickly hashes the file which is being written using a lightweight hash algorithm. Second, a dual-level node structure is proposed to improve the file update efficacy for deduplicated files, saving more storage space against the file re-splitting. Besides, we implement a superlink call to make deduplication process compatible with file-based encryption. These methods are combined to realize a transparent file deduplication (TFDedup) approach, which eliminates redundant data and reduces the associated cost of file update. Experimental results show that TFDedup succeeds to lower the space consumption when serving file updates by 55.6% and accelerate the deduplication process by 50.3%.
Junbin Ren, Cheng Ji 0002, Weiwei Jin, Weichao Guo, Yajuan Du, Zongwei Zhu
ICPADS7
2023 Ability-aware knowledge distillation for resource-constrained embedded devices
abstract
Deep Neural Network (DNN) models have notably improved the efficiency of machine learning tasks. However, their high storage and computational costs restrict their deployment on resource-limited embedded devices. Knowledge distillation (KD) has emerged as a promising approach for compressing DNN models. However, two challenges in KD, namely the capacity gap problem and the time-consuming redundancy problem, have hindered its performance and efficiency in compression. To alleviate these challenges, this paper proposes a novel framework, called Ability-Aware Knowledge Distillation (AAKD). AAKD introduces a knowledge sample selection strategy and an adaptive teacher switching strategy based on the dynamic awareness of the student’s ability. This enables the framework to automatically select suitable knowledge samples and teacher networks according to the increasing representation ability of students. Extensive experiments on different datasets and models have demonstrated that AAKD can enhance the performance of compact student models, significantly improve the efficiency of distillation, and lead to higher compression rates.
Yi Xiong 0003, Wenjie Zhai, Xueyong Xu, Jinchen Wang, Zongwei Zhu, Cheng Ji 0002
J. Syst. Archit.5
2023 iAware: Interaction Aware Task Scheduling for Reducing Resource Contention in Mobile Systems
abstract
To ensure the user experience of mobile systems, the foreground application can be differentiated to minimize the impact of background applications. However, this article observes that system services in the kernel and framework layer, instead of background applications, are now the major resource competitors. Specifically, these service tasks tend to be quiet when people rarely interact with the foreground application and active when interactions become frequent, and this high overlap of busy times leads to contention for resources. This article proposes iAware, an interaction-aware task scheduling framework in mobile systems. The key insight is to make use of the previously ignored idle period and schedule service tasks to run at that period. iAware quantify the interaction characteristic based on the screen touch event, and successfully stagger the periods of frequent user interactions. With iAware, service tasks tend to run when few interactions occur, for example, when the device’s screen is turned off, instead of when the user is frequently interacting with it. iAware is implemented on real smartphones. Experimental results show that the user experience is significantly improved with iAware. Compared to the state-of-the-art, the application launching speed and frame rate are enhanced by 38.89% and 7.97% separately, with no more than 1% additional battery consumption.
Yongchun Zheng, Changlong Li 0006, Yi Xiong 0003, Weihong Liu, Cheng Ji 0002, Zongwei Zhu, Lichen Yu
ACM Trans. Embed. Comput. Syst.6
2022 Sniper: cloud-edge collaborative inference scheduling with neural network similarity modeling
abstract
The cloud-edge collaborative inference demands scheduling the artificial intelligence (AI) tasks efficiently to the appropriate edge smart device. However, the continuously iterative deep neural networks (DNNs) and heterogeneous devices pose great challenges for inference tasks scheduling. In this paper, we propose a self-update cloud-edge collaborative inference scheduling system (Sniper) with time awareness. At first, considering that similar networks exhibit similar behaviors, we develop a non-invasive performance characterization network (PCN) based on neural network similarity (NNS) to accurately predict the inference time of DNNs. Moreover, PCN and time-based scheduling algorithms can be flexibly combined into the scheduling module of Sniper. Experimental results show that the average relative error of network inference time prediction is about 8.06%. Compared with the traditional method without time awareness, Sniper can reduce the waiting time by 52% on average while achieving a stable increase in throughput.
Weihong Liu, Jiawei Geng, Zongwei Zhu, Zirui Lian
DAC3
2022 FedNorm: An Efficient Federated Learning Framework with Dual Heterogeneity Coexistence on Edge Intelligence Systems
abstract
Federated learning (FL) is an emerging distributed learning paradigm, which aims to train machine learning models on geo-decentralized edge devices while keeping the training data stored locally. However, due to the scattered and diverse properties of edge devices, FL is often accompanied by typical heterogeneous features. One of the key challenges is statistical heterogeneity (aka non-independent identically distributed data, Non-IID), which leads to severe client-drift problem and unstable convergence. Moreover, the computational heterogeneity of devices can result in large computation time variation and thus exacerbate client-drift through inconsistent local training steps. The previous studies either ignore the client-drift problem or ignore the scatter in local gradient information, causing limited optimization effect. This paper proposes FedNorm framework to enable training Non-IID data on heterogeneous devices efficiently. First, a local model consistency update method is introduced to mitigate client-drift by allowing heterogeneous edge devices to implement different local training steps. Next, a federated gradient normalization method is introduced to reduce gradient scattering and achieves stable convergence of the model by balancing the gradient information of each edge device. We conducted extensive ablation experiments on different training tasks and training platforms with dual heterogeneity. The experimental results show that FedNorm achieves 1.52 × -3.52× speedup on convergence ratio and 7.38%-13.90% improvement in accuracy, compared to the state-of-the-art frameworks on CIFAR10.
Zirui Lian, Weihong Liu, Zongwei Zhu, Xuehai Zhou
ICCD4
2022 Task-aware swapping for efficient DNN inference on DRAM-constrained edge systems
abstract
Object detection at the edge side is a common task in various environments. The deployment of convolutional neural networks in intelligent edge systems is very challenging because of the highly constrained main-memory space. This study aims at operating neural networks with a reduced memory requirement. The basic idea is that tasks of the same type would involve the same critical subnetwork. We propose identifying the critical network connections by considering the importance of channels. During runtime, the proposed method detects the task types and timely swaps the model parameters of the critical subnetworks from the external storage into dynamic random access memory (DRAM). Compared with conventional network pruning, the proposed approach further reduced the DRAM requirement by 34.6% while maintaining a high inference accuracy.
Cheng Ji 0002, Zongwei Zhu, Xianmin Wang, Wenjie Zhai, Xuemei Zong, Mingliang Zhou 0001
Int. J. Intell. Syst.2
2022 Module Against Power Consumption Attacks for Trustworthiness of Vehicular AI Chips in Wide Temperature Range
abstract
Power consumption attacks monitoring on artificial intelligence (AI) chips play a critical role in the vehicular AI systems. However, most of the current monitoring and management methods focus on the trustworthiness of industrial equipment instead of resource-constrained edge devices. To address the above problem, a closed-loop module for monitoring and management of vehicular AI chips based on fitting and filtering to resist power consumption attacks is proposed in this paper. First, considering the characteristics of power, we propose a raw data correction approach for power monitoring to monitor abnormal power consumption. Second, we address the challenging problem of precision temperature monitoring to monitor the abnormal temperature of the chip, especially in a wide temperature range. Finally, the established method is applied to attack surveillance and transformed into a power consumption management problem solved by dynamic voltage and frequency scaling (DVFS) technology. As the experimental results reveal, compared with existing methods of power and temperature monitoring and power consumption control in wide temperature, our method can achieve significantly improved monitoring and managing performance.
Zongwei Zhu, Jiawei Geng, Mingliang Zhou 0001, Bin Fang 0001
Int. J. Pattern Recognit. Artif. Intell.1
2022 Juggler-ResNet: A Flexible and High-Speed ResNet Optimization Method for Intrusion Detection System in Software-Defined Industrial Networks
abstract
ResNetsare widely used in the intrusion detection system (IDS) of software-defined industrial network to construct accurate intelligence detection of network attacks. However, the IDS based on ResNets has a long detecting interval because of the fine-grained operator and intermediate outcomes of the multi-branch architecture of ResNets. To address this problem, in this article, we propose Juggler-ResNet with a fusible residual structure that preserves the feature extraction ability of the residual structure and enables equivalent transformation to linear topology to support low latency inference service in the industrial application (e.g., malicious network behavior detection, fault diagnosis, etc.). First, we propose a fusible multibranch residual structure to avoid gradient vanishing problems in the training phase. Second, we convert it to linear-topology by using a set of equivalent fusion operators. Finally, the linear-topology model is deployed to accelerate inference speed. Our experimental results on CIFAR-10 and CIFAR-100 show that fusible residual structure can achieve 2.08-4.3x acceleration with state-of-the-art level accuracy performance.
Zongwei Zhu, Wenjie Zhai, Huanghe Liu, Jiawei Geng, Mingliang Zhou 0001, Cheng Ji 0002, Gangyong Jia
IEEE Trans. Ind. Informatics1
2021 SAP-SGD: Accelerating Distributed Parallel Training with High Communication Efficiency on Heterogeneous Clusters
abstract
Due to rapid product iterations and high prices, the phenomenon that GPUs in clusters have heterogeneous configurations is widespread. However, existing parallel training mechanisms perform poorly on heterogeneous clusters. The synchronous parallel mechanism can cause fast GPUs to wait for the slowest GPU for synchronization, thus wasting their computing power. The asynchronous parallel mechanism has communication bottlenecks and can exacerbate the straggler problem, causing increased training iterations and even incorrect convergence. In this paper, we introduce a communication-efficient semi-asynchronous parallel mechanism (SAP-SGD), which can take full advantage of the acceleration effect of asynchronous strategy on heterogeneous training and can constrain the straggler problem by using interval global synchronization. A novel weighted aggregation strategy is proposed to aggregate the model parameters with different versions. Experimental results show that our proposed strategy can achieve up to $6.74\times$ speedup on training time, with almost no accuracy decrease.
Zongwei Zhu, Xuehai Zhou
CLUSTER2
2021 HADFL: Heterogeneity-aware Decentralized Federated Learning Framework
abstract
Federated learning (FL) supports training models on geographically distributed devices. However, traditional FL systems adopt a centralized synchronous strategy, putting high communication pressure and model generalization challenge. Existing optimizations on FL either fail to speedup training on heterogeneous devices or suffer from poor communication efficiency. In this paper, we propose HADFL, a framework that supports decentralized asynchronous training on heterogeneous devices. The devices train model locally with heterogeneity-aware local steps using local data. In each aggregation cycle, they are selected based on probability to perform model synchronization and aggregation. Compared with the traditional FL system, HADFL can relieve the central server’s communication pressure, efficiently utilize heterogeneous computing power, and can achieve a maximum speedup of 3.15x than decentralized-FedAvg and 4.68x than Pytorch distributed training scheme, respectively, with almost no loss of convergence accuracy.
Zirui Lian, Weihong Liu, Zongwei Zhu, Cheng Ji 0002
DAC4
2021 AGQFL: Communication-efficient Federated Learning via Automatic Gradient Quantization in Edge Heterogeneous Systems
abstract
With the widespread use of artificial intelligent (AI) applications and dramatic growth in data volumes from edge devices, there are currently many works that place the training of AI models onto edge devices. The state-of-the-art edge training framework, federated learning (FL), requires to transfer of a large amount of data between edge devices and the central server, which causes heavy communication overhead. To alleviate the communication overhead, gradient compression techniques are widely used. However, the bandwidth of the edge devices is usually different, causing communication heterogeneity. Existing gradient compression techniques usually adopt a fixed compression rate and do not take the straggler problem caused by the communication heterogeneity into account. To address these issues, we propose AGQFL, an automatic gradient quantization method consisting of three modules: quantization indicator module, quantization strategy module and quantization optimizer module. The quantization indicator module automatically determines the adjustment direction of quantization precision by measuring the convergence ability of the current model. Following the indicator and the physical bandwidth of each node, the quantization strategy module adjusts the quantization precision at run-time. Furthermore, the quantization optimizer module designs a new optimizer to reduce the training bias and eliminate the instability during the training process. Experimental results show that AGQFL can greatly speed up the training process in edge AI systems while maintaining or even improving model accuracy.
Zirui Lian, Yanru Zuo, Weihong Liu, Zongwei Zhu
ICCD5
2021 Memory-efficient deep learning inference with incremental weight loading and data layout reorganization on edge systems
Cheng Ji 0002, Zongwei Zhu, Li-Pin Chang, Huanghe Liu, Wenjie Zhai
J. Syst. Archit.3
2020 Machine learning assisted OSP approach for improved QoS performance on 3D charge-trap based SSDs
abstract
Three-dimensional (3D) charge-trap based solid-state-drivers (SSDs) have become an emerging storage solution in recent years. One-shot-programming in 3D charge-trap based SSDs could deliver a maximized system input/output (I/O) throughput at the cost of degraded Quality-of-Service (QoS) performance. This paper proposes reinforcement-learning based one-shot-programming (RLOSP), a reinforcement learning based approach to improve the QoS performance for 3D charge-trap based SSDs. By learning the I/O patterns of the workload environments as well as the device internal status, the proposed approach could properly choose requests in the device queue, and allocate physical addresses for these requests during one-shot-programming. In this manner, the storage device could deliver an improved QoS performance. Experimental results reveal that the proposed approach could reduce the worst-case latency at the 99.9th percentile by 37.5%–59.2%, with an optimal system I/O throughput.
Zongwei Zhu, Chao Wu 0006, Cheng Ji 0002, Xianmin Wang
Int. J. Intell. Syst.1
2020 Modified DenseNet for Automatic Fabric Defect Detection With Edge Computing for Minimizing Latency
abstract
As an essential step in quality control, fabric defect detection plays an important role in the textile manufacturing industry. The traditional manual detection method is inaccurate and incurs a high cost; as a result, it is gradually being replaced by deep learning algorithms based on cloud computing. However, a high data transmission latency between end devices and the cloud has a significant impact on textile production efficiency. In contrast, edge computing, which provides services near end devices by deploying network, computing and storage facilities at the edge of the Internet, can effectively solve the above-mentioned problem. In this article, we propose a deep-learning-based fabric defect detection method for edge computing scenarios. First, this article modifies the structure of DenseNet to better suit a resource-constrained edge computing scenario. To better assess the proposed model, an optimized cross-entropy loss function is also formulated. Afterward, six feasible expansion schemes are utilized to enhance the data set according to the characteristics of various defects in fabric samples. To balance the distribution of samples, proportions of various defect types are used to determine the number of enhancements. Finally, a fabric defect detection system is established to test the performance of the optimized model used on edge devices in a real-world textile industry scenario. Experimental results demonstrate that compared with the conventional convolutional neural network (CNN), the proposed optimized model attains an average improvement of 18% in the area under the curve (AUC) metric for 11 defects. Data transmission is reduced by approximately 50% and latency is reduced by 32% in the Cambricon 1H8 platform compared with a cloud platform.
Zongwei Zhu, Guangjie Han, Gangyong Jia, Lei Shu 0001
IEEE Internet Things J.1
2020 PHDFS: Optimizing I/O performance of HDFS in deep learning cloud computing platform
Zongwei Zhu, Luchao Tan, Yinzhen Li, Cheng Ji 0002
J. Syst. Archit.1
2020 Inspection and Characterization of App File Usage in Mobile Devices
abstract
While the computing power of mobile devices has been quickly evolving in recent years, the growth of mobile storage capacity is, however, relatively slower. A common problem shared by budget-phone users is that they frequently run out of storage space. This article conducts a deep inspection of file usage of mobile applications and their potential implications on user experience. Our major findings are as follows: First, mobile applications could rapidly consume storage space by creating temporary cache files, but these cache files quickly become obsolete after being re-used for a short period of time. Second, file access patterns of large files, especially executable files, appear highly sparse and random, and therefore large portions of file space are never visited. Third, file prefetching brings an excessive amount of file data into page cache but only a few prefetched data are actually used. The unnecessary memory pressure causes premature memory reclamation and prolongs application launching time. Through the feasibility study of two preliminary optimizations, we demonstrated a high potential to eliminate unnecessary storage and memory space consumption with a minimal impact on user experience.
Cheng Ji 0002, Riwei Pan, Li-Pin Chang, Liang Shi 0001, Zongwei Zhu, Yu Liang 0004, Tei-Wei Kuo, Chun Jason Xue
ACM Trans. Storage5
2018 Delayed Wake-Up Mechanism Under Suspend Mode of Smartphone
Bo Chen 0010, Xi Li 0003, Xuehai Zhou, Zongwei Zhu
CollaborateCom4
2014 Behavior Gaps and Relations between Operating System and Applications on Accessing DRAM
abstract
Detailed analyses of the behaviors of operating system and applications are significant for taking full advantage of the precious hardware resources and improving performance. This paper focus on their DRAM access behaviors based on access proportion and row-buffer miss ratio (RBM). The access proportions of Kernel and User vary greatly in different stages throughout the lifetime of a process. Most of the row-buffer misses are caused by the one having higher access proportion. By analyzing the RBM series through ARMA model, we found that User's DRAM accesses only have short-term influences on its behavior, while the Kernel's influences are relatively deeper. The ARMA model for the RBM series is able to predict the future RBMs, which are profound basis to schedule the DRAM access commands. The results of Gaussian Fitting show that Kernel and User are tightly correlated on accessing DRAM, especially in the steady stage and the end stage of a process's life cycle. Based on this close relation, it is possible to estimate the DRAM access behaviors of the other one according to the one whose behaviors have been known. System-calls that obviously affect the access proportions and RBMs are also revealed in this paper.
Beilei Sun, Xi Li 0003, Zongwei Zhu, Xuehai Zhou
ICECCS3
2014 A Thread Behavior-Based Memory Management Framework on Multi-core Smartphone
abstract
Memory management systems have significantly affected the overall performance of modern multi-core smartphone systems. Android, as one of the most popular smartphone operating systems, adopts a global buddy system with the FCFS (first come, first served) principle for memory allocation, and releases requests to manage external fragmentations and maintain the memory allocation efficiency. However, extensive experimental study on thread behaviors indicates that memory external fragmentation is no longer the crucial bottleneck in most Android applications. Specifically, a thread usually allocates or releases memory in bursts, resulting in serious memory locks and inefficient memory allocation. Furthermore, the pattern of such bursting behaviors varies throughout the life cycle of a thread. The conventional FCFS policy of Android buddy system fails to adapt to such variations and thus suffers from performance degradation. In this paper, we propose a novel memory management framework, called Memory Management Based on Thread Behaviors (MMBTB), for multi-core smartphone systems. It adapts to various thread behaviors through targeted optimizations to provide efficient memory allocation. The efficiency and effectiveness of this new memory management scheme on multicore architecture is proved by a theoretical emulation model. Our experimental studies on the real Android system show that MMBTB can improve the efficiency of memory allocation by 12%-20%, confirming the theoretical analysis results.
Zongwei Zhu, Xi Li 0003, Hengchang Liu, Cheng Ji 0002, Xuehai Zhou, Beilei Sun
ICECCS1
2014 Kernel-User Space Separation in DRAM Memory
abstract
Performance of software is increasingly restricted by the Memory Wall instead of CPU. Many studies focus on alleviating the DRAM latency by improving the row-buffer hit rate. But most of them treat the Kernel and User equally. Data used by Operating System and User applications spread in different rows of the same bank, leading to the contentions for the row-buffer when they access the bank successively. We find that contentions between Kernel and User make up of a great proportion of all the row-buffer misses. To alleviate the contentions between Kernel and User, we divide the united DRAM memory space into Kernel-Space and User-Space. A new page-allocation-system, the K/U-Aware page-allocation-system, is proposed to manage Kernel-Space and User-Space in DRAM memory in different address mapping schemes of DRAM memory controller. In the new system, pages are allocated from different spaces according to applicants (Kernel or User). Sizes of the two spaces increase and decrease dynamically as required. For benchmarks in PARSEC suites, the proposed system reduces the contentions of Kernel and User effectively, producing significant improvements of row-buffer hit rate. The execution time is reduced by 9.45% (max. 20.45%) and 6.51% (max. 18.05%) respectively in two typical address mapping schemes.
Xi Li 0003, Beilei Sun, Zongwei Zhu, Chao Wang 0003, Xuehai Zhou
ISPA3
2014 Memory power optimization on different memory address mapping schemas
abstract
Since memory accounts for a large and increasing fraction of the energy consumed by computers, memory manufacturers have developed memory devices with different power/work modes. For taking full advantage of these modes, more and more creditable hardware or software power mode control algorithms have been proposed. In this paper, by analyzing the effects of power mode control polices on different memory address mapping schemas (schema is used to translate a given physical address to a specific memory cell in DRAM system), we find that most previous power mode control policies are sensitive to mapping schemas. Therefore, in order to manage these power modes on different mapping schemas more effectively, we divide them into two categories: high-bit multi-access cross memory (HMCM) and low-bit multi-access cross memory (LMCM), and then take a targeted optimization. For the former schema, a rank-sensitive buddy system (RS-Buddy) was proposed to cluster pages together to prolong memory modules' low power time. For the latter, we introduce a comprehensive solution named as MSPA. It adopts a memory address segmentation module (MASM) to split memory into many regions configured as different mapping schemas. And with the help of an OS power-aware memory allocator (PAMA), MSPA can dynamically allocate one application's memory from its preferred region to balance power and performance. By performing extensive experiments on practical platform for HMCM while on simulator for LMCM, the results of HMCM show that RS-Buddy can optimize the power efficiency from 2% to 22%. Furthermore, the simulation results of LMCM demonstrate that MSPA can further improve the power efficiency from 3% to 17% when combined with other previous state-of-the-art studies.
Zongwei Zhu, Xi Li 0003, Chao Wang 0003, Xuehai Zhou
RTCSA1
2013 Power-aware buddy system and task group scheduler
abstract
Memory is responsible for a large and increasing fraction of the energy consumed by computers. To address this challenge, memory manufacturers have developed memory devices with different power states. In order to more effectively manage the power states in the operating system, in this paper, we propose a rank-sensitive buddy system (RS-Buddy) which clusters pages together to prolong the idle time of memory ranks without breaking defragmentation characteristics. For the purpose of decreasing unnecessary frequent mode transitions, we introduce a power-aware task group scheduler (PATGS) that groups the threads which access the same rank together to schedule while sustaining system fairness. Finally, we integrate state-of-the-art mode control policies with our RS-Buddy and PATGS, with experimental results demonstrating that our algorithms can improve the power efficiency from 25.31% to 27.35% compared with state-of-the-art studies.
Xi Li 0003, Zongwei Zhu, Gangyong Jia, Xuehai Zhou
ISCAS2
2012 Cache Promotion Policy Using Re-reference Interval Prediction
abstract
The last-level cache (LLC) mitigates the long latencies of memory access in today's chip multi-core processor (CMP). The promotion policy in the LLC largely affects cache efficiency, while an inappropriate promotion policy may lead useless blocks to remain in the cache longer than necessary, in turn result into inefficiency. Currently state-of-the-art promotion policies are unaware of the re-reference interval of cache accesses. Applications that exhibit a long re-reference interval perform poorly with these promotion policies. In this paper, we propose a promotion policy that uses re-reference interval prediction (RRIP) information. Such technique requires minor hardware modification over the least-recently-used (LRU) replacement policy. Our evaluation shows that RRIP improves IPCsumby 2.58%, Weighted Speedup by 3.54% and IPCnorm_hmeanby 6.2% on average over single-step promotion policy.
Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu
CLUSTER5
2012 Memory Affinity: Balancing Performance, Power, Thermal and Fairness for Multi-core Systems
abstract
Main memory is expected to grow significantly in both speed and capacity for it is a major shared resource among cores in a multi-core system, which will lead to increasing power consumption. Therefore, it is critical to address the power issue without seriously decreasing performance in the memory subsystem. In this paper, we firstly propose memory affinity which retains the active and low power memory ranks as long as possible to avoid frequently switching between active and low power status, and then present a memory affinity aware scheduling (MAS) to balance performance, power, thermal and fairness for multi-core systems. Experimental results demonstrate our memory affinity aware scheduling algorithms well adapt to system loading to maximize power saving and avoid memory hotspot at the same time while sustaining the system bandwidth demand and preserving fairness among threads.
Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu
CLUSTER5
2012 Share memory aware scheduler: balancing performance and fairness
abstract
Optimizing system performance through scheduling has received a lot of attention. However, none of the existing approaches can balance the system performance improvement and the fair share of CPU time among threads. We present in this paper a share memory aware scheduler (SMAS). The key idea is to adopt thread group scheduling which partitions threads based on memory address space to reduce switching overhead and to give each thread a fair chance to occupy CPU time. There are three main contributions: 1) SMAS does well in balancing system performance and fairness among all threads; 2) to our knowledge, this is the first attempt to use share memory aware scheduler for system performance improvement; 3) we implement SMAS both in testbed and simulator for evaluation. The testbed results on a 2-core processor show that our proposed scheduler can improve performance of different performance parameters with neglected overhead in fairness, which reduced 0.128% in cache miss rate, 2.62% in run time, 13.15% in DTBL misses, 31.68% in ITLB misses and 46.15% in ITLB flushes maximum. Furthermore, our extensive simulation results for 4 and 8 cores demonstrate that SMAS is highly scalable.
Xi Li 0003, Gangyong Jia, Zongwei Zhu, Xuehai Zhou
ACM Great Lakes Symposium on VLSI4
2012 Behavior Aware Data Locality for Caches
abstract
Optimizing cache performance through improving data locality has been receiving a lot of attention. However, none of the existing approaches can combine each task's behavior to optimize data locality for caches. We present a behavior aware data locality (BADL) to optimize cache performance in this paper. The key idea is to add each task's behavior when allocating memory, which can take advantage of each task's different locality to optimize cache performance. There are five main contributions: 1. to our best knowledge, this is the first attempt to improve cache performance through combining task behavior, 2. BADL detailed analyzes low performance derived from internal of the cache line, which is more fine-grained than the current state-of-the-art fine-grained in hardware angle, 3. BADL optimizes the cache performance through improving internal of cache line efficiency, 4. we implement BADL both in single-threaded application and multi-threaded applications scenarios, 5. BADL can be combined to most of the cache optimizing researches. The experiment results show our proposed BADL can improve 18.6% performance on average in single-threaded application situation and improve 20.8% performance on average in multi-threaded application situation.
Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu
ICPADS5
2012 Frequency Affinity: Analyzing and Maximizing Power Efficiency in Multi-core Systems
abstract
Performance optimization and energy efficiency are the major challenges in multi-core system design. Of the state-of-the-art approaches, cache affinity aware scheduling and techniques based on dynamic voltage frequency scaling (DVFS) are widely applied to improve performance and save energy consumptions respectively. In modern operating systems, schedulers exploit high cache affinity by allocating a process on a recently used processor whenever possible. When a process runs on a high-affinity processor it will find most of its states already in the cache and will thus achieve more efficiency. However, most state-of-the-art DVFS techniques do not concentrate on the cost analysis for DVFS mechanism. In this paper, we firstly propose frequency affinity which retains the voltage frequency as long as possible to avoid frequently switching, and then present a frequency affinity aware scheduling (FAS) to maximize power efficiency for multi-core systems. Experimental results demonstrate our frequency affinity aware scheduling algorithms are much more power efficient than single-ISA heterogeneous multi-core processors.
Gangyong Jia, Xi Li 0003, Chao Wang 0003, Xuehai Zhou, Zongwei Zhu
MASCOTS5