Jing Wang 0055

dblp:02/736-55 · DBLP profile ↗
← Back
66ranked-venue papers
12as first author
46since 2021 · last 2026
0000-0003-3653-7013ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 52 · 9 first-author · 33 since 2021Computer networks · 7 · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert Offloading
abstract
Mixture-of-experts (MoE) architectures enable scalable Large Language Models (LLMs) with reduced computational overhead, yet their deployment on memory-constrained edge devices is hindered by substantial memory demands. Traditional expert-offloading techniques mitigate memory constraints but often significantly increase inference latency. We introduce MoE-APEX, an Adaptive Precision EXpert offloading system that optimizes MoE inference for edge architectures by dynamically managing expert precision. Our core innovation is to replace less critical cache-miss experts with low-precision variants, reducing loading latency while maintaining accuracy. MoE-APEX introduces three innovative techniques that map the natural hierarchy of MoE computation: (1) a token-level dynamic expert loading mechanism, (2) a layer-level adaptive expert prefetching technique, and (3) a sequence-level cost-aware expert caching policy. These innovations enable MoE-APEX to leverage the benefits of mixed-precision expert inference fully. Implemented atop Llama.cpp, MoE-APEX achieves decoding speedups ranging from 1.34x to 9.75x compared to state-of-the-art MoE offloading systems across diverse edge devices, offering a robust solution for efficient MoE deployment in resource-constrained environments.
Jiacheng Liu 0001, Xiaofeng Hou, Yi-Fei Pu, Jing Wang 0055, Pheng-Ann Heng, Chao Li 0009, Minyi Guo
ASPLOS (2)5
2026 LiveGraph: High-Performance On-FPGA Dynamic Graph Updating Framework
abstract
Dynamic graphs are ubiquitous in real-world scenarios, demanding both timely updates and low-latency responses. However, existing FPGA-based solutions still face three major limitations in supporting such workloads: • Heavy CPU dependence [3] • Inadequate support for dynamic graphs [1] , [2] • Poor support for irregular updates [4]
Yufeng Luo, Peikun Hong, Jing Wang 0055, Feiyang Wu, Chao Li 0009, Minyi Guo
FCCM3
2026 AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM Serving
abstract
Generative AI, especially LLM, is driving a fundamental shift in software paradigms, prompting cloud providers to build more efficient serving infrastructures. To meet the computational demands of emerging software, modern CPU processors are integrating Accelerator Units (AU) in the pipeline to accelerate key operations, such as Intel AMX for matrix multiplication. Current practices that dedicate AU-enabled CPU exclusively to LLM serving lead to significant resource waste and inferior efficiency. To this end, sharing AU-enabled CPU with general workloads is necessary to harvest redundant resources and improve platform performance-per-watt. However, perfectly sharing AU can be challenging since they introduce three-dimensional variations: variable usage patterns, compulsory frequency interferences, and dissimilar resource bounds. Existing resource managers are oblivious to complex Accelerator Unit Variations (AUV), resulting in performance and efficiency degradations of up to 50 % in shared environments. Therefore, this paper introduces AUM, a novel AU-aware resource manager designed to handle AUV and maximize the efficiency of shared processors. AUM has two cooperative components with three stages for three-dimensional AUV. The background profiler characterizes the usage, frequency, and resource information into a discrete model, guiding the runtime controller to analyze usage-aware requirements, select frequency-aware divisions, and make bound-aware resource decisions. Through extensive evaluations on production AU-enabled CPUs, we show that AUM improves CPU efficiency by$4.7-8.8 \%$while maintaining high-performance AU applications by reducing SLO violations by$\mathbf{7 - 1 1 \%}$compared with state-of-the-art resource managers.
Xinkai Wang 0003, Chao Li 0009, Yiming Zhuansun, Jinyang Guo 0001, Xiaofeng Hou, Jing Wang 0055, Weigao Chen, Liping Zhang 0013, Minyi Guo
HPCA6
2026 OpScope: Exploiting Operation-Driven Visual Scope for QoS-Stable Cloud Gaming
Taolei Wang, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo
IWQoS3
2026 MI-LLM: Multiplier-Free LLM Inference on Commodity Processing-in-Memory Hardware
abstract
Large language models (LLMs) are prominent for their superior ability in language understanding and generation. However, a notorious problem for LLM inference is low computational utilization caused by the memory bottleneck, since it typically requires large memory capacity and high bandwidth to process neural weights. By integrating processing cores into memory, Processing-In-Memory (PIM) architecture excels at alleviating memory bottleneck; with the recent release of the first commodity near-bank PIM hardware (NBP), PIM becomes off-the-shelf and shows great potential for accelerating LLM inference practically. However, simply shoehorning LLM inference on NBP can not achieve satisfactory performance due to its inherent limitations: weak compute performance, frequent cache misses caused by the limited working memory capacity, and poor inter-PIM-core communication bandwidth. To address these limitations, we propose MI-LLM, an efficient system deploying LLM inference on NBP hardware. Its key idea is to build NBP-aware Lookup Tables (LUTs) and completely replace multiplications with lookups on LUTs, thereby mitigating the limitation of weak compute performance. 1) To reduce the model accuracy drop caused by the use of LUT, MI-LLM tailors a learning-based LUT construction method to maintain the model accuracy. 2) To cope with frequent cache misses caused by LUT sizes far exceeding PIM working memory capacity, MI-LLM introduces the design of PIM-aware linear kernel, with the optimization of intra-row and inter-row reordering enabled, to enhance LUT lookup locality. 3) MI-LLM further proposes a model partitioning scheme to minimize inter-PIM-core communication. Kernel-level benchmarks reveal that MI-LLM achieves a 9% throughput improvement and an 11% increase in energy efficiency over GPU implementations. Compared to FP8 quantization, MI-LLM incurs only a 0.24 times increase in perplexity, demonstrating minimal accuracy degradation. Moreover, in our end-to-end evaluation, MI-LLM requires 80% fewer ALU operation ticks per output token than the GPU baseline.
Puyun Hu, Minhui Xie, Linjiang Li, Kuiyaohui Zhang, Erge Xiang, Jing Wang 0055, Size Zheng 0001, Xiao Zhang 0001, Yunpeng Chai
IEEE Trans. Computers6
2026 Enabling Learning-Based Efficiency Optimizer With Shadow Cycles in Resource-Constrained Autonomous Embedded Systems
abstract
The emerging trend of autonomous embedded systems (AES) is promising to minimize human intervention in critical tasks. In the pursuit of maximal per-watt performance, the complex hardware and software of AES require intelligent energy efficiency optimizers (EO), and the stochastic runtime variances require continuous EO. However, deploying the desirable ondevice EO causes severe performance slowdown due to contention on limited computing power with the AES pipeline. We find that there are ignored and underutilized heterogeneous resources within AES for costly EO, which results from unbalanced accelerator behaviors and misaligned parallel inference executions. We experimentally and theoretically analyze theShadow Cycleswithin the realistic autonomous Bird’s Eye View pipeline on commercial embedded platforms, categorizing them into vertical and horizontal types with distinct properties.In this paper, we introduceSHEEO+, a continuous and intelligent energy efficiency optimizer that utilizes ignored heterogeneous shadow cycles. It achieves continuous and lightweight AES monitoring with the observation module, as well as intelligent and efficient AES power management with the optimization module. On the one hand,SHEEO+ observes both the internal runtime status and external environment variance with portable interfaces to capture shadow cycles and real-time states. On the other hand,SHEEO+ optimizes power configurations per iteration based on deep reinforcement learning (DRL) methods. It tailors DRL for two types of shadow cycles and invocates optimization processes based on resource availability. To extensively evaluateSHEEO+, we implement a prototype and deploy it on realistic edge platforms. The evaluation results show thatSHEEO+ utilizes up to 74.2% shadow cycles and achieves up to 18.6% energy efficiency improvements compared to state-of-the-art energy efficiency optimizers with negligible deployment overheads.
Xinkai Wang 0003, Chao Li 0009, Xiaofeng Hou, Jing Wang 0055, Minyi Guo, Yaqian Zhao
IEEE Trans. Computers7
2026 Multi-Task Parallel Execution-Oriented Content Caching, Computation Offloading and Channel Allocation in UAV-Assisted MEC Network
abstract
Leveraging flexible deployment, extensive coverage, and reliable communication links, unmanned aerial vehicle (UAV)-assisted mobile edge computing offers new opportunities to support mobile devices with heavy computational tasks. Considering the energy constraint of UAVs and delay requirement of tasks, many efforts should be devoted to pursuing lower service latency, which however is under explored in this innovational architecture. In this paper, with the purpose of minimizing the overall network service duration, departing from traditional serial task execution, we first design a multi-task parallel execution paradigm, and then investigate a joint optimization problem encompassing content caching, computation offloading, and channel allocation. To address this intractable problem involving large state and action spaces, we decompose it into two subproblems, i.e., an intra content caching and computation offloading optimization of each UAV, and an inter channel allocation of all UAVs. We then propose a reinforcement learning-based two-layer optimization scheme that integrates the efficient representation of DQN learning and the comprehensive exploration of regret minimization learning. Specifically, in the lower layer, a DQN-based algorithm is developed to solve the intra subproblem, and in the upper layer, a regret minimization-based algorithm is designed to tackle the inter subproblem. Through nested optimization between the two layers, optimal strategies for content caching, computation offloading and channel allocation can be achieved. Numerical results demonstrate that the proposed scheme significantly reduces service latency compared to various baseline methods.
Chaoqiong Fan, Jichao Zhan, Jing Wang 0055, Shiwen Mao
IEEE Trans. Mob. Comput.3
2025 Accelerating Large-Scale Out-of-GPU-Core GNN Training with Two-Level Historical Caching
Jing Wang 0055, Taolei Wang, Juntao Huang, Xinkai Wang 0003, Marius Kreutzer, Chao Li 0009, Minyi Guo
APPT1
2025 AsymServe: Demystifying and Optimizing LLM Serving Efficiency on CPU Acceleration Units
Xinkai Wang 0003, Yiming Zhuansun, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo
APPT4
2025 DiTAC: Discrete Teamwork Abstraction for Ad Hoc Collaboration
abstract
Training autonomous agents to collaborate with unknown teammates in cooperative multi-agent environments remains a fundamental challenge in ad hoc teamwork research. Conventional approaches rely heavily on online interactions with arbitrary teammates under the assumption of full observability. However, in real-world scenarios, teammate policies are often inaccessible, making historical trajectory rollouts a more practical alternative. We propose DiTAC, a method that learns discrete teamwork abstractions for ad hoc collaboration by automatically extracting latent cooperation patterns from short trajectory segments and adapting effectively to diverse teammate behaviors. To mitigate the out-of-distribution challenge, we constrain learned representations within a discrete code-book. Furthermore, we employ a masked bidirectional transformer architecture to infer teammate behaviors from local observations, thereby relaxing the full observability assumption. Empirical results demonstrate that DiTAC significantly outperforms existing baselines and its variants across widely-used ad hoc teamwork tasks.
Jing Wang 0055, Pengjie Gu, Mengchen Zhao, Guangyong Chen, Furui Liu, Pheng-Ann Heng
ECAI1
2025 TriCooling-Sim: Efficient Thermal Simulation for High-Density Micro AI Data Centers
Jinyang Guo 0001, Xinkai Wang 0003, Jing Wang 0055, Xiaofeng Hou, Chao Li 0009, Minyi Guo
NPC (2)3
2025 CGO: Cloud Game Orchestration via Resource Preception and CODEC Optimization
Taolei Wang, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo
NPC (2)3
2025 MMBypass: Towards efficient multi-modal AI computing with adaptive bypass network
Yi-Fei Pu, Xinfeng Xia, Xiaofeng Hou, Jiacheng Liu 0001, Jing Wang 0055, Minyi Guo, Jingling Yuan, Chao Li 0009
J. Parallel Distributed Comput.7
2025 NANI: Energy-efficient Neuron-Aware hardware Noise Injection for adversarial defense using undervolting
Qiyu Wan, Jing Wang 0055, Mingsong Chen 0001, Lu Peng 0001, Xin Fu 0001
J. Syst. Archit.3
2025 Enhancing High-Throughput GPU Random Walks Through Multi-Task Concurrency Orchestration
abstract
Random walk is a powerful tool for large-scale graph learning, but its high computational demand presents a challenge. While GPUs can accelerate random walk tasks, current frameworks fail to fully utilize GPU parallelism due to memory-to-compute bandwidth imbalance. In this article, CoWalker, an efficient GPU framework, is proposed to facilitate concurrent execution of random walks for high overall throughput. CoWalker features three novel designs. First, it incorporates a multi-level execution model that effectively orchestrates diverse walk tasks and reduces GPU stalls based on multiple graph characteristics. Second, it collaboratively manages graph data and streaming multiprocessors to minimize memory access interference and maximize core utilization under concurrent tasks. Finally, a multi-dimensional scheduler selects compatible random walk task combinations based on memory footprints to achieve maximum throughput. CoWalker significantly improves throughput over state-of-the-art baselines by mitigating concurrency overheads and effectively harnessing GPU parallelism. Our extensive evaluations on real-world workloads demonstrate that CoWalker achieves 2.75× higher overall system throughput compared with commercial tools and 1.56× over the SOTA academic system.
Chao Li 0009, Xiaofeng Hou, Junyi Mei, Jing Wang 0055, Pengyu Wang 0003, Shixuan Sun, Minyi Guo, Baoping Hao
ACM Trans. Archit. Code Optim.5
2025 P3ID: A Privacy-Preserving Person Identification Framework Towards Multi-Environments Based on Transfer Learning
abstract
Concerns surrounding privacy leakages caused by prevalent vision-based person identifications are countless. A promising privacy-preserving solution is to identify the wireless signals reflecting persons, which, however, faces a major challenge of losing efficacy in multi-environments. In this paper, we work on person identification based on wireless signals using transfer learning, toward tackling the performance deterioration across environments. We investigate the feature variations induced by environmental shifts based on data measurements. Lay our foundation on the feature alignment concept, we propose a novel wireless-based person identification framework using transfer learning. In the framework, we integrate a series of signal processing methods including signal selection, pre-processing, and augmentation, where the first includes a reference environment to assist the feature extraction while the latter two respectively reduce the data noise and improve the data diversity. We also propose a model generalization method where a neural network is employed to align features from different environments, which facilitates the extraction of environment-independent features while incorporating both person and environment information. On a real wireless testbed consisting of an Impulse Radio Ultra-WideBand (IR-UWB) radar, we build and publicly release a dataset with 22,264 samples of ten individuals from three environments, varying in testing distance and obstruction condition. Extensive experimental evaluations demonstrate that the proposed framework can improve the identification accuracy across environments, and surpasses state-of-the-art methods by up to 18.06%.
Hanxiang He, Xintao Huan, Jing Wang 0055, Yong Luo 0002, Han Hu 0003, Jianping An
IEEE Trans. Mob. Comput.3
2024 Improving the Efficiency of Serverless Computing via Core-Level Power Management
abstract
Serverless computing has recently become a significant application paradigm in data centers. However, existing power management methods focus on optimizations at the coarse-grained server level, making them unable to handle the characteristics of these short-lived, dynamic serverless functions. In this context, the unawareness of function-level characteristics by the existing power management systems can severely degrade the energy efficiency of the data centers. To address this challenge, we design a function-level power management system. Instead of relying on server-level schedulers, we propose a novel core-level scheduling policy for serverless functions that can efficiently allocate functions to the most suitable CPU core. Additionally, we propose a power management mechanism for serverless computing that can reduce system power consumption with functions’ QoS guaranteed. Our evaluation shows that our system achieves a maximum power saving of 8.5% and an average power saving of 8% across the majority of loads without incurring any loss in tail latency, as compared to the conventional server-level scheduling system.
Du Liu, Jing Wang 0055, Xinkai Wang 0003, Chao Li 0009, Lu Zhang 0049, Xiaofeng Hou, Xiaoxiang Shi, Minyi Guo
CCGrid2
2024 M2SN: Adaptive and Dynamic Multi-modal Shortcut Network Architecture for Latency-Aware Applications
abstract
Multi-modal neural networks have demonstrated exceptional performance by merging information across modalities, surpassing the state-of-the-art uni-modal DNNs. However, this accuracy improvement comes at the cost of increased computation, leading to higher inference latency. This defect significantly limits the practical value of multi-modal DNNs, especially for latency-aware applications. Therefore, we propose an adaptive and efficient multi-modal shortcut architecture called M2SN to reduce the execution latency with accuracy guarantees. It skips ineffective network layers to reduce computational costs as well as alleviate the overfitting problem adaptive to specific models and scenarios. The key contributions of M2SN are twofold: 1) We design and insert shortcuts into each uni-modal network to perform adaptive computing. 2) We design a navigator to dynamically choose the optimal shortcuts. Unlike previous approaches, M2SN features high generality as it does not rely on any prior knowledge. The experimental results show that M2SN can reduce 28.3% average latency while obtaining the same or higher accuracy compared with SOTA baselines.
Yi-Fei Pu, Xiaofeng Hou, Jiacheng Liu 0001, Jing Wang 0055, Minyi Guo, Chao Li 0009
ICME6
2024 CoCG: Fine-grained Cloud Game Co-location on Heterogeneous Platform
abstract
Cloud games have received widespread attention and exponential growth recently as a key technology for building metaverse. Unlike general tasks in the cloud, the scene-complex, latency-critical, and interaction-intensive features make it challenging for cloud game co-deployment on heterogeneous platforms. Game-grained resource allocation leads to low resource effectiveness. Although previous work tries to explore individual game partitioning methods, they still face the problem of inefficient game hosting decisions and ultimately QoS violations. In this paper, we propose a fine-grained game characteristic and scheduling strategy to co-locate games together for high resource usage effectiveness. First, we fully explore the relationships between game scenes and resource usage behaviors by breaking the cloud game into stages with multiple frames and clustering them. We adopt machine learning methods to predict game resource consumption in real-time. To further improve multi-game parallelism, we co-locate games in a complementary way and steal time from the loading stage to avoid oversubscribing. The evaluation shows that our work increased the throughput of the cloud game deployments by 23.7% with low overhead compared to previous work.
Taolei Wang, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo
IPDPS3
2024 PrimKD: Primary Modality Guided Multimodal Fusion for RGB-D Semantic Segmentation
abstract
The recent advancements in cross-modal transformers have demonstrated their superior performance in RGB-D segmentation tasks by effectively integrating information from both RGB and depth modalities. However, existing methods often overlook the varying levels of informative content present in each modality, treating them equally and using models of the same architecture. This oversight can potentially hinder segmentation performance, especially considering that RGB images typically contain significantly more information than depth images. To address this issue, we propose PrimKD, a knowledge distillation based approach that focuses on guided multimodal fusion, with an emphasis on leveraging the primary RGB modality. In our approach, we utilize a model trained exclusively on the RGB modality as the teacher, guiding the learning process of a student model that fuses both RGB and depth modalities. To prioritize information from the primary RGB modality while leveraging the depth modality, we incorporate primary focused feature reconstruction and a selective alignment scheme. This integration enhances the overall freature fusion, resulting in improved segmentation results. We evaluate our proposed method on the NYU Depth V2 and SUN-RGBD datasets, and the experimental results demonstrate the effectiveness of PrimKD. Specifically, our approach achieves mIoU scores of 57.8 and 52.5 on these two datasets, respectively, surpassing existing counterparts by 1.5 and 0.4 mIoU. The code is available at https://github.com/xiaoshideta/PrimKD.
Zhiwei Hao 0001, Zhongyu Xiao, Yong Luo 0002, Jianyuan Guo, Jing Wang 0055, Li Shen 0008, Han Hu 0003
ACM Multimedia5
2024 Boosting Data Center Performance via Intelligently Managed Multi-backend Disaggregated Memory
abstract
Existing disaggregated memory (DM) systems face a problem of underutilized far memory bandwidth, which greatly limits the data throughput when processing data-intensive applications. Specifically, prior works all target runtime design for a single PCIe-based secondary memory device (i.e., single-backend far memory) with low data bandwidth and high system overhead. In this work, we take the first step to realize a well-crafted, multi-backend DM system with scale-out far memory paths. We propose xDM, a novel DM management scheme that can dynamically build and implicitly select appropriate far memory access paths. As part of xDM, we devise a smart far memory configuration strategy that can further optimize bandwidth usage effectiveness by tuning a wide set of key parameters based on synthesized information of application page data. Our design shows up to $3.9 \times$ data swap performance speedup, $2.8 \times$ data throughput increase, and $5.1 \times$ data center task throughput improvement compared with state-of-the-art works.
Jing Wang 0055, Hanzhang Yang, Chao Li 0009, Yiming Zhuansun, Wang Yuan, Xiaofeng Hou, Minyi Guo, Yang Hu 0001, Yaqian Zhao
SC1
2024 FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk Framework
abstract
Dynamic graph random walk (DGRW) emerges as a practical tool for capturing structural relations within a graph. Effectively executing DGRW on GPU presents certain challenges. First, existing sampling methods demand a pre-processing buffer, causing substantial space complexity. Moreover, the power-law distribution of graph vertex degrees introduces workload imbalance issues, rendering DGRW embarrassed to parallelize. In this paper, we propose FlowWalker, a GPU-based dynamic graph random walk framework. FlowWalker implements an efficient parallel sampling method to fully exploit the GPU parallelism and reduce space complexity. Moreover, it employs a sampler-centric paradigm alongside a dynamic scheduling strategy to handle the huge amounts of walking queries. FlowWalker stands as a memory-efficient framework that requires no auxiliary data structures in GPU global memory. We examine the performance of FlowWalker extensively on ten datasets, and experiment results show that FlowWalker achieves up to 752.2×, 72.1×, and 16.4× speedup compared with existing CPU, GPU, and FPGA random walk frameworks, respectively. Case study shows that FlowWalker diminishes random walk time from 35% to 3% in a pipeline of ByteDance friend recommendation GNN training.
Junyi Mei, Shixuan Sun, Chao Li 0009, Cheng Chen 0008, Jing Wang 0055, Cheng Zhao 0001, Xiaofeng Hou, Minyi Guo, Bingsheng He, Xiaoliang Cong
Proc. VLDB Endow.7
2024 Enhancing Neural Network Reliability: Insights From Hardware/Software Collaboration With Neuron Vulnerability Quantization
abstract
Ensuring the reliability of deep neural networks (DNNs) is paramount in safety-critical applications. Although introducing supplementary fault-tolerant mechanisms can augment the reliability of DNNs, an efficiency tradeoff may be introduced. This study reveals the inherent fault tolerance of neural networks, where individual neurons exhibit varying degrees of fault tolerance, by thoroughly exploring the structural attributes of DNNs. We thereby develop a hardware/software collaborative method that guarantees the reliability of DNNs while minimizing performance degradation. We introduce the neuron vulnerability factor (NVF) to quantify the susceptibility to soft errors. We propose two efficient methods that leverage the NVF to minimize the negative effects of soft errors on neurons. First, we present a novel computational scheduling scheme. By prioritizing error-prone neurons, the expedited completion of their computations is facilitated to mitigate the risk of neural computing errors that arise from soft errors without sacrificing efficiency. Second, we propose the NVF-guided heterogeneous memory system. We employ variable-strength error-correcting codes and tailor their error-correction mechanisms to the vulnerability profile of specific neurons to ensure a highly targeted approach for error mitigation. Our experimental results demonstrate that the proposed scheme enhances the neural network accuracy by 18% on average, while significantly reducing the fault-tolerance overhead.
Jing Wang 0055, Jinbin Zhu, Xin Fu 0001, Di Zang, Keyao Li, Weigong Zhang
IEEE Trans. Computers1
2024 Interference-Aware Online Optimization for Cellular-Connected Multiple UAV Networks With Energy Constraints
abstract
The incorporation of Unmanned Aerial Vehicles (UAVs) into cellular networks opens up new possibilities to enhance their ubiquitous operations and establish superior performance owing to the high probability of line-of-sight (LoS) for air-to-ground channels. However, this also results in the UAV inducing more significant uplink interference to non-associated Base Stations (BSs). This paper explores the online design policy in cellular-connected multiple UAV communications in the absence of channel conditions, focusing on wireless resource allocation and dynamic three-dimensional (3-D) path planning. Our objective is to maximize the minimum uplink throughput for all UAVs while considering the energy constraints of the UAVs. First, we implement an online design utilizing the achievable rate based on the estimated instantaneous channel state information (CSI) for the current time slot, and the expected data rate for future time slots based on channel distribution information (CDI). Our solution employs the exact penalty method along with alternating optimization and successive convex optimization methods. Second, we formulate an online design by merely using the achievable rate based on the estimated instantaneous CSI for the current time slot. We introduce an energy-triggered penalty term to regulate the energy consumption of the UAVs, resulting in a low-complexity solution even if the CDI is unavailable before the flight. Lastly, we conduct extensive simulations to corroborate our findings and provide comprehensive comparisons with other baseline schemes to underline the effectiveness of the proposed designs.
Cheng Zhan, Han Hu 0003, Zhi Liu 0002, Jing Wang 0055, Rongfei Fan
IEEE Trans. Mob. Comput.4
2024 Tradeoff Between Age of Information and Operation Time for UAV Sensing Over Multi-Cell Cellular Networks
abstract
Unmanned aerial vehicles (UAVs) have a significant potential for sensing applications in further cellular networks due to their extensive coverage and flexible deployment. In this paper, we consider a multi-cell cellular network with a cellular-connected UAV, which senses data with onboard sensors and uploads sensory data to the ground base stations (BSs). To evaluate the freshness of sensory data, we employ the concept of age of information (AoI), which is defined as the time elapsed since the latest successful transmission of sensory data. A lower AoI implies fresher sensory data, which may lead to the increase of UAV operation time. To balance such tradeoff, we aim to minimize the weighted sum of operation time and total AoI for the UAV by jointly optimizing transmission scheduling, BS association, as well as UAV trajectory. The problem is formulated as a mixed-integer nonlinear programming (MINLP) problem, which is difficult to solve due to the time-varying propagation channels. To this end, we first characterize the average communication performance with statistic channel information, and then develop a search algorithm to obtain the optimal solution via employing the optimal structure as well as convex optimization techniques, while a low-complexity Double Graph based Algorithm (DGA) is developed to obtain a suboptimal solution. Then, by taking into account the site-specific performance and making fast decisions online, we propose a Deep reinforcement Learning Algorithm (DLA). Compared to DGA, DLA can adapt to the specific local environment and obtain a solution more rapidly once the training process is completed. Simulation results show that the proposed algorithms outperform the benchmarks about 30%, and achieve flexible tradeoff between operation time and AoI of UAV sensing, which is not available by considering just one objective.
Cheng Zhan, Han Hu 0003, Jing Wang 0055, Zhi Liu 0002, Shiwen Mao
IEEE Trans. Mob. Comput.3
2024 Aerial Video Streaming Over 3D Cellular Networks: An Environment and Channel Knowledge Map Approach
abstract
Aerial video streaming is a promising application of unmanned aerial vehicles (UAVs), which extends video service from ground to three-dimensional (3D) airspaces. However, high data rates and smooth transmission are required along with ubiquitous and environment-aware communications. To this end, we study the quality of experience (QoE) maximization problem in this paper for aerial video streaming over 3D cellular networks in urban environments with building avoidance. Different from the typical channel model based optimization in prior works, we tackle the joint design of 3D UAV trajectory and transmission scheduling as well as playback rate adaption with an environment and channel knowledge map (ECKM) approach, which provides rich information about the location-specific channel for enabling environment-aware communications. Specifically, we first consider the scenario with perfect ECKM, and propose efficient algorithms to obtain suboptimal solutions by utilizing two graph models and the iterative parameter-enabled block coordinate descent method. For the scenario without such map information, we propose a dueling Deep Q-learning (DQL) solution with map construction such that the learning process can be facilitated for path planning. Simulation results are provided to demonstrate the improvement in QoE by the proposed solutions over baseline schemes, as well as a tradeoff between video quality and rate variation.
Cheng Zhan, Han Hu 0003, Zhi Liu 0002, Jing Wang 0055, Nan Cheng 0001, Shiwen Mao
IEEE Trans. Wirel. Commun.4
2023 Not All Resources are Visible: Exploiting Fragmented Shadow Resources in Shared-State Scheduler Architecture
abstract
With the rapid development of cloud computing, the increasing scale of clusters and task parallelism put forward higher requirements on the scheduling capability at scale. To this end, the shared-state scheduler architecture has emerged as the popular solution for large-scale scheduling due to its high scalability and utilization. In such an architecture, a central resource state view periodically updates the global cluster status to distributed schedulers for parallel scheduling. However, the schedulers obtain broader resource views at the cost of intermittently stale states, rendering resources released invisible to schedulers until the next view update. These fleeting resource fragments are referred to as shadow resources in this paper. Current shared-state solutions overlook or fail to systematically utilize the shadow resources, leaving a void in fully exploiting these invisible resources.
Xinkai Wang 0003, Yuancheng Li 0001, Chao Li 0009, Xiaofeng Hou, Jing Wang 0055, Quan Chen 0002, Jingwen Leng, Minyi Guo, Leibo Wang
SoCC6
2023 STAR: An Efficient Softmax Engine for Attention Model with RRAM Crossbar
abstract
RRAM crossbars have been studied to construct in-memory accelerators for neural network applications due to their in-situ computing capability. However, prior RRAM-based accelerators show efficiency degradation when executing the popular attention models. We observed that the frequent softmax operations arise as the efficiency bottleneck and also are insensitive to computing precision. Thus, we propose STAR, which boosts the computing efficiency with an efficient RRAM-based softmax engine and a fine-grained global pipeline for the attention models. Specifically, STAR exploits the versatility and flexibility of RRAM crossbars to trade off the model accuracy and hardware efficiency. The experimental results evaluated on several datasets show STAR achieves up to 30.63×and 1.31× computing efficiency improvements over the GPU and the state-of-the-art RRAM-based attention accelerators, respectively.
Yifeng Zhai, Bonan Yan, Jing Wang 0055
DATE4
2023 A Bit Level Acceleration of Mixed Precision Neural Network
abstract
With the growth of the convolutional neural network (CNN) parameters, the hardware resources become limited when deploying CNN models. Single bit-width quantization may lead to degradation of accuracy, while mixed-precision quantization models maintain higher accuracy. However, current mixed-precision quantization accelerators only consider the algorithm level quantization and fixed-bit-width processing elements (PEs) without fully utilizing the resources. Therefore, we propose a neural network acceleration architecture based on bit-level computational units to improve resource utilization and throughput of mixed-precision quantization accelerators. The 2-bit and 3-bit low-bit computational units are designed to implement the high-bit quantization. We also propose spatio-temporal fusion to satisfy the unique bit widths in each layer with mixed precision quantization. In particular, we use the 2-bit and 3-bit computational units to achieve dynamic layer-level quantization. We also discuss various combinations of different bit widths, which can be dynamically implemented according to the requirements of accuracy, execution time, etc. The proposed acceleration architecture is implemented in Verilog, verified using three network models: ResNet18, VGG7 and LeNet-5, and tested against three accelerators: Eyeriss, Stripes and Bit Fusion. The experimental results show that our accelerators provide more accuracy and the grouping operation reduces the area overhead. It provides a 3.08 to 3.54 times acceleration ratio and 2.4 to 2.57 times reduction in energy consumption on the three networks.
Dehui Qiu, Jing Wang 0055, Weigong Zhang, Lan Gao 0004
ICPADS2
2023 Accelerating Look-Up Table based Matrix Multiplication on GPUs
abstract
Multiplying matrices is among the most fundamental and compute-intensive operations in machine learning. Approximated Matrix Multiplication (AMM) based on table look-ups can significantly reduce the pressure on computing units and memory bandwidth, and has great potential in large-scale machine learning applications. In this work, we speed up table look-ups on GPUs to improve the performance of matrix multiplication. To avoid random memory accesses in table look-ups, we propose a novel warp-wide data sharing execution model. With this execution model, we develop a GPU AMM library to speed up MADDNESS (the state-of-the-art AMM), named GPU-MADDNESS. The experimental results show that GPU-MADDNESS improves the performance by 103X on average, and outperforms the tiling implementation by up to 42%.
Lan Gao 0004, Weigong Zhang, Jing Wang 0055, Dehui Qiu
ICPADS4
2023 NAS-SE: Designing A Highly-Efficient In-Situ Neural Architecture Search Engine for Large-Scale Deployment
abstract
The emergence of Neural Architecture Search (NAS) enables an automated neural network development process that potentially replaces manually-enabled machine learning expertise. A state-of-the-art NAS method, namely One-Shot NAS, has been proposed to drastically reduce the lengthy search time for a wide spectrum of conventional NAS methods. Nevertheless, the search cost is still prohibitively expensive for practical large-scale deployment with real-world applications. In this paper, we reveal that the fundamental cause for inefficient deployment of One-Shot NAS in both single-device and large-scale scenarios originates from the massive redundant off-chip weight access during the numerous DNN inference in sequential searching. Inspired by its algorithmic characteristics, we depart from the traditional CMOS-based architecture designs and propose a promising processing-in-memory design alternative to perform in-situ architecture search, which helps fundamentally address the redundancy issue. Moreover, we further discovered two major performance challenges of directly porting the searching process onto the existing PIM-based accelerators: severe pipeline contention and resource under-utilization. By leveraging these insights, we propose the first highly-efficient in-situ One-Shot NAS search engine design, named NAS-SE, for both single-device and large-scale deployment scenarios. NAS-SE is equipped with a two-phased network diversification strategy for eliminating resource contention, and a novel hardware mapping scheme for boosting the resource utilization by an order of magnitude. Our extensive evaluation demonstrates that NAS-SE significantly outperforms the state-of-the-art digital-based customized NAS accelerator (NASA) with an average speedup of 8.8 × and energy-efficiency improvement of 2.05 ×.
Qiyu Wan, Jing Wang 0055, Shuaiwen Song, Xin Fu 0001
MICRO3
2023 RLMixer: A Reinforcement Learning Approach for Integrated Ranking with Contrastive User Preference Modeling
Jing Wang 0055, Mengchen Zhao, Wei Xia 0001, Zhenhua Dong, Ruiming Tang, Rui Zhang 0003, Jianye Hao, Guangyong Chen, Pheng-Ann Heng
PAKDD (3)1
2023 High-Throughput GPU Random Walk with Fine-Tuned Concurrent Query Processing
abstract
Random walk serves as a powerful tool in dealing with large-scale graphs, reducing data size while preserving structural information. Unfortunately, existing system frameworks all focus on the execution of a single walker task in serial. We propose CoWalker, a high-throughput GPU random walk framework tailored for concurrent random walk tasks. It introduces a multi-level concurrent execution model to allow concurrent random walk tasks to efficiently share GPU resources with low overhead. Our system prototype confirms that the proposed system could outperform (up to 54%) the state-of-the-art in a wide spectral of scenarios.
Chao Li 0009, Pengyu Wang 0003, Xiaofeng Hou, Jing Wang 0055, Shixuan Sun, Minyi Guo, Dongbai Chen, Xiangwen Liu
PPoPP5
2023 Fargraph+: Excavating the parallelism of graph processing workload on RDMA-based far memory system
Jing Wang 0055, Chao Li 0009, Taolei Wang, Junyi Mei, Lu Zhang 0049, Pengyu Wang 0003, Minyi Guo
J. Parallel Distributed Comput.1
2023 Enabling High-Efficient ReRAM-Based CNN Training Via Exploiting Crossbar-Level Insignificant Writing Elimination
abstract
Convolutional neural networks (CNNs) have been widely adopted in many deep learning applications. However, training a deep CNN requests intensive data transfer, which is both time and energy consuming. Using resistive random-access memory (ReRAM) to process data locally in memory is an emerging solution to eliminate the massive data movement. However, training cannot be efficiently supported with current ReRAM-based PIM accelerators because of the frequent and high-cost ReRAM writing operations from the delay, energy, and ReRAM lifetime perspectives. In this paper, we observe that activation induced and weight updating induced writing operations dominate the training energy on ReRAM-based accelerators. We then exploit and leverage a new angle in intermediate data (e.g., activations and errors) sparsity that fits the unique computation pattern in ReRAM crossbars to effectively eliminate the insignificant ReRAM writings, thus, enabling highly efficient CNN training without hurting the training accuracy. The experiment results show our proposed scheme achieves averagely$4.97\times$($19.23\times$) energy saving and$1.38\times$($30.08\times$) speedup compared to the state-of-the-art ReRAM-based accelerator (GPU). Our scheme also achieves$4.55\times$lifetime enhancement compared to the state-of-the-art ReRAM accelerator.
Qiyu Wan, Peixun Ma, Jing Wang 0055, Mingsong Chen 0001, Shuaiwen Song, Xin Fu 0001
IEEE Trans. Computers4
2023 Optimizing GPU-Based Graph Sampling and Random Walk for Efficiency and Scalability
abstract
Graph sampling and random walk algorithms are playing increasingly important roles today because they can significantly reduce graph size while preserving structural information, thus enabling computationally intensive tasks on large-scale graphs. Current frameworks designed for graph sampling and random walk tasks are generally not efficient in terms of memory requirement and throughput. Not to mention that some of them result in biased results. To solve the above problems, we introduce Skywalker+, a high-performance graph sampling and random walk framework on multiple GPUs supporting multiple algorithms. Skywalker+ makes four key contributions: First, it realizes highly paralleled alias method on GPUs. Second, it applies finely adjusted workload-balancing techniques and locality-aware execution modes to present a highly efficient execution engine. Third, it optimizes the GPU memory usage with efficient buffering and data compression schemes. Last, it scales to multi-GPU to further enhance the system throughput. Abundant experiments show that Skywalker+ exhibits significant advantage over the baselines both in performance and utility.
Pengyu Wang 0003, Chao Li 0009, Jing Wang 0055, Taolei Wang, Lu Zhang 0049, Xiaofeng Hou, Minyi Guo
IEEE Trans. Computers4
2023 DRAGON: Dynamic Recurrent Accelerator for Graph Online Convolution
abstract
Despite the extraordinary applicative potentiality that dynamic graph inference may entail, its practical-physical implementation has been a topic seldom explored in literature. Although graph inference through neural networks has received plenty of algorithmic innovation, its transfer to the physical world has not found similar development. This is understandable since the most preeminent Euclidean acceleration techniques from CNN have little implication in the non-Euclidean nature of relational graphs. Instead of coping with the challenges arising from forcing naturally sparse structures into more inflexible stochastic arrangements, in DRAGON, we embrace this characteristic in order to promote acceleration. Inspired by high-performance computing approaches like Parallel Multi-moth Flame Optimization for Link Prediction (PMFO-LP), we propose and implement a novel efficient architecture, capable of producing similar speed-up and performance than baseline but at a fraction of its hardware requirements and power consumption. We leverage the hidden parallelistic capacity of our previously developed static graph convolutional processor ACE-GCN and expanded it with RNN structures, allowing the deployment of a multi-processing network referenced around a common pool of proximity-based centroids. Experimental results demonstrate outstanding acceleration. In comparison with the fastest CPU-based software implementation available in the literature, DRAGON has achieved roughly 191× speed-up. Under the largest configuration and dataset, DRAGON was also able to overtake a more power-hungry PMFO-LP by almost 1.59× in speed, and at around 89.59% in power efficiency. More importantly than raw acceleration, we demonstrate the unique functional qualities of our approach as a flexible and fault-tolerant solution that makes it an interesting alternative for an anthology of applicative scenarios.
José Romero Hung, Chao Li 0009, Taolei Wang, Jinyang Guo 0001, Pengyu Wang 0003, Chuanming Shao, Jing Wang 0055, Guoyong Shi, Xiangwen Liu
ACM Trans. Design Autom. Electr. Syst.7
2022 HyFarM: Task Orchestration on Hybrid Far Memory for High Performance Per Bit
abstract
Tapping into secondary memory resources, i.e., far memory (FM), has shown huge potential to improve the cost-efficiency of data centers. Recent advances in both storage-based vertical FM and network-based horizontal FM have raised new questions about leveraging hybrid FM tiers to achieve the best performance per bit of memory. It is still unclear how to efficiently place tasks when far memory access is enabled.In this work, we propose HyFarM, a novel task management strategy for hybrid FM clusters. We analyze FM sensitivity and cooperatively co-locate tasks to enable high utilization and scalability. Further, by tapping into dynamic memory adaption within and across servers, our strategy allows one to consistently deliver high performance on memory-intensive tasks. We evaluate our design with a heavily instrumented testbench. Compared with the state-of-the-art designs, HyFarM respectively improves memory utilization and the overall performance per bit (PPB) by up to 17.6% and 20.5%, with minor overhead.
Jing Wang 0055, Chao Li 0009, Junyi Mei, Taolei Wang, Pengyu Wang 0003, Lu Zhang 0049, Minyi Guo, Dongbai Chen, Xiangwen Liu
ICCD1
2022 Energy-Efficient Trajectory Optimization for Aerial Video Surveillance under QoS Constraints
abstract
Surveillance drones are unmanned aerial vehicles (UAVs) that are utilized to collect video recordings of targets. In this paper, we propose a novel design framework for aerial video surveillance in urban areas, where a cellular-connected UAV captures and transmits videos to the cellular network that services users. Fundamental challenges arise due to the limited onboard energy and quality of service (QoS) requirements over environment-dependent air-to-ground cellular links, where UAVs are usually served by the sidelobes of base stations (BSs). We aim to minimize the energy consumption of the UAV by jointly optimizing the mission completion time and UAV trajectory as well as transmission scheduling and association, subject to QoS constraints. The problem is formulated as a mixed-integer nonlinear programming (MINLP) problem by taking into account building blockage and BS antenna patterns. We first consider the average performance for uncertain local environments, and obtain an efficient sub-optimal solution by employing graph theory and convex optimization techniques. Next, we investigate the site-specific performance for specific urban local environments. By reformulating the problem as a Markov decision process (MDP), a deep reinforcement learning (DRL) algorithm is proposed by employing a dueling deep Q-network (DQN) neural network model with only local observations of sampled rate measurements. Simulation results show that the proposed solutions achieve significant performance gains over baseline schemes.
Cheng Zhan, Han Hu 0003, Shiwen Mao, Jing Wang 0055
INFOCOM4
2022 Excavating the Potential of Graph Workload on RDMA-based Far Memory Architecture
abstract
Disaggregated architecture brings new opportunities to memory -consuming applications like graph processing. It allows one to outspread memory access pressure from local to far memory, providing an attractive alternative to disk-based processing. Although existing works on general-purpose far mem-ory platforms show great potentials for application expansion, it is unclear how graph processing applications could benefit from disaggregated architecture, and how different optimization methods influence the overall performance. In this paper, we take the first step to analyze the impact of graph processing workload on disaggregated architecture by extending the GridGraph framework on top of the RDMA-based far memory system. We design Fargraph, a far memory coordi-nation strategy for enhancing graph processing workload. Specif-ically, Fargraph reduces the overall data movement through a well-crafted, graph-aware data segment offloading mechanism. In addition, we use optimal data segment splitting and asynchronous data buffering to achieve graph iteration-friendly far memory access. We show that Fargraph achieves near-oracle performance for typical in-local-memory graph processing systems. Fargraph shows up to 8.3 x speedup compared to Fastswap, the state-of-the-art, general-purpose far memory platform.
Jing Wang 0055, Chao Li 0009, Taolei Wang, Lu Zhang 0049, Pengyu Wang 0003, Junyi Mei, Minyi Guo
IPDPS1
2022 Oversubscribing GPU Unified Virtual Memory: Implications and Suggestions
abstract
Recent GPU architectures support unified virtual memory (UVM), which offers great opportunities to solve larger problems by memory oversubscription. Although some studies are concerned over the performance degradation under UVM oversubscription, the reasons behind workloads' diverse sensitivities to oversubscription is still unclear. In this work, we take the first step to select various benchmark applications and conduct rigorous experiments on their performance under different oversubscription ratios. Specifically,we take into account the variety of memory access patterns and explain applications' diverse sensitivities to oversubscription. We also consider prefetching and UVM hints, and discover their complex impact under different oversubscription ratios. Moreover, the strengths and pitfalls of UVM's multi-GPU support are discussed. We expect that this paper will provide useful experiences and insights for UVM system design.
Chuanming Shao, Jinyang Guo 0001, Pengyu Wang 0003, Jing Wang 0055, Chao Li 0009, Minyi Guo
ICPE4
2022 Adaptive Contention Management for Fine-Grained Synchronization on Commodity GPUs
abstract
As more emerging applications are moving to GPUs, fine-grained synchronization has become imperative. However, their performance can be severely impaired in case of frequent synchronization failures caused by high data contention. Differently from CPUs, GPUs own thousands of hardware threads and adopt single instruction multiple threads paradigm, making it impractical to deploy the CPU contention management mechanisms directly on GPUs. In this article, we design a Software Warp Controlling Framework (SWCF), which employs producer-consumer execution model and leverages GPU hardware barriers to dynamically control the execution of warps at runtime. On the basis of SWCF, we propose a contention management strategy to decrease frequent synchronization failures while avoiding the over-reducing of parallelism. We evaluate SWCF and the proposed strategy on commodity GPUs using a set of applications with fine-grained synchronization. The results show that on V100 GPU our contention management achieves a 4.7X speedup and outperforms the conventional GPU software backoff solution by 42% on average.
Lan Gao 0004, Jing Wang 0055, Weigong Zhang
ACM Trans. Archit. Code Optim.2
2022 Tapping into NFV Environment for Opportunistic Serverless Edge Function Deployment
abstract
Even with Network Function Virtualization (NFV), many commodity network servers have spare cycles. Despite that they are small and irregularly occur, spare cycles are fit for deploying short-lived serverless computing functions at the network edge. In this work, we perform detailed analyses of the benefits and limitations of co-locating serverless functions on NFV-ready servers. We proposeNEMO, a novel platform that enables efficient serverless edge function deployment in the NFV environment. NEMO can intelligently harvest spare cycles of network functions to warm up the serverless functions and speed up the function invocation in an agile manner. Besides, NEMO can judiciously manage the thread conflict in a resource-limited environment. We build a prototype of NEMO. Our thorough evaluations show that NEMO can harvest up to 41% spare cycles and achieve about 12.5$\sim$25X performance improvement compared with straightforward co-location.
Lu Zhang 0049, Weiqi Feng, Chao Li 0009, Xiaofeng Hou, Pengyu Wang 0003, Jing Wang 0055, Minyi Guo
IEEE Trans. Computers6
2021 Skywalker: Efficient Alias-Method-Based Graph Sampling and Random Walk on GPUs
abstract
Graph sampling and random walk operations, capturing the structural properties of graphs, are playing an important role today as we cannot directly adopt computing-intensive algorithms on large-scale graphs. Existing system frameworks for these tasks are not only spatially and temporally inefficient, but many also lead to biased results. This paper presents Skywalker, a high-throughput, quality-preserving random walk and sampling framework based on GPUs. Skywalker makes three key contributions: first, it takes the first step to realize efficient biased sampling with the alias method on a GPU. Second, it introduces well-crafted load-balancing techniques to effectively utilize the massive parallelism of GPUs. Third, it accelerates alias table construction and reduce the GPU memory requirement with efficient memory management scheme. We show that Skywalker greatly outperforms the state-of-the-art CPU-based and GPU-based baselines, in a wide spectrum of workload scenarios.
Pengyu Wang 0003, Chao Li 0009, Jing Wang 0055, Taolei Wang, Lu Zhang 0049, Jingwen Leng, Quan Chen 0002, Minyi Guo
PACT3
2021 Grus: Toward Unified-memory-efficient High-performance Graph Processing on GPU
abstract
Today’s GPU graph processing frameworks face scalability and efficiency issues as the graph size exceeds GPU-dedicated memory limit. Although recent GPUs can over-subscribe memory with Unified Memory (UM), they incur significant overhead when handling graph-structured data. In addition, many popular processing frameworks suffer sub-optimal efficiency due to heavy atomic operations when tracking the active vertices. This article presents Grus, a novel system framework that allows GPU graph processing to stay competitive with the ever-growing graph complexity. Grus improves space efficiency through a UM trimming scheme tailored to the data access behaviors of graph workloads. It also uses a lightweight frontier structure to further reduce atomic operations. With easy-to-use interface that abstracts the above details, Grus shows up to 6.4× average speedup over the state-of-the-art in-memory GPU graph processing framework. It allows one to process large graphs of 5.5 billion edges in seconds with a single GPU.
Pengyu Wang 0003, Jing Wang 0055, Chao Li 0009, Jianzong Wang, Haojin Zhu, Minyi Guo
ACM Trans. Archit. Code Optim.2
2021 ACE-GCN: A Fast Data-driven FPGA Accelerator for GCN Embedding
abstract
ACE-GCN is a fast and resource/energy-efficient FPGA accelerator for graph convolutional embedding under data-driven and in-place processing conditions. Our accelerator exploits the inherent power law distribution and high sparsity commonly exhibited by real-world graphs datasets. Contrary to other hardware implementations of GCN, on which traditional optimization techniques are employed to bypass the problem of dataset sparsity, our architecture is designed to take advantage of this very same situation. We propose and implement an innovative acceleration approach supported by our “implicit-processing-by-association” concept, in conjunction with a dataset-customized convolutional operator. The computational relief and consequential acceleration effect arise from the possibility of replacing rather complex convolutional operations for a faster embedding result estimation. Based on a computationally inexpensive and super-expedited similarity calculation, our accelerator is able to decide from the automatic embedding estimation or the unavoidable direct convolution operation. Evaluations demonstrate that our approach presents excellent applicability and competitive acceleration value. Depending on the dataset and efficiency level at the target, between 23× and 4,930× PyG baseline, coming close to AWB-GCN by 46% to 81% on smaller datasets and noticeable surpassing AWB-GCN for larger datasets and with controllable accuracy loss levels. We further demonstrate the unique hardware optimization characteristics of our approach and discuss its multi-processing potentiality.
José Romero Hung, Chao Li 0009, Pengyu Wang 0003, Chuanming Shao, Jinyang Guo 0001, Jing Wang 0055, Guoyong Shi
ACM Trans. Reconfigurable Technol. Syst.6
2020 Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture Design
abstract
In recent years, the CNNs have achieved great successes in the image processing tasks, e.g., image recognition and object detection. Unfortunately, traditional CNN's classification is found to be easily misled by increasingly complex image features due to the usage of pooling operations, hence unable to preserve accurate position and pose information of the objects. To address this challenge, a novel neural network structure called Capsule Network has been proposed, which introduces equivariance through capsules to significantly enhance the learning ability for image segmentation and object detection. Due to its requirement of performing a high volume of matrix operations, CapsNets have been generally accelerated on modern GPU platforms that provide highly optimized software library for common deep learning tasks. However, based on our performance characterization on modern GPUs, CapsNets exhibit low efficiency due to the special program and execution features of their routing procedure, including massive unshareable intermediate variables and intensive synchronizations, which are very difficult to optimize at software level. To address these challenges, we propose a hybrid computing architecture design named PIM-CapsNet. It preserves GPU's on-chip computing capability for accelerating CNN types of layers in CapsNet, while pipelining with an off-chip in-memory acceleration solution that effectively tackles routing procedure's inefficiency by leveraging the processing-in-memory capability of today's 3D stacked memory. Using routing procedure's inherent parallellization feature, our design enables hierarchical improvements on CapsNet inference efficiency through minimizing data movement and maximizing parallel processing in memory. Evaluation results demonstrate that our proposed design can achieve substantial improvement on both performance and energy savings for CapsNet inference, with almost zero accuracy loss. The results also suggest good performance scalability in optimizing the routing procedure with increasing network size.
Xingyao Zhang 0002, Shuaiwen Song, Chenhao Xie 0001, Jing Wang 0055, Weigong Zhang, Xin Fu 0001
HPCA4
2020 Multi-dimensional optimization for approximate near-threshold computing
abstract
The demise of Dennard’s scaling has created both power and utilization wall challenges for computer systems. As transistors operating in the near-threshold region are able to obtain flexible trade-offs between power and performance, it is regarded as an alternative solution to the scaling challenge. A reduction in supply voltage will nevertheless generate significant reliability challenges, while maintaining an error-free system that generates high costs in both performance and energy consumption. The main purpose of research on computer architecture has therefore shifted from performance improvement to complex multi-objective optimization. In this paper, we propose a three-dimensional optimization approach which can effectively identify the best system configuration to establish a balance among performance, energy, and reliability. We use a dynamic programming algorithm to determine the proper voltage and approximate level based on three predictors: system performance, energy consumption, and output quality. We propose an output quality predictor which uses a hardware/software co-design fault injection platform to evaluate the impact of the error on output quality under near-threshold computing (NTC). Evaluation results demonstrate that our approach can lead to a 28% improvement in output quality with a 10% drop in overall energy efficiency; this translates to an approximately 20% average improvement in accuracy, power, and performance.
Jing Wang 0055, Wei-wei Liang, Yuehua Niu, Lan Gao 0004, Weigong Zhang
Frontiers Inf. Technol. Electron. Eng.1
2020 Enabling Energy-Efficient and Reliable Neural Network via Neuron-Level Voltage Scaling
abstract
With the platforms of running deep neural networks (DNNs) move from large-scale data centers to handheld devices, power emerge as one of the most significant obstacles. Voltage scaling is a promising technique that enables power saving. Nevertheless, it raises reliability and performance concerns that may undesirably deteriorate NNs accuracy and performance. Consequently, an energy-efficient and reliable scheme is required for NNs to balance the above three aspects with satisfied user experience. To this end, we propose a neuron-level voltage scaling framework called NN-APP to model the impact of supply voltages on NNs from output accuracy (A), power (P), and performance (P) perspectives. We analyze the error propagation characteristics in NNs at both inter- and intra-network layers to precisely model the impact of voltage scaling on the final output accuracy at neuron-level. Furthermore, we combine a voltage clustering method and the multi-objective optimization to identify the optimal voltage islands and apply the same voltage to neurons with similar fault tolerance capability. We perform three case studies to demonstrate the efficacy of the proposed techniques.
Jing Wang 0055, Xin Fu 0001, Lan Gao 0004, Weigong Zhang
IEEE Trans. Computers1
2019 Reliability Enhancement of Neural Networks via Neuron-Level Vulnerability Quantization
Keyao Li, Jing Wang 0055, Xin Fu 0001, Xiufeng Sui, Weigong Zhang
ICA3PP (2)2
2019 Neuron Fault Tolerance Capability Based Computation Reuse in DNNs
Pengnian Qi, Jing Wang 0055, Weigong Zhang
ICA3PP (2)2
2019 Enabling Energy-Efficient and Reliable Neural Network via Neuron-Level Voltage Scaling
abstract
As the application scope of deep neural networks (DNNs) moves from large-scale data centers to small-scale mobile devices, power wall has become one of the most important obstacles. Voltage scaling is a typical technique enables power saving, but it causes reliability and performance challenges. Therefore, an energy-efficient and reliable scheme for NNs is required to balance above three aspects according to users' requirements for excellent user experience. In this paper, we innovatively propose neuron-level voltage scaling framework called NN-APP to model the impact of supply voltages on NNs from output accuracy (A), power (P), and performance (P) perspectives. We analyze the error propagation in NNs and precisely model the impact of voltage scaling on the final output accuracy at neuron-level. Multi-objective optimization and clustering method are combined to find the optimal voltage islands. Finally, we conduct experiment to demonstrate the efficacy of the proposed technique.
Jing Wang 0055, Xin Fu 0001, Xingyao Zhang 0002, Lan Gao 0004, Weigong Zhang, Tao Li 0006
ICPADS3
2019 DR Refresh: Releasing DRAM Potential by Enabling Read Accesses Under Refresh
abstract
Emerging data analytic workloads such as graph processing, neural network and edge data preprocesing desire efficient memory read operations. Unfortunately, due to the necessity of dynamic refresh, modern DRAM systems have to stall access during refresh cycles. As DRAM device density continues to grow, refresh operations can be a crucial throughput bottleneck. To fully unleash memory access performance, we revisit conventional refresh mechanism and DRAM architecture. We propose DR refresh, a specific refresh mechanism that enable read and refresh operations to be done simultaneously. We devise DR DRAM, a specific memory hardware system that can efficiently deploy DR refresh. Unlike traditional refresh, DR explores device refresh that only refreshes a designated device at a time. Meanwhile, DR increases read efficiency by recovering the inaccessible data that resides on a device under refreshing. We also propose Hybrid Refresh Main Memory (HRMM) which can designate refresh schemes (DR or traditional refresh) in a specific memory space. We expect that our design can benefit many real-life tasks such as SPEC CPU2006, CNN, LLT and PageRank.
Yuhai Cao, Chao Li 0009, Jing Wang 0055, Weigong Zhang, Quan Chen 0002, Jingwen Leng, Bin Yao 0002, Minyi Guo
IEEE Trans. Computers3
2018 In-Situ AI: Towards Autonomous and Incremental Deep Learning for IoT Systems
abstract
Recent years have seen an exploration of data volumes from a myriad of IoT devices, such as various sensors and ubiquitous cameras. The deluge of IoT data creates enormous opportunities for us to explore the physical world, especially with the help of deep learning techniques. Traditionally, the Cloud is the option for deploying deep learning based applications. However, the challenges of Cloud-centric IoT systems are increasing due to significant data movement overhead, escalating energy needs, and privacy issues. Rather than constantly moving a tremendous amount of raw data to the Cloud, it would be beneficial to leverage the emerging powerful IoT devices to perform the inference task. Nevertheless, the statically trained model could not efficiently handle the dynamic data in the real in-situ environments, which leads to low accuracy. Moreover, the big raw IoT data challenges the traditional supervised training method in the Cloud. To tackle the above challenges, we propose In-situ AI, the first Autonomous and Incremental computing framework and architecture for deep learning based IoT applications. We equip deep learning based IoT system with autonomous IoT data diagnosis (minimize data movement), and incremental and unsupervised training method (tackle the big raw IoT data generated in ever-changing in-situ environments). To provide efficient architectural support for this new computing paradigm, we first characterize the two In-situ AI tasks (i.e. inference and diagnosis tasks) on two popular IoT devices (i.e. mobile GPU and FPGA) and explore the design space and tradeoffs. Based on the characterization results, we propose two working modes for the In-situ AI tasks, including Single-running and Co-running modes. Moreover, we craft analytical models for these two modes to guide the best configuration selection. We also develop a novel two-level weight shared In-situ AI architecture to efficiently deploy In-situ tasks to IoT node. Compared with traditional IoT systems, our In-situ AI can reduce data movement by 28-71%, which further yields 1.4X-3.3X speedup on model update and contributes to 30-70% energy saving.
Mingcong Song, Kan Zhong, Jiaqi Zhang 0002, Yang Hu 0001, Duo Liu 0002, Weigong Zhang, Jing Wang 0055, Tao Li 0006
HPCA7
2018 DR DRAM: Accelerating Memory-Read-Intensive Applications
abstract
Today, many data analytic workloads such as graph processing and neural network desire efficient memory read operation. The need for preprocessing various raw data also demands enhanced memory read bandwidth. Unfortunately, due to the necessity of dynamic refresh, modern DRAM system has to stall memory access during each refresh cycle. As DRAM device density continues to grow, the refresh time also needs to extend to cover more memory rows. Consequently, DRAM refresh operation can be a crucial throughput bottleneck for memory read intensive (MRI) data processing tasks. To fully unleash the performance of these applications, we revisit conventional DRAM architecture and refresh mechanism. We propose DR DRAM, an application-specific memory design approach that makes a novel tradeoff between read and write performance. Simply put, DR has two layers of meaning: device refresh and data recovery. It aims at eliminating stall by enabling read and refresh operations to be done simultaneously. Unlike traditional schemes, DR explores device refresh that only refreshes a specific device at a time. Meanwhile, DR increases read efficiency by recovering the inaccessible data that resides on a device under refreshing. Our design can be implemented on existing redundant data storage area on DRAM. In this paper we detail DR's architecture and protocol design. We evaluate it on a cycle accurate simulator. Our results show that DR can nearly eliminate refresh overhead for memory read operation and brings up to 12% extra maximum read bandwidth and 50~60% latency improvement on present DRR4 device.
Yuhai Cao, Chao Li 0009, Quan Chen 0002, Jingwen Leng, Minyi Guo, Jing Wang 0055, Weigong Zhang
ICCD6
2018 CounterMiner: Mining Big Performance Data from Hardware Counters
abstract
Modern processors typically provide a small number of hardware performance counters to capture a large number of microarchitecture events. These counters can easily generate a huge amount (e.g., GB or TB per day) of data, which we call big performance data in cloud computing platforms with more than thousands of servers and millions of complex workloads running in a "24/7/365" manner. The big performance data provides a precious foundation for root cause analysis of performance bottlenecks, architecture and compiler optimization, and many more. However, it is challenging to extract value from the big performance data due to: 1) the many unperceivable errors (e.g., outliers and missing values); and 2) the difficulty of obtaining insights, e.g., relating events to performance. In this paper, we propose CounterMiner, a rigorous methodology that enables the measurement and understanding of big performance data by using data mining and machine learning techniques. It includes three novel components: 1) using data cleaning to improve data quality by replacing outliers and filling in missing values; 2) iteratively quantifying, ranking, and pruning events based on their importance with respect to performance; 3) quantifying interaction intensity between two events by residual variance. We use sixteen benchmarks (eight from CloudSuite and eight from the Spark version of HiBench) to evaluate CounterMiner. The experimental results show that CounterMiner reduces the average error from 28.3% to 7.7% when multiplexing 10 events on 4 hardware counters. We also conduct a real-world case study, showing that identifying important configuration parameters of Spark programs by event importance is much faster than directly ranking the importance of these parameters.
Yirong Lv, Qingyi Luo, Jing Wang 0055, Zhibin Yu 0001, Xuehai Qian
MICRO4
2018 Towards Memory Friendly Long-Short Term Memory Networks (LSTMs) on Mobile GPUs
abstract
Intelligent Personal Assistants (IPAs) with the capability of natural language processing (NLP) are increasingly popular in today's mobile devices. Recurrent neural networks (RNNs), especially one of their forms – Long-Short Term Memory networks (LSTMs), are becoming the core machine learning technique applied in the NLP-based IPAs. With the continuously improved performance of mobile GPUs, local processing has become a promising solution to the large data transmission and privacy issues induced by the cloud-centric computations of IPAs. However, LSTMs exhibit quite inefficient memory access pattern when executed on mobile GPUs due to the redundant data movements and limited off-chip bandwidth. In this study, we aim to explore the memory friendly LSTM on mobile GPUs by hierarchically reducing the off-chip memory accesses. To address the redundant data movements, we propose inter-cell level optimizations that intelligently parallelize the originally sequentially executed LSTM cells (basic units in RNNs, corresponding to neurons in CNNs) to improve the data locality across cells with negligible accuracy loss. To relax the pressure on limited off-chip memory bandwidth, we propose intra-cell level optimizations that dynamically skip the loads and computations of rows in the weight matrices with trivial contribution to the outputs. We also introduce a light-weighted module to the GPUs architecture for the runtime row skipping in weight matrices. Moreover, our techniques are equipped with thresholds which provide a unique tunning space for performance-accuracy trade-offs directly guided by the user preferences. The experimental results show our optimizations achieves substantial improvements on both performance and power with user-imperceptible accuracy loss. And our optimizations exhibit the strong scalability with the increasing input data set. Our user study also shows that our designed system delivers the excellent user experience.
Xingyao Zhang 0002, Chenhao Xie 0001, Jing Wang 0055, Xin Fu 0001
MICRO3
2018 QIG: Quantifying the Importance and Interaction of GPGPU Architecture Parameters
abstract
Graphic processing units (GPUs) are widely used for general-purpose computing-so-called GPGPU computing. GPUs feature a large number of architecture parameters, resulting in a huge design space. To quickly explore this design space and identify the optimum architecture for a group of widely used computing kernels, it is critical to know how important each parameter is and how strongly these parameters interact with each other. This paper proposes an ensemble-learning-based approach, called quantifying the importance and interaction of Gpgpu architecture parameters (QIG), to quantify the importance of architecture parameters and their interactions with respect to performance. QIG employs a stochastic gradient boosted regression tree to construct performance models using performance data from a random set of GPU architectures. Leveraging these models, QIG observes the impact of each architecture parameter on performance, and calculates its importance and interaction intensity with other parameters. Using 25 widely used GPGPU kernels, we demonstrate that QIG accurately ranks the importance and interaction of GPU architecture parameters while the previously proposed Plackett-Burman design does not. Moreover, we show that QIG leads to a substantially more accurate performance model compared to prior work, including Starchart and approaches using artificial neural networks and supported vector machines: average error of 4.2% for QIG versus 23+% for prior work. Finally, QIG reveals a number of interesting insights for GPU architectures running GPGPU workloads.
Zhibin Yu 0001, Jing Wang 0055, Lieven Eeckhout, Cheng-Zhong Xu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2017 Processing-in-Memory Enabled Graphics Processors for 3D Rendering
abstract
The performance of 3D rendering of Graphics Processing Unit that converts 3D vector stream into 2D frame with 3D image effects significantly impacts users gaming experience on modern computer systems. Due to its high texture throughput requirement, main memory bandwidth becomes a critical obstacle for improving the overall rendering performance. 3D-stacked memory systems such as Hybrid Memory Cube provide opportunities to significantly overcome the memory wall by directly connecting logic controllers to DRAM dies. Although recent works have shown promising improvement in performance by utilizing HMC to accelerate special-purpose applications, a critical challenge of how to effectively leverage its high internal bandwidth and computing capability in GPU for 3D rendering remains unresolved. Based on the observation that texel fetches greatly impact off-chip memory traffic, we propose two architectural designs to enable Processing-In-Memory based GPU for efficient 3D rendering. Additionally, we employ camera angles of pixels to control the performance-quality tradeoff of 3D rendering. Extensive evaluation across several real-world games demonstrates that our design can significantly improve the performance of texture filtering and 3D rendering by an average of 3.97X (up to 6.4X) and 43% (up to 65%) respectively, over the baseline GPU. Meanwhile, our design provides considerable memory traffic and energy reduction without sacrificing rendering quality.
Chenhao Xie 0001, Shuaiwen Song, Jing Wang 0055, Weigong Zhang, Xin Fu 0001
HPCA3
2017 Data re-allocation enabled cache locking for embedded systems
Chun Jason Xue, Keni Qiu, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Mengying Zhao
J. Syst. Archit.4
2017 On the Implication of NTC versus Dark Silicon on Emerging Scale-Out Workloads: The Multi-Core Architecture Perspective
abstract
The end of Dennard's scaling poses computer systems, especially the datacenters, in front of both power and utilization walls. One possible solution to combat the power and utilization walls is dark silicon where transistors are under-utilized in the chip, but this will result in a diminishing performance. Another solution is Near-Threshold Voltage Computing (NTC) which operates transistors in the near-threshold region and provides much more flexible tradeoffs between power and performance. However, prior efforts largely focus on a specific design option based on the legacy desktop applications, therefore, lacking comprehensive analysis of emerging scale-out applications with multiple design options when dark silicon and/or NTC are/is applied. In this paper, we characterize different perspectives including performance, energy efficiency and reliability in the context of NTC/dark silicon cloud processors running emerging scale-out workloads on various architecture designs. We find NTC is generally an effective way to alleviate the power challenge over scale-out applications compared with dark silicon, it can improve performance by 1.6X, energy efficiency by 50 percent and the reliability problem can be relieved by ECC. Meanwhile, we also observe tiled-OoO architecture improves the performance by 20~370 percent and energy efficiency by 40~600 percent over alternative architecture designs, making it a preferable design paradigm for scale-out workloads. We believe that our observations will provide insights for the design of cloud processors under dark silicon and/or NTC.
Jing Wang 0055, Xin Fu 0001, Weigong Zhang, Keni Qiu, Tao Li 0006
IEEE Trans. Parallel Distributed Syst.1
2016 Refresh-aware loop scheduling for high performance low power volatile STT-RAM
abstract
The highlighted advantages of low leakage power, high storage density and immunity to electronic magnetic radiation make STT-RAM a promising candidate to build cache, SPM or main memory in embedded systems. However, write operations on STT-RAM have considerably longer latency and higher energy consumption than conventional SRAM. To solve this problem, researchers have proposed to relax STT-RAM's non-volatility and to have it work in a fast and low power mode. Under this volatile mode, refresh operations are needed to guarantee data correctness if their lifespan is larger than the retention time. It is observed that this refresh overhead is significant for data in stencil loops with the characteristic of constant read and write dependencies. This paper proposes a loop scheduling technique which can traverse loops in a new direction such that data lifespan can be greatly shortened. Therefore, overall refresh overhead can be efficiently mitigated so as to improve performance and reduce power consumption. The experimental results indicate that access latency and dynamic energy can be improved by 21.4~96.0% and 22.0~95.5% respectively by the proposed scheduling scheme.
Keni Qiu, Junpeng Luo, Zhiyao Gong, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Tao Li 0006, Chun Jason Xue
ICCD5
2016 An adaptive Non-Uniform Loop Tiling for DMA-based bulk data transfers on many-core processor
abstract
Mesh Network-on-Chip (NoC) is a key fabric to interconnect many cores with desirable scalability, reliability and interoperability. We observe that DMA-based bulk data block transfer exhibits non-negligible NoC latency due to heavy congestions. Loop tiling is an effective way to partition data space for SPM+DMA-based data block transfer. Nevertheless, we observe that the unbalanced NoC latency can degrade the effectiveness of loop tiling in a uniform fashion. In this paper, we propose a NoC-aware Non-Uniform Loop Tiling (NULT) scheme to improve DMA performance. A NULT framework is built on the proposed model to adaptively hide DMA latency into computation time and reduce the overall execution time. The framework first groups cores into different families taking into account their distance-to-data in NoC. Then a heuristic method is presented to solve the near optimal tiling factors for each core family. In this way, different core families are assigned non-uniform tiling sizes. We evaluate the NULT scheme on the NIRGAM platform. Compared to the traditional uniform tiling approach, the proposed NULT technique shows more benefit to overlap memory access time and computation time and thus reduce the overall execution time of a loop nest.
Keni Qiu, Yuanhui Ni, Weigong Zhang, Jing Wang 0055, Chun Jason Xue, Tao Li 0006
ICCD4
2016 Exploring Variation-Aware Fault-Tolerant Cache under Near-Threshold Computing
abstract
Near threshold voltage computing enables transistor voltage scaling to continue with Moore's Law projection and dramatically improves power and energy efficiency. However, reducing the supply voltage to near-threshold level significantly increases the susceptibility of on-chip caches to process variations, leading to the high error rate. Most existing fault-tolerant schemes significantly sacrifice cache capacity and performance. In this paper, we propose a novel fault-tolerant cache architecture at near-threshold computing, which is suitable for high error rate memories. We first propose a variation-aware skewed-associative cache, and then redirect the faulty blocks to the error-free blocks based on it to explore the fault-tolerance cache design. Unlike previous cache reconfiguration schemes for the fault tolerance, our cache design does not need to sacrifice or disable any fault-free blocks to form a completely functional set. We use all error-free blocks and have the least cache capacity waste. More importantly, since the aging impact could also cause cell failures, our skewed cache takes the aggregated process variation and aging impact into the consideration. Last but not least, our skewed cache design avoids the complex remapping from faulty blocks to the error-free blocks and minimizes the hardware overheads. Our evaluation results show that our variation-aware fault-tolerant cache design exhibits strong capability to tolerate the high error rate, and more excitingly, its effectiveness on reducing the cache miss rate and improving the performance is even more obvious as the supply voltage scales down to the near-threshold region.
Jing Wang 0055, Yanjun Liu 0005, Weigong Zhang, Kezhong Lu, Keni Qiu, Xin Fu 0001, Tao Li 0006
ICPP1
2013 SPIRE: improving dynamic binary translation through SPC-indexed indirect branch redirecting
abstract
Dynamic binary translation system must perform an address translation for every execution of indirect branch instructions. The procedure to convert Source binary Program Counter (SPC) address to Translated Program Counter (TPC) address always takes more than 10 instructions, becoming a major source of performance overhead. This paper proposes a novel mechanism called SPc-Indexed REdirecting (SPIRE), which can significantly reduce the indirect branch handling overhead. SPIRE doesn't rely on hash lookup and address mapping table to perform address translation. It reuses the source binary code space to build a SPC-indexed redirecting table. This table can be indexed directly by SPC address without hashing. With SPIRE, the indirect branch can jump to the originally SPC address without address translation. The trampoline residing in the SPC address will redirect the control flow to related code cache. Only 2-6 instructions are needed to handle an indirect branch execution. As part of the source binary would be overwritten, a shadow page mechanism is explored to keep transparency of the corrupt source binary code page. Online profiling is adopted to reduce the memory overhead.
Ning Jia 0004, Jing Wang 0055, Dong Tong 0001
VEE3
2010 Cache Management with Partitioning-Aware Eviction and Thread-Aware Insertion/Promotion Policy
abstract
With recent advances of processor technology, the LRU based shared last-level cache (LLC) has been widely employed in modern Chip Multi-processors (CMP). However, past research indicates that the cache performance of the LLC and further of the CMP processors may be degraded severely by LRU under the occurrence of the inter-thread interference or the excess of the working set size over the cache size. Existing approaches tackling this performance degradation problem have limited improvement of an overall cache performance because they usually focus on a single type of memory access behavior and thus lack full consideration of tradeoffs among different types of memory access behaviors. In this paper, we propose a unified cache management policy called Partitioning-Aware Eviction and Thread-aware Insertion/Promotion policy (PAE-TIP) that can effectively enhance capacity management, adaptive insertion/promotion, and further improve the overall cache performance. Specifically, PAE-TIP employs an adaptive mechanism to decide the position where to put the incoming lines or to move the hit lines, and chooses a victim line based on the target partitioning given by utility-based cache partitioning (UCP). In our study, we show that PAE-TIP can cover a variety of memory access behaviors simultaneously and provide a good tradeoff for overall cache performance improvement while retaining competitively low hardware and design overhead. The evaluation conducted on 4-way CMPs shows that the PAE-TIP-managed LLC can improve overall performance by19.3% on average over the LRU policy. Furthermore, the performance benefit of PAE-TIP is 1.09x compared to PIPP, 1.11x compared to TADIP and 1.12x compared to UCP.
Junmin Wu, Xiufeng Sui, Jing Wang 0055, Guoliang Chen 0001
ISPA5