VLDB 2026 Research / reviewers in the wild / expert
Mengying Zhao
dblp:73/11467
· DBLP profile ↗
93ranked-venue papers
13as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 77 · 13 first-author · 25 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 4 since 2021Computer networks · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SysVCoder: An LLM-Driven Framework for Systematic Generation of System-Level Design
Jian Zuo, Junzhe Liu, Xianyong Wang, Navya Goli, Umamaheswara Rao Tida, Zhenge Jia, Zhaoyan Shen, Mengying Zhao |
APPT | 9 |
| 2026 | Fault-tolerance Mapping of Spiking Neural Networks to RRAM-based Neuromorphic HardwareabstractSpiking neural networks (SNNs) have been widely used in artificial intelligence applications. Resistive random-access memory (RRAM) based neuromorphic hardware can be used for low-power and high-speed inference of SNNs. However, RRAM devices suffer from stuck-at-fault (SAF) defects due to the immature fabrication process. SAF defects can lead to incorrect weights of SNNs and thus severely degrade the inference accuracy. In this paper, we propose a fault-tolerance synaptic-to-RRAM mapping scheme to deploy SNNs while protecting the inference accuracy. We first explore how different weights affect the model accuracy in SNNs and find that the frequency of spikes plays a vital role. Motivated by this, we develop an SNN-oriented metric to evaluate the importance of weights. Then we propose a defect-aware mapping scheme based on a simulated-annealing framework to efficiently map synapses to RRAM to mitigate the impact of SAFs. Evaluation shows that the proposed strategy can improve the accuracy by 18.74% on average compared to existing mapping strategy. Yuqing Xiong, Cao Xiao, Mengying Zhao |
DATE | 5 |
| 2026 | Unity is Power: Semi-Asynchronous Collaborative Training of Large-Scale Models With Structured Pruning in Resource-Limited ClientsabstractIn this work, we study to release the potential of massive heterogeneous weak computing power to collaboratively train large-scale models on dispersed datasets. In order to improve both efficiency and accuracy in resource-adaptive collaborative learning, we take the first step to consider the unstructured pruning, varying submodel architectures, knowledge loss, and straggler challenges simultaneously. We propose a novel semiasynchronous collaborative training framework, namely Co-S2P, with data distribution-aware structured pruning and cross-block knowledge transfer mechanism to address the above concerns. Furthermore, we provide theoretical proof that Co-S2P can achieve asymptotic optimal convergence rate of O(1/√ N∗EQ). Finally, we conduct extensive experiments on two types of tasks with a real-world hardware testbed including diverse IoT devices. The experimental results demonstrate that Co-S2P improves accuracy by up to 8.8% and resource utilization by up to 1.2× compared to state-of-the-art methods, while reducing memory consumption by approximately 22% and training time by about 24% on all resource-limited devices. Xiao Zhang 0015, Feng Chen 0005, Yuan Yuan 0040, Yifei Zou, Mengying Zhao, Jianbo Lu 0001, Dongxiao Yu |
IEEE Trans. Mob. Comput. | 8 |
| 2025 | Exploiting Large Language Models for Software-Defined Solid-State Drives Design
Tianren Zhou, Zhenge Jia, Mengying Zhao, Zhaoyan Shen |
APPT | 6 |
| 2025 | Enabling On-Tiny-Device Model Personalization via Gradient Condensing and Alternant Partial UpdateabstractOn-device training enables the model to adapt to user-specific data by fine-tuning a pre-trained model locally. As embedded devices become ubiquitous, on-device training is increasingly essential since users can benefit from the personalized model without transmitting data and model parameters to the server. Despite significant efforts toward efficient training, ondevice training still faces a major challenge: The prohibitive cost of multi-layer backpropagation strains the limited resources of tiny devices. In this paper, we propose an algorithm-system cooptimization framework TinyMP that enables self-adaptive on-tiny-device model personalization. To mitigate backpropagation costs, we introduce Gradient Condensing to condense the gradient map structure, significantly reducing the computational complexity and memory consumption of backpropagation while preserving model performance. To further reduce computation overhead, we propose Alternant Partial Update, a mechanism that locally and alternatively selects essential parameters to update without requiring retraining or offline evolutionary search. Our framework is evaluated through extensive experiments using various CNN models (e.g., MobileNetV2, MCUNet) on embedded devices with minimal resources (e.g., OpenMV-H7 with less than 1MB SRAM and 2 MB Flash). Experimental results show that our framework achieves up to $2.4 \times$ speedup, 80.8% memory saving, and 30.3% accuracy improvement on downstream tasks, outperforming SOTA approaches. Zhenge Jia, Yiyang Shi, Zeyu Bao, Xin Pang, Huiguo Liu, Zhaoyan Shen, Mengying Zhao |
DAC | 9 |
| 2025 | Routability-aware Packing for High-density Nonvolatile FPGAsabstractNonvolatile field-programmable gate arrays (NVFPGAs) can use multi-level cell (MLC) nonvolatile memories (NVMs) to enhance their logic density. However, the highdensity design of NVFPGAs degrades the intra-routability of configurable logic blocks (CLBs), which significantly prolongs the time consumed by the packing process in the computer-aided design (CAD) flow. To relieve the efficiency degradation, in this paper, we propose a routability-aware re-pair stage to adjust the logical-physical look-up table (LUT) assignments to mitigate the congestion and improve their intra-routability, thereby reducing the packing time. In addition, exploiting the structural equivalence of MLC LUTs, we remove unnecessary intra-routing attempts from packing to further improve efficiency. Evaluation shows the proposed strategies reduce packing time by $41.48 \%$ on average. Index Terms-nonvolatile memory (NVM), multi-level cell (MLC), field-programmable gate array (FPGA), computer-aided design (CAD), packing. Huichuan Zheng, Yuqing Xiong, Jian Zuo, Zhenge Jia, Mengying Zhao |
DAC | 6 |
| 2025 | A Deep Dive into Protocol Design: How to Improve IPFS Performance without Sacrificing DecentralizationabstractThe InterPlanetary File System (IPFS) is a prominent decentralized storage solution; however, it struggles with performance issues. To address this challenge, the IPFS team has patched a series of centralized components, resulting in improved performance while giving more significant roles to specific entities. Balancing speed and decentralization has always posed a complex dilemma for storage systems. In this paper, we conduct a series of experiments and analyses to identify the performance advantages and constraints associated with the IPFS decentralized protocol. Based on thorough analysis, we propose a novel scheme named XIPFS to optimize IPFS. This scheme includes facilitating parallel block exchange across multiple nodes, refining content Publication strategies, and improving node selection algorithms for content routing. Our goal is to maximize the benefits of decentralized multi-source parallel downloading while minimizing the negative impact of decentralized indexing on execution time. Compared to previous approaches, the proposed optimizations are lightweight and fully compatible with the decentralized protocol, enabling autonomous execution and utility realization at each node. Experimental results demonstrate that XIPFS significantly improves node performance without compromising decentralization. Zhaoyan Shen, Mengying Zhao, Dongxiao Yu, Bingzhe Li |
ICDE | 3 |
| 2025 | Demo: Real-Time Inference on GPU-Based Heterogeneous SoCs with GPU Cache Locking
Kehao Ma, Wei Zhang 0173, Mengying Zhao, Lei Ju 0001 |
RTCSA | 3 |
| 2025 | FedMQ+: Towards efficient heterogeneous federated learning with multi-grained quantization
Mei Cao, Yuan Yuan 0040, Jianbo Lu 0001, Xiaojun Cai, Dongxiao Yu, Mengying Zhao |
J. Syst. Archit. | 7 |
| 2025 | CoaCAD: Correlation-Assisted Computer-Aided Design for Nonvolatile FPGAsabstractNonvolatile field-programmable gate arrays (FPGAs) offer advantages in terms of high logic density and near-zero leakage power when contrasted with conventional static random access memory-based FPGAs. However, they have a lifetime issue. To deal with this problem, a series of configuration files can be generated with various logical-to-physical mappings. This enables intensive writing to be distributed across different physical regions for wear leveling. Currently, the configuration files are independently generated, which is time consuming. In this article, we propose to investigate correlations and use them to assist the computer-aided design (CAD) flow to speed up the procedure of generating configuration files. First, we develop dynamic probabilities to drive the swapping of placement stage in CAD flow, so as to push components to locate appropriate positions quickly. Second, we design the congestion information inheritance strategy to adjust routing parameters in the routing stage, aiming to reduce the number of routing attempts. Evaluation shows that the proposed schemes can deliver 44.15% decrease in placement and routing runtime, while maintaining comparable performance and lifetime, when compared with existing strategies. Mengying Zhao, Yuqing Xiong, Huichuan Zheng, Dongxiao Yu, Zhaoyan Shen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Cache-aware Task Decomposition for Efficient Intermittent Computing SystemsabstractEnergy harvesting offers a scalable and cost-effective power solution for IoT devices, but it introduces the challenge of frequent and unpredictable power failures due to the unstable environment. To address this, intermittent computing has been proposed, which periodically backs up the system state to non-volatile memory (NVM), enabling robust and sustainable computing even in the face of unreliable power supplies. In modern processors, write back cache is extensively utilized to enhance system performance. However, it poses a challenge during backup operations as it buffers updates to memory, potentially leading to inconsistent system states. One solution is to adopt a write-through cache, which avoids the inconsistency issue but incurs increased memory access latency for each write reference. Some existing work enforces a cache flushing before backups to maintain a consistent system state, resulting in significant backup overhead. In this paper, we point out that although cache delays updates to the main memory, it may preserve a recoverable system state in the main memory. Leveraging this characteristic, we propose a cache-aware task decomposition method that divides an application into multiple tasks, ensuring that no dirty cache lines are evicted during their execution. Furthermore, the cache-aware task decomposition maintains an unchanged memory state during the execution of each task, enabling us to parallelize the backup process with task execution and effectively hide the backup latency. Experimental results with different power traces demonstrate the effectiveness of the proposed system. Wei Zhang 0173, Mengying Zhao, Zimeng Zhou, Lei Ju 0001 |
DAC | 3 |
| 2024 | Decentralized Federated Learning in Partially Connected Networks with Non-IID DataabstractFederated learning is a promising paradigm to enable joint model training across distributed data while preserving data privacy. The distributed data are usually not identically and in-dependently distributed (Non-IID), which brings great challenges for federated learning. There have been existing work proposing to guide model aggregation between similar clients to deal with Non-IID data. But they typically assume a fully connected network topology, while new design issues need to be considered when it comes to a partially connected topology. In this work, we propose a probability-driven gossip framework for partially connected network topology with Non-IID data. The main idea is to discover similarity relationship between non-adjacent clients and guide the model exchange to encourage aggregation between similar clients. We explore cross-node similarity assessment and define probability to guide the model exchange and aggregation. Both similarity and communication cost are considered in the probability-driven gossip. Evaluation shows that the proposed scheme can achieve 13.04%-14.24% improvement in model accuracy, when compared with related work. Xiaojun Cai, Nanxiang Yu, Mengying Zhao, Mei Cao, Jianbo Lu 0001 |
DATE | 3 |
| 2024 | Towards High-Throughput Neural Network Inference with Computational BRAM on Nonvolatile FPGAsabstractField-programmable gate arrays (FPGAs) have been widely used in artificial intelligence applications. As the capacity requirements of both computation and memory resources continuously increase, emerging nonvolatile memory has been proposed to replace static random access memory (SRAM) in FPGAs to build nonvolatile FPGAs (NV-FPGAs), which have advantages of high density and near-zero leakage power. Features of emerging nonvolatile memory should be fully explored to improve performance, energy efficiency as well as lifetime of NV-FPGAs. In this paper, we study an intrinsic characteristic of emerging nonvolatile memory, i.e., computing-in-memory, in nonvolatile block random access memory (BRAM) of NV-FPGAs. Specifically, we present a computational BRAM architecture (C-BRAM), and propose a computational density aware operator allocation strategy to fully utilize C-BRAM. Neural network inference is taken as an example to evaluate the proposed architecture and strategy, showing 68% and 62% improvement in computational density compared to traditional SRAM-based FPGA and existing NV-FPGA, respectively. Mengying Zhao, Huichuan Zheng, Yuqing Xiong, Yuhao Zhang 0006, Zhaoyan Shen |
DATE | 2 |
| 2024 | Towards Efficient Reconfiguration through Lightweight Input Inversion for MLC NVFPGAsabstractNonvolatile field programmable gate arrays (NVFP-GAs) have been proposed to address the challenges raised by artificial intelligence and big data related applications, since nonvolatile memories (NVMs) introduce advantages of high storage density, low leakage power, and high system robustness. In addition, multi-level cell (MLC), which can store multiple bits within one memory cell, further improves the logic density of NVFPGAs. However, the inefficient write operation of MLC NVM significantly increases the reconfiguration cost in aspects of energy, latency, and lifetime. In this paper, we focus on the reconfiguration cost of MLC LUTs in NVFPGA and propose a lightweight input inversion based scheme to reduce the reconfiguration cost. Inversion flexibility is defined and modeled for LUT inputs to guide the proposed scheme. We also discuss how the proposed scheme can be combined with other existing write reduction strategies. Evaluation shows the proposed scheme can reduce reconfiguration cost by 10.01 % with negligible overhead. Huichuan Zheng, Mengying Zhao, Yuqing Xiong, Xiaojun Cai, Zhiping Jia |
DATE | 2 |
| 2024 | Federation-Paced Learning: Towards Efficient Federated Learning with Synchronized PaceabstractFederated learning (FL) is a distributed machine learning approach that allows multiple devices or computing nodes to jointly train models without sharing raw data. However, in real-world application scenarios, FL usually encounters a critical challenge of data heterogeneity. Recent studies have revealed that the client’s model suffers severe bias between the local model and global model, leading to global performance degradation. Improving the generalization of local learning would inherently reduce bias. It has been proved that self-paced learning on a single device can greatly achieve a better generalization result. However, it is not well explored how it can be applied to federated learning with a number of distributed nodes working cooperatively. Specifically, self-paced learning suggests using easy data and then gradually difficult data during model training. It is not straightforward to differentiate “easy” and “difficult” data at the local since global data distribution is not available, especially with severe data heterogeneity. To address the above issues, we propose a novel federated learning framework, Federation-Paced Learning (FedPL), which enables a self-paced process in federated learning and effectively improves the model performance. First, we propose schemes to analyze the data characteristics in terms of difficulty. Then we define a stage controller to synchronize the learning process across cooperative nodes to follow the easy-to-hard rule. Finally, we propose a client selection strategy to further improve the learning efficacy. We evaluate the performance of FedPL on several generic public datasets. Experiment results show that the proposed FedPL outperforms existing methods by up to 13.50% in terms of accuracy. Code is available at https://github.com/tnghua/FedPL. Mei Cao, Zhenge Jia, Jianbo Lu 0001, Zhaoyan Shen, Dongxiao Yu, Mengying Zhao |
ECAI | 7 |
| 2024 | CSFL: Enhancing Splitfed Learning with Clustering on Non-IID DataabstractDistributed machine learning methods are gaining significant attention for their ability to enhance computational efficiency and safeguard privacy. Federated learning and split learning are two prominent approaches in this domain. Recently, splitfed learning, a hybrid of both methods, was introduced to address their individual limitations. However, splitfed learning overlooks the non-IID (non-Independent and Identically Distributed) problem commonly encountered in distributed environments, which can lead to substantial degradation in model performance. In this paper, we introduce Clustered Splitfed Learning (CSFL), a novel approach that integrates clustering with splitfed learning. We propose two training processes tailored to the degree of data heterogeneity: Non-Personalized Clustered Splitfed Learning (NPCSFL) and Personalized Clustered Splitfed Learning (PCSFL). Our experimental results demonstrate that CSFL significantly improves both model accuracy and convergence rates. Jianbo Lu 0001, Mei Cao, Mengying Zhao |
HPCC | 6 |
| 2024 | FedMQ: Multi-grained Quantization for Heterogeneous Federated Learning
Mei Cao, Jianbo Lu 0001, Zhaoyan Shen, Mengying Zhao |
WASA (2) | 6 |
| 2024 | An improved classification diagnosis approach for cervical images based on deep neural networks
Mengying Zhao, Chengyi Xia |
Pattern Anal. Appl. | 2 |
| 2024 | Branch Predictor Design for Energy Harvesting Powered Nonvolatile ProcessorsabstractNon-volatile processors are proposed for ambient energy harvesting systems to enable accumulative computing across power failures. They employ nonvolatile memory for processor status backup before power outage and resume the system after power recovers. A straightforward backup policy is to back up all volatile data in processors, but it induces high backup cost. In this paper, we focus on branch predictor, an important component in processor, and propose efficient backup schemes to reduce backup cost while maintaining its prediction ability. We first analyze the modules in both traditional and artificial intelligence (AI) assisted designs of branch predictor, and accordingly propose three backup mechanisms pertaining to saturation-driven, locality-driven and maturity-driven backup. On the basis of these mechanisms, adaptive backup branch predictors are designed. Evaluation shows that, with traditional Tournament architecture, the proposed design achieves 15.9% and 54.1% energy reduction when compared with no-backup and all-backup strategy. For AI assisted branch predictor, the proposed design achieves 27.5% and 82.2% energy saving. Mengying Zhao, Lihao Dong, Chun Jason Xue, Dongxiao Yu, Xiaojun Cai, Zhiping Jia |
IEEE Trans. Computers | 1 |
| 2024 | Implementing Neural Networks on Nonvolatile FPGAs With ReprogrammingabstractNV-FPGAs have attracted significant attention in research due to their high density, low leakage power, and reduced error rates. The nonvolatile memory (NVM) crossbar’s compute-in-memory (CiM) capability further enables NV-FPGAs to execute high-efficiency, high-throughput neural network (NN) inference tasks. However, with the rapid increase in network size and considering that the parameter size often exceeds the memory capacity of the field programmable gate array (FPGA), implementing the entire network on a single FPGA chip becomes impractical. In this article, we utilize FPGA’s inherent run time reprogramming feature to implement oversized NNs on NV-FPGAs. This approach splits NN models into multiple tasks for the cyclical execution. Specifically, we propose a performance-driven task adapter (PD-Adapter), which aims to achieve high-performance NN inference by employing the task deployment to optimize settings, such as processing element size and quantity, and the task switching to select the most suitable switching type for each task. We integrate the proposed PD-Adapter into an open-source toolchain and evaluate it. Experimental results demonstrate that the PD-Adapter can achieve a run time reduction of 85.37% and 76.12% compared to the baseline and execution-time-first policy, respectively. Hao Zhang 0145, Jian Zuo, Huichuan Zheng, Meihan Luo, Mengying Zhao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Introduction to Special Issue on In/Near Memory and Storage Computing for Embedded Systems
Liang Shi 0001, Jingtong Shi, Hussam Amrouch, Kuan-Hsun Chen, Mengying Zhao, Weichen Liu 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2024 | Introduction to the Special Issue on Embedded System Software/Tools
Ganapati Bhat, Biresh Kumar Joardar, Mengying Zhao |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2023 | Reinforcement Learning-Assisted Management for Convertible SSDsabstractConvertible SSDs, which allow flash cells to convert between different types of flash cells (e.g., SLC/MLC/TLC/QLC), are designed for achieving both high performance and high density. However, previous designs with two types of flash cells encounter a performance cliff degradation once the flash cells of single bit mode are consumed. In this work, we propose a novel level-based convertible SSD (e.g., including SLC-MLC-QLC), named RL-cSSD, that adopts an intermediate layer (e.g., MLC) as a performance cushion. A reinforcement learning-assisted device management scheme is designed to coordinate the data allocation, garbage collection and flash conversion processes considering both the SSD internal status and workload patterns. We evaluated RL-cSSD with various real-world workloads based on simulation. The experimental results show that the proposed RL-cSSD provides 72.98% higher performance on average compared with state-of-the-art schemes. Zhiping Jia, Mengying Zhao, Zhaoyan Shen, Bingzhe Li |
DAC | 4 |
| 2023 | Correlation-guided Placement for Nonvolatile FPGAsabstractNonvolatile FPGAs have advantages of high density and near-zero leakage power compared with traditional SRAM-based FPGAs. However, they have lifetime issue. To deal with this problem, a series of configuration files can be generated with various logical-to-physical mappings so that intensive writes can be distributed to different physical regions for wear leveling. Currently, the configuration files are independently generated, which is time-consuming. In this paper, we propose to investigate correlations between components and use them to guide the computer-aided design (CAD) flow to speed up the procedure of deriving configuration files. Specifically, we develop dynamic probabilities to drive the swapping of placement step in the CAD flow to push components to locate appropriate positions quickly. Evaluation shows that the proposed schemes can deliver 36.32% reduction in number of swappings when compared with existing strategies, while maintaining comparable performance and lifetime. Mengying Zhao, Fanjin Xu, Huichuan Zheng, Yuqing Xiong, Zhiping Jia, Xiaojun Cai |
DAC | 1 |
| 2023 | FedQL: Q-Learning Guided Aggregation for Federated Learning
Mei Cao, Mengying Zhao, Nanxiang Yu, Jianbo Lu 0001 |
ICA3PP (1) | 2 |
| 2022 | C2S: Class-aware client selection for effective aggregation in federated learningabstractFederated learning is proposed to train distributed data in a safe manner by avoiding to send data to server. The server maintains a global model and sends it to clients in each communication round, and then aggregates the updated local models to derive a new global model. Traditionally, the clients are randomly selected in each round and aggregation is based on weighted averaging. Researches show that the performance on IID data is satisfactory while significant accuracy drop can be observed for Non-IID data. In this paper, we explore the reasons and propose a novel aggregation approach for Non-IID data in federated learning. Specifically, we propose to group the clients according to classes of data they have, and select one set in each communication round. Local models from the same set are averaged as usual and the updated global model is sent to next group of clients for further training. In this way, the parameters are only averaged on similar clients and passed among different groups. Evaluation shows that the proposed scheme has advantages in terms of model accuracy and convergence speed with highly unbalanced data distribution and complex models. Mei Cao, Yujie Zhang 0007, Zezhong Ma, Mengying Zhao |
High Confid. Comput. | 4 |
| 2022 | Lifetime improvement through adaptive reconfiguration for nonvolatile FPGAs
Hao Zhang 0145, Huichuan Zheng, Shuangliang Li, Mengying Zhao, Xiaojun Cai |
J. Syst. Archit. | 5 |
| 2022 | Deep Reinforcement-Learning-Guided Backup for Energy Harvesting Powered SystemsabstractEnergy harvesting technology has been widely developed as a promising alternative of battery to power embedded systems. However, energy harvesting powered embedded systems may have potential frequent power interruptions due to unstable energy supply. Nonvolatile processors (NVPs) are proposed to survive power failures by saving volatile data to nonvolatile memory (NVM) upon power failures and resuming them after power comes back. Traditionally, backup is triggered immediately when an energy warning occurs. However, it is also possible to more aggressively utilize the residual energy for program execution to improve forward progress. In this work, we propose a deep reinforcement-learning-guided backup strategy to improve forward progress in energy harvesting powered intermittent embedded systems. The experimental results show an average of 8.3%, 51.6%, and 325.3% improved forward progress compared with$Q$-learning, the related work ALD, and traditional instant backup, respectively. Weifan Sun, Mengying Zhao, Weining Song, Xiaojun Cai, Tiantian Liu 0001, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | Adaptive Mode Transformation for Wear Leveling in Nonvolatile FPGAsabstractNowadays, field programmable gate arrays (FPGAs) have been widely adopted to serve as accelerators in artificial intelligence and big data related applications. Since the static random access memory (SRAM)-based FPGA is suffering from limited density and high leakage power, nonvolatile FPGAs have been proposed, where SRAM is replaced with emerging nonvolatile memories (NVMs). Multilevel cell (MLC), which can store multiple bits within one memory cell, further improves the density of nonvolatile FPGAs and shows great potential to enable large on-chip memory. However, it suffers from limited lifetime. In this article, we propose a wear leveling scheme to improve lifetime of MLC-based nonvolatile FPGAs. Instead of generating a series of configuration files for runtime reconfiguration, we propose to identify write-heavy MLC regions and dynamically transform them to durable single-level cell (SLC) mode. Specifically, we propose three modules: 1) pertaining to write behavior monitor; 2) approximate cost calculator; and 3) mode transformation manager to achieve adaptive mode transformations. We consider FPGA features to design these modules, which is different from implementations for MLC-SLC transformation in CPU architecture. Evaluation shows that the proposed scheme can improve lifetime for MLC nonvolatile FPGAs by$6.03\times $, at cost of 12.5% storage overhead. Huichuan Zheng, Fanjin Xu, Mengying Zhao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Fast-convergent federated learning with class-weighted aggregation
Zezhong Ma, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
J. Syst. Archit. | 2 |
| 2021 | A lightweight online backup manager for energy harvesting powered nonvolatile processor systemsabstractWith the explosive growth of battery-free and energy-harvesting devices, the energy harvesting powered system has gained more attentions and been widely used in different fields. However, unstable harvested energy is a challenge of energy-harvesting devices since the program execution would be interrupted frequently. Non-volatile processor (NVP) is proposed to back up volatile logics before energy depletion and recover the system status after energy resumption . This paper takes the challenge of forward progress improvement issue in NVP system and proposes a lightweight online backup strategy which tries to aggressively use energy in capacitor after receiving energy warnings. We also build a flexible and accurate simulation tool for NVP system evaluation. The experimental results show an average of 24.2% and 13.7% improved forward progress compared with instant backup method and the most related work, respectively. Weining Song, Xiaojun Cai, Mengying Zhao, Zhaoyan Shen, Zhiping Jia |
J. Syst. Archit. | 3 |
| 2021 | Pearl: Performance-Aware Wear Leveling for Nonvolatile FPGAsabstractSince static random access memory (SRAM)-based field-programmable gate array (FPGA) has limited density and comparatively high leakage power, researchers have proposed FPGA architectures based on emerging nonvolatile memories (NVMs) to satisfy the requirements of data-intensive and low-power applications. Among all components, block random access memory (BRAM) has the severest endurance problem in FPGA. Unluckily, traditional wear leveling (TWL) strategies cannot be directly applied to nonvolatile FPGA because it may induce large performance overhead. In this article, we propose performance-aware wear leveling schemes for nonvolatile FPGA to improve its lifetime. Two strategies pertaining to coarse-grained wear leveling (C-Pearl) and fine-grained wear leveling (F-Pearl) are developed to balance inter-BRAM and intra-BRAM writes. Procedures, including static analysis, wear leveling-guided placement, and reconfiguration are discussed. A supportive circuit design is proposed, too. The evaluation shows that C-Pearl and F-Pearl can achieve 34% and 46% higher lifetime improvement and simultaneously 8% and 11% lower performance overhead than TWL. Mengying Zhao, Zhaoyan Shen, Xiaojun Cai, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | PattPIM: A Practical ReRAM-Based DNN Accelerator by Reusing Weight Pattern RepetitionsabstractWeight sparsity has been explored to achieve energy efficiency for Resistive Random-access Memory (ReRAM) based DNN accelerators. However, most existing ReRAM-based DNN accelerators are based on an overidealized crossbar architecture and mainly focus on compressing zero weights. In this paper, we propose a novel ReRAM-based accelerator — PattPIM, to achieve space compression and computation reuse by studying DNN weight patterns based on practical ReRAM crossbars. We first thoroughly analyze the weight distribution characteristics of several typical DNN models and observe many non-zero weight pattern repetitions (WPRs). Thus, in PattPIM, we propose a WPR-aware DNN engine and a WPR-to-OU mapping scheme to save both space and computation resources. Furthermore, we adopt an approximate weight pattern transform algorithm to improve the DNN WPRs ratio to enhance the reuse efficiency with negligible inference accuracy loss. Our evaluation with 6 DNN models shows that the proposed PattPIM delivers significant performance improvement, ReRAM resources efficiency and energy saving. Yuhao Zhang 0006, Zhiping Jia, Yungang Pan, Hongchao Du, Zhaoyan Shen, Mengying Zhao, Zili Shao |
DAC | 6 |
| 2020 | Q-learning Based Backup for Energy Harvesting Powered Embedded SystemsabstractNon-volatile processors (NVPs) are used in energy harvesting powered embedded systems to preserve data across interruptions. In NVP systems, volatile data are backed up to non-volatile memory upon power failures and resumed after power comes back. Traditionally, backup is triggered immediately when energy warning occurs. However, it is also possible to more aggressively utilize the residual energy for program execution to improve forward progress. In this work, we propose a Q-learning based backup strategy to achieve maximal forward progress in energy harvesting powered intermittent embedded systems. The experimental results show an average of 307.4% and 43.4% improved forward progress compared with traditional instant backup and the most related work, respectively. Yujie Zhang 0007, Weining Song, Mengying Zhao, Zhaoyan Shen, Zhiping Jia |
DATE | 4 |
| 2020 | Maximizing CNN Throughput on FPGA ClustersabstractField Programmable Gate Array (FPGA) platform has been a popular choice for deploying Convolutional Neural Networks (CNNs) as a result of its high parallelism and low energy consumption. Due to the limitation of on-chip resources on a single board, FPGA clusters become promising solutions to improve the throughput of CNNs. In this paper, we firstly put forward strategies to optimize the resource allocation intra and inter FPGA boards. Then we model the multi-board cluster problem and design algorithms based on knapsack problem and dynamic programming to calculate the optimal topology of the FPGA clusters. We also give a quantitative analysis of the inter-board data transmission bandwidth requirement. To make our design accommodate for more situations, we provide solutions for deploying fully connected layers and special convolution layers with large memory requirement. Experimental results show that typical well-known CNNs with the proposed topology of FPGA clusters could obtain a higher throughput per board than single-board solutions and other multi-board solutions. Ruihao Li 0002, Mengying Zhao, Zhaoyan Shen, Xiaojun Cai, Zhiping Jia |
FPGA | 3 |
| 2020 | Design Insights of Non-volatile Processors and Accelerators in Energy Harvesting SystemsabstractThere is growing interest in deploying energy harvesting processors and accelerators in Internet of Things (IoT). Energy harvesting harnesses the energy scavenged from the environment to power a system. Although it has many advantages over battery-operated systems such as lightweight, compact size, and no necessity of recharging and maintenance, it may suffer frequently power-down and a fluctuating power supply even with power on. Non-volatile processor (NVP) is a promising architecture for effective computing in energy harvesting scenarios. Recently, non-volatile accelerators (NVA) have been proposed to perform computations of deep learning algorithms. In this paper, we overview the recent studies of NVP and NVA across the layers of hardware, architecture, software and their co-design. Especially, we present the design insights of how the state-of-the-art works adapt their specific designs to the intermittent and fluctuating power conditions with the energy harvesting technology. Finally, we discuss recent trends using NVP and NVA in energy harvesting scenarios. Keni Qiu, Mengying Zhao, Zhenge Jia, Jingtong Hu, Chun Jason Xue, Kaisheng Ma, Xueqing Li 0002, Yongpan Liu, Narayanan Vijaykrishnan |
ACM Great Lakes Symposium on VLSI | 2 |
| 2020 | ResiRCA: A Resilient Energy Harvesting ReRAM Crossbar-Based Accelerator for Intelligent Embedded ProcessorsabstractMany recent works have shown substantial efficiency boosts from performing inference tasks on Internet of Things (IoT) nodes rather than merely transmitting raw sensor data. However, such tasks, e.g., convolutional neural networks (CNNs), are very compute intensive. They are therefore challenging to complete at sensing-matched latencies in ultra-low-power and energy-harvesting IoT nodes. ReRAM crossbar-based accelerators (RCAs) are an ideal candidate to perform the dominant multiplication-and-accumulation (MAC) operations in CNNs efficiently, but conventional, performance-oriented RCAs, while energy-efficient, are power hungry and ill-optimized for the intermittent and unstable power supply of energy-harvesting IoT nodes. This paper presents the ResiRCA architecture that integrates a new, lightweight, and configurable RCA suitable for energy harvesting environments as an opportunistically executing augmentation to a baseline sense-and-transmit battery-powered IoT node. To maximize ResiRCA throughput under different power levels, we develop the ResiSchedule approach for dynamic RCA reconfiguration. The proposed approach uses loop tiling-based computation decomposition, model duplication within the RCA, and inter-layer pipelining to reduce RCA activation thresholds and more closely track execution costs with dynamic power income. Experimental results show that ResiRCA together with ResiSchedule achieve average speedups and energy efficiency improvements of 8× and 14× respectively compared to a baseline RCA with intermittency-unaware scheduling. Keni Qiu, Nicholas Jao, Mengying Zhao, Cyan Subhra Mishra, Gulsum Gudukbay Akbulut, Sethu Jose, Jack Sampson, Mahmut T. Kandemir, Narayanan Vijaykrishnan |
HPCA | 3 |
| 2020 | A Highly Parallelized PIM-Based Accelerator for Transaction-Based Blockchain in IoT EnvironmentabstractBlockchain has gained a lot of attention from both academia and industry. However, traditional standard blockchains, such as Bitcoin and Ethereum, suffer from low throughput, high computation overhead, and large transaction fee, which is not suitable for Internet of Things (IoT) transactions. Recently, transaction-based approaches, such as Tangle structure which is based on a directed acyclic graph (DAG), have emerged to solve blockchain scalability issues for IoT environment. In transaction-based blockchain, for a transaction, namely, a node, to be attached to the Tangle, it needs to verify two other transactions. However, with the Tangle expanding, this attaching process consumes huge computational resources and energy, which severely limits the performance of the transaction-based blockchain. In this article, we present Re-Tangle, a highly parallelized processing-in-memory (PIM)-based accelerator for transaction-based blockchain. Re-Tangle is composed of a random walking module, a transaction validation module, and a PoW module, to improve the Tangle system performance. These modules transfer Tangle functions, such as fast exponentiation and modular, into ReRAM-based logic analog computation units. In the random walking module, Re-Tangle maintains an exponentiation table to reduce its design complexity and improve its computation efficiency. In the transaction validation module, Re-Tangle further proposes a highly parallel modular unit to accelerate the validation of different tags in a transaction. In the PoW module, we decompose the Curl hash function into basic logic OR, AND, SHIFT, and XOR operations, and map these logic operations to ReRAM crossbars in parallel to accelerate the working process. The experimental results show that Re-Tangle distinguishes itself from other architectures with significant performance improvement and energy saving. The throughput of Re-Tangle is about 22.4× and 2.38× higher compared with CPU and GPU, respectively, and the energy consumption of Re-Tangle is 83.5× and 5.77× less for equal workload. Qian Wang 0042, Zhiping Jia, Tianyu Wang 0009, Zhaoyan Shen, Mengying Zhao, Renhai Chen, Zili Shao |
IEEE Internet Things J. | 5 |
| 2020 | Applying Multiple Level Cell to Non-volatile FPGAsabstractStatic random access memory– (SRAM) based field programmable gate arrays (FPGAs) are currently facing challenges of limited capacity and high leakage power. To solve this problem, non-volatile memory (NVM) is proposed as the alternative to build non-volatile FPGAs (NVFPGAs). Even though the feasibility of NVFPGA has been confirmed, the utilization of multiple level cells (MLCs) has not been fully exploited yet. In this article, we study architecture of MLC-based NVFPGAs, and propose five cluster structures. To give detailed comparisons and extensive discussions, we conduct experiments for area, performance and leakage power evaluation. Based on explorations of the characteristics of MLC-based NVFPGAs, we further present MLC-aware timing-driven packing method to improve delay. In critical paths, our proposed method reduces the overhead of the additional delay in slow MLC cells. Experiments show that, compared to SRAM-based FPGAs, the proposed architecture with the proposed CAD flow can reduce the area, critical path delay and leakage power by 31%, 10%, and 95%, respectively. Mengying Zhao, Lei Ju 0001, Zhiping Jia, Jingtong Hu, Chun Jason Xue |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2019 | Performance-aware Wear Leveling for Block RAM in Nonvolatile FPGAsabstractField programmable gate arrays (FPGAs) have been widely adopted in both high-performance servers and embedded systems. Since static random access memory (SRAM) has limited density and comparatively high leakage power, researchers have proposed FPGA architectures based on emerging non-volatile memories (NVMs) to satisfy the requirements of data-intensive and low-power applications. Block RAM is on-chip memory of FPGAs, when it is implemented with NVM, it will face the challenge of limited endurance. Traditional wear leveling strategy cannot be directly applied to block RAM because it may induce large performance overhead. In this paper, we propose a performance-aware wear leveling scheme for block RAM in FPGAs to improve its lifetime. The placement strategy is improved by injecting wear leveling guidance. The evaluation shows that 29.75% lifetime enhancement is achieved with 16.32% performance improvement at the same time, compared with traditional wear leveling. Shuo Huai, Weining Song, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
DAC | 3 |
| 2019 | Re-Tangle: A ReRAM-based Processing-in-Memory Architecture for Transaction-based BlockchainabstractBlockchain has gained a lot of attentions from both academic and industry. Transaction-based approaches such like Tangle structure, which is based on a DAG (Directed Acyclic Graph), are emerging to solve the blockchain scalability issues for IoT environment. In transaction-based blockchain, for a transaction, namely a node, to be attached to the Tangle, it needs to verify two other transactions. However, with the Tangle expanding, this attaching process consumes huge computational resources and energy, which severely limits the performance of the transaction-based blockchain. In this paper, we present Re-Tangle, a novel transaction-based blockchain acceleration architecture that explores the opportunity of performing massive parallel operations with low hardware and energy cost. Re-Tangle consists of a random walking module and a transaction validation module, which transfer Tangle functions into ReRAM-based logic analog computation units. In the random walking module, Re-Tangle maintains a exponentiation translator to reduce its design complexity and improve its computation efficiency for exponentiation. In the transaction validation module, Re-Tangle further proposes a highly parallel modular unit to accelerate the validation of different tags in a transaction. The experience results show that Re-Tangle distinguishes itself from other architectures, with significant performance improvement and energy saving. The throughput of Re-Tangle is about 19.4× and 2.13× higher compared with CPU and GPU, respectively, and the energy consumption of Re-Tangle is 63.35 × and 4.92 × less. Qian Wang 0042, Tianyu Wang 0009, Zhaoyan Shen, Zhiping Jia, Mengying Zhao, Zili Shao |
ICCAD | 5 |
| 2019 | EMC: Energy-Aware Morphable Cache Design for Non-Volatile ProcessorsabstractWearable, implantable and Internet of Things devices are attracting increasing attention from both research and industry fields. Energy harvesting is a promising alternative of battery to power these embedded systems. However, the intrinsic instability of energy harvesting systems leads to potential frequent power interruptions. In traditional volatile processor, all the status will be lost at power failures and the program needs to re-start after power resumes. In order to survive the power failures and enable accumulative execution, non-volatile processor (NVP) is proposed to back up volatile information before power depletion and recover the system status after power resumes. Non-volatile memory (NVM) is typically attached for cache and main memory backup. There are researches working on optimization of the backup. However, little of them involve multiple level cell (MLC) NVM. In this work, we first discuss the benefit of applying MLC NVM for cache backup and the architecture of morphable hybrid cache, and then propose a three-stage energy-aware cache management strategy to improve the system performance and energy utilization while guaranteeing successful backups. Backup-aware cache replacement policies are also developed for backup optimization. Evaluation shows that the proposed EMC scheme can achieve 10.6 percent performance improvement and simultaneous 25.2 percent energy reduction when compared with the single level cell (SLC) based hybrid cache. Weining Song, Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Zhiping Jia |
IEEE Trans. Computers | 3 |
| 2019 | Checkpointing-Aware Loop Tiling for Energy Harvesting Powered Nonvolatile ProcessorsabstractAs power failures often occur in energy harvesting powered nonvolatile processors (NVPs), checkpointing is needed during program execution. It is observed that checkpointing is implemented with high overhead in applications with loops, because a large amount of data needs backup during loop execution. As such, we are motivated to reduce the amount of checkpointing data by analyzing data locality and shortening data lifetime in loops. This paper proposes a checkpointing-aware loop tiling technique which targets to reduce the checkpointing and recovering overheads for loops. Specifically, we first derive the optimal tile size for nested loops considering checkpointing distance and data dependencies. Then, the implementations of checkpointing and recovering for tiled loops are presented. Finally, the experiments are conducted to evaluate the effectiveness of the proposed method. The experimental results show that compared to the no-tiling method, the checkpointing-aware loop tiling method reduces the checkpointing and recovering data by 36.2% on average and reduces the total execution time and dynamic energy for checkpointing and recovering by 27.2% and 22.9% on average, respectively. Keni Qiu, Mengying Zhao, Jingtong Hu, Yongpan Liu, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Set variation-aware shared LLC management for CPU-GPU heterogeneous architectureabstractHeterogeneous CPU-GPU multiprocessor systems-on-chip (HMPSoC) becomes a popular architecture choice for high performance embedded systems, where shared last-level cache (LLC) management becomes a critical design consideration. We observe that within a sampling period, CPU and GPU may have distinct access behaviors over various LLC sets. In this work, we propose a light-weighted and fined-grained cache management policy to cope with the CPU-GPU access behavior variation among cache sets. In particular, CPU and GPU requests are prioritized disparately in each LLC set during cache block insertion and promotion, based on the per-core utility behaviors and a per-set CPU-GPU miss counter. Experimental results show that our LLC management scheme outperforms the two state-of-the-art schemes TAP-RRIP and LSP by 12.6% and 10.01%, respectively. Zhaoying Li 0004, Lei Ju 0001, Hongjun Dai, Mengying Zhao, Zhiping Jia |
DATE | 5 |
| 2018 | H2-RAID: A Novel Hybrid RAID Architecture Towards High Reliability
Tianyu Wang 0009, Zhiyong Zhang 0006, Mengying Zhao, Zhiping Jia, Jianping Yang, Yang Wu 0003 |
ICA3PP (4) | 3 |
| 2018 | NVM-Based FPGA Block RAM With Adaptive SLC-MLC ConversionabstractThe capacity of SRAM-based FPGA block RAM (BRAM) is restrained by the low density and high leakage power of the current CMOS technology. In this paper, we propose a nonvolatile memory (NVM)-based BRAM architecture which enables flexible conversions between single-level cell (SLC) and multilevel cell (MLC) states. We show that despite the high per-access latency and power consumption, MLC-based BRAM blocks reduce the routing cost between logic units and on-chip data storages, which potentially leads to a smaller critical path delay and power consumption. Therefore, we propose an NVM BRAM architecture and an EDA framework which adaptively packs data into SLC- or MLC-state BRAMs during FPGA design flow in order to achieve better system performance. This paper illustrates that a simple memory device replacement from SRAM to NVM leads to nonoptimal system performance. On the other hand, compared with operating all NVM BRAM blocks in the SLC state with better per-access latency and power consumption, the proposed hybrid SLC-MLC architecture and design flow improves the critical path delay by 18.51%, with a system power reduction of 25.83% at the same time. Moreover, compared with the traditional “fast” SRAM-based BRAM blocks under the same BRAM area constraint, our hybrid NVM BRAM architecture improves the critical path delay by 8.55% on average, with an average system power reduction of 54.34% at the same time. Lei Ju 0001, Xiaojin Sui, Shiqing Li, Mengying Zhao, Chun Jason Xue, Jingtong Hu, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2018 | Shared Last-Level Cache Management and Memory Scheduling for GPGPUs with Hybrid Main MemoryabstractMemory intensive workloads become increasingly popular on general purpose graphics processing units (GPGPUs), and impose great challenges on the GPGPU memory subsystem design. On the other hand, with the recent development of non-volatile memory (NVM) technologies, hybrid memory combining both DRAM and NVM achieves high performance, low power, and high density simultaneously, which provides a promising main memory design for GPGPUs. In this article, we explore the shared last-level cache management for GPGPUs with consideration of the underlying hybrid main memory. To improve the overall memory subsystem performance, we exploit the characteristics of both the asymmetric read/write latency of the hybrid main memory architecture, as well as the memory coalescing feature of GPGPUs. In particular, to reduce the average cost of L2 cache misses, we prioritize cache blocks from DRAM or NVM based on observations that operations to NVM part of main memory have a large impact on the system performance. Furthermore, the cache management scheme also integrates the GPU memory coalescing and cache bypassing techniques to improve the overall system performance. To minimize the impact of memory divergence behaviors among simultaneously executed groups of threads, we propose a hybrid main memory and warp aware memory scheduling mechanism for GPGPUs. Experimental results show that in the context of a hybrid main memory system, our proposed L2 cache management policy and memory scheduling mechanism improve performance by 15.69% on average for memory intensive benchmarks, whereas the maximum gain can be up to 29% and achieve an average memory subsystem energy reduction of 21.27%. Chuanqi Zang, Lei Ju 0001, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2018 | Avoiding Data Inconsistency in Energy Harvesting Powered Embedded SystemsabstractEnergy harvesting is becoming a favorable alternative to power future generation embedded systems, as it is more environmentally and user friendly. However, energy harvesting powered embedded systems suffer from frequent execution interruption due to unstable energy supply. To tackle this problem, nonvolatile memory has been deployed to save the whole volatile state for computation. When power resumes, the processor can restore the state back to volatile memories and continue execution. However, without careful consideration, the process of checkpointing and resuming could cause inconsistency between volatile and nonvolatile memories, which leads to irreversible errors. In this article, we propose a consistency-aware adaptive checkpointing scheme that ensures correctness for all checkpoints. The proposed technique efficiently identifies all possible inconsistency positions in programs and inserts auxiliary code to ensure correctness by offline analysis. In addition, adaptive checkpointing assisted register file profiling and online tracking techniques further reduce the overhead of each checkpoint. Evaluation results show that the proposed checkpointing strategy can successfully eliminate inconsistency errors and greatly reduce the checkpointing overhead. Mimi Xie, Mengying Zhao, Yongpan Liu, Chun Jason Xue, Jingtong Hu |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2017 | Cooperative DVFS for energy-efficient HEVC decoding on embedded CPU-GPU architectureabstractThe next generation video coding standard High Efficiency Video Coding (HEVC) provides better compression rate for high resolution videos, at the cost of substantially higher computational complexity. While some latest off-the-shelf consumer electronics support HEVC via ASIC solutions, software implementation of real-time HEVC remains an open challenge for resource-constraint embedded systems. In this work, we present an HEVC decoder design on a low-power embedded heterogeneous multiprocessor System-on-Chip (HMPSoC) with CPU and GPU. Our analysis shows that the massive parallel architecture of GPU leads to a relatively smooth fluctuation on the processing time between video frames. Moreover, the dynamic workload of each frame has a monotonic correlation with a particular coding parameter that can be obtained at decoding time. Based on these observations, we propose an application-specific userspace CPU-GPU DVFS scheme which effectively saves the energy consumption for HEVC decoding. Furthermore, given our accurate workload prediction, only a small frame buffer is required to ensure real-time video decoding. Fan Gong, Lei Ju 0001, Deshan Zhang, Mengying Zhao, Zhiping Jia |
DAC | 4 |
| 2017 | Maximizing Forward Progress with Cache-aware Backup for Self-powered Non-volatile ProcessorsabstractEnergy harvesting is replacing battery to power embedded systems such as Internet of Things and wearable devices. Unstable energy supply brings challenges to energy harvesting powered system, resulting in frequent interruptions. Non-volatile processor is proposed to back up volatile logics before energy depletion and recover the system status after energy resumes. The backup efficiency of memory content significantly affects program performance. There are existing researches focusing on backup optimizations, but they did not fully consider cache behaviors. In this paper, we introduce cache persistence analysis into memory backup for self-powered non-volatile processors. The evaluation shows that the proposed cache-aware backup delivers on average 45.6% improvement in forward progress, achieving 40.2% and 12.7% higher system performance compared with instant and cache-unaware backup. Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Zhiping Jia |
DAC | 2 |
| 2017 | Shared last-level cache management for GPGPUs with hybrid main memoryabstractMemory intensive workloads become increasingly popular on general purpose graphics processing units (GPGPUs), and impose great challenges on the GPGPU memory subsystem design. On the other hand, with the recent development of non-volatile memory (NVM) technologies, hybrid memory combining both DRAM and NVM achieves high performance, low power and high density simultaneously, which provides a promising main memory design for GPGPUs. In this work, we explore the shared last-level cache management for GPGPUs with consideration of the underlying hybrid main memory. In order to improve the overall memory subsystem performance, we exploit the characteristics of both the asymmetric read/write latency of the hybrid main memory architecture, as well as the memory coalescing feature of GPGPU. In particular, to reduce the average cost of L2 cache misses, we prioritize cache blocks from DRAM or NVM based on observation that operations to NVM part of main memory have large impact on the system performance. Furthermore, the cache management scheme also integrates the GPU memory coalescing and cache bypassing techniques to improve the overall cache hit ratio. Experimental results show that in the context of a hybrid main memory system, our proposed L2 cache management policy improves performance against the traditional LRU policy and a state-of-the-art GPU cache strategy EABP [20] by up to 27.76% and 14%, respectively. Xiaojun Cai, Lei Ju 0001, Chuanqi Zang, Mengying Zhao, Zhiping Jia |
DATE | 5 |
| 2017 | Design Exploration for Multiple Level Cell Based Non-Volatile FPGAsabstractStatic random access memory (SRAM) based field programmable gate arrays (FPGAs) are currently facing challenges of limited capacity and high leakage power. To solve this problem, non-volatile memory (NVM) is proposed as the alternative to build non-volatile FPGAs (NVFPGAs). Even though the feasibility of NVFPGA has been confirmed, the utilization of multiple level cells (MLC) has not been fully exploited yet. In this paper, we study architecture of MLC based NVFPGAs, and propose five cluster structures, as well as the corresponding working mode supported by MLC based clusters. To give detailed comparisons and extensive discussions, we conduct experiments for area, performance and leakage power evaluation. Experiments show that, compared to SRAM based FPGAs, the proposed architecture can reduce the area, latency and leakage power by 32.66% and 7.45%, and 96.13%, respectively. Mengying Zhao, Lei Ju 0001, Zhiping Jia, Chun Jason Xue, Jingtong Hu |
ICCD | 2 |
| 2017 | Unified nvTCAM and sTCAM architecture for improving packet matching performanceabstractSoftware-Defined Networking (SDN) allows controlling applications to install fine-grained forwarding policies in the underlying switches. Ternary Content Addressable Memory (TCAM) enables fast lookups in hardware switches with flexible wildcard rule patterns. However, the performance of packet processing is severely constrained by the capacity of TCAM, which aggravates the processing burden and latency issues. In this paper, we propose a hybrid TCAM architecture which consists of NVM-based TCAM (nvTCAM) and SRAM-based TCAM (sTCAM), utilizing nvTCAM to cache the most popular rules to improve cache-hit-ratio while relying on a very small-size sTCAM to handle cache-miss traffic to effectively decrease update latency. Considering the special rule dependency, we present an efficient Rule Migration Replacement (RMR) policy to make full utilization of both nvTCAM and sTCAM to obtain better performance. Experimental results show that the proposed architecture outperforms current TCAM architectures. Xianzhong Ding, Zhiyong Zhang 0006, Zhiping Jia, Lei Ju 0001, Mengying Zhao, Huawei Huang |
LCTES | 5 |
| 2017 | An empirical study of F2FS on mobile devicesabstractFlash Friendly File System (F2FS) is getting popular among mobile devices. However, lack of empirical and comprehensive analysis for characteristics of F2FS prohibits better application of F2FS. In this paper, we present a set of comprehensive experimental studies on mobile devices and show several counterintuitive observations on F2FS, including imprecise hot/cold data separation, unexpected trigger condition of background GC, impact of fragmentation on read performance and impact of readahead by fragments and available space. Based on these observations, we further provide several pilot solutions to improve the performance of these mobile devices. The objective is to inspire researchers and users to pay attention to F2FS characteristics, and further optimize its performance. Yu Liang 0004, Chenchen Fu, Yajuan Du, Aosong Deng, Mengying Zhao, Liang Shi 0001, Chun Jason Xue |
RTCSA | 5 |
| 2017 | Energy-aware morphable cache management for self-powered non-volatile processorsabstractWearable, implantable and Internet of Things devices are attracting increasing attention from both research and industry. Energy harvesting is a promising alternative of battery to power these embedded systems. However, the intrinsic instability of energy harvesting systems leads to potential frequent power interruptions. In order to survive the power failures, non-volatile processor (NVP) is proposed to back up volatile information before power depletion and recover the system status after power resumes. Non-volatile memory (NVM) is typically attached for cache and main memory backup. There are researches working on optimization of the backup, however, little of them involve multiple level cell (MLC) NVM. In this work, we first discuss the benefit of applying MLC NVM for cache backup, and then propose a three-stage energy-aware cache management strategy to improve the system performance and energy utilization while guaranteeing successful backups. Evaluation shows that the proposed scheme can achieve 16.7% energy reduction with comparative performance with the single level cell (SLC) based hybrid cache. Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Xin Li 0001, Zhiping Jia |
RTCSA | 2 |
| 2017 | Data re-allocation enabled cache locking for embedded systems
Chun Jason Xue, Keni Qiu, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Mengying Zhao |
J. Syst. Archit. | 6 |
| 2017 | Data Backup Optimization for Nonvolatile SRAM in Energy Harvesting Sensor NodesabstractNonvolatile static random access memory (nvSRAM) has been widely investigated as a promising on-chip memory architecture in energy harvesting sensor nodes, due to zero standby power, resilience to power failures, and fast read/write operations. However, conventional approaches back up all data from static random access memory into nonvolatile memory when power failures happen. It leads to significant energy overhead and peak inrush current, which has a negative impact on the system performance and circuit reliability. This paper proposes a holistic data backup optimization to mitigate these problems in nvSRAM, consisting of a partial backup algorithm and a run-time adaptive write policy. A statistic dead-block predictor is employed to achieve dead block identification with trivial hardware overhead. An adaptive policy is used to switch between write-back and write-through strategy to reduce the rollback induced by backup failures. Experimental results show that the proposed scheme improves the performance by 4.6% on average while the backup power consumption and the inrush current are reduced by 38.1% and 54% on average compared to the full backup scheme. What is more, the backup capacitor size for energy buffer can be reduced by 40% on average under the same performance constraint. Yongpan Liu, Jinshan Yue, Hehe Li, Qinghang Zhao, Mengying Zhao, Chun Jason Xue, Guangyu Sun 0003, Meng-Fan Chang, Huazhong Yang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Asymmetric Error Rates of Cell States Exploration for Performance Improvement on Flash Memory Based Storage SystemsabstractRecent studies show that a multilevel cell flash cell in different states suffers from diverse error patterns in varying degrees. That is, the error rates of each page are highly dependent on the data content. Consequently, pages with different data will exhibit quite different error rates. However, existing technologies equipped with one uniform error correction code (ECC) scheme for all pages in a flash memory do not take the different error rates of pages into consideration. In this paper, we propose to exploit the asymmetric error rates of flash memory exhibited by the flash pages with different data for performance improvement. Before a page is programmed, its specific error rates, called content-dependent bit error rates (CDBERs), are estimated according to the content of the page. The margin between the CDBER of a page and the maximal error rates correctable by the uniform ECC code is exploited for performance improvement. On one hand, a faster and suitable write operation is selected to speed up the progress of programming while the increased speed induced CDBER does not exceed the maximal correctable error rates. On the other hand, a light-weight ECC scheme can be chosen for a faster read operation since the page decoding process of a light-weight ECC scheme incurs less time overhead. Finally, a state mapping scheme, which further reduces the CDBER through mapping high error rate states to the low error rate states of a page, is proposed. Simulation results show that the proposed approaches lead to significant write and read performance improvement. Edwin H.-M. Sha, Congming Gao, Liang Shi 0001, Kaijie Wu 0001, Mengying Zhao, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2017 | Stack-Size Sensitive On-Chip Memory Backup for Self-Powered Nonvolatile ProcessorsabstractWearable devices gain increasing popularity since they can collect important information for healthcare and well-being purposes. Compared with battery, energy harvesting is a better power source for these wearable devices due to many advantages. However, harvested energy is naturally unstable and program execution will be interrupted frequently. Nonvolatile processors demonstrate promising advantages to back up volatile state before the system energy is depleted. However, it also introduces non-negligible energy and area overhead. In this paper, we aim to reduce the amount of data that need to be backed up during a power failure. Based on the observation that stack size varies along program execution, we propose to analyze the application program and identify efficient backup positions, by which the stack content to back up can be significantly reduced. The evaluation results show an average of 45.7% reduction on nonvolatile stack size for stack backup, with 0.58% storage overhead. In the mean time, with the proposed schemes, the energy utilization and program forward progress can be greatly improved compared with instant backup. Mengying Zhao, Chenchen Fu, Qing'an Li, Mimi Xie, Yongpan Liu, Jingtong Hu, Zhiping Jia, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | State Asymmetry Driven State Remapping in Phase Change MemoryabstractPhase change memory (PCM) is one of the most promising candidates to replace DRAM as main memory in deep submicron regime. Regardless of single-level or multiple-level cells, the programming costs to each state exhibit significant asymmetries in latency, energy and endurance. In this paper, we exploit the potential of reducing programming costs in terms of latency, energy, and endurance for PCM through state remapping. First, quantitative programming models are constructed for cost assessments. Then, both dynamic and static remapping schemes are analyzed and compared. The observation that the efficacy of dynamic state remappings is instable motivates us to propose a static remapping technique, which outperforms previous work in cost reduction within much lower implementation overhead. The optimality of the proposed static state remapping is also proved. The evaluation results confirm the efficacy of the proposed state remapping technique in delivering a stable and promising cost reduction in latency, energy, and wear. Mengying Zhao, Jingtong Hu, Chengmo Yang, Tiantian Liu 0001, Zhiping Jia, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | Write-back aware shared last-level cache management for hybrid main memoryabstractHybrid main memory with both DRAM and emerging non-volatile memory (NVM) becomes a promising solution for high performance and energy-efficient embedded systems. Cache plays an important role and highly affects the number of write backs to NVM and DRAM blocks. However, existing cache policies fail to fully address the significant asymmetry between NVM operations (especially writes) and DRAM operations, leading to non-optimal system designs. We propose a write-back aware last-level cache management scheme for the hybrid main memory, which improves the cache hit ratio of NVM memory blocks and minimizes write-backs to NVM. Experimental results show that our proposed framework leads to better performance and energy saving compared with the state-of-the-art cache management scheme for hybrid main memory architecture. Deshan Zhang, Lei Ju 0001, Mengying Zhao, Xiang Gao 0012, Zhiping Jia |
DAC | 3 |
| 2016 | Multipath Load Balancing in SDN/OSPF Hybrid Network
Xiangshan Sun, Zhiping Jia, Mengying Zhao, Zhiyong Zhang 0006 |
NPC | 3 |
| 2016 | Redesigning software and systems for non-volatile processors on self-powered devicesabstractWearable devices gain increasing popularity since they can collect important information for healthcare and well-being purposes. Compared with battery, energy harvesting is a better power source for these wearable devices due to many advantages. However, harvested energy is naturally unstable and program execution will be interrupted frequently. Nonvolatile processor (NVP) demonstrates promising advantages to back up volatile state before the system energy is depleted. Due to the backup and resumption procedures resulted from frequent power failures, non-volatile processor exhibits different characteristics from traditional processors, necessitating a set of adaptive design and optimization strategies. Recently, there have been both hardware and software researches aiming to develop correct and efficient non-volatile processors. In this paper, we summarize the software-level techniques for NVP, covering error-correctness schemes, backup timing determination, backup content optimization, adaptive software modifications and NVP simulators and tools, to provide an overview of state-of-the-art NVP research from the software and system level. Mengying Zhao, Keni Qiu, Yuan Xie 0001, Jingtong Hu, Chun Jason Xue |
VLSI-SoC | 1 |
| 2016 | Retention Trimming for Lifetime Improvement of Flash Memory Storage SystemsabstractNAND flash memory has been widely deployed in embedded systems, personal computers, and data centers. While recent technology scaling and density improvement have reduced its price, they have also significantly shortened its endurance. In this paper, with the understanding of the relationship between data retention time and flash wearing, a retention trimming approach, which trims data retention time based on the data lifetime, is proposed to reduce the wearing of flash memory, and hence improve the endurance of flash memory. Extensive experimental results show that the proposed technique achieves significant endurance improvements. Liang Shi 0001, Kaijie Wu 0001, Mengying Zhao, Chun Jason Xue, Duo Liu 0002, Edwin H.-M. Sha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Exploiting Process Variation for Write Performance Improvement on NAND Flash Memory Storage SystemsabstractThe write performance of flash memory has been degraded significantly due to the recent density-oriented advancements of flash technology. Techniques have been proposed to improve the write performance by exploiting the varying strength of a flash block in its different worn-out stages. A block is written with a faster speed when it is new and strong, and gradually will be written with slower speeds as it is aging and becomes weak. Motivated by these works, this brief proposes a new technique by exploiting the significant process variation among flash blocks introduced by the advanced technology scaling. First, a write speed detection approach is proposed to identify the strength of each block. Then, a heuristic approach is proposed to exploit the speed variation among blocks for write performance improvement. A series of trace-driven simulations shows that the proposed approach generates substantial write performance improvement over state-of-the-art approaches by 30% on average. Liang Shi 0001, Yejia Di, Mengying Zhao, Chun Jason Xue, Kaijie Wu 0001, Edwin H.-M. Sha |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Minimizing MLC PCM write energy for free through profiling-based state remappingabstractPhase change memory is becoming one of the most promising candidates to replace DRAM as main memory in deep sub-micron regime. Multi-level cell (MLC) PCM outperforms single level cell (SLC) PCM in terms of storage capacity but requires an iterative programming-and-verifying scheme to program cells to different resistance levels. The energy consumed in programming different MLC states varies significantly, thus motivating a state remapping technique to minimize the overall write energy. In this paper, we first compare dynamic and static state remapping strategies in terms of their efficacy in reducing energy, and then propose an effective and low-cost static state remapping algorithm. The experimental studies show 10.6% average (up to 16.9%) reduction in MLC PCM write energy, achieved within negligible hardware and performance overhead. Compared with the most related work, the proposed scheme saves more write energy on average, with near-zero performance, area and energy overhead. Mengying Zhao, Chengmo Yang, Chun Jason Xue |
ASP-DAC | 1 |
| 2015 | Compiler directed automatic stack trimming for efficient non-volatile processorsabstractWearable devices are becoming increasingly important in our daily lives. Energy harvesting instead of battery is a better power source for these wearable devices due to many advantages. However, harvested energy is often unstable and program execution will be frequently interrupted. Non-volatile processors demonstrate promising advantages to back up volatile state before the system energy is depleted. But Non-volatile processors require additional memory for backing up, thus introducing non-negligible overhead in terms of energy, runtime as well as chip area. In this work, we target at non-volatile register reduction for energy harvesting based wearable devices. This paper proposes to stack trimming the memory footprint via a novel compiler directed method. The evaluation results deliver on average 28.6% reduction of non-volatile register files for backing up stack area, with ultra low runtime overhead. Qing'an Li, Mengying Zhao, Jingtong Hu, Yongpan Liu, Yanxiang He, Chun Jason Xue |
DAC | 2 |
| 2015 | Fixing the broken time machine: consistency-aware checkpointing for energy harvesting powered non-volatile processorabstractEnergy harvesting has become a favorable alternative to batteries for wearable embedded systems since it is more environmental and user friendly. However, harvested energy is intrinsically unstable, which could frequently interrupt a processor's execution. To tackle this problem, non-volatile processors have been proposed to checkpoint the whole volatile processor state into attached non-volatile memories periodically. When power resumes, the processor can copy the checkpointed state back to volatile memories and continue execution. However, without careful consideration, the process of checkpointing and resuming could cause inconsistency among different memory addresses and lead to irreversible errors. In this paper, we present a consistency aware checkpointing scheme that ensures correctness for all checkpoints. The proposed technique efficiently identifies all possible inconsistency positions in programs and inserts auxiliary code to ensure correctness. Evaluation results show that the proposed checkpointing technique can successfully eliminate inconsistency errors and greatly reduce the checkpointing overhead. Mimi Xie, Mengying Zhao, Jingtong Hu, Yongpan Liu, Chun Jason Xue |
DAC | 2 |
| 2015 | Software assisted non-volatile register reduction for energy harvesting based cyber-physical system
Mengying Zhao, Qing'an Li, Mimi Xie, Yongpan Liu, Jingtong Hu, Chun Jason Xue |
DATE | 1 |
| 2015 | C3: Cooperative Code Positioning and Cache Locking for WCET MinimizationabstractWorst-case execution time (WCET) is an important metric for designing real-time systems. Previous work such as code positioning and cache locking has been proposed for WCET reduction. Traditionally, these two techniques have been applied independently, which cannot derive desirably tight WCET. In this paper, a cooperative code positioning and static instruction cache locking (C3) framework is proposed to minimize WCET for real-time systems. We first propose the locking-basic-blocks selection heuristic to choose the most beneficial basic blocks on the longest path to improve WCET with locking. Then the code positioning technique is proposed to generate the memory layouts to reduce cache conflicts that influence WCET among these selected basic blocks or their functions. Based on the cache-locking-aware memory layouts, static cache locking is implemented to identify the most appropriate memory layout and also minimize WCET. The experiments show that compared to previous work, C3 reduces WCET significantly. Mengying Zhao, Chun Jason Xue |
RTCSA | 2 |
| 2015 | Modular Performance Analysis of Energy-Harvesting Real-Time Networked SystemsabstractThis paper studies the performance analysis problem of energy-harvesting real-time network systems in the Real-Time Calculus (RTC) framework. The behavior of an energy-harvesting node turns out to be a generalization of two known components in RTC: it behaves like an AND connector if the capacitor used to temporally store surplus energy has unlimited capacity and there is no energy loss, while it behaves like a greedy processing component (GPC) if the size of the capacitor is zero and thus surplus energy is lost or passed to other nodes immediately. In this paper, methods are developed to analyze the worst-case performance, in terms of delay and backlog, of energy-harvesting nodes as well as compute upper/lower bounds of their data and energy outputs. Moreover, with the proposed analysis methods, we disclose some interesting properties of the worst-case behaviors of energy-harvesting systems, which provide useful information to guide system design. Experiments are conducted to evaluate our theoretical contributions and also confirm that the disclosed properties are not just the result of our analysis, but indeed hold in realistic system behaviors. Nan Guan, Mengying Zhao, Chun Jason Xue, Yongpan Liu, Wang Yi 0001 |
RTSS | 2 |
| 2015 | Wear Relief for High-Density Phase Change Memory Through Cell Morphing Considering Process VariationabstractDue to the scalability and large leakage power, dynamic random-access memory (DRAM) has a lot of challenges in scaling. As an alternative, phase change memory (PCM) has demonstrated promising potential to serve as the main memory in deep submicrometer regime. The broad resistance range of PCM cells enables several cell modes with various densities, pertaining to multiple level cell (MLC), triple state cell (TSC), and single level cell (SLC). High-density mode outperforms low-density ones in terms of capacity and cost-per-bit, but suffers from a weaker cell endurance. Wear leveling strategies are proposed to enhance the memory endurance but encounter more challenges with the aggravating process variation. Due to endurance variations, physical domains are fabricated with irregular tenacity. As a result, balanced write traffic, which is the objective of traditional wear leveling, cannot fully exploit the PCM endurance since the weak parts will be worn out sooner than others. In this paper, considering process variation, we propose a cell morphing based wear leveling scheme. Cell morphing refers to the cell mode transformation between high density (e.g., MLC) and low densities (e.g., TSC and SLC). Instead of redistributing write operations, the proposed wear leveling scheme dynamically transforms weak and frequently written portions into low-density mode for endurance benefits. Multitier cell morphing schemes are proposed to support mode transformation among multiple density levels. The experimental results show 236% endurance improvement for single-tier cell morphing and 209% for two-tier cell morphing with 2% low-density page percentage, when compared with the most related work. Mengying Zhao, Lei Jiang 0001, Liang Shi 0001, Youtao Zhang, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Joint WCET and Update Activity Minimization for Cyber-Physical SystemsabstractA cyber-physical system (CPS) is a desirable computing platform for many industrial and scientific applications, such as industrial process monitoring, environmental monitoring, chemical processes, and battlefield surveillance. The application of CPSs has two challenges: First, CPSs often include a number of sensor nodes. Update of preloaded code on remote sensor nodes powered by batteries is extremely energy consuming. The code update issue in the energy-sensitive CPS must be carefully considered. Second, CPSs are often real-time embedded systems with real-time properties. Worst-case execution time (WCET) is one of the most important metrics in real-time system design. Whereas existing works only consider one of these two challenges at a time, in this article, a compiler optimization—joint WCET and update-conscious compilation, or WUCC—is proposed to jointly consider WCET and code update for CPSs. The novelty of the proposed approach is that the WCET problem and code update problem are considered concurrently such that a balanced solution with minimal WCET and minimal code difference can be achieved. The experimental results show that the proposed technique can minimize WCET and code difference effectively. Yazhi Huang, Mengying Zhao, Chun Jason Xue |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Joint Profit and Process Variation Aware High Level Synthesis With Speed BinningabstractAs integrated circuits continuously scale up, process variation plays an increasingly significant role in system design and semiconductor economic return. In this paper, we explore the potential of profit improvement under the inherent semiconductor variability based on the speed binning technique. We aim to develop a set of high level synthesis (HLS) solutions, for which purpose heuristic techniques, including allocation, scheduling, and resource binding, are proposed. The goal is to construct designs that maximize the number of chips that can be sold at the most advantageous price, leading to the maximization of the overall profit. In addition, a genetic algorithm-based formulation is constructed for HLS solutions. Then, we complement the HLS techniques with near-optimal bin placement strategies for further profit improvement. Experimental results confirm the superiority of the HLS results and the associated improvement in profit margins. Mengying Zhao, Alex Orailoglu, Chun Jason Xue |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Retention Trimming for Wear Reduction of Flash Memory Storage SystemsabstractNAND flash memory has been widely applied in embedded systems, personal computer systems, and data centers. However, with the development of flash memory, including its technology scaling and density improvement, the endurance of flash memory becomes a bottleneck. In this work, with the understanding of the relationship between data retention time and flash wearing, a retention trimming approach, which trims data retention time based on the time intervals between data updating, is proposed to reduce the wearing of flash memory. Reduced wearing of flash memory will improve the endurance of the flash memory. Extensive experimental results show that the proposed technique achieves significant wearing reduction for flash memory through retention trimming. Liang Shi 0001, Kaijie Wu 0001, Mengying Zhao, Chun Jason Xue, Edwin H.-M. Sha |
DAC | 3 |
| 2014 | SLC-enabled Wear Leveling for MLC PCM Considering Process VariationabstractPhase change memory is becoming one of the most promising candidates to replace DRAM as main memory in deep silicon regime. Multi-level cell (MLC) PCM outperforms single level cell (SLC) in terms of capacity while suffering from a weaker cell endurance. Wear leveling strategies are proposed to enhance the endurance but encounters more challenges with the aggravating process variation. Due to endurance variations, balanced write traffic cannot fully exploit the PCM endurance since the weak parts will be worn out sooner than others. In this work, considering process variation, we propose an SLC-enabled wear leveling scheme through dynamic and adaptive mode transformation from MLC to SLC. Instead of redistributing write operations, the proposed scheme dynamically transforms weak and write-dense parts into SLC mode for endurance benefits. The experimental results show that the proposed scheme can improve the endurance by 215% with 4% storage overhead while maintaining the capacity advantage of MLC, compared with the most related work. Mengying Zhao, Lei Jiang 0001, Youtao Zhang, Chun Jason Xue |
DAC | 1 |
| 2014 | Leveling to the last mile: Near-zero-cost bit level wear leveling for PCM-based main memoryabstractPhase change memory (PCM) has demonstrated great potential as an alternative of DRAM to serve as main memory due to its favorable characteristics of non-volatility, scalability and near-zero leakage power. However, the comparatively poor endurance of PCM largely limits its adoption. Wear leveling strategies targeting to even write distributions have been proposed at different granularities and on various memory hierarchies for PCM endurance enhancement. Write operations are distributed across the memory through migrating data from heavily written locations to less burdened ones, which is usually guided by counters recording the number of writes. However, evenly distributing writes at a coarse granularity cannot deliver the best endurance results as write distributions are highly imbalanced even at the bit level. In this work, we propose a near-zero-cost bit-level wear leveling strategy to improve PCM endurance. The proposed technique can be combined with various coarse-grained wear leveling strategies. Experiment results show 102% endurance enhancement on average, which is 34% higher than the most related work, with significantly lower storage, performance and energy overheads. Mengying Zhao, Liang Shi 0001, Chengmo Yang, Chun Jason Xue |
ICCD | 1 |
| 2014 | Sleep-aware variable partitioning for energy-efficient hybrid PRAM and DRAM main memoryabstractEnergy consumption of memories is always a significant issue for computing systems. Recently, hybrid PRAM and DRAM memory architectures have been proposed. It combines the advantages of DRAM and PRAM, such as low leakage power in PRAM and short write latency in DRAM. However, the leakage power in DRAM is still considerable in hybrid memories. The leakage power can only be reduced by turning DRAM into sleep state. In this paper, a novel proximity concept is proposed to guide the variable partitioning to maximize the possibility of turning DRAM into sleep mode. A novel Sleep-Aware Variable Partition Algorithm (SAVPA) is then proposed with the objective of maximizing the sleep time of DRAM while satisfying the performance and endurance constraints. The experiment results show that SAVPA reduces the energy consumption by 11.25% in average (up to 15.84%) compared to the state-of-art work with simple sleep technique. Chenchen Fu, Mengying Zhao, Chun Jason Xue, Alex Orailoglu |
ISLPED | 2 |
| 2014 | Exploiting parallelism in I/O scheduling for access conflict minimization in flash-based solid state drivesabstractSolid state drives (SSDs) have been widely deployed in personal computers, data centers, and cloud storages. In order to improve performance, SSDs are usually constructed with a number of channels with each channel connecting to a number of NAND flash chips. Despite the rich parallelism offered by multiple channels and multiple chips per channel, recent studies show that the utilization of flash chips (i.e. the number of flash chips being accessed simultaneously) is seriously low. Our study shows that the low chip utilization is caused by the access conflict among I/O requests. In this work, we propose Parallel Issue Queuing (PIQ), a novel I/O scheduler at the host system, to minimize the access conflicts between I/O requests. The proposed PIQ schedules I/O requests without conflicts into the same batch and I/O requests with conflicts into different batches. Hence the multiple I/O requests in one batch can be fulfilled simultaneously by exploiting the rich parallelism of SSD. And because PIQ is implemented at the host side, it can take advantage of rich resource at host system such as main memory and CPU, which makes the overhead negligible. Extensive experimental results show that PIQ delivers significant performance improvement to the applications that have heavy access conflicts. Congming Gao, Liang Shi 0001, Mengying Zhao, Chun Jason Xue, Kaijie Wu 0001, Edwin H.-M. Sha |
MSST | 3 |
| 2014 | Migration-Aware Loop Retiming for STT-RAM-Based Hybrid Cache in Embedded SystemsabstractRecently hybrid cache architecture consisting of both spin-transfer torque RAM (STT-RAM) and SRAM has been proposed for energy efficiency. In hybrid caches, migration-based techniques have been proposed. A migration technique dynamically moves write-intensive and read-intensive data between STT-RAM and SRAM to explore the advantages of hybrid cache. Meanwhile, migrations also introduce extra reads and writes during data movements. For stencil loops with read and write data dependencies, we observe that migration overhead is significant, and migrations closely correlate to the interleaved read and write memory access pattern in a memory block. This paper proposes a loop retiming framework during compilation to reduce the migration overhead by changing the interleaved memory access pattern. With the proposed loop retiming technique, the interleaved memory accesses can be significantly reduced so that migration overhead is mitigated, and energy efficiency of hybrid cache is significantly improved. The experimental results have shown that, with the proposed methods, on average, the migration number is reduced up to 27.1% and the cache dynamic energy is reduced up to 14.0%. Keni Qiu, Mengying Zhao, Qing'an Li, Chenchen Fu, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2014 | Error Model Guided Joint Performance and Endurance Optimization for Flash MemoryabstractAs flash memory has better performance than hard disks, it has been widely applied in embedded systems, personal computers, and data centers as storage components. However, endurance and write performance are the two key challenges in the deployment of flash memory. In this paper, with the awareness of errors induced from write operations, endurance, and retention time, a stage-based optimization approach is proposed to improve the write performance and endurance at different usage stages of flash memory. A series of trace-driven simulations show that the proposed approach outperforms a set of state-of-the-art approaches in terms of write performance and lifetime. Liang Shi 0001, Keni Qiu, Mengying Zhao, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2014 | Branch Prediction-Directed Dynamic Instruction Cache Locking for Embedded SystemsabstractCache locking is a cache management technique to preclude the replacement of locked cache contents. Cache locking is often adopted to improve cache access predictability in Worst-Case Execution Time (WCET) analysis. Static cache locking methods have been proposed recently to improve Average-Case Execution Time (ACET) performance. This article presents an approach, Branch Prediction-directed Dynamic Cache Locking (BPDCL), to improve system performance through cache conflict miss reduction. In the proposed approach, the control flow graph of a program is first partitioned into disjoint execution regions, then memory blocks worth locking are determined by calculating the locking profit for each region. These two steps are conducted during compilation time. At runtime, directed by branch predictions, locking routines are prefetched into a small high-speed buffer. The predetermined cache locking contents are loaded and locked at specific execution points during program execution. Experimental results show that the proposed BPDCL method exhibits an average improvement of 25.9%, 13.8%, and 8.0% on cache miss rate reduction in comparison to cases with no cache locking, the static locking method, and the dynamic locking method, respectively. Keni Qiu, Mengying Zhao, Chun Jason Xue, Alex Orailoglu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2014 | Compiler-Assisted STT-RAM-Based Hybrid Cache for Energy Efficient Embedded SystemsabstractHybrid caches consisting of static RAM (SRAM) and spin-torque transfer (STT)-RAM have been proposed recently for energy efficiency. To explore the advantages of hybrid cache, most of the management strategies for hybrid caches employ migration-based techniques to dynamically move write-intensive data from STT-RAM to SRAM. These techniques involve additional access operations, and thus lead to extra overheads. In this paper, we propose two compilation-based approaches to improve the energy efficiency and performance of STT-RAM-based hybrid cache by reducing the migration overheads. The first approach, migration-aware data layout, is proposed to reduce the migrations by rearranging the data layout. The second approach, migration-aware cache locking, is proposed to reduce the migrations by locking migration-intensive memory blocks into SRAM part of hybrid cache. Furthermore, experiments show that these two methods can be combined to reduce more migrations. The reduction of migration overheads can improve the energy efficiency and performance of STT-RAM-based hybrid cache. Experimental results show that, combining these two methods, on average, the number of write operations on STT-RAM is reduced by 17.6%, the number of migrations is reduced by 38.9%, the total dynamic energy is reduced by 15.6%, and the total access latency is reduced by 13.8%. Qing'an Li, Jianhua Li 0003, Liang Shi 0001, Mengying Zhao, Chun Jason Xue, Yanxiang He |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | Migration-aware loop retiming for STT-RAM based hybrid cache for embedded systemsabstractIn hybrid cache architecture consisting of both STT-RAM and SRAM, migration based techniques have been proposed. The migration technique dynamically moves write-intensive and read-intensive data between STT-RAM and SRAM to explore the advantage of hybrid cache. Meanwhile, migrations induce extra read and write overhead during data movements. For loops with intensive data array operations, we observe that migration overhead is significant and migrations closely correlate to the interleaved read and write access pattern in a memory block. This paper proposes a loop retiming framework to reduce the migration overhead by changing the interleaved memory access pattern. The experimental results show that with the proposed method, migrations are significantly reduced without any hardware modification. As a result, energy efficiency and performance of hybrid cache can be improved. Keni Qiu, Mengying Zhao, Chenchen Fu, Liang Shi 0001, Chun Jason Xue |
ASAP | 2 |
| 2013 | WUCC: Joint WCET and Update Conscious Compilation for cyber-physical systemsabstractThe cyber-physical system (CPS) is a desirable computing platform for many industrial and scientific applications. However, the application of CPSs has two challenges: First, CPSs often include a number of sensor nodes. Update of preloaded code on remote sensor nodes powered by batteries is extremely energy-consuming. The code update issue in the energy sensitive CPS must be carefully considered; Second, CPSs are often real-time embedded systems with real-time properties. Worst-Case Execution Time (WCET) is one of the most important metrics in real-time system design. While existing works only consider one of these two challenges at a time, in this paper, a compiler-level optimization, Joint WCET and Update Conscious Compilation (WUCC), is proposed to jointly consider WCET and code update for cyber-physical systems. The novelty of the proposed approach is that the WCET problem and code update problem are considered concurrently such that a balanced solution with minimal WCET and minimal code difference can be achieved. The experimental results show that the proposed technique can minimize WCET and code difference effectively. Yazhi Huang, Mengying Zhao, Chun Jason Xue |
ASP-DAC | 2 |
| 2013 | Profit maximization through process variation aware high level synthesis with speed binningabstractAs integrated circuits continuously scale up, process variation plays an increasingly significant role in system design and semiconductor economic return. In this paper, we explore the potential of profit improvement under the inherent semiconductor variability based on the speed binning technique. We first accordingly propose a set of high level synthesis techniques, including allocation, scheduling and resource binding, thus essentially constructing designs that maximize the number of chips that can be sold at the most advantageous price, leading to the maximization of the overall profit. We explore subsequently the optimal bin placement strategy for further profit improvement. Experimental results confirm the superiority of the high level synthesis results and the associated improvement in profit margins. Mengying Zhao, Alex Orailoglu, Chun Jason Xue |
DATE | 1 |
| 2013 | Branch Prediction directed Dynamic instruction Cache Locking for embedded systemsabstractCache locking is a cache management technique to preclude the replacement of locked cache contents. Cache locking is often used to improve cache access predictability in Worst-Case Execution Time (WCET) analysis. Static cache locking methods have been proposed recently to improve average system performance. This paper presents an approach, Branch Prediction directed Dynamic Cache Locking (BPDCL), to improve average system performance through effective cache conflict miss reduction in different execution regions. In this proposed approach, the control flow graph of a program is partitioned into regions and memory blocks worth locking for each region are calculated during compilation time. At runtime, directed by branch predictions, locking routines are prefetched into a high-speed buffer. The pre-determined cache locking contents are loaded and locked at specific execution points during program execution. Experimental results show that the proposed BPDCL method exhibits an average improvement of 21.8% and 10.3% on cache miss rate reduction in comparison to the case with no cache locking and the static locking method respectively. Keni Qiu, Mengying Zhao, Chun Jason Xue, Alex Orailoglu |
RTCSA | 2 |
| 2013 | Data re-allocation enabled cache locking for embedded systemsabstractCache locking is a cache management technique to preclude the replacement of locked contents. Recently, instruction cache locking has been applied to improve average-case execution time (ACET). However, we observe that the prior instruction cache locking method shows very limited performance improve-ment for data cache. The main reason lies in that, data access similarity in data memory blocks is weaker than that in code memory blocks. This paper proposes a data re-allocation enabled cache locking approach which can significantly enhance locking efficiency for data cache and thus improve system performance. The experimental results show that with the proposed approach, on average, the miss rate is reduced by 9.1% and execution cycles are reduced by 9.4% across a suite of benchmarks. Keni Qiu, Mengying Zhao, Chenchen Fu, Chun Jason Xue |
VLSI-SoC | 2 |
| 2012 | Quality-retaining OLED dynamic voltage scaling for video streaming applications on mobile devicesabstractThis paper developed a dynamic voltage scaling (DVS) technique for the power management of the OLED display on mobile devices in video streaming applications. An optimal voltage control scheme is proposed under input constraints. Fine-grained DVS technique is applied to maximize the power saving by leveraging the locality of the display content. The display quality is retained by monitoring structural-similarity-index (SSIM) during the optimization, subject to the hardware constraints like voltage regulator response time. Simulation results on four typical video test benchmarks show that the proposed technique saves 19.05%~49.05% OLED power on average while maintaining a high display quality (SSIM > 0.98) all the time. The power saving efficiency of the proposed technique varies at different display resolutions, refresh rates, and display contents. Xiang Chen 0010, Yiran Chen 0001, Mengying Zhao, Chun Jason Xue |
DAC | 4 |
| 2012 | Mobile devices user - The subscriber and also the publisher of real-time OLED display power management planabstractOLED (Organic Light Emitting Diode) technology has already been adopted in many modern smart mobile devices, including cellphones, tablets, laptop etc. However, the power dissipation of displays in some applications like real-time video streaming, significantly limits the smart mobile devices' battery life and influences user experience. In this work, we applied a set of power management techniques that based on the dynamic voltage scaling (DVS) to minimize the power consumption of OLED display. Circuit inventions are also introduced to enable the local voltage scaling of AMOLED display panel while the human vision reception criteria can still be met. We then apply the DVS-based power management techniques to online video streaming, which is the most energy-hungry application of mobile devices: For any known type of mobile devices with AMOLED displays, the DVS power management scheme of a specific video stream can be pre-analyzed and shared on network, i.e., the Cloud or local party that provide the video. The DVS scheme can be downloaded in the real-time simultaneously when the video is broadcasted to the end users. By doing so, the power and computation overheads of DVS optimization are amortized among all the users, achieving a high service quality. Furthermore, the optimization can be executed locally (at user end) and shared by all the users. Yiran Chen 0001, Xiang Chen 0010, Mengying Zhao, Chun Jason Xue |
ICCAD | 3 |
| 2012 | Active compensation technique for the thin-film transistor variations and OLED aging of mobile device displaysabstractOLED is becoming the main stream display for mobile devices. The process variations of thin-film transistors (TFT) and the aging degradation of OLED devices severely impact the display quality and the user experience on mobile devices throughout lifetime. In this paper, we quantitatively study the nonuniformity of OLED display panels incurred by the TFT variations and OLED cell aging effect. Furthermore, we develop a pixel level sensing circuit that detects and quantifies the nonuniformity condition and corresponding compensating technique. This proposed technique can be actively invoked based on various conditions with flexible configuration and minimal extra overhead, which is suitable to be integrated with mobile displays. The proposed technique's performance is simulated with different display contents. Experiments show that: for the proposed sensing circuit, the error rate for TFT process variation evaluation is only 4.7%~7.9%, and 0.29%~7.6% for OLED aging degradation. And for typical mobile display content, the compensation rate reaches 94.04%~100%. After applying the proposed techniques, the nonuniformity is unrecognizable and the OLED display panel's lifespan is highly extended. Xiang Chen 0010, Beiye Liu, Yiran Chen 0001, Mengying Zhao, Chun Jason Xue |
ICCAD | 4 |
| 2012 | WCET-aware re-scheduling register allocation for real-time embedded systems with clustered VLIW architectureabstractWorst-Case Execution Time (WCET) is one of the most important metrics in real-time embedded system design. For embedded systems with clustered VLIW architecture, register allocation, instruction scheduling, and cluster assignment are three key activities to pursue code optimization which have profound impact on WCET. At the same time, these three activities exhibit a phase ordering problem: Independently performing register allocation, scheduling and cluster assignment could have a negative effect on the other phases, thereby generating sub-optimal compiled codes. In this paper, a compiler level optimization, namely WCET-aware Re-scheduling Register Allocation (WRRA), is proposed to achieve WCET minimization for real-time embedded systems with clustered VLIW architecture. The novelty of the proposed approach is that the effects of register allocation, instruction scheduling and cluster assignment on the quality of generated code are taken into account for WCET minimization. These three compilation processes are integrated into a single phase to obtain a balanced result. The proposed technique is implemented in Trimaran 4.0. The experimental results show that the proposed technique can reduce WCET effectively, by 33% on average. Yazhi Huang, Mengying Zhao, Chun Jason Xue |
LCTES | 2 |
| 2012 | Compiler-assisted preferred caching for embedded systems with STT-RAM based hybrid cacheabstractAs technology scales down, energy consumption is becoming a big problem for traditional SRAM-based cache hierarchies. The emerging Spin-Torque Transfer RAM (STT-RAM) is a promising replacement for large on-chip cache due to its ultra low leakage power and high storage density. However, write operations on STT-RAM suffer from considerably higher energy consumption and longer latency than SRAM. Hybrid cache consisting of both SRAM and STT-RAM has been proposed recently for both performance and energy efficiency. Most management strategies for hybrid caches employ migration-based techniques to dynamically move write-intensive data from STT-RAM to SRAM. These techniques lead to extra overheads. In this paper, we propose a compiler-assisted approach, preferred caching, to significantly reduce the migration overhead by giving migration-intensive memory blocks the preference for the SRAM part of the hybrid cache. Furthermore, a data assignment technique is proposed to improve the efficiency of preferred caching. The reduction of migration overhead can in turn improve the performance and energy efficiency of STT-RAM based hybrid cache. The experimental results show that, with the proposed techniques, on average, the number of migrations is reduced by 21.3%, the total latency is reduced by 8.0% and the total dynamic energy is reduced by 10.8%. Qing'an Li, Mengying Zhao, Chun Jason Xue, Yanxiang He |
LCTES | 2 |