EDBT 2026 Demo / reviewers in the wild / expert
Zhiping Jia
dblp:99/6654
· DBLP profile ↗
101ranked-venue papers
1as first author
31since 2021 · last 2024
0000-0002-7769-4771ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 66 · 1 first-author · 24 since 2021Security and privacy · 12 · 1 since 2021Computer networks · 10 · 2 since 2021Software engineering, systems software and programming languages · 7 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Efficient Reconfiguration through Lightweight Input Inversion for MLC NVFPGAsabstractNonvolatile field programmable gate arrays (NVFP-GAs) have been proposed to address the challenges raised by artificial intelligence and big data related applications, since nonvolatile memories (NVMs) introduce advantages of high storage density, low leakage power, and high system robustness. In addition, multi-level cell (MLC), which can store multiple bits within one memory cell, further improves the logic density of NVFPGAs. However, the inefficient write operation of MLC NVM significantly increases the reconfiguration cost in aspects of energy, latency, and lifetime. In this paper, we focus on the reconfiguration cost of MLC LUTs in NVFPGA and propose a lightweight input inversion based scheme to reduce the reconfiguration cost. Inversion flexibility is defined and modeled for LUT inputs to guide the proposed scheme. We also discuss how the proposed scheme can be combined with other existing write reduction strategies. Evaluation shows the proposed scheme can reduce reconfiguration cost by 10.01 % with negligible overhead. Huichuan Zheng, Mengying Zhao, Yuqing Xiong, Xiaojun Cai, Zhiping Jia |
DATE | 6 |
| 2024 | An Access Pattern-aware Hybrid Learning-based and Conventional Mapping for Solid-State DrivesabstractLearning-based address mapping strategy has been proposed as an alternative to the conventional one-to-one mapping strategy for logical to physical address translation in flash-based Solid-State Drives (SSDs). This strategy employs a machine learning (ML)-based index to replace the memory-intensive mappings, thereby saving memory space. However, current learning-based mapping strategies mainly focus on the sequential characteristics of I/O workloads, making them less effective in scenarios with frequent updates. In this work, we introduce an access pattern-aware hybrid learning-based and conventional mapping scheme for SSDs, called APH-FTL. This approach features a machine learning-driven data classifier that employs distinct address mapping strategies for data requests with varying access patterns. Additionally, we develop a benefit-guided space regulation strategy to manage in-device memory space effectively and provide a solution to the challenge of mapping policies transformation. We evaluate APH-FTL through simulations using various real-world workloads. Experimental results demonstrate that APH-FTL significantly improves performance, reducing response time by up to 76.59% and write amplification (WA) by up to 82.97% compared with state-of-the-art schemes. Xiaosu Guo, Jie Wang 0128, Zhaoyan Shen, Dongxiao Yu, Zhiping Jia, Bingzhe Li |
ICCAD | 6 |
| 2024 | ASHL: An Adaptive Multi-Stage Distributed Deep Learning Training Scheme for Heterogeneous EnvironmentsabstractWith the increment of data sets and models sizes, distributed deep learning has been proposed to accelerate training and improve the accuracy of DNN models. The parameter server framework is a popular collaborative architecture for data-parallel training, which works well for homogeneous environments by properly aggregating the computation/communication capabilities of different workers. However, in heterogeneous environments, the resources of different workers vary a lot. Some stragglers may seriously limit the whole speed, which impacts the overall training process. In this paper, we propose an adaptive multi-stage distributed deep learning training framework, named ASHL, for heterogeneous environments. First, a profiling scheme is proposed to capture the capabilities of each worker to reasonably plan the training and communication tasks on each worker, and lay the foundation for the formal training. Second, a hybrid-mode training scheme (i.e., coarse-grained and fined-grained training) is proposed to balance the model accuracy and training speed. The coarse-grained training scheme (named AHL) adopts an asynchronous communication strategy, which involves less frequent communications. Its main goal is to make the model quickly converge to a certain level. The fine-grained training stage (named SHL) uses a semi-asynchronous communication strategy and adopts a high communication frequency. Its main goal is to improve the model convergence effect. Finally, a compression-based communication scheme is proposed to further increase the communication efficiency of the training process. Our experimental results show that ASHL reduces the overall training time by more than 35% to converge to the same degree and has better generalization ability compared with state-of-the-art schemes like ADSP. Zhaoyan Shen, Qingxiang Tang, Tianren Zhou, Yuhao Zhang 0006, Zhiping Jia, Dongxiao Yu, Zhiyong Zhang 0006, Bingzhe Li |
IEEE Trans. Computers | 5 |
| 2024 | Branch Predictor Design for Energy Harvesting Powered Nonvolatile ProcessorsabstractNon-volatile processors are proposed for ambient energy harvesting systems to enable accumulative computing across power failures. They employ nonvolatile memory for processor status backup before power outage and resume the system after power recovers. A straightforward backup policy is to back up all volatile data in processors, but it induces high backup cost. In this paper, we focus on branch predictor, an important component in processor, and propose efficient backup schemes to reduce backup cost while maintaining its prediction ability. We first analyze the modules in both traditional and artificial intelligence (AI) assisted designs of branch predictor, and accordingly propose three backup mechanisms pertaining to saturation-driven, locality-driven and maturity-driven backup. On the basis of these mechanisms, adaptive backup branch predictors are designed. Evaluation shows that, with traditional Tournament architecture, the proposed design achieves 15.9% and 54.1% energy reduction when compared with no-backup and all-backup strategy. For AI assisted branch predictor, the proposed design achieves 27.5% and 82.2% energy saving. Mengying Zhao, Lihao Dong, Chun Jason Xue, Dongxiao Yu, Xiaojun Cai, Zhiping Jia |
IEEE Trans. Computers | 7 |
| 2024 | A Semantic-Integrated LSM-Tree-Based Key-Value Storage Engine for Blockchain SystemsabstractBlockchain systems play an important role in distributed ledgers, database systems, etc. As more and more blocks are mined, the storage burden of blockchain system is significantly increased. The current blockchain system uniformly transforms all its data into key-value (KV) items and stores them to the underlying Log-Structure Merged tree (LSM-tree) storage engine ignoring the software semantics. Consequently, it not only aggravates the write amplification effect of the storage engine, but also increases the redundancy of data query steps, resulting in the performance bottleneck of blockchain system. In this paper, we propose a semantic-integrated LSM-tree based Key-Value storage engine for blockchain systems, called Block-LSM, which significantly improves the data synchronization and data query efficiency of blockchain system. Specifically, we first design a shared prefix scheme to transform blockchain data into ordered KV pairs to alleviate the key range overlaps of different levels in the underlying LSM-tree based storage engine. Moreover, we propose to maintain several semantic-orientated memory buffers to isolate different kinds of blockchain data, and implement memory buffer space management strategy to further improve memory efficiency. To save space overhead, Block-LSM further aggregates multiple blocks into a group and assigns the same prefix to all KV items from the same block group. We also reduce step redundancy in transaction queries by modifying the body data storage format. Finally, we implement Block-LSM in a real blockchain environment and conduct a series of comparative experiments with the typical blockchain system Ethereum. The evaluation results show that Block-LSM significantly reduces up to 7.56× storage write amplification and increases throughput by 8.64× compared with the original Ethereum design. In terms of data lookups (i.e. transaction and account lookup), Block-LSM improves the throughput by 50% compared to the original Ethereum design. Yuhao Zhang 0006, Xiaojun Cai, Zhiping Jia, Zhaoyan Shen, Yi Wang 0003, Zili Shao, Bingzhe Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Incremental Model Evolution for Power System Security Early Warning Based on Knowledge Distillation and Active LearningabstractSecurity early warning is crucial for resisting power system risk. To deal with renewable uncertainty, offline trained early-warning models need to be evolved incrementally. Catastrophic forgetting is main obstacle of model evolution. With the evolution times of early-warning models increasing, the capability of existing methods to mitigate catastrophic forgetting gradually decreases. Knowledge distillation can transfer the learned knowledge from the previous model to the updated model. To mitigate catastrophic forgetting, an incremental model evolution method for security early warning based on knowledge distillation and active learning is proposed. First, an incremental style-based generative adversarial network is formulated to generate renewable power scenarios, which utilizes knowledge distillation to retain previously learned knowledge. An improved concept drift detection method is proposed to determine model evolution moment. Then, an incremental deep regression model based on stacked target-related denoising autoencoder and knowledge distillation is constructed to assess system security index. Finally, active learning is deployed to select informative new samples and solve the objective conflict problem of knowledge distillation. Simulation results of a provincial grid in China demonstrate that proposed method can constantly improve early warning accuracy and effectively retain knowledge learned by previous samples. The proposed method can improve the adaptability of early-warning models to actual power system operating conditions, which enhances the intelligence level of security analysis to copy with renewable uncertainty. Jiongcheng Yan, Changgang Li, Yutian Liu 0002, Dongxiao Yu, Zhiping Jia |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | Reinforcement Learning-Assisted Management for Convertible SSDsabstractConvertible SSDs, which allow flash cells to convert between different types of flash cells (e.g., SLC/MLC/TLC/QLC), are designed for achieving both high performance and high density. However, previous designs with two types of flash cells encounter a performance cliff degradation once the flash cells of single bit mode are consumed. In this work, we propose a novel level-based convertible SSD (e.g., including SLC-MLC-QLC), named RL-cSSD, that adopts an intermediate layer (e.g., MLC) as a performance cushion. A reinforcement learning-assisted device management scheme is designed to coordinate the data allocation, garbage collection and flash conversion processes considering both the SSD internal status and workload patterns. We evaluated RL-cSSD with various real-world workloads based on simulation. The experimental results show that the proposed RL-cSSD provides 72.98% higher performance on average compared with state-of-the-art schemes. Zhiping Jia, Mengying Zhao, Zhaoyan Shen, Bingzhe Li |
DAC | 3 |
| 2023 | Correlation-guided Placement for Nonvolatile FPGAsabstractNonvolatile FPGAs have advantages of high density and near-zero leakage power compared with traditional SRAM-based FPGAs. However, they have lifetime issue. To deal with this problem, a series of configuration files can be generated with various logical-to-physical mappings so that intensive writes can be distributed to different physical regions for wear leveling. Currently, the configuration files are independently generated, which is time-consuming. In this paper, we propose to investigate correlations between components and use them to guide the computer-aided design (CAD) flow to speed up the procedure of deriving configuration files. Specifically, we develop dynamic probabilities to drive the swapping of placement step in the CAD flow to push components to locate appropriate positions quickly. Evaluation shows that the proposed schemes can deliver 36.32% reduction in number of swappings when compared with existing strategies, while maintaining comparable performance and lifetime. Mengying Zhao, Fanjin Xu, Huichuan Zheng, Yuqing Xiong, Zhiping Jia, Xiaojun Cai |
DAC | 6 |
| 2023 | Analysis of Augmentations in Contrastive Learning for Parkinson's Disease Diagnosis
Shuangyi Wang, Tianren Zhou, Zhaoyan Shen, Zhiping Jia |
ICANN (4) | 4 |
| 2023 | Runtime Row/Column Activation Pruning for ReRAM-based Processing-in-Memory DNN AcceleratorsabstractResistive random access memory (ReRAM)-based processing-in-memory (PIM) DNN accelerators have shown great potential in improving model efficiency and saving energy. To further improve memory and computation efficiency, model weight sparsity has been widely explored in ReRAM-based accelerator designs. However, these optimized accelerators rarely touched the model activation sparsity. In this paper, we observe that there exist plenty sparse rows/columns in the DNN model activation matrix which have negligible effect on accuracy, termed as insensitive rows/columns. Pruning them has little impact on model accuracy but would have significant potential to improve the performance and energy efficiency of DNN accelerators. Therefore, we propose a new ReRAM-based PIM accelerator, named as RapPIM, to take advantage of the model activation sparsity. In RapPIM, we first propose an insensitive activation rows/columns pruning method to search and prune the insensitive rows/columns. Then, we present an activation low-bits skipping strategy and a forward propagation delay hiding strategy to further improve model performance and minimize the latency of activation pruning on forward propagation. Our evaluations with several well-known DNN models show that the RapPIM achieves up to 2.40× speedup and 44.82% power reduction compared with the state-of-the-art ReRAM-based accelerator. Xikun Jiang, Zhaoyan Shen, Siqing Sun, Ping Yin, Zhiping Jia, Lei Ju 0001, Zhiyong Zhang 0006, Dongxiao Yu |
ICCAD | 5 |
| 2023 | PRAP-PIM: A weight pattern reusing aware pruning method for ReRAM-based PIM DNN acceleratorsabstractResistive Random-Access Memory (ReRAM) based Processing-in-Memory (PIM) frameworks are proposed to accelerate the working process of DNN models by eliminating the data movement between the computing and memory units. To further mitigate the space and energy consumption, DNN model weight sparsity and weight pattern repetition are exploited to optimize these ReRAM-based accelerators. However, most of these works only focus on one aspect of this software/hardware co-design framework and optimize them individually, which makes the design far from optimal. In this paper, we propose PRAP-PIM, which jointly exploits the weight sparsity and weight pattern repetition by using a weight pattern reusing aware pruning method. By relaxing the weight pattern reusing precondition, we propose a similarity-based weight pattern reusing method that can achieve a higher weight pattern reusing ratio. Experimental results show that PRAP-PIM achieves 1.64× performance improvement and 1.51× energy efficiency improvement in popular deep learning benchmarks, compared with the state-of-the-art ReRAM-based DNN accelerators. Zhaoyan Shen, Jinhao Wu, Xikun Jiang, Yuhao Zhang 0006, Lei Ju 0001, Zhiping Jia |
High Confid. Comput. | 6 |
| 2023 | ChainKV: A Semantics-Aware Key-Value Store for Ethereum SystemabstractThe Log-Structure Merged tree (LSM-tree) based key-value (KV) store has been widely adopted as the storage engine for blockchain systems, such as Ethereum, in which blockchain data are uniformly transformed into randomly distributed KV items for persistence. However, blockchain semantics are ignored during this process, making the blockchain storage suffer from heavy read/write amplification problems. Moreover, as the Ethereum network scales up, tremendous data further exacerbates its storage burden. Until now, most studies have focused on sharding, data archiving, decentralized distributed storage, etc., to mitigate the burden of the storage layer. However, the incompatibility between Ethereum semantics and the characteristics of the storage engine is ignored. In this paper, we present ChainKV, a new semantics-aware storage paradigm to improve the storage management performance for the Ethereum system. Firstly, based on Ethereum blockchain semantics, ChainKV separately stores different types of data in multiple storage zones in the KV store to mitigate the read/write amplification problem. Secondly, following the mechanism of the verification process in the authenticated data structure (ADS), a new ADS data transformer is proposed to exploit the data locality when persisting ADS. Moreover, a new space gaming caching policy is adopted to coordinate the cache space management for two independent storage zones. Finally, we propose an optional lightweight node crash recovery mechanism to eliminate functional redundancy between the Ethereum protocol and the storage engine. The experimental results indicate that ChainKV outperforms the prior Ethereum systems by up to 1.99× and 4.20× for synchronization and query operations, respectively Bingzhe Li, Xiaojun Cai, Zhiping Jia, Lei Ju 0001, Zili Shao, Zhaoyan Shen |
Proc. ACM Manag. Data | 4 |
| 2023 | A Multiagent Reinforcement Learning-Assisted Cache Cleaning Scheme for DM-SMRabstractTo support nonsequential writes, persistent cache (PC) is constructed in drive managed SMR (DM-SMR) drive. However, PC cleaning introduces drastic performance degradation and enlarges tail latencies. In this article, we propose to utilize reinforcement learning (RL) to mitigate the long-tail latency of PC cleaning. Our scheme uses the lightweight$Q$-learning method to monitor and learn the idle time of I/O workloads, based on which PC cleaning is intelligently guided, thus maximally exploit idle time between requests and hiding tail latency from normal requests. In addition, a multiagent RL scheme with clustering algorithm is adopted to further mitigate the tail latencies and adapt to variable workloads. We emulate a DM-SMR drive inside a Linux device driver to implement our proposed scheme. According to the experimental results, our scheme can effectively reduce the tail latency by 59.45% at the 99.9th percentile and the average latency by 48.75% compared with a typical shingled magnetic recording (SMR) design. Zhaoyan Shen, Yungang Pan, Yuhao Zhang 0006, Zhiping Jia, Xiaojun Cai, Bingzhe Li, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | Removing Double-Logging with Passive Data Persistence in LSM-tree based Relational Databases
Kecheng Huang, Zhaoyan Shen, Zhiping Jia, Zili Shao, Feng Chen 0005 |
FAST | 3 |
| 2022 | Re-LSM: A ReRAM-Based Processing-in-Memory Framework for LSM-Based Key-Value StoreabstractLog-structured merge (LSM) tree based key-value (KV) stores organize writes into hierarchical batches for high-speed writing. However, the notorious compaction process of LSM-tree severely hurts system performance. It not only involves huge I/O operations but also consumes tremendous computation and memory resources. In this paper, first we find that when compaction happens in the high levels (i.e., L0, L1) of the LSM-tree, it may saturate all system computation and memory resources, and eventually stall the whole system. Based on this observation, we present Re-LSM, a ReRAM-based Processing-in-Memory (PIM) framework for LSM-based Key-Value Store. Specifically, in Re-LSM, we propose to offload certain computation and memory-intensive tasks in the high levels of the LSM-tree to the ReRAM-based PIM space. A high parallel ReRAM compaction accelerator is designed by decomposing the three-phased compaction into basic logic operating units. Evaluation results based on db_bench and YCSB show that Re-LSM achieves 2.2× improvement on the throughput of random writes compared to RocksDB, and the ReRAM-based compaction accelerator speedups the CPU-based implementation by 64.3× and saves 25.5× energy. Zhaoyan Shen, Yiheng Tong, Zhiping Jia, Lei Ju 0001, Jiezhi Chen, Bingzhe Li |
ICCAD | 4 |
| 2022 | Understanding Characteristics and System Implications of DAG-Based Blockchain in IoT EnvironmentsabstractBlockchain is starting to be deployed in the Internet of Things (IoT) to enable autonomous device-to-device transactions. However, traditional block-based blockchain techniques, such as Bitcoin and Ethereum, are not suitable for IoT environments due to their low throughput, high computation overhead, and costly transaction fee. To satisfy the requirements of IoT environments, directed-acyclic-graph (DAG)-based approaches, aiming to provide cheap blockchain services with low latency and high throughput, are emerging. This article presents a set of comprehensive experimental studies on IOTA, a representative DAG-based blockchain. We aim to exhibit its unique characteristics mainly from three aspects: 1) performance; 2) security; and 3) system robustness. We have developed a series of benchmark tools and judiciously selected typical configurations to perform experimental examinations with a real private IOTA network. Our studies reveal several interesting findings: 1) the throughput of IOTA is higher than the traditional block-based blockchain but far less than the reported thousands of transactions per second (TPS) in its whitepaper, even with scaling-up configurations; 2) the database query heavily impacts the performance of IOTA, even more than its mining [i.e., Proof of Work (PoW)] process; and 3) the system robustness of IOTA is closely related to the frequency of the incoming transactions while the milestone sent by the centralized coordinator has little effect on the system robustness. We make our benchmark tools public and expect our works can inspire system architects, application designers, and practitioners with new optimization directions and potential application cases for further exploration. Tianyu Wang 0009, Qian Wang 0042, Zhaoyan Shen, Zhiping Jia, Zili Shao |
IEEE Internet Things J. | 4 |
| 2022 | PQ-PIM: A pruning-quantization joint optimization framework for ReRAM-based processing-in-memory DNN accelerator
Yuhao Zhang 0006, Xikun Jiang, Zhaoyan Shen, Zhiping Jia |
J. Syst. Archit. | 6 |
| 2022 | Prism-SSD: A Flexible Storage Interface for SSDsabstractThe rapid adoption of solid-state drives (SSDs) as a major storage component has been made possible, thanks to their ability to export a standard block I/O interface to the file system and application developers. Meanwhile, this high-level abstraction has been shown to limit the utilization of the devices and the performance of applications running on top of them. Indeed, many optimizations of performance-critical applications bypass the standard block interface and rely on low-level control over SSD internal processes. However, the need to directly manage the physical device significantly increases development complexity and cost, and reduces its portability. Thus, application developers must choose between two extreme options, eithereasy developmentoroptimal performance, without a real possibility to balance between these two objectives. To bridge this gap, we propose aflexible storage interfacethat exports the SSD hardware in three levels of abstraction: 1) as a raw flash media with its low-level details; 2) as a group of functions to manage flash capacity; or 3) as a configurable block device. This multilevel abstraction allows developers to choose the degree in which they desire to control the flash hardware in a manner that best suits the applications’ semantics and performance objectives. We demonstrate the usability of this new model withPrism-SSD—a prototype of this interface as a user-level library on the Open-Channel SSD platform. We use each of the interface’s three abstraction levels to modify the I/O module of three representative applications: 1) a key-value cache system; 2) a user-level file system; and 3) a graph processing engine. Prism-SSD improves application performance by 5%–27%, at varying development costs, between 200 and 3500 lines of code. Zhaoyan Shen, Feng Chen 0005, Gala Yadgar, Zhiping Jia, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Deep Reinforcement-Learning-Guided Backup for Energy Harvesting Powered SystemsabstractEnergy harvesting technology has been widely developed as a promising alternative of battery to power embedded systems. However, energy harvesting powered embedded systems may have potential frequent power interruptions due to unstable energy supply. Nonvolatile processors (NVPs) are proposed to survive power failures by saving volatile data to nonvolatile memory (NVM) upon power failures and resuming them after power comes back. Traditionally, backup is triggered immediately when an energy warning occurs. However, it is also possible to more aggressively utilize the residual energy for program execution to improve forward progress. In this work, we propose a deep reinforcement-learning-guided backup strategy to improve forward progress in energy harvesting powered intermittent embedded systems. The experimental results show an average of 8.3%, 51.6%, and 325.3% improved forward progress compared with$Q$-learning, the related work ALD, and traditional instant backup, respectively. Weifan Sun, Mengying Zhao, Weining Song, Xiaojun Cai, Tiantian Liu 0001, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2022 | A Practical Highly Paralleled ReRAM-Based DNN Accelerator by Reusing Weight Pattern RepetitionsabstractResistive random access memory (ReRAM)-based processing-in-memory (PIM) architecture has been designed to accelerate deep neural networks (DNNs) by concurring computation and memory barriers. To further improve memory and computation efficiency, the weight sparsity characteristic has been explored to optimize the ReRAM-based DNN accelerators. However, these designs only focus on compressing zero weights to eliminate ineffectual computation. In this article, we thoroughly analyze the weight distribution characteristics of several typical DNN models and observe many nonzero weight pattern repetitions (WPRs). Therefore, there is an opportunity to further improve the performance and energy efficiency by reusing these WPR. We propose a novel ReRAM-based accelerator—PattPIM, to achieve space compression and computation reuse by exploring DNN WPR based on practical ReRAM crossbars. In PattPIM, we propose a configurable WPR-aware DNN engine and a WPR-to-OU mapping scheme to save both space and computation resources. An intraprocessing engine (PE) pipeline is designed to improve the parallelism of the computation process. Furthermore, we adopt an approximate weight pattern transform algorithm to improve the DNN WPR ratio to enhance the reuse efficiency with negligible accuracy loss. Our evaluation with 6 DNN models shows that the proposed PattPIM delivers significant performance improvement, ReRAM resource efficiency and energy saving. Yuhao Zhang 0006, Zhiping Jia, Hongchao Du, Runzhen Xue, Zhaoyan Shen, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | A Survey of Blockchain Data Management SystemsabstractBlockchain has been widely deployed in various fields, such as finance, education, and public services. Blockchain has decentralized mechanisms with persistency and auditability and runs as an immutable distributed ledger, where transactions are jointly performed through cryptocurrency-based consensus algorithms by worldwide distributed nodes. There have been many survey papers reviewing the blockchain technologies from different perspectives, e.g., digital currencies, consensus algorithms, and smart contracts. However, none of them have focused on the blockchain data management systems. To fill in this gap, we have conducted a comprehensive survey on the data management systems, based on three typical types of blockchain, i.e., standard blockchain, hybrid blockchain, and DAG ( Directed Acyclic Graph )-based blockchain. We categorize their data management mechanisms into three layers: blockchain architecture, blockchain data structure, and blockchain storage engine, where block architecture indicates how to record transactions on a distributed ledger, blockchain data structure refers to the internal structure of each block, and blockchain storage engine specifies the storage form of data on the blockchain system. For each layer, the works advancing the state-of-the-art are discussed together with technical challenges. Furthermore, we lay out several possible future research directions for the blockchain data management systems. Bingzhe Li, Wanli Chang 0001, Zhiping Jia, Zhaoyan Shen, Zili Shao |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2021 | Reinforcement Learning-Assisted Cache Cleaning to Mitigate Long-Tail Latency in DM-SMRabstractDM-SMR adopts Persistent Cache (PC) to accommodate non-sequential write operations. However, the PC cleaning process induces severe long-tail latency. In this paper, we propose to mitigate the tail latency of PC cleaning by using Reinforcement Learning (RL). Specifically, a real-time lightweight Q-learning model is built to analyze the idle window of I/O workloads, based on which PC cleaning is judiciously scheduled, thereby maximally utilizing the I/O idle window and effectively hiding the tail latency from regular requests. We implement our technique inside a Linux device driver with an emulated SMR drive. Experimental results show that our technique can reduce the tail latency by 57.65% at 99.9th percentile and the average response time by 46.11% compared to a typical SMR design. Yungang Pan, Zhiping Jia, Zhaoyan Shen, Bingzhe Li, Wanli Chang 0001, Zili Shao |
DAC | 2 |
| 2021 | Accelerating DCNNs via Cooperative Weight/Activation Compression
Yuhao Zhang 0006, Xikun Jiang, Yudong Pan, Pusen Dong, Zhaoyan Shen, Zhiping Jia |
ICA3PP (3) | 8 |
| 2021 | Block-LSM: An Ether-aware Block-ordered LSM-tree based Key-Value Storage EngineabstractEthereum as one of the largest blockchain systems plays an important role in the distributed ledger, database systems, etc. As more and more blocks are mined, the storage burden of Ethereum is significantly increased. The current Ethereum system uniformly transforms all its data into key-value (KV) items and stores them to the underlying Log-Structure Merged tree (LSM-tree) storage engine ignoring the software semantics. Consequently, it not only exacerbates the write amplification effect of the storage engine but also hurts the performance of Ethereum. In this paper, we proposed a new Ethereum-aware storage model called Block-LSM, which significantly improves the data synchronization of the Ethereum system. Specifically, we first design a shared prefix scheme to transform Ethereum data into ordered KV pairs to alleviate the key range overlaps of different levels in the underlying LSM-tree based storage engine. Moreover, we propose to maintain several semantic-orientated memory buffers to isolate different kinds of Ethereum data. To save space overhead, Block-LSM further aggregates multiple blocks into a group and assigns the same prefix to all KV items from the same block group. Finally, we implement Block-LSM in the real Ethereum environment and conduct a series of experiments. The evaluation results show that Block-LSM significantly reduces up to 3.7× storage write amplification and increases throughput by 3× compared with the original Ethereum design. Bingzhe Li, Xiaojun Cai, Zhiping Jia, Zhaoyan Shen, Yi Wang 0003, Zili Shao |
ICCD | 4 |
| 2021 | Less is More: De-amplifying I/Os for Key-value Stores with a Log-assisted LSM-treeabstractIn recent years, Log-Structured Merge Tree (LSMtree) based key-value stores, such as LevelDB and RocksDB, have been widely adopted in data center systems. Though optimized for high-speed write processing, the severe I/O amplification remains a critical constraint that hinders them from reaching their maximum performance potential. Unfortunately, this problem is deeply rooted in the fundamental design of the LSMtree structure. A small number of frequently updated key-value items could quickly pollute the entire tree structure, causing repeated changes in the structure and quickly amplifying the amount of disk IOs across the levels in the tree. In this paper, we present a novel scheme, called Log-assisted LSM-tree (L2SM), to fundamentally address the long-existing I/O amplification problem. L2SM adopts a small-size, multi-level log structure to isolate selected key-value items that have a disruptive effect on the tree structure, accumulates and absorbs the repeated updates in a highly efficient manner, and removes obsolete and deleted key-value items at an early stage. We have prototyped the L2SM structure based on LevelDB. Our evaluation with the YCSB benchmark shows promising results by reducing the amount of disk IOs by up to 40.2%, increasing the throughput by up to 67.4%, and decreasing the average latency by up to 40.1%. Kecheng Huang, Zhiping Jia, Zhaoyan Shen, Zili Shao, Feng Chen 0005 |
ICDE | 2 |
| 2021 | DAP-Sketch: An accurate and effective network measurement sketch with Deterministic Admission Policy
Rui Wang 0075, Hongchao Du, Zhaoyan Shen, Zhiping Jia |
Comput. Networks | 4 |
| 2021 | Fast-convergent federated learning with class-weighted aggregation
Zezhong Ma, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
J. Syst. Archit. | 4 |
| 2021 | A lightweight online backup manager for energy harvesting powered nonvolatile processor systemsabstractWith the explosive growth of battery-free and energy-harvesting devices, the energy harvesting powered system has gained more attentions and been widely used in different fields. However, unstable harvested energy is a challenge of energy-harvesting devices since the program execution would be interrupted frequently. Non-volatile processor (NVP) is proposed to back up volatile logics before energy depletion and recover the system status after energy resumption . This paper takes the challenge of forward progress improvement issue in NVP system and proposes a lightweight online backup strategy which tries to aggressively use energy in capacitor after receiving energy warnings. We also build a flexible and accurate simulation tool for NVP system evaluation. The experimental results show an average of 24.2% and 13.7% improved forward progress compared with instant backup method and the most related work, respectively. Weining Song, Xiaojun Cai, Mengying Zhao, Zhaoyan Shen, Zhiping Jia |
J. Syst. Archit. | 5 |
| 2021 | An efficient highly parallelized ReRAM-based architecture for motion estimation of HEVC
Yuhao Zhang 0006, Zhiping Jia, Renhai Chen, Zhaoyan Shen |
J. Syst. Archit. | 3 |
| 2021 | Leveraging the Interplay of RAID and SSD for Lifetime Optimization of Flash-Based SSD RAIDabstractFlash-based SSD RAID arrays are increasingly being deployed in data centers. Compared with HDD arrays, SSD arrays drastically enhance I/O performance and density, and reduce power, cooling, and rack space. Nevertheless, SSDs suffer aging issues. Especially, an SSD has limited endurance and needs to be replaced when it reaches to the end of its lifetime. Although prior studies have been conducted to address this disadvantage, effective techniques of RAID/SSD controllers are urgently needed to extend the lifetime of SSD arrays. In this article, we propose a novel RAID architecture, called FreeRAID, to leverage the interplay of RAID and SSD controllers to optimize the lifespan of flash-based SSD arrays. FreeRAID adds a new exploitable phase to the life cycle of flash blocks. In FreeRAID, flash space is separated into normal space and exploitable space, and they are used to serve normal data and approximate data, respectively. We design a dual-space management scheme for RAID controllers to intelligently allocate SSD spaces based on their aging status. Inside an SSD, we propose an adaptive flash translation layer for the SSD controller to maintain the reliability and space efficiency of flash memories. We implemented a prototype of FreeRAID based on an SSD array simulator. Our experiments show that FreeRAID can significantly increase the lifetime by up to 3.07 × compared with conventional SSD-based RAID arrays. Zhaoyan Shen, Chenlin Ma, Zhiping Jia, Tao Li 0006, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Pearl: Performance-Aware Wear Leveling for Nonvolatile FPGAsabstractSince static random access memory (SRAM)-based field-programmable gate array (FPGA) has limited density and comparatively high leakage power, researchers have proposed FPGA architectures based on emerging nonvolatile memories (NVMs) to satisfy the requirements of data-intensive and low-power applications. Among all components, block random access memory (BRAM) has the severest endurance problem in FPGA. Unluckily, traditional wear leveling (TWL) strategies cannot be directly applied to nonvolatile FPGA because it may induce large performance overhead. In this article, we propose performance-aware wear leveling schemes for nonvolatile FPGA to improve its lifetime. Two strategies pertaining to coarse-grained wear leveling (C-Pearl) and fine-grained wear leveling (F-Pearl) are developed to balance inter-BRAM and intra-BRAM writes. Procedures, including static analysis, wear leveling-guided placement, and reconfiguration are discussed. A supportive circuit design is proposed, too. The evaluation shows that C-Pearl and F-Pearl can achieve 34% and 46% higher lifetime improvement and simultaneously 8% and 11% lower performance overhead than TWL. Mengying Zhao, Zhaoyan Shen, Xiaojun Cai, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | PattPIM: A Practical ReRAM-Based DNN Accelerator by Reusing Weight Pattern RepetitionsabstractWeight sparsity has been explored to achieve energy efficiency for Resistive Random-access Memory (ReRAM) based DNN accelerators. However, most existing ReRAM-based DNN accelerators are based on an overidealized crossbar architecture and mainly focus on compressing zero weights. In this paper, we propose a novel ReRAM-based accelerator — PattPIM, to achieve space compression and computation reuse by studying DNN weight patterns based on practical ReRAM crossbars. We first thoroughly analyze the weight distribution characteristics of several typical DNN models and observe many non-zero weight pattern repetitions (WPRs). Thus, in PattPIM, we propose a WPR-aware DNN engine and a WPR-to-OU mapping scheme to save both space and computation resources. Furthermore, we adopt an approximate weight pattern transform algorithm to improve the DNN WPRs ratio to enhance the reuse efficiency with negligible inference accuracy loss. Our evaluation with 6 DNN models shows that the proposed PattPIM delivers significant performance improvement, ReRAM resources efficiency and energy saving. Yuhao Zhang 0006, Zhiping Jia, Yungang Pan, Hongchao Du, Zhaoyan Shen, Mengying Zhao, Zili Shao |
DAC | 2 |
| 2020 | Q-learning Based Backup for Energy Harvesting Powered Embedded SystemsabstractNon-volatile processors (NVPs) are used in energy harvesting powered embedded systems to preserve data across interruptions. In NVP systems, volatile data are backed up to non-volatile memory upon power failures and resumed after power comes back. Traditionally, backup is triggered immediately when energy warning occurs. However, it is also possible to more aggressively utilize the residual energy for program execution to improve forward progress. In this work, we propose a Q-learning based backup strategy to achieve maximal forward progress in energy harvesting powered intermittent embedded systems. The experimental results show an average of 307.4% and 43.4% improved forward progress compared with traditional instant backup and the most related work, respectively. Yujie Zhang 0007, Weining Song, Mengying Zhao, Zhaoyan Shen, Zhiping Jia |
DATE | 6 |
| 2020 | Maximizing CNN Throughput on FPGA ClustersabstractField Programmable Gate Array (FPGA) platform has been a popular choice for deploying Convolutional Neural Networks (CNNs) as a result of its high parallelism and low energy consumption. Due to the limitation of on-chip resources on a single board, FPGA clusters become promising solutions to improve the throughput of CNNs. In this paper, we firstly put forward strategies to optimize the resource allocation intra and inter FPGA boards. Then we model the multi-board cluster problem and design algorithms based on knapsack problem and dynamic programming to calculate the optimal topology of the FPGA clusters. We also give a quantitative analysis of the inter-board data transmission bandwidth requirement. To make our design accommodate for more situations, we provide solutions for deploying fully connected layers and special convolution layers with large memory requirement. Experimental results show that typical well-known CNNs with the proposed topology of FPGA clusters could obtain a higher throughput per board than single-board solutions and other multi-board solutions. Ruihao Li 0002, Mengying Zhao, Zhaoyan Shen, Xiaojun Cai, Zhiping Jia |
FPGA | 6 |
| 2020 | Sequence-To-Subsequence Learning With Conditional Gan For Power DisaggregationabstractNon-intrusive load monitoring (a.k.a. power disaggregation) refers to identifying and extracting the consumption patterns of individual appliances from the mains which records the whole-house energy consumption. Recently, deep learning has been shown to be a promising method to solve this problem and many approaches based on it have been proposed. In this paper, we propose a sequence-to-subsequence learning method, which makes a trade-off between traditional sequence-to-sequence and sequence-to-point method, to balance the convergence difficulty in deep neural networks and the amount of computation in the inference period. We build our model based on conditional generative adversarial network that helps us avoid designing the loss function manually. In addition, we apply U-Net and Instance Normalization techniques to our model and demonstrate their effectiveness. Evaluations are performed on real-world data sets and we achieve the state-of-the-art performance. Yungang Pan, Zhaoyan Shen, Xiaojun Cai, Zhiping Jia |
ICASSP | 5 |
| 2020 | Optimizing Motion Estimation with an ReRAM-Based PIM Architecture
Zhaoyan Shen, Zhiping Jia, Xiaojun Cai |
WASA (1) | 3 |
| 2020 | A Highly Parallelized PIM-Based Accelerator for Transaction-Based Blockchain in IoT EnvironmentabstractBlockchain has gained a lot of attention from both academia and industry. However, traditional standard blockchains, such as Bitcoin and Ethereum, suffer from low throughput, high computation overhead, and large transaction fee, which is not suitable for Internet of Things (IoT) transactions. Recently, transaction-based approaches, such as Tangle structure which is based on a directed acyclic graph (DAG), have emerged to solve blockchain scalability issues for IoT environment. In transaction-based blockchain, for a transaction, namely, a node, to be attached to the Tangle, it needs to verify two other transactions. However, with the Tangle expanding, this attaching process consumes huge computational resources and energy, which severely limits the performance of the transaction-based blockchain. In this article, we present Re-Tangle, a highly parallelized processing-in-memory (PIM)-based accelerator for transaction-based blockchain. Re-Tangle is composed of a random walking module, a transaction validation module, and a PoW module, to improve the Tangle system performance. These modules transfer Tangle functions, such as fast exponentiation and modular, into ReRAM-based logic analog computation units. In the random walking module, Re-Tangle maintains an exponentiation table to reduce its design complexity and improve its computation efficiency. In the transaction validation module, Re-Tangle further proposes a highly parallel modular unit to accelerate the validation of different tags in a transaction. In the PoW module, we decompose the Curl hash function into basic logic OR, AND, SHIFT, and XOR operations, and map these logic operations to ReRAM crossbars in parallel to accelerate the working process. The experimental results show that Re-Tangle distinguishes itself from other architectures with significant performance improvement and energy saving. The throughput of Re-Tangle is about 22.4× and 2.38× higher compared with CPU and GPU, respectively, and the energy consumption of Re-Tangle is 83.5× and 5.77× less for equal workload. Qian Wang 0042, Zhiping Jia, Tianyu Wang 0009, Zhaoyan Shen, Mengying Zhao, Renhai Chen, Zili Shao |
IEEE Internet Things J. | 2 |
| 2020 | An Efficient Directory Entry Lookup Cache With Prefix-Awareness for Mobile DevicesabstractModern mobile devices, such as smartphones, maintain a directory cache (DCache) to accelerate directory and file accesses. However, the original DCache recursively walks through all the components of a path name for each directory entry (dentry) lookup operation, leading to low lookup efficiency. In this article, we first investigate intrinsic characteristics of dentry lookup operations in smartphones and make several interesting findings: 1) file path lookup operations are called frequently (up to 104times per second) by mobile applications; 2) file path lookup operations present high temporal and spatial locality; and 3) the latency of a file path lookup is linear to the depth of the path name. Based on our findings, we further propose an efficient directory entry lookup cache architecture, named dynamic skipping cache (DS-Cache), which adopts an ASCII-based hash table to simplify the path lookup complexity. In DS-Cache, a dynamic skip lookup algorithm is proposed to skip the common prefixes of different accessing paths. A rename algorithm and a delete algorithm are designed to promise the consistence of DS-Cache and the backend file system. We also design a prefix-aware cache replacement scheme to optimize the DS-Cache hit ratio. We have implemented and deployed DS-Cache on a Google Nexus 6P smartphone. The experimental results show that we can significantly reduce the latency of invoking system calls by up to 81%, and further reduce the completion time of real-world mobile applications by up to 67%. Zhaoyan Shen, Renhai Chen, Chenlin Ma, Zhiping Jia, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Scope-Aware Useful Cache Block Calculation for Cache-Related Pre-Emption Delay Analysis With Set-Associative Data CachesabstractTiming analysis of real-time systems must consider cache-related pre-emption delay (CRPD) costs when pre-emptive scheduling is used. While most previous work on CRPD analysis only considers instruction caches, the CRPD incurred on data caches is actually more significant. The state-of-the-art CRPD analysis methods are based on useful cache block (UCB) calculation. Unfortunately, as shown in this article, directly extending the existing UCB calculation techniques from instruction caches to data caches will lead to both unsoundness and significant imprecision. To solve these problems, we develop a new UCB calculation technique for data caches, which redefines the analysis unit (to address the unsoundness in the existing method) and precisely captures the dynamic cache access behavior by taking the temporal scopes of memory blocks into consideration. The experimental results show that our new technique yields substantially tighter CRPD estimations comparing with the state-of-the-art. Wei Zhang 0173, Nan Guan, Lei Ju 0001, Yue Tang 0001, Weichen Liu 0001, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | UniBuffer: Optimizing Journaling Overhead With Unified DRAM and NVM Hybrid Buffer CacheabstractJournaling techniques play an important role in addressing the reliability issue of filesystems caused by the volatile dynamic random access memory (DRAM)-based buffer cache. However, journaling techniques introduce a large number of extra storage writes, which greatly degrades system performance, especially for mobile devices. Emerging nonvolatile memory (NVM) technologies bring a new perspective of solving the write amplification issue caused by journaling. By adopting NVM as the buffer cache, the committed data can be maintained in NVM before being written back to the storage, thus, eliminating the journaling overhead. However, simply utilizing NVM as the buffer cache suffers from the limited lifetime and the relatively longer write latency of NVM. In this paper, we present a hybrid buffer cache architecture, named UniBuffer, by combing NVM with dynamic random access memory (DRAM) to reduce the journaling overhead and overcome the inherent constraints of NVM. In UniBuffer, we first propose a journaling-aware page management (JAPM) policy to smartly allocate data to DRAM and NVM pages. JAPM puts infrequently updated data in NVM to reduce the journaling overhead and frequently updated data in DRAM to improve the write performance and lifetime of the hybrid buffer cache. In addition, since data in one journaling transaction may be dispersed in NVM and DRAM simultaneously, different committing policies are required for different storage media, respectively. In order to guarantee the atomicity of the transaction execution in the hybrid cache architecture, a partial in-place commit (PIPC) journaling scheme is proposed to coordinate different committing patterns for NVM and DRAM. We have implemented the proposed techniques on Linux 3.14.52 and measured the performance with representative I/O-intensive benchmarks. The experimental results show that our scheme effectively improves the I/O performance compared with the ext4 filesystem and prolongs the lifetime of the hybrid buffer cache compared with the union of buffer cache and journaling (UBJ) scheme. Zhiyong Zhang 0006, Zhaoyan Shen, Zhiping Jia, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Applying Multiple Level Cell to Non-volatile FPGAsabstractStatic random access memory– (SRAM) based field programmable gate arrays (FPGAs) are currently facing challenges of limited capacity and high leakage power. To solve this problem, non-volatile memory (NVM) is proposed as the alternative to build non-volatile FPGAs (NVFPGAs). Even though the feasibility of NVFPGA has been confirmed, the utilization of multiple level cells (MLCs) has not been fully exploited yet. In this article, we study architecture of MLC-based NVFPGAs, and propose five cluster structures. To give detailed comparisons and extensive discussions, we conduct experiments for area, performance and leakage power evaluation. Based on explorations of the characteristics of MLC-based NVFPGAs, we further present MLC-aware timing-driven packing method to improve delay. In critical paths, our proposed method reduces the overhead of the additional delay in slow MLC cells. Experiments show that, compared to SRAM-based FPGAs, the proposed architecture with the proposed CAD flow can reduce the area, critical path delay and leakage power by 31%, 10%, and 95%, respectively. Mengying Zhao, Lei Ju 0001, Zhiping Jia, Jingtong Hu, Chun Jason Xue |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2019 | Performance-aware Wear Leveling for Block RAM in Nonvolatile FPGAsabstractField programmable gate arrays (FPGAs) have been widely adopted in both high-performance servers and embedded systems. Since static random access memory (SRAM) has limited density and comparatively high leakage power, researchers have proposed FPGA architectures based on emerging non-volatile memories (NVMs) to satisfy the requirements of data-intensive and low-power applications. Block RAM is on-chip memory of FPGAs, when it is implemented with NVM, it will face the challenge of limited endurance. Traditional wear leveling strategy cannot be directly applied to block RAM because it may induce large performance overhead. In this paper, we propose a performance-aware wear leveling scheme for block RAM in FPGAs to improve its lifetime. The placement strategy is improved by injecting wear leveling guidance. The evaluation shows that 29.75% lifetime enhancement is achieved with 16.32% performance improvement at the same time, compared with traditional wear leveling. Shuo Huai, Weining Song, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
DAC | 5 |
| 2019 | Accurate Network Flow Measurement with Deterministic Admission Policy
Hongchao Du, Rui Wang 0075, Zhaoyan Shen, Zhiping Jia |
ICA3PP (2) | 4 |
| 2019 | Re-Tangle: A ReRAM-based Processing-in-Memory Architecture for Transaction-based BlockchainabstractBlockchain has gained a lot of attentions from both academic and industry. Transaction-based approaches such like Tangle structure, which is based on a DAG (Directed Acyclic Graph), are emerging to solve the blockchain scalability issues for IoT environment. In transaction-based blockchain, for a transaction, namely a node, to be attached to the Tangle, it needs to verify two other transactions. However, with the Tangle expanding, this attaching process consumes huge computational resources and energy, which severely limits the performance of the transaction-based blockchain. In this paper, we present Re-Tangle, a novel transaction-based blockchain acceleration architecture that explores the opportunity of performing massive parallel operations with low hardware and energy cost. Re-Tangle consists of a random walking module and a transaction validation module, which transfer Tangle functions into ReRAM-based logic analog computation units. In the random walking module, Re-Tangle maintains a exponentiation translator to reduce its design complexity and improve its computation efficiency for exponentiation. In the transaction validation module, Re-Tangle further proposes a highly parallel modular unit to accelerate the validation of different tags in a transaction. The experience results show that Re-Tangle distinguishes itself from other architectures, with significant performance improvement and energy saving. The throughput of Re-Tangle is about 19.4× and 2.13× higher compared with CPU and GPU, respectively, and the energy consumption of Re-Tangle is 63.35 × and 4.92 × less. Qian Wang 0042, Tianyu Wang 0009, Zhaoyan Shen, Zhiping Jia, Mengying Zhao, Zili Shao |
ICCAD | 4 |
| 2019 | EMC: Energy-Aware Morphable Cache Design for Non-Volatile ProcessorsabstractWearable, implantable and Internet of Things devices are attracting increasing attention from both research and industry fields. Energy harvesting is a promising alternative of battery to power these embedded systems. However, the intrinsic instability of energy harvesting systems leads to potential frequent power interruptions. In traditional volatile processor, all the status will be lost at power failures and the program needs to re-start after power resumes. In order to survive the power failures and enable accumulative execution, non-volatile processor (NVP) is proposed to back up volatile information before power depletion and recover the system status after power resumes. Non-volatile memory (NVM) is typically attached for cache and main memory backup. There are researches working on optimization of the backup. However, little of them involve multiple level cell (MLC) NVM. In this work, we first discuss the benefit of applying MLC NVM for cache backup and the architecture of morphable hybrid cache, and then propose a three-stage energy-aware cache management strategy to improve the system performance and energy utilization while guaranteeing successful backups. Backup-aware cache replacement policies are also developed for backup optimization. Evaluation shows that the proposed EMC scheme can achieve 10.6 percent performance improvement and simultaneous 25.2 percent energy reduction when compared with the single level cell (SLC) based hybrid cache. Weining Song, Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Zhiping Jia |
IEEE Trans. Computers | 6 |
| 2018 | Set variation-aware shared LLC management for CPU-GPU heterogeneous architectureabstractHeterogeneous CPU-GPU multiprocessor systems-on-chip (HMPSoC) becomes a popular architecture choice for high performance embedded systems, where shared last-level cache (LLC) management becomes a critical design consideration. We observe that within a sampling period, CPU and GPU may have distinct access behaviors over various LLC sets. In this work, we propose a light-weighted and fined-grained cache management policy to cope with the CPU-GPU access behavior variation among cache sets. In particular, CPU and GPU requests are prioritized disparately in each LLC set during cache block insertion and promotion, based on the per-core utility behaviors and a per-set CPU-GPU miss counter. Experimental results show that our LLC management scheme outperforms the two state-of-the-art schemes TAP-RRIP and LSP by 12.6% and 10.01%, respectively. Zhaoying Li 0004, Lei Ju 0001, Hongjun Dai, Mengying Zhao, Zhiping Jia |
DATE | 6 |
| 2018 | H2-RAID: A Novel Hybrid RAID Architecture Towards High Reliability
Tianyu Wang 0009, Zhiyong Zhang 0006, Mengying Zhao, Zhiping Jia, Jianping Yang, Yang Wu 0003 |
ICA3PP (4) | 5 |
| 2018 | Mobility Analysis and Response for Software-Defined Internet of Things
Zhiyong Zhang 0006, Rui Wang 0075, Xiaojun Cai, Zhiping Jia |
ICA3PP (3) | 4 |
| 2018 | ETMRM: An Energy-efficient Trust Management and Routing Mechanism for SDWSNs
Rui Wang 0075, Zhiyong Zhang 0006, Zhiping Jia |
Comput. Networks | 4 |
| 2018 | NVM-Based FPGA Block RAM With Adaptive SLC-MLC ConversionabstractThe capacity of SRAM-based FPGA block RAM (BRAM) is restrained by the low density and high leakage power of the current CMOS technology. In this paper, we propose a nonvolatile memory (NVM)-based BRAM architecture which enables flexible conversions between single-level cell (SLC) and multilevel cell (MLC) states. We show that despite the high per-access latency and power consumption, MLC-based BRAM blocks reduce the routing cost between logic units and on-chip data storages, which potentially leads to a smaller critical path delay and power consumption. Therefore, we propose an NVM BRAM architecture and an EDA framework which adaptively packs data into SLC- or MLC-state BRAMs during FPGA design flow in order to achieve better system performance. This paper illustrates that a simple memory device replacement from SRAM to NVM leads to nonoptimal system performance. On the other hand, compared with operating all NVM BRAM blocks in the SLC state with better per-access latency and power consumption, the proposed hybrid SLC-MLC architecture and design flow improves the critical path delay by 18.51%, with a system power reduction of 25.83% at the same time. Moreover, compared with the traditional “fast” SRAM-based BRAM blocks under the same BRAM area constraint, our hybrid NVM BRAM architecture improves the critical path delay by 8.55% on average, with an average system power reduction of 54.34% at the same time. Lei Ju 0001, Xiaojin Sui, Shiqing Li, Mengying Zhao, Chun Jason Xue, Jingtong Hu, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2018 | Shared Last-Level Cache Management and Memory Scheduling for GPGPUs with Hybrid Main MemoryabstractMemory intensive workloads become increasingly popular on general purpose graphics processing units (GPGPUs), and impose great challenges on the GPGPU memory subsystem design. On the other hand, with the recent development of non-volatile memory (NVM) technologies, hybrid memory combining both DRAM and NVM achieves high performance, low power, and high density simultaneously, which provides a promising main memory design for GPGPUs. In this article, we explore the shared last-level cache management for GPGPUs with consideration of the underlying hybrid main memory. To improve the overall memory subsystem performance, we exploit the characteristics of both the asymmetric read/write latency of the hybrid main memory architecture, as well as the memory coalescing feature of GPGPUs. In particular, to reduce the average cost of L2 cache misses, we prioritize cache blocks from DRAM or NVM based on observations that operations to NVM part of main memory have a large impact on the system performance. Furthermore, the cache management scheme also integrates the GPU memory coalescing and cache bypassing techniques to improve the overall system performance. To minimize the impact of memory divergence behaviors among simultaneously executed groups of threads, we propose a hybrid main memory and warp aware memory scheduling mechanism for GPGPUs. Experimental results show that in the context of a hybrid main memory system, our proposed L2 cache management policy and memory scheduling mechanism improve performance by 15.69% on average for memory intensive benchmarks, whereas the maximum gain can be up to 29% and achieve an average memory subsystem energy reduction of 21.27%. Chuanqi Zang, Lei Ju 0001, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2017 | Cooperative DVFS for energy-efficient HEVC decoding on embedded CPU-GPU architectureabstractThe next generation video coding standard High Efficiency Video Coding (HEVC) provides better compression rate for high resolution videos, at the cost of substantially higher computational complexity. While some latest off-the-shelf consumer electronics support HEVC via ASIC solutions, software implementation of real-time HEVC remains an open challenge for resource-constraint embedded systems. In this work, we present an HEVC decoder design on a low-power embedded heterogeneous multiprocessor System-on-Chip (HMPSoC) with CPU and GPU. Our analysis shows that the massive parallel architecture of GPU leads to a relatively smooth fluctuation on the processing time between video frames. Moreover, the dynamic workload of each frame has a monotonic correlation with a particular coding parameter that can be obtained at decoding time. Based on these observations, we propose an application-specific userspace CPU-GPU DVFS scheme which effectively saves the energy consumption for HEVC decoding. Furthermore, given our accurate workload prediction, only a small frame buffer is required to ensure real-time video decoding. Fan Gong, Lei Ju 0001, Deshan Zhang, Mengying Zhao, Zhiping Jia |
DAC | 5 |
| 2017 | Maximizing Forward Progress with Cache-aware Backup for Self-powered Non-volatile ProcessorsabstractEnergy harvesting is replacing battery to power embedded systems such as Internet of Things and wearable devices. Unstable energy supply brings challenges to energy harvesting powered system, resulting in frequent interruptions. Non-volatile processor is proposed to back up volatile logics before energy depletion and recover the system status after energy resumes. The backup efficiency of memory content significantly affects program performance. There are existing researches focusing on backup optimizations, but they did not fully consider cache behaviors. In this paper, we introduce cache persistence analysis into memory backup for self-powered non-volatile processors. The evaluation shows that the proposed cache-aware backup delivers on average 45.6% improvement in forward progress, achieving 40.2% and 12.7% higher system performance compared with instant and cache-unaware backup. Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Zhiping Jia |
DAC | 5 |
| 2017 | Shared last-level cache management for GPGPUs with hybrid main memoryabstractMemory intensive workloads become increasingly popular on general purpose graphics processing units (GPGPUs), and impose great challenges on the GPGPU memory subsystem design. On the other hand, with the recent development of non-volatile memory (NVM) technologies, hybrid memory combining both DRAM and NVM achieves high performance, low power and high density simultaneously, which provides a promising main memory design for GPGPUs. In this work, we explore the shared last-level cache management for GPGPUs with consideration of the underlying hybrid main memory. In order to improve the overall memory subsystem performance, we exploit the characteristics of both the asymmetric read/write latency of the hybrid main memory architecture, as well as the memory coalescing feature of GPGPU. In particular, to reduce the average cost of L2 cache misses, we prioritize cache blocks from DRAM or NVM based on observation that operations to NVM part of main memory have large impact on the system performance. Furthermore, the cache management scheme also integrates the GPU memory coalescing and cache bypassing techniques to improve the overall cache hit ratio. Experimental results show that in the context of a hybrid main memory system, our proposed L2 cache management policy improves performance against the traditional LRU policy and a state-of-the-art GPU cache strategy EABP [20] by up to 27.76% and 14%, respectively. Xiaojun Cai, Lei Ju 0001, Chuanqi Zang, Mengying Zhao, Zhiping Jia |
DATE | 6 |
| 2017 | Energy-Balanced and Depth-Controlled Routing Protocol for Underwater Wireless Sensor Networks
Zhiyong Zhang 0006, Rui Wang 0075, Xiaojun Cai, Zhiping Jia |
ICA3PP | 5 |
| 2017 | ESD-WSN: An Efficient SDN-Based Wireless Sensor Network Architecture for IoT Applications
Zhiyong Zhang 0006, Rui Wang 0075, Zhiping Jia, Haijun Lei, Xiaojun Cai |
ICA3PP | 4 |
| 2017 | Design Exploration for Multiple Level Cell Based Non-Volatile FPGAsabstractStatic random access memory (SRAM) based field programmable gate arrays (FPGAs) are currently facing challenges of limited capacity and high leakage power. To solve this problem, non-volatile memory (NVM) is proposed as the alternative to build non-volatile FPGAs (NVFPGAs). Even though the feasibility of NVFPGA has been confirmed, the utilization of multiple level cells (MLC) has not been fully exploited yet. In this paper, we study architecture of MLC based NVFPGAs, and propose five cluster structures, as well as the corresponding working mode supported by MLC based clusters. To give detailed comparisons and extensive discussions, we conduct experiments for area, performance and leakage power evaluation. Experiments show that, compared to SRAM based FPGAs, the proposed architecture can reduce the area, latency and leakage power by 32.66% and 7.45%, and 96.13%, respectively. Mengying Zhao, Lei Ju 0001, Zhiping Jia, Chun Jason Xue, Jingtong Hu |
ICCD | 4 |
| 2017 | Unified nvTCAM and sTCAM architecture for improving packet matching performanceabstractSoftware-Defined Networking (SDN) allows controlling applications to install fine-grained forwarding policies in the underlying switches. Ternary Content Addressable Memory (TCAM) enables fast lookups in hardware switches with flexible wildcard rule patterns. However, the performance of packet processing is severely constrained by the capacity of TCAM, which aggravates the processing burden and latency issues. In this paper, we propose a hybrid TCAM architecture which consists of NVM-based TCAM (nvTCAM) and SRAM-based TCAM (sTCAM), utilizing nvTCAM to cache the most popular rules to improve cache-hit-ratio while relying on a very small-size sTCAM to handle cache-miss traffic to effectively decrease update latency. Considering the special rule dependency, we present an efficient Rule Migration Replacement (RMR) policy to make full utilization of both nvTCAM and sTCAM to obtain better performance. Experimental results show that the proposed architecture outperforms current TCAM architectures. Xianzhong Ding, Zhiyong Zhang 0006, Zhiping Jia, Lei Ju 0001, Mengying Zhao, Huawei Huang |
LCTES | 3 |
| 2017 | Scope-Aware Useful Cache Block Analysis for Data Cache Related Preemption DelayabstractStatic timing analysis is crucial for design of realtime systems. While the worst-case execution time of a task is typically computed or measured in a single task environment, the presence of caches imposes additional cache related preemption delay (CRPD) cost to the lower priority tasks in a preemptive multi-tasking system. In this work, we show that existing instruction CRPD analysis techniques cannot be straightforwardly extended for safe and precise data CRPD analysis. In order to capture the dynamic behavior of the data memory references, we introduce the notion of temporal scopes into the abstract cache state (ACS) to capture the data memory blocks that must or may reside in the cache during certain time intervals of program execution. Based on the improved ACS representation, we present a temporal scope aware useful cache block (UCB) calculation for safe and tight estimation of the data CRPD cost. Experimental results show that the proposed technique leads to substantially tighter CRPD estimation, and is applicable to programs with complex data reference patterns. Wei Zhang 0173, Fan Gong, Lei Ju 0001, Nan Guan, Zhiping Jia |
RTAS | 5 |
| 2017 | Energy-aware morphable cache management for self-powered non-volatile processorsabstractWearable, implantable and Internet of Things devices are attracting increasing attention from both research and industry. Energy harvesting is a promising alternative of battery to power these embedded systems. However, the intrinsic instability of energy harvesting systems leads to potential frequent power interruptions. In order to survive the power failures, non-volatile processor (NVP) is proposed to back up volatile information before power depletion and recover the system status after power resumes. Non-volatile memory (NVM) is typically attached for cache and main memory backup. There are researches working on optimization of the backup, however, little of them involve multiple level cell (MLC) NVM. In this work, we first discuss the benefit of applying MLC NVM for cache backup, and then propose a three-stage energy-aware cache management strategy to improve the system performance and energy utilization while guaranteeing successful backups. Evaluation shows that the proposed scheme can achieve 16.7% energy reduction with comparative performance with the single level cell (SLC) based hybrid cache. Mengying Zhao, Lei Ju 0001, Chun Jason Xue, Xin Li 0001, Zhiping Jia |
RTCSA | 6 |
| 2017 | Stack-Size Sensitive On-Chip Memory Backup for Self-Powered Nonvolatile ProcessorsabstractWearable devices gain increasing popularity since they can collect important information for healthcare and well-being purposes. Compared with battery, energy harvesting is a better power source for these wearable devices due to many advantages. However, harvested energy is naturally unstable and program execution will be interrupted frequently. Nonvolatile processors demonstrate promising advantages to back up volatile state before the system energy is depleted. However, it also introduces non-negligible energy and area overhead. In this paper, we aim to reduce the amount of data that need to be backed up during a power failure. Based on the observation that stack size varies along program execution, we propose to analyze the application program and identify efficient backup positions, by which the stack content to back up can be significantly reduced. The evaluation results show an average of 45.7% reduction on nonvolatile stack size for stack backup, with 0.58% storage overhead. In the mean time, with the proposed schemes, the energy utilization and program forward progress can be greatly improved compared with instant backup. Mengying Zhao, Chenchen Fu, Qing'an Li, Mimi Xie, Yongpan Liu, Jingtong Hu, Zhiping Jia, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2017 | State Asymmetry Driven State Remapping in Phase Change MemoryabstractPhase change memory (PCM) is one of the most promising candidates to replace DRAM as main memory in deep submicron regime. Regardless of single-level or multiple-level cells, the programming costs to each state exhibit significant asymmetries in latency, energy and endurance. In this paper, we exploit the potential of reducing programming costs in terms of latency, energy, and endurance for PCM through state remapping. First, quantitative programming models are constructed for cost assessments. Then, both dynamic and static remapping schemes are analyzed and compared. The observation that the efficacy of dynamic state remappings is instable motivates us to propose a static remapping technique, which outperforms previous work in cost reduction within much lower implementation overhead. The optimality of the proposed static state remapping is also proved. The evaluation results confirm the efficacy of the proposed state remapping technique in delivering a stable and promising cost reduction in latency, energy, and wear. Mengying Zhao, Jingtong Hu, Chengmo Yang, Tiantian Liu 0001, Zhiping Jia, Chun Jason Xue |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2017 | SLA-aware energy-efficient scheduling scheme for Hadoop YARN
Xiaojun Cai, Feng Li 0014, Lei Ju 0001, Zhiping Jia |
J. Supercomput. | 5 |
| 2016 | Write-back aware shared last-level cache management for hybrid main memoryabstractHybrid main memory with both DRAM and emerging non-volatile memory (NVM) becomes a promising solution for high performance and energy-efficient embedded systems. Cache plays an important role and highly affects the number of write backs to NVM and DRAM blocks. However, existing cache policies fail to fully address the significant asymmetry between NVM operations (especially writes) and DRAM operations, leading to non-optimal system designs. We propose a write-back aware last-level cache management scheme for the hybrid main memory, which improves the cache hit ratio of NVM memory blocks and minimizes write-backs to NVM. Experimental results show that our proposed framework leads to better performance and energy saving compared with the state-of-the-art cache management scheme for hybrid main memory architecture. Deshan Zhang, Lei Ju 0001, Mengying Zhao, Xiang Gao 0012, Zhiping Jia |
DAC | 5 |
| 2016 | Unified DRAM and NVM hybrid buffer cache architecture for reducing journaling overhead
Zhiyong Zhang 0006, Lei Ju 0001, Zhiping Jia |
DATE | 3 |
| 2016 | A flexible and scalable implementation of elliptic curve cryptography over GF(p) based on ASIPabstractPublic-key cryptography schemes are widely used due to their high level of security. As a very efficient one among public-key cryptosystems, elliptic curve cryptography (ECC) has been studied for years. Researchers used to improve the efficiency of ECC through point multiplication, which is the most important and complex operation of ECC. In our research, we use special families of curves and prime fields which have special properties. After that, we introduce the instruction set architecture (ISA) extension method to accelerate this algorithm (192-bit private key) and build an ECC_ASIP model with six new ECC custom instructions. Finally, the ECC_ASIP model is implemented in a field-programmable gate array (FPGA) platform. The persuasive experiments have been conducted to evaluate the performance of our new model in the aspects of the performance, the code storage space and hardware resources. Experimental results show that our processor improves 69.6% in the execution efficiency and requires only 6.2% more hardware resources. Zhiping Jia |
IPCCC | 3 |
| 2016 | Multipath Load Balancing in SDN/OSPF Hybrid Network
Xiangshan Sun, Zhiping Jia, Mengying Zhao, Zhiyong Zhang 0006 |
NPC | 2 |
| 2016 | Energy efficient task allocation for hybrid main memory architecture
Xiaojun Cai, Lei Ju 0001, Xin Li 0002, Zhiyong Zhang 0006, Zhiping Jia |
J. Syst. Archit. | 5 |
| 2016 | Data aggregation framework for energy-efficient WirelessHART networks
Feng Li 0014, Lei Ju 0001, Zhiping Jia |
J. Syst. Archit. | 3 |
| 2015 | A three-stage-write scheme with flip-bit for PCM main memoryabstractPhase-change memory (PCM) is a nonvolatile memory which suffers slow write performance and limited write endurance. Besides, writing a one to a PCM cell needs longer time but less electrical current than writing a zero. In traditional PCM schemes, zeros and ones in a word are written at the same time and word write time has to be the time to write a one, thus incurring time waste. In this paper, we propose a three-stage write scheme with flip-bit for PCM main memory to reduce the number of changed bits and write latency. In our scheme, write operation is divided into comparison, write-0 and write-1 stages. In the comparison stage, new data and old data are compared and the new data is re-encoded by a flip-bit to minimize changed bits. Then the flip-bit and re-encoded data are written to PCM cells in an accelerating manner. All zero bits and one bits are written separately in later two stages to avoid the time waste in traditional write. Our scheme shrinks time consumption and reduces bit changes caused by write operation over other existing schemes. The experimental results show that this scheme decreases 43.5% bit changes, 16.6% write time and 34.6% write energy consumption on average. Xin Li 0002, Lei Ju 0001, Zhiping Jia |
ASP-DAC | 4 |
| 2015 | Managing hybrid on-chip scratchpad and cache memories for multi-tasking embedded systemsabstractOn-chip memory management is essential in design of high performance and energy-efficient embedded systems. While many off-the-shelf embedded processors employ a hybrid on-chip SRAM architecture including both scratchpad memories (SPMs) and caches, many existing work on SPM management ignore the synergy between caches and SPMs. In this work, we propose a static SPM allocation strategy for the hybrid on-chip memory architecture in a multi-tasking environment, which minimizes the overall access latency and energy consumption of the instruction memory subsystem. We capture cache conflict misses via a fine-grained temporal cache behavior model. An integer linear programming (ILP) based formulation is proposed to generate an function-level SPM allocation scheme, where both intra- and inter-task cache interference as well as access frequency are captured for an optimal memory subsystem design. Compared with the state-of-the-art static SPM allocation strategy in a multitasking environment, experimental results show that our SPM management scheme achieves 30.51% further improvement in instruction memory subsystem performance, and up to 34.92% in terms of energy saving. Zimeng Zhou, Lei Ju 0001, Zhiping Jia, Xin Li 0002 |
ASP-DAC | 3 |
| 2015 | Reducing Journaling Overhead with Hybrid Buffer Cache
Zhiyong Zhang 0006, Lei Ju 0001, Zhiping Jia |
ICA3PP (4) | 3 |
| 2015 | Hybrid scratchpad and cache memory management for energy-efficient parallel HEVC encodingabstractThe next-generation video coding standard High Efficiency Video Coding (HEVC) provides better compression rates for high resolution videos compared with H.264, at the cost of significantly increased needs for computation power and memory bandwidth. Therefore, memory subsystem optimization is of paramount importance to support HEVC on resource and energy constrained embedded consumer electronics. In this paper, we present a hybrid on-chip memory architecture with both caches and scratchpad memories (SPMs) for parallel HEVC encoding. A run-time prediction algorithm is proposed to effectively identify the most-frequently accessed memory regions in the search window(s) for processing individual coding tree units (CTUs). Depending on their intra- and inter-core reuses, these regions are loaded into the private or shared SPMs for guaranteed on-chip memory accesses. On the other hand, a relatively small hardware-controlled cache is used for the rest of data accesses. Moreover, an adaptive power gating scheme is proposed to power off SPM sectors with expired load windows to further reduce the on-chip leakage power. Compared with the state-of-the-art solution, experimental results show that our proposed memory management framework supports high speed parallel HEVC processing with substantially smaller on-chip memory size, which achieves up to 76.23% on-chip leakage energy savings, and 33.31% energy saving for the overall memory subsystem. Lei Ju 0001, Zhiping Jia |
ICCD | 3 |
| 2015 | Dynamic malicious node detection with semi-supervised multivariate classification in cognitive wireless sensor networksabstractSummary Usually, wireless sensor networks are distributed massively with a number of nodes in an open large‐scale environment, and they are vulnerable to malicious attacks because the communications change dynamically and unpredictably. In this paper, we present a detection method based on multivariate classification to find out the malicious sensor nodes. It learns the features of a few type‐known node, classifies them with dynamical multivariate classification, and then establishes the sample space of all sensor nodes in the network activities to deduce the malicious nodes. The experiment results show that as long as the value of sensor node preferences and the number of active sensor nodes is stable, the false detection rate is stabilized below 0.5%. This proves that the algorithm can be used to the cognitive wireless sensor networks widely. Copyright © 2014 John Wiley & Sons, Ltd. Hongjun Dai, Huabo Liu, Zhiping Jia |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | A Novel OpenFlow-Based DDoS Flooding Attack Detection and Response Mechanism in Software-Defined NetworkingabstractSoftware-Defined Networking (SDN) and OpenFlow have brought a promising architecture for the future networks. However, there are still a lot of security challenges to SDN. To protect SDN from the Distributed denial-of-service (DDoS) flooding attack, this paper extends the flow entry counters and adds a mark action of OpenFlow, then proposes an entropy-based distributed attack detection model, a novel IP traceback and source filtering response mechanism in SDN with OpenFlow-based Deterministic Packet Marking. It achieves detecting the attack at the destination and filtering the malicious traffic at the source and can be easily implemented in SDN controller program, software or programmable switch, such as Open vSwitch and NetFPGA. The experimental results show that this scheme can detect the attack quickly, achieve a high detection accuracy with a low false positive rate, shield the victim from attack traffic and also avoid the attacker consuming resource and bandwidth on the intermediate links. Rui Wang 0075, Zhiyong Zhang 0006, Lei Ju 0001, Zhiping Jia |
Int. J. Inf. Secur. Priv. | 4 |
| 2015 | Temperature-Aware Data Allocation for Embedded Systems with Cache and Scratchpad MemoryabstractThe hybrid memory architecture that contains both on-chip cache and scratchpad memory (SPM) has been widely used in embedded systems. In this article, we explore this hybrid memory architecture by jointly optimizing time performance and temperature for embedded systems with loops. Our basic idea is to adaptively adjust the workload distribution between cache and SPM based on the current temperature. For a problem in which the workload can be estimated a priori, we present a nonlinear programming formulation to optimally minimize the total execution time of a loop under the constraints of SPM size and temperature. To solve a problem in which the workload is not known a priori, we propose a temperature-aware adaptive loop scheduling algorithm called TALS to dynamically allocate data to cache and SPM at runtime. The experimental results show that our algorithms can effectively achieve both performance and temperature optimization for embedded systems with cache and SPM. Zhiping Jia, Yi Wang 0003, Meng Wang 0005, Zili Shao |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2014 | An Improved Energy-Efficient Scheduling for Precedence Constrained Tasks in Multiprocessor Clusters
Xin Li 0002, Yanheng Zhao, Yibin Li 0002, Lei Ju 0001, Zhiping Jia |
ICA3PP (1) | 5 |
| 2014 | Reliable and Energy Efficient Routing Algorithm for WirelessHART
Feng Li 0014, Lei Ju 0001, Zhiping Jia, Zhaopeng Zhang |
ICA3PP (1) | 4 |
| 2014 | Energy efficient real-time task scheduling for embedded systems with hybrid main memoryabstractAvailable energy is the most critical limitation on the performance of embedded systems along with the increasing sophistication. Phase Change Memory (PCM), with high density and low idle power has recently been extensively studied as a promising alternative main memory of DRAM. In this paper, a hybrid PCM/DRAM main memory is utilized to leverage the low power of PCM and high performance of DRAM. We reconsider the real-time task scheduling problem of hybrid PCM/DRAM-based embedded systems. To maximize energy saving, two static scheduling algorithms under Rate-Monotonic(RM) and Earliest Deadline First (EDF) are proposed while guaranteeing the real-time constraints of all tasks. Since the actual execution time is much shorter than the worst-case execution time in real environment, we propose two dynamic mechanisms to optimize the energy consumption of our static solutions, so as to fully use the slack time produced by completed tasks. All the proposed algorithms minimize the number of task migrations from PCM to DRAM and ensure each task instance can be migrated at most once. Experimental results show our real-time scheduling algorithms reduce 25.7% to 47.2% of energy consumption on average. Zhiyong Zhang 0006, Lei Ju 0001, Zhiping Jia |
RTCSA | 4 |
| 2014 | On-demand gateway broadcast scheme for connecting mobile ad hoc networks to the InternetabstractGateway discovery algorithm is a fundamental protocol for interconnecting mobile ad hoc network (MANET) with the Internet. In most existing schemes, each gateway node broadcasts gateway advertisements to announce its presence. The decision of when to emit advertisements can influence the performance of the network. Traditional gateway discovery schemes adopt the method of periodically emitting advertisements with a time interval. However, this method does not fulfill the actual needs of the source nodes. This paper proposes a novel adaptive scheme for gateway discovery, in which the gateway broadcasts advertisements only on-demand instead of periodic emission. In order to obtain the network's actual demands for gateway advertisement, routes to the gateway are monitored. In particular, if any route is predicted to be broken, the source node requires fresh gateway advertisements to update routes, and then the gateway will be triggered to broadcast to fulfill such demands. We study the performance of on-demand gateway discovery scheme by a comparison approach. The results show that the proposed adaptive gateway discovery scheme greatly outperforms the conventional solutions: it is capable of achieving higher packet delivery ratio and lower end-to-end delay, while minimizing the routing overhead. Huaqiang Xu, Lei Ju 0001, Chongxian Guo, Zhiping Jia |
SMARTCOMP | 4 |
| 2014 | A High-Performance Distributed Certificate Revocation Scheme for Mobile Ad Hoc NetworksabstractMobile ad hoc networks (MANETs) are wireless networks which have a wide range applications due to their dynamic topologies and easy to deployment. However, such networks are also more vulnerable to attacks compared with traditional wireless networks. Certificate revocation is an effective mechanism for providing network security services. Existing schemes are not well suited for MANETs because of incurring much overhead or bring low accuracy on certificate revocation. Therefore, we propose a high-performance distributed certificate revocation scheme in which certificates of malicious nodes will be revoked quickly and accurately. Certificate revocation is the result of the collaborative effect of multiple accusations. For diluting damages to networks, one accusation is enough to limit the accusation function of the accused node. To enhance the accuracy of certificate revocation, our scheme requires nodes just accepting those accusations in which trust levels of accuser nodes are not less than accused nodes'. To guarantee the rapidity, we restore accusation functions of the falsely accused nodes after revoking certificates of all malicious nodes who ever accused them. Moreover, we design one mechanism to reward nodes who ever accused those malicious nodes, and in return, accusations made by them will accelerate the certificate revocation processes of other malicious nodes. Simulation results demonstrate the effectiveness and efficiency of our scheme in certificate revocation. In addition, our scheme achieves a great improvement of just limiting accusation functions of malicious nodes. Chongxian Guo, Huaqiang Xu, Lei Ju 0001, Zhiping Jia, Jihai Xu |
TrustCom | 4 |
| 2014 | High Performance FPGA Implementation of Elliptic Curve Cryptography over Binary FieldsabstractIn this paper, we propose a high performance hardware implementation architecture of elliptic curve scalar multiplication over binary fields. The proposed architecture is based on the Montgomery ladder method and uses polynomial basis for finite field (FF) arithmetic. A single Karatsuba multiplier runs with no idle cycle significantly increases the performance of FF multiplication while spending small amount of hardware resources, and other FF operations performed in parallel with the FF multiplier. The optimized circuits lead to a lesser area requirement compared to other high performance implementations. An implementation for the National Institute of Standards and Technology (NIST) recommended curve with degree 163 is shown, the proposed design can reach 121 MHz with 10,417 slices when implemented on Xilinx Virtex-4 XC4VLX200 FPGA device, the total time required for one elliptic curve scalar multiplication is 9.0 μs. Lei Ju 0001, Xiaojun Cai, Zhiping Jia, Zhiyong Zhang 0006 |
TrustCom | 4 |
| 2014 | Research of trust model based on fuzzy theory in mobile ad hoc networksabstractThe performance of ad hoc networks depends on the cooperative and trust nature of the distributed nodes. To enhance security in ad hoc networks, it is important to evaluate the trustworthiness of other nodes without central authorities. An information‐theoretic framework is presented, to quantitatively measure trust and build a novel trust model (FAPtrust) with multiple trust decision factors. These decision factors are incorporated to reflect trust relationship's complexity and uncertainty in various angles. The weight of these factors is set up using fuzzy analytic hierarchy process theory based on entropy weight method, which makes the model has a better rationality. Moreover, the fuzzy logic rules prediction mechanism is adopted to update a node's trust for future decision‐making. As an application of this model, a novel reactive trust‐based multicast routing protocol is proposed. This new trusted protocol provides a flexible and feasible approach in routing decision‐making, taking into account both the trust constraint and the malicious node detection in multi‐agent systems. Comprehensive experiments have been conducted to evaluate the efficiency of trust model and multicast trust enhancement in the improvement of network interaction quality, trust dynamic adaptability, malicious node identification, attack resistance and enhancements of system's security. Hui Xia 0001, Zhiping Jia, Edwin H.-M. Sha |
IET Inf. Secur. | 2 |
| 2014 | Applying link stability estimation mechanism to multicast routing in MANETs
Hui Xia 0001, Shoujun Xia, Jia Yu 0003, Zhiping Jia, Edwin H.-M. Sha |
J. Syst. Archit. | 4 |
| 2014 | Loop scheduling with memory access reduction subject to register constraints for DSP applicationsabstractSUMMARY Memory accesses introduce big‐time overhead and power consumption because of the performance gap between processors and main memory. This paper describes and evaluates a technique, loop scheduling with memory access reduction (LSMAR), that replaces hidden redundant load operations with register operations in loop kernels and performs partial scheduling for newly generated register operations subject to register constraints. By exploiting data dependence of memory access operations, the LSMAR technique can effectively reduce the number of memory accesses of loop kernels, thereby improving timing performance. The technique has been implemented into the Trimaran compiler and evaluated using a set of benchmarks from DSPstone and MiBench on the cycle‐accurate simulator of the Trimaran infrastructure. The experimental results show that when the LSMAR technique is applied, the number of memory accesses can be reduced by 18.47% on average over the benchmarks when it is not applied. The measurements also indicate that the optimizations only lead to an average 1.41% increase in code size. With such small code size expansion, the technique is more suitable for embedded systems compared with prior work.Copyright © 2013 John Wiley & Sons, Ltd. Yi Wang 0003, Zhiping Jia, Renhai Chen, Meng Wang 0005, Duo Liu 0002, Zili Shao |
Softw. Pract. Exp. | 2 |
| 2014 | A piecewise geometry method for optimizing the motion planning of data mule in tele-health wireless sensor networks
Hongjun Dai, Zhiping Jia, Meikang Qiu, Bin Wang 0002 |
Wirel. Networks | 3 |
| 2013 | Binarization-Based Human Detection for Compact FPGA Implementation
Shuai Xie, Yibin Li 0002, Zhiping Jia, Lei Ju 0001 |
APPT | 3 |
| 2013 | Binarization based implementation for real-time human detectionabstractHardware implementation of human detection is a challenging task for embedded designs. This paper presents a real-time image-based field-programmable gate array (FPGA) implementation of human detection. Our implementation is based on the histograms of oriented gradients (HOG) feature and linear support vector machine (SVM) classifier. The novelty of this work is that we replace normalization process of HOG with a modified binarization process. Therefore, during classification process with SVM classifier, all multiplication operations are replaced by addition operations. All these modifications result in reduction of hardware resource. Experimental evaluation reveals that 293 fps can be achieved on a low-end Xilinx Spartan-3e FPGA. Moreover, a detection accuracy of 1.97% miss rate and 1% false positive rate is achieved. For further demonstration, a prototype system is developed with an OV7670 camera device. Restricted to the speed of camera, a detection rate of 30 fps is achieved. Shuai Xie, Yibin Li 0002, Zhiping Jia, Lei Ju 0001 |
FPL | 3 |
| 2013 | Trust prediction and trust-based source routing in mobile ad hoc networks
Hui Xia 0001, Zhiping Jia, Xin Li 0002, Lei Ju 0001, Edwin H.-M. Sha |
Ad Hoc Networks | 2 |
| 2013 | Impact of trust model on on-demand multi-path routing in mobile ad hoc networks
Hui Xia 0001, Zhiping Jia, Lei Ju 0001, Xin Li 0002, Edwin H.-M. Sha |
Comput. Commun. | 2 |
| 2012 | A Multivariate Classification Algorithm for Malicious Node Detection in Large-Scale WSNsabstractWSN is a distributed network exposed to an open environment, which is vulnerable to malicious nodes. To find out malicious nodes among a WSN with mass sensor nodes, this paper presents a malicious detection method based on multi-variate classification. Given the types of a few sensor nodes, it extracts sensor nodes' preferences related with the known types of malicious node, establishes the sample space of all sensor nodes that participate in network activities. Then, according to the study on the type-known sensor nodes' samples based on the multivariate classification algorithm, a classifier is generated, and all of the unknown-type sensor nodes are classified. The experiment results show that as long as the value of sensor nodes preferences and the number of active sensor nodes is stable, the false detection rate is stabilized under 0.5%. Hongjun Dai, Huabo Liu, Zhiping Jia, Tianzhou Chen |
TrustCom | 3 |
| 2012 | The Research and Application of a Specific Instruction Processor for SMS4abstractSMS4 is a block cipher used in the Chinese National Standard for Wired Authentication and Privacy Infrastructure (WAPI). Like other cryptosystems, the SMS4 contains data-intensive computation, whose throughput makes a substantial contribution to the overall system performance. Design and implementation of the SMS4 cryptosystem to meet the real-time requirements of applications are challenge problems, especially for embedded network systems with limited computation resources. In this work, we present a systematic design approach of application-specific instruction-set processor (ASIP) for the SMS4 cryptographic algorithm, which exploits and compromises between the flexibility of software execution and the performance of the application-specific integrated circuit (ASIC) based implementation of the SMS4 cryptosystem. We identify and perform a design space exploration of the custom instructions found for the SMS4 algorithm, and extend the instruction set architecture (ISA) of a standard 32-bit RISC processor to accommodate them. We employ the Electronic System Level (ESL) methodology in the development of the proposed ASIP using the Xilinx Virtex5 LX110T FPGA platform. Results show that compared to the original RISC ISA, our ASIP for SMS4 achieves 2.93 times performance improvement and 52.4% less program memory utilization, with only 28.2% more resource required. Zhenzhou Li, Feng Li 0014, Zhiping Jia, Lei Ju 0001, Renhai Chen |
TrustCom | 3 |
| 2012 | Prediction-Based Algorithm for Event Detection in Wireless Sensor NetworksabstractEvent Detection in a special environment is an important application in the wireless sensor networks (WSNs). It's a hard work due to the sensor nodes are error-prone, and the energy is limited. The existing algorithms work when the events happen, which also assume that those events are spatially correlated. In this paper, we propose a prediction model. As an application of this model, we propose a novel prediction algorithm for the WSNs. We can get an accurate trend measurement for a sensor based on this new algorithm when the event is happening or just happened. In our algorithm, each node transmits its own status which is used to detect and correct the fault nodes to every neighbor. Our experimental results show that the algorithm can detect above 98% of the fault nodes and 92% area of the event region when 20% nodes are fault. Moreover, our algorithm is more energy efficient since it can get rid of frequent measurements exchange. Yongji Yu, Zhiping Jia, Ruihua Zhang |
TrustCom | 2 |
| 2012 | Link Stability Evaluation and Stability Based Multicast Routing Protocol in Mobile Ad Hoc NetworksabstractMobile ad hoc networks are more flexible than tradition networks since they do not require fixed infrastructure and allow all nodes move in a random trajectory, which leads frequent rerouting and degrades network performance. A better routing protocol can provide a more stable route to adapt to dynamic topology. In this paper, we propose a novel stability evaluation metric which is based on the relative mobility of each node pairs via sampling the received power levels. The metric is subsequently used as the routing standard in our proposed stability-based multicast routing protocol, termed as SMR, which can create bi-directional shared multicast trees composed of stable paths to decrease the links disconnections and improve network performance. Simulation results show the advantages of the SMR over other approaches in terms of the packet delivery ratio, routing packet overhead and multicast route lifetime. Zhiyong Zhang 0006, Zhiping Jia |
TrustCom | 2 |
| 2012 | Node trust evaluation in mobile ad hoc networks based on multi-dimensional fuzzy and Markov SCGM(1, 1) model
Feng Zhang 0002, Zhiping Jia, Hui Xia 0001, Xin Li 0002, Edwin H.-M. Sha |
Comput. Commun. | 2 |
| 2011 | Multicast Trusted Routing with QoS Multi-constraints in Wireless Ad Hoc NetworksabstractThe wireless ad hoc networks has been attracting increasing attention of researchers owing to its good performance and special applications. Multicast routing with QoS multi-constraints problem in this network has drawn wide spread attention from researchers who have been using different methods to solve it. One solution to this problem is to merge paths in accordance with a swarm intelligence algorithm after finding paths from source node to destination nodes in order to obtain a multicast tree that satisfies QoS multi-constraints. However, the security of routing paths is not guaranteed. To overcome this shortcoming, in this paper, we introduce the concept of trust into multicast routing problem which is made as another QoS constraint, and finally we propose a multicast trusted routing algorithm with QoS multi-constraints based on a modified ant colony algorithm. In this new algorithm, ants move constantly on the network to find an optimal constrained multicast security tree. Simulation results indicate that our algorithm can quickly find the feasible solution to solve constrained multicast security routing issues. Compared to CSTMAN, the packet delivery ratio of new algorithm is improved. Hui Xia 0001, Zhiping Jia, Lei Ju 0001, Youqin Zhu |
TrustCom | 2 |
| 2011 | Verification-Based Multi-backup Firmware Architecture, an Assurance of Trusted Boot Process for the Embedded SystemsabstractNAND flash has been widely used as the only non-volatile storage device in the embedded systems. However, it has high rates of bad block, which may lead the stored programs damaged. Especially for the firmware including bootloader and OS, this will lead the system crash immediately. This paper proposes a novel verification-based multi-backup firmware architecture (VMFA) to improve the reliability with the multiple copies of firmware in NAND flash. According to the theory of chain of trust, during the boot process, the integrity of one program should be checked before it gets the right to execute, and the program can be executed only on condition that its integrity is valid. Meanwhile, the system can automatically load and measure the backup copies and verify the integrity when the original program is damaged. Some experiments are taken on a real development platform and the VMFA is measured with time module to analyze the boot time. The results show that the system can work well with VMFA and the boot process can be ensured with the suitable verifications. Hongfei Yin, Hongjun Dai, Zhiping Jia |
TrustCom | 3 |
| 2010 | Node Trust Assessment in Mobile Ad Hoc Networks Based on Multi-dimensional Fuzzy Decision MakingabstractDue to the nature of distribution and self-organization, Mobile ad hoc networks rely on cooperation between nodes to transfer information. Therefore, one of the key factors to ensure high communication quality is an efficient assessment scheme for risks and trust of choosing next cooperative potential nodes. Trust model, an abstract psychological cognitive process, is one of the most complex concepts in social relationships, involving factors such as assumptions, expectations and behaviors. All above makes it very difficult to quantify and forecast trust accurately. In this paper, based on the theories of fuzzy recognition, we present a pattern of multi-dimensional fuzzy decision making with feedback. The analysis and experimental computation shows that this scheme is efficient in risk assessment of Ad hoc networks. Feng Zhang 0002, Zhiping Jia, Xin Li 0002, Hui Xia 0001 |
EUC | 2 |
| 2010 | Trust-based on-demand multipath routing in mobile ad hoc networksabstractA mobile ad hoc network (MANET) is a self-organised system comprised of mobile wireless nodes. All nodes act as both communicators and routers. Owing to multi-hop routing and absence of centralised administration in open environment, MANETs are vulnerable to attacks by malicious nodes. In order to decrease the hazards from malicious nodes, the authors incorporate the concept of trust to MANETs and build a simple trust model to evaluate neighbours’ behaviours – forwarding packets. Extended from the ad hoc on-demand distance vector (AODV) routing protocol and the ad hoc on-demand multipath distance vector (AOMDV) routing protocol, a trust-based reactive multipath routing protocol, ad hoc on-demand trusted-path distance vector (AOTDV), is proposed for MANETs. This protocol is able to discover multiple loop-free paths as candidates in one route discovery. These paths are evaluated by two aspects: hop counts and trust values. This two-dimensional evaluation provides a flexible and feasible approach to choose the shortest path from the candidates that meet the requirements of data packets for dependability or trust. Furthermore, the authors give a routing example in details to describe the procedures of route discovery and the differences among AODV, AOMDV and AOTDV. Several experiments have been conducted to compare these protocols and the results show that AOTDV improves packet delivery ratio and mitigates the impairment from black hole, grey hole and modification attacks. Xin Li 0002, Zhiping Jia, Peng Zhang 0008, Ruihua Zhang |
IET Inf. Secur. | 2 |
| 2009 | QoS-Aware Scheduling for Mixed Real-Time Queries over Data StreamsabstractData Stream Management Systems (DSMSs) usually need to satisfy multiple QoS requirements of applications, including timing and precision constraints. This paper focuses on the problem of real-time query model and scheduling in a DSMS to meet multiple performance objectives. At first, a mixed real-time query model is introduced which is composed of periodic, con-tinuous and one-time queries with deadlines. Moreover, an adaptive QoS-aware scheduling strategy, termed FC-TBS, is proposed to schedule these queries using feedback control mechanism. The objective of the FC-TBS strategy is to guaran-tee the deadlines of periodic queries and minimize the number of deadline violations for aperiodic queries. Besides the system tries to improve the overall query quality by adaptively adjusting CPU utilization factor for aperiodic queries according to work-load characteristics and application-defined relationship be-tween sample ratio and QoS. Experimental results show that the FC-TBS strategy is more effective than other mixed scheduling algorithms and can deal with workload fluctuations gracefully. Xin Li 0002, Zhiping Jia |
RTCSA | 2 |
| 2008 | Address assignment sensitive variable partitioning and scheduling for DSPS with multiple memory banksabstractMultiple memory banks design is employed in many high performance DSP processors. This architectural feature supports higher memory bandwidth by allowing multiple data memory access to be executed in parallel. Dedicated address generation units (AGUs) are commonly presented in DSPs to perform address arithmetic in parallel to the main datapath. Address assignment, optimization of memory layout of program variables to reduce address arithmetic instruction, has been studied extensively on single memory architecture. Make effective use of AGUs on multiple memory banks is a great challenge to compiler design and has not been studied previously. In this paper, we exploit address assignment with variable partitioning for scheduling on DSP architectures with multiple memory banks and AGUs. Our approach is built on novel graph models which capture both parallelism and serialism demands. An efficient scheduling algorithm, Address Assignment Sensitive Variable Partitioning (AASVP), is proposed to best leverage both multiple memory banks and AGUs. Experimental results show significant improvement compare to existing methods. Chun Jason Xue, Tiantian Liu 0001, Zili Shao, Jingtong Hu, Zhiping Jia, Weijia Jia 0001, Edwin H.-M. Sha |
ICASSP | 5 |