VLDB 2026 Research / reviewers in the wild / expert
Xiaojun Cai
dblp:157/2980
· DBLP profile ↗
31ranked-venue papers
3as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 3 first-author · 16 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FedMQ+: Towards efficient heterogeneous federated learning with multi-grained quantization
Mei Cao, Yuan Yuan 0040, Jianbo Lu 0001, Xiaojun Cai, Dongxiao Yu, Mengying Zhao |
J. Syst. Archit. | 5 |
| 2024 | Iterative Glycan Structure Generation by Component Localization and Classification from Mass SpectraabstractAnalyzing intact glycopeptides via mass spectrometry data elucidates the glycan composition and attachment sites, which are pivotal for investigating protein folding, stability, and function. GlycanFinder, an advanced technique, builds the glycan structure starting from the peptide (i.e., the root) and sequentially adds monosaccharide components (i.e., the leaves). However, its major limitation is the dependence on predefined rules with strong assumptions for leaf addition locations. To address this issue, we introduce MS2Glycan, a deep learning method for iteratively generating novel glycan structures from mass spectrometry data. It models component generation similarly to an object detection task, incorporating both localization and classification sub-tasks, and constructs the glycan structure component by component. MS2Glycan offers a key improvement over GlycanFinder by using a data-driven approach to predict component locations, rather than relying on predefined rules. By analyzing intact N-linked glycopeptides collected from five different mouse tissues, we assessed MS2Glycan through five-fold cross-validation, observing a notable enhancement in structural accuracy, which rose from 32% to 74%. Additionally, our findings indicate that our model more effectively mimics the principles of glycan structure construction and offers valuable insights into fully leveraging mass spectrometry data. The Python implementation can be accessed via the following link: https://github.com/xfcui/MS2Glycan. Defeng Li, Zizheng Nie, Yanmin Liu, Xiaojun Cai, Xuefeng Cui |
BIBM | 4 |
| 2024 | Decentralized Federated Learning in Partially Connected Networks with Non-IID DataabstractFederated learning is a promising paradigm to enable joint model training across distributed data while preserving data privacy. The distributed data are usually not identically and in-dependently distributed (Non-IID), which brings great challenges for federated learning. There have been existing work proposing to guide model aggregation between similar clients to deal with Non-IID data. But they typically assume a fully connected network topology, while new design issues need to be considered when it comes to a partially connected topology. In this work, we propose a probability-driven gossip framework for partially connected network topology with Non-IID data. The main idea is to discover similarity relationship between non-adjacent clients and guide the model exchange to encourage aggregation between similar clients. We explore cross-node similarity assessment and define probability to guide the model exchange and aggregation. Both similarity and communication cost are considered in the probability-driven gossip. Evaluation shows that the proposed scheme can achieve 13.04%-14.24% improvement in model accuracy, when compared with related work. Xiaojun Cai, Nanxiang Yu, Mengying Zhao, Mei Cao, Jianbo Lu 0001 |
DATE | 1 |
| 2024 | Towards Efficient Reconfiguration through Lightweight Input Inversion for MLC NVFPGAsabstractNonvolatile field programmable gate arrays (NVFP-GAs) have been proposed to address the challenges raised by artificial intelligence and big data related applications, since nonvolatile memories (NVMs) introduce advantages of high storage density, low leakage power, and high system robustness. In addition, multi-level cell (MLC), which can store multiple bits within one memory cell, further improves the logic density of NVFPGAs. However, the inefficient write operation of MLC NVM significantly increases the reconfiguration cost in aspects of energy, latency, and lifetime. In this paper, we focus on the reconfiguration cost of MLC LUTs in NVFPGA and propose a lightweight input inversion based scheme to reduce the reconfiguration cost. Inversion flexibility is defined and modeled for LUT inputs to guide the proposed scheme. We also discuss how the proposed scheme can be combined with other existing write reduction strategies. Evaluation shows the proposed scheme can reduce reconfiguration cost by 10.01 % with negligible overhead. Huichuan Zheng, Mengying Zhao, Yuqing Xiong, Xiaojun Cai, Zhiping Jia |
DATE | 5 |
| 2024 | ISVDA: An in Storage Processing Accelerator for Visual Data AnalysisabstractDeep neural networks(DNNs) have been widely applied in visual data analysis. To accelerate DNN-based visual data inference, contemporary visual data analytics systems employ strategies like model specialization and model cascading to alleviate the computational demands of DNN processes. In addition, some existing systems seek to reduce the overhead of data decoding by directly storing decoded RGB data on storage devices, albeit at the expense of significantly increased storage costs and data transfer volumes. In this paper, we introduce ISVDA, a visual data analysis accelerator based on In-storage computing. ISVDA proposes to offload the data-intensive video preprocessing tasks, encompassing data decoding and resizing, to the computational storage device, leveraging its ample internal bandwidth and computational advantages. Concurrently, ISVDA introduces a dynamic sampling algorithm that intelligently adjusts sampling strides based on spatiotemporal information encoded in motion vectors, thereby skipping more video frames for DNN inferences in both retrospective and streaming analysis scenarios. Experiments conducted on real hardware demonstrate that compared to CPU-based preprocessing systems, ISVDA can achieve average 6.8x speedup for image data analysis and 2.8x speedup for video data analysis. Zhaoyan Shen, Xiaojun Cai |
ICCD | 3 |
| 2024 | Branch Predictor Design for Energy Harvesting Powered Nonvolatile ProcessorsabstractNon-volatile processors are proposed for ambient energy harvesting systems to enable accumulative computing across power failures. They employ nonvolatile memory for processor status backup before power outage and resume the system after power recovers. A straightforward backup policy is to back up all volatile data in processors, but it induces high backup cost. In this paper, we focus on branch predictor, an important component in processor, and propose efficient backup schemes to reduce backup cost while maintaining its prediction ability. We first analyze the modules in both traditional and artificial intelligence (AI) assisted designs of branch predictor, and accordingly propose three backup mechanisms pertaining to saturation-driven, locality-driven and maturity-driven backup. On the basis of these mechanisms, adaptive backup branch predictors are designed. Evaluation shows that, with traditional Tournament architecture, the proposed design achieves 15.9% and 54.1% energy reduction when compared with no-backup and all-backup strategy. For AI assisted branch predictor, the proposed design achieves 27.5% and 82.2% energy saving. Mengying Zhao, Lihao Dong, Chun Jason Xue, Dongxiao Yu, Xiaojun Cai, Zhiping Jia |
IEEE Trans. Computers | 6 |
| 2024 | A Semantic-Integrated LSM-Tree-Based Key-Value Storage Engine for Blockchain SystemsabstractBlockchain systems play an important role in distributed ledgers, database systems, etc. As more and more blocks are mined, the storage burden of blockchain system is significantly increased. The current blockchain system uniformly transforms all its data into key-value (KV) items and stores them to the underlying Log-Structure Merged tree (LSM-tree) storage engine ignoring the software semantics. Consequently, it not only aggravates the write amplification effect of the storage engine, but also increases the redundancy of data query steps, resulting in the performance bottleneck of blockchain system. In this paper, we propose a semantic-integrated LSM-tree based Key-Value storage engine for blockchain systems, called Block-LSM, which significantly improves the data synchronization and data query efficiency of blockchain system. Specifically, we first design a shared prefix scheme to transform blockchain data into ordered KV pairs to alleviate the key range overlaps of different levels in the underlying LSM-tree based storage engine. Moreover, we propose to maintain several semantic-orientated memory buffers to isolate different kinds of blockchain data, and implement memory buffer space management strategy to further improve memory efficiency. To save space overhead, Block-LSM further aggregates multiple blocks into a group and assigns the same prefix to all KV items from the same block group. We also reduce step redundancy in transaction queries by modifying the body data storage format. Finally, we implement Block-LSM in a real blockchain environment and conduct a series of comparative experiments with the typical blockchain system Ethereum. The evaluation results show that Block-LSM significantly reduces up to 7.56× storage write amplification and increases throughput by 8.64× compared with the original Ethereum design. In terms of data lookups (i.e. transaction and account lookup), Block-LSM improves the throughput by 50% compared to the original Ethereum design. Yuhao Zhang 0006, Xiaojun Cai, Zhiping Jia, Zhaoyan Shen, Yi Wang 0003, Zili Shao, Bingzhe Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Learned Fingerprint Embedding for Large-Scale Peptide Mass Spectra RetrievalabstractTandem mass spectrometry (MS/MS) is a widely used technique for protein identification, post-translational modifications, immunotherapy, and other applications. As the amount of MS/MS spectra data increases, new computational methods are needed to efficiently search through these databases. This study introduces MS2VEC, a novel fingerprint embedding model designed to facilitate large-scale retrieval of peptide mass spectra. MS2VEC captures the relationships between distant peaks and incorporates position-aware fingerprint features from all peaks. To do this, dilated convolutions are used to capture remote relationships, and a novel position-aware multi-head attention pooling mechanism is used to abstract fingerprint features. The results demonstrate that MS2VEC achieves a top-1 retrieval accuracy of 0.810, outperforming existing methods by 5.1%. Interestingly, the precursor charge is not essential for the retrieval task, as the spectra itself contains enough information to accurately predict the charge. Additionally, the results suggest that weight-balanced fragment ions and water losses are important contributors to fingerprint features. Yongshuai Wang, Xiaojun Cai, Defeng Li, Shiwei Sun, Xuefeng Cui |
BIBM | 2 |
| 2023 | Correlation-guided Placement for Nonvolatile FPGAsabstractNonvolatile FPGAs have advantages of high density and near-zero leakage power compared with traditional SRAM-based FPGAs. However, they have lifetime issue. To deal with this problem, a series of configuration files can be generated with various logical-to-physical mappings so that intensive writes can be distributed to different physical regions for wear leveling. Currently, the configuration files are independently generated, which is time-consuming. In this paper, we propose to investigate correlations between components and use them to guide the computer-aided design (CAD) flow to speed up the procedure of deriving configuration files. Specifically, we develop dynamic probabilities to drive the swapping of placement step in the CAD flow to push components to locate appropriate positions quickly. Evaluation shows that the proposed schemes can deliver 36.32% reduction in number of swappings when compared with existing strategies, while maintaining comparable performance and lifetime. Mengying Zhao, Fanjin Xu, Huichuan Zheng, Yuqing Xiong, Zhiping Jia, Xiaojun Cai |
DAC | 7 |
| 2023 | MDCF: Multiple Dynamic Cuckoo Filters for LSM-Tree
Xingfei Yao, Taotao Xie, Zhaoyan Shen, Xiaojun Cai |
ICA3PP (6) | 5 |
| 2023 | ChainKV: A Semantics-Aware Key-Value Store for Ethereum SystemabstractThe Log-Structure Merged tree (LSM-tree) based key-value (KV) store has been widely adopted as the storage engine for blockchain systems, such as Ethereum, in which blockchain data are uniformly transformed into randomly distributed KV items for persistence. However, blockchain semantics are ignored during this process, making the blockchain storage suffer from heavy read/write amplification problems. Moreover, as the Ethereum network scales up, tremendous data further exacerbates its storage burden. Until now, most studies have focused on sharding, data archiving, decentralized distributed storage, etc., to mitigate the burden of the storage layer. However, the incompatibility between Ethereum semantics and the characteristics of the storage engine is ignored. In this paper, we present ChainKV, a new semantics-aware storage paradigm to improve the storage management performance for the Ethereum system. Firstly, based on Ethereum blockchain semantics, ChainKV separately stores different types of data in multiple storage zones in the KV store to mitigate the read/write amplification problem. Secondly, following the mechanism of the verification process in the authenticated data structure (ADS), a new ADS data transformer is proposed to exploit the data locality when persisting ADS. Moreover, a new space gaming caching policy is adopted to coordinate the cache space management for two independent storage zones. Finally, we propose an optional lightweight node crash recovery mechanism to eliminate functional redundancy between the Ethereum protocol and the storage engine. The experimental results indicate that ChainKV outperforms the prior Ethereum systems by up to 1.99× and 4.20× for synchronization and query operations, respectively Bingzhe Li, Xiaojun Cai, Zhiping Jia, Lei Ju 0001, Zili Shao, Zhaoyan Shen |
Proc. ACM Manag. Data | 3 |
| 2023 | A Multiagent Reinforcement Learning-Assisted Cache Cleaning Scheme for DM-SMRabstractTo support nonsequential writes, persistent cache (PC) is constructed in drive managed SMR (DM-SMR) drive. However, PC cleaning introduces drastic performance degradation and enlarges tail latencies. In this article, we propose to utilize reinforcement learning (RL) to mitigate the long-tail latency of PC cleaning. Our scheme uses the lightweight$Q$-learning method to monitor and learn the idle time of I/O workloads, based on which PC cleaning is intelligently guided, thus maximally exploit idle time between requests and hiding tail latency from normal requests. In addition, a multiagent RL scheme with clustering algorithm is adopted to further mitigate the tail latencies and adapt to variable workloads. We emulate a DM-SMR drive inside a Linux device driver to implement our proposed scheme. According to the experimental results, our scheme can effectively reduce the tail latency by 59.45% at the 99.9th percentile and the average latency by 48.75% compared with a typical shingled magnetic recording (SMR) design. Zhaoyan Shen, Yungang Pan, Yuhao Zhang 0006, Zhiping Jia, Xiaojun Cai, Bingzhe Li, Zili Shao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Lifetime improvement through adaptive reconfiguration for nonvolatile FPGAs
Hao Zhang 0145, Huichuan Zheng, Shuangliang Li, Mengying Zhao, Xiaojun Cai |
J. Syst. Archit. | 6 |
| 2022 | Deep Reinforcement-Learning-Guided Backup for Energy Harvesting Powered SystemsabstractEnergy harvesting technology has been widely developed as a promising alternative of battery to power embedded systems. However, energy harvesting powered embedded systems may have potential frequent power interruptions due to unstable energy supply. Nonvolatile processors (NVPs) are proposed to survive power failures by saving volatile data to nonvolatile memory (NVM) upon power failures and resuming them after power comes back. Traditionally, backup is triggered immediately when an energy warning occurs. However, it is also possible to more aggressively utilize the residual energy for program execution to improve forward progress. In this work, we propose a deep reinforcement-learning-guided backup strategy to improve forward progress in energy harvesting powered intermittent embedded systems. The experimental results show an average of 8.3%, 51.6%, and 325.3% improved forward progress compared with$Q$-learning, the related work ALD, and traditional instant backup, respectively. Weifan Sun, Mengying Zhao, Weining Song, Xiaojun Cai, Tiantian Liu 0001, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | ESD-PCM: Constructing Reliable Super Dense Phase Change Memory Under Write DisturbanceabstractPhase Change Memory (PCM) is an emerging Non-Volatile Memory (NVM) which has the characteristics of no data loss during power-off, generally no need to refresh, low power consumption, and high scalability. However, constructing super dense PCM based memory system will face Write Disturbance (WD) problem under 20nm technology node, which seriously affects data reliability and has become an urgent problem that should be solved. In this paper, we improve existing Super Dense Phase Change Memory (SD-PCM) scheme and propose Enhanced Super Dense Phase Change Memory (ESD-PCM) to mitigate WD errors in super dense PCM. ESD-PCM mainly includes the following three technical methods: first, Shared ECP based Correction makes more efficient use of Error-Correcting Pointers (ECP); second, Data Comparison based Read reduces latency, energy consumption, and space overhead during the process of Verify and Correct (VnC); third, Wear Leveling based (N:M)-Alloc achieves wear leveling and prolongs memory lifetime. Compared to basic VnC scheme, ESD-PCM reduces overhead and energy consumption by 13.7% and 14.5%, respectively. Moreover, ESD-PCM effectively reduces the probability of WD, enhances data reliability and improves system performance. Wenke Jin, Siqi Lu, Xiaojun Cai |
ETS | 3 |
| 2021 | Block-LSM: An Ether-aware Block-ordered LSM-tree based Key-Value Storage EngineabstractEthereum as one of the largest blockchain systems plays an important role in the distributed ledger, database systems, etc. As more and more blocks are mined, the storage burden of Ethereum is significantly increased. The current Ethereum system uniformly transforms all its data into key-value (KV) items and stores them to the underlying Log-Structure Merged tree (LSM-tree) storage engine ignoring the software semantics. Consequently, it not only exacerbates the write amplification effect of the storage engine but also hurts the performance of Ethereum. In this paper, we proposed a new Ethereum-aware storage model called Block-LSM, which significantly improves the data synchronization of the Ethereum system. Specifically, we first design a shared prefix scheme to transform Ethereum data into ordered KV pairs to alleviate the key range overlaps of different levels in the underlying LSM-tree based storage engine. Moreover, we propose to maintain several semantic-orientated memory buffers to isolate different kinds of Ethereum data. To save space overhead, Block-LSM further aggregates multiple blocks into a group and assigns the same prefix to all KV items from the same block group. Finally, we implement Block-LSM in the real Ethereum environment and conduct a series of experiments. The evaluation results show that Block-LSM significantly reduces up to 3.7× storage write amplification and increases throughput by 3× compared with the original Ethereum design. Bingzhe Li, Xiaojun Cai, Zhiping Jia, Zhaoyan Shen, Yi Wang 0003, Zili Shao |
ICCD | 3 |
| 2021 | Fast-convergent federated learning with class-weighted aggregation
Zezhong Ma, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
J. Syst. Archit. | 3 |
| 2021 | A lightweight online backup manager for energy harvesting powered nonvolatile processor systemsabstractWith the explosive growth of battery-free and energy-harvesting devices, the energy harvesting powered system has gained more attentions and been widely used in different fields. However, unstable harvested energy is a challenge of energy-harvesting devices since the program execution would be interrupted frequently. Non-volatile processor (NVP) is proposed to back up volatile logics before energy depletion and recover the system status after energy resumption . This paper takes the challenge of forward progress improvement issue in NVP system and proposes a lightweight online backup strategy which tries to aggressively use energy in capacitor after receiving energy warnings. We also build a flexible and accurate simulation tool for NVP system evaluation. The experimental results show an average of 24.2% and 13.7% improved forward progress compared with instant backup method and the most related work, respectively. Weining Song, Xiaojun Cai, Mengying Zhao, Zhaoyan Shen, Zhiping Jia |
J. Syst. Archit. | 2 |
| 2021 | Pearl: Performance-Aware Wear Leveling for Nonvolatile FPGAsabstractSince static random access memory (SRAM)-based field-programmable gate array (FPGA) has limited density and comparatively high leakage power, researchers have proposed FPGA architectures based on emerging nonvolatile memories (NVMs) to satisfy the requirements of data-intensive and low-power applications. Among all components, block random access memory (BRAM) has the severest endurance problem in FPGA. Unluckily, traditional wear leveling (TWL) strategies cannot be directly applied to nonvolatile FPGA because it may induce large performance overhead. In this article, we propose performance-aware wear leveling schemes for nonvolatile FPGA to improve its lifetime. Two strategies pertaining to coarse-grained wear leveling (C-Pearl) and fine-grained wear leveling (F-Pearl) are developed to balance inter-BRAM and intra-BRAM writes. Procedures, including static analysis, wear leveling-guided placement, and reconfiguration are discussed. A supportive circuit design is proposed, too. The evaluation shows that C-Pearl and F-Pearl can achieve 34% and 46% higher lifetime improvement and simultaneously 8% and 11% lower performance overhead than TWL. Mengying Zhao, Zhaoyan Shen, Xiaojun Cai, Zhiping Jia |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Maximizing CNN Throughput on FPGA ClustersabstractField Programmable Gate Array (FPGA) platform has been a popular choice for deploying Convolutional Neural Networks (CNNs) as a result of its high parallelism and low energy consumption. Due to the limitation of on-chip resources on a single board, FPGA clusters become promising solutions to improve the throughput of CNNs. In this paper, we firstly put forward strategies to optimize the resource allocation intra and inter FPGA boards. Then we model the multi-board cluster problem and design algorithms based on knapsack problem and dynamic programming to calculate the optimal topology of the FPGA clusters. We also give a quantitative analysis of the inter-board data transmission bandwidth requirement. To make our design accommodate for more situations, we provide solutions for deploying fully connected layers and special convolution layers with large memory requirement. Experimental results show that typical well-known CNNs with the proposed topology of FPGA clusters could obtain a higher throughput per board than single-board solutions and other multi-board solutions. Ruihao Li 0002, Mengying Zhao, Zhaoyan Shen, Xiaojun Cai, Zhiping Jia |
FPGA | 5 |
| 2020 | Sequence-To-Subsequence Learning With Conditional Gan For Power DisaggregationabstractNon-intrusive load monitoring (a.k.a. power disaggregation) refers to identifying and extracting the consumption patterns of individual appliances from the mains which records the whole-house energy consumption. Recently, deep learning has been shown to be a promising method to solve this problem and many approaches based on it have been proposed. In this paper, we propose a sequence-to-subsequence learning method, which makes a trade-off between traditional sequence-to-sequence and sequence-to-point method, to balance the convergence difficulty in deep neural networks and the amount of computation in the inference period. We build our model based on conditional generative adversarial network that helps us avoid designing the loss function manually. In addition, we apply U-Net and Instance Normalization techniques to our model and demonstrate their effectiveness. Evaluations are performed on real-world data sets and we achieve the state-of-the-art performance. Yungang Pan, Zhaoyan Shen, Xiaojun Cai, Zhiping Jia |
ICASSP | 4 |
| 2020 | Optimizing Motion Estimation with an ReRAM-Based PIM Architecture
Zhaoyan Shen, Zhiping Jia, Xiaojun Cai |
WASA (1) | 4 |
| 2019 | Performance-aware Wear Leveling for Block RAM in Nonvolatile FPGAsabstractField programmable gate arrays (FPGAs) have been widely adopted in both high-performance servers and embedded systems. Since static random access memory (SRAM) has limited density and comparatively high leakage power, researchers have proposed FPGA architectures based on emerging non-volatile memories (NVMs) to satisfy the requirements of data-intensive and low-power applications. Block RAM is on-chip memory of FPGAs, when it is implemented with NVM, it will face the challenge of limited endurance. Traditional wear leveling strategy cannot be directly applied to block RAM because it may induce large performance overhead. In this paper, we propose a performance-aware wear leveling scheme for block RAM in FPGAs to improve its lifetime. The placement strategy is improved by injecting wear leveling guidance. The evaluation shows that 29.75% lifetime enhancement is achieved with 16.32% performance improvement at the same time, compared with traditional wear leveling. Shuo Huai, Weining Song, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
DAC | 4 |
| 2018 | Mobility Analysis and Response for Software-Defined Internet of Things
Zhiyong Zhang 0006, Rui Wang 0075, Xiaojun Cai, Zhiping Jia |
ICA3PP (3) | 3 |
| 2018 | Shared Last-Level Cache Management and Memory Scheduling for GPGPUs with Hybrid Main MemoryabstractMemory intensive workloads become increasingly popular on general purpose graphics processing units (GPGPUs), and impose great challenges on the GPGPU memory subsystem design. On the other hand, with the recent development of non-volatile memory (NVM) technologies, hybrid memory combining both DRAM and NVM achieves high performance, low power, and high density simultaneously, which provides a promising main memory design for GPGPUs. In this article, we explore the shared last-level cache management for GPGPUs with consideration of the underlying hybrid main memory. To improve the overall memory subsystem performance, we exploit the characteristics of both the asymmetric read/write latency of the hybrid main memory architecture, as well as the memory coalescing feature of GPGPUs. In particular, to reduce the average cost of L2 cache misses, we prioritize cache blocks from DRAM or NVM based on observations that operations to NVM part of main memory have a large impact on the system performance. Furthermore, the cache management scheme also integrates the GPU memory coalescing and cache bypassing techniques to improve the overall system performance. To minimize the impact of memory divergence behaviors among simultaneously executed groups of threads, we propose a hybrid main memory and warp aware memory scheduling mechanism for GPGPUs. Experimental results show that in the context of a hybrid main memory system, our proposed L2 cache management policy and memory scheduling mechanism improve performance by 15.69% on average for memory intensive benchmarks, whereas the maximum gain can be up to 29% and achieve an average memory subsystem energy reduction of 21.27%. Chuanqi Zang, Lei Ju 0001, Mengying Zhao, Xiaojun Cai, Zhiping Jia |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2017 | Shared last-level cache management for GPGPUs with hybrid main memoryabstractMemory intensive workloads become increasingly popular on general purpose graphics processing units (GPGPUs), and impose great challenges on the GPGPU memory subsystem design. On the other hand, with the recent development of non-volatile memory (NVM) technologies, hybrid memory combining both DRAM and NVM achieves high performance, low power and high density simultaneously, which provides a promising main memory design for GPGPUs. In this work, we explore the shared last-level cache management for GPGPUs with consideration of the underlying hybrid main memory. In order to improve the overall memory subsystem performance, we exploit the characteristics of both the asymmetric read/write latency of the hybrid main memory architecture, as well as the memory coalescing feature of GPGPU. In particular, to reduce the average cost of L2 cache misses, we prioritize cache blocks from DRAM or NVM based on observation that operations to NVM part of main memory have large impact on the system performance. Furthermore, the cache management scheme also integrates the GPU memory coalescing and cache bypassing techniques to improve the overall cache hit ratio. Experimental results show that in the context of a hybrid main memory system, our proposed L2 cache management policy improves performance against the traditional LRU policy and a state-of-the-art GPU cache strategy EABP [20] by up to 27.76% and 14%, respectively. Xiaojun Cai, Lei Ju 0001, Chuanqi Zang, Mengying Zhao, Zhiping Jia |
DATE | 2 |
| 2017 | Energy-Balanced and Depth-Controlled Routing Protocol for Underwater Wireless Sensor Networks
Zhiyong Zhang 0006, Rui Wang 0075, Xiaojun Cai, Zhiping Jia |
ICA3PP | 4 |
| 2017 | ESD-WSN: An Efficient SDN-Based Wireless Sensor Network Architecture for IoT Applications
Zhiyong Zhang 0006, Rui Wang 0075, Zhiping Jia, Haijun Lei, Xiaojun Cai |
ICA3PP | 6 |
| 2017 | SLA-aware energy-efficient scheduling scheme for Hadoop YARN
Xiaojun Cai, Feng Li 0014, Lei Ju 0001, Zhiping Jia |
J. Supercomput. | 1 |
| 2016 | Energy efficient task allocation for hybrid main memory architecture
Xiaojun Cai, Lei Ju 0001, Xin Li 0002, Zhiyong Zhang 0006, Zhiping Jia |
J. Syst. Archit. | 1 |
| 2014 | High Performance FPGA Implementation of Elliptic Curve Cryptography over Binary FieldsabstractIn this paper, we propose a high performance hardware implementation architecture of elliptic curve scalar multiplication over binary fields. The proposed architecture is based on the Montgomery ladder method and uses polynomial basis for finite field (FF) arithmetic. A single Karatsuba multiplier runs with no idle cycle significantly increases the performance of FF multiplication while spending small amount of hardware resources, and other FF operations performed in parallel with the FF multiplier. The optimized circuits lead to a lesser area requirement compared to other high performance implementations. An implementation for the National Institute of Standards and Technology (NIST) recommended curve with degree 163 is shown, the proposed design can reach 121 MHz with 10,417 slices when implemented on Xilinx Virtex-4 XC4VLX200 FPGA device, the total time required for one elliptic curve scalar multiplication is 9.0 μs. Lei Ju 0001, Xiaojun Cai, Zhiping Jia, Zhiyong Zhang 0006 |
TrustCom | 3 |