VLDB 2026 Research / reviewers in the wild / expert
Dejun Jiang 0001
dblp:50/4908-1
· DBLP profile ↗
28ranked-venue papers
0as first author
14since 2021 · last 2026
0009-0001-0041-5957ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CETOFS: A High-Performance File System with Host-Server Collaboration for Remote Storage
Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
FAST | 2 |
| 2025 | PStore: End-to-End Integrity and High-Performance I/O for Cloud-Native DatabasesabstractEnsuring data integrity is critical for cloud-native databases (CNDs), which typically adopt a disaggregated architecture separating compute and storage. These systems rely on distributed file systems (DFS) for persistence, where the storage backend plays a key role in overall performance and reliability. However, existing checksum mechanisms either protect only local I/O scopes or impose significant overhead when applied end-to-end. In this paper, we present PStore, a high-performance, integrity-aware storage backend module that provides end-to-end protection through an optimized checksum architecture. By leveraging the structured I/O patterns of CNDs and employing deferred checksum recalculation, PStore significantly reduces both I/O and computational overhead. We implement PStore and evaluate it on real hardware. Results show that PStore improves performance by up to 168.2 % over existing integrity-aware solutions, while fault injection experiments confirm its robust error detection, demonstrating that strong integrity guarantees can coexist with high performance. Ying Wang 0001, Dejun Jiang 0001 |
ICPADS | 2 |
| 2025 | A High-Performance and Scalable Userspace Log-Structured File System for Modern SSDsabstractWe present AugeFS , a scalable userspace log-structured file system for modern SSDs. AugeFS re-architects the file system stack to address three critical challenges: inefficient control plane, limited metadata scalability, and underutilized device bandwidth. First, we propose a shared and protected address space within the userspace of accessing applications to run AugeFS , which enables the high-performance data plane and efficient control plane. Second, we design a scalable LSM-tree based key-value store called MetaDB to organize small-sized metadata in AugeFS . To improve the metadata scalability, MetaDB employs parallel request processing to reduce thread synchronization overhead and fine-grained parallel write-ahead log to eliminate false sharing in metadata persistence. Finally, AugeFS distributes files into different domains. To reduce contention, we maintain space management metadata for each domain independently, which helps scale data performance and improve the device IO utilization. Moreover, AugeFS designs an asynchronous IO stack for fsync to reduce the latency of synchronous writes. The evaluation results show that AugeFS significantly improves both metadata scalability and data scalability. Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
ACM Trans. Storage | 2 |
| 2024 | SuperMap: High-Performance and Flexible Memory-Mapped IO for Fast Storage DeviceabstractMemory-mapped IO offers several advantages over explicit read/write IO. It requires no system call, incurs minimal overhead in case of cache hits, and avoids extra data copies between user and kernel space. However, we still identify inefficiencies in current memory-mapped IO designs when meeting fast storage devices: i) the heavy IO stack in the page fault handler, ii) the suboptimal prefetching design, and iii) the inefficient eviction policy. To address these limitations, we present SuperMap, an alternative design for the memory-mapped IO in Linux, which specifically brings high performance and flexibility for fast devices. First, SuperMap designs a lightweight and asynchronous IO stack by directly accessing device, reducing software overhead significantly. Second, SuperMap introduces a fine-grained and application-customized prefetcher framework based on eBPF, further improving performance. Third, SuperMap proposes a hotness-aware eviction policy with the hardware assistance, trying to keep frequently accessed data in memory. Through evaluations using benchmarks and real-world applications, we demonstrate that SuperMap outperforms the state-of-the-art memory-mapped IO design (FastMap) up to 67%. Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
ICCD | 2 |
| 2024 | LBZ: A Lightweight Block Device for Supporting F2FS on ZNS SSDabstractZoned NameSpace SSD (ZNS) is emerging as a new paradigm for serving log-structured storage system (e.g. F2FS). However, existing ZNS SSD cannot directly support random write, which is still highly required by the metadata region of log-structured storage system. In this paper, we propose LBZ, a lightweight in-kernel block device running on ZND SSD with the log-structured design to allow random writes. Specifically, we design LBZ to hold the metadata area of F2FS such that F2FS can run on single ZNS SSD without using extra conventional SSD or conventional namespace of ZND SSD. LBZ first utilizes the out-of-band area of NAND flash chips to maintain the mappings between logical blocks and physical zone offsets. Since logical blocks of LBZ are sequentially written into ZNS SSD, LBZ incurs garbage collection overhead when logical blocks become invalid. In order to support efficient garbage collection in LBZ, we then propose multi-stream separation to distinguish different metadata areas of F2FS and place metadata with similar lifetimes into same zones. Meanwhile, LBZ executes three-phase checkpoint to get aware of invalid blocks and reclaim the blocks itself without modifying F2FS. We implement LBZ on real ZNS SSD and evaluate its basic performance as well as its garbage collection efficiency. The experiment results show LBZ can fully exploit the bandwidth of ZNS SSD. Moreover, LBZ provides efficient garbage collection by reducing the number of written blocks by up to 19x. Dejun Jiang 0001, Hao-Chiang Hsu, Zifeng Yang |
ICCD | 2 |
| 2024 | zQoS: Unleashing full performance capabilities of NVMe SSDs while enforcing SLOs in distributed storage systemsabstractNowadays, data centers consolidate latency-critical (LC) tenants and best-effort (BE) tenants on the same cloud platform to increase resource utilization and reduce costs. In such a scenario, the underlying distributed storage systems are responsible for guaranteeing SLOs for LC tenants while maximizing bandwidth for BE tenants. As high-performance NVMe SSDs are widely deployed, how to make full use of their performance capabilities and guarantee SLOs has become an urgent problem. However, current methods restrict the performance capabilities of NVMe SSDs based on a conservative offline model, and also ignore runtime changes in tenant loads and device states, which definitely affect the performance capabilities. Liuying Ma, Zhenqing Liu, Jin Xiong, Renhai Chen, Xi Peng 0006, Gong Zhang 0001, Dejun Jiang 0001 |
ICPP | 9 |
| 2024 | DiStore: A Fully Memory Disaggregation Friendly Key-Value Store with Improved Tail Latency and Space EfficiencyabstractMemory disaggregation decouples CPUs and memory in monolithic servers to form compute nodes (CNs) and memory nodes (MNs) for elastic and efficient memory scaling. Memory disaggregation benefits in-memory key-value stores (KVSs) that demand large memory capacity. Building a KVS with full functionality, low latency, and high space efficiency under memory disaggregation is critically required by real-world applications. However, we observe that existing disaggregated KVSs fail to achieve all the above demands simultaneously. In this paper, we present DiStore, a full-disaggregation-friendly KVS that accomplishes the above goals. We carefully divide the responsibilities of CNs and MNs when involved in index traversing, concurrency control, and memory management. We first design a two-layer indexing structure, separable adaptive linked array, to reduce network RTTs for index traversing and improve space efficiency. Then, DiStore introduces thread-context-based concurrency control to enable inter-thread collaborating to reduce stall time under multi-thread contention on CNs. Finally, we propose cachable disaggregated memory management to allow CNs to manage remote memory locally with marginal space overhead. We implement DiStore and the evaluation shows that DiStore can reduce P999 tail latency by up to 72.1% and improve space efficiency by up to 33.6%. Meanwhile, DiStore achieves comparable throughput as state-of-the-art disaggregated KVSs. Ziwei Xiong, Dejun Jiang 0001, Jin Xiong |
ICPP | 2 |
| 2023 | FastStore: A High-Performance RDMA-Enabled Distributed Key-Value Store with Persistent MemoryabstractDistributed persistent key-value store (KVS) plays an important role in today's storage infrastructure. The development of persistent memory (PM) and remote direct memory access (RDMA) allows to build distributed persistent KVS to provide fast data access. However, prior works focus on either PM-oriented or RDMA-oriented optimizations for key-value stores. We find these optimizations disallow a simple porting of RDMA-enabled KVS to PM or vice versa. This paper proposes FastStore, a high-performance distributed persistent KVS, by fully exploiting RDMA features and PM-friendly optimizations. First, FastStore utilizes RDMA-enabled PM exposure to establish direct indexing at the client side to reduce RTTs for reading values. Meanwhile, PM exposure allows PM sharing among cluster nodes, which helps to mitigate attribute-value skewness. Then, FastStore designs PM-friendly ownership transferring log and failure-atomic slotted-page allocator to achieve highly efficient PM management without PM leakage. Finally, FastStore proposes volatile search key to its B+tree indexing to reduce excessive PM accesses. We implement FastStore and the evaluation shows that FastStore outperforms the state-of-the-art ordered KVS Sherman by 2.8× higher throughput and 71.5% fewer RTTs. Ziwei Xiong, Dejun Jiang 0001, Jin Xiong |
ICDCS | 2 |
| 2023 | Exploiting Hybrid Index Scheme for RDMA-based Key-Value StoresabstractRDMA (Remote Direct Memory Access) is widely studied in building key-value stores to achieve ultra-low latency. In RDMA-based key-value stores, the indexing time takes a large fraction of the overall operation latency as RDMA enables fast data access. However, the single index structure used in existing RDMA-based key-value stores, either hash-based or sorted index, fails to support range queries efficiently while achieving high performance for singlepoint operations. In this paper, we explore the adoption of a hybrid index in the key-value stores based on RDMA, especially under the memory disaggregation architecture, to combine the benefits of a hash table and a sorted index. We propose HStore, an RDMA-based key-value store that uses a hash table for single-point lookups and leverages a skiplist for range queries to index the values stored in the memory pool. Guided by previous work on using RDMA for key-value services, HStore dedicatedly chooses different RDMA verbs to optimize the read and write performance. To efficiently keep the index structures within a hybrid index consistent, HStore asynchronously applies the updates to the sorted index by shipping the update log via two-sided verbs. Compared to state-of-the-art Sherman and Clover, HStore improves the throughput by up to 54.5% and 38.5% respectively under the YCSB benchmark. Shukai Han, Mi Zhang 0007, Dejun Jiang 0001, Jin Xiong |
SYSTOR | 3 |
| 2023 | A Survey of Non-Volatile Main Memory File Systems
Ying Wang 0001, Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
J. Comput. Sci. Technol. | 3 |
| 2023 | Dalea: A Persistent Multi-Level Extendible Hashing with Improved Tail Performance
Ziwei Xiong, Dejun Jiang 0001, Jin Xiong |
J. Comput. Sci. Technol. | 2 |
| 2022 | NapFS: A High-Performance NUMA-Aware PM File SystemabstractPersistent memory (PM) allows file systems to directly persist data on the memory bus. In order to expand the capacity of PM file system, building a file system across sockets with each attached PMs is attractive. However, accessing data across sockets incurs impacts of non-uniform memory access (NUMA) architecture, which degrade PM file system performance. In this paper, we first conduct experiments to understand the NUMA impacts on building PM file systems. We then propose four design principles for building a high-performance NUMA-aware PM file system NapFS. We architect NapFS with per-socket local PM file systems and per-socket dedicated IO thread pools. This not only allows applications to delegate data access to IO threads for avoiding remote PM access, but also fully reuses existing single-socket PM file systems to reduce implementation complexity. In addition, NapFS utilizes fast DRAM to accelerate performance by adding a global cache. We evaluate NapFS against other multi-socket PM file systems. The evaluation results show that NapFS achieves 2.2× and 1.0× throughput improvement for Filebench and RocksDB. Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
ICCD | 2 |
| 2022 | QWin: Core Allocation for Enforcing Differentiated Tail Latency SLOs at Shared Storage BackendabstractDistributed storage systems consolidate latency-critical (LC) tenants and best-effort (BE) tenants together to increase resources utilization. As a result, storage server processes require to guarantee differentiated tail latency SLOs for LC tenants, and meanwhile provide sustainable high bandwidth for BE tenants. With the adoption of high-performance NVMe SSDs, queuing time on storage server process contributes a large part to tail latency of LC tenant. Partitioning request queues, worker threads, and CPU cores on storage server processes among tenants helps to reduce interference and thus queuing time. This in turn requires careful core allocation to guarantee tail latency SLOs. In this paper, we argue that accurate core allocation is necessary for storage server processes to allocate the actual cores required by LC tenants. Thus, we propose QWin, a tail latency SLO-aware core allocation to enforce differentiated tail latency SLOs for multiple LC tenants. QWin first designs an SLO-to-core calculation model to accurately calculate the number of cores required by LC tenant. Bursty loads or fluctuated I/O latencies of storage devices can change core requirements of LC tenants. Thus, we design three core policies in QWin to adapt to the changing core requirements. We evaluate QWin by consolidating multiple LC and BE tenants together. The experiment results show that QWin outperforms the-state-of-the-art approaches in guaranteeing differentiated tail latency SLOs for LC tenants and meanwhile increasing bandwidth for BE tenants by 4x~21x. Liuying Ma, Zhenqing Liu, Jin Xiong, Dejun Jiang 0001 |
ICDCS | 4 |
| 2021 | Using Vectorized Execution to Improve SQL Query Performance on SparkabstractMapReduce-based SQL processing frameworks, such as Hive and Spark SQL, are widely used to support big data analytics. Currently these systems mainly adopt the record-at-a-time execution model, which is less efficient in terms of CPU utilization. In contrast, vectorized execution is able to make better use of CPU cache by bulk processing a record batch at a time. However, simply applying vectorized execution to MapReduce-based frameworks results in low efficient vectorized shuffle. Moreover, existing vectorized execution donot make full use of CPU cache for complex operators (e.g. Sort and Aggregation). In this paper, we present VEE, a thorough vectorized execution engine designed for SQL query processing on Spark. First, VEE designs compact in-memory data layout and serialization-aware assembling for vectorized shuffle to expedites shuffle execution, since they reduce shuffle data footprint and related computations. Secondly, VEE applies in-memory record batch rearrangement for Sort and Aggregation to greatly reduce random memory access and increase query performance. Thirdly, VEE carefully designs operator-aware batch length when handling different operators, which makes better utilization of CPU cache and increases query performance. We conduct extensive performance evaluations. The experiment results show that the performance speedup of VEE against Spark is up to 72.7% and 25.0% on average for OLAP workloads (TPC-H). The vectorized execution technologies in VEE are also applicable to other MapReduce-based data analytic frameworks to improve their query performance. Yijie Shen, Jin Xiong, Dejun Jiang 0001 |
ICPP | 3 |
| 2020 | HiLSM: an LSM-based key-value store for hybrid NVM-SSD storage systemsabstractIn order to ensure data durability and crash consistency, the LSM-tree based key-value stores suffer from high WAL synchronization overhead. Fortunately, the advent of NVM offers an opportunity to address this issue. However, NVM is currently too expensive to meet the demand of massive storage systems. Therefore, the hybrid NVM and SSD storage system provides a more cost-efficient solution. This paper proposes HiLSM, a key-value store for hybrid NVM-SSD storage systems. According to the characteristics of hybrid storage mediums, HiLSM adopts hybrid data structures consisting of the log-structured memory and the LSM-tree. Aiming at the issue of write stalls in write intensive scenario, a fine-grained data migration strategy is proposed to make the data migration start as early as possible. Aiming at the performance gap between NVM and SSD, a multi-threaded data migration strategy is proposed to make the data migration complete as soon as possible. Aiming at the LSM-tree's inherent issue of write amplification, a data filtering strategy is proposed to make data updates be absorbed in NVM as much as possible. We compare HiLSM with the state-of-the-art key-value stores via extensive experiments and the results show that HiLSM achieves 1.3x higher throughput for write, 10x higher throughput for read and 79% less write traffic under the skewed workload. Wen-Jie Li, Dejun Jiang 0001, Jin Xiong, Yungang Bao |
CF | 2 |
| 2020 | SplitKV: Splitting IO Paths for Different Sized Key-Value Items with Advanced Storage Devices
Shukai Han, Dejun Jiang 0001, Jin Xiong |
HotStorage | 2 |
| 2020 | Gecko: Guaranteeing Latency SLO on a Multi-Tenant Distributed Storage SystemabstractMeeting tail latency Service Level Objective (SLO) as well as achieving high resource utilization is important to distributed storage systems. Recent works adopt strict priority scheduling or constant rate limiting to provide SLO guarantee but cause under-utilization resources. To address this issue, we first analyze the relationship between workload burst and latency SLO. Based on burst patterns and latency SLOs, we classify tenants into two categories: Postponement-Tolerable tenant and Postponement-Intolerable tenant. We then explore the opportunity to improve resource utilization by carefully allocating resources to each tenant type. We design Rate-Limiting-Priority scheduling algorithm to limit the impact of high priority tenants on low priority ones. Meanwhile, we propose Postponement-Aware scheduling algorithm which allows Postponement-Intolerable tenants to preempt system capacity from Postponement-Tolerable tenants. This helps to increase resource utilization. We propose a latency SLO guarantee framework Gecko. Gecko guarantees multi-tenant latency SLOs via combining the two proposed scheduling algorithms together with an admission control strategy. We evaluate Gecko with real production traces and the results show that Gecko admits 44% more tenants on average than state-of-the-art techniques meanwhile guaranteeing latency SLO. Zhenyu Leng, Dejun Jiang 0001, Liuying Ma, Jin Xiong |
ICPADS | 2 |
| 2020 | SrSpark: Skew-resilient Spark based on Adaptive Parallel ProcessingabstractMapReduce-based SQL processing systems, e.g., Hive and Spark SQL, are widely used for big data analytic applications due to automatic parallel processing on large-scale machines. They provide high processing performance when loads are balanced across the machines. However, skew loads are not rare in real applications. Although many efforts have been made to address the skew issue in MapReduce-based systems, they can neither fully exploit all available computing resources nor handle skews in SQL processing. Moreover, none of them can expedite the processing of skew partitions in case of failures. In this paper, we present SrSpark, a MapReduce-based SQL processing system that can make full use of all computing resources for both non-skew loads and skew loads. To achieve this goal, SrSpark introduces fine-grained processing and work-stealing into the MapReduce framework. More specifically, SrSpark is implemented based on Spark SQL. In SrSpark, partitions are further divided into sub-partitions and processed in sub-partition granularity. Moreover, SrSpark adaptively uses both intra-node and inter-node parallel processing for skew loads according to available computing resources in realtime. Such adaptive parallel processing increases the degree of parallelism and reduces the interaction overheads among the cooperative worker threads. In addition, SrSpark checkpoints sub-partition's processing results periodically to ensure fast recovery from failures during skew partition processing. Our experiment results show that for skew loads, SrSpark outperforms Spark SQL by up to 3.5x, and 2.2x on average, while the performance overhead is only about 4% under non-skew loads. Yijie Shen, Jin Xiong, Dejun Jiang 0001 |
ICPADS | 3 |
| 2018 | Caching or Not: Rethinking Virtual File System for Non-Volatile Main Memory
Ying Wang 0001, Dejun Jiang 0001, Jin Xiong |
HotStorage | 2 |
| 2018 | H-Scheduler: Storage-Aware Task Scheduling for Heterogeneous-Storage Spark ClustersabstractA trend in nowadays data centers is that heterogeneous storage devices are deployed to meet different storage demands of various big data workloads. For example, many nodes are equipped with both SSDs and HDDs. And HDFS has introduced the heterogeneous-storage-aware feature to adapt to such hybrid storage clusters. However, current task scheduler on big data processing platforms (such as Hadoop and Spark) only considers the overhead of network data transmission by exploiting the data locality principle. On heterogeneous storage clusters, task completion time is also affected by the speed of storage devices (SSDs and HDDs) where the data are stored. Ignoring the different speed of storage devices results in poor utilization of high speed devices such as SSD. In this paper, we propose a task scheduling strategy for heterogeneous storage clusters called H-Scheduler. The key idea of H-Scheduler is to differentiate speeds of storage devices by storage types. It classifies the tasks by both data locality and storage types, and redefines the priorities of different classes of tasks by both storage device speed and data locality to reduce job execution time. We implemented H-Scheduler in Spark, and the experiment results show that H-Scheduler can reduce job execution time by up to 73.6 %, depending on the workload characteristics and data distribution among different types of storage devices. Fengfeng Pan, Jin Xiong, Yijie Shen, Tianshi Wang 0002, Dejun Jiang 0001 |
ICPADS | 5 |
| 2017 | HiKV: A Hybrid Index Key-Value Store for DRAM-NVM Memory Systems
Dejun Jiang 0001, Jin Xiong, Ninghui Sun |
USENIX ATC | 2 |
| 2017 | HAP: Hybrid-Memory-Aware Partition in Shared Last-Level CacheabstractData-center servers benefit from large-capacity memory systems to run multiple processes simultaneously. Hybrid DRAM-NVM memory is attractive for increasing memory capacity by exploiting the scalability of Non-Volatile Memory (NVM). However, current LLC policies are unaware of hybrid memory. Cache misses to NVM introduce high cost due to long NVM latency. Moreover, evicting dirty NVM data suffer from long write latency. We propose hybrid memory aware cache partitioning to dynamically adjust cache spaces and give NVM dirty data more chances to reside in LLC. Experimental results show Hybrid-memory-Aware Partition (HAP) improves performance by 46.7% and reduces energy consumption by 21.9% on average against LRU management. Moreover, HAP averagely improves performance by 9.3% and reduces energy consumption by 6.4% against a state-of-the-art cache mechanism. Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001 |
ACM Trans. Archit. Code Optim. | 2 |
| 2016 | S-RAC: SSD Friendly Caching for Data Center WorkloadsabstractCurrent data-center applications tend to process increasingly large volume of data sets. The caching effect of page cache is reduced by its limited capacity. Emerging flash-based solid state drives (SSD) have latency and price advantages compared to hard disk and DRAM. Thus, SSD-based caching is widely deployed in data centers. However, SSD caching faces two challenges. First, SSD has limited write endurance, which requires cache manager to reduce write amount to SSD. Second, data-center workloads exhibit a diverse I/O access patterns, which requires one to figure out SSD caching friendly access patterns. This paper first classifies 6 I/O access patterns among 32 data-center workloads using a cost-benefit analysis. We derive implications for designing SSD cache from analyzing the access patterns. We then propose an SSD cache manager S-RAC with re-adding blocks and ghost cache adaptation to retain SSD friendly blocks in SSD. The experimental evaluation shows the efficiency of S-RAC in reducing SSD write amount while improving/maintaining cache hit ratio. Yuanjiang Ni, Ji Jiang, Dejun Jiang 0001, Xiaosong Ma, Jin Xiong, Yuangang Wang |
SYSTOR | 3 |
| 2015 | Exploiting Program Semantics to Place Data in Hybrid MemoryabstractLarge-memory applications like data analytics and graph processing benefit from extended memory hierarchies, and hybrid DRAM/NVM (non-volatile memory) systems represent an attractive means by which to increase capacity at reasonable performance/energy tradeoffs. Compared to DRAM, NVMs generally have longer latencies and higher energies for writes, which makes careful data placement essential for efficient system operation. Data placement strategies that resort to monitoring all data accesses and migrating objects to dynamically adjust data locations incur high monitoring overhead and unnecessary memory copies due to mispredicted migrations. We find that program semantics (specifically, global access characteristics) can effectively guide initial data placement with respect to memory types, which, in turn, makes run-time migration more efficient. We study a combined offline/online placement scheme that uses access profiling information to place objects statically and then selectively monitors run-time behaviors to optimize placements dynamically. We present a software/hardware cooperative framework, 2PP, and evaluate it with respect to state-of-the-art migratory placement, finding that it improves performance by an average of 12.1%. Furthermore, 2PP improves energy efficiency by up to 51.8%, and by an average of 18.4%. It does so by reducing run-time monitoring and migration overheads. Dejun Jiang 0001, Sally A. McKee, Jin Xiong, Mingyu Chen 0001 |
PACT | 2 |
| 2015 | A Survey of Phase Change Memory Systems
Dejun Jiang 0001, Jin Xiong, Ninghui Sun |
J. Comput. Sci. Technol. | 2 |
| 2014 | HAP: Hybrid-memory-Aware Partition in shared Last-Level CacheabstractData-center servers require large capacity main memory to run multiple workloads simultaneously. However, the scalability and power consumption of DRAM limit its capability of constructing large capacity memory. Emerging non-volatile memories (e.g. PCM and STT-RAM) provide better scalability and lower power leakage than DRAM. Especially, hybrid memory consisting of DRAM and NVM is able to exploit advantages of different memory medias. However, NVMs have a few drawbacks, such as relatively longer read and write latency. Cache miss at the shared last level cache (LLC) suffers from longer latency if the missing data resides in NVM. Current LLC policies manage the cache space without being aware of the underlying heterogeneous medias. This results in cache performance degradation if a large number of missing data come from NVM. Taking the asymmetric cache miss cost into account, we first propose a new performance metric -TMPKI, which can exactly reflect the LLC performance on the top of hybrid memories. Then we propose a hybrid memory aware cache partitioning technique (HAP) to dynamically adjust the cache spaces for DRAM and NVM data based on TMPKI. Experimental results show that HAP improves performance against the traditional LRU policy by up to 54.3% (19.6% on average) while it incurs a little storage overhead (0.2%). Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001 |
ICCD | 2 |
| 2014 | Write-aware random page initialization for non-volatile memory systemsabstractDue to the high scalability and low power leakage, emerging non-volatile memories (NVMs) are promising to be integrated into memory hierarchy. However, NVMs have write issues, such as limited write endurance and high write energy. This paper observes that a large fraction of writes are caused by memory page initialization in the OS kernel stack. Thus, this paper proposes a write-aware random page initialization technique (WRPI) to reduce writes without sacrificing system security. Instead of initializing all bits of an allocated page, WRPI randomly initializes part bits. Moreover, WRPI sets the values of initialized bits to zeros or ones that require writing the least number of bits. The evaluation results show that WRPI can reduce writing bits to NVM memory by up to 21.0% and 11.3% on average, compared to the conventional page initialization. WRPI can also reduce the write energy consumption and the total energy consumption of NVM memory system by 14.0% and 5.8% on average, respectively. Dejun Jiang 0001, Jin Xiong, Ninghui Sun |
ICCD | 2 |
| 2014 | DWC: dynamic write consolidation for phase change memory systemsabstractPhase change memory (PCM) is promising to become an alternative main memory thanks to its better scalability and lower leakage than DRAM. However, the long write latency of PCM puts it at a severe disadvantage against DRAM. In this paper, we propose a Dynamic Write Consolidation (DWC) scheme to improve PCM memory system performance while reducing energy consumption. This paper is motivated by the observation that a large fraction of a cache line being written back to memory is not actually modified. DWC exploits the unnecessary burst writes of unmodified data to consolidate multiple writes targeting the same row into one write. By doing so, DWC enables multiple writes to be send within one. DWC incurs low implementation overhead and shows significant efficiency. The evaluation results show that DWC achieves up to 35.7% performance improvement, and 17.9% on average. The effective write latency are reduced by up to 27.7%, and 16.0% on average. Moreover, DWC reduces the energy consumption by up to 35.3%, and 13.9% on average. Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001, Lixin Zhang 0002, Ninghui Sun |
ICS | 2 |