Linbo Long

dblp:141/0607 · DBLP profile ↗
← Back
34ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0003-1966-0714ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 9 first-author · 23 since 2021Software engineering, systems software and programming languages · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Nemo: A Low-Write-Amplification Cache for Tiny Objects on Log-Structured Flash Devices
Xufeng Yang, Jingxin Hu, Congming Gao, Tianyang Jiang, Linbo Long, Yina Lv, Jiwu Shu
ASPLOS (2)8
2026 Optimizing F2FS performance with the inter-zone parallelism in small-zone ZNS SSDs
Linbo Long, Xinrui Dong, Ting Wu 0012, Jingcheng Shen, Kan Zhong
Future Gener. Comput. Syst.1
2026 CXL shared coherent memory simulation and cross-host synchronization mechanisms design for data sharing
Ting Wu 0012, Qingyuan Song, Xihong Huang, Linbo Long, Zhulin Ma, Weichen Liu 0001
J. Syst. Archit.4
2026 WPAlloc: An Efficient Wear-Leveling-Aware Parallel Allocator for Persistent Memory File Systems
abstract
Internet and IoT applications have generated increasing amounts of data that require efficient storage. Many persistent memory file systems have been designed to handle high-performance storage demands by fully exploiting the senior features of persistent memory (PM). However, PMs suffer from limited write endurance. Existing PM file systems achieve PM wear-leveling by designing wear-leveling-aware allocators. These allocators focus on providing higher-balanced writes to PMs while neglecting the overhead. Moreover, they cause serious request conflicts in parallel block requests by multiple threads in modern multiprocessor computer systems. In this paper, we propose an efficient wear-leveling-aware parallel allocator, WPAlloc, for persistent memory file systems to achieve wear-leveling of PM and high parallel performance. The essential idea of WPAlloc is to allocate blocks with lower write counters for each allocation request and to provide parallel block allocation and deallocation for multiple threads via one free list per logical processor. WPAlloc consists of two key techniques: the bucket sort-based range management scheme (BRMS) and the parallel allocation and deallocation scheme (PADS). First, we design the BRMS to obtain the less worn free blocks. Based on BRMS, an online wear range adjustment algorithm has been designed to adjust the wear range dynamically. Then, we present the PADS to avoid request conflicts by multiple threads. We implement WPAlloc based on PMFS. Experimental results show that WPAlloc can reduce the maximum write by 54.2%, 10.2%, and 55.7%, while achieving average performance improvements of 5.68%, 54.18%, and 11.28% compared to PMFS, DWARM, and WASA, respectively.
Ting Wu 0012, Linbo Long, Zhulin Ma, Duo Liu 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 SecDS: A security-aware DAG task scheduling strategy for edge computing
Linbo Long, Jingcheng Shen
Future Gener. Comput. Syst.1
2025 PIM-IoT: Enabling hierarchical, heterogeneous, and agile Processing-in-Memory in IoT systems
Kan Zhong, Qiao Li 0001, Ao Ren, Yujuan Tan, Xianzhang Chen, Linbo Long, Duo Liu 0002
Future Gener. Comput. Syst.6
2025 Overlapping Aware Data Placement Optimizations for LSM Tree-Based Store on ZNS SSDs
abstract
Solid State Drives (SSDs) based on the NVMe Zoned Namespaces (ZNS) interface can notably reduce the costs of address mapping, garbage collection, and over-provisioning by dividing the storage space into multiple zones for sequential writes and random reads. The Log-Structured Merge (LSM) tree, which is extensively used in key-value storage systems, converts random writes to sequential writes, hence a suitable scenario to utilize ZNS SSDs. However, LSM tree associated data significantly varies in lifetime due to the levels and merging mechanisms of the LSM tree. Therefore, without an accurate method to estimate data lifetime, data with disparate lifetimes may be placed in the same zone, thus causing low space utilization and high write amplification within the SSD. To address these issues, the article proposes two data overlapping aware optimizations to realize intelligent data placement: a zone allocation scheme and a garbage collection scheme. The key technique of these optimizations is an accurate data-lifetime estimation by considering both the associated tree level of the data and the data overlapping ratio between the data and those in the neighboring level. Using the estimation technique, the zone allocation optimization can place data with similar lifetimes in the same zone. Besides, the garbage collection optimization can reclaim zones in an adaptive manner based on overlapping ratios to reduce the amount of data migration. Experimental results demonstrate that the optimization schemes effectively reduce garbage collection-incurred data copy by average factors of 2.11× and 1.50× in comparison to a conventional work and a state-of-the-art work, respectively. Consequently, the proposed work successfully alleviates the write amplification effect by 18% and 6%, compared to the conventional work and the state-of-the-art work, respectively.
Jingcheng Shen, Linbo Long, Zhenhua Tan, Congming Gao, Kan Zhong, Masao Okita, Fumihiko Ino
ACM Trans. Archit. Code Optim.3
2024 Overlapping Aware Zone Allocation for LSM Tree-Based Store on ZNS SSDs
abstract
NVMe Zoned Namespace (ZNS) devices partition the storage space into sequential-write zones, notably reducing the costs of address mapping, garbage collection (GC), and overprovisioning. Log-Structured Merge (LSM) tree-based databases convert random writes into sequential writes and can thus be efficiently handled by ZNS devices. Efficient zone-allocation methods play a pivotal role in maximizing the performance of LSM tree-based store running on ZNS devices. However, existing zone-allocation methods encounter high write-amplification factors due to inaccurate lifetime estimation solely based on the LSM-tree levels. To address this, this paper proposes an overlapping-aware zone-allocation method, termed OAZA, which efficiently selects suitable zones to place data. First, OAZA estimates the data lifetime by considering both the LSM-tree level of the data and the relative data hotness within the same tree level. Secondly, OAZA intelligently selects an appropriate zone to store the data based on the estimated lifetime. Experimental results demonstrate that OAZA outperforms two zone-allocation methods that correlate data lifetime merely to the tree level. Specially, OAZA reduces the amount of GC-induced data copy by average factors of 2.7 × and 1.7× in comparison to the two methods, respectively. Additionally, OAZA achieves an impressively low write-amplification factor of 1.1 ×, outperforming the factors of 1.2× and 1.3× achieved by the two compared methods, respectively.
Jingcheng Shen, Linbo Long, Renping Liu 0002, Zhenhua Tan, Congming Gao
ASPDAC3
2024 Para-ZNS: Improving Small-Zone ZNS SSDs Parallelism Through Dynamic Zone Mapping
abstract
The emerging Zoned Namespace (ZNS) interface helps flash-based SSDs achieve high performance by dividing the logical space into fixed-size zones. Typically, a zone is mapped to blocks across multiple dies to achieve I/O parallelism. Small zones can make better use of space and are therefore widely studied. However, a small zone fails to be mapped to blocks residing on all dies, causing underutilized die-level parallelism. Meanwhile, a fine-grained (i.e., plane-level) parallelism is rarely exploited for ZNS SSDs due to a strict limitation mandating that only the same type of operation can be simultaneously performed on the same address across different planes within a die. To address these issues, this paper proposes a novel small-zone ZNS-SSD design with dynamic zone mapping, named Para-ZNS. First, a new parallel block grouping module is devised to group blocks across all planes from multiple dies as a basic unit to be mapped to a zone. Such a basic mapping unit achieves parallelism among multiple dies and plane-level parallelism. Then, a die-parallelism identification module is implemented to locate idle dies. Subsequently, to fully exploit the die-level parallelism, a dynamic zone mapping scheme is employed to intelligently map the basic mapping units on the identified idle dies to open zones. The evaluation results based on a widely-used I/O tester (FIO) demonstrate that Para-ZNS improves the bandwidth by 3.42× on average in comparison to state-of-the-art work.
Zhenhua Tan, Linbo Long, Jingcheng Shen, Congming Gao, Renping Liu 0002
DATE2
2024 Hi-ZNS: High Space Efficiency and Zero-Copy LSM-Tree Based Stores on ZNS SSDs
abstract
The Zoned Namespace (ZNS) SSD is a newly introduced storage device and provides several new ZNS commands to upper-level applications. Zone-reset command is one of the ZNS commands to erase all the flash blocks within a zone. Since data is grouped and erased in zone units, ZNS SSDs are widely used in LSM-tree-based stores. However, the basic invalidated unit in LSM-tree is an SST/WAL file, which mismatches the erasing unit of a ZNS SSD. Placing different SST/WAL files in the same zone, LSM-tree on ZNS SSDs faces dramatic space amplification and extensive data migration problems.
Renping Liu 0002, Peng Chen 0027, Linbo Long, Anping Xiong, Duo Liu 0002
ICPP4
2024 DPC: DPU-accelerated High-Performance File System Client
abstract
To achieve efficient file access to the file system backend, file system clients employ various intricate optimization techniques, such as local data/metadata caching and direct data access. However, these techniques impose a significant load on the host CPU, posing substantial challenges to the valuable CPU resources.
Kan Zhong, Zhiwang Yu, Qiao Li 0001, Xianqiang Luo, Linbo Long, Yujuan Tan, Ao Ren, Duo Liu 0002
ICPP5
2024 ZNS-Cleaner: Enhancing lifespan by reducing empty erase in ZNS SSDs
Renping Liu 0002, Peng Chen 0027, Linbo Long, Anping Xiong, Duo Liu 0002
J. Syst. Archit.5
2024 DAG-Order: An Order-Based Dynamic DAG Scheduling for Real-Time Networks-on-Chip
abstract
With the high-performance requirement of safety-critical real-time tasks, the platforms of many-core processors with high parallelism are widely utilized, where network-on-chip (NoC) is generally employed for inter-core communication due to its scalability and high efficiency. Unfortunately, large uncertainties are suffered on NoCs from both the overly parallel architecture and the distributed scheduling strategy (e.g., wormhole flow control), which complicates the response time upper bounds estimation (i.e., either unsafe or pessimistic). For DAG-based real-time parallel tasks, to solve this problem, we propose DAG-Order, an order-based dynamic DAG scheduling approach, which strictly guarantees NoC real-time services. First, rather than build the new analysis to fit the widely used best-effort wormhole NoC, DAG-Order is built upon a kind of advanced low-latency NoC with SLT ( S ingle-cycle L ong-range T raversal) to avoid the unpredictable parallel transmission on the shared source-destination link of wormhole NoCs. Second, DAG-Order is a non-preemptive dynamic scheduling strategy, which jointly considers communication as well as computation workloads, and fits SLT NoC. With such an order-based dynamic scheduling strategy, the provably bound safety is ensured by enforcing certain order constraints among DAG edges/vertices that eliminate the execution-timing anomaly at runtime. Third, the order constraints are further relaxed for higher average-case runtime performance without compromising bound safety. Finally, an effective heuristic algorithm seeking a proper schedule order is developed to tighten the bounds. Experiments on synthetic and realistic benchmarks demonstrate that DAG-Order performs better than the state-of-the-art related scheduling methods.
Peng Chen 0027, Hui Chen 0016, Weichen Liu 0001, Linbo Long, Wanli Chang 0001, Nan Guan
ACM Trans. Archit. Code Optim.4
2024 WA-Zone: Wear-Aware Zone Management Optimization for LSM-Tree on ZNS SSDs
abstract
ZNS SSDs divide the storage space into sequential-write zones, reducing costs of DRAM utilization, garbage collection, and over-provisioning. The sequential-write feature of zones is well-suited for LSM-based databases, where random writes are organized into sequential writes to improve performance. However, the current compaction mechanism of LSM-tree results in widely varying access frequencies (i.e., hotness) of data and thus incurs an extreme imbalance in the distribution of erasure counts across zones. The imbalance significantly limits the lifetime of SSDs. Moreover, the current zone-reset method involves a large number of unnecessary erase operations on unused blocks, further shortening the SSD lifetime. Considering the access pattern of LSM-tree, this article proposes a wear-aware zone-management technique, termed WA-Zone , to effectively balance inter- and intra-zone wear in ZNS SSDs. In WA-Zone, a wear-aware zone allocator is first proposed to dynamically allocate data with different hotness to zones with corresponding lifetimes, enabling an even distribution of the erasure counts across zones. Then, a partial-erase-based zone-reset method is presented to avoid unnecessary erase operations. Furthermore, because the novel zone-reset method might lead to an unbalanced distribution of erasure counts across blocks in a zone, a wear-aware block allocator is proposed. Experimental results based on the FEMU emulator demonstrate the proposed WA-Zone enhances the ZNS-SSD lifetime by 5.23×, compared with the baseline scheme.
Linbo Long, Shuiyong He, Jingcheng Shen, Renping Liu 0002, Zhenhua Tan, Congming Gao, Duo Liu 0002, Kan Zhong
ACM Trans. Archit. Code Optim.1
2024 Optimizing Garbage Collection for ZNS SSDs via In-storage Data Migration and Address Remapping
abstract
The NVMe Zoned Namespace (ZNS) is a high-performance interface for flash-based solid-state drives (SSDs), which divides the logical address space into fixed-size and sequential-write zones. Meanwhile, ZNS SSDs eliminate in-device garbage collection (GC) by shifting the responsibility of GC to the host. However, the host-side GC of ZNS SSDs is not efficient. On the one hand, data migration during GC first moves data to the host buffer and then writes back the transferred data to the new location in the SSD, resulting in an unnecessary end-to-end transfer overhead. On the other hand, due to the pre-configured mapping between zones and blocks, GC incurs a large block-to-block rewrite overhead, i.e., even if most of the data in a block of the victim zone is valid, the valid data will still be rewritten to another block in the target zone. To address these issues, this article proposes a novel ZNS SSD design that features dynamic zone mapping, termed Brick-ZNS . Brick-ZNS implements two key functionalities: in-storage data migration and address remapping. New ZNS commands are first designed to realize in-storage data migration to avoid the end-to-end transfer overhead of GC while ensuring performance predictability. Then, a remapping strategy exploiting parallel physical blocks is proposed to reduce the large block-to-block rewrite overhead while ensuring zone-level access parallelism. The basic idea of the strategy is to directly remap the parallel physical blocks with a sufficient amount of valid data in the victim zone to the target zone, hence avoiding the large block-to-block rewrite overhead. Based on a full-stack SSD emulator, the evaluation results show that Brick-ZNS improves write throughput by 25% and SSD lifetime by 1.41×.
Zhenhua Tan, Linbo Long, Jingcheng Shen, Renping Liu 0002, Congming Gao, Kan Zhong
ACM Trans. Archit. Code Optim.2
2024 Fair-ZNS: Enhancing Fairness in ZNS SSDs Through Self-Balancing I/O Scheduling
abstract
The NVMe Zoned Namespace (ZNS) is a new type of storage interface, which divides logical address space into fixed-size zones, and each zone strictly follows a sequential write constraint with a write pointer. Owing to the sequential write constraint of the ZNS, I/O requests would not be scheduled arbitrarily like the traditional SSDs with block interface. When multiple applications concurrently access one ZNS SSD hardware, the constraint deteriorates I/O blocking and causes huge unfairness. To resolve the problem, we propose a self-balance I/O scheduling dedicated for ZNS SSDs, called Fair-ZNS, to balance the slowdown among multiple applications and ensure fairness. Fair-ZNS identifies the unfair requests by the maximum slowdown value, and violently schedules these requests into the head of the queues overcoming the sequential write constraint. To eliminate the negative effect of the violent scheduling, Fair-ZNS deploys a self-balancing coordinator to fine-tune the order of the requests. Comprehensive evaluations show that Fair-ZNS alleviates I/O blocking and reduces average waiting time by 8.3×, increases fairness by 2.3×, and decreases the max slowdown by 5.1× averagely when compared to the current ZNS SSDs.
Renping Liu 0002, Zhenhua Tan, Linbo Long, Duo Liu 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2023 Optimizing Data Migration for Garbage Collection in ZNS SSDs
abstract
ZNS SSDs shift the responsibility of garbage collection (GC) to the host. However, data migration in GC needs to move data to the host's buffer first and write back to the new location, resulting in an unnecessary end-to-end transfer overhead. Moreover, due to the pre-configured mapping between zones and blocks, GC needs to perform a large number of unnecessary block-to-block data migrations between zones. To address these issues, this paper proposes a simple and efficient data migration method, called IS-AR, with in-storage data migration and address remapping. Based on a full-stack SSD emulator, our evaluation shows that IS-AR reduces GC latency by 6.78× and improves SSD lifetime by 1.17× on average.
Zhenhua Tan, Linbo Long, Renping Liu 0002, Congming Gao
DATE2
2023 ADAR: Application-Specific Data Allocation and Reprogramming Optimization for 3-D TLC Flash Memory
abstract
High bit-density flash memories, such as triple-level cell (TLC) and quad-level cell (QLC), have been widely used in flash memory-based storage systems, offering significantly high capacity. However, these high bit-density flash memories suffer from asymmetric access performance on the different pages that sharing the same physical cells. Meanwhile, 3-D flash memory adopts stacking technology to increase capacity and reduce cost per bit. The flash unit can be reprogrammed many times as long as the voltage increases. The reprogramming technology is also an effective solution for further increasing the 3-D flash capacity, allowing multiple program operations in an erase cycle. Considering the restrictions of reprogram operations, solid-state drives (SSDs) should capture the access pattern to perform more reprogramming operations to realize the joint optimization of read and write performance. In this work, we propose an application-specific data allocation and reprogramming technique named ADAR to enhance the read and write performance of 3-D TLC flash memory-based SSDs. The core idea is to allocate low-latency least significant bit (LSB) and central significant bit pages to frequently updated write data (termed hot write data) to improve the write performance, and reprogram the pages from high-latency pages (e.g., most significant bit page) to low-latency pages (e.g., LSB page mode) to enhance the read performance while initially storing frequently read data (termed hot read data) in high-latency pages. We explored data access patterns and designed an effective hotness identification method to present a new data allocation and reprogramming technique for 3-D TLC flash memory. Based on a modified 3-D TLC SSD simulator with typical workloads, our evaluation showed that our technique achieved 35.36% and 25.72% performance improvements in read and write latencies, respectively.
Linbo Long, Jinpeng Huang, Congming Gao, Duo Liu 0002, Renping Liu 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 A compression-based memory-efficient optimization for out-of-core GPU stencil computation
Jingcheng Shen, Linbo Long, Masao Okita, Fumihiko Ino
J. Supercomput.2
2022 Optimizing data placement and size configuration for morphable NVM based SPM in embedded multicore systems
Linbo Long, Jinpei Du, Xuxu Deng, Renping Liu 0002, Yan Wang 0022
Future Gener. Comput. Syst.1
2022 Performance-oriented cache management scheme based on a retention state for energy-harvesting nonvolatile processors
Yan Wang 0022, Henian Fang, Linbo Long
Future Gener. Comput. Syst.3
2022 Efficient persistent memory file systems using virtual superpages with multi-level allocator
Chaoshu Yang, Zhiwang Yu, Runyu Zhang 0002, Shun Nie, Hui Li 0046, Xianzhang Chen, Linbo Long, Duo Liu 0002
J. Syst. Archit.7
2022 Improving Fairness for SSD Devices through DRAM Over-Provisioning Cache Management
abstract
Modern NVMe SSDs have been widely deployed in multi-tenant cloud computing environments or multi-programming systems. When multiple applications concurrently access one SSD hardware, unfairness within the shared SSD will slow down the application significantly and lead to a violation of service level objectives. However, traditional data cache management within SSDs mainly focuses on improving cache hit ratio, which causes data cache contention and sacrifices fairness among multiple applications. In this paper, we propose a DRAM-based Over-Provisioning (OP) cache management mechanism, named Justitia, to reduce data cache contention and improve fairness for modern SSDs. Justitia consists of two stages includingStatic-OPstage andDynamic-OPstage. Through the novel OP mechanism in the two stages, Justitia reduces the max slowdown by$4.5\times$on average. At the same time, Justitia increases fairness by$20.6\times$and buffer hit ratio by$19.6\%$averagely, compared with the traditional shared mechanism.
Renping Liu 0002, Zhenhua Tan, Linbo Long, Yu Wu 0016, Yujuan Tan, Duo Liu 0002
IEEE Trans. Parallel Distributed Syst.3
2019 Towards Fast and Lightweight Checkpointing for Mobile Virtualization Using NVRAM
abstract
Checkpointing is a key enabler of hibernation, live migration and fault-tolerance for virtual machines (VMs) in mobile devices. However, checkpointing a VM is usually heavyweight: the VM's entire memory needs to be dumped to storage, which induces a significant amount of (slow) I/O operations, degrading system performance and user experience. In this paper, we propose FLIC, a fast and lightweight checkpointing machinery for virtualized mobile devices by taking advantages of recent byte-addressable, non-volatile memory (NVRAM). Instead of saving the VM's entire memory to storage, we store its working set pages in NVRAM, avoiding accessing slow flash memory (compared to server-grade SSDs). To further reduce the write activities to flash memory, we propose an energy-efficient data deduplication to eliminate redundant data in VM snapshot and save storage space. Experimental results based on an Exynos 5250 SoC show that our approach can effectively improve the performance of checkpointing in mobile virutalization and save energy.
Kan Zhong, Duo Liu 0002, Yunsong Wu, Linbo Long, Weichen Liu 0001, Jinting Ren, Renping Liu 0002, Liang Liang 0002, Zili Shao, Tao Li 0006
IEEE Trans. Parallel Distributed Syst.4
2017 Revisiting swapping in mobile systems with SwapBench
Duo Liu 0002, Liang Liang 0002, Kan Zhong, Linbo Long, Meikang Qiu, Zili Shao, Edwin H.-M. Sha
Future Gener. Comput. Syst.5
2016 FLIC: Fast, lightweight checkpointing for mobile virtualization using NVRAM
Kan Zhong, Duo Liu 0002, Liang Liang 0002, Linbo Long, Zili Shao
DATE4
2016 A compiler assisted wear leveling for morphable PCM in embedded systems
Linbo Long, Edwin H.-M. Sha, Duo Liu 0002, Liang Liang 0002, Kan Zhong
J. Syst. Archit.1
2016 Morphable Resistive Memory Optimization for Mobile Virtualization
abstract
Virtualization offers significant benefits, such as better isolation and security for mobile systems. However, the limited amount of memory and virtualization's memory-demanding nature make it challenging to virtualize mobile systems efficiently. In this paper, we utilize morphable resistive memories to design a high-performance mobile system with an extensible memory space. With morphable resistive memories, a simple and effective page management technique, Balloonfish, is proposed to convert the memory cell state between multilevel and single-level for achieving a balance between performance and memory space. First, an application-specific page allocation is proposed for managing morphable resistive memories in virtualized mobile systems. Besides, we use a balloon-style algorithm to balance memory allocation among multiple virtual machines. Our evaluation based on the Samsung Exynos 5250 system-on-chip with various real Android applications shows that our system achieves 28.63% performance improvement compared with the baseline scheme.
Linbo Long, Duo Liu 0002, Liang Liang 0002, Kan Zhong, Zili Shao, Edwin H.-M. Sha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2016 Energy-Efficient In-Memory Paging for Smartphones
abstract
Smartphones are becoming increasingly energy-hungry to support feature-rich applications, posing a lot of pressure on battery lifetime and making energy consumption a non-negligible issue. In particular, dynamic random access memory (DRAM)-based main memory subsystem is a major contributor to the energy consumption of mobile devices. In this paper, we propose direct read (DR). Swap, an energy-efficient in-memory paging design to reduce energy consumption in smartphones. In DR. Swap, we adopt emerging energy-efficient nonvolatile memory (NVM) and use it as the swap area. Utilizing NVMs byte-addressability, we propose DR which guarantees zero memory copy for read-only requests when accessing a page in swap area. To better understand the energy consumption of swapping, we build an energy model to analyze the energy consumption of different paging architectures. We evaluate DR. Swap based on the Google Nexus 5 smartphone, experimental results show that our technique can reduce more than 50% energy consumption compared to DRAM backed swapping.
Kan Zhong, Duo Liu 0002, Liang Liang 0002, Linbo Long, Yi Wang 0003, Edwin H.-M. Sha
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2015 Balloonfish: Utilizing morphable resistive memory in mobile virtualization
abstract
Virtualization offers significant benefits such as better isolation and security for mobile systems. However, the limited amount of memory and virtualization's memory-demanding nature makes it challenging to virtualize mobile systems efficiently. In this paper, we utilize morphable resistive memories to design a high-performance mobile system with extensible memory space. With morphable resistive memory, we convert the memory cell state between multi-level and single-level to achieve a balance between performance and memory space. Our evaluation based on the Samsung Exynos 5250 SoC with real Android applications shows that our system achieve 27% performance improvement compared with the baseline scheme.
Linbo Long, Duo Liu 0002, Kan Zhong, Zili Shao, Edwin H.-M. Sha
ASP-DAC1
2015 nCode: limiting harmful writes to emerging mobile NVRAM through code swapping
Kan Zhong, Duo Liu 0002, Linbo Long, Weichen Liu 0001, Qingfeng Zhuge, Edwin H.-M. Sha
DATE3
2014 Building high-performance smartphones via non-volatile memory: The swap approach
abstract
Smartphones are getting increasingly high-performance with advances in mobile processors and larger main memories to support feature-rich applications. However, the storage subsystem has always been a prohibitive factor that slows down the pace of reaching even higher performance while maintaining good user experience. Despite today's smartphones are equipped with larger-than-ever main memories, they consume more energy and still run out of memory. But the slow NAND flash based storage vetoes the possibility of swapping---an important technique to extend main memory---and leaves a system that constantly terminates user applications under memory pressure.
Kan Zhong, Tianzheng Wang 0001, Linbo Long, Duo Liu 0002, Weichen Liu 0001, Zili Shao, Edwin H.-M. Sha
EMSOFT4
2014 A space allocation and reuse strategy for PCM-based embedded systems
Linbo Long, Duo Liu 0002, Jingtong Hu, Shouzhen Gu, Qingfeng Zhuge, Edwin H.-M. Sha
J. Syst. Archit.1
2013 A space-based wear leveling for PCM-based embedded systems
abstract
Phase change memory (PCM) has emerged as a promising candidate to replace DRAM in embedded systems. However, it can only sustain a limited number of write operations. To solve this issue, this paper proposes a novel and effective wear-leveling technique in software level to prolong the lifetime of PCM-based embedded systems. A polynomial-time algorithm, Multi-Space Wear Leveling Algorithm (MWL), is proposed to achieve effective wear-leveling. The experimental results show our technique can greatly extend the lifetime of PCM-based embedded systems compared with the previous work. Compared with the method without adopting wear-leveling, it introduces no more than 0.7% extra writes and 0.6% running overhead.
Linbo Long, Duo Liu 0002, Jingtong Hu, Shouzhen Gu, Qingfeng Zhuge, Edwin H.-M. Sha
RTCSA1