Tong Zhang 0002

dblp:07/4227-2 · DBLP profile ↗
← Back
91ranked-venue papers
6as first author
11since 2021 · last 2026
0009-0009-8005-0043ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 73 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorComputer networks · 5 · 1 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
abstract
In 1987, Jim Gray and Gianfranco Putzolu introduced the five-minute rule, a simple, storage-memory-economics-based heuristic for deciding when data should live in DRAM rather than on storage. Subsequent revisits to the rule largely retained that economics-only view, leaving host costs, feasibility limits, and workload behavior out of scope. This paper revisits the rule from first principles, integrating host costs, DRAM bandwidth/capacity, and physics-grounded models of SSD performance and cost, and then embedding these elements in a constraint- and workload-aware framework that yields actionable provisioning guidance. We show that, for modern AI platforms, especially GPU-centric hosts paired with ultra-high-IOPS SSDs engineered for fine-grained random access, the DRAM$\leftrightarrow$flash caching threshold collapses from minutes to a few seconds. This shift reframes NAND flash memory as an \emph{active data tier} and exposes a broad research space across the hardware-software stack. We further introduce MQSim-Next, a calibrated SSD simulator that supports validation and sensitivity analysis and facilitates future architectural and system research. Finally, we present two concrete case studies that showcase the software system design space opened by such memory hierarchy paradigm shift. Overall, we turn a classical heuristic into an actionable, feasibility-aware analysis and provisioning framework and set the stage for further research on AI-era memory hierarchy.
Tong Zhang 0002, Vikram S. Mailthody, Linsen Ma, Chris J. Newburn, Teresa Zhang, Jiangpeng Li, Hao Zhong 0006, Wen-Mei W. Hwu
ISCA1
2026 Towards Encrypted Data Compression with Computational Storage Drives
abstract
Modern data center systems need to achieve several critical goals in security, performance, and cost efficiency. However, realizing these goals simultaneously is highly challenging. In secure data storage systems, a common practice is to first compress and encrypt data on the host side and then transmit it to the storage system using a log-based structure. This approach, unfortunately, leads to increased complexity and performance penalty. As an emerging storage technology, Computational Storage Drives (CSD) can not only offload heavy computation burdens to storage device hardware, but also provide a virtualized logical storage space, creating new optimization opportunities. In this paper, we showcase two unique opportunities enabled by the new CSD technology in data storage management. By replacing ordinary SSDs with CSDs, we can realize efficient one-to-one mapping from host-side blocks to storage-side CSD blocks, eliminating the need for a complex log-based structure and the associated heavy-cost operations, such as garbage collections (GC). Moreover, with a carefully redesigned data format in each compression unit, CSDs can transparently remove redundant data across encrypted snapshots in data-intensive environments, such as databases. We have developed a prototype and conducted experiments on ScaleFlux’s CSD 3000 devices to demonstrate the efficacy of these solutions. We hope that our system investigations in this work provide valuable insight into CSDs and inspire researchers and practitioners to explore additional cases for adopting CSDs to improve the performance and productivity of data center systems.
Linsen Ma, Rui Xie 0006, Feng Chen 0005, Xiaodong Zhang 0001, Tong Zhang 0002
SSDBM5
2026 TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
abstract
LLM inference is increasingly limited by memory bandwidth, and the bottleneck worsens at long context as the KV cache grows. CXL memory adds capacity to offload weights and KV, but its link and device-side DDR bandwidth are far below HBM, so decoding stalls once traffic shifts to the CXL tier. Many CXL controllers are starting to add genericlosslesscompression, yet applying commodity codecs directly to standard word-major LLM tensors is largely ineffective, especially for token-major KV streams. We propose TRACE (Traffic-Reduced Architecture for Compression and Elasticity), which preserves the unmodified CXL.mem interface but changes the device-internal representation. It stores tensors in a channel-major, disaggregated bit-plane layout, and applies a KV-specific transform before compression, converting mixed-field words into low-entropy plane streams that commodity codecs can compress. The same substrate enables precision-proportional fetch by reading only the required bit-planes. Across public LLMs, TRACE reduces BF16 weight footprint by 25.2% and BF16 KV footprint by 46.9% losslessly, with per-layer KV ratios peaking at 2.69×. In tracedriven system modeling, once KV spills to CXL, GPT-OSS-120B-MXFP4 improves throughput at 128k tokens from 16.28 to 68.99 tok/s (4.24×). DRAMSim3 shows up to 40.3% lower DRAM access energy under plane-aligned fetch. A 7nm SystemVerilog implementation sustains 256 GB/s device bandwidth. Relative to a CXL controller with generic inline lossless compression, TRACE only adds 7.2% area, 4.7% power, and 6.0% load-to-use latency at 2 GHz and 0.7V.
Rui Xie 0006, Asad Ul Haq, Yunhua Fang, Linsen Ma, Zirak Burzin Engineer, Liu Liu 0017, Tong Zhang 0002
IEEE Trans. Computers7
2025 HaSiS: A Hardware-assisted Single-index Store for Hybrid Transactional and Analytical Processing
Kecheng Huang, Zhaoyan Shen, Zili Shao, Feng Chen 0005, Tong Zhang 0002
FAST5
2024 Eliminating Storage Management Overhead of Deduplication over SSD Arrays Through a Hardware/Software Co-Design
abstract
This paper presents a hardware/software co-design solution to efficiently implement block-layer deduplication over SSD arrays. By introducing complex and varying dependency over the entire storage space, deduplication is infamously subject to high storage management overheads in terms of CPU/memory resource usage and I/O performance degradation. To fundamentally address this problem, one intuitive idea is to offload deduplication storage management from host into SSDs, which is motivated by the redundant dual address mapping in host-side deduplication layer and intra-SSD flash translation layer (FTL). The practical implementation of this idea is nevertheless challenging because of the array-wide deduplication vs. per-SSD FTL management scope mismatch. Aiming to tackle this challenge, this paper presents a solution, called ARM-Dedup, that makes SSD FTL deduplication-oriented and array-aware and accordingly re-architects deduplication software to achieve lightweight and high-performance deduplication over an SSD array. We implemented an ARM-Dedup prototype based on the Linux Dmdedup engine and mdraid software RAID over FEMU SSD emulators. Experimental results show that ARM-Dedup has good scalability and can improve system performance significantly, such as by up to 272% and 127% higher IOPS in synthetic and real-world workloads, respectively.
Yuhong Wen, You Zhou 0009, Tong Zhang 0002, Shangjun Yang, Changsheng Xie 0001, Fei Wu 0005
ASPLOS (2)4
2023 ZipKV: In-Memory Key-Value Store with Built-In Data Compression
abstract
This paper studies how to mitigate the speed performance loss caused by integrating block data compression into in-memory key-value ‍(KV) stores. Despite extensive prior research on in-memory KV stores, little focus has been given to memory usage reduction via block data compression (e.g., LZ4, ZSTD) due to potential performance degradation. This paper introduces design techniques to mitigate compression-induced performance degradation by utilizing decompression streaming, latency differences between compression and decompression, and data access locality in real-world workloads. These techniques can be incorporated into conventional hash or B+-tree indexing structures, enabling integration with most in-memory KV stores without altering their core indexing data structures. For demonstration, we implemented ZipKV that incorporates the developed design techniques. Compared with RocksDB ‍(in-memory mode) that employs the log-structured merge tree indexing data structure with natural support of block data compression, ZipKV realizes similar memory usage reduction via block data compression, reduces the point query latency by 68% ‍(LZ4) and 58% ‍(ZSTD), and achieves up to 3.8× ‍(LZ4) and 2.7× ‍(ZSTD) point query throughput.
Linsen Ma, Rui Xie 0006, Tong Zhang 0002
ISMM3
2023 Elastic RAID: Implementing RAID over SSDs with Built-in Transparent Compression
abstract
This paper studies how RAID (redundant array of independent disks) could take full advantage of modern SSDs (solid-state drives) with built-in transparent compression. In current practice, RAID users are forced to choose a specific RAID level (e.g., RAID 10 or RAID 5) with a fixed storage cost vs. speed performance trade-off. The commercial market is witnessing the emergence of a new family of SSDs that can internally perform hardware-based lossless compression on each 4KB LBA (logical block address) block, transparent to host OS and user applications. Beyond straightforwardly reducing the RAID storage cost, such modern SSDs make it possible to relieve RAID users from being locked into a fixed storage cost vs. speed performance trade-off. In particular, RAID systems could opportunistically leverage higher-than-expected runtime user data compressibility to enable dynamic RAID level conversion to improve the speed performance without compromising the effective storage capacity. This paper presents techniques to enable and optimize the practical implementation of such elastic RAID systems. We implemented a Linux software-based elastic RAID prototype that supports dynamic conversion between RAID 5 and RAID 10. Compared with a baseline software-based RAID 5, under sufficient runtime data compressibility that enables the conversion from RAID 5 to RAID 10 over 60% of user data, the elastic RAID could improve the 4KB random write IOPS (I/O per second) by 42% and 4KB random read IOPS in degraded mode by 46%, while maintaining the same effective storage capacity.
Jiangpeng Li, Yang Liu 0256, Tong Zhang 0002
SYSTOR5
2022 Closing the B+-tree vs. LSM-tree Write Amplification Gap on Modern Storage Hardware with Built-in Transparent Compression
Yifan Qiao 0003, Xubin Chen, Jiangpeng Li, Yang Liu 0256, Tong Zhang 0002
FAST6
2021 KallaxDB: A Table-less Hash-based Key-Value Store on Storage Hardware with Built-in Transparent Compression
abstract
This paper studies the design of a key-value (KV) store that can take full advantage of modern storage hardware with built-in transparent compression capability. Many modern storage appliances/drives implement hardware-based data compression, transparent to OS and applications. Moreover, the growing deployment of hardware-based compression in Cloud infrastructure leads to the imminent arrival of Cloud-based storage hardware with built-in transparent compression. By decoupling the logical storage space utilization efficiency from the true physical storage usage, transparent compression allows data management software to purposely waste logical storage space in return for simpler data structures and algorithms, leading to lower implementation complexity and higher performance. This work proposes a table-less hash-based KV store, where the basic idea is to hash the key space directly onto the logical storage space without using a hash table at all. With a substantially simplified data structure, this approach is subject to significant logical storage space under-utilization, which can be seamlessly mitigated by storage hardware with transparent compression. This paper presents the basic KV store architecture, and develops mathematical formulations to assist its configuration and analysis. We implemented such a KV store KallaxDB and carried out experiments on a commercial SSD with built-in transparent compression. The results show that, while consuming very little memory resource, it compares favorably with the other modern KV stores in terms of throughput, latency, and CPU usage.
Xubin Chen, Shukun Xu, Yifan Qiao 0003, Yang Liu 0256, Jiangpeng Li, Tong Zhang 0002
DaMoN7
2021 Implementing Flash-Cached Storage Systems Using Computational Storage Drive with Built-in Transparent Compression
abstract
This paper studies utilizing the growing family of solid-state drives (SSDs) with built-in transparent compression to simplify the data structure of cache design. Such storage hardware allows the user applications to intentionally under-utilize logical storage space (i.e., sparse LBA utilization, and sparse storage block content) without sacrificing the physical storage space. Accordingly, this work proposed an index-less cache management approach to largely simplify the flash-based cache management by leveraging SSDs with built-in transparent compression. We carried out various experiments to evaluate the write amplification and read performance of the proposed cache management, and the results show that our proposed indexless cache management can achieve comparable or much better performance than the conventional policies while consuming much less host computing and memory resources.
Jingpeng Hao, Xubin Chen, Yifan Qiao 0003, Tong Zhang 0002
NAS5
2021 Improving Relational Database Upon the Arrival of Storage Hardware with Built-in Transparent Compression
abstract
This paper presents an approach to enable relational database take full advantage of modern storage hardware with built-in transparent compression. Advanced storage appliances (e.g., all-flash array) and some latest SSDs (solid-state drives) can perform hardware-based data compression, transparently from OS and applications. Moreover, the growing deployment of hardware-based compression capability in Cloud storage infrastructure leads to the imminent arrival of cloud-based storage hardware with built-in transparent compression. To make relational database better leverage modern storage hardware, we propose to deploy a dual in-memory vs. on-storage page format: While pages in database cache memory retain the conventional row-based format, each page on storage devices has a column-based format so that it can be better compressed by storage hardware. We present design techniques that can further improve the on-storage page data compressibility through additional light-weight column data transformation. We the impact of compression algorithms on the selection of column data transformation techniques. We integrated the design techniques into MySQL/InnoDB by adding only about 600 lines of code, and ran Sysbench OLTP workloads on a commercial SSD with built-in transparent compression. The results show that the proposed solution can bring up to 45% additional reduction on the storage cost at only a few percentage of performance degradation.
Yifan Qiao 0003, Xubin Chen, Jingpeng Hao, Jiangpeng Li, Qi Wu 0006, Jingqiang Wang, Yang Liu 0256, Tong Zhang 0002
NAS8
2020 POLARDB Meets Computational Storage: Efficiently Support Analytical Workloads in Cloud-Native Relational Database
Yang Liu 0256, Zhushi Cheng, Linqiang Ouyang, Ray Kuan, Zhenjun Liu, Tong Zhang 0002
FAST13
2020 Re-think Data Management Software Design Upon the Arrival of Storage Hardware with Built-in Transparent Compression
Xubin Chen, Jiangpeng Li, Qi Wu 0006, Yang Liu 0256, Hao Zhong 0006, Tong Zhang 0002
HotStorage9
2019 Mitigate HDD Fail-Slow by Pro-actively Utilizing System-level Data Redundancy with Enhanced HDD Controllability and Observability
abstract
This paper presents a design framework aiming to mitigate occasional HDD fail-slow. Due to their mechanical nature, HDDs may occasionally suffer from spikes of abnormally high internal read retry rates, leading to temporarily significant degradation of speed (especially the read latency). Intuitively, one could expect that existing system-level data redundancy (e.g., RAID or distributed erasure coding) may be opportunistically utilized to mitigate HDD fail-slow. Nevertheless, current practice tends to use system-level redundancy merely as a safety net, i.e., reconstruct data sectors via system-level redundancy only after the costly intra-HDD read retry fails. This paper shows that one could much more effectively mitigate occasional HDD fail-slow by more pro-actively utilizing existing system-level data redundancy, in complement to (or even replacement of) intra-HDD read retry. To enable this, HDDs should support a higher degree of controllability and observability in terms of their internal read retry operations. Assuming a very simple form enhanced HDD controllability and observability, this paper presents design solutions and a mathematical formulation framework to facilitate the practical implementation of such pro-active strategy for mitigating occasional HDD fail-slow. Using RAID as a test vehicle, our experimental results show that the proposed design solutions can effectively mitigate the RAID read latency degradation even when HDDs suffer from read retry rates as high as 1% or 2%.
Jingpeng Hao, Xubin Chen, Tong Zhang 0002
MSST4
2019 Reducing Flash Memory Write Traffic by Exploiting a Few MBs of Capacitor-Powered Write Buffer Inside Solid-State Drives (SSDs)
abstract
To mitigate the long write latency of NAND flash memory, solid-state drives (SSDs) typically use capacitor-powered SRAM or DRAM to realize internal nonvolatile write buffering. Due to the cost and size constraints, intra-SSD capacitors can only power a very small amount (e.g., 8 MB or 16 MB) of nonvolatile write buffer. As a result, most commercial SSDs simply use the few MBs of capacitor-powered write buffer in the first-in first-out (FIFO) manner without employing any advanced data eviction policy. This paper presents a set of design strategies across the application and storage device levels that can effectively leverage the very small intra-SSD write buffer to noticeably reduce the amount of data being physically written to NAND flash memory. These cross-layer design strategies are primarily geared towards mainstream applications (e.g., database and filesystem) that heavily involve logging/journaling operations. This paper discusses different strategies for realizing flash memory write traffic reduction through nominal application-level modifications, and presents solutions to accordingly manage the write buffer at small processing and memory resource usage inside SSDs. To evaluate the potential effectiveness, we carried out case studies based upon popular open-source relational databases and filesystem. With only 8 MB of intra-SSD capacitor-powered write buffer, our experimental results show that the developed design solutions can reduce up to 39.7, 36.5, and 52.9 percent of total NAND flash memory write traffic for MySQL, ext4, and PostgreSQL, respectively.
Xubin Chen, Tong Zhang 0002
IEEE Trans. Computers3
2019 An Exploratory Study on Software-Defined Data Center Hard Disk Drives
abstract
This article presents a design framework aiming to reduce mass data storage cost in data centers. Its underlying principle is simple: Assume one may noticeably reduce the HDD manufacturing cost by significantly (i.e., at least several orders of magnitude) relaxing raw HDD reliability, which ensures the eventual data storage integrity via low-cost system-level redundancy. This is called system-assisted HDD bit cost reduction. To better utilize both capacity and random IOPS of HDDs, it is desirable to mix data with complementary requirements on capacity and random IOPS in each HDD. Nevertheless, different capacity and random IOPS requirements may demand different raw HDD reliability vs. bit cost trade-offs and hence different forms of system-assisted bit cost reduction. This article presents a software-centric design framework to realize data-adaptive system-assisted bit cost reduction for data center HDDs. Implementation is solely handled by the filesystem and demands only minor change of the error correction coding (ECC) module inside HDDs. Hence, it is completely transparent to all the other components in the software stack (e.g., applications, OS kernel, and drivers) and keeps fundamental HDD design practice (e.g., firmware, media, head, and servo) intact. We carried out analysis and experiments to evaluate its implementation feasibility and effectiveness. We integrated the design techniques into ext4 to further quantitatively measure its impact on system speed performance.
Xubin Chen, Jingpeng Hao, Tong Zhang 0002
ACM Trans. Storage5
2018 Realizing Low-Cost Flash Memory Based Video Caching in Content Delivery Systems
abstract
To implement caching devices in content delivery systems, flash memory is preferable to hard disk drives from the performance perspective. Nevertheless, the higher bit cost of flash memory is one major obstacle for the wide real-life deployment of flash-based video caching. This paper presents a set of design solutions to address this cost issue. First, we present a flash memory error tolerance design strategy customized for video data storage, which can enable the use of lower-cost less-reliable flash memory chips for video storage. The cost challenge can also be addressed by reducing the video storage footprint through on-the-fly transcoding. However, direct transcoding suffers from a high implementation cost. We propose two design techniques that can largely reduce the transcoding complexity at minimal storage overhead in flash memory. All the developed design solutions share the common feature of cohesively exploring the characteristics of video coding and flash memory device physics. Their effectiveness has been well demonstrated through experiments with 20-nm MLC NAND flash memory chips and extensive simulations with representative video sequences.
Danni Xiong, Kai Zhao 0005, Chang Wen Chen, Tong Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.5
2018 CrowdDBS: A Crowdsourced Brightness Scaling Optimization for Display Energy Reduction in Mobile Video
abstract
Mobile display has become one of the most power-hungry components in mobile video viewing. Currently, mobile devices can reduce the display energy by performing dynamic brightness scaling (DBS) under the distortion constraint of video signals. We observe that there is a pitfall preventing current practice from systematic display energy reduction. In particular, existing objective DBS schemes lack direct connection to the subjective human perception on DBS-enabled videos, which is the key to achieving human-centered energy-experience optimization. To overcome this pitfall, we present CrowdDBS, a crowdsourced display energy reduction framework for mobile video viewing. CrowdDBS is empowered by a set of crowdsourcing studies that uncover the relationship between human perception and DBS frequency, magnitude, and temporal consistency, respectively. Motivated by the insights obtained from these studies, CrowdDBS employs a suit of designs and a DBS optimization framework to optimize the energy-experience tradeoff in mobile video viewing. Comprehensive experimental results and user evaluations under a variety of practical settings show that CrowdDBS can achieve 37 percent device energy reduction on average while guaranteeing satisfactory user experience in mobile video viewing.
Zhisheng Yan, Qian Liu 0001, Tong Zhang 0002, Chang Wen Chen
IEEE Trans. Mob. Comput.3
2017 Facilitating Magnetic Recording Technology Scaling for Data Center Hard Disk Drives through Filesystem-Level Transparent Local Erasure Coding
Hao Wang 0042, Shafa Dahandeh, Tong Zhang 0002
FAST6
2017 Software Support Inside and Outside Solid-State Devices for High Performance and High Efficiency
abstract
In the past decade, flash memory has been in the spotlight across a variety of research communities from circuits to computer systems, and significant progress has been accomplished. This has enabled flash memory to become increasingly pervasive across the entire information technology infrastructure, from consumer electronics to cloud and supercomputing. This paper aims to provide a comprehensive survey on the important advancements and milestones in the domains across flash translation layer (FTL), operating systems, and applications. As the storage device hardware has been quickly commoditized, software becomes increasingly important to tap the potential of flash memory to its full extent. Therefore, a comprehensive survey with a focus on software aspects will be very valuable to the research community and industry. It is our hope that this survey paper will serve as a good reference for system practitioners and researchers.
Feng Chen 0005, Tong Zhang 0002, Xiaodong Zhang 0001
Proc. IEEE2
2017 Improving 3D DRAM Fault Tolerance Through Weak Cell Aware Error Correction
abstract
Although the emerging 3D DRAM products can significantly improve the computing system performance, the relatively high cost is one of the most critical issues that prevent their wide real-life adoption. Intuitively, a strong memory fault tolerance can be leveraged to reduce the fabrication cost of DRAM dies, and the total cost will reduce if the fabrication cost saving can off-set the cost overhead of memory fault tolerance. Nevertheless, such a simple concept can be a practically viable option only for 3D DRAM because: (1) The stacked logic die can solely implement memory fault tolerance inside 3D DRAM chips, obviating any changes on the host CPUs and CPU-DRAM interfaces. (2) With the total ownership of both the logic die and DRAM dies inside 3D DRAM chips, DRAM manufacturers can fully exploit the potential to truly minimize the 3D DRAM bit cost. Following this intuition, we developed a 3D DRAM fault tolerance design strategy. It can achieve a very strong tolerance to weak DRAM cells at very small redundancy and latency overhead. The key is to cohesively leverage the detectability of weak cells and runtime configurability of error correction code (ECC) decoding. In addition, this design strategy can gracefully embrace the inaccuracy of weak cell detection (e.g., weak cell miss-detection and false-detection). We carried out thorough mathematical analysis, and the results show that, under the redundancy overhead of 1:8 (same as today's ECC DIMM), this design strategy can tolerate the weak cell rate of as high as 10-4 and 6x10-5 if 100 and 90 percent of all the weak cells are known in prior. Using Micron's hybrid memory cube (HMC) 3D DRAM chips as the test vehicle, we evaluated the implementation cost and the results show that it only consumes less than 0.4 mm2 (45 nm node) on the logic die. Using CPU and DRAM simulators, we further carried out simulations over a variety of computing benchmarks and the results show that this design solution only incurs less than 2 percent performance degradation on average.
Hao Wang 0042, Kai Zhao 0005, Minjie Lv, Hongbin Sun 0001, Tong Zhang 0002
IEEE Trans. Computers6
2017 Realizing Transparent OS/Apps Compression in Mobile Devices at Zero Latency Overhead
abstract
Motivated by the significant storage footprint of OS/Apps in mobile devices, this paper studies the realization of OS/Apps transparent compression. In spite of its obvious advantage, this feature however is not widely available in commercial mobile devices, which is due to the justifiable concern on the read latency penalty. In conventional practice on implementing transparent compression, read latency overhead comes from two aspects, including read amplification and decompression computational latency. This paper presents simple yet effective design solutions to eliminate the read amplification at the filesystem level and eliminate the computational latency overhead at the computer architecture level. To demonstrate its practical feasibility, we first implemented a prototyping filesystem to empirically verify the realization of transparent compression with zero read amplification. We further demonstrated that the OS/Apps footprint can be reduced by up to 39 percent on a Nexus 7 tablet installed with Android 5.0. Through application-specific integrated circuit (ASIC) synthesis, we show that the proposed computer architecture level design solution can eliminate the decompression latency overhead with very small silicon cost.
Jiangpeng Li, Hao Wang 0042, Danni Xiong, Jerry Qu, Hyunsuk Shin, Jung Pill Kim, Tong Zhang 0002
IEEE Trans. Computers8
2016 Reducing Solid-State Storage Device Write Stress through Opportunistic In-place Delta Compression
Jiangpeng Li, Hao Wang 0042, Kai Zhao 0005, Tong Zhang 0002
FAST5
2016 Guest Editorial Channel Modeling, Coding and Signal Processing for Novel Physical Memory Devices and Systems
abstract
The digital universe is doubling every two years and expected to reach an unwieldy 44 zettabytes into the next decade. To cope with the ever increasing need for storing, transmitting and retrieving huge amounts of data, cloud storage, data centers and other massively distributed storage networks have emerged. These rely on efficient memory technologies at the physical level for speed, reliability and energy efficiency.
Shayan Garani Srinivasa, Tong Zhang 0002, Ravi Motwani, Haralampos Pozidis, Bane Vasic
IEEE J. Sel. Areas Commun.2
2016 Exploiting Intracell Bit-Error Characteristics to Improve Min-Sum LDPC Decoding for MLC NAND Flash-Based Storage in Mobile Device
abstract
A multilevel per cell (MLC) technique significantly improves the storage density, but also poses serious data integrity challenge for NAND flash memory. This consequently makes the low-density parity-check (LDPC) code and the soft-decision memory sensing become indispensable in the next-generation flash-based solid-state storage devices. However, the use of LDPC codes inevitably increases memory read latency and, hence, degrades speed performance. Motivated by the observation of intracell unbalanced bit error probability and data dependence in the MLC NAND flash memory, this paper proposes two techniques, i.e., intracell data placement interleaving and intracell data dependence aware LDPC decoding, to efficiently improve the LDPC decoding throughput and energy efficiency for the MLC NAND flash-based storage in a mobile device. Experimental results show that, by exploiting the intracell bit-error characteristics, the proposed techniques together can improve the LDPC decoding throughput by up to 84.6% and reduce the energy consumption by up to 33.2% while only incurring less than 0.2% silicon area overhead.
Hongbin Sun 0001, Minjie Lv, Guiqiang Dong, Nanning Zheng 0001, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.6
2016 An Information Theory Perspective for the Binary STT-MRAM Cell Operation Channel
abstract
Spin-torque transfer magnetic random access memory (STT-MRAM) has emerged as a promising nonvolatile memory technology, with advantages, such as scalability, speed, endurance, and power consumption. This paper presents an STT-MRAM cell operation channel model with write and read operations for information theorists and error correction code designers. This model considers the effects of process variations and thermal fluctuations and considers all principle flaws during the fabrication and operation processes. With this model, evaluations are not only made for the write channel, the read channel, but also the write and read channel with metrics, such as operation failure rate, bit error rate, channel ergodic capacity, and channel outage probability at certain outage capacity. Moreover, it is proved that the distributions of written-in bit states are not uniformly distributed and are proportional to their respective write success probabilities. Finally, simulation results show that practical code rates and code block lengths can guarantee reliable performances only if the operation success rate difference between state 1 and state 0 is small enough.
Jianxiao Yang, Benoit Geller, Meng Li 0012, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.4
2015 How Much Can Data Compressibility Help to Improve NAND Flash Memory Lifetime?
Jiangpeng Li, Kai Zhao 0005, Jun Ma 0012, Tong Zhang 0002
FAST6
2015 Leveraging Progressive Programmability of SLC Flash Pages to Realize Zero-overhead Delta Compression for Metadata Storage
Jiangpeng Li, Kai Zhao 0005, Hao Wang 0042, Tong Zhang 0002
HotStorage5
2015 Exploring QoE for Power Efficiency: A Field Study on Mobile Videos with LCD Displays
abstract
Display power consumption has become a major concern for both mobile users and design engineers, especially considering the prevalence of today's video-rich mobile services. The power consumption of liquid crystal display (LCD), a dominant mobile display technology, can be reduced by dynamic backlight scaling (DBS). However, such dynamic changes of screen brightness may degrade users' quality of experience (QoE) in viewing videos. How would QoE be impacted by different DBS strategies has not yet been understood clearly and thus obscures the way to achieve systematic power saving. In this paper, we take a first step to explore the QoE of DBS on smartphones and aim at maximally enhancing the display power performance without negatively impacting users' QoE. In particular, we conduct three motivational studies to uncover the inherent relationship between QoE and backlight scaling frequency, magnitude, and temporal consistency, respectively. Motivated by the findings of these studies, we design a suite of techniques to implement a comprehensive DBS strategy. We demonstrate an example application of the proposed DBS designs in a mobile video streaming system. Measurements and user evaluations show that more than 40% system power reduction, or equivalently, 20% more power savings than the non-QoE approaches, can be achieved without QoE impairment.
Zhisheng Yan, Qian Liu 0001, Tong Zhang 0002, Chang Wen Chen
ACM Multimedia3
2015 True-Damage-Aware Enumerative Coding for Improving nand Flash Memory Endurance
abstract
This brief presents a technique that can fully exploit the data dependency of flash memory cell damage to improve the program/erase (P/E) cycling endurance of nand flash memory. The key is to opportunistically leverage data lossless compressibility and utilize the compression gain to realize memory-damage-aware data manipulation to reduce the cycling-induced physical damage. Based upon experiments using commercial sub-22-nm MLC nand flash memory chips, we show that the proposed design technique can improve the P/E cycling endurance by 50%. We further carried out application-specific integrated circuit design to demonstrate the practical feasibility for implementing the proposed design technique.
Jiangpeng Li, Kai Zhao 0005, Jun Ma 0012, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.4
2015 Optimizing the Use of STT-RAM in SSDs Through Data-Dependent Error Tolerance
abstract
This brief presents a design strategy for spin-transfer torque (STT)-RAM to reduce the error-tolerance redundancy overhead and increase effective storage capacity without sacrificing its reliability. The key is to cohesively exploit the run-time data characteristics (e.g., access unit length and access frequency) and the fundamental read disturbance versus sensing error tradeoff in STT-RAM. It presents three specific data-dependent error-tolerance design techniques, and demonstrates their effectiveness in the context of using STT-RAM to replace DRAM in solid-state drives. Based on detailed modeling/simulations down to 22-nm node, we showed that these design solutions can increase the effective STT-RAM storage capacity by 26%, compared with conventional design practice.
Hao Wang 0042, Kai Zhao 0005, Jiangpeng Li, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.4
2014 Over-clocked SSD: Safely running beyond flash memory chip I/O clock specs
abstract
This paper presents a design strategy that enables aggressive use of flash memory chip I/O link over-clocking in solid-state drives (SSDs) without sacrificing storage reliability. The gradual wear-out and process variation of NAND flash memory makes the worst-case oriented error correction code (ECC) in SSDs largely under-utilized most of the time. This work proposes to opportunistically leverage under-utilized error correction strength to allow error-prone flash memory I/O link over-clocking. Its rationale and key design issues are presented and studied in this paper, and its potential effectiveness has been verified through hardware experiments and system simulations. Using sub-22nm NAND flash memory chips with I/O specs of 166MBps, we carried out extensive experiments and show that the proposed design strategy can enable SSDs safely operate with error-prone I/O link running at 275MBps. Trace-driven SSD simulations over a variety of workload traces show the system read response time can be reduced by over 20%.
Kai Zhao 0005, Kalyana S. Venkataraman, Jiangpeng Li, Tong Zhang 0002
HPCA6
2014 Proximate control stream assisted video transcoding for heterogeneous content delivery network
abstract
Video transcoding can be used to facilitate video streaming in content delivery network. A concept of control stream assisted transcoding has been recently proposed aiming to reduce transcoding computational complexity at the cost of data storage/transmission overhead, and the control stream is obtained by directly removing the residual information from the target video bitstream of transcoding. However, it is subject to relatively significant storage/transmission overhead, and more importantly does not well match to the increasingly heterogeneous networking environment with varying computation/storage/transmission resources at different nodes. This work presents a proximate control stream assisted transcoding design strategy that can reduce the control stream size and enable a large storage/transmission vs. computational complexity trade-off design space. Experiments demonstrate its effectiveness and noticeable advantages over other alternatives including simulcast and SVC (scalable video coding).
Jiangpeng Li, Kai Zhao 0005, Tong Zhang 0002
ICIP5
2014 Improving min-sum LDPC decoding throughput by exploiting intra-cell bit error characteristic in MLC NAND flash memory
abstract
Multi-level per cell (MLC) technique significantly improves storage density, but also poses new challenge to data integrity in NAND flash memory. Therefore, low-density parity-check (LDPC) code and soft-decision memory sensing have become indispensable in future NAND flash-based solid state drive design. However, these more powerful technologies inevitably increase the memory read latency and hence degrade the decoding throughput. Motivated by intra-cell unbalanced bit error probability and data dependency in MLC NAND flash memory, this paper proposes two techniques, i.e. intra-cell data placement interleaving and intra-cell data dependency aware min-sum decoding, to effectively improve the throughput of LDPC decoding. Experimental results show that, the proposed techniques used in an integrated way can improve the LDPC decoding throughput by up to 85% when the MLC NAND flash chip is heavily cycled, compared with conventional design practice.
Hongbin Sun 0001, Minjie Lv, Guiqiang Dong, Nanning Zheng 0001, Tong Zhang 0002
MSST6
2014 Enhanced Precision Through Multiple Reads for LDPC Decoding in Flash Memories
abstract
Multiple reads of the same Flash memory cell with distinct word-line voltages provide enhanced precision for LDPC decoding. In this paper, the word-line voltages are optimized by maximizing the mutual information (MI) of the quantized channel. The enhanced precision from a few additional reads allows frame error rate (FER) performance to approach that of full-precision soft information and enables an LDPC code to significantly outperform a BCH code. A constant-ratio constraint provides a significant simplification in the optimization with no noticeable loss in performance. For a well-designed LDPC code, the quantization that maximizes the mutual information also minimizes the FER in our simulations. However, for an example LDPC code with a high error floor caused by small absorbing sets, the MMI quantization does not provide the lowest frame error rate. The best quantization in this case introduces more erasures than would be optimal for the channel MI in order to mitigate the absorbing sets of the poorly designed code. The paper also identifies a trade-off in LDPC code design when decoding is performed with multiple precision levels; the best code at one level of precision will typically not be the best code at a different level of precision.
Kasra Vakilinia, Tsung-Yi Chen, Thomas A. Courtade, Guiqiang Dong, Tong Zhang 0002, Hari Shankar, Richard D. Wesel
IEEE J. Sel. Areas Commun.6
2014 OFWAR: Reducing SSD Response Time Using On-Demand Fast-Write-and-Rewrite
abstract
This paper presents a cross-layer design strategy to reduce SSD response time and its variation. The key is to cohesively exploit system-level run-time data access workload variation and temporal locality and device-level NAND flash memory write latency versus data retention time trade-off. The basic idea is simple: once write intensity of the workload increases and begins to degrade SSD response time, we speed up memory programming at the penalty of shorter data retention time, and rewrite these short-lifetime data later, if necessary, when workload write intensity drops. A scheduling solution is developed to effectively implement this design strategy. Simulations over various workloads were carried out and the results demonstrate that the cross-layer design strategy can reduce the average SSD response time by up to 52.3%.
Qi Wu 0006, Tong Zhang 0002
IEEE Trans. Computers2
2014 Using Lifetime-Aware Progressive Programming to Improve SLC NAND Flash Memory Write Endurance
abstract
This paper advocates a lifetime-aware progressive programming concept to improve single-level per cell NAND flash memory write endurance. NAND flash memory program/erase (P/E) cycling gradually degrades memory cell storage noise margin, and sufficiently strong fault tolerance must be used to ensure the memory P/E cycling endurance. As a result, the relatively large cell storage noise margin in early memory lifetime is essentially wasted in conventional design practice. This paper proposes to always fully utilize the available cell storage noise margin by adaptively adjusting the number of storage levels per cell, and progressively use these levels to realize multiple 1-bit programming operations between two consecutive erase operations. This simple progressive programming design concept is realized by two different implementation strategies, which are discussed and compared in detail. On the basis of an approximate NAND flash memory device model, we carried out simulations to quantitatively evaluate this design concept. The results show that it can improve the write endurance by 35.9% and in the meanwhile improve the average programming speed by 12% without sacrificing read speed.
Guiqiang Dong, Yangyang Pan, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2013 LDPC-in-SSD: making advanced error correction codes work effectively in solid state drives
Kai Zhao 0005, Hongbin Sun 0001, Tong Zhang 0002, Xiaodong Zhang 0001, Nanning Zheng 0001
FAST4
2013 Scheduling Algorithms for Handling Updates in Shingled Magnetic Recording
abstract
Shingled recording has recently emerged as one promising candidate to sustain the historical growth of magnetic recording storage areal density. However, since the convenient update-in-place feature is no longer available in shingled recording, many sectors must be read and written back in order to update one sector. This leads to a significant update-induced latency overhead and makes conventional hard disk drive scheduling algorithms perform poorly. This paper concerns with the development of appropriate scheduling algorithms for shingled recording based hard disk drives. We first present a simple partial-update scheduling algorithm that can naturally embrace the update latency issue and achieves significant gains over conventional scheduling algorithms. We enhance this algorithm by incorporating a shortest update first policy, which can further reduce the update response time on an average by 70%. Finally, motivated by abundant workload spatial and temporal locality, we develop a spatio-temporal band coalescing scheme that can achieve an additional reduction of update response time of up to 96.8%.
Kalyana Sundaram Venkataraman, Tong Zhang 0002, Hongbin Sun 0001, Nanning Zheng 0001
NAS2
2013 Using Quasi-EZ-NAND Flash Memory to Build Large-Capacity Solid-State Drives in Computing Systems
abstract
Future flash-based solid-state drives (SSDs) must employ increasingly powerful error correction code (ECC) and digital signal processing (DSP) techniques to compensate the negative impact of technology scaling on NAND flash memory device reliability. Currently, all the ECC and DSP functions are implemented in a central SSD controller. However, the use of more powerful ECC and DSP makes such design practice subject to significant speed performance degradation and complicated controller implementation. An EZ-NAND (Error Zero NAND) flash memory design strategy is emerging in the industry, which moves all the ECC and DSP functions to each memory chip. Although EZ-NAND flash can simplify controller design and achieve high system speed performance, its high silicon cost may not be affordable for large-capacity SSDs in computing systems. We propose a quasi-EZ-NAND design strategy that hierarchically distributes ECC and DSP functions on both NAND flash memory chips and the central SSD controller. Compared with EZ-NAND design concept, it can maintain almost the same speed performance while reducing silicon cost overhead. Assuming the use of low-density parity-check (LDPC) code and postcompensation DSP technique, trace-based simulations show that SSDs using quasi-EZ-NAND flash can realize almost the same speed as SSDs using EZ-NAND flash, and both can reduce the average SSD response time by over 90 percent compared with conventional design practice. Silicon design at 65 nm node shows that quasi-EZ-NAND can reduce the silicon cost overhead by up to 44 percent compared with EZ-NAND.
Yangyang Pan, Guiqiang Dong, Ningde Xie, Tong Zhang 0002
IEEE Trans. Computers4
2013 Using Multilevel Phase Change Memory to Build Data Storage: A Time-Aware System Design Perspective
abstract
This paper advocates a time-aware design methodology for using multilevel per cell (MLC) phase-change memory (PCM) in data storage systems such as solid-state disk and disk cache. It is well known that phase-change material resistance drift gradually reduces memory device noise margin and degrades the raw storage reliability. Intuitively, due to the time-dependent nature of resistance drift, if we can dynamically adjust storage system operations adaptive to the time and, hence, memory cell resistance drift, we may improve various PCM-based data storage system performance metrics. Under such an intuitive time-aware system design concept, we propose three specific design techniques, including time-aware variable-strength error correction code (ECC) decoding, time-aware partial rewrite, and time-aware read-&-refresh. Since PCM-based data storage systems have to use powerful ECC whose decoding can be energy-hungry, the first technique aims to minimize the ECC decoding energy consumption. The second technique improves the data retention limit when using partial rewrite in MLC PCM, and the third technique can further improve the efficiency of time-aware variable-strength ECC decoding. Using hypothetical 2-bit/cell PCM with device parameters from recent device research as a test vehicle, we carry out mathematical analysis and trace-based simulations, which show that these techniques can improve the data retention limit by few orders of magnitude, and enable up to 97 and 79 percent energy savings for PCM-based solid-state disk and PCM-based disk cache.
Qi Wu 0006, Wei Xu 0021, Tong Zhang 0002
IEEE Trans. Computers4
2013 Exploiting workload dynamics to improve SSD read latency via differentiated error correction codes
abstract
This article presents a cross-layer codesign approach to reduce SSD read response latency. The key is to cohesively exploit the NAND flash memory device write speed vs. raw storage reliability trade-off at the physical layer and runtime data access workload dynamics at the system level. Leveraging runtime data access workload variation, we can opportunistically slow down NAND flash memory write speed and hence improve NAND flash memory raw storage reliability. This naturally enables an opportunistic use of weaker error correction schemes that can directly reduce SSD read access latency. We develop a disk-level scheduling scheme to effectively smooth the write workload in order to maximize the occurrence of runtime opportunistic NAND flash memory write slowdown. Using 2 bits/cell NAND flash memory with BCH-based error correction correction as a test vehicle, we carry out extensive simulations over various workloads and demonstrate that this developed cross-layer co-design solution can reduce the average SSD read latency by up to 59.4% without sacrificing the write throughput performance.
Guanying Wu, Xubin He, Ningde Xie, Tong Zhang 0002
ACM Trans. Design Autom. Electr. Syst.4
2013 Error Rate-Based Wear-Leveling for nand Flash Memory at Highly Scaled Technology Nodes
abstract
This brief presents a NAND Flash memory wear-leveling algorithm that explicitly uses memory raw bit error rate (BER) as the optimization target. Although NAND Flash memory wear-leveling has been well studied, all the existing algorithms aim to equalize the number of programming/erase cycles among all the memory blocks. Unfortunately, such a conventional design practice becomes increasingly suboptimal as inter-block variation becomes increasingly significant with the technology scaling. This brief presents a dynamic BER-based greedy wear-leveling algorithm that uses BER statistics as the measurement of memory block wear-out pace, and guides dynamic memory block data swapping to fully maximize the wear-leveling efficiency. Simulations have been carried out to quantitatively demonstrate its advantages over existing wear-leveling algorithms.
Yangyang Pan, Guiqiang Dong, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2013 Exploring the Use of Emerging Nonvolatile Memory Technologies in Future FPGAs
abstract
As new nonvolatile memory technologies become increasingly mature, there has been a growing interest on investigating their use in future field-programmable gate arrays (FPGAs). Similar to existing FPGAs with embedded Flash memory, future FPGAs can embed these new nonvolatile memories to persistently store configuration data. By comparing with prior work, we first propose the more appropriate design style for new nonvolatile configuration data storage memory. Moreover, this brief studies a dynamic random-access memory (DRAM)-based FPGA design strategy enabled by high-density embedded nonvolatile memory. Existing FPGAs do not use on-chip DRAM cells for configuration data storage mainly because DRAM self-refresh involves destructive DRAM read. This problem can be solved, if we use embedded nonvolatile memory as primary FPGA configuration data storage and externally refresh on-chip DRAM cells. Analysis and simulations have been carried out to demonstrate the potential advantages of such a design strategy.
Yangyang Pan, Yiran Li 0001, Hongbin Sun 0001, Wei Xu 0021, Nanning Zheng 0001, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.6
2012 Quasi-nonvolatile SSD: Trading flash memory nonvolatility to improve storage system performance for enterprise applications
abstract
This paper advocates a quasi-nonvolatile solid-state drive (SSD) design strategy for enterprise applications. The basic idea is to trade data retention time of NAND flash memory for other system performance metrics including program/erase (P/E) cycling endurance and memory programming speed, and meanwhile use explicit internal data refresh to accommodate very short data retention time (e.g., few weeks or even days). We also propose SSD scheduling schemes to minimize the impact of internal data refresh on normal I/O requests. Based upon detailed memory cell device modeling and SSD system modeling, we carried out simulations that clearly show the potential of using this simple quasi-nonvolatile SSD design strategy to improve system cycling endurance and speed performance. We also performed detailed energy consumption estimation, which shows the energy consumption overhead induced by data refresh is negligible.
Yangyang Pan, Guiqiang Dong, Qi Wu 0006, Tong Zhang 0002
HPCA4
2012 Reducing data transfer latency of NAND flash memory with soft-decision sensing
abstract
With the aggressive technology scaling and use of multi-bit per cell storage, NAND flash memory is subject to continuous degradation of raw storage reliability and demands more and more powerful error correction codes (ECC). This inevitable trend makes conventional BCH code increasingly inadequate, and iterative coding solutions such as LDPC codes become very natural alternative options. However, these powerful coding solutions demand soft-decision memory sensing, which results in longer on-chip memory sensing latency and memory-to-controller data transfer latency. This paper presents two simple design techniques that can reduce the memory-to-controller data transfer latency. The key is to appropriately apply entropy coding to compress the memory sensing results. Simulation results show that the proposed design solutions can reduce the data transfer latency by up to 64% for soft-decision memory sensing.
Guiqiang Dong, Yuelin Zou, Tong Zhang 0002
ICC3
2012 Reducing DRAM Image Data Access Energy Consumption in Video Processing
abstract
This paper presents domain-specific techniques to reduce DRAM energy consumption for image data access in video processing. In mobile devices, video processing is one of the most energy-hungry tasks, and DRAM image data access energy consumption becomes increasingly dominant in overall video processing system energy consumption. Hence, it is highly desirable to develop domain-specific techniques that can exploit unique image data access characteristics to improve DRAM energy efficiency. Nevertheless, prior efforts on reducing DRAM energy consumption in video processing pale in comparison with that on reducing video processing logic energy consumption. In this work, we first apply three simple yet effective data manipulation techniques that exploit image data spatial/temporal correlation to reduce DRAM image data access energy consumption, then propose a heterogeneous DRAM architecture that can better adapt to unbalanced image access in most video processing to further improve DRAM energy efficiency. DRAM modeling and power estimation have been carried out to evaluate these domain-specific design techniques, and the results show that they can reduce DRAM energy consumption by up to 92%.
Yiran Li 0001, Tong Zhang 0002
IEEE Trans. Multim.2
2012 Estimating Information-Theoretical nand Flash Memory Storage Capacity and its Implication to Memory System Design Space Exploration
abstract
Today and future NAND flash memory will heavily rely on system-level fault-tolerance techniques such as error correction code (ECC) to ensure the overall system storage integrity. Since ECC demands the storage of coding redundancy and hence degrades effective cell storage efficiency, it is highly desirable to use more powerful coding solutions that can maintain the system storage reliability at less coding redundancy. This has motivated a growing interest in the industry to search for alternatives to BCH code being used in today. Regardless to specific ECCs, it is of great practical importance to know the theoretical limit on the achievable cell storage efficiency, which motivates this work. We first develop an approximate NAND flash memory channel model that explicitly incorporates program/erase (P/E) cycling effects and cell-to-cell interference, based on which we then develop strategies for estimating the information-theoretical bounds on cell storage efficiency. We show that it can readily reveal the tradeoffs among cell storage efficiency, P/E cycling endurance, and retention limit, which can provide important insights for system designers. Finally, motivated by the dynamics of P/E cycling effect revealed by the information-theoretical study, we propose two memory system design techniques that can improve the average NAND flash memory programming speed and increase the total amount of user data that can be stored in NAND flash cell over its entire lifetime.
Guiqiang Dong, Yangyang Pan, Ningde Xie, Chandra Varanasi, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.5
2012 Using Magnetic RAM to Build Low-Power and Soft Error-Resilient L1 Cache
abstract
Due to its great scalability, fast read access, low leakage power, and nonvolatility, magnetic random access memory (MRAM) appears to be a promising memory technology for on-chip cache memory in microprocessors. However, the write-to-MRAM process is relatively slow and results in high dynamic power consumption. Such inherent disadvantages of MRAM make researchers easily conclude that MRAM can only be used for low-level caches (e.g., L2 or L3 cache), where cache memories are less frequently accessed and slow write to MRAM can be more easily compensated using simple architectural techniques. By developing a hybrid cache architecture, this paper attempts to show that, with appropriate architecture design, MRAM can also be used in L1 cache to improve both the energy efficiency and soft error immunity. The basic idea is to supplement the MRAM L1 cache with several small SRAM buffers, which can substantially mitigate the performance degradation and dynamic energy overhead induced by MRAM write operations. Moreover, the proposed hybrid cache architecture is also an efficient solution to protect cache memory from radiation-induced soft errors, as MRAM is inherently invulnerable to emissive particles. Simulation results show that, with only less than 2% performance degradation, the proposed design approach can reduce the power consumption by up to 76.1% on average compared with the traditional SRAM L1 cache. In addition, the architectural vulnerability factor of L1 data cache is reduced from 28.3% to as low as 0.5%.
Hongbin Sun 0001, Chuanyin Liu, Wei Xu 0021, Jizhong Zhao, Nanning Zheng 0001, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.6
2011 Exploiting Memory Device Wear-Out Dynamics to Improve NAND Flash Memory System Performance
Yangyang Pan, Guiqiang Dong, Tong Zhang 0002
FAST3
2011 Design techniques to improve the device write margin for MRAM-based cache memory
abstract
As one promising non-volatile memory technology, magnetoresistive RAM (MRAM) based on magnetic tunneling junctions (MTJs) has recently attracted much attention. However, latest device research has discovered that, in order to maintain sufficient MTJ write margin to prevent device breakdown, MTJs will be subject to unconventionally high random write error rates (e.g., 10-3 and above) as memory cell size is being scaled down. This new discovery seriously threatens the scalability of MRAM, and the material/device research community is actively searching for solutions to largely reduce MTJ write error rates and meanwhile maintain sufficient device write margin. In this paper, we attempt to address this challenge from the architecture level when using MRAM to implement cache memory. In particular, we show that two simple cache architecture design techniques can be used to effectively tolerate high MTJ write error rates at small performance and implementation cost, which makes it much easier to maintain sufficient MTJ write margin and hence push the MRAM scalability envelope. Using the full system simulator PTLsim and a variety of benchmarks, we show that the proposed design techniques can readily accommodate MTJ write error rate up to 0.75% at the penalty of less than 4% processor performance degradation, less than 10% silicon area overhead, and 6% energy consumption overhead.
Hongbin Sun 0001, Chuanyin Liu, Nanning Zheng 0001, Tai Min, Tong Zhang 0002
ACM Great Lakes Symposium on VLSI5
2011 Exploiting Heat-Accelerated Flash Memory Wear-Out Recovery to Enable Self-Healing SSDs
Qi Wu 0006, Guiqiang Dong, Tong Zhang 0002
HotStorage3
2011 Using Lossless Data Compression in Data Storage Systems: Not for Saving Space
abstract
Lossless data compression for data storage has become less popular as mass data storage systems are becoming increasingly cheap. This leaves many files stored on mass data storage media uncompressed although they are losslessly compressible. This paper proposes to exploit the lossless compressibility of those files to improve the underlying storage system performance metrics such as energy efficiency and access speed, other than saving storage space as in conventional practice. The key idea is to apply runtime lossless data compression to enable an opportunistic use of a stronger error correction code (ECC) with more coding redundancy in data storage systems, and trade such opportunistic extra error correction capability to improve other system performance metrics in the runtime. Since data storage is typically realized in the unit of equal-sized sectors (e.g., 512 B or 4 KB user data per sector), we only apply this strategy to each individual sector independently in order to be completely transparent to the firmware, operating systems, and users. Using low-density parity check (LDPC) code as ECC in storage systems, this paper quantitatively studies the effectiveness of this design strategy in both hard disk drives and NAND flash memories. For hard disk drives, we use this design strategy to reduce average hard disk drive read channel signal processing energy consumption, and results show that up to 38 percent read channel energy saving can be achieved. For NAND flash memories, we use this design strategy to improve average NAND flash memory write speed, and results show that up to 36 percent write speed improvement can be achieved for 2 bits/cell NAND flash memories.
Ningde Xie, Guiqiang Dong, Tong Zhang 0002
IEEE Trans. Computers3
2011 Design Techniques to Facilitate Processor Power Delivery in 3-D Processor-DRAM Integrated Systems
abstract
As a promising option to address the memory wall problem, 3-D processor-DRAM integration has recently received many attentions. Since DRAM dies should be stacked between the processor die and package substrate, we have to fabricate a large number of through-DRAM through-silicon vias (TSVs) to connect the processor die and package for power and input/output (I/O) signal delivery. Although such through-DRAM TSVs will inevitably interfere with DRAM design and induce non-negligible power consumption overhead, little prior research has been done to study how to allocate these through-DRAM TSVs on the DRAM dies and analyze their impacts. To address this open issue, this paper first presents a through-DRAM TSV allocation strategy that well fits to the regular DRAM architecture. Meanwhile, due to the longer path between power/ground pads and processor die, power delivery integrity issue may become more serious in such 3-D processor-DRAM integrated systems. Decoupling capacitor insertion is the most popular method to deal with power delivery integrity issue in high-performance integrated circuits. This paper further proposes to use 3-D stacked DRAM dies to provide decoupling capacitors for the processor die. This can well leverage the superior capacitor fabrication ability of DRAM to reduce the area penalty of decoupling capacitor insertion on the processor die. For its practical implementation, a simple uniform decoupling capacitor network design strategy is presented. To demonstrate through-DRAM TSV allocation and decoupling capacitor insertion strategy and evaluate involved tradeoffs, circuit SPICE simulations and computer system simulations are carried out to quantitatively demonstrate the effectiveness and investigate various design tradeoffs.
Qi Wu 0006, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Design of Last-Level On-Chip Cache Using Spin-Torque Transfer RAM (STT RAM)
abstract
Because of its high storage density with superior scalability, low integration cost and reasonably high access speed, spin-torque transfer random access memory (STT RAM) appears to have a promising potential to replace SRAM as last-level on-chip cache (e.g., L2 or L3 cache) for microprocessors. Due to unique operational characteristics of its storage device magnetic tunneling junction (MTJ), STT RAM is inherently subject to a write latency versus read latency tradeoff that is determined by the memory cell size. This paper first quantitatively studies how different memory cell sizing may impact the overall computing system performance, and shows that different computing workloads may have conflicting expectations on memory cell sizing. Leveraging MTJ device switching characteristics, we further propose an STT RAM architecture design method that can make STT RAM cache with relatively small memory cell size perform well over a wide spectrum of computing benchmarks. This has been well demonstrated using CACTI-based memory modeling and computing system performance simulations using SimpleScalar. Moreover, we show that this design method can also reduce STT RAM cache energy consumption by up to 30% over a variety of benchmarks.
Wei Xu 0021, Hongbin Sun 0001, Xiaobin Wang, Yiran Chen 0001, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.5
2011 A Time-Aware Fault Tolerance Scheme to Improve Reliability of Multilevel Phase-Change Memory in the Presence of Significant Resistance Drift
abstract
Because of its promising scalability potential and support of multilevel per cell storage, phase-change memory has become a topic of great current interest. However, recent studies show that structural relaxation effect makes the resistance of phase-change material drift over the time, which can severely degrade multilevel per cell phase-change memory storage reliability. This makes powerful memory fault tolerance solutions indispensable, where error correction code (ECC) will play an essential role. This work aims to develop fault tolerance solutions that can effectively compensate memory cell resistance drift. First, based upon information-theoretical study, we show that conventional use of ECC, which is unaware of memory content lifetime, can only achieve the performance with a big gap from the information-theoretical bounds. This motivates us to study the potential of time-aware memory fault tolerance, where the basic idea is to keep track the memory content lifetime and use this lifetime information to accordingly adjust how memory cell resistance is quantized and interpreted for ECC decoding. Under this time-aware fault tolerance framework, we study the use of two types of ECCs, including classical codes such as BCH that only demand hard-decision input and advanced codes such as low-density parity-check (LDPC) codes that demand soft-decision probability input. Using hypothetical four-level per cell and eight-level per cell phase-change memory with BCH and LDPC codes as test vehicles, we carry out extensive analysis and simulations, which demonstrate very significant performance advantages of such time-aware memory fault tolerance strategy in the presence of significant memory cell resistance drift.
Wei Xu 0021, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.2
2010 A nondestructive self-reference scheme for Spin-Transfer Torque Random Access Memory (STT-RAM)
abstract
We proposed a novel self-reference sensing scheme for Spin-Transfer Torque Random Access Memory (STT-RAM) to overcome the large bit-to-bit variation of Magnetic Tunneling Junction (MTJ) resistance. Different from all the existing schemes, our solution is nondestructive: The stored value in the STT-RAM cell does NOT need to be overwritten by a reference value. And hence, long write-back operation (of the original stored value) is eliminated. The robustness analyses of the existing scheme and our proposed nondestructive scheme are also presented. The measurement results from a 16kb testing chip successfully confirmed the effectiveness of our technique.
Yiran Chen 0001, Hai Li 0001, Xiaobin Wang, Wenzhong Zhu, Wei Xu 0021, Tong Zhang 0002
DATE6
2010 DRAM-based FPGA enabled by three-dimensional (3d) memory stacking (abstract only)
abstract
Motivated by the emerging three-dimensional (3D integration technologies, this paper studies the potential of applying 3D memory stacking to enable FPGA devices use on-chip DRAM cells to store configuration data. In current design practice, FPGAs do not use on-chip DRAM cells for configuration data storage mainly because on-chip DRAM self-refreshing involves destructive DRAM read operations. This problem can be solved if we use a 3D stacked memory as primary FPGA configuration data storage and externally refresh on-chip DRAM cells. Since the 3D stacked memory can easily store multiple sets of configuration data, it can meanwhile enable high-speed FPGA dynamic reconfiguration. In this paper, we study such DRAM-based FPGA design enabled by 3D memory stacking and investigate potential design issues, and employ the VPR tool set to demonstrate that DRAM-based FPGAs can noticeably reduce FPGA die area and hence improve speed and energy consumption performance, compared their SRAM-based counterparts.
Yangyang Pan, Tong Zhang 0002
FPGA2
2010 Combined magnetic- and circuit-level enhancements for the nondestructive self-reference scheme of STT-RAM
abstract
A nondestructive self-reference read scheme (NSRS) was recently proposed to overcome the bit-to-bit variation in Spin-Transfer Torque Random Access Memory (STT-RAM). In this work, we introduced three magnetic- and circuit-level techniques, including 1) R-I curve skewing, 2) yield-driven sensing current selection, and 3) ratio matching to improve the sense margin and robustness of NSRS. The measurements of our 16Kb STT-RAM test chip show that compared to the original NSRS design, our proposed technologies successfully increased the sense margin by 2.5X with minimized impacts on the memory reliability and hardware cost.
Yiran Chen 0001, Hai Li 0001, Xiaobin Wang, Wenzhong Zhu, Wei Xu 0021, Tong Zhang 0002
ISLPED6
2010 DiffECC: Improving SSD Read Performance Using Differentiated Error Correction Coding Schemes
abstract
This paper presents a cross-layer co-design approach to reduce SSD read response latency. The key is to cohesively exploit the NAND flash memory device write speed vs. raw storage reliability trade-off at the physical layer and run-time data access workload variation at the system level. Leveraging run-time data access workload variation, we can opportunistically slow down NAND flash memory write speed and hence improve NAND flash memory raw storage reliability. This naturally enables an opportunistic use of weaker error correction schemes that can directly reduce SSD read access latency. We develop a disk-level scheduling scheme to effectively smooth the write workload in order to maximize the occurrence of run-time opportunistic NAND flash memory write slow down. Using 2 bits/cell NAND flash memory with BCH-based error correction correction as a test vehicle, we carry out extensive simulations over various workloads and demonstrate that this developed cross-layer co-design solution can reduce the average SSD read latency by up to 96%.
Guanying Wu, Xubin He, Ningde Xie, Tong Zhang 0002
MASCOTS4
2010 Exploiting three-dimensional (3D) memory stacking to improve image data access efficiency for motion estimation accelerators
Yiran Li 0001, Yang Liu 0016, Tong Zhang 0002
Signal Process. Image Commun.3
2010 Improving Multi-Level NAND Flash Memory Storage Reliability Using Concatenated BCH-TCM Coding
abstract
By storing more than one bit in each memory cell, multi-level per cell (MLC) NAND flash memories are dominating global flash memory market due to their appealing storage density advantage. However, continuous technology scaling makes MLC NAND flash memories increasingly subject to worse raw storage reliability. This paper presents a memory fault tolerance design solution geared to MLC NAND flash memories. The basic idea is to concatenate trellis coded modulation (TCM) with an outer BCH code, which can greatly improve the error correction performance compared with the current design practice that uses BCH codes only. The key is that TCM can well leverage the multi-level storage characteristic to reduce the memory bit error rate and hence relieve the burden of outer BCH code, at no cost of extra redundant memory cells. The superior performance of such concatenated BCH-TCM coding systems for MLC NAND flash memories has been well demonstrated through computer simulations. A modified TCM demodulation approach is further proposed to improve the tolerance to static memory cell defects. We also address the associated practical implementation issues in case of using either single-page or multi-page programming strategy, and demonstrate the silicon implementation efficiency through application-specific integrated circuit design at 65 nm node.
Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Computation Error Analysis in Digital Signal Processing Systems With Overscaled Supply Voltage
abstract
It has been recently demonstrated that digital signal processing systems may possibly leverage unconventional voltage overscaling (VOS) to reduce energy consumption while maintaining satisfactory signal processing performance. Due to the computation-intensive nature of most signal processing algorithms, the energy saving potential largely depends on the behavior of computer arithmetic units in response to overscaled supply voltage. This paper shows that different hardware implementations of the same computer arithmetic function may respond to VOS very differently and result in different energy saving potentials. Therefore, the selection of appropriate computer arithmetic architecture is an important issue in voltage-overscaled signal processing system design. This paper presents an analytical method to estimate the statistics of computer arithmetic computation errors due to supply voltage overscaling. Compared with computation-intensive circuit simulations, this analytical approach can be several orders of magnitude faster and can achieve a reasonable accuracy. This approach can be used to choose the appropriate computer arithmetic architecture in voltage-overscaled signal processing systems. Finally, we carry out case studies on a coordinate rotation digital computer processor and a finite-impulse-response filter to further demonstrate the importance of choosing proper computer arithmetic implementations.
Yang Liu 0016, Tong Zhang 0002, Keshab K. Parhi
IEEE Trans. Very Large Scale Integr. Syst.2
2010 Design of Spin-Torque Transfer Magnetoresistive RAM and CAM/TCAM with High Sensing and Search Speed
abstract
With a great scalability potential, nonvolatile magnetoresistive memory with spin-torque transfer (STT) programming has become a topic of great current interest. This paper addresses cell structure design for STT magnetoresistive RAM, content addressable memory (CAM) and ternary CAM (TCAM). We propose a new RAM cell structure design that can realize high speed and reliable sensing operations in the presence of relatively poor magnetoresistive ratio, while maintaining low sensing current through magnetic tunneling junctions (MTJs). We further apply the same basic design principle to develop new cell structures for nonvolatile CAM, and TCAM. The effectiveness of the proposed RAM, CAM and TCAM cell structures has been demonstrated by circuit simulation at 0.18 ¿m CMOS technology.
Wei Xu 0021, Tong Zhang 0002, Yiran Chen 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2009 Improving VLIW Processor Performance Using Three-Dimensional (3D) DRAM Stacking
abstract
This work studies the potential of using emerging 3D integration to improve embedded VLIW computing system. We focus on the 3D integration of one VLIW processor die with multiple high-capacity DRAM dies. Our proposed memory architecture employs 3D stacking technology to bond one die containing several processing clusters to multiple DRAM dies for a primary memory. The 3D technology also enables wide low-latency buses between clusters and memory and enable the latency of 3D DRAM L2 cache comparable to 2D SRAM L2 cache. These enable it to replace the 2D SRAM L2 cache with 3D DRAM L2 cache. The die area for 2D SRAM L2 cache can be re-allocated to additional clusters that can improve the performance of the system. From the simulation results, we find 3D stacking DRAM main memory can improve the system performance by 10%~80% than 2D off-chip DRAM main memory depending on different benchmarks. Also, for a similar logic die area, a four clusters system with 3D DRAM L2 cache and 3D DRAM main memory outperforms a two clusters system with 2D SRAM L2 cache and 3D DRAM main memory by about 10%.
Yangyang Pan, Tong Zhang 0002
ASAP2
2009 Improving STT MRAM storage density through smaller-than-worst-case transistor sizing
abstract
This paper presents a technique to improve the storage density of spin-torque transfer (STT) magnetoresistive random access memory (MRAM) in the presence of significant magnetic tunneling junction (MTJ) write current threshold variability. In conventional design practice, the nMOS transistor within each memory cell is sized to be large enough to carry a current larger than the worst-case MTJ write current threshold, leading to an increasing storage density penalty as the technology scales down. To mitigate such variability-induced storage density penalty, this paper presents a smaller-than-worst-case transistor sizing approach with the underlying theme of jointly considering memory cell transistor sizing and defect tolerance. Its effectiveness is demonstrated using 256Mb STT MRAM design at 45nm node as a test vehicle. Results show that, under a normalized write current threshold deviation of 20%, the overall memory die size can be reduced by more than 20% compared with the conventional worst-case transistor sizing design practice.
Wei Xu 0021, Yiran Chen 0001, Xiaobin Wang, Tong Zhang 0002
DAC4
2009 Improving multi-level NAND flash memory storage reliability using concatenated TCM-BCH coding
abstract
By storing multiple bits in each memory cell, multi-level per cell (MLC) NAND flash memories have been increasingly dominant in the flash memory market due to their obvious storage density advantage. However, MLC NAND flash memories are much more subject to storage reliability degradation as the technology continues to scale down. This paper presents an error correcting solution by concatenating trellis coded modulation (TCM) with an outer BCH code, which can greatly improve the performance compared with the current design practice that uses BCH codes only. The key is that TCM can well match to the multi-level storage characteristic in order to reduce the memory bit error rate and hence relieve the burden of outer BCH code, at no cost of extra redundant memory cells. The superior error correcting performance of such concatenated TCM-BCH coding systems for MLC NAND flash memories has been well demonstrated through computer simulations, and their silicon implementation efficiency has been evaluated through ASIC design at 65nm node.
Tong Zhang 0002
ACM Great Lakes Symposium on VLSI2
2009 Efficient implementation of decoupling capacitors in 3D processor-dram integrated computing systems
abstract
Three-dimensional (3D) integration of a single high performance microprocessor die and multiple DRAM dies has been considered as a viable option to tackle the looming memory wall problem. Meanwhile, on-chip decoupling capacitors are becoming increasingly important to ensure power delivery integrity, particularly for high-performance integrated circuits. Targeting at 3D processor-DRAM integrated computing systems, this paper proposes to use 3D stacked DRAM dies to provide decoupling capacitors for the processor die. This can well leverage the superior capacitor fabrication ability of DRAM to eliminate the area penalty of decoupling capacitor insertion on the processor die. For its practical implementation, a simple uniform decoupling capacitor network design strategy is presented, and circuit SPICE simulations and computer system simulations are carried out to quantitatively demonstrate the effectiveness and illustrate various design trade-offs.
Qi Wu 0006, Jian-Qiang Lu, Kenneth Rose, Tong Zhang 0002
ACM Great Lakes Symposium on VLSI4
2009 Candidate bit based bit-flipping decoding algorithm for LDPC codes
abstract
A novel hard-decision decoding algorithm for low-density parity-check (LDPC) codes is proposed in this paper. This algorithm employs the correlation information among the column vectors of the parity-check matrix and syndrome vector for decoding. It does not require soft information, and has low decoding complexity. Simulation results show that the proposed decoding algorithm could provide an effective tradeoff between error performance and decoding complexity.
Guiqiang Dong, Ningde Xie, Tong Zhang 0002, Huaping Liu 0002
ISIT4
2009 Data manipulation techniques to reduce phase change memory write energy
abstract
Due to its great scalability potential, phase change memory has become a topic of great current interest. However, high write energy consumption appears to be one of the biggest challenges to be tackled before phase change memory can be adopted as a mainstream memory technology. This paper presents architecture level technique to reduce phase change memory write energy consumption through data manipulations. Motivated by the fact that phase change memory read incurs much less energy than write and write of different value to a phase change memory cell incurs largely different energy, we present two memory write data manipulation techniques that can effectively reduce the overall memory write energy consumption. Their effectiveness has been demonstrated based on mathematical analysis and computer system simulation using phase change memory as the main memory in the computer memory hierarchy. Significant energy savings with up to more than 60% have been shown over a wide range of computer system benchmarks.
Wei Xu 0021, Jibang Liu, Tong Zhang 0002
ISLPED3
2009 3-D Data Storage, Power Delivery, and RF/Optical Transceiver - Case Studies of 3-D Integration From System Design Perspectives
abstract
Three-dimensional (3-D) integration of systems by vertically stacking and interconnecting multiple materials, technologies, and functional components offers a wide range of benefits, including speed, bandwidth and density increase, power reduction, small form factor, packaging reduction, yield and reliability increase, flexible heterogeneous integration with multifunctionality, and overall cost reduction. A new spectrum of opportunities and challenges arises for integrated system designers, which warrants rethinking and innovations from system design perspectives. By selecting three representative cases, i.e., solid-state data storage, power delivery, and hybrid radio-frequency/optical transceiver for distributed sensor networks, this paper intends to exemplify the potentials of exploiting the benefits of 3-D integration technology from system perspectives.
Tong Zhang 0002, Rino Micheloni, Guoyan Zhang, Z. Rena Huang, Jian-Qiang Lu
Proc. IEEE1
2009 Leveraging Access Locality for the Efficient Use of Multibit Error-Correcting Codes in L2 Cache
abstract
It is almost evident that SRAM-based cache memories will be subject to a significant degree of parametric random defects if one wants to leverage the technology scaling to its full extent. Although strong multibit error-correcting codes (ECC) appear to be a natural choice to handle a large number of random defects, investigation of their applications in cache remains largely missing arguably because it is commonly believed that multibit ECC may incur prohibitive performance degradation and silicon/energy cost. By developing a cost-effective L2 cache architecture using multibit ECC, this paper attempts to show that, with appropriate cache architecture design, this common belief may not necessarily hold true for L2 cache. The basic idea is to supplement a conventional L2 cache core with several special-purpose small caches/buffers, which can greatly reduce the silicon cost and minimize the probability of explicitly executing multibit ECC decoding on the cache read critical path, and meanwhile, maintain soft error tolerance. Experiments show that, at the random defect density of 0.5 percent, this design approach can maintain almost the same instruction per cycle (IPC) performance over a wide spectrum of benchmarks compared with ideal defect-free L2 cache, while only incurring less than 3 percent of silicon area overhead and 36 percent power consumption overhead.
Hongbin Sun 0001, Nanning Zheng 0001, Tong Zhang 0002
IEEE Trans. Computers3
2009 Design of Voltage Overscaled Low-Power Trellis Decoders in Presence of Process Variations
abstract
In hardware implementations of many signal processing functions, timing errors on different circuit signals may have largely different importance with respect to the overall signal processing performance. This motivates us to apply the concept ofunequalerrortoleranceto enable the use of voltage overscaling at minimal signal processing performance degradation. Realization of unequal error tolerance involves two main issues, including how to quantify the importance of each circuit signal and how to incorporate the importance quantification into signal processing circuit design. We developed techniques to tackle these two issues and applied them to two types of trellis decoders including Viterbi decoder for convolutional code decoding and max-log-maximum a posteriori (MAP) decoder for turbo code decoding. Simulation results demonstrated promising energy saving potentials of the proposed design solution on both trellis decoding computation and memory storage at small decoding performance degradation.
Yang Liu 0016, Tong Zhang 0002, Jiang Hu 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2008 Energy-efficient soft-output trellis decoder design using trellis quasi-reduction and importance-aware clock skew scheduling
abstract
Energy-efficient implementation of high-speed soft-output trellis decoders is of great practical importance. This paper first presents an algorithm-level technique, referred to as quasi-reduced-state trellis decoding, that enables the use of reduced-state trellis decoding concept to reduce the energy consumption of decoding data storage without incurring any speed penalty. Then we propose to integrate this algorithm-level technique with an importance-aware clock skew scheduling approach that enables the use of aggressive voltage overscaling in decoding computation datapath at the cost of small decoding performance degradation. The integration of these two techniques can provide a wide and flexible design space to explore the decoding performance vs. decoding energy consumption trade-off for very high-speed soft-output trellis decoder implementations. The effectiveness has been demonstrated through 1Gbps soft- output Viterbi algorithm (SOVA) decoder ASIC design at 65nm technology node.
Yang Liu 0016, Tong Zhang 0002
ISCAS3
2008 Spin-transfer torque magnetoresistive content addressable memory (CAM) cell structure design with enhanced search noise margin
abstract
This paper presents a new memory cell structure for content addressable memory (CAM) based on magnetic tunneling junction (MTJ). Each CAM cell uses a pair of differential MTJs as basic storage element and incorporates transistors to greatly improve the cell search noise margin at low sensing current. Using the same design principle, we further develop an area-efficient cell structure for ternary CAM (TCAM), which occupies about 25% less area compared with directly using two CAM cells to form one TCAM cell. The effectiveness of the proposed CAM and TCAM cell structures has been demonstrated by circuit simulation at 0.18 mum CMOS technology.
Wei Xu 0021, Tong Zhang 0002, Yiran Chen 0001
ISCAS2
2007 Hybrid resistor/FET-logic demultiplexer architecture design for hybrid CMOS/nanodevice circuits
abstract
Hybrid nanoelectronics are emerging as one viable option to sustain the Moorepsilas Law after the CMOS scaling limit is reached. One main design challenge in hybrid nanoelectronics is the interface (named as demux) between the highly dense nanowires in nanodevice crossbars and relatively coarse microwires in CMOS domain. The prior work on demux design use a single type of devices to realize the demultiplexing function, but hardly provides a satisfactory solution. This work proposes to combine resistor with FET to implement the demux, leading to the so-called hybrid resistor/FET-logic demux. Such hybrid demux architecture can make these two types of devices well complement each other to improve the overall demux design effectiveness. Furthermore, the effects of resistor conductance variability are analyzed and evaluated based on computer simulations.
Tong Zhang 0002
ICCD2
2007 Low power soft-output signal detector design for wireless MIMO communication systems
abstract
Energy-efficient realization of soft-output signal detection is of great importance in emerging high-speed multiple-input multiple-output (MIMO) wireless communication systems. This paper presents three algorithm-level complexity-reduction techniques for soft-output detector design to achieve significant energy savings. To demonstrate their effectiveness, we designed a soft-output detector for 4-4 MIMO with 64-QAM using 65nm CMOS technology. While achieving near-optimum detection performance, the detector can support over 100Mbps throughput with only 0.24mm2 silicon area and 11mw power, leading to a -10 improvement over the state of the art.
Sizhong Chen, Tong Zhang 0002
ISLPED2
2007 On the selection of arithmetic unit structure in voltage overscaled soft digital signal processing
abstract
A soft digital signal processing (DSP) design paradigm has been recently proposed to reduce the energy consumption of DSP systems through voltage overscaling. This paper shows that the selection of arithmetic unit structure can be an important and non-trivial issue in soft DSP system design. We present an optimal formulation and propose sub-optimal low-complexity approximations for selecting the appropriate arithmetic unit structure in voltage overscaled signal processing systems. We further present a case study on choosing the appropriate MAC (multiply-accumulate) structure in voltage overscaled FIR (finite impulse response) filter.
Yang Liu 0016, Tong Zhang 0002
ISLPED2
2007 Turbo- and LDPC-Coded MIMO-OFDM Systems: A Comparative Study
abstract
In this paper, we employ iteratively decodable codes in a turbolike receiver of a multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) communication system. With such a receiver, we compare the decoding complexity and performance of a turbo code with a low- density-parity-check (LDPC) code. For the same level of decoding complexity, we show that LDPC codes perform better than turbo codes in a typical high data rate MIMO-OFDM system.
Baoshen Tan, Yan Xin 0001, Syed Aon Mujtaba, Tong Zhang 0002
PIMRC4
2007 Relaxed K-Best MIMO Signal Detector Design and VLSI Implementation
abstract
Signal detector is a key element in a multiple-input multiple-output (MIMO) wireless communication receiver. It has been well demonstrated that nonlinear tree search MIMO detectors can achieve near-optimum detection performance, nevertheless their efficient high-speed VLSI implementations are not trivial. For example, the hardware design of hard- or soft- output detectors for a 4 times 4 MIMO system with 64 quadrature amplitude modulation (QAM) still remains missing in the open literature. As an attempt to tackle this challenge, this paper presents an implementation-oriented breadth-first tree search MIMO detector design solution. The key is to appropriately modify the conventional breadth-first tree search detection algorithm in order to largely improve the suitability for efficient hardware implementation, while maintaining good detection performance. To demonstrate the effectiveness of the proposed design solution, using 0.13-mum CMOS standard cell and memory libraries, we designed a soft-output signal detector for 4 times 4 MIMO with 64-QAM. With the silicon area of about 31 mm2, the detector can achieve above 100 Mb/s and realize the performance very close to that of the sphere decoding algorithm
Sizhong Chen, Tong Zhang 0002, Yan Xin 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2006 Relaxed tree search MIMO signal detection algorithm design and VLSI implementation
abstract
This paper presents an implementation-oriented breadth-first tree search MIMO detector design solution. Techniques at algorithm and VLSI architecture levels are developed to improve the implementation efficiency. Using Synopsys synthesis tool with 0.13 /spl mu/m CMOS technology, we designed soft-output detectors for 4 /spl times/ 4 MIMO channel with 64-QAM modulation. With the silicon areas less than 15 mm/sup 2/, the detectors can achieve up to about 80 Mbps and realize the performance very close to detectors using the sphere decoding algorithm.
Sizhong Chen, Tong Zhang 0002, Manish Goel
ISCAS2
2006 Multilevel flash memory on-chip error correction based on trellis coded modulation
abstract
This paper presents a multilevel (ML) flash memory on-chip error correction system design based on the concept of trellis coded modulation (TCM). This is motivated by the non-trivial modulation process in ML memory storage and the effectiveness of TCM on integrating coding with modulation to provide better performance. Using code storage 2bits/cell flash memory as a test vehicle, the effectiveness of TCM-based systems, in terms of error-correcting performance, coding redundancy, silicon cost, and operation latency, has been successfully demonstrated
Siddharth Devarajan, Kenneth Rose, Tong Zhang 0002
ISCAS4
2006 Low power state-parallel relaxed adaptive Viterbi decoder design and implementation
abstract
In this paper, we present an algorithm/architecture-level design solution for implementing state-parallel adaptive Viterbi decoders that, compared with their Viterbi counterparts, can achieve significant power savings and modest silicon area reduction, while maintaining almost the same decoding performance and throughput. The effectiveness of the proposed solution has been demonstrated using convolutional codes decoders as test vehicles, where Synopsys tools are used for synthesis, layout, and post-layout power estimation.
Tong Zhang 0002
ISCAS2
2006 High-rate quasi-cyclic LDPC codes for magnetic recording channel with low error floor
abstract
By implementing an FPGA-based simulator, we investigate the performance of high-rate quasi-cyclic (QC) LDPC codes for the magnetic recording channel at very low sector error rates. Results show that error-floor-free performance can be realized by randomly constructed high-rate regular QC-LDPC codes with column weight 4 for sector error rates as low as 10/sup -9/. We also conjecture several rules for designing randomly constructed high-rate regular QC-LDPC codes with low error floor. We also present a decoder architecture that is well suited to achieving high decoding throughput for these high-rate QC-LDPC codes with low error floor.
Hao Zhong 0006, Tong Zhang 0002, Erich F. Haratsch
ISCAS2
2006 Triple-rail MOS current mode logic for high-speed self-timed pipeline applications
abstract
High speed and low power is the dream of circuit designers. In this paper a novel self-timed logic family is presented for high-speed self-timed pipelining applications. We developed a novel triple-rail MOS current mode logic (Tr-MCML) logic family and integrated it seamlessly with self-timed pipelines. This self-timed pipeline is designed to realize power-on-demand operations that achieve both high speed and low power, which is appropriate for the design of bus drivers, asynchronous I/Os, and infinite impulse response (IIR) filters. The ripple-carry adder is used as a testbench for verification. Simulation shows that the energy-delay product (EDP) of large digital systems can be reduced significantly.
Kuan Zhou, Sizhong Chen, Allen Drake, John F. McDonald 0001, Tong Zhang 0002
ISCAS6
2005 Parallel Logic Simulation of Million-Gate VLSI Circuits
abstract
The complexity of today's VLSI chip designs makes verification a necessary step before fabrication. As a result, gate-level logic simulation has became an integral component of the VLSI circuit design process which verifies the design and analyzes its behavior. Since the designs constantly grow in size and complexity, there is a need for ever more efficient simulations to keep the gate-level logic verification time acceptably small. The focus of this paper is an efficient simulation of large chip designs. We present the design and implementation of a new parallel simulator, called DSIM, and demonstrate DSIM's efficiency and speed by simulating a million gate circuit using different numbers of processors.
Lijuan Zhu, Gilbert Chen, Boleslaw K. Szymanski, Carl Tropper, Tong Zhang 0002
MASCOTS5
2005 Parallel high-throughput limited search trellis decoder VLSI design
abstract
Limited search trellis decoding algorithms have great potentials of realizing low power due to their largely reduced computational complexity compared with the widely used Viterbi algorithm. However, because of the lack of operational parallelism and regularity in their original formulations, the limited search decoding algorithms have been traditionally ruled out for applications demanding very high throughput. We believe that, through appropriate algorithm and hardware architecture co-design, certain limited search trellis decoding algorithms can become serious competitors to the Viterbi algorithm for high-throughout applications. Focusing on the well-known T-algorithm, this paper presents techniques at the algorithm and VLSI architecture levels to design fully parallel T-algorithm limited search trellis decoders. We first develop a modified T-algorithm, called SPEC-T, to improve the algorithmic parallelism. Then, based on the conventional state-parallel register exchange Viterbi decoder, we develop a parallel SPEC-T decoder architecture that can effectively transform the reduced computational complexity at the algorithm level to the reduced switching activities in the hardware. We demonstrate the effectiveness of the SPEC-T design solution in the context of convolutional code decoding. Compared with state-parallel register exchange Viterbi decoders, the SPEC-T convolutional code decoders can achieve almost the same throughput and decoding performance, while realizing up to 56% power savings. For the first time, this work provides an approach to exploit the low power potential of the T-algorithm in very high throughput applications.
Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.2
2002 On the high-speed VLSI implementation of errors-and-erasures correcting reed-solomon decoders
abstract
Recently a novel algorithm transformation was proposed to reduce the critical path of Berlekamp-Massey algorithm implementation for errors-alone Reed-Solomon decoding. In this paper, we apply the same methodology to transform the Berlekamp-Massey algorithm for errors-and-erasures RS decoding. We present a regular hardware architecture to implement the reformulated Berlekamp-Massey algorithm, which can achieve high throughput. Moreover, an operation scheduling scheme is proposed to further reduce the hardware complexity without loss of throughput.
Tong Zhang 0002, Keshab K. Parhi
ACM Great Lakes Symposium on VLSI1
2001 High-performance, low-complexity decoding of generalized low-density parity-check codes
abstract
A class of pseudo-random compound error-correcting codes, called generalized low-density (GLD) parity-check codes, has been proposed recently. As a generalization of Gallager's low-density parity-check (LDPC) codes, GLD codes are also asymptotically good in the sense of minimum distance criterion and can be effectively decoded based on iterative soft-input soft-output (SISO) decoding of individual constituent codes. The code performance and decoding complexity of GLD codes are heavily dependent on the employed SISO decoding algorithm. In this paper, we show that Max-Log-MAP is an attractive SISO decoding algorithm for GLD coding scheme, considering the trade-off between performance and complexity in the practical implementations. A normalized Max-Log-MAP is presented to improve the GLD code performance significantly compared with using conventional Max-Log-MAP. Moreover, we propose two techniques, decoding task scheduling and reduced search Max-Log-MAP, to effectively reduce the decoding complexity without any performance degradation.
Tong Zhang 0002, Keshab K. Parhi
GLOBECOM1
2001 A class of efficient-encoding generalized low-density parity-check codes
abstract
In this paper, we investigate an efficient encoding approach for generalized low-density (GLD) parity check codes, a generalization of Gallager's (1962, 1963) low-density parity check (LDPC) codes. We propose a systematic approach to construct an approximate upper triangular GLD parity check matrix which defines a class of efficient-encoding GLD codes. It is shown that such GLD codes have equally good performance. By effectively exploiting structure sharing in the encoding process, we also present a hardware/software codesign for practical encoder implementation of these efficient-encoding GLD codes.
Tong Zhang 0002, Keshab K. Parhi
ICASSP1
2001 Systematic Design of Original and Modified Mastrovito Multipliers for General Irreducible Polynomials
abstract
This paper considers the design of bit-parallel dedicated finite field multipliers using standard basis. An explicit algorithm is proposed for efficient construction of Mastrovito product matrix, based on which we present a systematic design of Mastrovito multiplier applicable to GF(2/sup m/) generated by an arbitrary irreducible polynomial. This design effectively exploits the spatial correlation of elements in Mastrovito product matrix to reduce the complexity. Using a similar methodology, we propose a systematic design of modified Mastrovito multiplier, which is suitable for GF(2/sup m/) generated by high-Hamming weight irreducible polynomials. For both original and modified Mastrovito multipliers, the developed multiplier architectures are highly modular, which is desirable for VLSI hardware implementation. Applying the proposed algorithm and design approach, we study the Mastrovito multipliers for several special irreducible polynomials, such as trinomial and equally-spaced-polynomial, and the obtained complexity results match the best known results. Moreover, we have discovered several new special irreducible polynomials which also lead to low-complexity Mastrovito multipliers.
Tong Zhang 0002, Keshab K. Parhi
IEEE Trans. Computers1