Chao Zhang 0007

dblp:94/3019-7 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
0since 2021 · last 2020
0000-0003-0940-4709ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 4 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Memory systems · 68% Storage systems · 11% Energy-efficient computing · 8%
Network and information security
2 papers
Hardware security and side channels · 50% Cryptographic protocols and secure computation · 50%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
oblivious RAM
1.032020
Fork Path: Batching ORAM Requests to Remove Redundant Memory Accesses · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Shadow Block: Accelerating ORAM Accesses with Data Duplication · MICRO 2018
Fork path: improving efficiency of ORAM by removing redundant memory accesses · MICRO 2015
Hardware security and side channels › side-channel countermeasures
memory access pattern protection
0.522018
Shadow Block: Accelerating ORAM Accesses with Data Duplication · MICRO 2018
Fork path: improving efficiency of ORAM by removing redundant memory accesses · MICRO 2015
Cryptographic protocols and secure computation › oblivious data structures
oblivious random access machine
0.522018
Shadow Block: Accelerating ORAM Accesses with Data Duplication · MICRO 2018
Fork path: improving efficiency of ORAM by removing redundant memory accesses · MICRO 2015
Memory systems
non-volatile memory
0.522016
Statistical Cache Bypassing for Non-Volatile Memory · IEEE Trans. Computers 2016
Hi-fi playback: tolerating position errors in shift operations of racetrack memory · ISCA 2015
Energy-efficient computing
power management
0.312018
PM3: Power Modeling and Power Management for Processing-in-Memory · HPCA 2018
Memory systems
processing-in-memory
0.312018
PM3: Power Modeling and Power Management for Processing-in-Memory · HPCA 2018
Memory systems › cache management › cache insertion policy
cache bypassing
0.212016
Statistical Cache Bypassing for Non-Volatile Memory · IEEE Trans. Computers 2016
Memory systems
cache management
0.212016
Statistical Cache Bypassing for Non-Volatile Memory · IEEE Trans. Computers 2016
Memory systems › emerging memory technologies › spintronic memory
racetrack memory
0.212015
Hi-fi playback: tolerating position errors in shift operations of racetrack memory · ISCA 2015
Hardware reliability and fault tolerance
soft errors
0.212015
Hi-fi playback: tolerating position errors in shift operations of racetrack memory · ISCA 2015
Cloud and datacenter computing
cloud storage
0.112020
Fork Path: Batching ORAM Requests to Remove Redundant Memory Accesses · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems
cache coherence
0.112016
Statistical Cache Bypassing for Non-Volatile Memory · IEEE Trans. Computers 2016
Memory systems › cache
chip multiprocessor cache
0.112016
Statistical Cache Bypassing for Non-Volatile Memory · IEEE Trans. Computers 2016
Hardware reliability and fault tolerance › error correction
error-correcting codes
0.112015
Hi-fi playback: tolerating position errors in shift operations of racetrack memory · ISCA 2015

Methods — techniques the papers use, named apart from their topics

path merging · 0.9ORAM request scheduling · 0.9shadow blocks · 0.7data duplication · 0.7ORAM space partitioning · 0.7prefetching · 0.4caching · 0.4processing unit boost · 0.3power-aware subtask throttling · 0.3power sprinting · 0.3merging-aware caching · 0.2
YearPublicationVenuePosition
2020 Fork Path: Batching ORAM Requests to Remove Redundant Memory Accesses
abstract
Outsourcing data to a third-party cloud provider has become quite common with the increasing use of cloud computing. This brings convenience, as well as the concern for data security and privacy. It is believed that data encryption alone is often not enough to protect users' privacy from the cloud provider. According to previous work, the sequence of storage locations accessed by the client can leak up to 90% of the sensitive information, even with data encrypted. In this context, Oblivious RAM (ORAM) is proposed. ORAM algorithms allow the client to hide its access pattern from the service provider while introducing a lot of extra operations. Among all the prototypes, Path ORAM is one of the most promising designs. However, there are still redundant memory accesses that can be removed without harming the security of traditional ORAM as we observed. We came up with three optimization techniques, including path merging, ORAM request scheduling, and merging aware caching. We also propose a prefetching technique to further decreasing the access overhead. Moreover, we also illustrate the compatibility of Fork Path and some state-of-the-art Path ORAM optimizations. Compared to traditional Path ORAM approaches, our Fork Path ORAM can reduce overall performance overhead and power consumption of memory system by 65% and 44%, while the design overhead is trivial.
Jingchen Zhu, Guangyu Sun 0003, Xian Zhang 0001, Chao Zhang 0007, Yun Liang 0001, Tao Wang 0004, Yiran Chen 0001, Jia Di
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2018 Performance analysis on structure of racetrack memory
abstract
Racetrack Memory(RM) has attracted abundant attention of memory researchers recently. RM can achieve ultrahigh storage density, fast access velocity and non-volatility. Former research has demonstrated that RM has potential to serve as on-chip cache or main memory. However, RM has more flexibility and difficulty in design space of main memory because it has more device level design parameters. The layout of macro unit (MU) needs trade-off among area, access performance and energy consumption, and its shift operation introduces extra dimension of design space. In this paper, we explore these design parameters and analyze their relationship in memory design space in both device and system levels. Based on the results, we also propose a hybrid MU structure to further optimize read intensive applications. Experimental results demonstrated the existence of regularity between design parameters and performance features. The optimized layout of racetrack MU is suggested for application areas such as big-data and IoT which need cost-effective and energy-efficient memory respectively. Together with hybrid MU structures, RM can be designed with more flexibility so that specific structures are suitable for specific applications which make “All stack optimization” possible in memory structure level.
Chao Zhang 0007, Qingda Hu, Chengmo Yang, Jiwu Shu
ASP-DAC2
2018 PM3: Power Modeling and Power Management for Processing-in-Memory
abstract
Processing-in-Memory (PIM) has been proposed as a solution to accelerate data-intensive applications, such as real-time Big Data processing and neural networks. The acceleration of data processing using a PIM relies on its high internal memory bandwidth, which always comes with the cost of high power consumption. Consequently, it is important to have a comprehensive quantitative study of the power modeling and power management for such PIM architectures. In this work, we first model the relationship between the power consumption and the internal bandwidth of PIM. This model not only provides a guidance for PIM designs but also demonstrates the potential of power management via bandwidth throttling. Based on bandwidth throttling, we propose three techniques, Power-Aware Subtask Throttling (PAST), Processing Unit Boost (PUB), and Power Sprinting (PS), to improve the energy efficiency and performance. In order to demonstrate the universality of the proposed methods, we applied them to two kinds of popular PIM designs. Evaluations show that the performance of PIM can be further improved if the power consumption is carefully controlled. Targeting at the same performance, the peak power consumption of HMC-based PIM can be reduced from 20W to 15W. The proposed power management schemes improve the speedup of prior RRAM-based PIM from 69 × to 273 ×, after pushing the power usage from about 1W to 10W safely. The model also shows that emerging RRAM is more suitable for large processing-in-memory designs, due to its low power cost to store the data.
Chao Zhang 0007, Tong Meng, Guangyu Sun 0003
HPCA1
2018 Shadow Block: Accelerating ORAM Accesses with Data Duplication
abstract
Oblivious RAM (ORAM) is a cryptographic primitive designed to hide memory access patterns. To achieve this objective, the intended data block is loaded and evicted back together with other data blocks and dummy blocks in each ORAM access. To further protect the timing pattern, extra dummy ORAM accesses are triggered periodically. Such designs lead to huge memory access overheads. Many techniques have been proposed to mitigate this problem by reducing the total number of ORAM accesses and the number of blocks per access. However, the impact of the access order of intended data block in an ORAM access is not addressed yet. In this work, we argue that higher performance can be achieved by advancing the access to the intended data block in ORAM accesses. However, changing the access order of blocks directly compromises the ORAM security. To solve this problem, we propose a duplication method to advance the access to the intended data blocks without compromising the ORAM security. The method leverages dummy blocks to store extra copies of data blocks, to facilitate early access of intended data blocks. These dummy blocks with valid data duplications are called Shadow blocks in this work. We further introduce two data duplication techniques, called RD-Dup and HD-Dup, to reorder the data block access for different purposes. In addition, we propose ORAM space partitioning to make RD-Dup and HD-Dup cooperate with each other efficiently. Compared with state-of-the-art ORAMs, our design can achieve a 32% reduction in system execution time on average, with negligible hardware overheads.
Xian Zhang 0001, Guangyu Sun 0003, Peichen Xie, Chao Zhang 0007, Yannan Liu, Lingxiao Wei, Qiang Xu 0001, Chun Jason Xue
MICRO4
2016 Performance-centric register file design for GPUs using racetrack memory
abstract
The key to high performance for GPU architecture lies in massive threading to drive the large number of cores and enable overlapping of threading execution. However, in reality, the number of threads that can simultaneously execute is often limited by the size of the register file on GPUs. The traditional SRAM-based register file costs so large amount of chip area that it cannot scale to meet the increasing demand of massive threading for GPU applications. Racetrack memory is a promising technology for designing large capacity register file on GPUs due to its high data storage density. However, without careful deployment of registers, the lengthy shift operation of racetrack memory may hurt the performance. In this paper, we explore racetrack memory for designing high performance register file for GPU architecture. High storage density racetrack memory helps to improve the thread level parallelism, i.e., the number of threads that simultaneously execute. However, if the bits of the registers are not aligned to the ports, shift operations are required to move the bits to the ports. To mitigate the shift operation overhead problem, we develop a register file preshifting strategy and a compile-time managed register mapping algorithm. Experimental results demonstrate that our technique achieves up to 24% (19% on average) improvement in performance for a variety of GPU applications.
Shuo Wang 0009, Yun Liang 0001, Chao Zhang 0007, Xiaolong Xie, Guangyu Sun 0003, Yongpan Liu, Yu Wang 0002
ASP-DAC3
2016 Pin Tumbler Lock: A shift based encryption mechanism for racetrack memory
abstract
As various non-volatile memory (NVM) technologies have been adopted in different levels of memory hierarchy, the security issue of protecting information retained in NVM after power-off has become a new challenge, which results in extensive research on data encryption for NVM. Previous encryption approaches, however, have some limitations, such as high design complexity and non-trivial timing and energy overhead. Recently, an emerging NVM called racetrack memory (RM) has been widely investigated because of its advantages of ultra-high storage density and fast read/write speed. Besides these well-known advantages, we observe that the tape-like structure of RMcell and its unique shift operation can also be leveraged to facilitate NVM data encryption. Base on this observation, we propose an efficient shift based mechanism, named Pin Tumbler Lock (PTL), which completes encryption and decryption by shifting racetracks in several nanoseconds. Experimental results demonstrate that our design can achieve the same security strength of AES-128 with 3.1% performance overhead and 3.7% energy overhead and 1.56% storage cost and 1.6% area cost.
Chao Zhang 0007, Xian Zhang 0001, Guangyu Sun 0003, Jiwu Shu
ASP-DAC2
2016 Exploring Main Memory Design Based on Racetrack Memory Technology
abstract
Emerging non-volatile memories (NVMs), which include PC-RAM and STT-RAM, have been proposed to replace DRAM, mainly because they have better scalability and lower standby power. However, previous research has demonstrated that these NVMs cannot completely replace DRAM due to either lifetime/performance (PCRAM) or density (STT-RAM) issues. Recently, a new type of emerging NVM, called Racetrack Memory (RM), has attracted more and more attention of memory researchers because it has ultra-high density and fast access speed without the write cycle issue. However, there lacks research on how to leverage RM for main memory. To this end, we explore main memory design based on RM technology in both circuit and architecture levels. In the circuit level, we propose the structure of the RM based main memory and investigate different design parameters. In the architecture level, we design a simple and efficient shift-sense address mapping policy to reduce 95% shift operations for performance improvement and power saving. At the same time, we analyze the efficiency of existing optimization strategies for NVM main memory. Our experiments show that RM can outperform DRAM for main memory, in respect of density, performance, and energy efficiency.
Qingda Hu, Guangyu Sun 0003, Jiwu Shu, Chao Zhang 0007
ACM Great Lakes Symposium on VLSI4
2016 Statistical Cache Bypassing for Non-Volatile Memory
abstract
With the increasing data throughput requirement, non-volatile memories, such as STT-RAM, PCM and RRAM, have become very competitive designs as on-chip caches in chip-multi-processors (CMPs). Since the write operations are more expensive in an asymmetric-access cache, it is more valuable to justify the data allocation. However, the asymmetric-access property of non-volatile memory is not well addressed in prior bypassing approaches, which are not energy efficient and induce non-trivial operation overhead. In this paper, we propose cache-bypassing methods designed for non-volatile memory. The basic method, SBAC, is based on data locality statistics of the whole cache rather than a signature of each cache line. The multicore extensions, SBAC-C and SBAC-G, strengthen the SBAC by distinguishing data patterns in CMPs. We observe that the decision-making of SBAC and its multicore extensions is highly accurate. Experiments show that SBAC can reduce overall energy consumption by 22.3 percent, and reduce execution time by 8.3 percent on average. The energy consumption is reduced by 21.4 and 23.4 percent for SBAC-C and SBAC-G. And the performance is improved by 7.8 and 9.6 percent for SBAC-C and SBAC-G in multicore scenario. Compared to prior approaches, SBAC outperforms and induces trivial design overhead.
Guangyu Sun 0003, Chao Zhang 0007, Peng Li 0031, Tao Wang 0004, Yiran Chen 0001
IEEE Trans. Computers2
2015 Quantitative modeling of racetrack memory, a tradeoff among area, performance, and power
abstract
Recently, an emerging non-volatile memory called Racetrack Memory (RM) becomes promising to satisfy the requirement of increasing on-chip memory capacity. RM can achieve ultra-high storage density by integrating many bits in a tape-like racetrack, and also provide comparable read/write speed with SRAM. However, the lack of circuit-level modeling has limited the design exploration of RM, especially in the system-level. To overcome this limitation, we develop an RM circuit-level model, with careful study of device configurations and circuit layouts. This model introduces Macro Unit (MU) as the building block of RM, and analyzes the interaction of its attributes. Moreover, we integrate the model into NVsim to enable the automatic exploration of its huge design space. Our case study of RM cache demonstrates significant variance under different optimization targets, in respect of area, performance, and energy. In addition, we show that the cross-layer optimization is critical for adoption of RM as on-chip memory.
Chao Zhang 0007, Guangyu Sun 0003, Fan Mi, Hai Li 0001, Weisheng Zhao 0001
ASP-DAC1
2015 An energy efficient backup scheme with low inrush current for nonvolatile SRAM in energy harvesting sensor nodes
Hehe Li, Yongpan Liu, Qinghang Zhao, Yizi Gu, Xiao Sheng, Guangyu Sun 0003, Chao Zhang 0007, Meng-Fan Chang, Huazhong Yang
DATE7
2015 From device to system: cross-layer design exploration of racetrack memory
Guangyu Sun 0003, Chao Zhang 0007, Hehe Li, Yue Zhang 0010, Yizi Gu, Jacques-Olivier Klein, Dafine Ravelosona, Yongpan Liu, Weisheng Zhao 0001, Huazhong Yang
DATE2
2015 Hi-fi playback: tolerating position errors in shift operations of racetrack memory
abstract
Racetrack memory is an emerging non-volatile memory based on spintronic domain wall technology. It can achieve ultra-high storage density. Also, its read/write speed is comparable to that of SRAM. Due to the tape-like structure of its storage cell, a "shift" operation is introduced to access racetrack memory. Thus, prior research mainly focused on minimizing shift latency/energy of racetrack memory while leveraging its ultra-high storage density. Yet the reliability issue of a shift operation, however, is not well addressed. In fact, racetrack memory suffers from unsuccessful shift due to domain misalignment. Such a problem is called "position error" in this work. It can significantly reduce mean-time-to-failure (MTTF) of racetrack memory to an intolerable level. Even worse, conventional error correction codes (ECCs), which are designed for "bit errors", cannot protect racetrack memory from the position errors.
Chao Zhang 0007, Guangyu Sun 0003, Xian Zhang 0001, Weisheng Zhao 0001, Tao Wang 0004, Yun Liang 0001, Yongpan Liu, Yu Wang 0002, Jiwu Shu
ISCA1
2015 Perspectives of racetrack memory based on current-induced domain wall motion: From device to system
abstract
Current-induced domain wall motion (CIDWM) is regarded as a promising way towards achieving emerging high-density, high-speed and low-power non-volatile devices. Racetrack memory is an attractive concept based on this phenomenon, which can store and transfer a series of data along a magnetic nanowire. Although the first prototype has been successfully fabricated, its advancement is relatively arduous caused by certain technique and material limitations. Particularly, the storage capacity issue is one of the most serious bottlenecks hindering its application for practical systems. In this paper, we present two alternative solutions to improve the capacity of racetrack memory: magnetic field assistance and chiral domain wall (DW) motion. The former one can lower the current density for DW shifting; the latter one can utilize materials with low resistivity. Both of them are able to increase the nanowire length and allow higher feasibility of large-capacity racetrack memory. Furthermore, system level simulation shows that a racetrack memory based cache can improve system performance by about 15.8% and significantly reduces the energy consumption, compared to the SRAM counterpart.
Yue Zhang 0010, Chao Zhang 0007, Jacques-Olivier Klein, Dafine Ravelosona, Guangyu Sun 0003, Weisheng Zhao 0001
ISCAS2
2015 Fork path: improving efficiency of ORAM by removing redundant memory accesses
abstract
Oblivious RAM (ORAM) is a cryptographic primitive that can prevent information leakage in the access trace to untrusted external memory. It has become an important component in modern secure processors. However, the major obstacle of adopting an ORAM design is the significantly induced overhead in memory accesses. Recently, Path ORAM has attracted attentions from researchers because of its simplicity in algorithms and efficiency in reducing memory access overhead. However, we observe that there exist a lot of redundant memory accesses during the process of ORAM requests. Moreover, we further argue that these redundant memory accesses can be removed without harming security of ORAM. Based on this observation, we propose a novel Fork Path ORAM scheme. By leveraging three optimization techniques, namely, path merging, ORAM request scheduling, and merging-aware caching, Fork Path ORAM can efficiently remove these redundant memory accesses. Based on this scheme, a detailed ORAM controller architecture is proposed and comprehensive experiments are performed. Compared to traditional Path ORAM approaches, our Fork Path ORAM can reduce overall performance overhead and power consumption of memory system by 58% and 38%, respectively, with negligible design overhead.
Xian Zhang 0001, Guangyu Sun 0003, Chao Zhang 0007, Yun Liang 0001, Tao Wang 0004, Yiran Chen 0001, Jia Di
MICRO3
2014 SBAC: a statistics based cache bypassing method for asymmetric-access caches
abstract
Asymmetric-access caches with emerging technologies, such as STT-RAM and RRAM, have become very competitive designs recently. Since the write operations consume more time and energy than read ones, data should bypass an asymmetric-access cache unless the locality can justify the data allocation. However, the asymmetric-access property is not well addressed in prior bypassing approaches, which are not energy efficient and induce non-trivial operation overhead. To overcome these problems, we propose a cache bypassing method, SBAC, based on data locality statistics of the whole cache rather than a single cache line's signature. We observe that the decision-making of SBAC is highly accurate and the optimization technique for SBAC works efficiently for multiple applications running concurrently. Experiments show that SBAC cuts down overall energy consumption by 22.3%, and reduces execution time by 8.3%. Compared to prior approaches, the design overhead of SBAC is trivial.
Chao Zhang 0007, Guangyu Sun 0003, Peng Li 0031, Tao Wang 0004, Dimin Niu, Yiran Chen 0001
ISLPED1
2013 An efficient run-time encryption scheme for non-volatile main memory
abstract
Emerging non-volatile memories (NVMs) have been considered as promising alternatives of DRAM for future main memory design. The NVM main memory has advantages of low standby power, high density, and good scalability. Its non-volatility, however, induces a security design challenge that data retained in memory after power-off need to be protected from malicious attacks. Although several approaches have been proposed to solve this problem through data encryption, they have some limitations such as high design complexity and non-trivial timing/energy overhead. Moreover, these techniques decrease the lifetime of NVM main memory due to extra write operations caused by encryption. In order to overcome these limitations, we propose an efficient PAD-XOR based encryption scheme in this work. A novel PAD generator based on a randomizer and several sub-PAD tables is introduced. With the PAD generator, our encryption scheme can provide run-time data protection to all data in NVM memory with low timing and power overhead. In addition, the encryption process can co-operate with wear-leveling of NVM to reduce design complexity. More important, our encryption technique has no impact on lifetime because no extra writes are incurred. Experimental results demonstrate that, compared to prior approaches, our design can achieve the same security strength with substantial lower overhead in respect of timing, energy consumption, and design complexity.
Xian Zhang 0001, Chao Zhang 0007, Guangyu Sun 0003, Jia Di, Tao Zhang 0032
CASES2
2013 Asymmetric-access aware optimization for STT-RAM caches with process variations
abstract
STT-RAM (Spin Transfer Torque Random Access Memory) has been extensively researched as a potential replacement of SRAM (Static RAM) as on-chip caches. Prior work has shown that STT-RAM caches can improve performance and reduce power consumption because of its advantages of high density, fast read speed, low standby power, etc. However, under the impact of process variations, using worst-case design can induce significant performance and power overhead in STT-RAM caches. In order to overcome the problem of process variations, we propose to apply the variable-latency access method to STT-RAM caches by introducing a variation-aware LRU (Least Recently Used) policy. Moreover, we show that simply applying traditional variable-latency access method is inefficient due to the read/write asymmetry. First, we demonstrate that a write-oriented data migration is preferred. Second, a block remapping is necessary to prevent some cache sets from being significantly affected by process variations. After using our techniques, the experimental results show that the performance can be improved by 13.8% and power consumption can be reduced by 14.1% compared to a prior approach [3].
Chao Zhang 0007, Guangyu Sun 0003
ACM Great Lakes Symposium on VLSI2
2012 Spintronic memristor based temperature sensor design with CMOS current reference
abstract
As the technology scales down, the increased power density brings in significant system reliability issues. Therefore, the temperature monitoring and the induced power management become more and more critical. The thermal fluctuation effects of the recently discovered spintronic memristor make it a promising candidate as a temperature sensing device. In this paper, we carefully analyzed the thermal fluctuations of spintronic memristor and the corresponding design considerations. On top of it, we proposed a temperature sensing circuit design by combining spintronic memristor with the traditional CMOS current reference. Our simulation results show that the proposed design can provide high accuracy of temperature detection within a much smaller footprint compared to the traditional CMOS temperature sensor designs. As magnetic device scales down, the relatively high power consumption is expected to be reduced.
Xiuyuan Bi, Chao Zhang 0007, Hai Li 0001, Yiran Chen 0001, Robinson E. Pino
DATE2