Zhao Zhang 0010

dblp:87/6853-10 · DBLP profile ↗
← Back
38ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 35 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 6Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
17 papers
Memory systems · 36% Energy-efficient computing · 28% Hardware reliability and fault tolerance · 19%
Network and information security
1 paper
Cryptographic protocols and secure computation · 50% Authentication and access control · 50%

Topics — the 30 heaviest of 52, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
DRAM
0.662014
Mini-Rank: A Power-EfficientDDRx DRAM Memory Architecture · IEEE Trans. Computers 2014
Thermal Modeling and Management of DRAM Systems · IEEE Trans. Computers 2013
Software thermal management of dram memory for multicore systems · SIGMETRICS 2008
Energy-efficient computing › thermal management
dynamic thermal management
0.332013
Thermal Modeling and Management of DRAM Systems · IEEE Trans. Computers 2013
Software thermal management of dram memory for multicore systems · SIGMETRICS 2008
Thermal modeling and management of DRAM memory systems · ISCA 2007
Energy-efficient computing
thermal management
0.222013
Thermal Modeling and Management of DRAM Systems · IEEE Trans. Computers 2013
Software thermal management of dram memory for multicore systems · SIGMETRICS 2008
Hardware reliability and fault tolerance › error correction
error-correcting codes
0.212014
MemGuard: A low cost and energy efficient design to support and enhance memory system reliability · ISCA 2014
Hardware reliability and fault tolerance › memory fault tolerance
memory error detection
0.212014
MemGuard: A low cost and energy efficient design to support and enhance memory system reliability · ISCA 2014
Energy-efficient computing › power management
memory power management
0.212014
Mini-Rank: A Power-EfficientDDRx DRAM Memory Architecture · IEEE Trans. Computers 2014
Hardware reliability and fault tolerance › error correction › error-correcting codes
SECDED
0.212014
MemGuard: A low cost and energy efficient design to support and enhance memory system reliability · ISCA 2014
Memory systems
DRAM memory system
0.232009
Decoupled DIMM: building high-bandwidth memory system using low-speed DRAM devices · ISCA 2009
A Performance Comparison of DRAM Memory System Optimizations for SMT Processors · HPCA 2005
Fine-Grain Priority Scheduling on Multi-Channel Memory Systems · HPCA 2002
Storage systems › flash and SSD › flash memory management › flash translation layer
address mapping
0.212013
E3CC: A memory error protection scheme with novel address mapping for subranked and low-power memories · ACM Trans. Archit. Code Optim. 2013
Hardware reliability and fault tolerance
error-correcting codes for memory
0.212013
E3CC: A memory error protection scheme with novel address mapping for subranked and low-power memories · ACM Trans. Archit. Code Optim. 2013
Energy-efficient computing
power management
0.122014
Mini-rank: Adaptive DRAM architecture for improving memory power efficiency · MICRO 2008
Mini-Rank: A Power-EfficientDDRx DRAM Memory Architecture · IEEE Trans. Computers 2014
Processor architecture and microarchitecture
multicore design
0.122009
Enabling software management for multicore caches with a lightweight hardware support · SC 2009
Thermal modeling and management of DRAM memory systems · ISCA 2007
Authentication and access control › authentication
message authentication
0.112010
An Application-Level Data Transparent Authentication Scheme without Communication Overhead · IEEE Trans. Computers 2010
Cryptographic protocols and secure computation › secure message transmission
stream authentication
0.112010
An Application-Level Data Transparent Authentication Scheme without Communication Overhead · IEEE Trans. Computers 2010
Memory systems
cache
0.122008
Gaining insights into multicore cache partitioning: Bridging the gap between simulation and real systems · HPCA 2008
Cache-Optimal Methods for Bit-Reversals · SC 1999
Memory systems
cache management
0.122009
Enabling software management for multicore caches with a lightweight hardware support · SC 2009
Cacheminer: A Runtime Approach to Exploit Cache Locality on SMP · IEEE Trans. Parallel Distributed Syst. 2000
Processor architecture and microarchitecture › multicore design
multicore cache management
0.112009
Enabling software management for multicore caches with a lightweight hardware support · SC 2009
Memory systems › cache management
shared cache management
0.112009
Enabling software management for multicore caches with a lightweight hardware support · SC 2009
Memory systems › cache management
cache partitioning
0.112008
Gaining insights into multicore cache partitioning: Bridging the gap between simulation and real systems · HPCA 2008
Memory systems › DRAM
DRAM architecture
0.112008
Mini-rank: Adaptive DRAM architecture for improving memory power efficiency · MICRO 2008
Energy-efficient computing › power management › memory power management
DRAM power reduction
0.112008
Mini-rank: Adaptive DRAM architecture for improving memory power efficiency · MICRO 2008
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.112008
Software thermal management of dram memory for multicore systems · SIGMETRICS 2008
Performance modeling and evaluation
workload characterization
0.122005
A Performance Comparison of DRAM Memory System Optimizations for SMT Processors · HPCA 2005
A permutation-based page interleaving scheme to reduce row-buffer conflicts and exploit data locality · MICRO 2000
Operating systems › fault tolerance
checkpoint and rollback
0.112014
MemGuard: A low cost and energy efficient design to support and enhance memory system reliability · ISCA 2014
Energy-efficient computing
memory power
0.112014
Mini-Rank: A Power-EfficientDDRx DRAM Memory Architecture · IEEE Trans. Computers 2014
Hardware reliability and fault tolerance
memory reliability
0.112014
MemGuard: A low cost and energy efficient design to support and enhance memory system reliability · ISCA 2014
High-performance computing › collective communication
MPI collective communication
0.112005
Performance Modeling and Tuning Strategies of Mixed Mode Collective Communications · SC 2005
Energy-efficient computing › memory energy efficiency
low-power memory
0.012013
E3CC: A memory error protection scheme with novel address mapping for subranked and low-power memories · ACM Trans. Archit. Code Optim. 2013
Energy-efficient computing
thermal modeling
0.012013
Thermal Modeling and Management of DRAM Systems · IEEE Trans. Computers 2013
Memory systems › DRAM › DRAM architecture
cached DRAM
0.012004
Design and Optimization of Large Size and Low Overhead Off-Chip Caches · IEEE Trans. Computers 2004

Methods — techniques the papers use, named apart from their topics

simulation · 0.4non-cryptographic hash function · 0.4multiset hash function · 0.4log hashing · 0.4modeling-based analysis · 0.2mini-rank architecture · 0.2heterogeneous design · 0.2error-correcting codes · 0.2coordinated DVFS · 0.2cache design · 0.2adaptive core gating · 0.2hardware support · 0.1OS-based cache management · 0.1software-based cache partitioning · 0.1measurement · 0.1
YearPublicationVenuePosition
2021 Discreet-PARA: Rowhammer Defense with Low Cost and High Efficiency
abstract
DRAM rowhammer attack is a severe security concern on computer systems using DRAM memories. A number of defense mechanisms have been proposed, but all with short-coming in either performance overhead, storage requirement, or defense strength. In this paper, we present a novel design called discreet-PARA. It creatively integrates two new components, namely Disturbance Bin Counting (DBC) and PARA-cache, into the existing PARA (Probabilistic Adjacent Row Activation) defense. The two components only require small counter and cache storages but can eliminate or significantly reduce the performance overhead of PARA. Our evaluation using SPEC CPU2017 workloads confirms that discreet-PARA can achieve very high defense strength with a performance overhead much lower than the original PARA.
Yang Liu 0114, Peiyun Wu, Zhao Zhang 0010
ICCD4
2015 Memory design for selective error protection
abstract
Memory error protection is increasingly important as memory density and capacity continue to scale. This paper presents a memory SEP (Selective Memory Protection) design that enables SEP for commodity memory modules, with no change to the modules or devices. Memory error protection is provided through embedded ECC, a recently proposed, energy-efficient ECC memory organization. The memory SEP design splits the physical memory address space into two memory regions of adjustable sizes, one with error protection and one without. With this support, the OS can adjust the size ratio of the protected region and non-protected region based on the needs of applications. In this scheme, the mapping from a physical memory address to memory device addresses is no longer power-of-two based. New and efficient address mapping schemes based on the Chinese Remainder Mapping are proposed to avoid the use of complex Euclidean division. The simulation results show that the memory SEP design may retain memory performance and cut memory power increase, while providing the ECC protection to commodity memory modules.
Yanan Cao 0002, Zhao Zhang 0010
ICCD3
2015 Flexible memory: A novel main memory architecture with block-level memory compression
abstract
Main memory system is facing increasingly high pressure from the advances of multi-core processors. The simplicity of conventional memory architecture has helped minimize memory latency and reduce the design cost. However, in present multi-core era, it is increasingly attractive to adopt flexible and advanced memory organization to further improve memory bandwidth utilization, power efficiency, and reliability, despite an increase of memory system complexity. Motivated by the idea, we propose an innovative memory compression scheme with a flexible memory organization, used in combination with the recently proposed, power-efficient sub-ranked memory. Our detailed simulation show that the scheme may gain an average of 1.5× effective capacity gain, reduce the power consumption of memory subsystem by up to 45%, on average in the range from 13% to 16%, and yield moderate performance improvement.
Yanan Cao 0002, Zhao Zhang 0010
NAS3
2014 MemGuard: A low cost and energy efficient design to support and enhance memory system reliability
abstract
Memory system reliability is increasingly a concern as memory cell density and capacity continue to grow. The conventional approach is to use redundant memory bits for error detection and correction, with significant storage, cost and power overheads. In this paper, we propose a novel, system-level scheme called MemGuard for memory error detection. With OS-based checkpointing, it is also able to recover program execution from memory errors. The memory error detection of MemGuard is motivated by memory integrity verification using log hashes. It is much stronger than SECDED in error detection, incurs negligible hardware cost and energy overhead and no storage overhead, and is compatible with various memory organizations. It may play the role of ECC memory in consumer-level computers and mobile devices, without the shortcomings of ECC memory. In server computers, it may complement SECDED ECC or Chipkill Correct by providing even stronger error detection. We have comprehensively investigated and evaluated the feasibility and reliability of MemGuard. We show that using an incremental multiset hash function and a non-cryptographic hash function, the performance and energy overheads of Mem-Guard are negligible. We use the mathematical deduction and synthetic simulation to prove that MemGuard is robust and reliable.
Zhao Zhang 0010
ISCA2
2014 A Host-Based Approach for Unknown Fast-Spreading Worm Detection and Containment
abstract
The fast-spreading worm, which immediately propagates itself after a successful infection, is becoming one of the most serious threats to today’s networked information systems. In this article, we present WormTerminator, a host-based solution for fast Internet worm detection and containment with the assistance of virtual machine techniques based on the fast-worm defining characteristic. In WormTerminator, a virtual machine cloning the host OS runs in parallel to the host OS. Thus, the virtual machine has the same set of vulnerabilities as the host. Any outgoing traffic from the host is diverted through the virtual machine. If the outgoing traffic from the host is for fast worm propagation, the virtual machine should be infected and will exhibit worm propagation pattern very quickly because a fast-spreading worm will start to propagate as soon as it successfully infects a host. To prove the concept, we have implemented a prototype of WormTerminator and have examined its effectiveness against the real Internet worm Linux/Slapper. Our empirical results confirm that WormTerminator is able to completely contain worm propagation in real-time without blocking any non-worm traffic. The major performance cost of WormTerminator is a one-time delay to the start of each outgoing normal connection for worm detection. To reduce the performance overhead, caching is utilized, through which WormTerminator will delay no more than 6% normal outgoing traffic for such detection on average.
Songqing Chen, Lei Liu 0021, Xinyuan Wang 0005, Xinwen Zhang, Zhao Zhang 0010
ACM Trans. Auton. Adapt. Syst.5
2014 Mini-Rank: A Power-EfficientDDRx DRAM Memory Architecture
abstract
Memory power consumption has become a severe concern in multi-core computer platforms. As memory data rate, capacity and bandwidth are being pushed higher and higher, the power consumption of memory systems becomes a significant part in the overall system power profile. Conventional memory systems do not provide an efficient mechanism for managing its power and performance tradeoff. We propose a novel mini-rank architecture for DDRx memories to reduce memory power consumption by breaking each DRAM rank into multiple narrow mini-ranks and activating fewer devices for each request. We also propose a heterogeneous mini-rank design to further improve the performance-power tradeoff for each workload based on its memory access behavior and bandwidth requirement. The evaluation results show that homogeneous mini-rank significantly reduces memory power with small performance loss. For instance, using four-core multiprogramming workloads, a x32 mini-rank configuration reduces memory power by 19.5 percent with 1.3 percent performance loss on average for memory-intensive workloads. Heterogeneous mini-rank further improves the balance between the performance and power saving. For instance, it reduces the memory power by up to 38.0 percent with an average performance loss of 2.4 percent, compared with a conventional memory system. In comparison, the x32 homogeneous mini-rank reduces memory power by up to 25.4 percent; while the x8 homogeneous mini-rank incurs performance loss by up to 19.3 percent. Furthermore, heterogeneous mini-rank achieves consistently good performance-power tradeoff for workloads made by programs of diverse memory access behavior and bandwidth requirement.
Kun Fang 0005, Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Zhichun Zhu
IEEE Trans. Computers4
2014 Secure, Efficient and Fine-Grained Data Access Control Mechanism for P2P Storage Cloud
abstract
By combining cloud computing and Peer-to-Peer computing, a P2P storage cloud can be formed to offer highly available storage services, lowering the economic cost by exploiting the storage space of participating users. However, since cloud severs and users are usually outside the trusted domain of data owners, P2P storage cloud brings forth new challenges for data security and access control when data owners store sensitive data for sharing in the trusted domain. Moreover, there are no mechanisms for access control in P2P storage cloud. To address this issue, we design a ciphertext-policy attribute-based encryption (ABE) scheme and a proxy re-encryption scheme. Based on them, we further propose a secure, efficient and fine-grained data Access Control mechanism for P2P storage Cloud named ACPC. We enforce access policies based on user attributes, and integrate P2P reputation system in ACPC. ACPC enables data owners to delegate most of the laborious user revocation tasks to cloud servers and reputable system peers. Our security analysis demonstrates that ACPC is provably secure. The performance evaluation shows that ACPC is highly efficient under practical settings, and it significantly reduces the computation overheads brought to data owners and cloud servers during user revocation, compared with other state-of-the-art revocable ABE schemes.
Heng He, Ruixuan Li 0001, Xinhua Dong, Zhao Zhang 0010
IEEE Trans. Cloud Comput.4
2014 Automatic runtime frequency-scaling system for energy savings in parallel applications
Vaibhav Sundriyal, Masha Sosonkina, Zhao Zhang 0010
J. Supercomput.3
2014 MASTER: A Multicore Cache Energy-Saving Technique Using Dynamic Cache Reconfiguration
abstract
With increasing number of on-chip cores and CMOS scaling, the size of last-level caches (LLCs) is on the rise and hence, managing their leakage energy consumption has become vital for continuing to scale performance. In multicore systems, the locality of memory access stream is significantly reduced because of multiplexing of access streams from different running programs and hence, leakage energy-saving techniques such as decay cache, which rely on memory access locality, do not save a large amount of energy. The techniques based on way level allocation provide very coarse granularity and the techniques based on offline profiling become infeasible to use for large number of cores. We present a multicore cache energy saving technique using dynamic cache reconfiguration (MASTER) that uses online profiling to predict energy consumption of running programs at multiple LLC sizes. Using these estimates, suitable cache quotas are allocated to different programs using cache coloring scheme and the unused LLC space is turned off to save energy. Even for four core systems, the implementation overhead of MASTER is only 0.8% of L2 size. We evaluate MASTER using out-of-order simulations with multiprogrammed workloads from SPEC2006 and compare it with conventional cache leakage energy-saving techniques. The results show that MASTER gives the highest saving in energy and does not harm performance or cause unfairness. For twoand four-core simulations, the average savings in memory subsystem (which includes LLC and main memory) energy over shared baseline LLC are 15% and 11%, respectively. Also, the average values of weighted speedup and fair speedup are close to one (≥0.98).
Sparsh Mittal, Yanan Cao 0002, Zhao Zhang 0010
IEEE Trans. Very Large Scale Integr. Syst.3
2013 Free ECC: An efficient error protection for compressed last-level caches
abstract
Cache reliability is increasingly a concern as cache cell dimension shrinks and cache capacity grows. Conventionally, an extra, dedicated storage is appended to cache to store error correcting code. Recently, cache compression schemes have been proposed to increase the effective cache capacity of last-level cache (LLC), for which we found the conventional cache ECC design is inefficient. We propose Free ECC that utilizes the unused fragments in compressed cache design to store ECC. It not only reduces the chip overhead but also improves cache utilization and power efficiency. Additionally, we propose an efficient convergent cache allocation scheme to organize the compressed data blocks more effectively than existing schemes. Our evaluation using SPEC CPU2006 and PARSEC benchmarks shows that the Free ECC design improves cache capacity utilization and power efficiency significantly, with negligible overhead on overall performance. This new design makes compressed cache an increasingly viable choice for processors with requirements of high reliability.
Yanan Cao 0002, Zhao Zhang 0010
ICCD3
2013 FlexiWay: A cache energy saving technique using fine-grained cache reconfiguration
abstract
Recent trends of CMOS scaling and use of large last level caches (LLCs) have led to significant increase in the leakage energy consumption of LLCs and hence, managing their energy consumption has become extremely important in modern processor design. The conventional cache energy saving techniques require offline profiling or provide only coarse granularity of cache allocation. We present FlexiWay, a cache energy saving technique which uses dynamic cache reconfiguration. FlexiWay logically divides the cache sets into multiple (e.g. 16) modules and dynamically turns off suitable and possibly different number of cache ways in each module. FlexiWay has very small implementation overhead and it provides fine-grain cache allocation even with caches of typical associativity, e.g. an 8-way cache. Microarchitectural simulations have been performed using an x86-64 simulator and workloads from SPEC2006 suite. Also, FlexiWay has been compared with two conventional energy saving techniques. The results show that FlexiWay provides largest energy saving and incurs only small loss in performance. For single, dual and quad core systems, the average energy saving using FlexiWay are 26.2%, 25.7% and 22.4%, respectively.
Sparsh Mittal, Zhao Zhang 0010, Jeffrey S. Vetter
ICCD2
2013 Achieving energy efficiency during collective communications
abstract
SUMMARY Energy consumption has become a major design constraint in modern computing systems. With the advent of petaflops architectures, power‐efficient software stacks have become imperative for scalability. Techniques such as dynamic voltage and frequency scaling (called DVFS) and CPU clock modulation (called throttling) are often used to reduce the power consumption of the compute nodes. To avoid significant performance losses, these techniques should be used judiciously during parallel application execution. For example, its communication phases may be good candidates to apply the DVFS and CPU throttling without incurring a considerable performance loss. They are often considered as indivisible operations although little attention is being devoted to the energy saving potential of their algorithmic steps. In this work, two important collective communication operations, all‐to‐all and allgather, are investigated as to their augmentation with energy saving strategies on theper‐callbasis. The experiments prove the viability of such a fine‐grain approach. They also validate a theoretical power consumption estimate for multicore nodes proposed here. While keeping the performance loss low, the obtained energy savings were always significantly higher than those achieved when DVFS or throttling were switched on across the entire application run. Copyright © 2012 John Wiley & Sons, Ltd.
Vaibhav Sundriyal, Masha Sosonkina, Zhao Zhang 0010
Concurr. Comput. Pract. Exp.3
2013 Energy saving strategies for parallel applications with point-to-point communication phases
Vaibhav Sundriyal, Masha Sosonkina, Alexander Gaenko, Zhao Zhang 0010
J. Parallel Distributed Comput.4
2013 E3CC: A memory error protection scheme with novel address mapping for subranked and low-power memories
abstract
This study presents and evaluates E 3 CC (Enhanced Embedded ECC), a full design and implementation of a generic embedded ECC scheme that enables power-efficient error protection for subranked memory systems. It incorporates a novel address mapping scheme called Biased Chinese Remainder Mapping (BCRM) to resolve the address mapping issue for memories of page interleaving, plus a simple and effective cache design to reduce extra ECC traffic. Our evaluation using SPEC CPU2006 benchmarks confirms the performance and power efficiency of the E 3 CC scheme for subranked memories as well as conventional memories.
Yanan Cao 0002, Zhao Zhang 0010
ACM Trans. Archit. Code Optim.3
2013 Thermal Modeling and Management of DRAM Systems
abstract
With increasing data rate and power density, high-performance memories have started to require dynamic thermal management (DTM), following the trend of processor and hard drive. There are also lack of a memory thermal model and simulation tools to facilitate the research of memory DTM. This study investigates the approach of coordinating processor, which is the source of memory access requests, and memory to improve system performance and/or power efficiency during memory thermal emergency. Two such schemes, namely adaptive core gating (DTM-ACG) and coordinated DVFS (DTM-CDVFS), are proposed and evaluated on a real server platform. DTM-ACG gates processor cores and DTM-CDVFS scales down the frequency and voltage level of processor cores according to memory thermal emergency level. Their combination, namely DTM-COMB, is also evaluated. The experimental results show that the two schemes, while successfully controlling memory activities and handling thermal emergencies, improve performance significantly under the given thermal envelope. The measurement results from an Intel SR1500AL server testbed show that on average, DTM-ACG and DTM-CDVFS improve performance by 6.7 and 15.3 percent, respectively, over a prior memory bandwidth throttling scheme. DTM-CDVFS also reduces the processor power rate by 15.5 percent and system (including processor and memory) energy by 22.7 percent. Additionally, we propose a DRAM thermal model and validate it with measurement on the instrumented server platform. We find that our proposed model faithfully catches the dynamic DRAM temperature changes; the average difference between the modeled and measured temperature is less than $(1^{\circ}{\rm C})$.
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Zhao Zhang 0010
IEEE Trans. Computers4
2011 Memory Architecture for Integrating Emerging Memory Technologies
abstract
Current main memory system design is severely limited by the decades-old synchronous DRAM architecture, which requires the memory controller to track the internal status of memory devices (chips) and schedule the timing of all device operations. This rigidity has become an obstacle of integrating emerging memory technologies such as PCM into existing memory systems, because their timing requirements are vastly different. Furthermore, with the trend of embedding memory controllers into processors, it is crucial to have interoperability among general-purpose processors and diverse memory modules. To address this issue, we propose a new memory architecture framework called universal memory architecture (UniMA). It enables the interoperability by decoupling the scheduling of device operations from memory controller, using a bridge chip at each memory module to perform local scheduling. The new architecture may also help improve memory scalability, power efficiency, and bandwidth as previously proposed decoupled memory organizations. A major focus of this study is to evaluate the performance impact of local scheduling of device operations. We present a prototype implementation of UniMA on top of DDRx memory bus, and then evaluate its efficiency with different workloads. The simulation results show that UniMA actually improves memory system efficiency for memory-intensive workloads due to increased parallelism among memory modules. The overall performance improvement over the conventional DDRx memory architecture is 3.1% on average. The performance of other workloads is reduced slightly, by 1.0% on average, due to a small increase of memory idle latency. In short, the prototype and evaluation demonstrate that it is possible to integrate diverse memory technologies into a single memory architecture with virtually no loss of overall performance.
Kun Fang 0005, Zhao Zhang 0010, Zhichun Zhu
PACT3
2010 An Application-Level Data Transparent Authentication Scheme without Communication Overhead
abstract
With abundant aggregate network bandwidth, continuous data streams are commonly used in scientific and commercial applications. Correspondingly, there is an increasing demand of authenticating these data streams. Existing strategies explore data stream authentication by using message authentication codes (MACs) on a certain number of data packets (a data block) to generate a message digest, then either embedding the digest into the original data, or sending the digest out-of-band to the receiver. Embedding approaches inevitably change the original data, which is not acceptable under some circumstances (e.g., when sensitive information is included in the data). Sending the digest out-of-band incurs additional communication overhead, which consumes more critical resources (e.g., power in wireless devices for receiving information) besides network bandwidth. In this paper, we propose a novel strategy, DaTA, which effectively authenticates data streams by selectively adjusting some interpacket delay. This authentication scheme requires no change to the original data and no additional communication overhead. Modeling-based analysis and experiments conducted on an implemented prototype system in an LAN and over the Internet show that our proposed scheme is efficient and practical.
Songqing Chen, Shiping Chen 0003, Xinyuan Wang 0005, Zhao Zhang 0010, Sushil Jajodia
IEEE Trans. Computers4
2009 Soft-OLP: Improving Hardware Cache Performance through Software-Controlled Object-Level Partitioning
abstract
Performance degradation of memory-intensive programs caused by the LRU policy's inability to handle weak-locality data accesses in the last level cache is increasingly serious for two reasons. First, the last-level cache remains in the CPU's critical path, where only simple management mechanisms, such as LRU, can be used, precluding some sophisticated hardware mechanisms to address the problem. Second, the commonly used shared cache structure of multi-core processors has made this critical path even more performance-sensitive due to intensive inter-thread contention for shared cache resources. Researchers have recently made efforts to address the problem with the LRU policy by partitioning the cache using hardware or OS facilities guided by run-time locality information. Such approaches often rely on special hardware support or lack enough accuracy. In contrast, for a large class of programs, the locality information can be accurately predicted if access patterns are recognized through small training runs at the data object level. To achieve this goal, we present a system-software framework referred to as Soft-OLP (Software-based Object-Level cache Partitioning). We first collect per-object reuse distance histograms and inter-object interference histograms via memory-trace sampling. With several low-cost training runs, we are able to determine the locality patterns of data objects. For the actual runs, we categorize data objects into different locality types and partition the cache space among data objects with a heuristic algorithm, in order to reduce cache misses through segregation of contending objects. The object-level cache partitioning framework has been implemented with a modified Linux kernel, and tested on a commodity multi-core processor. Experimental results show that in comparison with a standard L2 cache managed by LRU, Soft-OLP significantly reduces the execution time by reducing L2 cache misses across inputs for a set of single- and multi-threaded programs from the SPEC CPU2000 benchmark suite, NAS benchmarks and a computational kernel set.
Qingda Lu, Jiang Lin, Xiaoning Ding, Zhao Zhang 0010, Xiaodong Zhang 0001, P. Sadayappan
PACT4
2009 Run-Time Detection of Malwares via Dynamic Control-Flow Inspection
abstract
Conventional approach of detecting malwares relies on static scanning of malware signature. However, it may not work on the malwares that use software protection methods such as encryption and packing with run-time decryption and unpacking. We propose a hardware-assisted malware detection system that detects malwares during program run time to complement the conventional approach. It searches for control flow-based signature of malware during program execution, therefore bypassing the protection method used by those malwares. A new hardware design is used to assist the collection of control flow information. We have implemented and evaluated a prototype system on top of a full-system simulator based on the Intel x86 architecture. The experimental results show that the system can successfully distinguish all 30 malware variants and other benign programs that we have randomly collected, and that the overall run-time performance overhead is negligible. In short, the study demonstrates that it is a viable approach to detect malware in run time using control flow-based signature.
Yong-Joon Park, Zhao Zhang 0010, Songqing Chen
ASAP2
2009 Decoupled DIMM: building high-bandwidth memory system using low-speed DRAM devices
abstract
The widespread use of multicore processors has dramatically increased the demands on high bandwidth and large capacity from memory systems. In a conventional DDR2/DDR3 DRAM memory system, the memory bus and DRAM devices run at the same data rate. To improve memory bandwidth, we propose a new memory system design called decoupled DIMM that allows the memory bus to operate at a data rate much higher than that of the DRAM devices. In the design, a synchronization buffer is added to relay data between the slow DRAM devices and the fast memory bus; and memory access scheduling is revised to avoid access conflicts on memory ranks. The design not only improves memory bandwidth beyond what can be supported by current memory devices, but also improves reliability, power efficiency, and cost effectiveness by using relatively slow memory devices. The idea of decoupling, precisely the decoupling of bandwidth match between memory bus and a single rank of devices, can also be applied to other types of memory systems including FB-DIMM.
Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Zhichun Zhu
ISCA3
2009 Enabling software management for multicore caches with a lightweight hardware support
abstract
The management of shared caches in multicore processors is a critical and challenging task. Many hardware and OS-based methods have been proposed. However, they may be hardly adopted in practice due to their non-trivial overheads, high complexities, and/or limited abilities to handle increasingly complicated scenarios of cache contention caused by many-cores.
Jiang Lin, Qingda Lu, Xiaoning Ding, Zhao Zhang 0010, Xiaodong Zhang 0001, P. Sadayappan
SC4
2008 Gaining insights into multicore cache partitioning: Bridging the gap between simulation and real systems
abstract
Cache partitioning and sharing is critical to the effective utilization of multicore processors. However, almost all existing studies have been evaluated by simulation that often has several limitations, such as excessive simulation time, absence of OS activities and proneness to simulation inaccuracy. To address these issues, we have taken an efficient software approach to supporting both static and dynamic cache partitioning in OS through memory address mapping. We have comprehensively evaluated several representative cache partitioning schemes with different optimization objectives, including performance, fairness, and quality of service (QoS). Our software approach makes it possible to run the SPEC CPU2006 benchmark suite to completion. Besides confirming important conclusions from previous work, we are able to gain several insights from whole-program executions, which are infeasible from simulation. For example, giving up some cache space in one program to help another one may improve the performance of both programs for certain workloads due to reduced contention for memory bandwidth. Our evaluation of previously proposed fairness metrics is also significantly different from a simulation-based study. The contributions of this study are threefold. (1) To the best of our knowledge, this is a highly comprehensive execution- and measurement-based study on multicore cache partitioning. This paper not only confirms important conclusions from simulation-based studies, but also provides new insights into dynamic behaviors and interaction effects. (2) Our approach provides a unique and efficient option for evaluating multicore cache partitioning. The implemented software layer can be used as a tool in multicore performance evaluation and hardware design. (3) The proposed schemes can be further refined for OS kernels to improve performance.
Jiang Lin, Qingda Lu, Xiaoning Ding, Zhao Zhang 0010, Xiaodong Zhang 0001, P. Sadayappan
HPCA4
2008 Memory Access Scheduling Schemes for Systems with Multi-Core Processors
abstract
On systems with multi-core processors, the memory access scheduling scheme plays an important role not only in utilizing the limited memory bandwidth but also in balancing the program execution on all cores. In this study, we propose a scheme, called ME-LREQ, which considers the utilization of both processor cores and memory subsystem. It takes into consideration both the long-term and short-term gains of serving a memory request by prioritizing requests hitting on the row buffers and from the cores that can utilize memory more efficiently and have fewer pending requests. We have also thoroughly evaluated a set of memory scheduling schemes that differentiate and prioritize requests from different cores. Our simulation results show that for memory-intensive, multiprogramming workloads, the new policy improves the overall performance by 10.7% on average and up to 17.7% on a four-core processor, when compared with scheme that serves row buffers hit memory requests first and allows memory reads bypassing writes; and by up to 9.2% (6.4% on average) when compared with the scheme that serves requests from the core with the fewest pending requests first.
Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Zhichun Zhu
ICPP3
2008 BotTracer: Execution-Based Bot-Like Malware Detection
Lei Liu 0021, Songqing Chen, Guanhua Yan, Zhao Zhang 0010
ISC4
2008 Mini-rank: Adaptive DRAM architecture for improving memory power efficiency
abstract
The widespread use of multicore processors has dramatically increased the demand on high memory bandwidth and large memory capacity. As DRAM subsystem designs stretch to meet the demand, memory power consumption is now approaching that of processors. However, the conventional DRAM architecture prevents any meaningful power and performance trade-offs for memory-intensive workloads. We propose a novel idea called mini-rank for DDRx (DDR/DDR2/DDR3) DRAMs, which uses a small bridge chip on each DRAM DIMM to break a conventional DRAM rank into multiple smaller mini-ranks so as to reduce the number of devices involved in a single memory access. The design dramatically reduces the memory power consumption with only a slight increase on the memory idle latency. It does not change the DDRx bus protocol and its configuration can be adapted for the best performance-power trade-offs. Our experimental results using four-core multiprogramming workloads show that using x32 mini-ranks reduces memory power by 27.0% with 2.8% performance penalty and using x16 mini-ranks reduces memory power by 44.1% with 7.4% performance penalty on average for memory-intensive workloads, respectively.
Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Eugene Gorbatov, Howard David, Zhichun Zhu
MICRO3
2008 Software thermal management of dram memory for multicore systems
abstract
Thermal management of DRAM memory has become a critical issue for server systems. We have done, to our best knowledge, the first study of software thermal management for memory subsystem on real machines. Two recently proposed DTM (Dynamic Thermal Management) policies have been improved and implemented in Linux OS and evaluated on two multicore servers, a Dell PowerEdge 1950 server and a customized Intel SR1500AL server testbed. The experimental results first confirm that a system-level memory DTM policy may significantly improve system performance and power efficiency, compared with existing memory bandwidth throttling scheme. A policy called DTM-ACG (Adaptive Core Gating) shows performance improvement comparable to that reported previously. The average performance improvements are 13.3% and 7.2% on the PowerEdge 1950 and the SR1500AL (vs. 16.3% from the previous simulation-based study), respectively. We also have surprising findings that reveal the weakness of the previous study: the CPU heat dissipation and its impact on DRAM memories, which were ignored, are significant factors. We have observed that the second policy, called DTM-CDVFS (Coordinated Dynamic Voltage and Frequency Scaling), has much better performance than previously reported for this reason. The average improvements are 10.8% and 15.3% on the two machines (vs. 3.4% from the previous study), respectively. It also significantly reduces the processor power by 15.5% and energy by 22.7% on average.
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Eugene Gorbatov, Howard David, Zhao Zhang 0010
SIGMETRICS6
2007 Thermal modeling and management of DRAM memory systems
abstract
With increasing speed and power density, high-performance memories, including FB-DIMM (Fully Buffered DIMM) and DDR2 DRAM, now begin to require dynamic thermal management(DTM) as processors and hard drives did. The DTM of memories, nevertheless, is different in that it should take the processor performance and power consumption into consideration. Existing schemes have ignored that. In this study, we investigate a new approach that controls the memory thermal issues from the source generating memory activities - the processor. It will smooth the program execution when compared with shutting down memory abruptly, and therefore improve the overall system performance and power efficiency. For multicore systems, we propose two schemes called adaptive core gating and coordinated DVFS. The first scheme activates clock gating on selected processor cores and the second one scales down the frequency and voltage levels of processor cores when the memory is to be over-heated. They can successfully control the memory activities and handle thermal emergency. More importantly, they improve performance significantly under the given thermal envelope. Our simulation results show that adaptive coregating improves performance by up to 23.3% (16.3% on average) on a four-core system with FB-DIMM when compared with DRAM thermal shutdown; and coordinated DVFS with control-theoretic methods improves the performance by up to 18.5% (8.3% on average).
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Howard David, Zhao Zhang 0010
ISCA5
2007 DRAM-Level Prefetching for Fully-Buffered DIMM: Design, Performance and Power Saving
abstract
We have studied DRAM-level prefetching for the fully buffered DIMM (FB-DIMM) designed for multi-core processors. FB-DIMM has a unique two-level interconnect structure, with FB-DIMM channels at the first-level connecting the memory controller and Advanced Memory Buffers (AMBs); and DDR2 buses at the second-level connecting the AMBs with DRAM chips. We propose an AMB prefetching method that prefetches memory blocks from DRAM chips to AMBs. It utilizes the redundant bandwidth between the DRAM chips and AMBs but does not consume the crucial channel bandwidth. The proposed method fetches K memory blocks of L2 cache block sizes around the demanded block, where K is a small value ranging from two to eight. The method may also reduce the DRAM power consumption by merging some DRAM precharges and activations. Our cycle-accurate simulation shows that the average performance improvement is 16% for single-core and multi-core workloads constructed from memory-intensive SPEC2000 programs with software cache prefetching enabled; and no workload has negative speedup. We have found that the performance gain comes from the reduction of idle memory latency and the improvement of channel bandwidth utilization. We have also found that there is only a small overlap between the performance gains from the AMB prefetching and the software cache prefetching. The average of estimated power saving is 15%.
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Zhao Zhang 0010, Howard David
ISPASS4
2005 A Performance Comparison of DRAM Memory System Optimizations for SMT Processors
abstract
Memory system optimizations have been well studied on single-threaded systems; however, the wide use of simultaneous multithreading (SMT) techniques raises questions over their effectiveness in the new context. In this study, we thoroughly evaluate contemporary multi-channel DDR SDRAM and Rambus DRAM systems in SMT systems, and search for new thread-aware DRAM optimization techniques. Our major findings are: (1) in general, increasing the number of threads tends to increase the memory concurrency and thus the pressure on DRAM systems, but some exceptions do exist; (2) the application performance is sensitive to memory channel organizations, e.g. independent channels may outperform ganged organizations by up to 90%; (3) the DRAM latency reduction through improving row buffer hit rates becomes less effective due to the increased bank contentions; and (4) thread-aware DRAM access scheduling schemes may improve performance by up to 30% on workload mixes of memory-intensive applications. In short, the use of SMT techniques has somewhat changed the context of DRAM optimizations but does not make them obsolete.
Zhichun Zhu, Zhao Zhang 0010
HPCA2
2005 Performance Characterization of Java Applications on SMT Processors
abstract
As Java is emerging as one of the major programming languages in software development, studying how Java applications behave on recent SMT processors is of great interest. This paper characterizes the performance of Java applications on an Intel Pentium 4 hyper-threading processor. Using the performance counters provided by Pentium 4, we quantitatively evaluate micro-architecture metrics while running various types of Java applications. The experimental results reveal that: (1) Hyper-threading can indeed improve the performance of multithreaded Java programs; (2) The resource contentions within Pentium 4 are the major reason of pipeline inefficiency, which prevents better performance promised by SMT; (3) The static partition design of hyper-threading causes considerable performance loss for many single-thread Java programs; (4) Most multiprogrammed Java benchmarks can achieve decent combined speedups on hyper-threading processors
Wei Huang 0032, Jiang Lin, Zhao Zhang 0010, J. Morris Chang
ISPASS3
2005 Towards Pairing Java Applications on SMT Processors
abstract
This paper investigates various issues of pairing Java applications for multithreaded execution on Intel's hyper-threading Pentium 4 processor. We first quantify the overall performance of multiprogrammed Java applications using a metric called combined speedup. Using the performance counters provided by the Pentium 4, we then quantitatively evaluate the performance of underneath micro-architecture components and their implications to the combined speedup. A statistical model is proposed to analyze the collected data. This novel approach reveals that trace cache is the major factor determining the pairing performance. In particular, we find that the trace cache miss rates of Java applications can be utilized to predict the combined speedups. Three new scheduling strategies are proposed based on these observations and then evaluated. The experimental results show that the proposed strategies have better performance than the conventional round-robin scheduling scheme. Overall, our best strategy enables a reduction in execution time of 10.5% over the serial execution, comparing with a reduction of 5.92% achieved by the round-robin scheduling. The improvement will be increasingly significant on future SMT processors.
Wei Huang 0032, Jiang Lin, Zhao Zhang 0010, J. Morris Chang
MASCOTS3
2005 Performance Modeling and Tuning Strategies of Mixed Mode Collective Communications
abstract
On SMP clusters, mixed mode collective MPI communications, which use shared memory communications within SMP nodes and point-to-point communications between SMP nodes, are more efficient than conventional implementations. In a previous study, we proposed several new methods that made mixed mode collective communications significantly faster than the pure point-to-point ones. However, the optimal performance required the tuning of many parameters, which was done by testing every possible setting and was very time consuming. In this study, we propose a new performance model that considers the special characteristics of mixed mode collective communications. The model provides good predictions to reduce most settings without testing by execution. It considers both shared-memory and point-to-point communications, while existing performance models only consider the point-to-point ones. Based on this model, we develop a number of tuning strategies that reduce the overall tuning time to only 10% of previous tuning time.
Meng-Shiou Wu, Ricky A. Kendall, Kyle Wright, Zhao Zhang 0010
SC4
2004 Design and Optimization of Large Size and Low Overhead Off-Chip Caches
abstract
Large off-chip L3 caches can significantly improve the performance of memory-intensive applications. However, conventional L3 SRAM caches are facing two issues as those applications require increasingly large caches. First, an SRAM cache has a limited size due to the low density and high cost of SRAM and, thus, cannot hold the working sets of many memory-intensive applications. Second, since the tag checking overhead of large caches is nontrivial, the existence of L3 caches increases the cache miss penalty and may even harm the performance of some memory-intensive applications. To address these two issues, we present a new memory hierarchy design that uses cached DRAM to construct a large size and low overhead off-chip cache. The high density DRAM portion in the cached DRAM can hold large working sets, while the small SRAM portion exploits the spatial locality appearing in L2 miss streams to reduce the access latency. The L3 tag array is placed off-chip with the data array, minimizing the area overhead on the processor for L3 cache, while a small tag cache is placed on-chip, effectively removing the off-chip tag access overhead. A prediction technique accurately predicts the hit/miss status of an access to the cached DRAM, further reducing the access latency. Conducting execution-driven simulations for a 2 GHz 4-way issue processor and with 11 memory-intensive programs from the SPEC 2000 benchmark, we show that a system with a cached DRAM of 64 MB DRAM and 128 KB on-chip SRAM cache as the off-chip cache outperforms the same system with an 8 MB SRAM L3 off-chip cache by up to 78 percent measured by the total execution time. The average speedup of the system with the cached-DRAM off-chip cache is 25 percent over the system with the L3 SRAM cache.
Zhao Zhang 0010, Zhichun Zhu, Xiaodong Zhang 0001
IEEE Trans. Computers1
2002 Fine-Grain Priority Scheduling on Multi-Channel Memory Systems
abstract
Configurations of contemporary DRAM memory systems become increasingly complex. A recent study shows that the application performance is highly sensitive to choices of configurations. In this study we show that, by utilizing fine-grain priority access scheduling, we are able to find a workload independent configuration that achieves optimal performance on a multichannel memory system. Our approach can well utilize the available high concurrency and high bandwidth on such memory systems, and effectively reduce the memory stall time of memory-intensive applications. Conducting execution-driven simulation of a 4-way issue, a 2 GHz processor, we show that the average performance improvement for fifteen memory-intensive SPEC2000 programs by using an optimized fine-grain priority scheduling is about 13% and 8% for a 2-channel and a 4-channel Direct Rambus DRAM memory system, respectively, compared with gang scheduling. Compared with burst scheduling, the average performance improvement is 16% and 14% for the 2-channel and 4-channel memory systems, respectively.
Zhichun Zhu, Zhao Zhang 0010, Xiaodong Zhang 0001
HPCA2
2000 A permutation-based page interleaving scheme to reduce row-buffer conflicts and exploit data locality
abstract
DRAM row-buffer conflicts occur when a sequence of requests on different rows goes to the same memory bank, causing much higher memory access latency than requests to the same row or to different banks. We analyze the sources of row-buffer conflicts in the context of superscalar processors, and propose a permutation based page interleaving scheme to reduce row-buffer conflicts and to exploit data access locality in the row-buffer. Compared with several existing schemes, we show that the permutation based scheme dramatically increases the hit rates on DRAM row-buffers and reduces memory stall time of the SPEC95 and TPC-C workloads. The memory stall times of the workloads are reduced up to 68% and 50%, compared with the conventional cache line and page interleaving schemes, respectively.
Zhao Zhang 0010, Zhichun Zhu, Xiaodong Zhang 0001
MICRO1
2000 Cacheminer: A Runtime Approach to Exploit Cache Locality on SMP
abstract
Exploiting cache locality of parallel programs at runtime is a complementary approach to a compiler optimization. This is particularly important for those applications with dynamic memory access patterns. We propose a memory-layout oriented technique to exploit cache locality of parallel loops at runtime on Symmetric Multiprocessor (SMP) systems. Guided by application-dependent and targeted architecture-dependent hints, our system, called Cacheminer, reorganizes and partitions a parallel loop using the memory-access space of its execution. Through effective runtime transformations, our system maximizes the data reuse in each partitioned data region assigned in a cache, and minimizes the data sharing among the partitioned data regions assigned to all caches. The executions of tasks in the partitions are scheduled in an adaptive and locality-presented way to minimize the execution time of programs by trading off load balance and locality. We have implemented the Cacheminer runtime library on two commercial SMP servers and an SimCS simulated SMP. Our simulation and measurement results show that our runtime approach can achieve comparable performance with the compiler optimizations for programs with regular computation and memory-access patterns, whose load balance and cache locality can be well optimized by the tiling and other program transformations. However, our experimental results show that our approach is able to significantly improve the memory performance for the applications with irregular computation and dynamic memory access patterns. These types of programs are usually hard to optimize by static compiler optimizations.
Yong Yan 0003, Xiaodong Zhang 0001, Zhao Zhang 0010
IEEE Trans. Parallel Distributed Syst.3
1999 Cache-Optimal Methods for Bit-Reversals
abstract
Bit-reversals are representative and important data reordering operations in many scientific computations. Performance degradation is mainly caused by cache conflict misses. Bit-reversals are often repeatedly used as fundamental subroutines for many scientific programs. Thus, in order to gain the best performance, cache-optimal methods and their implementations should be carefully and precisely done at the programming level. This type of performance programming for some special programs, such as the data reorderings, may significantly outperform an optimization from an automatic tool, such as a compiler. In this paper, we examine different methods using techniques of blocking, buffering, and padding for efficient implementations. We evaluate the merits and limits of each technique and their application and architecture-dependent conditions for developing cache-optimal methods. We present two contributions in this paper: (1) Our integrated blocking methods, which match cache associativity and TLB cache size and which fully use the available registers are cache-optimal and fast. (2) We show that our padding methods outperform other software oriented methods, and believe they are the fastest in terms of minimizing both CPU and memory access cycles. Since the padding methods are almost independent of hardware, they could be widely used on many uniprocessor workstations and SMP multiprocessors. 1
Zhao Zhang 0010, Xiaodong Zhang 0001
SC1
1998 A memory-layout oriented run-time technique for locality optimization
abstract
Exploiting locality at run-time is a complementary approach to a compiler approach for those applications with dynamic memory access patterns. This paper proposes a memory-layout oriented approach to exploit cache locality for parallel loops at run-time on symmetric multi-processor (SMP) systems. Guided by application-dependent hints and the targeted cache architecture, it reorganizes and partitions a parallel loop through shrinking and partitioning the memory-access space of the loop at run-time. In the generated task partitions, the data sharing among partitions is minimized and the data reuse in a partition is maximized. The execution of tasks in partitions is scheduled in an adaptive and locality-preserved way to achieve balanced execution, for minimizing the execution time of applications by trading off load balance and locality. Based on simulation and measurement, we show our run-time approach can achieve comparable performance with the compiler optimizations for two applications, whose load balance and cache locality can be well optimized by the tiling and other program transformations. However our experimental results also show that our approach is able to significantly improve the memory performance for the applications with dynamic memory access patterns. This type of programs are usually hard to be optimized by compilers.
Yong Yan 0003, Xiaodong Zhang 0001, Zhao Zhang 0010
ICPP3